Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

The complete nucleotide sequence of the domestic dog (Canis familiaris) mitochondrial genome.

The complete nucleotide sequence of the mitochondrial genome of the domestic dog, Canis familiaris, was determined. The length of the sequence was 16,728 bp; however, the length was not absolute due to the variation (heteroplasmy) caused by differing numbers of the repetitive motif, 5'-GTACACGT(A/G)C-3', in the control region. The genome organization, gene contents, and codon usage conformed to those of other mammalian mitochondrial genomes. Although its features were unknown, the "CTAGA" duplication event which followed the translational stop codon of the COII gene was not observed in other mammalian mitochondrial genomes. In order to determine the possible differences between mtDNAs in carnivores, two rRNA and 13 protein-coding genes from the cat, dog, and seal were compared. The combined molecular differences, in two rRNA genes as well as in the inferred amino acid sequences of the mitochondrial 13 protein-coding genes, suggested that there is a closer relationship between the dog and the seal than there is between either of these species and the cat. Based on the molecular differences of the mtDNA, the evolutionary divergence between the cat, the dog, and the seal was dated to approximately 50 +/- 4 million years ago. The degree of difference between carnivore mtDNAs varied according to the individual protein-coding gene applied, showing that the evolutionary relationships of distantly related species should be presented in an extended study based on ample sequence data like complete mtDNA molecules.

Animals↗

Molecular evolution of an imprinted gene: repeatability of patterns of evolution within the mammalian insulin-like growth factor type II receptor.

The repeatability of patterns of variation in Ka/Ks and Ks is expected if such patterns are the result of deterministic forces. We have contrasted the molecular evolution of the mammalian insulin-like growth factor type II receptor (Igf2r) in the mouse-rat comparison with that in the human-cow comparison. In so doing, we investigate explanations for both the evolution of genomic imprinting and for Ks variation (and hence putatively for mutation rate evolution). Previous analysis of Igf2r, in the mouse-rat comparison, found Ka/Ks patterns that were suggested to be contrary to those expected under the conflict theory of imprinting. We find that Ka/Ks variation is repeatable and hence confirm these patterns. However, we also find that the molecular evolution of Igf2r signal sequences suggests that positive selection, and hence conflict, may be affecting this region. The variation in Ks across Igf2r is also repeatable. To the best of our knowledge this is the first demonstration of such repeatability. We consider three explanations for the variation in Ks across the gene: (1) that it is the result of mutational biases, (2) that it is the result of selection on the mutation rate, and (3) that it is the product of selection on codon usage. Explanations 2 and 3 predict a Ka-Ks correlation, which is not found. Explanation 3 also predicts a negative correlation between codon bias and Ks, which is also not found. However, in support of explanation 1 we do find that in rodents the rate of silent C --> T mutations at CpG sites does covary with Ks, suggesting that methylation-induced mutational patterns can explain some of the variation in Ks. We find evidence to suggest that this CpG effect is due to both variation in CpG density, and to variation in the frequency with which CpGs mutate. Interestingly, however, a GC4 analysis shows no covariance with Ks, suggesting that to eliminate methyl-associated effects CpG rates themselves must be analyzed. These results suggest that, in contrast to previous studies of intragenic variation, Ks patterns are not simply caused by the same forces responsible for Ka/Ks correlations.

Animals↗

Asparaginyl-tRNA synthetase from Thermus thermophilus HB8. Sequence of the gene and crystallization of the enzyme expressed in Escherichia coli.

The gene for the asparaginyl-tRNA synthetase, a class IIb enzyme, from the extreme thermophile Thermus thermophilus HB8 has been cloned and sequenced. Sequence analysis revealed an open reading frame that codes for a protein of 438 amino acid residues (50875 Da). Codon usage in the asparaginyl-tRNA synthetase gene (asnS) is similar to the characteristic usage in the genes for proteins from bacteria of the genus Thermus, and the G+C content in the third position of the codons is as high as 94%. The amino acid sequence of asparaginyl-tRNA synthetase from T. thermophilus shows high similarity with other bacterial asparaginyl-tRNA synthetase sequences (30-55% identity). By expression of the T. thermophilus asnS gene in Escherichia coli, the thermostable enzyme was overproduced and purified to homogeneity by heat treatment and two chromatography steps. The protein obtained is remarkably thermostable and retains 50% of its initial tRNA aminoacylation activity after 1 h of incubation at 90 degrees C or 21 h at 85 degrees C. Crystals of the enzyme were obtained from polyethylene glycol 6000 solutions by vapour diffusion techniques. The crystals diffract X-rays beyond 2.8 A.

Amino Acid Sequence↗

Analyses of the gene and amino acid sequence of the Prevotella (Bacteroides) ruminicola 23 xylanase reveals unexpected homology with endoglucanases from other genera of bacteria.

The DNA sequence for the xylanase gene from Prevotella (Bacteroides) ruminicola 23 was determined. The xylanase gene encoded for a protein with a molecular weight of 65,740. An apparent leader sequence of 22 amino acids was observed. The promoter region for expression of the xylanase gene in Bacteroides species was identified with a promoterless chloramphenicol acetyltransferase gene. A region of high amino acid homology was found with the proposed catalytic domain of endoglucanases from several organisms, including Butyrivibrio fibrisolvens, Ruminococcus flavefaciens, and Clostridium thermocellum. The cloned xylanase was found to exhibit endoglucanase activity against carboxymethyl cellulose. Analysis of the codon usage for the xylanase gene found a bias towards G and C in the third position in 16 of 18 amino acids with degenerate codons.

Amino Acid Sequence↗

The difference in the type of codon-anticodon base pairing at the ribosomal P-site is one of the determinants of the translational rate.

By utilizing an enzymatically reconstructed tRNA variant containing an altered anticodon sequence, we have examined the different biochemical behavior of translation between the Watson-Crick type and the wobble type base pair interactions at the first anticodon position. We have found that the Watson-Crick type base pair has an advantage in translation in contrast to the wobble type base pair by comparing the efficiency of transpeptidation of native tRNA(Phe) (anticodon; GmAA) with its variant tRNA (anticodon; AAA) in the poly(U)-programmed ribosome system. Thomas et al. [Proc. Natl. Acad. Sci. U.S. (1988) 85, 4242-4246] showed that the wobble codon at the ribosomal A-site accepted its cognate tRNA less efficiently than the Watson-Crick base pairing codon. We report here that the wobble interaction at the ribosomal P-site also affected the rate of translation. This variable translational rate may be a mechanism of gene regulation through preferential codon usage.

Anticodon↗

A comparative mitogenomic analysis of the potential adaptive value of Arctic charr mtDNA introgression in brook charr populations (Salvelinus fontinalis Mitchill).

Wild brook charr populations (Salvelinus fontinalis) completely introgressed with the mitochondrial genome (mtDNA) of arctic charr (Salvelinus alpinus) are found in several lakes of northeastern Québec, Canada. Mitochondrial respiratory enzymes of these populations are thus encoded by their own nuclear DNA and by arctic charr mtDNA. In the present study we performed a comparative sequence analysis of the whole mitochondrial genome of both brook and arctic charr to identify the distribution of mutational differences across these two genomes. This analysis revealed 47 amino acid replacements, 45 of which were confined to subunits of the NADH dehydrogenase complex (Complex I), one in the cox3 gene (Complex IV), and one in the atp8 gene (Complex V). A cladistic approach performed with brook charr, arctic charr, and two other salmonid fishes (rainbow trout [Oncorhynchus mykiss] and Atlantic salmon [Salmo salar]) revealed that only five amino acid replacements were specific to the charr comparison and not shared with the other two salmonids. In addition, five amino acid substitutions localized in the nad2 and nad5 genes denoted negative scores according to the functional properties of amino acids and, therefore, could possibly have an impact on the structure and functional properties of these mitochondrial peptides. The comparison of both brook and arctic charr mtDNA with that of rainbow trout also revealed a relatively constant mutation rate for each specific gene among species, whereas the rate was quite different among genes. This pattern held for both synonymous and nonsynonymous nucleotide positions. These results, therefore, support the hypothesis of selective constraints acting on synonymous codon usage.

Amino Acid Substitution↗

Compositional bias and size of genomes of human DNA viruses.

Genomes of 144 human DNA viruses were analyzed in the aspect of their compositional asymmetry. DNA viruses were divided into two groups according to their genome sizes. The analysis revealed that the level of guanine and cytosine (GC content) in the coding sequences of small genome DNA viruses was significantly lower than that of large genome DNA viruses. Because small genome viruses replicate their genomes using cellular enzymes, while large genome viruses use their own enzymes for genome replication, the two groups of viruses may be under different mutational bias and/or selection pressure. In these viruses, GC content at the third codon position correlated with GC content at the first and second codon position. However, the relationship in small genome DNA viruses was weaker than that in large genome DNA viruses, suggesting that their genome composition may be more strongly influenced by codon usage preference or restriction on amino acid composition.

Base Composition↗

Several distinct genes encode nearly identical to 16 kDa proteolipids of the vacuolar H(+)-ATPase from Arabidopsis thaliana.

To understand the subcellular roles and the regulation of vacuolar H(+)-ATPases, we have begun to identify the genes encoding the major subunits and to determine their patterns of expression in Arabidopsis thaliana. Two distinct cDNAs (AVA-P1 and AVA-P2) and one genomic sequence (AVA-P3) encoding the 16 kDa subunit have been isolated. The 16 kDa proteolipid is a major component of the membrane integral sector that forms the proton conductance pathway and is required for assembly of the V-ATPase complex. Interestingly, the open reading frame of one full-length cDNA (AVA-P1) and a genomic sequence (AVA-P3) encoded an identical polypeptide of 164 amino acids with a molecular mass of 16,570. The deduced amino acid sequences of the two cDNAs were nearly identical (99%) and hydropathy plots suggested a molecule with four membrane-spanning domains characteristic of V-ATPase proteolipids. The three genes differed mainly in their codon usage and in their 3'-untranslated regions. The coding region of the genomic sequence, AVA-P3, was interrupted by two introns located at the codons for Cys-26 and Arg-121. The presence of additional 16 kDa proteolipid genes was suggested from several polymerase chain reaction (PCR)-amplified fragments that differed from one another in the size of the second intron. PCR 1 had an intron of ca. 800 bp and its identity as AVA-P4, a fourth member of the gene family, was confirmed from sequence analyses of an EST cDNA. The mRNAs of three genes (AVA-P1, AVA-P2 and AVA-P3) were detected in Arabidopsis leaf, root, flower and silique; yet expression of AVA-P1 and AVA-P2 was lower in roots. All three genes were expressed in light- or dark-grown seedlings; however mRNA levels of AVA-P2 were enhanced in etiolated plants. Arabidopsis thaliana, therefore, has at least four distinct genes encoding nearly identical 16 kDa proteolipids, and the enhanced expression of AVA-P2 transcript in etiolated seedlings suggests that an increase in V-ATPase could accompany cell expansion.

Adaptation, Biological↗

The ribosomal protein gene cluster of Mycoplasma capricolum.

The DNA sequence of the part of the Mycoplasma capricolum genome that contains the genes for 20 ribosomal proteins and two other proteins has been determined. The organization of the gene cluster is essentially the same as that in the S10 and spc operons of Escherichia coli. The deduced amino acid sequence of each protein is also well conserved in the two bacteria. The G + C content of the M. capricolum genes is 29%, which is much lower than that of E. coli (51%). The codon usage pattern of M. capricolum is different from that of E. coli and extremely biased to use of A and U(T): about 91% of codons have A or U in the third position. UGA, which is a stop codon in the "universal" code, is used more abundantly than UGG to dictate tryptophan.

Amino Acid Sequence↗

Nucleotide sequence of the phosphoenolpyruvate carboxylase gene of the cyanobacterium Anacystis nidulans.

Nucleotide sequence of the open reading frame (ORF) for the phosphoenolpyruvate carboxylase gene (ppc) of the cyanobacterium Anacystis nidulans was determined. The ORF consists of 3159 bp and codes for 1053 amino acid (aa) residues. The codon usage of the ppc of A. nidulans is not so markedly different from that of the Escherichia coli ppc, yet, in A. nidulans the preferred codons are AAG for lysine and CCC for proline, whereas those are seldom used in the E. coli ppc.

Amino Acid Sequence↗

Improvement of human interferon HUIFNalpha2 and HCV core protein expression levels in Escherichia coli but not of HUIFNalpha8 by using the tRNA(AGA/AGG).

High-level expression from one particular heterologous gene in Escherichia coli generally requires the optimization of codon usage. Genes encoding for Hepatitis C virus core protein (HCcAg), human interferon alpha2 and 8 subtypes (HUIFNalpha2 and HUIFNalpha8) show a high content of AGA/AGG codons. These are encoded by the product of the dnaY gene in E. coli. The proteins used in this work have a high therapeutic value and were used as models for studying the effects of these rare codons on the efficiency of heterologous gene expression in E. coli. Expression plasmids were constructed to express any of these proteins and the dnaY gene product simultaneously in E. coli. After dnaY gene expression, HCcAg, and HUIFNalpha2 expression levels increased 5 and 3 times, respectively. However, HUIFNalpha8 expression was barely detected either supplying or not the additional dnaY gene product. These results suggest that the high frequency of AGA/AGG codons present in the HCcAg and HUIFNalpha2 genes could be one of the factors limiting its expression in E. coli. Nevertheless, for HUIFNalpha8 it seems that other factors prevail upon the lack of dnaY product. Data presented here for HCcAg and HUIFNalpha2 expressions proved the value of this approach to obtain therapeutic proteins in E. coli.

Antiviral Agents↗

Nucleotide sequence of the rubella virus capsid protein gene reveals an unusually high G/C content.

The nucleotide sequence of the rubella virus capsid protein (C) gene has been determined from a cDNA clone derived from the 40S genomic RNA. The sequence covers the coding region of the C protein (831 nucleotides), 70 nucleotides of the 5' untranslated region, and the 5' end of the downstream E2 membrane protein gene. The capsid gene is unusually rich in C (41.6%) and G (31.2%) residues (G + C 72.8%), and poor in A (15.4%) and U residues (11.8%). There are regions with long runs of up to 45% C or 35% G residues. The codon usage is non-random, with a strong preference for C and G residues in the third position. Starting from two in-frame AUG codons (seven amino acid residues apart) an open reading frame (ORF) was identified that extended in frame into the ORF coding for the downstream E2 membrane protein gene. Since the amino terminus of the capsid protein is blocked, we could not determine which of the AUGs serve as the initiating codon. To verify that the deduced ORF was correct, we have determined the amino acid sequence of 13 tryptic peptides corresponding to one-third of the C protein. Our data show that the C protein is about 277 residues in length (Mr about 30750). It is very hydrophilic and rich in prolines (14.1%) and arginines (14.4%). Clusters of these amino acids are concentrated in the amino-terminal third of the C protein. No sequence homology to the capsid protein of several alphaviruses was observed. Together with our previous sequence data we have now completed the sequence of the genes coding for the structural proteins C, E2 and E1 of rubella virus.

Amino Acid Sequence↗

High-level expression of the human tumor necrosis factor-alpha.

Five expression plasmids for a total chemically synthesized gene of tumor necrosis factor-alpha (TNF-alpha) were constructed, which differ in two aspects: 1) distance (D) between the SD and the initiation codon ATG; 2) energy-delta G0f298 released to form the stable secondary structure in the translation initiation region. The plasmid with a (D) of 6bp has the lowest delta G0f298 (absolute value) and showed the highest expression level, up to 60% of the total bacterial proteins. Codons usage will impose much influence on expression level, so it is reasonable and understandable that the total chemically synthesized TNF-alpha gene with the codons preferable to E. coli use gave much higher expression level than the sc-TNF-alpha in which only partial N-terminal codons were modified.

Base Sequence↗

Statistical analysis of L-tuple frequencies in eubacteria and organelles.

This work is an attempt to study the structural features and evolutionary patterns of nucleotide sequences by analyzing their 1- through 4-plet frequencies and statistical relations between them. We present mathematical apparatus for this analysis. In particular, we introduce criteria to estimate the degree of homogeneity of L-plet composition in a given set of sequences and the dependence of the L-plet frequencies on the composition of lower orders. We apply these criteria to the study of eubacteria, mitochondria and chloroplasts. We demonstrate that L-plet frequencies are quite useful for revealing evolutionary relationship between DNA sequences and that the non-random distribution is more typical for doublets than to triplets. Non-randomness of triplet composition is more characteristic to coding than to non-coding regions, while no significant differences in dinucleotide composition can be observed. The obtained results can be used for revealing possible mechanisms of the codon usage phenomena.

Analysis of Variance↗

Frequencies of codons in histones, tubulins and fibrinogen: bias due to interference between transcription signals and protein function.

The distribution of codons was studied in 65 proteins: 48 histones, 14 tubulins, and three fibrinogens, With the methodology used, (1) we confirmed that the preterminator state of a codon has no detectable effect on codon bias. (2) The well-known effect of CG suppression was visible. We also found that (3) some codons which are very rare, are equal to parts of known transcription signals. Thus, we advanced that to avoid signal interference, the use of these codons is suppressed when a synonymous codon is available. In addition we found that in the whole series of codons, transcription signals are less frequent than in a random sequence of equal composition. Finally we observed (4) that tryptophan is absent in histones. This absence was related not to the TGG codon itself, but to characteristics of the amino acid. We conclude that the functional constraints of a protein can influence, at least for synonymous codon usage, the evolution of its own coding sequence.

Animals↗

Nucleotide sequence of the Azospirillum brasilense Sp7 glutamine synthetase structural gene.

The complete nucleotide sequence of the glnA gene, encoding the glutamine synthetase subunit of Azospirillum brasilense Sp7, was established. This is the first Azospirillum gene sequenced. The gene encodes a 468 residue polypeptide of MW 51,917. The similarity coefficient (SAB) between the polypeptidic sequence of Azospirillum and Anabaena 7120, which is the only other glnA sequence available, is 58%. No significant homology with E. coli canonical and ntr promoters, or with the promoter region of the Anabaena glnA gene was found. When fused to an E. coli promoter, the gene could be translated in E. coli, despite a very biased codon usage and an atypical Shine-Dalgarno sequence.

Amino Acid Sequence↗

Modulation of lambda integrase synthesis by rare arginine tRNA.

Lambda's int gene contains an anomalously high frequency of the rare arginine codons AGA and AGG when compared to genes of Escherichia coli or to the rest of phage lambda. These are the least frequent codons in genes of E. coli and are recognized by the rarest tRNAs. The presence of these codons reduces the translation rate and, depending on the context, this can strongly modulate translational efficiency by a variety of mechanisms. In this study, we show that expression of the natural int gene may also be modulated by rare arginine codon usage, and we explore this mechanism.

Amino Acid Sequence↗

Transcription attenuation in Salmonella typhimurium: the significance of rare leucine codons in the leu leader.

The leucine operon of Salmonella typhimurium is controlled by a transcription attenuation mechanism. Four adjacent leucine codons within a 160-nucleotide leu leader RNA are thought to play a central role in this mechanism. Three of the four codons are CUA, a rarely used leucine codon within enteric bacteria. To determine whether the nature of the leucine codon affects the regulation of the leucine operon, we used oligonucleotide-directed mutagenesis to first convert one CUA of the leader to CUG and then convert all three CUA codons to CUG. CUG is the most frequently used leucine codon in enteric bacteria. A mutant having (CUA)2CUGCUC in place of (CUA)3CUC has an altered response to leucine limitation, requiring a slightly higher degree of limitation to effect derepression. Changing (CUA)3CUC to (CUG)3CUC has more dramatic effects upon operon expression. First, the basal level of expression is lowered to the point that the mutant grows more slowly than the parent in a minimal medium lacking leucine. Second, the response of the mutant to a leucine limitation is dramatically altered such that even a strong limitation elicits only a modest degree of derepression. If the mutant is grown under conditions of leucyl-tRNA limitation rather than leucine limitation, complete derepression can be achieved, but only at a much higher degree of limitation than for the wild-type operon. These results provide a clear-cut example of codon usage having a dramatic effect upon gene expression.

Codon↗