Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

How mitochondria redefine the code.

Annotated, complete DNA sequences are available for 213 mitochondrial genomes from 132 species. These provide an extensive sample of evolutionary adjustment of codon usage and meaning spanning the history of this organelle. Because most known coding changes are mitochondrial, such data bear on the general mechanism of codon reassignment. Coding changes have been attributed variously to loss of codons due to changes in directional mutation affecting the genome GC content (Osawa and Jukes 1988), to pressure to reduce the number of mitochondrial tRNAs to minimize the genome size (Anderson and Kurland 1991), and to the existence of transitional coding mechanisms in which translation is ambiguous (Schultz and Yarus 1994a). We find that a succession of such steps explains existing reassignments well. In particular, (1) Genomic variation in the prevalence of a codon's third-position nucleotide predicts relative mitochondrial codon usage well, though GC content does not. This is because A and T, and G and C, are uncorrelated in mitochondrial genomes. (2) Codons predicted to reach zero usage (disappear) do so more often than expected by chance, and codons that do disappear are disproportionately likely to be reassigned. However, codons predicted to disappear are not significantly more likely to be reassigned. Therefore, low codon frequencies can be related to codon reassignment, but appear to be neither necessary nor sufficient for reassignment. (3) Changes in the genetic code are not more likely to accompany smaller numbers of tRNA genes and are not more frequent in smaller genomes. Thus, mitochondrial codons are not reassigned during demonstrable selection for decreased genome size. Instead, the data suggest that both codon disappearance and codon reassignment depend on at least one other event. This mitochondrial event (leading to reassignment) occurs more frequently when a codon has disappeared, and produces only a small subset of possible reassignments. We suggest that coding ambiguity, the extension of a tRNA's decoding capacity beyond its original set of codons, is the second event. Ambiguity can act alone but often acts in concert with codon disappearance, which promotes codon reassignment.

Base Composition↗

Structure and expression of the gene encoding the periplasmic arylsulfatase of Chlamydomonas reinhardtii.

Chlamydomonas reinhardtii produces a periplasmic arylsulfatase in response to sulfur deprivation. We have isolated and sequenced arylsulfatase cDNAs from a lambda gt11 expression library. The amino acid sequence of the protein, as deduced from the nucleotide sequence, has features characteristic of secreted proteins, including a signal sequence and putative glycosylation sites. The gene has a broad codon usage with seven codons, all having A residues in the third position, not previously observed in C. reinhardtii genes. Arylsulfatase transcription is tightly regulated by sulfur availability. The approximately 2.7 kb arylsulfatase transcript is very susceptible to degradation, disappearing in less than an hour after sulfur starved cells are administered either sulfate or alpha-amanitin. The accumulation of the arylsulfatase transcript is also suppressed by the addition of cycloheximide. Transcription initiation from the arylsulfatase gene occurs approximately 100 bp upstream of the initiation codon, in a region that is 5' to a 43 bp imperfect inverted repeat. Preceding the transcription start site are sequences similar to those present in promoter regions of other genes from C. reinhardtii.

Amino Acid Sequence↗

Development of Polymorphic EST Markers Suitable for Genetic Linkage Mapping of Catfish.

: Expressed sequence tag (EST) markers are important for gene mapping and for marker-assisted selection (MAS). To develop EST markers for use in catfish gene mapping, 100 randomly picked complementary DNAs from the channel catfish (Ictalurus punctatus) pituitary library were sequenced. The EST sequences were used to design primers to amplify channel catfish and blue catfish (I. furcatus) genomic DNAs. Polymerase chain reaction products of the ESTs were analyzed to determine length polymorphism between the channel catfish and blue catfish. Eleven polymorphic EST markers were identified. Five of the 11 EST markers were from known genes and the other six were from unidentified ESTs. Seven ESTs were found to be associated with microsatellite sequences. Analysis of channel catfish gene sequences indicated highly biased codon usage, with 16 codons being preferably used. These codons were more preferably used in highly expressed ribosomal protein genes and in highly expressed pituitary hormone genes. G/C-rich codons are less used in channel catfish than those in other vertebrates suggesting AT-richness of the channel catfish genome.

Journal Article↗

Catalyzing bacterial speciation: correlating lateral transfer with genetic headroom.

Unlike crown eukaryotic species, microbial species are created by continual processes of gene loss and acquisition promoted by horizontal genetic transfer. The amounts of foreign DNA in bacterial genomes, and the rate at which this is acquired, are consistent with gene transfer as the primary catalyst for microbial differentiation. However, the rate of successful gene transfer varies among bacterial lineages. The heterogeneity in foreign DNA content is directly correlated with amount of genetic headroom intrinsic to a bacterial species. Genetic headroom reflects the amount of potentially dispensable information--reflected in codon usage bias and codon context bias--that can be transiently sacrificed to allow experimentation with functions introduced by gene transfer. In this way, genetic headroom offers a potential metric for assessing the propensity of a lineage to speciate.

Bacteria↗

Molecular population genetics and evolution of a prion-like protein in Saccharomyces cerevisiae.

The prion-like behavior of Sup35p, the eRF3 homolog in the yeast Saccharomyces cerevisiae, mediates the activity of the cytoplasmic nonsense suppressor known as [PSI(+)]. Sup35p is divided into three regions of distinct function. The N-terminal and middle (M) regions are required for the induction and propagation of [PSI(+)] but are not necessary for translation termination or cell viability. The C-terminal region encompasses the termination function. The existence of the N-terminal region in SUP35 homologs of other fungi has led some to suggest that this region has an adaptive function separate from translation termination. To examine this hypothesis, we sequenced portions of SUP35 in 21 strains of S. cerevisiae, including 13 clinical isolates. We analyzed nucleotide polymorphism within this species and compared it to sequence divergence from a sister species, S. paradoxus. The N domain of Sup35p is highly conserved in amino acid sequence and is highly biased in codon usage toward preferred codons. Amino acid changes are under weak purifying selection based on a quantitative analysis of polymorphism and divergence. We also conclude that the clinical strains of S. cerevisiae are not recently derived and that outcrossing between strains in S. cerevisiae may be relatively rare in nature.

Amino Acid Sequence↗

Suppression of the negative effect of minor arginine codons on gene expression; preferential usage of minor codons within the first 25 codons of the Escherichia coli genes.

AGA and AGG codons for arginine are the least used codons in Escherichia coli, which are encoded by a rare tRNA, the product of the dnaY gene. We examined the positions of arginine residues encoded by AGA/AGG codons in 678 E. coli proteins. It was found that AGA/AGG codons appear much more frequently within the first 25 codons. This tendency becomes more significant in those proteins containing only one AGA or AGG codon. Other minor codons such as CUA, UCA, AGU, ACA, GGA, CCC and AUA are also found to be preferentially used within the first 25 codons. The effects of the AGG codon on gene expression were examined by inserting one to five AGG codons after the 10th codon from the initiation codon of the lacZ gene. The production of beta-galactosidase decreased as more AGG codons were inserted. With five AGG codons, the production of beta-galactosidase (Gal-AGG5) completely ceased after a mid-log phase of cell growth. After 22 hr induction of the lacZ gene, the overall production of Gal-AGG5 was 11% of the control production (no insertion of arginine codons). When five CGU codons, the major arginine codon were inserted instead of AGG, the production of beta-galactosidase (Gal-CGU5) continued even after stationary phase and the overall production was 66% of the control. The negative effect of the AGG codons on the Gal-AGG5 production was found to be dependent upon the distance between the site of the AGG codons and the initiation codon. As the distance was increased by inserting extra sequences between the two codons, the production of Gal-AGG5 increased almost linearly up to 8 fold. From these results, we propose that the position of the minor codons in an mRNA plays an important role in the regulation of gene expression possibly by modulating the stability of the initiation complex for protein synthesis.

Amino Acid Sequence↗

The complete mitochondrial genome of Tupaia belangeri and the phylogenetic affiliation of scandentia to other eutherian orders.

The complete mitochondrial genome of Tupaia belangeri, a representative of the eutherian order Scandentia, was determined and compared with full-length mitochondrial sequences of other eutherian orders described to date. The complete mitochondrial genome is 16, 754 nt in length, with no obvious deviation from the general organization of the mammalian mitochondrial genome. Thus, features such as start codon usage, incomplete stop codons, and overlapping coding regions, as well as the presence of tandem repeats in the control region, are within the range of mammalian mitochondrial (mt) DNA variation. To address the question of a possible close phylogenetic relationship between primates and Tupaia, the evolutionary affinities among primates, Tupaia and bats as representatives of the Archonta superorder, ferungulates, guinea pigs, armadillos, rats, mice, and hedgehogs were examined on the basis of the complete mitochondrial DNA sequences. The opossum sequence was used as an outgroup. The trees, estimated from 12 concatenated genes encoded on the mitochondrial H-strand, add further molecular evidence against an Archonta monophyly. With the new data described in this paper, most of both the mitochondrial and the nuclear data point away from Scandentia as the closest extant relatives to primates. Instead, the complete mitochondrial data support a clustering of Scandentia with Lagomorpha connecting to the branch leading to ferungulates. This closer phylogenetic relationship of Tupaia to rabbits than to primates first received support from several analyses of nuclear and partial mitochondrial DNA data sets. Given that short sequences are of limited use in determining deep mammalian relationships, the partial mitochondrial data available to date support this hypothesis only tentatively. Our complete mitochondrial genome data therefore add considerably more evidence in support of this hypothesis.

Animals↗

Shortening of the symptom-free period in rhesus macaques is associated with decreasing nonsynonymous variation in the env variable regions of simian immunodeficiency virus SIVsm during passage.

During six blood passages of simian immunodeficiency virus SIVsm in rhesus macaques, the asymptomatic period shortened from 18 months to 1 month. To study SIVsm envelope gene (env) evolution during passage in rhesus macaques, the C1 to CD4 binding regions of multiple clones were sequenced at seroconversion and again at death. The env variation found during adaptation was almost completely confined to the variable regions. Intrasample sequence variation among clones at seroconversion was lower than the variation among clones at death. Intrasample variation among clones from a single time point as well as intersample variation decreased during the passage. In the variable regions, the mean number of intrasample nonsynonymous nucleotide substitutions decreased from the first passage (5.26 x 10(-2) +/- 0.6 x 10(-2) per site) to the fifth passage (2.24 x 10(-2) +/- 0.4 x 10(-2) per site), whereas in the constant regions, the mean number of intrasample nonsynonymous nucleotide substitutions differed less between the first and fifth passages (1. 14 x 10(-2) +/- 0.27 x 10(-2) and 0.80 x 10(-2) +/- 0.24 x 10(-2) per site). Shortening of the asymptomatic period coincided with a rise in the Ks/Ka ratio (ratio between the number of synonymous [Ks] and the number of nonsynonymous [Ka] substitutions) from 1.080 in passage one to 1.428 in passage five and mimicked the difference seen in the intrahost evolution between asymptomatic and fast-progressing individuals infected with human immunodeficiency virus type 1. The distribution of nonsynonymous substitutions was biphasic, with most of the adaptation of env variable regions occurring in the first three passages. This phase, in which the symptom-free period fell to 4 months, was followed by a plateau phase of apparently reduced adaptation. Analysis of codon usage revealed decreased codon redundancy in the variable regions. Overall, the results suggested a biphasic pattern of adaptation and evolution, with extremely rapid selection in the first three passages followed by an equilibrium or stabilization of the variation between env clones at different time points in passages four to six.

Animals↗

Cluster analysis of the codon use frequency of MHC genes from different species.

The relative synonymous codon use frequency of 135 MHC genes from four mammal species (Homo sapiens, Pan troglodyte, Macaca mulanta and Rattus norvegicus) is analyzed using a hierarchical cluster method. The result suggests that gene function is the dominant factor that determines codon usage bias, while species is a minor factor that determines further difference in codon usage bias for genes with similar functions. The conclusion may be useful in gene classification and gene function prediction.

Animals↗

Intercodon dinucleotides affect codon choice in plant genes.

In this work, 710 CDSs corresponding to over 290 000 codons equally distributed between Brassica napus, Arabidopsis thaliana, Lycopersicon esculentum, Nicotiana tabacum, Pisum sativum, Glycine max, Oryza sativa, Triticum aestivum, Hordeum vulgare and Zea mays were considered. For each amino acid, synonymous codon choice was determined in the presence of A, G, C or T as the initial nucleotide of the subsequent triplet; data were statistically analysed under the hypothesis of an independent assortment of codons. In 33.4% of cases, a frequency significantly (P: = 0.01) different from that expected was recorded. This was mainly due to a pervasive intercodon TpA and CpG deficiency. As a general rule, intercodon TpAs and CpGs were preferably replaced by CpAs and TpGs, respectively. In several instances, codon frequencies were also modified to avoid homotetramer and homotrimer formation, to reduce intercodon ApCs downstream (1,2) GG or AG dinucleotides, as well as to increase GpA or ApG intercodons under certain contexts. Since TpA, CpG and homotetra(tri)mer deficiency directly or indirectly accounted for 77% of significant variation in the codon frequency, it can be concluded that codon usage mirrors precise needs at the DNA structure level. Plant species exhibited a phylogenetically-related adaptation to structural constraints. Codon usage flexibility was reflected in strikingly different arrays of optimum codons for probe design.

Base Composition↗

Primary structure of the tolC gene that codes for an outer membrane protein of Escherichia coli K12.

We present the nucleotide sequence of the tolC gene of Escherichia coli K12, and the amino acid sequence of the TolC protein (an outer membrane protein) as deduced from it. The mature TolC protein comprises 467 amino acid residues, and, as previously reported (1), a signal sequence of 22 amino acid residues is attached to the N-terminus. The C-terminus of the gene is followed by a stem-loop structure (8 base pair stem, 4 base loop) which may be a rho-independent termination signal. The codon usage of the gene is nonrandom; the major isoaccepting species of tRNA are preferentially utilised, or, among synonomous codons recognized by the same tRNA, those codons are used which can interact better with the anticodon (2,3). In contrast to the codon usage for other outer membrane proteins of E. coli (4) the rare arginine codons AGA and AGG are used once and twice respectively.

Bacterial Outer Membrane Proteins↗

[Role of the code redundancy in determining cotranslational protein folding].

It has been demonstrated earlier in our laboratory that rare codon clusters can determine the boundaries of the polypeptide chain fragments of the same secondary structure type during the co-translational protein folding. According to this data, co-translational protein folding can occur under condition of a correlation between the frequency of codon choice in mRNAs and the relative abundance of their isoaccepting tRNAs. The alterations in the spectrum and concentrations of the isoaccepting tRNAs in different cells were demonstrated by many authors. The existence of a mechanism of the coordinate regulation of the levels (activities) of the isoaccepting tRNAs, corresponding aminoacyl-tRNA synthetases and mRNAs predominantly translated at a given moment of time can be suggested. Such a mechanism can ensure the needed accuracy of the protein folding process. Analysis of gene sequences of various pro- and eukaryotic organisms carried out in the present work revealed that the codon usage frequency spectra of simultaneously synthesized proteins are similar. The relative appearance of the most rare and frequent codons in investigated gene sequences displays a high degree of conservatism. It has also been found that structural-homologous proteins from different organisms (cytochromes c, myoglobins) have very similar codon frequency distribution profiles. This property retains despite the significant variations in the codon usage spectra in the investigated gene sequences. The data obtained indicate that the codon distribution in mRNAs whose diversity is mainly conditioned by the genetic code redundance is a program that determines translational rates of different mRNA parts thus controlling the spatial folding of the synthesized peptide chain.

Animals↗

Comparison of the nucleoside sequence of trpA and sequences immediately beyond the trp operon of Klebsiella aerogenes. Salmonella typhimurium and Escherichia coli.

The nucleotide sequence of trpA of Klebsiella aerogenes is presented and compared with the trpA sequences of Salmonella typhimurium and Escherichia coli. The majority of the approximately 200 differences between each pair of trpA's are single nucleotide pair changes that do not alter the amino acid sequence. Codon usage conforms to the general patterns revealed by examination of other prokaryotic gene sequences. However, codon usage in K. aerogenes trpA reflects the high G+C content of the genome of this organism. The DNA sequences just beyond trpA, the presumed transcription termination region, are also compared for the three species. Perusal of these sequences indicates that the secondary structure of the transcript segment just beyond trpA has been preserved, while the primary sequence has diverged appreciably.

Amino Acid Sequence↗

Complete nucleotide sequence and genetic organization of the bacteriocinogenic plasmid, pIP404, from Clostridium perfringens.

The complete nucleotide sequence of the bacteriocinogenic plasmid, pIP404, from Clostridium perfringens has been determined. The plasmid genome comprises 10,207 bp and has a dA + dT content of 75%. Functions have been tentatively assigned to 6 of the 10 open reading frames and an origin-like region of repeated sequence identified. The codon usage of this extremely dA + dT rich plasmid is highly unusual and displays a pronounced preference for codons with the lowest dG + dC content. Only one of the genes from pIP404 was expressed at a significant level in Escherichia coli, suggesting that the atypical codon usage could represent a major obstacle to heterologous gene expression.

Base Sequence↗

[Transformation of Chlamydomonas reinhardtii CW-15 with the hygromycin phosphotransferase gene as a selective marker].

To transform Chlamydomonas reinhardtii Dang. Cells, plasmid pCTVHyg was constructed with the use of the Escherichia coli hygromycin phosphotransferase gene (hpt) controlled by the SV40 early promoter. Cells of the CW-15 mutant strain were transformed by electroporation, with the yield reaching 10(3) hygromycin-resistant (HygR) clones per 10(6) recipient cells. The exogenous DNA integrated in the Ch. reinhardtii nuclear genome showed stable transmission for approximately 350 cell generations, while hygromycin resistance was expressed as an unstable character. Codon usage was compared for the hpt gene and Ch. reinhardtii nuclear genes. The results testified that codon usage bias, which is characteristic of Ch. reinhardtii, is not the major factor affecting foreign gene expression. The advantages of the selective system for studying Ch. reinhardtii transformation with heterologous genes are discussed.

Animals↗

Preferential use of A- and U-rich codons for Mycoplasma capricolum ribosomal proteins S8 and L6.

The nucleotide sequence of the 1.3 kilobase-pair DNA segment, which contains the genes for ribosomal proteins S8 and L6, and a part of L18 of Mycoplasma capricolum, has been determined and compared with the corresponding sequence in Escherichia coli (Cerretti et al., Nucl. Acids Res. 11, 2599, 1983). Identities of the predicted amino acid sequences of S8 and L6 between the two organisms are 54% and 42%, respectively. The A + T content of the M. capricolum genes is 71%, which is much higher than that of E. coli (49%). Comparisons of codon usage between the two organisms have revealed that M. capricolum preferentially uses A- and U-rich codons. More than 90% of the codon third positions and 57% of the first positions in M. capricolum is either A or U, whereas E. coli uses A or U for the third and the first positions at a frequency of 51% and 36%, respectively. The biased choice of the A- and U-rich codons in this organism has been also observed in the codon replacements for conservative amino acid substitutions between M. capricolum and E. coli. These facts suggest that the codon usage of M. capricolum is strongly influenced by the high A + T content of the genome.

Adenine↗

The sequence of the chloroplast atpB gene and its flanking regions in Chlamydomonas reinhardtii.

The chloroplast (cp)-encoded CF1 ATPase beta-subunit gene (atpB) of Chlamydomonas reinhardtii and its flanking regions have been sequenced. The derived amino acid (aa) sequence is highly homologous to that of the beta-subunit gene in Escherichia coli, bovine heart mitochondria, and higher plant cp. In contrast to all other cp genomes, the CF1 epsilon subunit gene (atpE) does not lie at the 3' end of the atpB gene but maps to a position 92 kb away in the other single-copy region. Northern blots confirm that the beta subunit is not encoded as part of a dicistronic message as it is in higher plants. The region just upstream from the atpB gene in C. reinhardtii contains two small open reading frames (ORFs) and not the gene for the large subunit of ribulose-1,5-bisphosphate carboxylase/oxygenase as is found in cp genomes of higher plants. No transcripts for either ORF were detected, but the codon usage in these ORFs as well as in the atpB gene follows the unique pattern of codon usage previously seen in other cp genes in C. reinhardtii.

Amino Acid Sequence↗