Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Reduced synonymous substitution rate at the start of enterobacterial genes.

Synonymous codon usage is less biased at the start of Escherichia coli genes than elsewhere. The rate of synonymous substitution between E.coli and Salmonella typhimurium is substantially reduced near the start of the gene, which suggests the presence of an additional selection pressure which competes with the selection for codons which are most rapidly translated. Possible competing sources of selection are the presence of secondary ribosome binding sites downstream from the start codon, the avoidance of mRNA secondary structure near the start of the gene and the use of sub-optimal codons to regulate gene expression. We provide evidence against the last of these possibilities. We also show that there is a decrease in the frequency of A, and an increase in the frequency of G along the E.coli genes at all three codon positions. We argue that these results are most consistent with selection to avoid mRNA secondary structure.

Base Composition↗

The nucleotide sequence of the rat cytoplasmic beta-actin gene.

The nucleotide sequence of the rat beta-actin gene was determined. The gene codes for a protein identical to the bovine beta-actin. It has a large intron in the 5' untranslated region 6 nucleotides upstream from the initiator ATG, and 4 introns in the coding region at codons specifying amino acids 41/42, 121/122, 267, and 327/328. Unlike the skeletal muscle actin gene and many other actin genes, the beta-actin gene lacks the codon for Cys between the initiator ATG and the codon for the N-terminal amino acid of the mature protein. The usage of synonymous codons in the beta-actin gene is nonrandom, and is similar to that in the rat skeletal muscle and other vertebrate actin genes, but differs from the codon usage in yeast and soybean actin genes.

Actins↗

Codon usage in Tetrahymena and other ciliates.

Codon usage in ciliates was examined by analyzing the coding regions of 22 ciliate genes corresponding to a total of 26,142 nucleotides (8,714 codons). It was found that Tetrahymena, Paramecium and the hypotrichs (Oxytricha and Stylonychia) differed in which synonymous codons were used most frequently by their genes. In fact, the codon choices in highly expressed Tetrahymena genes were more similar to those of yeast genes than those of Paramecium genes. The ciliates do not appear to have unusually strong biases in codon usage frequency when compared to other protists such as yeast. The analysis of the Tetrahymena genes indicated that genes which are highly expressed during normal cell growth have a stronger bias towards using the "preferred" codons than those expressed at lower levels during growth or for brief periods during processes such as conjugation. This conforms to what is found in other protists.

Amino Acids↗

Why the rate of silent codon substitutions is variable within a vertebrate's genome.

Different genes within the murine genome are diverging at different rates. The rate of synonymous codon substitutions in these genes is related to their base composition. It is proposed that the variabilities of rate of mutation accumulation, of codon choice and of average GC content within vertebrate genomes are caused by differences between DNA synthesis in repair and replication, as far as the frequency and compositional bias of mutations introduced by these systems are concerned. DNA repair contributes substantially to the evolution of the DNA domains, which are actively repaired in germline cells and which correspond to regions available for transcription in these cells and to Giemsa-negative bands in stained chromosomes.

Animals↗

Primary structure of the ompF gene that codes for a major outer membrane protein of Escherichia coli K-12.

The nucleotide sequence of the ompF gene coding for a major outer membrane protein of Escherichia coli K-12 has been determined and the amino acid sequence of the OmpF protein was deduced from it. The OmpF protein contains 340 amino acid residues, and is produced from a precursor having 22 extra amino acid residues, the signal peptide, at the amino terminus. The expected secondary structure of the OmpF protein had a high beta-sheet content with a low alpha-helix content. The promoter region and the transcription termination region of the ompF gene had a significantly high AT content, while the AT content of the coding region was about the same as the average AT content of the E. coli chromosome. Following the termination codon, a typical rho-independent transcription termination signal was observed. The codon usage in the ompF gene was highly nonrandom; the codons preferably utilized are those recognized by the most abundant species of isoaccepting tRNAs or those, among synonymous codons recognized by the same tRNA, that can interact more properly with the anticodon.

Amino Acid Sequence↗

Codon reading scheme in Mycoplasma pneumoniae revealed by the analysis of the complete set of tRNA genes.

The 33 genes encoding the complete set of tRNA species in Mycoplasma pneumoniae have been cloned and sequenced. They are organized into 5 clusters in addition to 9 single genes. No redundant gene was found, indicating that 33 tRNAs correspond to 32 different anticodons and decode all 62 codons used in this organism. There is only one single tRNA for each of the Ala, Leu, Pro, and Val family boxes. Therefore, a simplified decoding system resembling that recently described for Mycoplasma capricolum (1) has to also exist in M.pneumoniae. However, analysis of the anticodon set and codon usage revealed features characteristic of the latter: (i) there is no obvious preference toward AT rich synonymous codons, (ii) CGG codons are assigned for arginine and are translated by tRNA Arg(UCG), and (iii) CNN or GNN anticodons are encountered in the Ser, Thr, Arg, and Gly family boxes. We thus propose that this codon-anticodon recognition pattern has emerged in the 'M.pneumoniae cluster' under a genomic economization strategy but without the influence of AT pressure.

Amino Acid Sequence↗

Rates of aminoacyl-tRNA selection at 29 sense codons in vivo.

We have placed aminoacyl-tRNA selection at individual codons in competition with a frameshift that is assumed to have a uniform rate. By assaying a reporter in the shifted frame, relative rates for association of the 29 YNN codons and their cognate aminoacyl-tRNAs were obtained during logarithmic growth in Escherichia coli. For five codons, three beginning with C and two with U, these relative rates agree with relative in vitro rates for elongation factor Tu-mediated aminoacyl-tRNA binding to ribosomes and subsequent GTP hydrolysis. Therefore, the frameshift assay probably measures this process in vivo. Observed rates for aminoacyl-tRNA selection span a 25-fold range. Therefore, the time required to transit different codons in vivo probably differs substantially. Codons very frequently used in highly expressed genes generally select aminoacyl-tRNAs more quickly than do rarely used codons. This suggests that speed of aminoacyl-tRNA selection is a significant factor determining biased use of synonymous codons. However, the preferential use of codons appears to be marked only for codons with the highest rates of aminoacyl-tRNA selection. Rapid selection in vivo is usually effected by elevation of the tRNA concentration for codons with moderate intrinsic speed (rate constant), not by choosing intrinsically fast codons. Despite a preference for high rate, there are quickly translated codons that are not commonly used, and common codons that are translated relatively slowly. Other factors are therefore more important than speed for some codons. Strong preference for rapid aminoacyl-tRNA selection is not observed in weakly expressed genes. Instead, there is a slight preference for slower aminoacyl-tRNA selection. The rate of aminoacyl-tRNA selection by a YNC codon is always greater than the rate of the corresponding YNU codon even though in many YNC/U pairs both codons react with the same elongation factor Tu/GTP/aminoacyl-tRNA complex. Thus, for these tRNAs, the differences between in vivo rate constants of tRNAs are dependent on the nature of anticodon base-pairing. However, no more general relationship is evident between codon/anticodon composition and rate of aminoacyl-tRNA selection. The frameshift method can be extended to all codons.

Base Sequence↗

A functional significance for codon third bases.

Most amino acids are specified by more than one trinucleotide codon. Here we show that amino acids of differing functional importance may be distinguished by the pattern of synonymous codon usage. GC-rich genes tend to be of a greater transcriptional (p<0.01) and mitogenic (p<0.0001) significance than AT-rich genes, consistent with GC-->AT mutational drift in methylated genomic regions. Third-base GC retention also identifies critical amino acids within individual proteins, as indicated by non-random patterns of codon variation between gene homologs and also by differential sequelae of site-directed mutagenesis. Sequence analysis of human receptor tyrosine kinase genes confirms that functionally important transmembrane hydrophobic amino acids are specified by codons containing GC third bases more often than are transmembrane neutral amino acids (chi(2)=134.2). Amino acids encoded by GC third bases thus appear more tightly linked to cell function and survival than are those encoded by AT third bases.

Amino Acids↗

Evolution of base composition in the insulin and insulin-like growth factor genes.

The genomes of homeothermic (warm-blooded) vertebrates are mosaic interspersions of homogeneously GC-rich and GC-poor regions (isochores). Evolution of genome compartmentalization and GC-rich isochores is hypothesized to reflect either selective advantages of an elevated GC content or chromosome location and mutational pressure associated with the timing of DNA replication in germ cells. To address the present controversy regarding the origins and maintenance of isochores in homeothermic vertebrates, newly obtained as well as published nucleotide sequences of the insulin and insulin-like growth factor (IGF) genes, members of a well-characterized gene family believed to have evolved by repeated duplication and divergence, were utilized to examine the evolution of base composition in nonconstrained (flanking) and weakly constrained (introns and fourfold degenerate sites) regions. A phylogeny derived from amino acid sequences supports a common evolutionary history for the insulin/IGF family genes. In cold-blooded vertebrates, insulin and the IGFs were similar in base composition. In contrast, insulin and IGF-II demonstrate dramatic increases in GC richness in mammals, but no such trend occurred in IGF-I. Base composition of the coding portions of the insulin and IGF genes across vertebrates correlated (r = 0.90) with that of the introns and flanking regions. The GC content of homologous introns differed dramatically between insulin/IGF-II and IGF-I genes in mammals but was similar to the GC level of noncoding regions in neighboring genes. Our findings suggest that the base composition of introns and flanking regions is determined by chromosomal location and the mutational pressure of the isochore in which the sequences are embedded. An elevated GC content at codon third positions in the insulin and the IGF genes may reflect selective constraints on the usage of synonymous codons.

Animals↗

The relation between codon usage, base correlation and gene expression level in Escherichia coli and yeast.

Based on the investigation of the relation between gene expression and the usage of synonymous codons, a method of classifying and predicting the gene expression level is proposed which is called the Self-consistent Information Clustering (SCIC). Using the modified Codon Adaption Index (CAI) values, we have accomplished the linear regression analysis on the relation between base composition, base correlation and gene expression level in Escherichia coli and yeast. The assumption of Expression-Enhancing-Network Site (EENS) is proposed, the existence of which can be demonstrated by the linear equations between gene expression and base correlations in a codon, in adjacent codons and in non-adjacent codons. The modes of base correlation of E. coli and yeast which are important to gene expression have been found and listed in this paper.

Base Sequence↗

Comparison of the nucleotide sequences of the meta-cleavage pathway genes of TOL plasmid pWW0 from Pseudomonas putida with other meta-cleavage genes suggests that both single and multiple nucleotide substitutions contribute to enzyme evolution.

TOL plasmid pWW0 from Pseudomonas putida mt-2 encodes catabolic enzymes required for the oxidation of toluene and xylenes. The structural genes for these catabolic enzymes are clustered into two operons, the xylCMABN operon, which encodes a set of enzymes required for the transformation of toluene/xylenes to benzoate/toluates, and the xylXYZLTEGFJQKIH operon, which encodes a set of enzymes required for the transformation of benzoate/toluates to Krebs cycle intermediates. The latter operon can be divided physically and functionally into two parts, the xylXYZL cluster, which is involved in the transformation of benzoate/toluates to (methyl)catechols, and the xylTEGFJQKIH cluster, which is involved in the transformation of (methyl)catechols to Krebs cycle intermediates. Genes isofunctional to xylXYZL are present in Acinetobacter calcoaceticus, and constitute a benzoate-degradative pathway, while xylTEGFJQKIH homologous encoding enzymes of a methylphenol-degradative pathway and a naphthalene-degradative pathway are present on plasmid pVI150 from P. putida CF600, and on plasmid NAH7 from P. putida PpG7, respectively. Comparison of the nucleotide sequences of the xylXYZLTEGFJQKIH genes with other isofunctional genes suggested that the xylTEGFJQKIH genes on the TOL plasmid diverged from these homologues 20 to 50 million years ago, while the xylXYZL genes diverged from the A. calcoaceticus homologues 100 to 200 million years ago. In codons where amino acids are not conserved, the substitutions rate in the third base was higher than that in synonymous codons. This result was interpreted as indicating that both single and multiple nucleotide substitutions contributed to the amino acid-substituting mutations, and hence to enzyme evolution. This observation seems to be general because mammalian globin genes exhibit the same tendency.

Amino Acid Sequence↗

Selection pressures on codon usage in the complete genome of bacteriophage T7.

We searched the complete 39,936 base DNA sequence of bacteriophage T7 for nonrandomness that might be attributed to natural selection. Codon usage in the 50 genes of T7 is nonrandom, both over the whole code and among groups of synonymous codons. There is a great excess of purine- any base-pyrimidine (RNY) codons. Codon usage varies between genes, but from the pooled data for the whole genome (12,145 codons) certain putative selective constraints can be identified. Codon usage appears to be influenced by host tRNA abundance (particularly in highly expressed genes), tRNA-mRNA (one such interaction being perhaps responsible for maintaining the excess of RNY codons) and a lack of short palindromes. This last constraint is probably due to selection against host restriction enzyme recognition sites; this is the first report of an effect of this kind on codon usage. Selection against susceptibility to mutational damage does not appear to have been involved.

Base Sequence↗

Codon usage in plant peroxidase genes.

Codon preference and asymmetry in usage in the DNA sequences encoding the mature enzyme protein of 24 plant peroxidases from 12 different species were examined. Codon usage in highly conserved/non-conserved areas of the sequences was analysed, as well as possible deficiency/excess in CpG dinucleotides in the pairs of codon positions. Sequence relationships displayed by overall codon usage, dinucleotide frequencies within codons, and amino acid sequences were also studied. The main findings were: (1) Monocots clustered separately from dicots for overall codon usage and dinucleotide frequencies in codon positions 2 and 3, with six and seven clusters respectively discernible among these 24 peroxidase sequences. The monocot/dicot distinction disappeared in the four clusters among the mature protein amino acid sequences. Overall codon usage in sequences from monocotyledon and dicotyledon species differed, the monocots favouring codons with C or G in the third position. (2) Codon usage was biassed in many sequences, asymmetry was particularly noticeable in the monocots. (3) For repeated amino acids within conserved areas, codon preference appeared dependent on the order in which the repeated amino acid occurred, so that its usage of synonymous codons frequently balanced out.

Amino Acid Sequence↗

How optimized is the translational machinery in Escherichia coli, Salmonella typhimurium and Saccharomyces cerevisiae?

The optimization of the translational machinery in cells requires the mutual adaptation of codon usage and tRNA concentration, and the adaptation of tRNA concentration to amino acid usage. Two predictions were derived based on a simple deterministic model of translation which assumes that elongation of the peptide chain is rate-limiting. The highest translational efficiency is achieved when the codon recognized by the most abundant tRNA reaches the maximum frequency. For each codon family, the tRNA concentration is optimally adapted to codon usage when the concentration of different tRNA species matches the square-root of the frequency of their corresponding synonymous codons. When tRNA concentration and codon usage are well adapted to each other, the optimal content of all tRNA species carrying the same amino acid should match the square-root of the frequency of the amino acid. These predictions are examined against empirical data from Escherichia coli, Salmonella typhimurium, and Saccharomyces cerevisiae.

Codon↗

Codon usage in nucleopolyhedroviruses.

Phylogenetic analyses based on baculovirus polyhedrin nucleotide and amino acid sequences revealed two major nucleopolyhedrovirus (NPV) clades, designated Group I and Group II. Subsequent phylogenetic analyses have revealed three Group II subclades, designated A, B and C. Variations in amino acid frequencies determine the extent of dissimilarity for divergent but structurally and functionally conserved genes and therefore significantly influence the analysis of phylogenetic relationships. Hence, it is important to consider variations in amino acid codon usage. The Genome Hypothesis postulates that genes in any given genome use the same coding pattern with respect to synonymous codons and that genes in phylogenetically related species generally show the same pattern of codon usage. We have examined codon usage in six genes from six NPVs and found that: (1) there is significant variation in codon use by genes within the same virus genome; (2) there is significant variation in the codon usage of homologous genes encoded by different NPVs; (3) there is no correlation between the level of gene expression and codon bias in NPVs; (4) there is no correlation between gene length and codon bias in NPVs; and (5) that while codon use bias appears to be conserved between viruses that are closely related phylogenetically, the patterns of codon usage also appear to be a direct function of the GC-content of the virus-encoded genes.

Animals↗

Genetic code and optimal resistance to the effects of mutations.

This paper deals with the notion of resistance of the genetic code to the effects of mutations. We measure the resistance of a group of t codons as the number of pairs of those which differ from each other in only one of their three bases. We find for each value of t the maximum possible value of the resistance and we describe some groups of codons giving this value. Important examples of such configurations are found in the genetic code, among these are the groups of synonymous codons, as observed elsewhere, and the cluster of codons which have an hydrophobic amino acid for translation.

Base Sequence↗

Codon usage bias and base composition in MHC genes in humans and common chimpanzees.

Codon bias and base composition in major histocompatibility complex (MHC) sequences have been studied for both class I and II loci in Homo sapiens and Pan troglodytes. There is low to moderate codon bias for the MHC of humans and chimpanzees. In the class I loci, the same level of moderate codon bias is seen for HLA-B, HLA-C, Patr-A, Patr-B, and Patr-C, while at HLA-A the level of codon bias is lower. There is a correlation between codon usage bias and G+C content in the A and B loci in humans and chimps, but not at the C locus. To examine the effect of diversifying selection on codon bias, we subdivided class I alleles into antigen recognition site (ARS) and non-ARS codons. ARS codons had lower bias than non-ARS codons. This may indicate that the constraint of codon bias on nucleotide substitution may be selected against in ARS codons. At the class II loci, there are distinct differences between alpha and beta chain genes with respect to codon usage, with the beta chain genes being much more biased. Species-specific differences in base composition were seen in exon 2 at the DRB1 locus, with lower GC content in chimpanzees. Considering the complex evolutionary history of MHC genes, the study of codon usage patterns provides us with a better understanding of both the evolutionary history of these genes and the evolution of synonymous codon usage in genes under natural selection.

Animals↗

Translation rate modification by preferential codon usage: intragenic position effects.

We present a model for calculating the protein production rate as a function of the translation rate. The model takes into account that the elongation rate along an mRNA molecule is non-uniform as a result of different tRNA availabilities for different codons. Initiation of ribosomes on an mRNA is normally the rate-limiting step in the translation process, and blocking of the initiation site can be avoided if the codons closest to this site allow fast translation by the ribosome. Hence, different selective forces may act on the choice of synonymous codons in the initiation region than elsewhere on a given mRNA. We show that the elongation rate along the whole mRNA influences the production rate of abundant proteins, whereas only the elongation rate in the initiation region is of importance for the production rate of rare proteins. We also present an analysis of the codon distribution along known mRNAs coding for abundant and rare proteins.

Bacterial Proteins↗