Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Translational selection on codon usage in Xenopus laevis.

A correspondence analysis of codon usage in Xenopus laevis revealed that the first axis is strongly correlated with the base composition at third codon positions. The second axis discriminates between putatively highly expressed genes and the other coding sequences, with expression levels being confirmed by the analysis of Expressed sequence tag frequencies. The comparison of codon usage of the sequences displaying the extreme values on the second axis indicates that several codons are statistically more frequent among the highly expressed (mainly housekeeping) genes. Translational selection appears, therefore, to influence synonymous codon usage in Xenopus.

Amino Acids↗

Evolution of base composition in the insulin and insulin-like growth factor genes.

The genomes of homeothermic (warm-blooded) vertebrates are mosaic interspersions of homogeneously GC-rich and GC-poor regions (isochores). Evolution of genome compartmentalization and GC-rich isochores is hypothesized to reflect either selective advantages of an elevated GC content or chromosome location and mutational pressure associated with the timing of DNA replication in germ cells. To address the present controversy regarding the origins and maintenance of isochores in homeothermic vertebrates, newly obtained as well as published nucleotide sequences of the insulin and insulin-like growth factor (IGF) genes, members of a well-characterized gene family believed to have evolved by repeated duplication and divergence, were utilized to examine the evolution of base composition in nonconstrained (flanking) and weakly constrained (introns and fourfold degenerate sites) regions. A phylogeny derived from amino acid sequences supports a common evolutionary history for the insulin/IGF family genes. In cold-blooded vertebrates, insulin and the IGFs were similar in base composition. In contrast, insulin and IGF-II demonstrate dramatic increases in GC richness in mammals, but no such trend occurred in IGF-I. Base composition of the coding portions of the insulin and IGF genes across vertebrates correlated (r = 0.90) with that of the introns and flanking regions. The GC content of homologous introns differed dramatically between insulin/IGF-II and IGF-I genes in mammals but was similar to the GC level of noncoding regions in neighboring genes. Our findings suggest that the base composition of introns and flanking regions is determined by chromosomal location and the mutational pressure of the isochore in which the sequences are embedded. An elevated GC content at codon third positions in the insulin and the IGF genes may reflect selective constraints on the usage of synonymous codons.

Animals↗

The relation between codon usage, base correlation and gene expression level in Escherichia coli and yeast.

Based on the investigation of the relation between gene expression and the usage of synonymous codons, a method of classifying and predicting the gene expression level is proposed which is called the Self-consistent Information Clustering (SCIC). Using the modified Codon Adaption Index (CAI) values, we have accomplished the linear regression analysis on the relation between base composition, base correlation and gene expression level in Escherichia coli and yeast. The assumption of Expression-Enhancing-Network Site (EENS) is proposed, the existence of which can be demonstrated by the linear equations between gene expression and base correlations in a codon, in adjacent codons and in non-adjacent codons. The modes of base correlation of E. coli and yeast which are important to gene expression have been found and listed in this paper.

Base Sequence↗

Comparison of the nucleotide sequences of the meta-cleavage pathway genes of TOL plasmid pWW0 from Pseudomonas putida with other meta-cleavage genes suggests that both single and multiple nucleotide substitutions contribute to enzyme evolution.

TOL plasmid pWW0 from Pseudomonas putida mt-2 encodes catabolic enzymes required for the oxidation of toluene and xylenes. The structural genes for these catabolic enzymes are clustered into two operons, the xylCMABN operon, which encodes a set of enzymes required for the transformation of toluene/xylenes to benzoate/toluates, and the xylXYZLTEGFJQKIH operon, which encodes a set of enzymes required for the transformation of benzoate/toluates to Krebs cycle intermediates. The latter operon can be divided physically and functionally into two parts, the xylXYZL cluster, which is involved in the transformation of benzoate/toluates to (methyl)catechols, and the xylTEGFJQKIH cluster, which is involved in the transformation of (methyl)catechols to Krebs cycle intermediates. Genes isofunctional to xylXYZL are present in Acinetobacter calcoaceticus, and constitute a benzoate-degradative pathway, while xylTEGFJQKIH homologous encoding enzymes of a methylphenol-degradative pathway and a naphthalene-degradative pathway are present on plasmid pVI150 from P. putida CF600, and on plasmid NAH7 from P. putida PpG7, respectively. Comparison of the nucleotide sequences of the xylXYZLTEGFJQKIH genes with other isofunctional genes suggested that the xylTEGFJQKIH genes on the TOL plasmid diverged from these homologues 20 to 50 million years ago, while the xylXYZL genes diverged from the A. calcoaceticus homologues 100 to 200 million years ago. In codons where amino acids are not conserved, the substitutions rate in the third base was higher than that in synonymous codons. This result was interpreted as indicating that both single and multiple nucleotide substitutions contributed to the amino acid-substituting mutations, and hence to enzyme evolution. This observation seems to be general because mammalian globin genes exhibit the same tendency.

Amino Acid Sequence↗

Selection pressures on codon usage in the complete genome of bacteriophage T7.

We searched the complete 39,936 base DNA sequence of bacteriophage T7 for nonrandomness that might be attributed to natural selection. Codon usage in the 50 genes of T7 is nonrandom, both over the whole code and among groups of synonymous codons. There is a great excess of purine- any base-pyrimidine (RNY) codons. Codon usage varies between genes, but from the pooled data for the whole genome (12,145 codons) certain putative selective constraints can be identified. Codon usage appears to be influenced by host tRNA abundance (particularly in highly expressed genes), tRNA-mRNA (one such interaction being perhaps responsible for maintaining the excess of RNY codons) and a lack of short palindromes. This last constraint is probably due to selection against host restriction enzyme recognition sites; this is the first report of an effect of this kind on codon usage. Selection against susceptibility to mutational damage does not appear to have been involved.

Base Sequence↗

Correlation between sequence conservation of the 5' untranslated region and codon usage bias in Mus musculus genes.

The codon adaptation index (CAI) values of all protein-coding sequences of the full-length cDNA libraries of Mus musculus were computed based on the RIKEN mouse full-length cDNA library. We have also computed the extent of consensus in flanking sequences of the initiator ATG codon based on the 'relative entropy' values of respective nucleotide positions (from -20 to +12 bp relative to the initiator ATG codon) for each group of genes classified by CAI values. With regard to the two nucleotides positions (-3 and +4) known to be highly conserved in Kozak's consensus sequence, a clear correlation between CAI values and relative entropy values was observed at position -3 but this was not significant at position +4, although a significant correlation was found at position -1 of the consensus sequence. Further, although no correlation was observed at any additional positions, relative entropy values were very high at positions -4, -6, and -8 in genes with high CAI values. These findings suggest that the extent of conservation in the flanking sequence of the initiator ATG codon including Kozak's consensus sequence was an important factor in modulation of the translation efficiency as well as synonymous codon usage bias particularly in highly expressed genes.

5' Untranslated Regions↗

Codon usage in plant peroxidase genes.

Codon preference and asymmetry in usage in the DNA sequences encoding the mature enzyme protein of 24 plant peroxidases from 12 different species were examined. Codon usage in highly conserved/non-conserved areas of the sequences was analysed, as well as possible deficiency/excess in CpG dinucleotides in the pairs of codon positions. Sequence relationships displayed by overall codon usage, dinucleotide frequencies within codons, and amino acid sequences were also studied. The main findings were: (1) Monocots clustered separately from dicots for overall codon usage and dinucleotide frequencies in codon positions 2 and 3, with six and seven clusters respectively discernible among these 24 peroxidase sequences. The monocot/dicot distinction disappeared in the four clusters among the mature protein amino acid sequences. Overall codon usage in sequences from monocotyledon and dicotyledon species differed, the monocots favouring codons with C or G in the third position. (2) Codon usage was biassed in many sequences, asymmetry was particularly noticeable in the monocots. (3) For repeated amino acids within conserved areas, codon preference appeared dependent on the order in which the repeated amino acid occurred, so that its usage of synonymous codons frequently balanced out.

Amino Acid Sequence↗

How optimized is the translational machinery in Escherichia coli, Salmonella typhimurium and Saccharomyces cerevisiae?

The optimization of the translational machinery in cells requires the mutual adaptation of codon usage and tRNA concentration, and the adaptation of tRNA concentration to amino acid usage. Two predictions were derived based on a simple deterministic model of translation which assumes that elongation of the peptide chain is rate-limiting. The highest translational efficiency is achieved when the codon recognized by the most abundant tRNA reaches the maximum frequency. For each codon family, the tRNA concentration is optimally adapted to codon usage when the concentration of different tRNA species matches the square-root of the frequency of their corresponding synonymous codons. When tRNA concentration and codon usage are well adapted to each other, the optimal content of all tRNA species carrying the same amino acid should match the square-root of the frequency of the amino acid. These predictions are examined against empirical data from Escherichia coli, Salmonella typhimurium, and Saccharomyces cerevisiae.

Codon↗

Codon usage in nucleopolyhedroviruses.

Phylogenetic analyses based on baculovirus polyhedrin nucleotide and amino acid sequences revealed two major nucleopolyhedrovirus (NPV) clades, designated Group I and Group II. Subsequent phylogenetic analyses have revealed three Group II subclades, designated A, B and C. Variations in amino acid frequencies determine the extent of dissimilarity for divergent but structurally and functionally conserved genes and therefore significantly influence the analysis of phylogenetic relationships. Hence, it is important to consider variations in amino acid codon usage. The Genome Hypothesis postulates that genes in any given genome use the same coding pattern with respect to synonymous codons and that genes in phylogenetically related species generally show the same pattern of codon usage. We have examined codon usage in six genes from six NPVs and found that: (1) there is significant variation in codon use by genes within the same virus genome; (2) there is significant variation in the codon usage of homologous genes encoded by different NPVs; (3) there is no correlation between the level of gene expression and codon bias in NPVs; (4) there is no correlation between gene length and codon bias in NPVs; and (5) that while codon use bias appears to be conserved between viruses that are closely related phylogenetically, the patterns of codon usage also appear to be a direct function of the GC-content of the virus-encoded genes.

Animals↗

Genetic code and optimal resistance to the effects of mutations.

This paper deals with the notion of resistance of the genetic code to the effects of mutations. We measure the resistance of a group of t codons as the number of pairs of those which differ from each other in only one of their three bases. We find for each value of t the maximum possible value of the resistance and we describe some groups of codons giving this value. Important examples of such configurations are found in the genetic code, among these are the groups of synonymous codons, as observed elsewhere, and the cluster of codons which have an hydrophobic amino acid for translation.

Base Sequence↗

Codon usage bias and base composition in MHC genes in humans and common chimpanzees.

Codon bias and base composition in major histocompatibility complex (MHC) sequences have been studied for both class I and II loci in Homo sapiens and Pan troglodytes. There is low to moderate codon bias for the MHC of humans and chimpanzees. In the class I loci, the same level of moderate codon bias is seen for HLA-B, HLA-C, Patr-A, Patr-B, and Patr-C, while at HLA-A the level of codon bias is lower. There is a correlation between codon usage bias and G+C content in the A and B loci in humans and chimps, but not at the C locus. To examine the effect of diversifying selection on codon bias, we subdivided class I alleles into antigen recognition site (ARS) and non-ARS codons. ARS codons had lower bias than non-ARS codons. This may indicate that the constraint of codon bias on nucleotide substitution may be selected against in ARS codons. At the class II loci, there are distinct differences between alpha and beta chain genes with respect to codon usage, with the beta chain genes being much more biased. Species-specific differences in base composition were seen in exon 2 at the DRB1 locus, with lower GC content in chimpanzees. Considering the complex evolutionary history of MHC genes, the study of codon usage patterns provides us with a better understanding of both the evolutionary history of these genes and the evolution of synonymous codon usage in genes under natural selection.

Animals↗

Translation rate modification by preferential codon usage: intragenic position effects.

We present a model for calculating the protein production rate as a function of the translation rate. The model takes into account that the elongation rate along an mRNA molecule is non-uniform as a result of different tRNA availabilities for different codons. Initiation of ribosomes on an mRNA is normally the rate-limiting step in the translation process, and blocking of the initiation site can be avoided if the codons closest to this site allow fast translation by the ribosome. Hence, different selective forces may act on the choice of synonymous codons in the initiation region than elsewhere on a given mRNA. We show that the elongation rate along the whole mRNA influences the production rate of abundant proteins, whereas only the elongation rate in the initiation region is of importance for the production rate of rare proteins. We also present an analysis of the codon distribution along known mRNAs coding for abundant and rare proteins.

Bacterial Proteins↗

A cluster of vitellogenin genes in the Mediterranean fruit fly Ceratitis capitata: sequence and structural conservation in dipteran yolk proteins and their genes.

Four genes encoding the major egg yolk polypeptides of the Mediterranean fruit fly Ceratitis capitata, vitellogenins 1 and 2 (VG1 and VG2), were cloned, characterized and partially sequenced. The genes are located on the same region of chromosome 5 and are organized in pairs, each encoding the two polypeptides on opposite DNA strands. Restriction and nucleotide sequence analysis indicate that the gene pairs have arisen from an ancestral pair by a relatively recent duplication event. The transcribed part is very similar to that of the Drosophila melanogaster yolk protein genes Yp1, Yp2 and Yp3. The Vg1 genes have two introns at the same positions as those in D. melanogaster Yp3; the Vg2 genes have only one of the introns, as do D. melanogaster Yp1 and Yp2. Comparison of the five polypeptide sequences shows extensive homology, with 27% of the residues being invariable. The sequence similarity of the processed proteins extends in two regions separated by a nonconserved region of varying size. Secondary structure predictions suggest a highly conserved secondary structure pattern in the two regions, which probably correspond to structural and functional domains. The carboxy-end domain of the C. capitata proteins shows the same sequence similarities with triacyglycerol lipases that have been reported previously for the D. melanogaster yolk proteins. Analysis of codon usage shows significant differences between D. melanogaster and C. capitata vitellogenins with the latter exhibiting a less biased representation of synonymous codons.

Alleles↗

Codon usage in Cryptosporidium parvum differs from that in other Eimeriorina.

Codon usage of Crytosporidium parvum was compared with those of other Eimeriorina Toxoplasma gondii and Eimeria tenella and revealed a biased use of synonymous codons with a preference for NNU (40.0%) and NNA (33.4%). There was no close resemblance of the codon usage of C. parvum to T. gondii (correlation coefficient, r = 0.14) or E. tenella (r = 0.14) but it was similar to Entamoeba histolytica (r = 0.75) and Plasmodium falciparum (r = 0.5). Analysis of the codon usage in homologous gene sequences (actin, beta-tubulin) also failed to reveal a close relationship between C. parvum and T. gondii or E. tenella. The low usage codons in C. parvum were most frequently used codons in T. gondii and E. tenella. These observations are consistent with 18S rRNA sequence analysis which shows no close relationship of Cryptosporidium with other Eimeriorina (Sarcocystis, Toxoplasma and Eimeria) and questions the validity of the current classification of C. parvum.

Actins↗

Molecular cloning of the chicken melanocortin 2 (ACTH)-receptor gene.

The chicken melanocortin 2-receptor (MC2-R) gene was isolated. It is found to be a single copy gene encoding a 357 amino acid protein, sharing 65.8-68.7% identity with mammalian counterparts. The chicken MC2-R mRNA is expressed in the adrenal and spleen, suggesting that the receptor mediates both endocrine and immunoregulatory functions of ACTH in the chicken. The amino acid sequence of the chicken MC2-R is collinear with those of other subtypes of MC-R, whereas all cloned mammalian MC2-Rs contain a gap in the third intracellular loop, suggesting that mammalian MC2-R molecules have evolved by lacking a part of the domain which determines the specificity of signal transduction in G-protein coupled receptors. Interestingly, the codon usage differs dramatically between MC1-R and MC2-R in the chicken; the GC-contents at the third codon position in MC1-R and MC2-R are 94.6 and 50.6%, respectively. It may reflect selective constraints on the usage of synonymous codons.

Adrenal Glands↗

Studies of codon usage and tRNA genes of 18 unicellular organisms and quantification of Bacillus subtilis tRNAs: gene expression level and species-specific diversity of codon usage based on multivariate analysis.

We examined codon usage in Bacillus subtilis genes by multivariate analysis, quantified its cellular levels of individual tRNAs, and found a clear constraint of tRNA contents on synonymous codon choice. Individual tRNA levels were proportional to the copy number of the respective tRNA genes. This indicates that the tRNA gene copy number is an important factor to determine in cellular tRNA levels, which is common with Escherichia coli and yeast Saccharomyces cerevisiae. Codon usage in 18 unicellular organisms whose genomes have been sequenced completely was analyzed and compared with the composition of tRNA genes. The 18 organisms are as follows: yeast S. cerevisiae, Aquifex aeolicus, Archaeoglobus fulgidus, B. subtilis, Borrelia burgdorferi, Chlamydia trachomatis, E. coli, Haemophilus influenzae, Helicobacterpylori, Methanococcusjannaschii, Methanobacterium thermoautotrophicum, Mycobacterium tuberculosis, Mycoplasma genitalium, Mycoplasma pneumoniae, Pyrococcus horikoshii, Rickettsia prowazekii, Synechocystis sp., and Treponema pallidum. Codons preferred in highly expressed genes were related to the codons optimal for the translation process, which were predicted by the composition of isoaccepting tRNA genes. Genes with specific codon usage are discussed in connection with their evolutionary origins and functions. The origin and terminus of replication could be predicted on the basis of codon usage when the usage was analyzed relative to the transcription direction of individual genes.

Archaea↗

Codon usage and bias among individual genes of the coccidia and piroplasms.

Codon usage has been analysed in individual gene sequences, derived from a variety of parasitic protozoa in the class Sporozoa of the phylum Apicomplexa using metric multidimensional scaling. The two groups of codon usage patterns detected reflect the two main subgroups of organisms studied (the coccidia and the piroplasms), and it is the pattern of usage of synonymous codons that has the largest influence on overall codon usage in the individual genes, rather than being the pattern of amino acid composition of the gene product. The magnitude of the codon usage bias in the sequences was determined using three commonly used indices-NC, GC3S and B. In general, although relatively low levels of codon usage bias were detected in these gene sequences, codon usage bias does explain at least some of the codon usage patterns observed. Codon usage bias was observed to be dependent on the overall base composition of the genes analysed, which in turn was reflected in the types of codons that were either over- or under-represented in the nucleotide sequences. In keeping with observations on prokaryotic organisms, it is speculated that the codon usage patterns detected in these parasitic protozoa are the result of directional mutation pressure on the base composition of the genomic DNA.

Animals↗

Codon usage and base composition in Rickettsia prowazekii.

Codon usage and base composition in sequences from the A + T-rich genome of Rickettsia prowazekii, a member of the alpha Proteobacteria, have been investigated. Synonymous codon usage patterns are roughly similar among genes, even though the data set includes genes expected to be expressed at very different levels, indicating that translational selection has been ineffective in this species. However, multivariate statistical analysis differentiates genes according to their G + C contents at the first two codon positions. To study this variation, we have compared the amino acid composition patterns of 21 R. prowazekii proteins with that of a homologous set of proteins from Escherichia coli. The analysis shows that individual genes have been affected by biased mutation rates to very different extents: genes encoding proteins highly conserved among other species being the least affected. Overall, protein coding and intergenic spacer regions have G + C content values of 32.5% and 21.4%, respectively. Extrapolation from these values suggests that R. prowazekii has around 800 genes and that 60-70% of the genome may be coding.

Base Composition↗