Search PubMedSearch

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

De Novo Assembly and Comparative Analysis of the Complete Mitochondrial Genome of Mesenchytraeus (Annelida, Enchytraeidae).

The Changbai Mountain range is one of the key glacial refugia in Northeast Asia. Mesenchytraeus exhibits high species diversity, strong endemism, and widespread cryptic species in this region, for which mitogenomes provide useful molecular markers for exploring cryptic species complexes. This makes Mesenchytraeus an ideal model for studying mitogenome evolution among closely related lineages; however, no mitogenome data have been reported for this genus to date. In this study, we performed de novo assembly, annotation, and comparative analysis of the mitogenomes of 13 Mesenchytraeus species (14 individuals) from Changbai Mountain. All mitogenomes are typical circular molecules containing 37 genes, but putative control regions are rearranged and consistently located between ATP6 and trnR. All species exhibit annelid-specific strand nucleotide biases, characterized by negative GC skew and near-zero AT skew. Codon usage analysis reveals that codon families with wobble U are significantly biased toward mtDNA codons, whereas those with wobble C or G are biased toward non-mtDNA codons, suggesting a conserved mitochondrial codon usage pattern in annelids. All tRNAs form typical cloverleaf secondary structures except trnS2, which lacks the D-stem and the dihydrouridine (DHU) arm in some species. The putative control regions commonly contain complex palindromic repeats, hairpins, and repetitive elements, and may harbor dual replication origins. Phylogenetic analyses support the monophyly of Mesenchytraeus and reveal significant molecular divergence among morphologically cryptic species. This study provides the first mitogenome dataset for Mesenchytraeus and offers new insights into the evolution and replication mechanisms of mitogenomes in Clitellata and broader Annelida.

Mesenchytraeus

Usage of the three termination codons: compilation and analysis of the known eukaryotic and prokaryotic translation termination sequences.

The published translation termination sequences have been compiled and analysed to aid the interpretation of experiments on termination codon usage in the Xenopus oocyte (Bienz et al. 1981). There are significant differences between prokaryotes and eukaryotes concerning the usage of the three termination codons and of tandem stops. In addition viruses show termination strategies that differ from those of their hosts. Preferred context sequences flanking termination codons are described. Contexts vary within the last codon according to the nature of the termination codon, but are uniform within the first triplet following the terminators.

Animals

Multilayered nucleotide organization reveals purifying selection and host-driven adaptation in CPV and FPV.

Since feline panleukopenia virus (FPV) is considered the most likely ancestor of canine parvovirus (CPV), comprehensive comparisons of nucleotide organization in corresponding viral genes between CPV and FPV may provide novel insights into the evolutionary dynamics underlying the divergence of these two viruses. Here, we characterize the evolutionary patterns of CPV and FPV genes across multiple levels of nucleotide organization. Both viruses exhibited highly conserved nucleotide usage at nonsynonymous sites, with Ka/Ks patterns consistent with strong purifying selection, whereas synonymous sites showed greater variability. CpG dinucleotides were markedly underrepresented across all four viral genes, suggesting host-associated selective pressure and/or intrinsic nucleotide compositional constraints. Extensive nonrandom biases in synonymous codon usage, codon neighboring nucleotide context, and codon pair usage further revealed fine-scale genomic optimization shaped by natural selection and nucleotide compositional constraints. Structural protein genes (VP1 and VP2) displayed stronger codon usage bias and higher tRNA adaptation than nonstructural genes. Moreover, CPV genes showed greater translational adaptation to feline hosts than to canine hosts. These findings highlight how closely related parvoviruses exploit flexible nucleotide organization to facilitate host adaptation while maintaining essential protein functions.

Animals

On the rate of DNA sequence evolution in Drosophila.

Analysis of the rate of nucleotide substitution at silent sites in Drosophila genes reveals three main points. First, the silent rate varies (by a factor of two) among nuclear genes; it is inversely related to the degree of codon usage bias, and so selection among synonymous codons appears to constrain the rate of silent substitution in some genes. Second, mitochondrial genes may have evolved only as fast as nuclear genes with weak codon usage bias (and two times faster than nuclear genes with high codon usage bias); this is quite different from the situation in mammals where mitochondrial genes evolve approximately 5-10 times faster than nuclear genes. Third, the absolute rate of substitution at silent sites in nuclear genes in Drosophila is about three times higher than the average silent rate in mammals.

Animals

Beta tubulin gene of the parasitic protozoan Leishmania mexicana.

A genomic DNA library was generated with Sau3A cut DNA derived from promastigotes of Leishmania mexicana amazonensis and the lambda vector EMBL3. The library was screened for beta tubulin clones using 32P-labeled heterologous probe of chicken beta tubulin cDNA. From the various genomic clones the one designated 23.1, which gave the simplest hybridization banding pattern, was further characterized. The leishmanial insert DNA was subcloned into plasmid vectors and the resulting clones were designated as T11, T28 and T50. Using these clones leishmanial beta tubulin coding region was sequenced by the dideoxy method. The result shows that the beta tubulin has 445 amino acids, a carboxyl terminal tyrosine, and no intron. Leishmanial beta tubulin has 93% amino acid sequence similarity with that of trypanosome and 82% with that of man: and there is a strong bias in codon usage for codons possessing guanine or cytosine in the third base.

Amino Acid Sequence

On the relationship between preferred termination codon contexts and nonsense suppression in human cells.

The nucleotide sequences 3' to the translational termination codons in a collection of human genes have been analysed for evidence of a preferred 3' context for natural UAG codons. The aim was to see whether human UAG contexts can be related to the recent demonstration of the effects of 3' context on nonsense suppression in human cells. Since mammalian genomes are known to consist of a patchwork of blocks of sequences or 'isochores' with different G+C contents, the collection of genes was split into 5 classes containing genes with similar frequencies of G+C at the 3rd position of synonymous codons. This analysis revealed that the frequency of bases 3' to UAG varies with the G+C frequency of the gene, and that these changes were mirrored by changes in the patterns of bases in GN and AGN strings. The identity of the next 3' base appears therefore to be determined by genome wide changes in G+C composition, rather than selection to maintain a particular tetranucleotide stop signal. These findings argue strongly that the failure to find bias in the patterns of bases used in human coding sequences is an insensitive guide for the existence of codon usage or codon context effects during translation in human cells.

Base Composition

DNA sequences from the str operon of Escherichia coli.

The str operon at 72 min on the Escherichia coli chromosome contains genes for ribosomal proteins (r-proteins) S12 (str or rpsL) and S7 (rpsG) and elongation factors G (fus) and Tu (tufA). The sequence of the entire S12 gene, the S12-S7 intercistronic region, and the beginning of the S7 gene is reported. Also, the sequence of the end of the S7 gene, the S7-G intercistronic region, and the beginning of the elongation fractor G gene is reported. The S12-S7 intercistronic region is 96 base pairs long, in contrast to other intercistronic regions in r-protein operons which have been found to vary from 3 to 66 base pairs. The S7-G intercistronic region is only 27 bases long, supporting the previous conclusion that r-protein and elongation factor genes are co-transcribed. A comparison of translation initiation sites of the S12 and S7 genes, and other examples of co-transcribed r-protein genes, reveals no obvious features that could account for equimolar synthesis of all r-proteins. The codon usage in the S12 and S7 genes follows the pattern observed in other r-protein genes; that is, there is a highly preferential usage of codons recognized by the most abundant of isoaccepting tRNA species. This pattern could reflect the cell's need for efficient translation or minimal errors, or both, in r-protein synthesis.

Amino Acid Sequence

The primary structure of the alcohol dehydrogenase gene from the fission yeast Schizosaccharomyces pombe.

We have cloned and sequenced the alcohol dehydrogenase gene of the fission yeast Schizosaccharomyces pombe. The gene was isolated by transformation and complementation of a Saccharomyces cerevisiae strain which lacked functional alcohol dehydrogenase with an S. pombe gene bank constructed in the autonomously replicating yeast plasmid YEp13. Southern hybridization analysis indicates that S. pombe contains only one alcohol dehydrogenase gene. The structural region of the gene is 50% homologous to the alcohol dehydrogenase encoding genes of the budding yeast S. cerevisiae. The gene exhibits a very strong codon usage bias; with the set of predominantly used codons generally resembling that which S. cerevisiae employs preferentially. All of the differences in codon usage bias between S. pombe and S. cerevisiae are in the direction of greater G + C content in S. pombe codons. It is argued that this observation supports the hypothesis that selection toward uniform codon-anticodon binding energies contributes to codon usage bias and that the optimum binding energy is, on the average, higher in S. pombe than S. cerevisiae.

Alcohol Dehydrogenase

Complementary DNA and amino acid sequence of rat liver microsomal, xenobiotic epoxide hydrolase.

The coding nucleotide sequence for rat liver microsomal, xenobiotic epoxide hydrolase was determined from two overlapping cDNA clones, which together contain 1750 nucleotides complementary to epoxide hydrolase mRNA. The single open reading frame of 1365 nucleotides codes for a 455 amino acid polypeptide with a molecular weight of 52,581. The deduced amino acid composition agrees well with those determined by direct amino acid analysis of the rat protein, and the amino acid sequence is 81% identical to that of rabbit epoxide hydrolase. Analysis of codon usage for epoxide hydrolase, and that of rabbit epoxide hydrolase. Analysis of codon usage for epoxide hydrolase, and comparison to codon usage for NADPH-cytochrome P-450 oxidoreductase and cytochromes P-450b, P-450d, and P-450PCN, suggest that epoxide hydrolase is more conserved than cytochromes P-450b and P-450PCN; comparison of the extent of sequence conservation for 12 homologous proteins between the rat and rabbit, including cytochrome P-450b, supports this hypothesis, and indicates that much of epoxide hydrolase is constrained to maintain its hydrophobic character, consistent with its intramembranous location. The predicted membrane topology of epoxide hydrolase delineates 6 membrane-spanning segments, less than the 8 or 10 predicted for two cytochrome P-450 isozymes; the lower number of membrane-spanning segments predicted for epoxide hydrolase correlates with its lesser dependence on the membrane for maintenance of its tertiary structure and catalytic activity.

Amino Acid Sequence

Codon bias in actin multigene families and effects on the reconstruction of phylogenetic relationships.

Codon usage patterns and phylogenetic relationships in the actin multigene family have been analyzed for three dipteran species--Drosophila melanogaster, Bactrocera dorsalis, and Ceratitis capitata. In certain phylogenetic tree reconstructions, using synonymous distances, some gene relationships are altered due to a homogenization phenomenon. We present evidence to show that this homogenization phenomenon is due to codon usage bias. A survey of the pattern of synonymous codon preferences for 11 actin genes from these three species reveals that five out of the six Drosophila actin genes show high degrees of codon bias as indicated by scaled chi 2 values. In contrast to this, four out of the five actin genes from the other species have low codon bias values. A Monte Carlo contingency test indicates that for those Drosophila actin genes which exhibit codon bias, the patterns of codon usage are different compared to actin genes from the other species. In addition, the genes exhibiting codon bias also appear to have reduced rates of synonymous substitution. The homogenization phenomenon seen in terms of synonymous substitutions is not observed for nonsynonymous changes. Because of this homogenization phenomenon, "trees" constructed based on synonymous substitutions will be affected. These effects can be overt in the case of multigene families, but similar distortions may underlie reconstructions based on single-copy genes which exhibit codon usage bias.

Actins

Selection on silent sites in the rodent H3 histone gene family.

Selection promoting differential use of synonymous codons has been shown for several unicellular organisms and for Drosophila, but not for mammals. Selection coefficients operating on synonymous codons are likely to be extremely small, so that a very large effective population size is required for selection to overcome the effects of drift. In mammals, codon-usage bias is believed to be determined exclusively by mutation pressure, with differences between genes due to large-scale variation in base composition around the genome. The replication-dependent histone genes are expressed at extremely high levels during periods of DNA synthesis, and thus are among the most likely mammalian genes to be affected by selection on synonymous codon usage. We suggest that the extremely biased pattern of codon usage in the H3 genes is determined in part by selection. Silent site G + C content is much higher than expected based on flanking sequence G + C content, compared to other rodent genes with similar silent site base composition but lower levels of expression. Dinucleotide-mediated mutation bias does affect codon usage, but the affect is limited to the choice between G and C in some fourfold degenerate codons. Gene conversion between the two clusters of histone genes has not been an important force in the evolution of the H3 genes, but gene conversion appears to have had some effect within the cluster on chromosome 13.

Animals

Influence of the codon following the initiation codon on the expression of the lacZ gene in Saccharomyces cerevisiae.

A set of 32 different codons were introduced in a lacZ expression vector (pPTK400) immediately 3' from the AUG initiation codon. Expression of the lacZ gene was determined in Saccharomyces cerevisiae by measuring the amount of beta-galactosidase fusion protein using immuno-gel electrophoresis. A 5.3-fold difference in expression was found among the various constructs. It was found that there was no preference for a certain nucleotide in any position of the second codon and there was no distinct correlation between the level of tRNA corresponding to any particular second codon and expression. No correlation could be found between the local secondary structure and expression. When the overall codon usage in yeast and the codon usage in the second position of the mRNA is compared, there is no obvious significant difference in preference. This indicates that in yeast, in contrast to Escherichia coli, the codon choice at the beginning of the mRNA does not deviate from the one further downstream and is determined by the requirements for optimal translation elongation. Important determinants of the optimal context for an initiation codon in yeast therefore must be located mainly 5' from this codon.

Amino Acid Sequence

Interaction of silent and replacement changes in eukaryotic coding sequences.

We examined the codon usages in well-conserved and less-well-conserved regions of vertebrate protein genes and found them to be similar. Despite this similarity, there is a statistically significant decrease in codon bias in the less-well-conserved regions. Our analysis suggests that although those codon changes initially fixed under amino acid replacements tend to follow the overall codon usage pattern, they also reduce the bias in codon usage. This decrease in codon bias leads one to predict that the rate of change of synonymous codons should be greater in those regions that are less well conserved at the amino acid level than in the better-conserved regions. Our analysis supports this prediction. Furthermore, we demonstrate a significantly elevated rate of change of synonymous codons among the adjacent codons 5' to amino acid replacement positions. This provides further support for the idea that there are contextual constraints on the choice of synonymous codons in eukaryotes.

Cell Physiological Phenomena

Codon catalog usage and the genome hypothesis.

Frequencies for each of the 61 amino acid codons have been determined in every published mRNA sequence of 50 or more codons. The frequencies are shown for each kind of genome and for each individual gene. A surprising consistency of choices exists among genes of the same or similar genomes. Thus each genome, or kind of genome, appears to possess a "system" for choosing between codons. Frameshift genes, however, have widely different choice strategies from normal genes. Our work indicates that the main factors distinguishing between mRNA sequences relate to choices among degenerate bases. These systematic third base choices can therefore be used to establish a new kind of genetic distance, which reflects differences in coding strategy. The choice patterns we find seem compatible with the idea that the genome and not the individual gene is the unit of selection. Each gene in a genome tends to conform to its species' usage of the codon catalog; this is our genome hypothesis.

Animals

Origin and evolution of genes specifying resistance to macrolide, lincosamide and streptogramin antibiotics: data and hypotheses.

Resistance to macrolide, lincosamide and streptogramin antibiotics is due to alteration of the target site or detoxification of the antibiotic. Postranscriptional methylation of 23S ribosomal rRNA confers resistance to macrolide (M), lincosamide (L) and streptogramin (S) B-type antibiotics, the so-called MLSB phenotype. Several classes of rRNA methylases conferring resistance to MLSB antibiotics have been characterized in Gram-positive cocci, in Bacillus spp, and in strains of actinomycetes producing erythromycin. The enzymes catalyze N6-dimethylation of an adenine residue situated in a highly conserved region of prokaryotic 23S rRNA. In this review, we compare the amino acid sequences of the rRNA methylases and analyze the codon usage in the corresponding erm (erythromycin resistance methylase) genes. The homology detected at the protein level is consistent with the notion that an ancestor of the erm genes was implicated in erythromycin resistance in a producing strain. However, the rRNA methylases of producers and non-producers present substantial sequence diversity. In Gram-positive bacteria the preferential codon usage in the erm genes reflects the guanosine plus cytosine content of the chromosome of the host. These observations suggest that the presence of erm genes in these micro-organisms is ancient. By contrast, it would appear that enterobacteria have acquired only recently an rRNA methylase gene of the ermB class from a Gram-positive coccus since the genes isolated in Escherichia coli and in Gram-positive cocci are highly homologous (homology greater than 98%) and present a codon usage typical of the latter micro-organisms. As opposed to the MLSB phenotype which results from a single biochemical mechanism, inactivation of structurally related antibiotics of the MLS group involves synthesis of various other enzymes. In enterobacteria, resistance to erythromycin and oleandomycin is due to production of erythromycin esterases which hydrolyze the lactone ring of the 14-membered macrolides. We recently reported the nucleotide sequence of ereA and ereB (erythromycin resistance esterase) genes which encode erythromycin esterases type I and II, respectively. The amino acid sequences of the two isozymes do not exhibit statistically significant homology. Analysis of codon usage in both genes suggests that esterase type I is indigenous to E. coli, whereas the type II enzyme was acquired by E. coli from a phylogenetically remote micro-organism. Inactivation of lincosamides, first reported in staphylococci and lactobacilli of animal origin, was also recently detected in Gram-positive cocci isolated from humans.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence

Statistical method for predicting protein coding regions in nucleic acid sequences.

Protein coding regions of a genome fragment can be mathematically predicted by studying variations in the statistical properties or by searching the signals characteristic of the junctions between the coding and non-coding regions. We propose here a new statistical method using correspondence analysis. This method does not use any reference codon set but takes into account the codon usage homogeneity along the studied genome fragment. Comparison with previously published methods especially the 'codon usage method' of Staden has been made, and two examples are presented here. Applications to analysis of prokaryotic operon and eukaryotic split genes are also discussed. Use of the method has also shown two structures not previously described: i) in the human prt gene, a strong triplet structure exists in a non-coding region; ii) in the human tp-a codon usage is not uniform between the different exons.

Algorithms

Directional mutation pressure and transfer RNA in choice of the third nucleotide of synonymous two-codon sets.

Bacterial species have diverged into a series of families, some with high G + C content in their DNA, and other with high A + T content, resulting, respectively, from G.C- and A.T-directional mutation pressures. Such mutation pressure (G.C/A.T pressure) may be an important determinant for codon usage. It has also been suggested that tRNA acts as a selective constraint for determining codon usage. We have studied the relation between G.C/A.T pressure and tRNA constraints in determining choice of the third nucleotide of eight two-codon sets, using codon usage data obtained from protein genes in four bacterial species, Mycoplasma capricolum, Bacillus subtilis, Escherichia coli, and Micrococcus luteus, and in liverwort (Marchantia polymorpha) chloroplasts. The genomic G + C contents of these range from 25% to 74%. The results demonstrate that tRNA levels act additively to A.T and G.C pressure in affecting contents of A (pairing with *UNN anticodons, in which *U indicates a 2-thiouridine derivative) and C (pairing with GNN anticodons) or G (pairing with CNN anticodons), respectively, in third nucleotide positions of codons.

Chloroplasts

Compositional properties of nuclear genes from Plasmodium falciparum.

We have analyzed the compositional distributions of coding sequences and their different codon positions, as well as the codon usage of the nuclear genes of Plasmodium falciparum, a parasite characterized by an extremely GC-poor genome. As expected, coding sequences are AT-rich, codon usage is strongly biased towards A or T in third codon positions, and some particular amino acids (aa) are especially abundant in the encoded proteins. Remarkably, however, no difference was detected between housekeeping (HK) and antigen (Ag) genes, in spite of differences in expression level and evolutionary constraints. Moreover, all the features found in P. falciparum are very similar to those found in a bacterium characterized by a very GC-poor genome, Staphylococcus aureus. These findings stress the importance of compositional constraints in determining codon usage and aa utilisation.

Amino Acids