Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Complementary DNA and amino acid sequence of rat liver microsomal, xenobiotic epoxide hydrolase.

The coding nucleotide sequence for rat liver microsomal, xenobiotic epoxide hydrolase was determined from two overlapping cDNA clones, which together contain 1750 nucleotides complementary to epoxide hydrolase mRNA. The single open reading frame of 1365 nucleotides codes for a 455 amino acid polypeptide with a molecular weight of 52,581. The deduced amino acid composition agrees well with those determined by direct amino acid analysis of the rat protein, and the amino acid sequence is 81% identical to that of rabbit epoxide hydrolase. Analysis of codon usage for epoxide hydrolase, and that of rabbit epoxide hydrolase. Analysis of codon usage for epoxide hydrolase, and comparison to codon usage for NADPH-cytochrome P-450 oxidoreductase and cytochromes P-450b, P-450d, and P-450PCN, suggest that epoxide hydrolase is more conserved than cytochromes P-450b and P-450PCN; comparison of the extent of sequence conservation for 12 homologous proteins between the rat and rabbit, including cytochrome P-450b, supports this hypothesis, and indicates that much of epoxide hydrolase is constrained to maintain its hydrophobic character, consistent with its intramembranous location. The predicted membrane topology of epoxide hydrolase delineates 6 membrane-spanning segments, less than the 8 or 10 predicted for two cytochrome P-450 isozymes; the lower number of membrane-spanning segments predicted for epoxide hydrolase correlates with its lesser dependence on the membrane for maintenance of its tertiary structure and catalytic activity.

Amino Acid Sequence

Influence of the codon following the initiation codon on the expression of the lacZ gene in Saccharomyces cerevisiae.

A set of 32 different codons were introduced in a lacZ expression vector (pPTK400) immediately 3' from the AUG initiation codon. Expression of the lacZ gene was determined in Saccharomyces cerevisiae by measuring the amount of beta-galactosidase fusion protein using immuno-gel electrophoresis. A 5.3-fold difference in expression was found among the various constructs. It was found that there was no preference for a certain nucleotide in any position of the second codon and there was no distinct correlation between the level of tRNA corresponding to any particular second codon and expression. No correlation could be found between the local secondary structure and expression. When the overall codon usage in yeast and the codon usage in the second position of the mRNA is compared, there is no obvious significant difference in preference. This indicates that in yeast, in contrast to Escherichia coli, the codon choice at the beginning of the mRNA does not deviate from the one further downstream and is determined by the requirements for optimal translation elongation. Important determinants of the optimal context for an initiation codon in yeast therefore must be located mainly 5' from this codon.

Amino Acid Sequence

Interaction of silent and replacement changes in eukaryotic coding sequences.

We examined the codon usages in well-conserved and less-well-conserved regions of vertebrate protein genes and found them to be similar. Despite this similarity, there is a statistically significant decrease in codon bias in the less-well-conserved regions. Our analysis suggests that although those codon changes initially fixed under amino acid replacements tend to follow the overall codon usage pattern, they also reduce the bias in codon usage. This decrease in codon bias leads one to predict that the rate of change of synonymous codons should be greater in those regions that are less well conserved at the amino acid level than in the better-conserved regions. Our analysis supports this prediction. Furthermore, we demonstrate a significantly elevated rate of change of synonymous codons among the adjacent codons 5' to amino acid replacement positions. This provides further support for the idea that there are contextual constraints on the choice of synonymous codons in eukaryotes.

Cell Physiological Phenomena

Codon catalog usage and the genome hypothesis.

Frequencies for each of the 61 amino acid codons have been determined in every published mRNA sequence of 50 or more codons. The frequencies are shown for each kind of genome and for each individual gene. A surprising consistency of choices exists among genes of the same or similar genomes. Thus each genome, or kind of genome, appears to possess a "system" for choosing between codons. Frameshift genes, however, have widely different choice strategies from normal genes. Our work indicates that the main factors distinguishing between mRNA sequences relate to choices among degenerate bases. These systematic third base choices can therefore be used to establish a new kind of genetic distance, which reflects differences in coding strategy. The choice patterns we find seem compatible with the idea that the genome and not the individual gene is the unit of selection. Each gene in a genome tends to conform to its species' usage of the codon catalog; this is our genome hypothesis.

Animals

Origin and evolution of genes specifying resistance to macrolide, lincosamide and streptogramin antibiotics: data and hypotheses.

Resistance to macrolide, lincosamide and streptogramin antibiotics is due to alteration of the target site or detoxification of the antibiotic. Postranscriptional methylation of 23S ribosomal rRNA confers resistance to macrolide (M), lincosamide (L) and streptogramin (S) B-type antibiotics, the so-called MLSB phenotype. Several classes of rRNA methylases conferring resistance to MLSB antibiotics have been characterized in Gram-positive cocci, in Bacillus spp, and in strains of actinomycetes producing erythromycin. The enzymes catalyze N6-dimethylation of an adenine residue situated in a highly conserved region of prokaryotic 23S rRNA. In this review, we compare the amino acid sequences of the rRNA methylases and analyze the codon usage in the corresponding erm (erythromycin resistance methylase) genes. The homology detected at the protein level is consistent with the notion that an ancestor of the erm genes was implicated in erythromycin resistance in a producing strain. However, the rRNA methylases of producers and non-producers present substantial sequence diversity. In Gram-positive bacteria the preferential codon usage in the erm genes reflects the guanosine plus cytosine content of the chromosome of the host. These observations suggest that the presence of erm genes in these micro-organisms is ancient. By contrast, it would appear that enterobacteria have acquired only recently an rRNA methylase gene of the ermB class from a Gram-positive coccus since the genes isolated in Escherichia coli and in Gram-positive cocci are highly homologous (homology greater than 98%) and present a codon usage typical of the latter micro-organisms. As opposed to the MLSB phenotype which results from a single biochemical mechanism, inactivation of structurally related antibiotics of the MLS group involves synthesis of various other enzymes. In enterobacteria, resistance to erythromycin and oleandomycin is due to production of erythromycin esterases which hydrolyze the lactone ring of the 14-membered macrolides. We recently reported the nucleotide sequence of ereA and ereB (erythromycin resistance esterase) genes which encode erythromycin esterases type I and II, respectively. The amino acid sequences of the two isozymes do not exhibit statistically significant homology. Analysis of codon usage in both genes suggests that esterase type I is indigenous to E. coli, whereas the type II enzyme was acquired by E. coli from a phylogenetically remote micro-organism. Inactivation of lincosamides, first reported in staphylococci and lactobacilli of animal origin, was also recently detected in Gram-positive cocci isolated from humans.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence

Statistical method for predicting protein coding regions in nucleic acid sequences.

Protein coding regions of a genome fragment can be mathematically predicted by studying variations in the statistical properties or by searching the signals characteristic of the junctions between the coding and non-coding regions. We propose here a new statistical method using correspondence analysis. This method does not use any reference codon set but takes into account the codon usage homogeneity along the studied genome fragment. Comparison with previously published methods especially the 'codon usage method' of Staden has been made, and two examples are presented here. Applications to analysis of prokaryotic operon and eukaryotic split genes are also discussed. Use of the method has also shown two structures not previously described: i) in the human prt gene, a strong triplet structure exists in a non-coding region; ii) in the human tp-a codon usage is not uniform between the different exons.

Algorithms

Directional mutation pressure and transfer RNA in choice of the third nucleotide of synonymous two-codon sets.

Bacterial species have diverged into a series of families, some with high G + C content in their DNA, and other with high A + T content, resulting, respectively, from G.C- and A.T-directional mutation pressures. Such mutation pressure (G.C/A.T pressure) may be an important determinant for codon usage. It has also been suggested that tRNA acts as a selective constraint for determining codon usage. We have studied the relation between G.C/A.T pressure and tRNA constraints in determining choice of the third nucleotide of eight two-codon sets, using codon usage data obtained from protein genes in four bacterial species, Mycoplasma capricolum, Bacillus subtilis, Escherichia coli, and Micrococcus luteus, and in liverwort (Marchantia polymorpha) chloroplasts. The genomic G + C contents of these range from 25% to 74%. The results demonstrate that tRNA levels act additively to A.T and G.C pressure in affecting contents of A (pairing with *UNN anticodons, in which *U indicates a 2-thiouridine derivative) and C (pairing with GNN anticodons) or G (pairing with CNN anticodons), respectively, in third nucleotide positions of codons.

Chloroplasts

Comparison of three actin-coding sequences in the mouse; evolutionary relationships between the actin genes of warm-blooded vertebrates.

We have determined the sequences of three recombinant cDNAs complementary to different mouse actin mRNAs that contain more than 90% of the coding sequences and complete or partial 3' untranslated regions (3'UTRs): pAM 91, complementary to the actin mRNA expressed in adult skeletal muscle (alpha sk actin); pAF 81, complementary to an actin mRNA that is accumulated in fetal skeletal muscle and is the major transcript in adult cardiac muscle (alpha c actin); and pAL 41, identified as complementary to a beta nonmuscle actin mRNA on the basis of its 3'UTR sequence. As in other species, the protein sequences of these isoforms are highly (greater than 93%) conserved, but the three mRNAs show significant divergence (13.8-16.5%) at silent nucleotide positions in their coding regions. A nucleotide region located toward the 5' end shows significantly less divergence (5.6-8.7%) among the three mouse actin mRNAs; a second region, near the 3' end, also shows less divergence (6.9%), in this case between the mouse beta and alpha sk actin mRNAs. We propose that recombinational events between actin sequences may have homogenized these regions. Such events distort the calculated evolutionary distances between sequences within a species. Codon usage in the three actin mRNAs is clearly different, and indicates that there is no strict relation between the tissue type, and hence the tRNA precursor pool, and codon usage in these and other muscle mRNAs examined. Analysis of codon usage in these coding sequences in different vertebrate species indicates two tendencies: increases in bias toward the use of G and C in the third codon position in paralogous comparisons (in the order alpha c less than beta less than alpha sk), and in orthologous comparisons (in the order chicken less than rodent less than man). Comparison of actin-coding sequences between species was carried out using the Perler method of analysis. As one moves backward in time, changes at silent sites first accumulate rapidly, then begin to saturate after -(30-40) million years (MY), and actually decrease between -400 and -500 MY. Replacements or silent substitutions therefore cannot be used as evolutionary clocks for these sequences over long periods. Other phenomena, such as gene conversion or isochore compartmentalization, probably distort the estimated divergence time.

Actins

Nucleotide sequence of the structural gene for tryptophanase of Escherichia coli K-12.

The tryptophanase structural gene, tnaA, of Escherichia coli K-12 was cloned and sequenced. The size, amino acid composition, and sequence of the protein predicted from the nucleotide sequence agree with protein structure data previously acquired by others for the tryptophanase of E. coli B. Physiological data indicated that the region controlling expression of tnaA was present in the cloned segment. Sequence data suggested that a second structural gene of unknown function was located distal to tnaA and may be in the same operon. The pattern of codon usage in tnaA was intermediate between codon usage in four of the ribosomal protein structural genes and the structural genes for three of the tryptophan biosynthetic proteins.

Amino Acid Sequence

Selective differences among translation termination codons.

The frequency of use of the three alternative translation termination codons has been examined in 165 Escherichia coli, 52 Bacillus subtilis and 106 Saccharomyces cerevisiae genes. Genes were first categorised according to their degree of bias in sense codon usage. In each species there is a very strong bias in favour of UAA (over UAG and UGA) in genes where sense codon usage is highly biased. This bias declines, principally with an increase in the use of UGA, in genes with lower sense codon bias. It appears that selection operating during translation may maintain the bias in stop codon usage. Such selection could result from the greater availability of UAA-cognate release factor(s), or from a lower frequency of translational readthrough at UAA.

Bacillus subtilis

Correlations between the coding and non-coding regions in DNA.

In this paper various aspects of codon usage and k-tuple correlations in the DNA are compared. It is shown that the correlation structures of the coding and the non-coding regions are very similar and that codon usage is reasonably specific for large groups of organisms. These results suggest that the origin of codon usage is related to the origin and structure of the DNA.

Amino Acid Sequence

Contextual constraints on synonymous codon choice.

We have studied the statistical constraints on synonymous codon choice to evaluate various proposals regarding the origin of the bias in synonymous codon usage observed by Fiers et al. (1975), Air et al. (1976), Grantham et al. (1980) and others. We have determined the statistical dependence of the degenerate third base on either of its nearest neighbors in mitochondrial, prokaryotic, and eukaryotic coding sequences. We noted an increasing dependence of the third base on its nearest neighbors in moving from mitochondria to prokaryotes to eukaryotes. A statistical model assuming random equiprobable selection of synonymous codons was found grossly adequate for the mitochondria, but totally inadequate for prokaryotes and eukaryotes. A model assuming selection of synonymous codons reflecting a genomic strategy, i.e. the genome hypothesis of Grantham et al. (1980), gave a good approximation of the mitochondrial sequences. A statistical model which exactly maintains codon frequency, but allows the position of corresponding synonymous codons to vary was only grossly adequate for prokaryotes and totally inadequate for eukaryotes. The results of these simulations are consistent with the measures on experimental sequences and suggest that a "frequency constraint" model such as that of Grantham et al. (1980) may be an adequate explanation of the codon usage in mitochondria. However, in addition to this frequency constraint, there may be constraints on synonymous codon choice in prokaryotes due to codon context. Furthermore, any proposal to explain codon usage in eukaryotes must involve a constraint on the context of a codon in the sequence.

Amino Acid Sequence

BIGPROBE: a computer program that predicts the sequence of long oligonucleotide probes with high reliability.

We have written a computer program, BIGPROBE, which facilitates the design of long nucleic acid probes from the partial or complete amino acid sequence of a protein. BIGPROBE relies upon information on codon usage, intercodon dinucleotide frequency, and potential probe self-complementarity. We have examined the accuracy with which the program predicts coding sequences using sample human and rat genes and probe lengths of 30-60 nucleotides. Rat probe sequences selected by BIGPROBE using either codon usage or dinucleotide frequency data alone averaged 86-92% homology with the known exons of the corresponding gene sequences. Predictive accuracy with rat gene probes could be improved to 89-94%, depending upon probe length, by applying codon usage and dinucleotide frequency data in combination. Similar accuracy was achieved for human genes.

Amino Acid Sequence

Codon equilibrium I: Testing for homogeneous equilibrium.

We present theoretical considerations that suggest that synonymous-codon usage might be expected to be close to an equilibrium distribution given a very homogeneous process of silent substitution. By homogeneous we mean that substitution depends only on the two bases involved, so that 12 base-substitution rates completely describe the silent substitution process. We have developed a method of statistically testing for such homogeneous equilibrium and applied it to reported data on the codon usages of different classes of organisms. Weakly expressed bacterial sequences and both mammalian and nonmammalian eukaryotic sequences deviate significantly from a random pattern of codon usage, in the direction of homogeneous equilibrium. On the other hand, highly expressed bacterial sequences do not exhibit homogeneous equilibrium, which may be correlated with recent experimental results showing that they are optimized to accept the most abundant tRNAs. To examine the effect of amino acid replacements on the homogeneous model of silent substitution, we divided the amino acids with degenerate codes into two classes, those with high mutabilities and those with low, and performed the same analysis on bacterial and eukaryotic data sets. The codon sets of the highly mutable class of amino acids are not further from homogeneous equilibrium than are the codon sets of the class with low mutabilities. We also found for the eukaryotic data that these independent classes of codon sets show very similar equilibrium patterns. The various results suggest a high level of uniformity in the process of silent fixation in the different synonymous-codon sets, especially in eukaryotes.

Amino Acid Sequence

Molecular cloning, heterologous expression, and primary structure of the structural gene for the copper enzyme nitrous oxide reductase from denitrifying Pseudomonas stutzeri.

The nos genes of Pseudomonas stutzeri are required for the anaerobic respiration of nitrous oxide, which is part of the overall denitrification process. A nos-coding region of ca. 8 kilobases was cloned by plasmid integration and excision. It comprised nosZ, the structural gene for the copper-containing enzyme nitrous oxide reductase, genes for copper chromophore biosynthesis, and a supposed regulatory region. The location of the nosZ gene and its transcriptional direction were identified by using a series of constructs to transform Escherichia coli and express nitrous oxide reductase in the heterologous background. Plasmid pAV5021 led to a nearly 12-fold overexpression of the NosZ protein compared with that in the P. stutzeri wild type. The complete sequence of the nosZ gene, comprising 1,914 nucleotides, together with 282 nucleotides of 5'-flanking sequences and 238 nucleotides of 3'-flanking sequences was determined. An open reading frame coded for a protein of 638 residues (Mr, 70,822) including a presumed signal sequence of 35 residues for protein export. The presequence is in conformity with the periplasmic location of the enzyme. Another open reading frame of 2,097 nucleotides, in the opposite transcriptional direction to that of nosZ, was excluded by several criteria from representing the coding region for nitrous oxide reductase. Codon usage for nosZ of P. stutzeri showed a high G + C content in the degenerate codon position (83.9% versus an average of 60.2%) and relaxed codon usage for the Glu codon, characteristic features of Pseudomonas genes from other species. E. coli nitrous oxide reductase was purified to homogeneity. It had the Mr of the P. stutzeri enzyme but lacked the copper chromophore.

Amino Acid Sequence

Chromosomal localization of the human hexabrachion (tenascin) gene and evidence for recent reduplication within the gene.

Using analysis of rodent-human somatic cell hybrids as well as in situ hybridization of hexabrachion cDNA probes to normal human metaphase chromosomes, we have localized the human hexabrachion gene to chromosome 9, bands q32-q34. We also put forward the hypothesis that there has been a recent reduplication of a small segment of the human hexabrachion gene. We support this hypothesis by comparison of codon usage in this segment of the gene to codon usage in the remainder of the gene. This hypothesis is also supported by comparison of the sequence of human hexabrachion to that of the chicken hexabrachion. In addition, the latter comparison shows that the reduplication most likely occurred after the divergence of mammalian and avian species.

Amino Acid Sequence

Synonymous codon preferences in bacteriophage T4: a distinctive use of transfer RNAs from T4 and from its host Escherichia coli.

Codon usage data of bacteriophage T4 genes were compiled and synonymous codon preferences were investigated in comparison with tRNA availabilities in an infected cell. Since the genome of T4 is highly AT rich and its codon usage pattern is significantly different from that of its host Escherichia coli, certain codons of T4 genes need to be translated by appropriate host transfer RNAs present in minor amounts. To avoid this predicament, T4 phage seems to direct the synthesis of its own tRNA molecules and these phage tRNAs are suggested to supplement the host tRNA population with isoacceptors that are normally present in minor amounts. A positive correlation was found in that the frequency of E. coli optimal codons in T4 genes increases as the number of protein monomers per phage particle increases. A negative correlation was also found between the number of protein monomers per phage and the frequency of "T4 optimal codons", which are defined as those codons that are efficiently recognized by T4 tRNAs. From these observations it was proposed that tRNAs from the host are predominantly used for translation of highly expressed T4 genes while tRNAs from T4 tend to be used for translation of weakly expressed T4 genes. This distinctive tRNA-usage in T4 may be an optimization of translational efficiency, and an adjustment of T4-encoded tRNAs to the synonymous codon preferences, which are largely influenced by the high genomic AT-content, would have occurred during evolution.

Bacteriophage T4

[Regularities of the nucleotide sequence at the 5'-end of the codon in Escherichia coli genes].

The frequencies of occurrence of nucleotides at the 5' side of codons have been determined in highly and weakly expressed genes from E. coli. Significant constraints on the nucleotide 5' to some codons were found in highly expressed genes. Certain rules of synonymous codon usage depending on the amino acid 3' of the codon were established. E. g., codon possessing quanosine in the third position (NNG) are preferred over NNA if the next amino acid is lysine (P less than 10(-5)). On the other hand, rules of synonymous codon usage in relation to 5' flanking nucleotide were found. For example, when coding for aspartic acid, GAC codon is preferred over GAU (P less than 0.001) if uridine is 5' to codon and on the contrary GAU is favoured (P less than 0.0001) if quanosine is at the 5' side of aspartic acid codon. These rules can be used in the chemical synthesis of genes designed for expression in E. coli.

Base Sequence