Search PubMedSearch

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Statistical method for predicting protein coding regions in nucleic acid sequences.

Protein coding regions of a genome fragment can be mathematically predicted by studying variations in the statistical properties or by searching the signals characteristic of the junctions between the coding and non-coding regions. We propose here a new statistical method using correspondence analysis. This method does not use any reference codon set but takes into account the codon usage homogeneity along the studied genome fragment. Comparison with previously published methods especially the 'codon usage method' of Staden has been made, and two examples are presented here. Applications to analysis of prokaryotic operon and eukaryotic split genes are also discussed. Use of the method has also shown two structures not previously described: i) in the human prt gene, a strong triplet structure exists in a non-coding region; ii) in the human tp-a codon usage is not uniform between the different exons.

Algorithms

Directional mutation pressure and transfer RNA in choice of the third nucleotide of synonymous two-codon sets.

Bacterial species have diverged into a series of families, some with high G + C content in their DNA, and other with high A + T content, resulting, respectively, from G.C- and A.T-directional mutation pressures. Such mutation pressure (G.C/A.T pressure) may be an important determinant for codon usage. It has also been suggested that tRNA acts as a selective constraint for determining codon usage. We have studied the relation between G.C/A.T pressure and tRNA constraints in determining choice of the third nucleotide of eight two-codon sets, using codon usage data obtained from protein genes in four bacterial species, Mycoplasma capricolum, Bacillus subtilis, Escherichia coli, and Micrococcus luteus, and in liverwort (Marchantia polymorpha) chloroplasts. The genomic G + C contents of these range from 25% to 74%. The results demonstrate that tRNA levels act additively to A.T and G.C pressure in affecting contents of A (pairing with *UNN anticodons, in which *U indicates a 2-thiouridine derivative) and C (pairing with GNN anticodons) or G (pairing with CNN anticodons), respectively, in third nucleotide positions of codons.

Chloroplasts

Compositional properties of nuclear genes from Plasmodium falciparum.

We have analyzed the compositional distributions of coding sequences and their different codon positions, as well as the codon usage of the nuclear genes of Plasmodium falciparum, a parasite characterized by an extremely GC-poor genome. As expected, coding sequences are AT-rich, codon usage is strongly biased towards A or T in third codon positions, and some particular amino acids (aa) are especially abundant in the encoded proteins. Remarkably, however, no difference was detected between housekeeping (HK) and antigen (Ag) genes, in spite of differences in expression level and evolutionary constraints. Moreover, all the features found in P. falciparum are very similar to those found in a bacterium characterized by a very GC-poor genome, Staphylococcus aureus. These findings stress the importance of compositional constraints in determining codon usage and aa utilisation.

Amino Acids

Comparison of three actin-coding sequences in the mouse; evolutionary relationships between the actin genes of warm-blooded vertebrates.

We have determined the sequences of three recombinant cDNAs complementary to different mouse actin mRNAs that contain more than 90% of the coding sequences and complete or partial 3' untranslated regions (3'UTRs): pAM 91, complementary to the actin mRNA expressed in adult skeletal muscle (alpha sk actin); pAF 81, complementary to an actin mRNA that is accumulated in fetal skeletal muscle and is the major transcript in adult cardiac muscle (alpha c actin); and pAL 41, identified as complementary to a beta nonmuscle actin mRNA on the basis of its 3'UTR sequence. As in other species, the protein sequences of these isoforms are highly (greater than 93%) conserved, but the three mRNAs show significant divergence (13.8-16.5%) at silent nucleotide positions in their coding regions. A nucleotide region located toward the 5' end shows significantly less divergence (5.6-8.7%) among the three mouse actin mRNAs; a second region, near the 3' end, also shows less divergence (6.9%), in this case between the mouse beta and alpha sk actin mRNAs. We propose that recombinational events between actin sequences may have homogenized these regions. Such events distort the calculated evolutionary distances between sequences within a species. Codon usage in the three actin mRNAs is clearly different, and indicates that there is no strict relation between the tissue type, and hence the tRNA precursor pool, and codon usage in these and other muscle mRNAs examined. Analysis of codon usage in these coding sequences in different vertebrate species indicates two tendencies: increases in bias toward the use of G and C in the third codon position in paralogous comparisons (in the order alpha c less than beta less than alpha sk), and in orthologous comparisons (in the order chicken less than rodent less than man). Comparison of actin-coding sequences between species was carried out using the Perler method of analysis. As one moves backward in time, changes at silent sites first accumulate rapidly, then begin to saturate after -(30-40) million years (MY), and actually decrease between -400 and -500 MY. Replacements or silent substitutions therefore cannot be used as evolutionary clocks for these sequences over long periods. Other phenomena, such as gene conversion or isochore compartmentalization, probably distort the estimated divergence time.

Actins

Nucleotide sequence of the structural gene for tryptophanase of Escherichia coli K-12.

The tryptophanase structural gene, tnaA, of Escherichia coli K-12 was cloned and sequenced. The size, amino acid composition, and sequence of the protein predicted from the nucleotide sequence agree with protein structure data previously acquired by others for the tryptophanase of E. coli B. Physiological data indicated that the region controlling expression of tnaA was present in the cloned segment. Sequence data suggested that a second structural gene of unknown function was located distal to tnaA and may be in the same operon. The pattern of codon usage in tnaA was intermediate between codon usage in four of the ribosomal protein structural genes and the structural genes for three of the tryptophan biosynthetic proteins.

Amino Acid Sequence

Selective differences among translation termination codons.

The frequency of use of the three alternative translation termination codons has been examined in 165 Escherichia coli, 52 Bacillus subtilis and 106 Saccharomyces cerevisiae genes. Genes were first categorised according to their degree of bias in sense codon usage. In each species there is a very strong bias in favour of UAA (over UAG and UGA) in genes where sense codon usage is highly biased. This bias declines, principally with an increase in the use of UGA, in genes with lower sense codon bias. It appears that selection operating during translation may maintain the bias in stop codon usage. Such selection could result from the greater availability of UAA-cognate release factor(s), or from a lower frequency of translational readthrough at UAA.

Bacillus subtilis

Correlations between the coding and non-coding regions in DNA.

In this paper various aspects of codon usage and k-tuple correlations in the DNA are compared. It is shown that the correlation structures of the coding and the non-coding regions are very similar and that codon usage is reasonably specific for large groups of organisms. These results suggest that the origin of codon usage is related to the origin and structure of the DNA.

Amino Acid Sequence

Contextual constraints on synonymous codon choice.

We have studied the statistical constraints on synonymous codon choice to evaluate various proposals regarding the origin of the bias in synonymous codon usage observed by Fiers et al. (1975), Air et al. (1976), Grantham et al. (1980) and others. We have determined the statistical dependence of the degenerate third base on either of its nearest neighbors in mitochondrial, prokaryotic, and eukaryotic coding sequences. We noted an increasing dependence of the third base on its nearest neighbors in moving from mitochondria to prokaryotes to eukaryotes. A statistical model assuming random equiprobable selection of synonymous codons was found grossly adequate for the mitochondria, but totally inadequate for prokaryotes and eukaryotes. A model assuming selection of synonymous codons reflecting a genomic strategy, i.e. the genome hypothesis of Grantham et al. (1980), gave a good approximation of the mitochondrial sequences. A statistical model which exactly maintains codon frequency, but allows the position of corresponding synonymous codons to vary was only grossly adequate for prokaryotes and totally inadequate for eukaryotes. The results of these simulations are consistent with the measures on experimental sequences and suggest that a "frequency constraint" model such as that of Grantham et al. (1980) may be an adequate explanation of the codon usage in mitochondria. However, in addition to this frequency constraint, there may be constraints on synonymous codon choice in prokaryotes due to codon context. Furthermore, any proposal to explain codon usage in eukaryotes must involve a constraint on the context of a codon in the sequence.

Amino Acid Sequence

BIGPROBE: a computer program that predicts the sequence of long oligonucleotide probes with high reliability.

We have written a computer program, BIGPROBE, which facilitates the design of long nucleic acid probes from the partial or complete amino acid sequence of a protein. BIGPROBE relies upon information on codon usage, intercodon dinucleotide frequency, and potential probe self-complementarity. We have examined the accuracy with which the program predicts coding sequences using sample human and rat genes and probe lengths of 30-60 nucleotides. Rat probe sequences selected by BIGPROBE using either codon usage or dinucleotide frequency data alone averaged 86-92% homology with the known exons of the corresponding gene sequences. Predictive accuracy with rat gene probes could be improved to 89-94%, depending upon probe length, by applying codon usage and dinucleotide frequency data in combination. Similar accuracy was achieved for human genes.

Amino Acid Sequence

Codon equilibrium I: Testing for homogeneous equilibrium.

We present theoretical considerations that suggest that synonymous-codon usage might be expected to be close to an equilibrium distribution given a very homogeneous process of silent substitution. By homogeneous we mean that substitution depends only on the two bases involved, so that 12 base-substitution rates completely describe the silent substitution process. We have developed a method of statistically testing for such homogeneous equilibrium and applied it to reported data on the codon usages of different classes of organisms. Weakly expressed bacterial sequences and both mammalian and nonmammalian eukaryotic sequences deviate significantly from a random pattern of codon usage, in the direction of homogeneous equilibrium. On the other hand, highly expressed bacterial sequences do not exhibit homogeneous equilibrium, which may be correlated with recent experimental results showing that they are optimized to accept the most abundant tRNAs. To examine the effect of amino acid replacements on the homogeneous model of silent substitution, we divided the amino acids with degenerate codes into two classes, those with high mutabilities and those with low, and performed the same analysis on bacterial and eukaryotic data sets. The codon sets of the highly mutable class of amino acids are not further from homogeneous equilibrium than are the codon sets of the class with low mutabilities. We also found for the eukaryotic data that these independent classes of codon sets show very similar equilibrium patterns. The various results suggest a high level of uniformity in the process of silent fixation in the different synonymous-codon sets, especially in eukaryotes.

Amino Acid Sequence

Molecular cloning, heterologous expression, and primary structure of the structural gene for the copper enzyme nitrous oxide reductase from denitrifying Pseudomonas stutzeri.

The nos genes of Pseudomonas stutzeri are required for the anaerobic respiration of nitrous oxide, which is part of the overall denitrification process. A nos-coding region of ca. 8 kilobases was cloned by plasmid integration and excision. It comprised nosZ, the structural gene for the copper-containing enzyme nitrous oxide reductase, genes for copper chromophore biosynthesis, and a supposed regulatory region. The location of the nosZ gene and its transcriptional direction were identified by using a series of constructs to transform Escherichia coli and express nitrous oxide reductase in the heterologous background. Plasmid pAV5021 led to a nearly 12-fold overexpression of the NosZ protein compared with that in the P. stutzeri wild type. The complete sequence of the nosZ gene, comprising 1,914 nucleotides, together with 282 nucleotides of 5'-flanking sequences and 238 nucleotides of 3'-flanking sequences was determined. An open reading frame coded for a protein of 638 residues (Mr, 70,822) including a presumed signal sequence of 35 residues for protein export. The presequence is in conformity with the periplasmic location of the enzyme. Another open reading frame of 2,097 nucleotides, in the opposite transcriptional direction to that of nosZ, was excluded by several criteria from representing the coding region for nitrous oxide reductase. Codon usage for nosZ of P. stutzeri showed a high G + C content in the degenerate codon position (83.9% versus an average of 60.2%) and relaxed codon usage for the Glu codon, characteristic features of Pseudomonas genes from other species. E. coli nitrous oxide reductase was purified to homogeneity. It had the Mr of the P. stutzeri enzyme but lacked the copper chromophore.

Amino Acid Sequence

Optimization of the synthesis of porcine somatotropin in Escherichia coli.

We report on the influence of choice of promoter and RNA polymerase, 5'-untranslated regions and ribosome binding sites, codon usage, leader peptide coding sequences and poly A tail in the 3'-untranslated region on the synthesis of porcine somatotropin (PST) in Escherichia coli. A total of 12 different constructs were tested in this study for the production of porcine somatotropin (PST) in E. coli. Several factors have significant effects on PST synthesis. In the presence of a strong promoter and a strong ribosome binding site, the next most important factor seems to be the combination of sequences at the 5'-end of the mRNA including both the 5'-untranslated region and the start of the coding sequence. Codon usage in the 5'-coding sequence per se is not important in determining the level of PST synthesis where high level expression is achieved from a strong ribosome binding site. However, where low level synthesis of recombinant PST (rPST) is achieved, codon usage in the 5'-coding sequence is important in determining the level of PST synthesis. Leader sequences dramatically reduce the level of PST synthesis. The presence of a poly A tail in the 3'-untranslated region has no significant effect on PST synthesis.

Animals

Chromosomal localization of the human hexabrachion (tenascin) gene and evidence for recent reduplication within the gene.

Using analysis of rodent-human somatic cell hybrids as well as in situ hybridization of hexabrachion cDNA probes to normal human metaphase chromosomes, we have localized the human hexabrachion gene to chromosome 9, bands q32-q34. We also put forward the hypothesis that there has been a recent reduplication of a small segment of the human hexabrachion gene. We support this hypothesis by comparison of codon usage in this segment of the gene to codon usage in the remainder of the gene. This hypothesis is also supported by comparison of the sequence of human hexabrachion to that of the chicken hexabrachion. In addition, the latter comparison shows that the reduplication most likely occurred after the divergence of mammalian and avian species.

Amino Acid Sequence

Synonymous codon preferences in bacteriophage T4: a distinctive use of transfer RNAs from T4 and from its host Escherichia coli.

Codon usage data of bacteriophage T4 genes were compiled and synonymous codon preferences were investigated in comparison with tRNA availabilities in an infected cell. Since the genome of T4 is highly AT rich and its codon usage pattern is significantly different from that of its host Escherichia coli, certain codons of T4 genes need to be translated by appropriate host transfer RNAs present in minor amounts. To avoid this predicament, T4 phage seems to direct the synthesis of its own tRNA molecules and these phage tRNAs are suggested to supplement the host tRNA population with isoacceptors that are normally present in minor amounts. A positive correlation was found in that the frequency of E. coli optimal codons in T4 genes increases as the number of protein monomers per phage particle increases. A negative correlation was also found between the number of protein monomers per phage and the frequency of "T4 optimal codons", which are defined as those codons that are efficiently recognized by T4 tRNAs. From these observations it was proposed that tRNAs from the host are predominantly used for translation of highly expressed T4 genes while tRNAs from T4 tend to be used for translation of weakly expressed T4 genes. This distinctive tRNA-usage in T4 may be an optimization of translational efficiency, and an adjustment of T4-encoded tRNAs to the synonymous codon preferences, which are largely influenced by the high genomic AT-content, would have occurred during evolution.

Bacteriophage T4

[Regularities of the nucleotide sequence at the 5'-end of the codon in Escherichia coli genes].

The frequencies of occurrence of nucleotides at the 5' side of codons have been determined in highly and weakly expressed genes from E. coli. Significant constraints on the nucleotide 5' to some codons were found in highly expressed genes. Certain rules of synonymous codon usage depending on the amino acid 3' of the codon were established. E. g., codon possessing quanosine in the third position (NNG) are preferred over NNA if the next amino acid is lysine (P less than 10(-5)). On the other hand, rules of synonymous codon usage in relation to 5' flanking nucleotide were found. For example, when coding for aspartic acid, GAC codon is preferred over GAU (P less than 0.001) if uridine is 5' to codon and on the contrary GAU is favoured (P less than 0.0001) if quanosine is at the 5' side of aspartic acid codon. These rules can be used in the chemical synthesis of genes designed for expression in E. coli.

Base Sequence

Rare codons in E. coli and S. typhimurium signal sequences.

Codon usage has been examined in the signal sequences of 27 genes encoding proteins which possess leader peptides, and are inner-membrane located or exported. The results have been compared with codon usage in the corresponding coding sequences of most of the mature proteins. A bias is observed in the usage of rare codons for two of the three hydrophobic amino acids for which there are rare codons. Since hydrophobic residues are predominant in leader peptides, we suggest that a resulting concentration of rare codons in the signal sequence may play a role (or have played a role in the evolutionary past) in the secretion process by delaying translation.

Base Sequence

DNA Translator and Aligner: HyperCard utilities to aid phylogenetic analysis of molecules.

DNA Translator and Aligner are molecular phylogenetics HyperCard stacks for Macintosh computers. They manipulate sequence data to provide graphical gene mapping, conversions, translations and manual multiple-sequence alignment editing. DNA Translator is able to convert documented GenBank or EMBL documented sequences into linearized, rescalable gene maps whose gene sequences are extractable by clicking on the corresponding map button or by selection from a scrolling list. Provided gene maps, complete with extractable sequences, consist of nine metazoan, one yeast, and one ciliate mitochondrial DNAs and three green plant chloroplast DNAs. Single or multiple sequences can be manipulated to aid in phylogenetic analysis. Sequences can be translated between nucleic acids and proteins in either direction with flexible support of alternate genetic codes and ambiguous nucleotide symbols. Multiple aligned sequence output from diverse sources can be converted to Nexus, Hennig86 or PHYLIP format for subsequent phylogenetic analysis. Input or output alignments can be examined with Aligner, a convenient accessory stack included in the DNA Translator package. Aligner is an editor for the manual alignment of up to 100 sequences that toggles between display of matched characters and normal unmatched sequences. DNA Translator also generates graphic displays of amino acid coding and codon usage frequency relative to all other, or only synonymous, codons for approximately 70 select organism-organelle combinations. Codon usage data is compatible with spreadsheet or UWGCG formats for incorporation of additional molecules of interest. The complete package is available via anonymous ftp and is free for non-commercial uses.

Amino Acid Sequence

"Silent" sites in Drosophila genes are not neutral: evidence of selection among synonymous codons.

The patterns of synonymous codon usage in 91 Drosophila melanogaster genes have been examined. Codon usage varies strikingly among genes. This variation is associated with differences in G+C content at silent sites, but (unlike the situation in mammalian genes) these differences are not correlated with variation in intron base composition and so are not easily explicable in terms of mutational biases. Instead, those genes with high G+C content at silent sites, resulting from a strong "preference" for a particular subset of the codons that are mostly C-ending, appear to be the more highly expressed genes. This suggests that G+C content is reduced in sequences where selective constraints are weaker, as indeed seen in a pseudogene. These and other data discussed are consistent with the effects of translational selection among synonymous codons, as seen in unicellular organisms. The existence of selective constraints on silent substitutions, which may vary in strength among genes, has implications for the use of silent molecular clocks.

Animals