Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Substitution rates in Drosophila nuclear genes: implications for translational selection.

The relationships between synonymous and nonsynonymous substitution rates and between synonymous rate and codon usage bias are important to our understanding of the roles of mutation and selection in the evolution of Drosophila genes. Previous studies used approximate estimation methods that ignore codon bias. In this study we reexamine those relationships using maximum-likelihood methods to estimate substitution rates, which accommodate the transition/transversion rate bias and codon usage bias. We compiled a sample of homologous DNA sequences at 83 nuclear loci from Drosophila melanogaster and at least one other species of Drosophila. Our analysis was consistent with previous studies in finding that synonymous rates were positively correlated with nonsynonymous rates. Our analysis differed from previous studies, however, in that synonymous rates were unrelated to codon bias. We therefore conducted a simulation study to investigate the differences between approaches. The results suggested that failure to properly account for multiple substitutions at the same site and for biased codon usage by approximate methods can lead to an artifactual correlation between synonymous rate and codon bias. Implications of the results for translational selection are discussed.

Animals↗

Selection on the codon bias of chloroplast and cyanelle genes in different plant and algal lineages.

In the plant chloroplast genome the codon usage of the highly expressed psbA gene is unique and is adapted to the tRNA population, probably due to selection for translation efficiency. In this study the role of selection on codon usage in each of the fully sequenced chloroplast genomes, in addition to Chlamydomonas reinhardtii, is investigated by measuring adaptation to this pattern of codon usage. A method is developed which tests selection on each gene individually by constructing sequences with the same amino acid composition as the gene and randomly assigning codons based on the nucleotide composition of noncoding regions of that genome. The codon bias of the actual gene is then compared to a distribution of random sequences. The data indicate that within the algae selection is strong in Cyanophora paradoxa, affecting a majority of genes, of intermediate intensity in Odontella sinensis, and weaker in Porphyra purpurea and Euglena gracilis. In the plants, selection is found to be quite weak in Pinus thunbergii and the angiosperms but there is evidence that an intermediate level of selection exists in the liverwort Marchantia polymorpha. The role of selection is then further investigated in two comparative studies. It is shown that average relative codon bias is correlated with expression level and that, despite saturation levels of substitution, there is a strong correlation among the algae genomes in the degree of codon bias of homologous genes. All of these data indicate that selection for translation efficiency plays a significant role in determining the codon bias of chloroplast genes but that it acts with different intensities in different lineages. In general it is stronger in the algae than the higher plants, but within the algae Euglena is found to have several unusual features which are noted. The factors that might be responsible for this variation in intensity among the various genomes are discussed.

Chloroplasts↗

How optimized is the translational machinery in Escherichia coli, Salmonella typhimurium and Saccharomyces cerevisiae?

The optimization of the translational machinery in cells requires the mutual adaptation of codon usage and tRNA concentration, and the adaptation of tRNA concentration to amino acid usage. Two predictions were derived based on a simple deterministic model of translation which assumes that elongation of the peptide chain is rate-limiting. The highest translational efficiency is achieved when the codon recognized by the most abundant tRNA reaches the maximum frequency. For each codon family, the tRNA concentration is optimally adapted to codon usage when the concentration of different tRNA species matches the square-root of the frequency of their corresponding synonymous codons. When tRNA concentration and codon usage are well adapted to each other, the optimal content of all tRNA species carrying the same amino acid should match the square-root of the frequency of the amino acid. These predictions are examined against empirical data from Escherichia coli, Salmonella typhimurium, and Saccharomyces cerevisiae.

Codon↗

Gene synthesis, bacterial expression and purification of the Rickettsia prowazekii ATP/ADP translocase.

The Rickettsia prowazekii ATP/ADP translocase (Tlc) is the first member of a new family of ATP/ADP exchangers that includes both prokaryotic and eukaryotic proteins. We optimized the codon usage for expression of tlc in Escherichia coli by means of gene synthesis, expressed the synthetic gene in E. coli, and purified a modified Tlc that contained a C-terminal tag of 10 consecutive histidine residues by immobilized metal affinity chromatography. Although codon usage in R. prowazekii is very different from E. coli, the optimization of the codon usage by itself was insufficient to improve expression. However, the change of the cloning vector from pET11a to pT7-5 led to a 3-10-fold increase in the specific ATP transport rate by cells expressing the synthetic construct. The authenticity of the purified protein was confirmed by N-terminal amino acid sequencing and a matrix assisted laser desorption/ionization mass spectrometry.

Amino Acid Sequence↗

Contextual constraints on synonymous codon choice.

We have studied the statistical constraints on synonymous codon choice to evaluate various proposals regarding the origin of the bias in synonymous codon usage observed by Fiers et al. (1975), Air et al. (1976), Grantham et al. (1980) and others. We have determined the statistical dependence of the degenerate third base on either of its nearest neighbors in mitochondrial, prokaryotic, and eukaryotic coding sequences. We noted an increasing dependence of the third base on its nearest neighbors in moving from mitochondria to prokaryotes to eukaryotes. A statistical model assuming random equiprobable selection of synonymous codons was found grossly adequate for the mitochondria, but totally inadequate for prokaryotes and eukaryotes. A model assuming selection of synonymous codons reflecting a genomic strategy, i.e. the genome hypothesis of Grantham et al. (1980), gave a good approximation of the mitochondrial sequences. A statistical model which exactly maintains codon frequency, but allows the position of corresponding synonymous codons to vary was only grossly adequate for prokaryotes and totally inadequate for eukaryotes. The results of these simulations are consistent with the measures on experimental sequences and suggest that a "frequency constraint" model such as that of Grantham et al. (1980) may be an adequate explanation of the codon usage in mitochondria. However, in addition to this frequency constraint, there may be constraints on synonymous codon choice in prokaryotes due to codon context. Furthermore, any proposal to explain codon usage in eukaryotes must involve a constraint on the context of a codon in the sequence.

Amino Acid Sequence↗

Synonymous and nonsynonymous substitution rates in diatoms: a comparison between chloroplast and nuclear genes.

Rates of synonymous and nonsynonymous nucleotide substitutions and codon usage bias (ENC) were estimated for a number of nuclear and chloroplast genes in a sample of centric and pennate diatoms. The results suggest that DNA evolution has taken place, on an average, at a slower rate in the chloroplast genes than in the nuclear genes: a rate variation pattern similar to that observed in land plants. Synonymous substitution rates in the chloroplast genes show a negative association with the degree of codon usage bias, suggesting that genes with a higher degree of codon usage bias have evolved at a slower rate. While this relationship has been shown in both prokaryotes and multicellular eukaryotes, it has not been demonstrated before in diatoms.

Cell Nucleus↗

Coding sequence divergence between two closely related plant species: Arabidopsis thaliana and Brassica rapa ssp. pekinensis.

To characterize the coding-sequence divergence of closely related genomes, we compared DNA sequence divergence between sequences from a Brassica rapa ssp. pekinensis EST library isolated from flower buds and genomic sequences from Arabidopsis thaliana. The specific objectives were (i) to determine the distribution of and relationship between K(a) and K(s), (ii) to identify genes with the lowest and highest K(a): K(s) values, and (iii) to evaluate how codon usage has diverged between two closely related species. We found that the distribution of K(a): K(s) was unimodal, and that substitution rates were more variable at nonsynonymous than synonymous sites, and detected no evidence that K(a) and K(s) were positively correlated. Several genes had K(a): K(s) values equal to or near zero, as expected for genes that have evolved under strong selective constraint. In contrast, there were no genes with K(a): K(s) >1 and thus we found no strong evidence that any of the 218 sequences we analyzed have evolved in response to positive selection. We detected a stronger codon bias but a lower frequency of GC at synonymous sites in A. thaliana than B. rapa. Moreover, there has been a shift in the profile of most commonly used synonymous codons since these two species diverged from one another. This shift in codon usage may have been caused by stronger selection acting on codon usage or by a shift in the direction of mutational bias in the B. rapa phylogenetic lineage.

Arabidopsis↗

BIGPROBE: a computer program that predicts the sequence of long oligonucleotide probes with high reliability.

We have written a computer program, BIGPROBE, which facilitates the design of long nucleic acid probes from the partial or complete amino acid sequence of a protein. BIGPROBE relies upon information on codon usage, intercodon dinucleotide frequency, and potential probe self-complementarity. We have examined the accuracy with which the program predicts coding sequences using sample human and rat genes and probe lengths of 30-60 nucleotides. Rat probe sequences selected by BIGPROBE using either codon usage or dinucleotide frequency data alone averaged 86-92% homology with the known exons of the corresponding gene sequences. Predictive accuracy with rat gene probes could be improved to 89-94%, depending upon probe length, by applying codon usage and dinucleotide frequency data in combination. Similar accuracy was achieved for human genes.

Amino Acid Sequence↗

Codon equilibrium I: Testing for homogeneous equilibrium.

We present theoretical considerations that suggest that synonymous-codon usage might be expected to be close to an equilibrium distribution given a very homogeneous process of silent substitution. By homogeneous we mean that substitution depends only on the two bases involved, so that 12 base-substitution rates completely describe the silent substitution process. We have developed a method of statistically testing for such homogeneous equilibrium and applied it to reported data on the codon usages of different classes of organisms. Weakly expressed bacterial sequences and both mammalian and nonmammalian eukaryotic sequences deviate significantly from a random pattern of codon usage, in the direction of homogeneous equilibrium. On the other hand, highly expressed bacterial sequences do not exhibit homogeneous equilibrium, which may be correlated with recent experimental results showing that they are optimized to accept the most abundant tRNAs. To examine the effect of amino acid replacements on the homogeneous model of silent substitution, we divided the amino acids with degenerate codes into two classes, those with high mutabilities and those with low, and performed the same analysis on bacterial and eukaryotic data sets. The codon sets of the highly mutable class of amino acids are not further from homogeneous equilibrium than are the codon sets of the class with low mutabilities. We also found for the eukaryotic data that these independent classes of codon sets show very similar equilibrium patterns. The various results suggest a high level of uniformity in the process of silent fixation in the different synonymous-codon sets, especially in eukaryotes.

Amino Acid Sequence↗

Divergent evolutionary constraints on mitochondrial and nuclear genomes of malaria parasites.

Genetic variation among malaria parasites has important consequences with regard to drug resistance, pathogenicity, immunity, transmission, and speciation. In this regard, malaria parasites have been shown to display a high degree of inter- and intra-species genetic divergence. The nuclear genomes of Plasmodium falciparum, Plasmodium yoelii, and Plasmodium gallinaceum are vastly divergent yet share a similar codon usage and total A/T content of approximately 82%. This is in contrast to other primate-specific species including P. vivax which have an A/T content of approximately 67%. To assess the effects of this evolutionary divergence on the conservation of gene content, organization, and codon usage in the mitochondrial DNA (mtDNA) of malaria parasites, we have cloned and sequenced the mitochondrial genome of Plasmodium vivax, and compared it with the mtDNAs of P. falciparum, P. yoelii, and P. gallinaceum. The P. vivax mitochondrial genome was found to be 5990 base pairs in length, and displayed a gene organization identical to that of P. falciparum, P. yoelii, and P. gallinaceum. Furthermore, there was a remarkable 90% conservation of sequence identity between the mitochondrial genomes of all four species. As an example of intra-species conservation, comparison of mtDNAs from two independently cloned P. falciparum isolates, Malay Camp and C10, revealed only a single nucleotide substitution. A/T content of the P. vivax mitochondrial genome was found to be identical to other species of Plasmodium, hence, we have postulated that the mitochondrial genomes of malaria parasites were refractory to the evolutionary shifts in nucleotide content seen among the nuclear genomes of malaria parasites. Among different Plasmodium species, the second position of mitochondrial codons were found to be the least prone to substitutions and displayed a significant bias in pyrimidines. These aspects of mitochondrial codon usage were distinct from the nuclear genome and may reflect functional aspects of decoding by the mitochondrial translational system.

Amino Acid Sequence↗

Effect of a rare leucine codon, TTA, on expression of a foreign gene in Streptomyces lividans.

Streptomyces are bacteria with a very high chromosomal G+C composition (> 70 mol%) and extremely biased codon usage. In order to investigate the relationship between codon usage and gene expression in Streptomyces, we used ssi (Streptomyces subtilisin inhibitor) as a reporter gene and monitored its secretory expression in S. lividans. In consequence of alteration of the native codons of Leu, Lys and Ser of ssi to minor ones by site-directed mutagenesis, i.e., Leu79-Leu80: CTG-CTC to TTA-TTA, Lys89: AAG to AAA, Ser108-Ser109: TCG-AGC to TCT-TCT, respectively, the production of SSI was reduced remarkably in the case of TTA codons, while it was slightly increased in the case of AAA and almost the same in TCT codons. This conspicuous decrease found for Leu codon replacement was probably due to the low availability of intracellular tRNA(Leu) (UUA), a product of bldA which has been reported to be expressed only during the late stage of growth.

Amino Acid Sequence↗

Molecular cloning, heterologous expression, and primary structure of the structural gene for the copper enzyme nitrous oxide reductase from denitrifying Pseudomonas stutzeri.

The nos genes of Pseudomonas stutzeri are required for the anaerobic respiration of nitrous oxide, which is part of the overall denitrification process. A nos-coding region of ca. 8 kilobases was cloned by plasmid integration and excision. It comprised nosZ, the structural gene for the copper-containing enzyme nitrous oxide reductase, genes for copper chromophore biosynthesis, and a supposed regulatory region. The location of the nosZ gene and its transcriptional direction were identified by using a series of constructs to transform Escherichia coli and express nitrous oxide reductase in the heterologous background. Plasmid pAV5021 led to a nearly 12-fold overexpression of the NosZ protein compared with that in the P. stutzeri wild type. The complete sequence of the nosZ gene, comprising 1,914 nucleotides, together with 282 nucleotides of 5'-flanking sequences and 238 nucleotides of 3'-flanking sequences was determined. An open reading frame coded for a protein of 638 residues (Mr, 70,822) including a presumed signal sequence of 35 residues for protein export. The presequence is in conformity with the periplasmic location of the enzyme. Another open reading frame of 2,097 nucleotides, in the opposite transcriptional direction to that of nosZ, was excluded by several criteria from representing the coding region for nitrous oxide reductase. Codon usage for nosZ of P. stutzeri showed a high G + C content in the degenerate codon position (83.9% versus an average of 60.2%) and relaxed codon usage for the Glu codon, characteristic features of Pseudomonas genes from other species. E. coli nitrous oxide reductase was purified to homogeneity. It had the Mr of the P. stutzeri enzyme but lacked the copper chromophore.

Amino Acid Sequence↗

Relationships between transcriptional and translational control of gene expression in Saccharomyces cerevisiae: a multiple regression analysis.

Natural selection for an increased translation efficiency has been proposed as the main determinant for the bias in codon usage observed in many genes of Saccharomyces cerevisiae. Recently, the efficiency of transcription of a large number of yeast genes has been determined, based on the cellular content of the respective mRNAs: this provides an additional dimension to the study of the multisep process of gene expression. Using a representative set of yeast genes with a known level of transcription, the relationship between transcriptional and translational steps was evaluated by a multiple linear regression model. This analysis demonstrated a positive correlation between the amount of transcript, given as the number of mRNA copies per cell for each individual gene, and indices evaluating the effects of translational selection on the corresponding codon usage pattern. This finding suggests a close association of the cellular mRNA content, regulated also at the transcriptional level, to its efficiency of translation, mediated by a fine-tuning of codon usage strategy. Moreover, multiple regression analysis demonstrated that the transcription level of a gene can be approximately predicted using indices of bias deriving from its nucleotide sequence. This allowed for an extensive investigation of uncharacterized regions of the complete genome sequence of S. cerevisiae, to detect new potential short protein coding genes that were not considered by previous searching procedures. Several small open reading frames exhibiting a statistically significant coding potential were thus identified as good candidates for functional analysis.

Amino Acid Sequence↗

Optimization of the synthesis of porcine somatotropin in Escherichia coli.

We report on the influence of choice of promoter and RNA polymerase, 5'-untranslated regions and ribosome binding sites, codon usage, leader peptide coding sequences and poly A tail in the 3'-untranslated region on the synthesis of porcine somatotropin (PST) in Escherichia coli. A total of 12 different constructs were tested in this study for the production of porcine somatotropin (PST) in E. coli. Several factors have significant effects on PST synthesis. In the presence of a strong promoter and a strong ribosome binding site, the next most important factor seems to be the combination of sequences at the 5'-end of the mRNA including both the 5'-untranslated region and the start of the coding sequence. Codon usage in the 5'-coding sequence per se is not important in determining the level of PST synthesis where high level expression is achieved from a strong ribosome binding site. However, where low level synthesis of recombinant PST (rPST) is achieved, codon usage in the 5'-coding sequence is important in determining the level of PST synthesis. Leader sequences dramatically reduce the level of PST synthesis. The presence of a poly A tail in the 3'-untranslated region has no significant effect on PST synthesis.

Animals↗

Chromosomal localization of the human hexabrachion (tenascin) gene and evidence for recent reduplication within the gene.

Using analysis of rodent-human somatic cell hybrids as well as in situ hybridization of hexabrachion cDNA probes to normal human metaphase chromosomes, we have localized the human hexabrachion gene to chromosome 9, bands q32-q34. We also put forward the hypothesis that there has been a recent reduplication of a small segment of the human hexabrachion gene. We support this hypothesis by comparison of codon usage in this segment of the gene to codon usage in the remainder of the gene. This hypothesis is also supported by comparison of the sequence of human hexabrachion to that of the chicken hexabrachion. In addition, the latter comparison shows that the reduplication most likely occurred after the divergence of mammalian and avian species.

Amino Acid Sequence↗

Synonymous codon preferences in bacteriophage T4: a distinctive use of transfer RNAs from T4 and from its host Escherichia coli.

Codon usage data of bacteriophage T4 genes were compiled and synonymous codon preferences were investigated in comparison with tRNA availabilities in an infected cell. Since the genome of T4 is highly AT rich and its codon usage pattern is significantly different from that of its host Escherichia coli, certain codons of T4 genes need to be translated by appropriate host transfer RNAs present in minor amounts. To avoid this predicament, T4 phage seems to direct the synthesis of its own tRNA molecules and these phage tRNAs are suggested to supplement the host tRNA population with isoacceptors that are normally present in minor amounts. A positive correlation was found in that the frequency of E. coli optimal codons in T4 genes increases as the number of protein monomers per phage particle increases. A negative correlation was also found between the number of protein monomers per phage and the frequency of "T4 optimal codons", which are defined as those codons that are efficiently recognized by T4 tRNAs. From these observations it was proposed that tRNAs from the host are predominantly used for translation of highly expressed T4 genes while tRNAs from T4 tend to be used for translation of weakly expressed T4 genes. This distinctive tRNA-usage in T4 may be an optimization of translational efficiency, and an adjustment of T4-encoded tRNAs to the synonymous codon preferences, which are largely influenced by the high genomic AT-content, would have occurred during evolution.

Bacteriophage T4↗

[Regularities of the nucleotide sequence at the 5'-end of the codon in Escherichia coli genes].

The frequencies of occurrence of nucleotides at the 5' side of codons have been determined in highly and weakly expressed genes from E. coli. Significant constraints on the nucleotide 5' to some codons were found in highly expressed genes. Certain rules of synonymous codon usage depending on the amino acid 3' of the codon were established. E. g., codon possessing quanosine in the third position (NNG) are preferred over NNA if the next amino acid is lysine (P less than 10(-5)). On the other hand, rules of synonymous codon usage in relation to 5' flanking nucleotide were found. For example, when coding for aspartic acid, GAC codon is preferred over GAU (P less than 0.001) if uridine is 5' to codon and on the contrary GAU is favoured (P less than 0.0001) if quanosine is at the 5' side of aspartic acid codon. These rules can be used in the chemical synthesis of genes designed for expression in E. coli.

Base Sequence↗

The contributions of replication orientation, gene direction, and signal sequences to base-composition asymmetries in bacterial genomes.

Asymmetries in base composition between the leading and the lagging strands have been observed previously in many prokaryotic genomes. Since a majority of genes is encoded on the leading strand in these genomes, previous analyses have not been able to determine the relative contribution to the base composition skews of replication processes and transcriptional and/or translational forces. Using qualitative graphical presentations and quantitative statistical analyses (analysis of variance), we have found that a significant proportion of the GC and AT skews can be attributed to replication orientation, i.e., the sequence of a gene is influenced by whether it is encoded on the leading or lagging strand. This effect of replication orientation on skews is independent of, and can be opposite in sign to, the effects of transcriptional or translational processes, such as selection for codon usage, amino acid preferences, expression levels (inferred from codon adaptation index), or potential short signal sequences (e.g., chi sequences). Mutational differences between the leading and the lagging strands are the most likely explanation for a significant proportion of the base composition skew in these bacterial genomes. The finding that base composition skews due to replication orientation are independent of those due to selection for function of the encoded protein may complicate the interpretation of phylogenetic relationships, conserved positions in nucleotide or amino acid sequence alignments, and codon usage patterns.

Analysis of Variance↗