Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Evolutionary dynamics of the chloroplast genome in Abutilon (Malvoideae, Malvaceae).

The genus Abutilon Mill. (Malvaceae) comprises approximately 178 species distributed across tropical and subtropical regions, many of which hold significant ornamental, economic, and medicinal value; yet its taxonomic classification remains challenging. In this study, six species were sequenced from herbarium specimens, and the chloroplast (cp.) genomes of ten additional species were assembled de novo from publicly available raw data. Three previously reported cp. genomes were also incorporated to characterise cp. genome structure, identify polymorphic loci, and perform phylogenetic analyses. The cp. genomes ranged from 159,458 to 160,454 bp and exhibited the typical quadripartite structure, with each genome containing 112 unique genes (78 protein-coding, 30 tRNA, and 4 rRNA) that showed conserved content and organisation. These genomes exhibited high similarity in GC content, inverted repeat boundaries, relative synonymous codon usage, amino acid frequencies, and substitution patterns. However, notable variation was observed in the total number of simple sequence repeats, ranging from 70 to 97 per genome. Selection analyses indicated predominant purifying selection, with evidence of episodic positive selection detected in rpoC2, rbcL, and ycf1. Two codons in rbcL were clade-specific and provided phylogenetic signal distinguishing Australian and Old World pantropical species. Nucleotide diversity analysis identified six highly polymorphic intergenic spacers (trnH-psbA, rps19-rpl2, psbT-pbf1, psaC-ndhD, trnR-atpA, and ndhJ-ndhK) that may be suitable for taxonomic studies. The phylogeny from maximum likelihood (ML) and Bayesian inference (BI) resolved two major clades: one comprising an exclusively Australian lineage occurring predominantly in arid and semi-arid environments, and the other a pantropical lineage spanning multiple continents. Abutilon grandifolium was recovered as sister to the remaining sampled Abutilon taxa in both ML and BI analyses, although no biogeographic origin inference can be drawn from this placement pending broader taxon sampling and integration of nuclear genomic data. These findings provide insights into the evolutionary dynamics of the cp. genome in Abutilon and offer a foundational genomic framework for refining Abutilon taxonomy.

Genome, Chloroplast↗

Evolutionary rates and expression level in Chlamydomonas.

In many biological systems, especially bacteria and unicellular eukaryotes, rates of synonymous and nonsynonymous nucleotide divergence are negatively correlated with the level of gene expression, a phenomenon that has been attributed to natural selection. Surprisingly, this relationship has not been examined in many important groups, including the unicellular model organism Chlamydomonas reinhardtii. Prior to this study, comparative data on protein-coding sequences from C. reinhardtii and its close noninterfertile relative C. incerta were very limited. We compiled and analyzed protein-coding sequences for 67 nuclear genes from these taxa; the sequences were mostly obtained from the C. reinhardtii EST database and our C. incerta EST data. Compositional and synonymous codon usage biases varied among genes within each species but were highly correlated between the orthologous genes of the two species. Relative rates of synonymous and nonsynonymous substitution across genes varied widely and showed a strong negative correlation with the level of gene expression estimated by the codon adaptation index. Our comparative analysis of substitution rates in introns of lowly and highly expressed genes suggests that natural selection has a larger contribution than mutation to the observed correlation between evolutionary rates and gene expression level in Chlamydomonas.

Animals↗

Cloning and nucleotide sequencing of the genes for ribosomal proteins S9 (rpsI) and L13 (rplM) of Escherichia coli.

The genes for the ribosomal proteins S9 (rpsI) and L13 (rplM) of Escherichia coli have been cloned into a lambda phage vector termed L47.1. The two genes were identified by infecting UV-light irradiated cells with the resultant phages and analyzing the protein products by two-dimensional gel electrophoresis. Suitable DNA fragments of the isolate were cloned subsequently into M13 phage vectors and their nucleotide sequence was determined by the dideoxy method. It is evident that the two genes form a transcriptional unit, the rplM gene being promoter-proximal. There is a typical signal sequence for transcriptional termination after the rpsI gene. The codon usage pattern in the two genes is similar to other ribosomal protein genes of E. coli.

Amino Acid Sequence↗

Limitations of codon adaptation index and other coding DNA-based features for prediction of protein expression in Saccharomyces cerevisiae.

The relationship between codon usage and protein/mRNA expression in S. cerevisiae has been extensively studied. Recently, protein expression data for the whole yeast genome was published. We investigate which properties of coding DNA sequences can be used to predict expression levels. The new algorithm by Carbone et al. for computing dominating codon bias in a genome is evaluated. It is concluded that it works at least as well as existing methods, and eliminates the need to arbitrarily choose a set of highly expressed genes. Also, the hypothesis that information on codon pair frequencies can be used to predict expression is investigated. Our conclusion is that codon pairs do not contribute more information than do single codon frequencies. Overall correlation between predicted and actual expression data using properties of coding DNA sequences is around 0.65. Hence, while being a useful source of information, the expression levels predicted by these methods should only be used as a rule of thumb.

Algorithms↗

Reduced synonymous substitution rate at the start of enterobacterial genes.

Synonymous codon usage is less biased at the start of Escherichia coli genes than elsewhere. The rate of synonymous substitution between E.coli and Salmonella typhimurium is substantially reduced near the start of the gene, which suggests the presence of an additional selection pressure which competes with the selection for codons which are most rapidly translated. Possible competing sources of selection are the presence of secondary ribosome binding sites downstream from the start codon, the avoidance of mRNA secondary structure near the start of the gene and the use of sub-optimal codons to regulate gene expression. We provide evidence against the last of these possibilities. We also show that there is a decrease in the frequency of A, and an increase in the frequency of G along the E.coli genes at all three codon positions. We argue that these results are most consistent with selection to avoid mRNA secondary structure.

Base Composition↗

Nucleotide sequence, transcriptional analysis, and glucose regulation of the phenoxazinone synthase gene (phsA) from Streptomyces antibioticus.

The nucleotide sequence of a 2.3-kb SphI fragment containing the structural gene (phsA) for phenoxazinone synthase (PHS) of Streptomyces antibioticus was determined. The sequence was found to contain an open reading frame (ORF) with a G+C content of 71.5% oriented in the direction of transcription that was confirmed by primer extension. The ORF encodes a protein with an M(r) of 70,223 consisting of 642 amino acids and is preceded by a potential ribosome-binding site. The codon usage pattern is in agreement with the general pattern for streptomycete genes, with a 92.5 mol% G+C content in the third position. The N-terminal sequence of the mature PHS subunit corresponds exactly to that predicted from the nucleotide sequence. Neither ATG nor GTG initiator codons were identified for the protein. However, a TTG codon was located near the amino terminus of the mature protein and is a good candidate for the initiator codon. The transcriptional start point of phsA was located 36 bp upstream of the start codon by primer extension. The -10 region of the putative promoter showed some similarity to the consensus sequence for the major class of prokaryotic promoters, but the -35 region was less similar. Comparison of the primary amino acid sequence of PHS of S. antibioticus with other amino acid sequences indicated that PHS is a blue copper protein with copper binding domains in the N-terminal and C-terminal regions of the polypeptide chain. A BsrBI fragment containing the promoter region of phsA and a portion of the ORF was shown to promote xylE expression when cloned in the streptomycete promoter probe vector pIJ2843. This phsA promoter-dependent xylE expression could be repressed by glucose in S. antibioticus when the organism was grown on glucose or galactose plus glucose. Thus, the cloned promoter region appears to contain the sequences responsible for catabolite repression of PHS production.

Amino Acid Sequence↗

DNA sequence of the early E3 transcription unit of adenovirus 5.

The DNA sequence of the early E3 transcription unit of adenovirus 5 (Ad5) has been determined and it has been compared to Ad2, as published previously [J. Hérissé, G. Courtois, and F. Galibert (1980), Nucl. Acids Res. 8, 2173-2192; J. Hérissé and F. Galibert (1981), Nucl. Acids Res. 9, 1229-1240]. The E3 regions of Ad5 and Ad2 are quite homologous despite being nonessential for Ad growth in cultured cells. The major differences are "gaps" that exist either in Ad5 or Ad2 in intergenic regions. The conservation of sequences suggests that E3 plays a beneficial role in natural infection of humans. E3 appears to encode about seven to nine proteins; based on sequence, seven of these may be membrane proteins. Thus, E3 may be a transcription unit devoted to the synthesis of membrane proteins. The E3 genes lie essentially one after the other along the genome, and which gene is expressed from a given primary transcript is determined by the choice of the 3' end site and the 5' and 3' splice sites. Almost all E3 mRNAs contain nonfunctional AUGs that are 5' to the initiation codon. Codon usage is nonrandom. Although the CG dinucleotide frequency is low, CG clusters exist in the promoter and other regions.

Adenoviruses, Human↗

RNA primary sequence or secondary structure in the translational initiation region controls expression of two variant interferon-beta genes in Escherichia coli.

Efficient expression in Escherichia coli (E. coli) of the human interferon-beta gene (IFN-beta) gene and of a chemically synthesized IFN-beta gene variant (506 base pairs; synIFN-beta) adapted to the E. coli codon usage, both fused to the E. coli atpE ribosome-binding site, is controlled either by primary sequence or by mRNA secondary-structure in the translational initiation region. High level expression of the natural human atpE/IFN-beta gene fusion is governed by the nucleotide composition preceding the initiator codon AUG. A single U----C exchange in the -2 or -1 position preceding the initiator codon AUG reduces the translational efficiency from 18% of total cellular protein to only 8% or 4%, respectively, while both U----C substitutions reduce IFN-beta expression below 1%. These sequence alterations interfere with efficient ribosome binding as revealed by toeprinting. They provide further evidence for the influence of the anticodon-flanking regions of tRNA(fMet) upon the initiation rate of translation. In contrast, translation of the synthetic variant atpE/synIFN-beta gene fusion is controlled by a moderately stable stem-loop structure (delta G = -4 kcal/mol; 37 degrees C) located within the coding region and overlapping the 30 S ribosomal subunit attachment site. That the stability of the hairpin interferes with the initiation of translation is inferred from site-directed mutagenesis and toeprint analyses. mRNA half-life in these variants is positively correlated with the rate of translation and involves two major endonucleolytic cleavage site 5'-upstream of the Shine-Dalgarno region.

Base Sequence↗

Translation and stability of an Escherichia coli beta-galactosidase mRNA expressed under the control of pyruvate kinase sequences in Saccharomyces cerevisiae.

Plasmids were assembled in which the coding region of the pyruvate kinase (PYK) gene of Saccharomyces cerevisiae was replaced by that of the B-galactosidase (LacZ) gene from Escherichia coli. Analysis of the resultant, chimaeric transcripts from low copy number, centromeric plasmids indicated that this substitution caused a dramatic reduction in the steady-state level of the messenger RNA (mRNA). This fluctuation cannot be wholly accounted for by the 2-fold decrease in mRNA stability observed. This is consistent with the existence of a transcriptional Downstream Activation Site (DAS) within the PYK coding region, analogous to the DAS reported within the yeast phosphoglycerate kinase gene (PGK; Kingsman, S M et al. (1985) Biotech. Gen. Eng. Rev. 3, 377). At these low levels of heterologous gene expression, comparison of the distribution of PYK and PYK/LacZ transcripts across polysome gradients revealed no significant effect mediated by their striking disparity in codon usage. Nevertheless, upon increasing B-galactosidase mRNA levels, via manipulation of plasmid copy number, a distinct decline in ribosome loading was observed for the heterologous PYK/LacZ transcript which was not mirrored by either endogenous PYK transcripts or other yeast mRNAs of high (Ribosomal protein 1) or moderate (Actin) codon bias. However, high levels of the PYK/LacZ mRNA did affect the translation of an endogenous mRNA with poor codon bias (TRP2). The possible basis for this phenomenon is discussed.

Bacterial Proteins↗

Structure, evolution and expression of the mitochondrial ADP/ATP translocator gene from Chlamydomonas reinhardtii.

The first AUG in the Chlamydomonas reinhardtii ADP/ATP translocator (CRANT) mRNA initiates an open reading frame (ORF) which is very similar (51-79% amino acid identity) to other ANT proteins. In contrast to higher plants, no evidence for a long amino-terminal extension was obtained. The 5' non-transcribed region of the single-copy CRANT gene contains sequence motifs present in other C. reinhardtii nuclear genes. Four introns, whose positions are not conserved in other ANT genes, interrupt the protein coding region. A short heat shock specifically reduces CRANT mRNA levels. CRANT mRNA levels were unaffected by a mutation in photosynthesis. In a dark/light regime CRANT mRNA levels are high in the dark phase and low in the early light phase. Data on translation initiation sites, splice junctions and the codon preferences of C. reinhardtii nuclear genes were compiled. With the exception of two rare codons, ACA and GGA, the CRANT gene exhibits the biased codon usage of C. reinhardtii nuclear genes that are highly expressed during normal vegetative growth.

Amino Acid Sequence↗

Doublet preference and gene evolution.

Doublet preference analysis was carried out on coding and noncoding regions of Escherichia coli, Saccharomyces cerevisiae, and human mitochondrial and nuclear DNA. The preference pattern in 1-2 and 2-3 doublets in E. coli and S. cerevisiae correlated with that in noncoding regions. The 3-1 doublet preference in E. coli genes with low optimal codon frequency and in S. cerevisiae genes also showed a correlation with each of their noncoding doublet preference. A mechanism to explain these double preference correlations in doublet preference is presented: mutational biases, the origin of the noncoding region doublet preference, evolved so as to maintain the 1-2 and 2-3 doublet preference, which is determined by codon usage. These biases then acted on the 3-1 doublet, which was almost free of coding constraints, resulting in a similar preference in this doublet.

Base Composition↗

Sequence and organization of the Trichoplusia ni ascovirus 2c (Ascoviridae) genome.

The complete Trichoplusia ni ascovirus 2c (TnAV-2c) genome sequence was determined. The circular genome contains 174,059 bp with 165 open reading frames (ORFs) of greater than 180 bp and two major homologous regions (hrs). The genome is quite A+T rich at 64.6%. Fifty-four ORFs had homologues in other insect viruses, such as ascoviruses, iridoviruses, baculoviruses and entomopoxviruses; 30 ORFs showed low identities with those from different parasitic protozoa and 12 ORFs were unique to TnAV-2c. TnAV-2c has 15 ORFs that could be grouped into six gene families. Three major conserved repeating sequences were identified and were interspersed in two regions. BLAST analyses revealed that there were 16 enzymes involved in gene transcription, DNA replication, and nucleotide metabolism. TnAV-2c has 12 and 25 ORFs sharing high identities with ascovirus and iridovirus homologues, respectively. The codon usage bias appears to be more similar to Spodoptera frugiperda ascovirus 1a than to iridoviruses.

Animals↗

Nucleotide sequence of the gene for aqualysin I (a thermophilic alkaline serine protease) of Thermus aquaticus YT-1 and characteristics of the deduced primary structure of the enzyme.

Aqualysin I is an alkaline serine protease which is secreted into the culture medium by Thermus aquaticus YT-1, an extreme thermophile [Matsuzawa, H., Hamaoki, M. & Ohta, T. (1983) Agric. Biol. Chem. 47, 25-28]. The gene encoding aqualysin I was cloned into Escherichia coli using synthetic oligodeoxyribonucleotides as hybridization probes. The nucleotide sequence of the cloned DNA was determined. The primary structure of aqualysin I, deduced from the nucleotide sequence, agreed with the NH2-terminal sequence previously reported and the determined amino acid sequences, including the COOH-terminal sequence, of the tryptic peptides derived from aqualysin I. Aqualysin I comprised 281 amino acid residues and its molecular mass was determined to be 28,350. On alignment of the whole amino acid sequence, aqualysin I showed high sequence homology with the subtilisin-type serine proteases, and 43% identity with proteinase K, 37-39% with subtilisins and 34% with thermitase. Extremely high sequence identity was observed in the regions containing the active-site residues, corresponding to Asp32, His64 and Ser221 of subtilisin BPN'. The nucleotide sequence of the cloned DNA (1105 nucleotides) revealed that it contains the entire gene encoding aqualysin I and one open reading frame without a translational stop codon. Therefore, aqualysin I was considered to be produced as a large precursor, which contains a NH2-terminal portion, the protease and a COOH-terminal portion. The G + C content of the coding region for aqualysin I was 64.6%, which is lower than those of other Thermus genes (68-74%). The codon usage in the aqualysin I gene was rather random in comparison with that in other Thermus genes.

Amino Acid Sequence↗

Strength of the purifying selection against different categories of the point mutations in the coding regions of the human genome.

Using available Information on the total absolute size of the coding region of the human genome, data on codon usage and pseudogene-derived mutation rates for different single nucleotide substitutions we have estimated, for the human genome, the potential numbers of mutation events capable to produce: (1) nonsense; (2) missense (radical and conservative); (3) silent; (4) splice; and (5) protein-elongating (those changing wild-type stop codon into an amino acid encoding codon) mutations. We used the NCBI dbSNP database to retrieve data on the observed number of polymorphisms of each category. The fraction of polymorphisms in each category among all potential events in the genome depends on the strength of selection: the higher the rate of polymorphism, the weaker the selection. We used nonsense mutations as a referent group. Compared with nonsense mutations, we found that the relative selection coefficient against protein-elongating mutations was 21%, and the relative selection was 12% against missense mutations. Radical missense mutations were found to be four times more deleterious compared to conservative ones. Surprisingly, we found that silent mutations on average are not neutral; with the average harmfulness of 3% of nonsense mutations. Silent mutations may be deleterious when they affect splicing by creating cryptic donor-acceptor sites or by disturbing exonic splicing enhancers (ESESs). The average selection coefficient against splice mutations was 48% of that against nonsense mutations. Converting the relative selection coefficients into absolute ones using data on loss-of-function mutations in Saccharomyces cerevisiae and Caenorhabditis elegans, or by analysis of the expected frequency of mutations in the human genome, suggested that genetic drift could play a role in population dynamics of conservative missense and silent mutations.

Computational Biology↗

Reduction of wobble-position GC bases in Corynebacteria genes and enhancement of PCR and heterologous expression.

Corynebacteria codon usage exhibits an overall GC content of 67%, and a wobble-position GC content of 88%. Escherichia coli, on the other hand has an overall GC content of 51%, and a wobble-position GC content of 55%. The high GC content of Corynebacteria genes results in an unfavorable codon preference for heterologous expression, and can present difficulties for polymerase-based manipulations due to secondary-structure effects. Since these characteristics are due primarily to base composition at the wobble-position, synthetic genes can, in principle, be designed to eliminate these problems and retain the wild-type amino acid sequence. Such genes would obviate the need for special additives or bases during in vitro polymerase-based manipulation and mutant host strains containing uncommon tRNA's for heterologous expression. We have evaluated synthetic genes with reduced wobble-position G/C content using two variants of the enzyme 2,5-diketo-D-gluconic acid reductase (2,5-DKGR A and B) from Corynebacterium. The wild-type genes are refractory to polymerase-based manipulations and exhibit poor heterologous expression in enteric bacteria. The results indicate that a subset of codons for five amino acids (alanine, arginine, glutamate, glycine and valine) contribute the greatest contribution to reduction in G/C content at the wobble-position. Furthermore, changes in codons for two amino acids (leucine and proline) enhance bias for expression in enteric bacteria without affecting the overall G/C content. The synthetic genes are readily amplified using polymerase-based methodologies, and exhibit high levels of heterologous expression in E. coli.

Base Composition↗

Phylogenetic relationships of the liverworts (Hepaticae), a basal embryophyte lineage, inferred from nucleotide sequence data of the chloroplast gene rbcL.

Sequence data from the chloroplast-encoded gene rbcL were obtained for 24 liverworts, a basal group of embryophytes. Maximum likelihood and parsimony analyses of these data, along with data from other major green plant lineages, confirm hypotheses based on morphological data, such as the paraphyly of bryophytes, and the basal position of liverworts. Molecular data corroborate the deep separation between the complex thalloid and leafy/simple thalloid liverworts implied by morphological data, but the monophyly of liverworts could not be rejected. The effects of accounting for site-to-site rate heterogeneity in these data were examined using maximum likelihood methods. Comparison of trees obtained with and without rate heterogeneity showed that simply allowing for heterogeneity had a greater improvement on likelihood score than optimization of transition/transversion bias. Incorporation of site-to-site rate heterogeneity in the larger analysis, however, did not necessarily change which topology was favored. Properties of rbcL sequences from the two liverwort groups were compared. Significantly different substitution rates were found between leafy/simple thalloid and complex thalloid liverwort taxa, with rates of rbcL sequence evolution in leafy/simple thalloid taxa being higher and more indicative of those of vascular plants, and with those of complex thalloid taxa (such as Marchantia) being slower. Codon usage in rbcL in complex thalloid liverworts was biased toward NNU and NNA, compared to the leafy/simple thalloid liverworts. Although base composition and relative substitution rates differed between the two groups, no significant differences were detected within each of the two groups of liverworts. The signal present in first and second codon sites versus third codon sites was compared. While the third codon positions in rbcL across this taxon sampling are highly variable (with only 15 constant sites of 439), the trees obtained were in general agreement with trees from the entire data set and with trees obtained from independent sources of data. The presence of signal in third codon positions across greater than 400 MY of plant evolution means that definitions of saturation based on pair-wise comparisons of sequences inadequately assess phylogenetic signal.

Chloroplasts↗

Characterization of the str operon genes from Spirulina platensis and their evolutionary relationship to those of other prokaryotes.

A 5.3 kb DNA segment containing the str operon (ca. 4.5 kb) of the cyanobacterium Spirulina platensis has been sequenced. The str operon includes the structural genes rpsL (ribosomal protein S12), rpsG (ribosomal protein S7), fus (translation elongation factor EF-G) and tuf (translation elongation factor EF-Tu). From the nucleotide sequence of this operon, the primary structures of the four gene products have been derived and compared with the available corresponding structures from eubacteria, archaebacteria and chloroplasts. Extensive homologies were found in almost all cases and in the order S12 greater than EF-Tu greater than EF-G greater than S7; the largest homologies were generally found between the cyanobacterial proteins and the corresponding chloroplast gene products. Overall codon usage in S. platensis was found to be rather unbiased.

Amino Acid Sequence↗

Expression of a synthetic gene encoding a Tribolium castaneum carboxylesterase in Pichia pastoris.

This is the first report of an insect esterase efficiently expressed in the methylotrophic yeast Pichia pastoris (so far insect esterases have been produced only in the baculovirus system). Having isolated a Tribolium castaneum carboxylesterase cDNA (TCE), we were initially unable to express it in Escherichia coli or P. pastoris despite significant transcription levels. As codon usage bias is different in T. castaneum and P. pastoris, we assumed this was a possible explanation for the translational barrier observed in yeast. Accordingly, we designed and constructed by recursive PCR a synthetic TCE gene (synTCE) optimized for heterologous expression in P. pastoris, i.e., a gene in which certain TCE codons are replaced with synonymous codons 'preferred' in P. pastoris. When the altered gene was placed under the control of either the P. pastoris glyceraldehyde-3-phosphate dehydrogenase (GAP) promoter or the inducible alcohol oxidase (AOX1) promoter and introduced on an expression vector into P. pastoris, its product was produced intracellularly. We also successfully explored the possibility of obtaining a secreted product: P. pastoris cells expressing an in-frame fusion of synTCE with the alpha-factor secretion signal under the control of the GAP promoter were found to secrete the recombinant esterase into the external medium (to a concentration of 7 mg/L). In addition to this demonstration of TCE production in yeast, our results suggest that the GAP promoter could advantageously replace the AOX1 promoter as a driver of synTCE expression. TCE specific activity was approximately 5 U/mg when p-nitrophenyl acetate was used as substrate.

Animals↗