Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Synonymous mutations in the human dopamine receptor D2 (DRD2) affect mRNA stability and synthesis of the receptor.

Although changes in nucleotide sequence affecting the composition and the structure of proteins are well known, functional changes resulting from nucleotide substitutions cannot always be inferred from simple analysis of DNA sequence. Because a strong synonymous codon usage bias in the human DRD2 gene, suggesting selection on synonymous positions, was revealed by the relative independence of the G+C content of the third codon positions from the isochoric G+C frequencies, we chose to investigate functional effects of the six known naturally occurring synonymous changes (C132T, G423A, T765C, C939T, C957T, and G1101A) in the human DRD2. We report here that some synonymous mutations in the human DRD2 have functional effects and suggest a novel genetic mechanism. 957T, rather than being 'silent', altered the predicted mRNA folding, led to a decrease in mRNA stability and translation, and dramatically changed dopamine-induced up-regulation of DRD2 expression. 1101A did not show an effect by itself but annulled the above effects of 957T in the compound clone 957T/1101A, demonstrating that combinations of synonymous mutations can have functional consequences drastically different from those of each isolated mutation. C957T was found to be in linkage disequilibrium in a European-American population with the -141C Ins/Del and TaqI 'A' variants, which have been reported to be associated with schizophrenia and alcoholism, respectively. These results call into question some assumptions made about synonymous variation in molecular population genetics and gene-mapping studies of diseases with complex inheritance, and indicate that synonymous variation can have effects of potential pathophysiological and pharmacogenetic importance.

Dopamine↗

Correlations between Shine-Dalgarno sequences and gene features such as predicted expression levels and operon structures.

This work assesses relationships for 30 complete prokaryotic genomes between the presence of the Shine-Dalgarno (SD) sequence and other gene features, including expression levels, type of start codon, and distance between successive genes. A significant positive correlation of the presence of an SD sequence and the predicted expression level of a gene based on codon usage biases was ascertained, such that predicted highly expressed genes are more likely to possess a strong SD sequence than average genes. Genes with AUG start codons are more likely than genes with other start codons, GUG or UUG, to possess an SD sequence. Genes in close proximity to upstream genes on the same coding strand in most genomes are significantly higher in SD presence. In light of these results, we discuss the role of the SD sequence in translation initiation and its relationship with predicted gene expression levels and with operon structure in both bacterial and archaeal genomes.

Archaeal Proteins↗

Detection of transposable elements by their compositional bias.

BACKGROUND: Transposable elements (TE) are mobile genetic entities present in nearly all genomes. Previous work has shown that TEs tend to have a different nucleotide composition than the host genes, either considering codon usage bias or dinucleotide frequencies. We show here how these compositional differences can be used as a tool for detection and analysis of TE sequences. RESULTS: We compared the composition of TE sequences and host gene sequences using probabilistic models of nucleotide sequences. We used hidden Markov models (HMM), which take into account the base composition of the sequences (occurrences of words n nucleotides long, with n ranging here from 1 to 4) and the heterogeneity between coding and non-coding parts of sequences. We analyzed three sets of sequences containing class I TEs, class II TEs and genes respectively in three species: Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana. Each of these sets had a distinct, homogeneous composition, enabling us to distinguish between the two classes of TE and the genes. However the particular base composition of the TEs differed in the three species studied. CONCLUSIONS: This approach can be used to detect and annotate TEs in genomic sequences and complements the current homology-based TE detection methods. Furthermore, the HMM method is able to identify the parts of a sequence in which the nucleotide composition resembles that of a coding region of a TE. This is useful for the detailed annotation of TE sequences, which may contain an ancient, highly diverged coding region that is no longer fully functional.

Animals↗

Genome variability and capsid structural constraints of hepatitis a virus.

The number of synonymous mutations per synonymous site (K(s)), the number of nonsynonymous mutations per nonsynonymous site (K(a)), and the codon usage statistic (N(c)) were calculated for several hepatitis A virus (HAV) isolates. While K(s) was similar to those of poliovirus (PV) and foot-and-mouth disease virus (FMDV), K(a) was 1 order of magnitude lower. The N(c) parameter provides information on codon usage bias and decreases when bias increases. The N(c) value in HAV was about 38, while in PV and FMDV, it was about 53. The emergence of 22 rare codons in front of 8 in PV and 7 in FMDV was detected. Most of the conserved rare codons of the P1 region were strategically located at the carboxy borders of beta barrels and alpha helices, their potential function being the assurance of proper folding of the capsid proteins through a decrease in the translation speed. This strategic location was not observed for amino acids encoded by the conserved rare codons of the 3D region. The percentage of bases with low pairing number values was higher in the latter region, suggesting a role of the conserved rare codons in the maintenance of RNA structure. Many of the rare codons in HAV are among the most frequent in humans, unlike in PV or in FMDV. This fact may be explained by the lack of cellular shutoff in HAV. One hypothesis is that HAV has evolved in order to avoid competition with its host for cellular tRNAs.

Amino Acid Sequence↗

Gene expression and molecular evolution.

The combination of complete genome sequence information and estimates of mRNA abundances have begun to reveal causes of both silent and protein sequence evolution. Translational selection appears to explain patterns of synonymous codon usage in many prokaryotes as well as a number of eukaryotic model organisms (with the notable exception of vertebrates). Relationships between gene length and codon usage bias, however, remain unexplained. Intriguing correlations between expression patterns and protein divergence suggest some general mechanisms underlying protein evolution.

Animals↗

Codon optimization, expression, and characterization of an internalizing anti-ErbB2 single-chain antibody in Pichia pastoris.

Anti-ErbB2 antibodies are used as convenient tools in exploration of ErbB2 functional mechanisms and in treatment of ErbB2-overexpressing tumors. When we employed the yeast Pichia pastoris to express an anti-ErbB2 single-chain antibody (scFv) derived from the tumor-inhibitory monoclonal antibody A21, the yield did not exceed 1-2 mg/L in shake flask cultures. As we considered that the poor codon usage bias may be one limiting factor leading to the inefficient translation and scFv production, we designed and synthesized the full-length scFv gene by choosing the P. pastoris preferred codons while keeping the G+C content at relatively low level. Codon optimization increased the scFv expression level 3- to 5-fold and up to 6-10 mg/L. Northern blotting further confirmed that the increase of scFv expression was mainly due to the enhancement of translation efficiency. Investigation of culture conditions revealed that the maximal cell growth and scFv expression were achieved at pH 6.5-7.0 with 2% casamino acids after 72 h methanol induction. Secreted scFv was easily purified (>95% homogeneous product) from culture supernatants in one step by using Ni2+ chelating affinity chromatography. The yield was approximately 10-15 mg/L. Functional studies showed that the A21 scFv could be internalized with high efficiency after binding to the ErbB2-overexpressing cells, suggesting this regent may prove especially useful for ErbB2-targeted immunotherapy.

Amino Acid Sequence↗

The influence of translational selection on codon usage in fishes from the family Cyprinidae.

In this paper, the main factors shaping codon usage in three species of fishes that belong to the family Cyprinidae (namely Brachidanio rerio, Cyprinus carpio, and Carassius auratus) are reported. Correspondence analysis (COA), a commonly used multivariate statistical approach, was used to analyze codon usage bias. Our results show that the main trend is strongly correlated with the GC(3) content at silent sites of each sequence. On the other hand, the second axis discriminates between presumed highly and lowly expressed genes, a result that is confirmed by the distribution of matching expressed sequence tags (ESTs) along that axis. Translational selection appears, therefore, to influence synonymous codon usage in these fishes. The comparison of codon usages of the sequences displaying the extreme values on the second axis indicates that several codons are significantly incremented among the heavily expressed sequences. Interestingly, several of these triplets are not only shared by the three fishes but also by Xenopus laevis, another cold-blooded vertebrate in which translational selection influences codon choices. We postulate that natural selection was operative for codon usage in the last common ancestor of these fishes and Xenopus, and will probably be detected in cold-blooded vertebrates in general. Finally, we raise the possibility that the same phenomena will be found among warm-blooded vertebrates.

Amino Acids↗

High-level expression of staphylococcal nuclease R gene in Escherichia coli.

Staphylococcal nuclease R, an analogue of nuclease A, was overproduced under the transcriptional control of the bacteriophage lambda PRPL promoters regulated by temperature sensitive repressors. The expression level reached 200-300 mg l-1 and showed little host dependence in different strains. The investigations of the recombinant nuclease R have revealed that the amino terminal formyl methionine residue of the nuclease is precisely processed, the protein consists of 155 amino acid residues. The experiment shows that the pBV221-DH5 alpha is a quite suitable vector-host system for high-level expression and precise processing of heterologous genes in Escherichia coli. The comparative studies between the codons used in the staphylococcal nuclease R gene and the optimal codon usage in E. coli indicate that high level expression of heterologous genes in E. coli may not always require a high degree of codon usage bias.

Amino Acid Sequence↗

Optimality of codon usage in Escherichia coli due to load minimization.

The canonical genetic code is known to be highly efficient in minimizing the effects of mistranslational errors and point mutations, an ability which in term is designated "load minimization". One parameter involved in calculating the load minimizing property of the genetic code is codon usage. In most bacteria, synonymous codons are not used with equal frequencies. Different factors have been proposed to contribute to codon usage preference. It has been shown that the codon preference is correlated with the composition of the tRNA pool. Selection for translational efficiency and translational accuracy both result in such a correlation. In this work, it is shown that codon usage bias in Escherichia coli works so as to minimize the consequences of translational errors, i.e. optimized for load minimization.

Codon↗

Sequence, transcription and translation of a late gene of the Autographa californica nuclear polyhedrosis virus encoding a 34.8K polypeptide.

A 1.4 kb region downstream of the DNA polymerase gene of Autographa californica nuclear polyhedrosis virus was sequenced. Two open reading frames (ORFs) were identified of 927 and 474 bases in length. The 927 base ORF encodes a 34.8K protein as determined by in vitro translation of both hybrid-selected RNA and RNA synthesized in vitro from a 927 base ORF template. The predicted amino acid sequence of the 34.8K polypeptide (p34.8) reveals a hydrophobic N terminus, two potential N-glycosylation sites, and potential sites for phosphorylation by casein kinase I and protein kinase C. The p34.8 gene has a strong codon usage bias which is strikingly different from that of the polyhedrin gene. The two 5' ends of the 927 base ORF transcripts initiate from an ATAAG sequence and a GTAAG sequence 11 and 87 bases upstream of the ATG codon respectively. A short upstream reading frame is present in the leader sequence of the longer RNA. The transcripts have multiple 3' ends; the most proximal endpoint correlates with a polyadenylation signal overlapping the translational termination codon of the 927 base ORF. Transcripts of the latter were not observed early in the infection cycle but appeared 6 h after infection and were maximally expressed at 12 to 24 h post-infection. The late nature of these transcripts was confirmed by their sensitivity to aphidicolin and cycloheximide, inhibitors of DNA replication and protein synthesis respectively. Attempts to construct viral mutants carrying a deletion of the p34.8 gene and fusion with the beta-galactosidase gene suggest that the former gene is essential for viral replication.

Amino Acid Sequence↗

Codon optimization markedly improves doxycycline regulated gene expression in the mouse heart.

Tetracycline regulated gene expression in transgenic animals is potentially a very powerful technique (Furth et al., 1994; Gossen & Bujard 1992). We have utilized this system in an attempt to overcome the perinatal lethality resulting from constitutive transgenic expression in the heart (Valencik & McDonald, Am J Physiol Heart Circ Physiol 280: H361-H367). We found that compound hemizygous animals created by mating selected reverse tetracycline transactivator (rtTA) and transresponder (TR) lines display tightly regulated TR expression in the heart. However, we identified two fundamental problems. First, codon usage bias appeared to severely limit the expression of the rtTA driven by the cardiac alpha-myosin heavy chain promoter. Second, co-injection of rtTA and TR transgenes led to compound hemizygous animals that exhibited unregulated TR gene expression. Codon optimization of the rtTA construct leads to marked improvement (increasing the average induction from 20-fold to 832-fold) in cardiac myocyte expression. The resulting opt-rtTA lines can be bred to homozygosity, facilitating rapid screening of F0 TR animals for doxycycline regulated transgene expression.

Animals↗

Sequence and organization of the Trichoplusia ni ascovirus 2c (Ascoviridae) genome.

The complete Trichoplusia ni ascovirus 2c (TnAV-2c) genome sequence was determined. The circular genome contains 174,059 bp with 165 open reading frames (ORFs) of greater than 180 bp and two major homologous regions (hrs). The genome is quite A+T rich at 64.6%. Fifty-four ORFs had homologues in other insect viruses, such as ascoviruses, iridoviruses, baculoviruses and entomopoxviruses; 30 ORFs showed low identities with those from different parasitic protozoa and 12 ORFs were unique to TnAV-2c. TnAV-2c has 15 ORFs that could be grouped into six gene families. Three major conserved repeating sequences were identified and were interspersed in two regions. BLAST analyses revealed that there were 16 enzymes involved in gene transcription, DNA replication, and nucleotide metabolism. TnAV-2c has 12 and 25 ORFs sharing high identities with ascovirus and iridovirus homologues, respectively. The codon usage bias appears to be more similar to Spodoptera frugiperda ascovirus 1a than to iridoviruses.

Animals↗

Heterogeneity in regional GC content and differential usage of codons and amino acids in GC-poor and GC-rich regions of the genome of Apis mellifera.

The honeybee (Apis mellifera) has a genome with a wide variation in GC content showing 2 clear modal GC values, in some ways reminiscent of an isochore-like structure. To gain insight into causes and consequences of this pattern, we used a comparative approach to study the genome-wide alignment of primarily coding sequence of A. mellifera with Drosophila melanogaster and Anopheles gambiae. The latter 2 species show a higher average GC content than A. mellifera and no indications of bimodality, suggesting that the GC-poor mode is a derived condition in honeybee. In A. mellifera, synonymous sites of genes generally adopt the GC content of the region in which they reside. A large proportion of genes in GC-poor regions have not been assigned to the honeybee assembly because of the low sequence complexity of their genome neighborhood. The synonymous substitution rate between A. mellifera and the other species is very close to saturation, but analyses of nonsynonymous substitutions as well as amino acid substitutions indicate that the GC-poor regions are not evolving faster than the GC-rich regions. We describe the codon usage and amino acid usage and show that they are remarkably heterogeneous within the honeybee genome between the 2 different GC regions. Specifically, the genes located in GC-poor regions show a much larger deviation in both codon usage bias and amino acid usage from the Dipterans than the genes located in the GC-rich regions.

Amino Acids↗

Comparison of codon usage and tRNAs in mitochondrial genomes of Candida species.

To gain insight into the nature of the mitochondrial genomes (mtDNA) of different Candida species, the synonymous codon usage bias of mitochondrial protein coding genes and the tRNAs in C. albicans, C. parapsilosis, C. stellata, C. glabrata and the closely related yeast Saccharomyces cerevisiae were analyzed. Common features of the mtDNA in Candida species are a strong A+T pressure on protein coding genes, and insufficient mitochondrial tRNA species are encoded to perform protein synthesis. The wobble site of the anticodon is always U for the NNR (NNA and NNG) codon families, which are dominated by A-ending codons, and always G for the NNY (NNC and NNU) codon families, which is dominated by U-ending codons, and always U for the NNN (NNA, NNU, NNC and NNG) codon families, which are dominated by A-ending codons and U-ending codons. Patterns of synonymous codon usage of Candida species can be classified into three groups: (1) optimal codon-anticodon usage, Glu, Lys, Leu (translated by anti-codon UAA), Gln, Arg (translated by anti-codon UCU) and Trp are containing NNR codons. NNA, whose corresponding tRNA is encoded in the mtDNA, is used preferentially. (2) Non-optimal codon-anticodon usage, Cys, Asp, Phe, His, Asn, Ser (translated by anti-codon GCU) and Tyr are containing NNY codons. The NNU codon, whose corresponding tRNA is not encoded in the mtDNA, is used preferentially. (3) Combined codon-anticodon usage, Ala, Gly, Leu (translated by anti-codon UAG), Pro, Ser (translated by anti-codon UGA), Thr and Val are containing NNN codons. NNA (tRNA encoded in the mtDNA) and NNU (tRNA not encoded in the mtDNA) are used preferentially. In conclusion, we propose that in Candida species, codons containing A or U at third position are used preferentially, regardless of whether corresponding tRNAs are encoded in the mtDNA. These results might be useful in understanding the common features of the mtDNA in Candida species and patterns of synonymous codon usage.

Algorithms↗

Evolutionary rates and expression level in Chlamydomonas.

In many biological systems, especially bacteria and unicellular eukaryotes, rates of synonymous and nonsynonymous nucleotide divergence are negatively correlated with the level of gene expression, a phenomenon that has been attributed to natural selection. Surprisingly, this relationship has not been examined in many important groups, including the unicellular model organism Chlamydomonas reinhardtii. Prior to this study, comparative data on protein-coding sequences from C. reinhardtii and its close noninterfertile relative C. incerta were very limited. We compiled and analyzed protein-coding sequences for 67 nuclear genes from these taxa; the sequences were mostly obtained from the C. reinhardtii EST database and our C. incerta EST data. Compositional and synonymous codon usage biases varied among genes within each species but were highly correlated between the orthologous genes of the two species. Relative rates of synonymous and nonsynonymous substitution across genes varied widely and showed a strong negative correlation with the level of gene expression estimated by the codon adaptation index. Our comparative analysis of substitution rates in introns of lowly and highly expressed genes suggests that natural selection has a larger contribution than mutation to the observed correlation between evolutionary rates and gene expression level in Chlamydomonas.

Animals↗

Evidence for horizontal transfer from Streptococcus to Escherichia coli of the kfiD gene encoding the K5-specific UDP-glucose dehydrogenase.

Capsular polysaccharides are important virulence factors both in Gram-positive and Gram-negative bacteria. A similar cluster organization of the genes involved in the synthesis of bacterial exopolysaccharides has been postulated in both cases, suggesting that these clusters evolved by module assembly. Horizontal gene transfer has been postulated to explain the polymorphism found in these cellular polymers. The cap1 K and cap3A genes coding for the pneumococcal type 1 and type 3 UDP-glucose dehydrogenases, respectively, have been compared with other UDP-sugar dehydrogenases. We have observed that the evolutionary distance between Cap1K and Cap3A is approximately equal to that found between Cap1K (or Cap3A) and other UDP-GlcDH of families evolutionarily distant like KfiD, the dehydrogenase from Escherichia coli K5. On the basis of comparisons of G + C content, patterns of synonymous and nonsynonymous substitutions, dinucleotide frequencies, and codon usage bias, we conclude that the kfiD gene has been introduced into E. coli from an exogenous source, probably from a streptococcal species.

Bacterial Capsules↗

CpsK of Streptococcus agalactiae exhibits alpha2,3-sialyltransferase activity in Haemophilus ducreyi.

Streptococcus agalactiae (GBS) is a major cause of serious newborn bacterial infections. Crucial to GBS evasion of host immunity is the production of a capsular polysaccharide (CPS) decorated with sialic acid, which inactivates the alternative complement pathway. The CPS operons of serotypes Ia and III GBS have been described, but the CPS sialyltransferase gene was not identified. We identified cpsK, an open reading frame in the CPS operon of most serotypes, which was homologous to the lipooligosaccharide (LOS) sialyltransferase gene, lst, of Haemophilus ducreyi. To determine if cpsK might encode a sialyltransferase, we complemented a H. ducreyi lst mutant with cpsK. CpsK was expressed in H. ducreyi and LOS was isolated and analysed for sialic acid content by SDS-PAGE and high-performance liquid chromatography (HPLC). Sialo-LOS was seen in the wild-type, cpsK- or lst-complemented mutant strains, but not in the mutant without cpsK. Addition of Neu5Ac to the LOS was confirmed by mass spectroscopy. Lectin binding studies detected terminal Neu5Ac(alpha 2-->3)Gal(beta 1- on LOS produced by the wild-type, cpsK or lst-complemented mutant strain LOS, compared with the mutant alone. Our data characterize the first sialyltransferase gene from a Gram- positive bacterium and provide compelling evidence that its product catalyses the alpha2,3 addition of Neu5Ac to H. ducreyi LOS and therefore the terminal side-chain of GBS CPS. Phylogenetic studies further indicated that lst and cpsK are related but distinct from sialyltransferases of most other bacteria and, along with their similar codon usage bias and G + C content, suggests acquisition by lateral transfer from an ancestral low G + C organism.

Amino Acid Sequence↗

Metabolic efficiency and amino acid composition in the proteomes of Escherichia coli and Bacillus subtilis.

Biosynthesis of an Escherichia coli cell, with organic compounds as sources of energy and carbon, requires approximately 20 to 60 billion high-energy phosphate bonds [Stouthamer, A. H. (1973) Antonie van Leeuwenhoek 39, 545-565]. A substantial fraction of this energy budget is devoted to biosynthesis of amino acids, the building blocks of proteins. The fueling reactions of central metabolism provide precursor metabolites for synthesis of the 20 amino acids incorporated into proteins. Thus, synthesis of an amino acid entails a dual cost: energy is lost by diverting chemical intermediates from fueling reactions and additional energy is required to convert precursor metabolites to amino acids. Among amino acids, costs of synthesis vary from 12 to 74 high-energy phosphate bonds per molecule. The energetic advantage to encoding a less costly amino acid in a highly expressed gene can be greater than 0.025% of the total energy budget. Here, we provide evidence that amino acid composition in the proteomes of E. coli and Bacillus subtilis reflects the action of natural selection to enhance metabolic efficiency. We employ synonymous codon usage bias as a measure of translation rates and show increases in the abundance of less energetically costly amino acids in highly expressed proteins.

Amino Acids↗