Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Expression of a synthetic gene encoding a Tribolium castaneum carboxylesterase in Pichia pastoris.

This is the first report of an insect esterase efficiently expressed in the methylotrophic yeast Pichia pastoris (so far insect esterases have been produced only in the baculovirus system). Having isolated a Tribolium castaneum carboxylesterase cDNA (TCE), we were initially unable to express it in Escherichia coli or P. pastoris despite significant transcription levels. As codon usage bias is different in T. castaneum and P. pastoris, we assumed this was a possible explanation for the translational barrier observed in yeast. Accordingly, we designed and constructed by recursive PCR a synthetic TCE gene (synTCE) optimized for heterologous expression in P. pastoris, i.e., a gene in which certain TCE codons are replaced with synonymous codons 'preferred' in P. pastoris. When the altered gene was placed under the control of either the P. pastoris glyceraldehyde-3-phosphate dehydrogenase (GAP) promoter or the inducible alcohol oxidase (AOX1) promoter and introduced on an expression vector into P. pastoris, its product was produced intracellularly. We also successfully explored the possibility of obtaining a secreted product: P. pastoris cells expressing an in-frame fusion of synTCE with the alpha-factor secretion signal under the control of the GAP promoter were found to secrete the recombinant esterase into the external medium (to a concentration of 7 mg/L). In addition to this demonstration of TCE production in yeast, our results suggest that the GAP promoter could advantageously replace the AOX1 promoter as a driver of synTCE expression. TCE specific activity was approximately 5 U/mg when p-nitrophenyl acetate was used as substrate.

Animals↗

Molecular Evolution and Expression Analysis of the ADH Gene Family in Apple Bud Mutants.

Alcohol dehydrogenase (ADH) catalyzes the reduction of aldehydes to alcohols, key precursor substrates for volatile ester biosynthesis, which determines the characteristic aroma of apple fruit. However, a comprehensive genome-wide investigation of the ADH gene family in apple has been lacking. In this study, we systematically identified ADH genes in the apple genome using integrated bioinformatics approaches, including phylogenetic analysis, synteny evaluation, promoter cis-element prediction, codon usage bias assessment, and protein interaction network modeling. Expression patterns were examined through transcriptomic data and validated by RT-qPCR analysis across different organs and among 'Red Delicious' and its four bud mutant lines. We identified 44 ADH genes, with 12 forming a prominent cluster on chromosome 1. RT-qPCR analysis revealed that MdADH20 was dramatically upregulated in the 'Red Chief' mutant (relative expression of 59.38), suggesting its pivotal role. Phylogenetic analysis revealed a close evolutionary relationship with wild strawberry. The encoded proteins were generally stable and predominantly localized to the cytoplasm. Promoter analysis showed enrichment of growth/development-related and ARE elements, while codon usage analysis identified AGA, GCU, GUU, and CUU as preferred codons. Protein interaction prediction suggested MdADH19 and MdADH20 as hub proteins. Expression profiling and RT-qPCR further identified MdADH20 as a core candidate gene, characterized by its stable and high expression, particularly in the 'Red Delicious' mutant. Its central position in the predicted protein-protein interaction network suggests a potential regulatory role in the aroma biosynthesis pathway of apple fruit. This study provides the first systematic genome-wide characterization of the apple ADH gene family, establishing a theoretical groundwork for deciphering aroma biosynthesis mechanisms and offering potential target genes for flavor improvement through bud mutation breeding strategies.

ADH gene family↗

Multiplicative versus additive selection in relation to genome evolution: a simulation study.

The evolution of molecular quantitative traits, such as codon usage bias or base frequencies, can be explained as the result of mutational biases alone, or as the result of mutation and selection. Whereas mutation models can be investigated easily, realistic modelling of selection-directed genome evolution is analytically intractable, and numerical calculations require substantial computer resources. We investigated the evolution of optimal codon frequency under additive and multiplicative effects of selected linked codons. We show that additive selective effects of many linked sites cannot be effective in genomes when the number of selected sites is greater than the effective population size, a realistic assumption according to current molecular data. We then discuss the implications of these results for isochore evolution in vertebrates.

Codon↗

[Transformation of Chlamydomonas reinhardtii CW-15 with the hygromycin phosphotransferase gene as a selective marker].

To transform Chlamydomonas reinhardtii Dang. Cells, plasmid pCTVHyg was constructed with the use of the Escherichia coli hygromycin phosphotransferase gene (hpt) controlled by the SV40 early promoter. Cells of the CW-15 mutant strain were transformed by electroporation, with the yield reaching 10(3) hygromycin-resistant (HygR) clones per 10(6) recipient cells. The exogenous DNA integrated in the Ch. reinhardtii nuclear genome showed stable transmission for approximately 350 cell generations, while hygromycin resistance was expressed as an unstable character. Codon usage was compared for the hpt gene and Ch. reinhardtii nuclear genes. The results testified that codon usage bias, which is characteristic of Ch. reinhardtii, is not the major factor affecting foreign gene expression. The advantages of the selective system for studying Ch. reinhardtii transformation with heterologous genes are discussed.

Animals↗

Another putative heat-shock gene and aminoacyl-tRNA synthetase gene are located upstream from the grpE-like and dnaK-like genes in Chlamydia trachomatis.

The 4.1-kb sequence of genomic DNA located upstream from the Chlamydia trachomatis grpE-like and dnaK-like heat shock (HS) genes was determined. Another putative HS gene was located just 5' to grpE along with an inverted repeat (IR) sequence proposed to be involved in HS regulation. The overall organization of this locus in Chlamydia resembles that of Bacillus subtilis, rather than Escherichia coli. Two other open reading frames (ORFs) were found in the sequence, one of which has homology to aminoacyl-tRNA synthetases. The other ORF has no significant homology to reported genes. We also examined the codon usage bias for these newly identified chlamydial ORFs and for previously reported chlamydial genes, and found them to be different from E. coli.

Amino Acid Sequence↗

R1 and R2 retrotransposable elements of Drosophila evolve at rates similar to those of nuclear genes.

The non-long-terminal repeat retrotransposable elements, R1 and R2, insert at unique locations in the 28S ribosomal RNA genes of insects. Based on the nucleotide sequences of these elements in the eight members of the melanogaster species subgroup of the genus Drosophila, they have been maintained by vertical germline transmission for the 17-20 million year history of this subgroup. The stable inheritance of R1 and R2 within these species has enabled a determination of their nucleotide substitution rates. The sequence of the R1 and R2 elements from D. ambigua, a member of the obscura species group, has also been determined to enable an extrapolation of this rate over an estimated 45-60 million years. The mean rate of substitutions at synonymous sites (Ks) was 6.6 and 9.6 times the rate at replacement sites (Ka) in the R1 and R2 elements, respectively. Both elements appear to have been under selective pressure to maintain their open reading frames and thus their ability to retrotranspose for most of their evolution in these lineages. Using the rate of change at synonymous sites (Ks) as the best indicator of the nucleotide substitution rate, the mean Ks values for R1 and R2 were 2.3 and 2.2 times that of the alcohol dehydrogenase (Adh) genes. However, this faster rate is a result of the lower codon usage bias of R1 and R2 compared with that of Adh. When the Ks rates of R1 and R2 were compared with that of a larger number of nuclear genes available from at least two of the nine species under investigation, R1 and R2 were found to evolve in most lineages at rates similar to that of nuclear genes with low codon bias. The ability of R1 and R2 to maintain their presence in this species subgroup by retrotransposition while exhibiting rates of nucleotide evolution similar to nuclear genes suggests these transposition events are rare or not as error prone as that of retroviruses.

Amino Acid Sequence↗

Codon usage and selection on proteins.

Selection pressures on proteins are usually measured by comparing homologous nucleotide sequences (Zuckerkandl and Pauling 1965). Recently we introduced a novel method, termed volatility, to estimate selection pressures on proteins on the basis of their synonymous codon usage (Plotkin and Dushoff 2003; Plotkin et al. 2004). Here we provide a theoretical foundation for this approach. Under the Fisher-Wright model, we derive the expected frequencies of synonymous codons as a function of the strength of selection on amino acids, the mutation rate, and the effective population size. We analyze the conditions under which we can expect to draw inferences from biased codon usage, and we estimate the time scales required to establish and maintain such a signal. We find that synonymous codon usage can reliably distinguish between negative selection and neutrality only for organisms, such as some microbes, that experience large effective population sizes or periods of elevated mutation rates. The power of volatility to detect positive selection is also modest--requiring approximately 100 selected sites--but it depends less strongly on population size. We show that phenomena such as transient hyper-mutators can improve the power of volatility to detect selection, even when the neutral site heterozygosity is low. We also discuss several confounding factors, neglected by the Fisher-Wright model, that may limit the applicability of volatility in practice.

Algorithms↗

Codon optimization of Caenorhabditis elegans GluCl ion channel genes for mammalian cells dramatically improves expression levels.

Organisms use synonymous codons in a highly non-random fashion. These codon usage biases sometimes frustrate attempts to express high levels of exogenous genes in hosts of widely divergent species. The Caenorhabditis elegans GluClalpha1 and GluClbeta genes form a functional glutamate and ivermectin-gated chloride channel when expressed in Xenopus oocytes, but expression is weak in mammalian cells. We have constructed synthetic genes that retain the amino acid sequence of the wild-type GluCl channel proteins, but use codons that are optimal for mammalian cell expression. We have tagged the native and codon-optimized GluCl cDNAs with enhanced yellow fluorescent protein (EYFP, GluClalpha1 subunit) and enhanced cyan fluorescent protein (EFCP, GluClbeta subunit), expressed the channels in E18 rat hippocampal neurons and measured the relative expression levels of the two genes with fluorescence microscopy as well as with electrophysiology. Codon optimization provides a 6- to 9-fold increase in expression, allowing the conclusions that the ivermectin-gated channel has an EC(50) of 1.2 nM and a Hill coefficient of 1.9. We also confirm that the Y182F mutation in the codon-optimized beta subunit results in a heteromeric channel that retains the response to ivermectin while reducing the response to 100 microM glutamate by 7-fold. The engineered GluCl channel is the first codon-optimized membrane protein expressed in mammalian cells and may be useful for selectively silencing specific neuronal populations in vivo.

Animals↗

Determinants of DNA sequence divergence between Escherichia coli and Salmonella typhimurium: codon usage, map position, and concerted evolution.

The nature and extent of DNA sequence divergence between homologous protein-coding genes from Escherichia coli and Salmonella typhimurium have been examined. The degree of divergence varies greatly among genes at both synonymous (silent) and nonsynonymous sites. Much of the variation in silent substitution rates can be explained by natural selection on synonymous codon usage, varying in intensity with gene expression level. Silent substitution rates also vary significantly with chromosomal location, with genes near oriC having lower divergence. Certain genes have been examined in more detail. In particular, the duplicate genes encoding elongation factor Tu, tufA and tufB, from S. typhimurium have been compared to their E. coli homologues. As expected these very highly expressed genes have high codon usage bias and have diverged very little between the two species. Interestingly, these genes, which are widely spaced on the bacterial chromosome, also appear to be undergoing concerted evolution, i.e., there has been exchange between the loci subsequent to the divergence of the two species.

Base Sequence↗

An analysis of determinants of amino acids substitution rates in bacterial proteins.

The variation of amino acid substitution rates in proteins depends on several variables. Among these, the protein's expression level, functional category, essentiality, or metabolic costs of its amino acid residues may play an important role. However, the relative importance of each variable has not yet been evaluated in comparative analyses. To this aim, we made regression analyses combining data available on these variables and on evolutionary rates, in two well-documented model bacteria, Escherichia coli and Bacillus subtilis. In both bacteria, the level of expression of the protein in the cell was by far the most important driving force constraining the amino acids substitution rate. Subsequent inclusion in the analysis of the other variables added little further information. Furthermore, when the rates of synonymous substitutions were included in the analysis of the E. coli data, only the variable expression levels remained statistically significant. The rate of nonsynonymous substitution was shown to correlate with expression levels independently of the rate of synonymous substitution. These results suggest an important direct influence of expression levels, or at least codon usage bias for translation optimization, on the rates of nonsynonymous substitutions in bacteria. They also indicate that when a control for this variable is included, essentiality plays no significant role in the rate of protein evolution in bacteria, as is the case in eukaryotes.

Amino Acids↗

Genetics of lactobacilli: plasmids and gene expression.

This paper reviews the present knowledge of the structure and properties of small (< 5 kb) plasmids present in Lactobacillus spp. The data show that plasmids from Lactobacillus spp., like many plasmids from other Gram-positive bacteria, display a modular organization and replicate by a mechanism of rolling circle replication. Structurally, plasmids from lactobacilli are closely related to plasmids from other Gram-positive bacteria. They contain elements (plus- and minus origin of replication, element(s) for control of plasmid replication, mobilization function) showing extensive similarity to analogous elements in plasmids from these other organisms. It is believed that lactobacilli have acquired such elements by intra- and/or intergenic transfer mechanisms. The first part of the review is concluded with a description of plasmid vectors with a Lactobacillus replicon and integrative vectors, including data concerning their structural and segregational stability. In the second part of this review we describe the progress that has been made during the last few years in identifying and characterizing elements that control expression of genetic information in lactobacilli. Based on the sequence of eleven identified and twenty presumed promoters, some preliminary conclusions can be drawn regarding the structure of Lactobacillus promoters. A typical Lactobacillus promoter shows significant similarity to promoters from E. coli and B. subtilis. An analysis of published sequences of seventy genes indicates that the region encompassing the translation start codon AUG also shows extensive similarity to that of E. coli and B. subtilis. Codon usage of Lactobacillus genes is not random and shows interspecies as well as intraspecies heterogeneity. Interspecies differences may, in part, be explained by differences in G+C content of different lactobacilli. Differences in gene expression levels can, to a large extent, account for intraspecies differences of codon usage bias. Finally, we review the knowledge that has become available concerning protein secretion and heterologous gene expression in lactobacilli. This part is concluded with a compilation of data on the expression in Lactobacillus of heterologous genes under the control of their own promoter or under control of a Lactobacillus promoter.

Bacillus subtilis↗

The quality of merC, a module of the mer mosaic.

We examined a region of high variability in the mosaic mercury resistance (mer) operon of natural bacterial isolates from the primate intestinal microbiota. The region between the merP and merA genes of nine mer loci was sequenced and either the merC, the merF, or no gene was present. Two novel merC genes were identified. Overall nucleotide diversity, pi (per 100 sites), of the merC gene was greater (49.63) than adjacent merP (35.82) and merA (32.58) genes. However, the consequences of this variability for the predicted structure of the MerC protein are limited and putative functional elements (metal-binding ligands and transmembrane domains) are strongly conserved. Comparison of codon usage of the merTP, merC, and merA genes suggests that several merC genes are not coeval with their flanking sequences. Although evidence of homologous recombination within the very variable merC genes is not apparent, the flanking regions have higher homologies than merC, and recombination appears to be driving their overall sequence identities higher. The synonymous codon usage bias (EN(C)) values suggest greater variability in expression of the merC gene than in flanking genes in six different bacterial hosts. We propose a model for the evolution of MerC as a host-dependent, adventitious module of the mer operon.

Amino Acid Sequence↗

Variable rates of evolution among Drosophila opsin genes.

DNA sequences and chromosomal locations of four Drosophila pseudoobscura opsin genes were compared with those from Drosophila melanogaster, to determine factors that influence the evolution of multigene families. Although the opsin proteins perform the same primary functions, the comparisons reveal a wide range of evolutionary rates. Amino acid identities for the opsins range from 90% for Rh2 to more than 95% for Rh1 and Rh4. Variation in the rate of synonymous site substitution is especially striking: the major opsin, encoded by the Rh1 locus, differs at only 26.1% of synonymous sites between D. pseudoobscura and D. melanogaster, while the other opsin loci differ by as much as 39.2% at synonymous sites. Rh3 and Rh4 have similar levels of synonymous nucleotide substitution but significantly different amounts of amino acid replacement. This decoupling of nucleotide substitution and amino acid replacement suggests that different selective pressures are acting on these similar genes. There is significant heterogeneity in base composition and codon usage bias among the opsin genes in both species, but there are no consistent relationships between these factors and the rate of evolution of the opsins. In addition to exhibiting variation in evolutionary rates, the opsin loci in these species reveal rearrangements of chromosome elements.

Amino Acid Sequence↗

Recombination and base composition: the case of the highly self-fertilizing plant Arabidopsis thaliana.

BACKGROUND: Rates of recombination can vary among genomic regions in eukaryotes, and this is believed to have major effects on their genome organization in terms of base composition, DNA repeat density, intron size, evolutionary rates and gene order. In highly self-fertilizing species such as Arabidopsis thaliana, however, heterozygosity is expected to be strongly reduced and recombination will be much less effective, so that its influence on genome organization should be greatly reduced. RESULTS: Here we investigated theoretically the joint effects of recombination and self-fertilization on base composition, and tested the predictions with genomic data from the complete A. thaliana genome. We show that, in this species, both codon-usage bias and GC content do not correlate with the local rates of crossing over, in agreement with our theoretical results. CONCLUSIONS: We conclude that levels of inbreeding modulate the effect of recombination on base composition, and possibly other genomic features (for example, transposable element dynamics). We argue that inbreeding should be considered when interpreting patterns of molecular evolution.

Arabidopsis↗

Distribution of potential type II restriction sites (palindromes) in prokaryotes.

Restriction-modification systems are used as a defensive mechanism against inappropriate invasion of foreign DNA. The recognition sequences for the common type II restriction enzymes and their corresponding methylases are usually palindromes. In this study, we identified the most over- and underrepresented words in DNA of four bacteria: Escherichia coli, Bacillus subtilis, Clostridium perfringens, and Pseudomonas aeruginosa. Using maximum order Markov chain analysis, we found that palindromic words were most often more underrepresented than their non-palindromic counterparts. No strict rule for the intragenic palindrome content could be derived, but for three of the bacteria there was a weak correlation between codon usage bias and palindrome content. A clear drop in palindrome counts was observed in the Shine-Dalgarno region for B. subtilis and C. perfringens, but not in E. coli or P. aeruginosa. It was also shown that palindromes in eubacteria and archaebacteria seem to occur slightly more infrequently than expected on the basis of the genomic GC-content, but some exceptions to this principle exist.

Bacteria↗

Peptidase D gene (pepD) of Escherichia coli K-12: nucleotide sequence, transcript mapping, and comparison with other peptidase genes.

The nucleotide sequence of a 2.3-kilobase-pair DNA fragment of Escherichia coli that contains the transcription signals and the coding region of the pepD gene specifying aminopeptidase D was determined. The location and extent of the open reading frame were verified by partial amino acid sequencing of the purified pepD product. By use of a promoter-screening vector, initiation signals for pepD transcription were located in the 5'-flanking region of the open reading frame. Analysis of pepD transcripts by S1 mapping, primer extension, and Northern (RNA) hybridization revealed two species of monocistronic mRNA with different 5' ends and a common 3' end. Calculation of the degree of codon usage bias in the coding region suggested that the efficiency of pepD translation is relatively low. As deduced from the predicted amino acid sequence, peptidase D is a slightly hydrophilic protein of 485 amino acid residues that contains no extended domains of marked hydrophobicity. Structural and functional features of the pepD gene are discussed and compared with other already sequenced peptidase genes of E. coli.

Amino Acid Sequence↗

High level expression of a synthetic gene encoding Peniophora lycii phytase in methylotrophic yeast Pichia pastoris.

Phytase is widespread in nature. It has been used as a cereal feed additive that can enhance the phosphorus and mineral absorption in monogastric animals to reduce the level of phosphorus output in manure. Phytase of Peniophora lycii is a 6'-phytase, which owns high specific activity. To achieve a high expression level of 6'-phytase in Pichia pastoris, the 1,230-bp phytase gene of P. lycii was synthesized and optimized for codon usage, G+C content, as well as mRNA secondary structures. The gene constructs containing wild type or modified phytase gene coding sequences under the control of the highly-inducible alcohol oxidase gene (AOX1) promoter, the synthetic signal peptide (designated MF4I), which is a codon-modified Saccharomyces cerevisiae mating factor alpha-prepro-leader sequence, were used to transform P. pastoris. The P. pastoris strain that expressed the modified phytase gene (phy-pl-sh) with MF4I sequence produced 12.2 g phytase per liter of fluid culture, with the phytase activity of 10,540 U ml(-1). The yield of the modified phytase gene, with bias codon usage and MF4I signal, is 4.4 times higher than that of the wild type gene with MF4I signal and 13.6 times higher than that of the wild type gene with wild type S. cerevisiae signal. The recombinant phytase had one optimum pH (pH 4.5) and an optimum temperature of 50 degrees C. The P. pastoris strain expressed the modified 6-phytase gene, with the MF4I signal peptide showing great potential as a commercial phytase production system.

6-Phytase↗

Evolution of the Adh locus in the Drosophila willistoni group: the loss of an intron, and shift in codon usage.

We report here the DNA sequence of the alcohol dehydrogenase gene (Adh) cloned from Drosophila willistoni. The three major findings are as follows: (1) Relative to all other Adh genes known from Drosophila, D. willistoni Adh has the last intron precisely deleted; PCR directly from total genomic DNA indicates that the deletion exists in all members of the willistoni group but not in any other group, including the closely related saltans group. Otherwise the structure and predicted protein are very similar to those of other species. (2) There is a significant shift in codon usage, especially compared with that in D. melanogaster Adh. The most striking shift is from C to U in the wobble position (both third and first position). Unlike the codon-usage-bias pattern typical of highly biased genes in D. melanogaster, including Adh, D. willistoni has nearly 50% G + C in the third position. (3) The phylogenetic information provided by this new sequence is in agreement with almost all other molecular and morphological data, in placing the obscura group closer to the melanogaster group, with the willistoni group farther distant but still clearly within the subgenus Sophophora.

Alcohol Dehydrogenase↗