Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Sequence comparisons of developmentally regulated collagen genes of Caenorhabditis elegans.

Collagen genes col-6, col-7 (partial), col-8, col-14 and col-19 from the nematode Caenorhabditis elegans were sequenced, and compared to the previously sequenced genes col-1 and col-2. The genes are between 1.0 and 1.2 kb in length, and each includes one or two short introns. The presumptive promoter regions contain sequences similar to the eukaryotic TATA promoter element. Two distinct, conserved sequences were found in the presumptive promoter regions of, respectively, the dauer larva-specific genes col-2 and col-6, and the primarily adult-specific genes col-7 and col-19. The domain structures of the collagen polypeptides are similar: each polypeptide contains two triple-helix forming (Gly-X-Y)n domains, one of 30-33 amino acids (aa), and the other of 127-132 aa. The latter domain is interrupted by one to three short (2-8 aa) non-(Gly-X-Y)n segments that occur at relatively conserved locations in each polypeptide. Sets of cysteine residues flank the (Gly-X-Y)n domains in all of the polypeptides. The genes can be placed into three families based upon amino acid sequence similarities. Genes within a family do not always exhibit similar developmental expression programs, suggesting that structural and regulatory regions of the genes have evolved separately. The codon usage in the genes is highly asymmetrical, with adenine appearing in the third position of 85% of the glycine codons, and 93% of the proline codons.

Amino Acid Sequence↗

Complete sequence of omp1, the structural gene encoding the 40-kDa outer membrane protein of Fusobacterium nucleatum strain Fev1.

The sequence of the omp1 gene coding for the 40-kDa outer membrane protein (OMP) of the Gram- oral bacterium, Fusobacterium nucleatum strain Fev1, has been determined. Degenerate oligodeoxyribonucleotide primers were used to prime the amplification of a 120-mer sequence of the gene. This sequence was successively used for constructing new primers applied in asymmetrical, symmetrical, and inverse polymerase chain reaction using as template genomic DNA, self-ligated DNA fragments, or fragments ligated into either pGEM-7Zf+ or pACYC184. The codon usage of the gene was unusual in that A or T was used as the third base in the codon triplets in all cases, except for those amino acids (aa) which have only one or two possible codon choices. Only 35 of the 61 sense codons were used. The aa sequence of the protein was deduced; it consisted of 348 aa (M(r) 39,954), which is in good agreement with the 40-kDa size estimated from electrophoretic analyses. The mature protein was preceded by a 20-aa signal peptide.

Amino Acid Sequence↗

Optimized gene synthesis and high expression of human interleukin-18.

Human interleukin-18 (hIL-18), originally known as an IFN-gamma-inducing factor, is a recently cloned cytokine that is secreted by Kupffer cells of the liver and by stimulated macrophages. We have previously established a method of expression and purification of IL-18. The yield however remains low and the insufficient expression of a heterologous protein could be due to skewed codon usage between the expression host and the cDNA donor. The sequence of mature hIL-18 has 37 a.a. rare codons for Escherichia coli in a total of 157 a.a. To overcome this problem, gene synthesis was performed with optimized codons for the expression host E. coli. The final yield of the hIL-18 protein with optimized codons was about five times higher than the yield with the native sequence. Using a minimal medium, this system produces large quantities of labeled proteins that can be used in NMR analysis. Our simple and efficient production system can be applied to the production of other cytokines for new structural and therapeutic use.

Amino Acid Sequence↗

The effect of queuosine on tRNA structure and function.

Computational modeling was performed to determine the potential function of the queuosine modification of tRNA found in wobble position 34 of tRNAasp, tRNAasn, tRNAhis, and tRNAtyr. Using the crystal structure of tRNAasp and a tRNA-tRNA-mRNA complex model, we show that the queuosine modification serves as a structurally restrictive base for tRNA anticodon loop flexibility. An extended intraresidue and intramolecular hydrogen bonding network is established by queuosine. The quaternary amine of the 7-aminomethyl side chain hydrogen bonds with the base's carbonyl oxygen. This positions the dihydroxycyclopentenediol ring of queuosine in proper orientation for hydrogen bonding with the backbone of the neighboring uridine 33 residue. The interresidue association stabilizes the formation of a cross-loop hydrogen bond between the uridine 33 base and the phosphoribosyl backbone of the cytosine at position 36. Additional interactions between RNAs in the translation complex were studied with regard to potential codon context and codon bias effects. Neither steric nor electrostatic interaction occurs between aminoacyl- and peptidyl-site tRNA anticodon loops that are modified with queuosine. However, there is a difference in the strength of anticodon/codon associations (codon bias) based on the presence or lack of queuosine in the wobble position of the tRNA. Unmodified (guanosine-containing) tRNAasp forms a very stable association with cytosine (GAC), but is much less stable in complex with a uridine-containing codon (GAU). Queuosine-modified tRNAasp exhibits no bias for either of cognate codons GAC or GAU and demonstrates a lower binding energy similar to the wobble pairing of guanosine-containing tRNA with a GAU codon. This is proposed to be due to the inflexibility of the queuosine-modified anticodon loop to accommodate proper positioning for optimal Watson-Crick type associations. A preliminary survey of codon usage patterns in oncodevelopmental versus housekeeping gene transcripts suggests a significant difference in bias for the queuosine-associated codons. Therefore, the queuosine modification may have the potential to influence cellular growth and differentiation by codon bias-based regulation of protein synthesis for discrete mRNA transcripts.

Anticodon↗

Compositional heterogeneity and patterns of molecular evolution in the Drosophila genome.

The rates and patterns of molecular evolution in many eukaryotic organisms have been shown to be influenced by the compartmentalization of their genomes into fractions of distinct base composition and mutational properties. We have examined the Drosophila genome to explore relationships between the nucleotide content of large chromosomal segments and the base composition and rate of evolution of genes within those segments. Direct determination of the G + C contents of yeast artificial chromosome clones containing inserts of Drosophila melanogaster DNA ranging from 140-340 kb revealed significant heterogeneity in base composition. The G + C content of the large segments studied ranged from 36.9% G + C for a clone containing the hunchback locus in polytene region 85, to 50.9% G + C for a clone that includes the rosy region in polytene region 87. Unlike other organisms, however, there was no significant correlation between the base composition of large chromosomal regions and the base composition at fourfold degenerate nucleotide sites of genes encompassed within those regions. Despite the situation seen in mammals, there was also no significant association between base composition and rate of nucleotide substitution. These results suggest that nucleotide sequence evolution in Drosophila differs from that of many vertebrates and does not reflect distinct mutational biases, as a function of base composition, in different genomic regions. Significant negative correlations between codon-usage bias and rates of synonymous site divergence, however, provide strong support for an argument that selection among alternative codons may be a major contributor to variability in evolutionary rates within Drosophila genomes.

Animals↗

Codon preference reflects mistranslational constraints: a proposal.

Following the observation of lysine for arginine misincorporation at the poor choice codon arg-AGA, a comparison of codon usage patterns for highly expressed mRNA's in E. coli provides a basis for the proposal that the major codon preference is subject to mistranslational constraints. In addition, the codons are utilized, as well as arranged, to provide a hydropathically conservative amino acid as the most probable replacement resulting from a mistranslational event.

Arginine↗

Sequence analysis of the lpdV gene for lipoamide dehydrogenase of branched-chain-oxoacid dehydrogenase of Pseudomonas putida.

The production of two lipoamide dehydrogenases by Pseudomonas is so far unique. One, LPD-val, is the specific E3 component of the branched-chain-oxoacid dehydrogenase and the second, LPD-glc, is the E3 component of 2-oxoglutarate dehydrogenase and the L-factor of the glycine oxidation system. The objective of the present research was to determine the nucleotide sequence of the structural gene for LPD-val in order to compare its deduced amino acid structure with that of other redox-active disulfide flavoproteins. Northern blots using mRNA isolated from P. putida grown in media with branched-chain amino acids identified a transcript of 6.2 kb which is long enough to encode all the structural genes for the complex. The nucleotide sequence of the structural gene for LPD-val, lpdV, was determined and consists of 459 codons plus the stop codon. The open reading frame begins two bases after the stop codon for the E2 subunit and is composed of 66.3% G + C. Codon usage is characteristic of moderately strongly expressed genes. There is a ribosome-binding site preceding the ATG start codon and a strong candidate for a rho-independent terminator at the 3' end of the reading frame. The Mr of the protein encoded is 48,164 and when the Mr of FAD is added, the total Mr is 48,949, which is very close to the value of 49,000 obtained by SDS-polyacrylamide gel electrophoresis. Similarity comparisons of LPD-val with sequences of three other lipoamide dehydrogenases showed that LPD-val was somewhat more distantly related. It is probable that the lipoamide dehydrogenases and the glutathione and mercuric reductases evolved from a common ancestral flavoprotein.

3-Methyl-2-Oxobutanoate Dehydrogenase (Lipoamide)↗

Bradyrhizobium japonicum does not require alpha-ketoglutarate dehydrogenase for growth on succinate or malate.

The sucA gene, encoding the E1 component of alpha-ketoglutarate dehydrogenase, was cloned from Bradyrhizobium japonicum USDA110, and its nucleotide sequence was determined. The gene shows a codon usage bias typical of non-nif and non-fix genes from this bacterium, with 89.1% of the codons being G or C in the third position. A mutant strain of B. japonicum, LSG184, was constructed with the sucA gene interrupted by a kanamycin resistance marker. LSG184 is devoid of alpha-ketoglutarate dehydrogenase activity, indicating that there is only one copy of sucA in B. japonicum and that it is completely inactivated in the mutant. Batch culture experiments on minimal medium revealed that LSG184 grows well on a variety of carbon substrates, including arabinose, malate, succinate, beta-hydroxybutyrate, glycerol, formate, and galactose. The sucA mutant is not a succinate auxotroph but has a reduced ability to use glutamate as a carbon or nitrogen source and an increased sensitivity to growth inhibition by acetate, relative to the parental strain. Because LSG184 grows well on malate or succinate as its sole carbon source, we conclude that B. japonicum, unlike most other bacteria, does not require an intact tricarboxylic acid (TCA) cycle to meet its energy needs when growing on the four-carbon TCA cycle intermediates. Our data support the idea that B. japonicum has alternate energy-yielding pathways that could potentially compensate for inhibition of alpha-ketoglutarate dehydrogenase during symbiotic nitrogen fixation under oxygen-limiting conditions.

Amino Acid Sequence↗

Mycobacterial codon optimization enhances antigen expression and virus-specific immune responses in recombinant Mycobacterium bovis bacille Calmette-Guérin expressing human immunodeficiency virus type 1 Gag.

Although its potential for vaccine development is already known, the introduction of recombinant human immunodeficiency virus (HIV) genes to Mycobacterium bovis bacille Calmette-Guérin (BCG) has thus far elicited only limited responses. In order to improve the expression levels, we optimized the codon usage of the HIV type 1 (HIV-1) p24 antigen gene of gag (p24 gag) and established a codon-optimized recombinant BCG (rBCG)-p24 Gag which expressed a 40-fold-higher level of p24 Gag than did that of nonoptimized rBCG-p24 Gag. Inoculation of mice with the codon-optimized rBCG-p24 Gag elicited effective immunity, as evidenced by virus-specific lymphocyte proliferation, gamma interferon ELISPOT cell induction, and antibody production. In contrast, inoculation of animals with the nonoptimized rBCG-p24 Gag induced only low levels of immune responses. Furthermore, a dose as small as 0.01 mg of the codon-optimized rBCG per animal proved capable of eliciting immune responses, suggesting that even low doses of a codon-optimized rBCG-based vaccine could effectively elicit HIV-1-specific immune responses.

AIDS Vaccines↗

Cloning and analysis of the DNA polymerase-encoding gene from Thermus filiformis.

The gene encoding Thermus filiformis (Tfi) DNA polymerase was cloned and its nucleotide sequence was determined. The primary structure of Tfi DNA polymerase was deduced from its nucleotide sequence. Tfi DNA polymerase is comprised of 833 amino acid residues and its molecular mass was determined to be 93,890 Da. The deduced amino acid sequence of Tfi DNA polymerase showed a high sequence homology to E. coli DNA polymerase I-like DNA polymerases: 78.5% homology to Taq DNA polymerase, 78.4% to Tca DNA polymerase, and 41.8% to E. coli DNA polymerase I. An extremely high sequence identity was observed in the region containing polymerase activity. The G + C content of the coding region for the Tfi DNA polymerase gene was 68.5%, which was higher than that of the chromosomal DNA (65%). The G + C contents in the first, second, and third positions of the codons used were 71.8%, 40.9%, and 92.7% respectively. Codon usage in Tfi DNA polymerase was heavily biased towards the use of G + C in the third position. Rare codons with U or A as the third base were sometimes used to avoid using GA(A/T) TC and TCGA sequences, as they are recognition sites for the restriction endonucleases TfiI and TaqI.

Amino Acids↗

A study of the purine/pyrimidine codon occurrence with a reduced centered variable and an evaluation compared to the frequency statistic.

With the three-letter alphabet [R,Y,N] (R = purine, Y = pyrimidine, N = R or Y), there are 26 codons (NNN being excluded): RNN,...,NNY (six codons at two unspecified bases N), RRN,...,NYY (12 codons at one unspecified base N), RRR,...,YYY (eight specified codons). A statistical methodology that uses the codon frequency and a reduced centered variable leads to similar results for a codon occurrence study, regardless of gene function and regardless of a particular protein coding gene taxonomic population. Therefore, this variable can be considered a new codon usage index, whose use removes certain nonsignificant results found with the frequency statistic. This methodology identifies the common and rare codons (i.e., the codons having the highest and lowest occurrence) and leads to a model of codon evolution at three successive states: RNN, then RNY, and finally RYY. Some biological relations between this model and the YRY(N)6YRY preferential occurrence are also presented.

Base Sequence↗

Codon optimization of Candida rugosa lip1 gene for improving expression in Pichia pastoris and biochemical characterization of the purified recombinant LIP1 lipase.

An important industrial enzyme, Candida rugosa lipase (CRL) possesses several different isoforms encoded by the lip gene family (lip1-lip7), in which the recombinant LIP1 is the major form of the CRL multigene family. Previously, 19 of the nonuniversal serine codons (CTG) of the lip1 gene hav been successfully converted into universal serine codons (TCT) by overlap extension PCR-based multiple-site-directed mutagenesis to express an active recombinant LIP1 in the yeast Pichia pastoris. To improve the expression efficiency of recombinant LIP1 in P. pastoris, a regional synthetic gene fragment of lip1 near the 5' end of a transcript has been constructed to match P. pastoris-preferred codon usage for simple scale-up fermentation. The present results show that the production level (152 mg/L) of coLIP1 (codon-optimized LIP1) has an overall improvement of 4.6-fold relative to that (33 mg/L) of non-codon-optimized LIP1 with only half the cultivation time of P. pastoris. This finding demonstrates that the regional codon optimization the lip1 gene fragment at the 5' end can greatly increase the expression level of recombinant LIP1 in the P. pastoris system. More distinct biochemical properties of the purified recombinant LIP1 for further industrial applications are also determined and discussed in detail.

Base Sequence↗

Molecular structure of the Frankia spp. nifD-K intergenic spacer and design of Frankia genus compatible primer.

The nifD-K intergenic spacer (IGS) of ArI3 and ACoN24d were found to have a length 265 and 199 nucleotides, respectively. They are markedly less conserved than the two neighbouring genes and have, in some instances, a repeated structure reminiscent of an insertion event. The repeated sequence and the IGSs have no detectable homology with sequences in DNA databanks. The IGS has a stem-loop structure with a low folding energy, lower than that between nifH and nifD. No convincing alignment of IGS sequences could be obtained among Frankia strains. Only between ACoN24d and ArI3, which belong to the same genomic species, was the alignment good enough to permit detection of a doubly repeated structure. No promoter could be detected in the IGSs. The putative nifK open reading frame (ORF) in Frankia strain ArI3 has a length of 1587 nucleotides, starting with a GTG codon, preceded by a ribosome binding site of a structure similar to that of nifH (GGAGGN7). The codon usage was similar to that of previously sequenced Frankia genes with a strong bias toward G- and C-ending codons except in the case of glycine where GGT is frequent. Alignment of the three Frankia nifK sequences (EUN1f; ArI3 and ACoN24d) with those of other nitrogen-fixing bacteria permitted detection of a sequence conserved among the three Frankia strains but absent in the other sequences. A primer targeted to that region in combination with FGPD807-85 amplified the nifD-KIGS sequences of all Frankia strains (except the non-nitrogen-fixing Frankia strains CN3 and AgB1-9) and yet failed to amplify DNA of all other nitrogen-fixing bacteria.(ABSTRACT TRUNCATED AT 250 WORDS)

Actinomycetales↗

Molecular considerations in the evolution of bacterial genes.

Synonymous and nonsynonymous substitution rates at the loci encoding glyceraldehyde-3-phosphate dehydrogenase (gap) and outer membrane protein 3A (ompA) were examined in 12 species of enteric bacteria. By examining homologous sequences in species of varying degrees of relatedness and of known phylogenetic relationships, we analyzed the patterns of synonymous and nonsynonymous substitutions within and among these genes. Although both loci accumulate synonymous substitutions at reduced rates due to codon usage bias, portions of the gap and ompA reading frames show significant deviation in synonymous substitution rates not attributable to local codon bias. A paucity of synonymous substitutions in portions of the ompA gene may reflect selection for a novel mRNA secondary structure. In addition, these studies allow comparisons of homologous protein-coding sequences (gap) in plants, animals, and bacteria, revealing differences in evolutionary constraints on this glycolytic enzyme in these lineages.

Amino Acid Sequence↗

Selection on the codon bias of Chlamydomonas reinhardtii chloroplast genes and the plant psbA gene.

Plant chloroplast genes have a codon use that reflects the genome compositional bias of a high A+T content with the single exception of the highly translated psbA gene which codes for the photosystem II D1 protein. The codon usage of plant psbA corresponds more closely to the limited tRNA population of the chloroplast and is very similar to the codon use observed in the chloroplast genes of the green alga Chlamydomonas reinhardtii. This pattern of codon use may be an adaptation for increased translation efficiency. A correspondence between codon use of plant psbA and Chlamydomonas chloroplast genes and the tRNAs coded by the chloroplast genome, however, is not observed in all synonymous codon groups. It is shown here that the degree of correspondence between codon use and tRNA population in different synonymous groups is correlated with the second codon position composition. Synonymous groups with an A or T at the second codon position have a high representation of codons for which a complementary tRNA is coded by the chloroplast genome. Those with a G or C at the second position have an increased representation of codons that bind a chloroplast tRNA by wobble. It is proposed that the difference between synonymous groups in terms of codon adaptation to the tRNA population in plant psbA and Chlamydomonas chloroplast genes may be the result of differences in second position composition.

Animals↗

The oligopeptide permease (Opp) of the plant pathogen Xanthomonas axonopodis pv. citri.

The oligopeptide permease (Opp), a protein-dependent ABC transporter, has been found in the genome of Xanthomonas axonopodis pv. citri ( Xac), but not in Xanthomonas campestris pv. campestris ( Xcc). Sequence analysis indicated that 4 opp genes ( oppA, oppB, oppC, oppD/F), located in a 33.8-kbp DNA fragment present only in the Xac genome, are arranged in an operon-like structure and share highest sequence similarities with Streptomyces roseofulvus orthologs. Nonetheless, analyses of the GC content, codon usage, and transposon positioning suggested that the Xac opp operon does not have an exogenous origin. The presence of a stop codon at one of the ATP-binding domains of OppD/F would render the uptake system nonfunctional, but detection of a single polycistronic mRNA and periplasmic OppA in actively growing bacteria suggests that the Opp permease is active and could contribute to the distinct nutritional requirements and host specificities of the two Xanthomonas species.

Bacterial Proteins↗

Effects of rare codon clusters on high-level expression of heterologous proteins in Escherichia coli.

Within Escherichia coli and other species, a clear codon bias exists among the 61 amino acid codons found within the population of mRNA molecules, and the level of cognate tRNA appears directly proportional to the frequency of codon usage. Given this situation, one would predict translational problems with an abundant mRNA species containing an excess of rare low tRNA codons. Such a situation might arise after the initiation of transcription of a cloned heterologous gene in the E. coli host. Recent studies suggest clusters of AGG/AGA, CUA, AUA, CGA or CCC codons can reduce both the quantity and quality of the synthesized protein. In addition, it is likely that an excess of any of these codons, even without clusters, could create translational problems.

Arginine↗