Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Codon bias variation in Staphylococcus aureus.

BACKGROUND: Staphylococcus aureus causes a multiplicity of human diseases acquired in community and healthcare settings alike around the globe. While most studies focus on coding changes to assess genome evolution and study genetic adaptation, interrogation of silent mutations in the form of synonymous codon usage bias is less well-studied. As such, understanding of patterns in codon bias at the gene and genome levels, and how codon bias impacts protein expression in S. aureus remains incomplete. METHODS: The codon bias of 2,565 protein encoding genes from NCTC 8325 was queried against all publicly available closed S. aureus genomes. Using public BioSample data, genomes were sorted by disease state, submitting institution, and collection site. Codon bias was assessed at the level of gene and genome using the codon adaptation index (CAI), calculated using 30S and 50S ribosomal genes. Gene set enrichment analysis was applied to determine associations between physiological functions, CAI gene scores, and interquartile ranges. CAI scores were also compared to an in vitro S. aureus proteomics database to correlate codon bias and protein expression. RESULTS: CAI scores varied within and between isolates at the gene and genome levels. Genes with ribosome-associated functions were most enriched among high CAI genes, and had low CAI interquartile ranges (IQR), suggesting selective pressure to maintain high expression of these genes across all S. aureus isolates. Genome sequences submitted by Aga Khan University Hospital, Nairobi, Kenya were most different from others. For the LAC USA 300 strain, CAI and protein expression were moderately positively correlated (cor&#x2009;=&#x2009;0.534, p&#x2009;<&#x2009;2.2e-16). CONCLUSIONS: Codon bias in S. aureus was shown to vary between gene, and to be a source of genetic variation between isolates; CAI and in vitro protein expression were positively correlated.

Staphylococcus aureus↗

Expression of enterovirus 70 capsid protein VP1 in Escherichia coli.

The VP1 gene of enterovirus 70 (EV70) possesses a large number of Escherichia coli low-usage codons (11.0%) and a bacterial ribosome binding site complementary sequence (RBSCS) 5'-UGUCUCCUUUUC-3' flanking the codon 139. Plasmids containing EV70 cDNA encoding the full-length VP1 failed to express in E. coli (BL21(DE3), Rosetta 2(DE3) or Rosetta (DE3)pLysS). High expression (>8% of total protein) of recombinant VP1 (rVP1m) in E. coli required engineering of the encoding cDNA (conserved modification of the native cDNA) by simultaneous substitution of a rare-codon cluster located between codons 103 and 132, and replacement of the RBSCS-TCCTTT sequence. The rare-codon frequencies of the cDNAs encoding VP1 non-overlapping terminal fragments N138 (1-138 aa) and C170 (141-310 aa) are similar (10.9 and 11.2%, respectively). However, in E. coli, high expression of recombinant C170 (rC170) required no modification of the native cDNA whereas high expression of recombinant N138 (rN138m) required minimal synonymous substitution of the above rare-codon cluster. The rare-codon cluster of EV70 VP1 gene has five least-usage arginine codons (AGG/AGA) and three tandem rare-codon pairs (AGGAGG, CUAAGG, and AGACUA). Our results suggest that the rare-codon cluster (its rare codon arrangement per se and/or its related mRNA secondary structure(s)) and the RBSCS in EV70 VP1 gene, not the rare-codon frequency, constitute the key elements that suppress its expression in E. coli.

Binding Sites↗

Effects of consecutive AGG codons on translation in Escherichia coli, demonstrated with a versatile codon test system.

A system for testing the effects of specific codons on gene expression is described. Tandem test and control genes are contained in a transcription unit for bacteriophage T7 RNA polymerase in a multicopy plasmid, and nearly identical test and control mRNAs are generated from the primary transcript by RNase III cleavages. Their coding sequences, derived from T7 gene 9, are translated efficiently and have few low-usage codons of Escherichia coli. The upstream test gene contains a site for insertion of test codons, and the downstream control gene has a 45-codon deletion that allows test and control mRNAs and proteins to be separated by gel electrophoresis. Codons can be inserted among identical flanking codons after codon 13, 223, or 307 in codon test vectors pCT1, pCT2, and pCT3, respectively, the third site being six codons from the termination codon. The insertion of two to five consecutive AGG (low-usage) arginine codons selectively reduced the production of full-length test protein to extents that depended on the number of AGG codons, the site of insertion, and the amount of test mRNA. Production of aberrant proteins was also stimulated at high levels of mRNA. The effects occurred primarily at the translational level and were not produced by CGU (high-usage) arginine codons. Our results are consistent with the idea that sufficiently high levels of the AGG mRNA can cause essentially all of the tRNA(AGG) in the cell to become sequestered in translating peptidyl-tRNA(AGG) -mRNA-ribosome complexes stalled at the first of two consecutive AGG codons and that the approach of an upstream translating ribosome stimulates a stalled ribosome of frameshift, hop, or terminate translation.

Arginine↗

Primary structure of the reaction center from Rhodopseudomonas sphaeroides.

The reaction center is a pigment-protein complex that mediates the initial photochemical steps of photosynthesis. The amino-terminal sequences of the L, M, and H subunits and the nucleotide and derived amino acid sequences of the L and M structural genes from Rhodopseudomonas sphaeroides have previously been determined. We report here the sequence of the H subunit, completing the primary structure determination of the reaction center from R. sphaeroides. The nucleotide sequence of the gene encoding the H subunit was determined by the dideoxy method after subcloning fragments into single-stranded M13 phage vectors. This information was used to derive the amino acid sequence of the corresponding polypeptide. The termini of the primary structure of the H subunit were established by means of the amino and carboxy terminal sequences of the polypeptide. The data showed that the H subunit is composed of 260 residues, corresponding to a molecular weight of 28,003. A molecular weight of 100,858 for the reaction center was calculated from the primary structures of the subunits and the cofactors. Examination of the genes encoding the reaction center shows that the codon usage is strongly biased towards codons ending in G and C. Hydropathy analysis of the H subunit sequence reveals one stretch of hydrophobic residues near the amino terminus; the L and M subunits contain five such stretches. From a comparison of the sequences of homologous proteins found in bacterial reaction centers and photosystem II of plants, an evolutionary tree was constructed. The analysis of evolutionary relationships showed that the L and M subunits of reaction centers and the D1 and D2 proteins of photosystem II are descended from a common ancestor, and that the rate of change in these proteins was much higher in the first billion years after the divergence of the reaction center and photosystem II than in the subsequent billion years represented by the divergence of the species containing these proteins.

Amino Acid Sequence↗

Spinach holo-acyl carrier protein: overproduction and phosphopantetheinylation in Escherichia coli BL21(DE3), in vitro acylation, and enzymatic desaturation of histidine-tagged isoform I.

Spinach ACP isoform I was overexpressed in Escherichia coli BL21(DE3) using a gene synthesized from codons associated with high-level expression in E. coli. The synthetic gene has extensive changes in codon usage (23 of 77 total codons) relative to that of the originally synthesized plant gene (P. D. Beremand et al., 1987, Arch. Biochem. Biophys. 256, 90-100). After expression of the new synthetic gene, purified ACP and ACP-His6 were obtained in yields of up to 70 mg L-1 of culture medium, compared to approximately 1-6 mg L-1 of purified ACP obtained from the gene composed of predicted spinach codons. In either shaken flask or fermentation culture, approximately 15% conversion to holo-ACP or holo-ACP-His6 was obtained regardless of the level of protein expression. However, coexpression of ACP-His6 with E. coli holo-ACP synthase in E. coli BL21(DE3) during pH- and dissolved O2-controlled fermentation routinely yielded greater than 95% conversion to holo-ACP-His6. Electrospray ionization mass spectrometric analysis of the purified recombinant ACPs revealed that the amino terminal Met was efficiently removed, but only if the bacterial cell lysates were prepared in the absence of EDTA. This observation is consistent with the inhibition of endogenous Met-aminopeptidase by removal of catalytically essential Co(II) and introduces the importance of considering the catalytic properties of host enzymes providing ad hoc posttranslational modification of recombinant proteins. Stearoyl-ACP-His6 was shown to be indistinguishable from stearoyl-ACP as a substrate for enzymatic acylation and desaturation. In combination, these studies provide a coordinated scheme to produce and characterize quantities of acyl-ACPs sufficient to support expanded biophysical and structural studies.

Acyl Carrier Protein↗

Isolation and characterization of a Ustilago maydis glyceraldehyde-3-phosphate dehydrogenase-encoding gene.

The complete nucleotide sequence of the glyceraldehyde-3-phosphate dehydrogenase gene from the corn smut fungus Ustilago maydis is reported. The gene encodes a 337-amino acid protein, parts of which show sequence identity to corresponding regions of GAPDH-encoding genes from other organisms. A single, putative 407-bp intron interrupts the tenth codon. Codon usage is highly biased for codons ending in cytosine.

Amino Acid Sequence↗

Enhanced readthrough of opal (UGA) stop codons and production of Mycoplasma pneumoniae P1 epitopes in Escherichia coli.

Expression of mycoplasma sequences in Escherichia coli is often hindered by an unusual mycoplasmal codon usage pattern: the UGA stop codon is utilized for tryptophan. This may result in the truncation of cloned proteins and may prevent the detection of products of many cloned genes. To circumvent this translation barrier, we have developed an expression system for the production of mycoplasma proteins in E. coli. The efficiency of an opal suppressor tRNA (trpT176) was augmented with other suppressor mutations (prfB3 or rrsB(SuUGA-delta C1054)) which influence termination events. System efficacy was analyzed by employing suppressor mutations in the expression of TGA-containing sequences from the P1 protein-encoding gene of Mycoplasma pneumoniae.

Adhesins, Bacterial↗

Molecular cloning of, and phylogenetic analysis of, an actin in Naegleria fowleri.

We cloned and sequenced an intronless actin gene from the amoebo-flagellate Naegleria fowleri, LEE strain, an opportunistic pathogen of man. Codon usage and third-position-codon nucleotide frequency were significantly different from Acanthamoeba, another amoeba genus which also includes opportunistic pathogens of man. Between the two amoebae, actin peptide sequences were 92.8% similar, while nucleotide sequences were only 70% similar. A phylogenetic reconstruction of actin amino acid sequences, using a distance method, placed Naegleria in a cluster with Plasmodium and Entamoeba.

Actins↗

Association of the phi nucleotide with codon bias, amino acid usage and expressivity: differences between Bacillus subtilis and Escherichia coli.

By measuring the non-randomness in Shine-Dalgarno regions it was recently shown that the compositional non-randomness peaks approximately 10 nucleotides upstream of the start codons. This position, termed the phi position, was furthermore shown to be associated with certain characteristics of the gene/protein and start codon usage. This raises the question whether codon usage in general is associated with the phi position. In this study, the connection between the phi nucleotide and general codon usage, both gene-wide and at the level of individual amino acids, was studied in Eschericia coli and Bacillus subtilis. E. coli but not B. subtilis shows a strong general association between the phi position and codon usage bias. In both species, the genes with higher expressivity show stronger conservation in the Shine-Dalgarno region compared to the genes with lower expressivity.

Amino Acids↗

Cloning and sequencing of a phospholipase C gene of Clostridium perfringens.

The gene encoding phospholipase C (alpha-toxin) of Clostridium perfringens was cloned into lambda gt10. The maximal size of the coding region was 1.4 kb and the minimum was 1.1 kb as determined by subcloning into the vector pBR322 and testing for activity. The nucleotide sequence of this region contained a single open reading frame of 1194 bp corresponding to a protein of Mr 45473 with a possible N-terminal signal sequence of 28 amino acids which when removed, would give a mature protein of Mr 42521. This is in good agreement with the reported size of 43 kDa. The coding region has a dG + dC content of 33.7%, and the codon usage displays a pronounced preference for codons with the lowest dG + dC content.

Amino Acid Sequence↗

Characterization of catalase transcripts and their differential expression in maize.

In maize, the three unlinked catalase (EC 1.11.1.6) structural genes (Cat1, Cat2 and Cat3) are differentially expressed temporally, spatially and in response to environmental signals in the developing seedling. In order to understand more fully the molecular mechanisms involved in catalase gene expression, full-length cDNA clones representing the maize Cat1, Cat2 and Cat3 transcripts were isolated and characterized. DNA sequence analysis confirmed that each cDNA encodes a unique catalase protein. Gene-specific probes for the three maize catalase cDNAs were isolated and used to probe blots of poly(A)+ RNA isolated from various maize tissues. Cat1 mRNA was found in scutella, milky endosperm of immature kernels, leaves and epicotyls. The Cat2 mRNA was present primarily in post-germinative scutella, with lower levels in leaves and epicotyls. Cat3 mRNA was detected primarily in epicotyls and, to a lesser extent, in leaves and scutella. The gene-specific probes hybridized with maize genomic DNA blots in simple, but unique patterns, indicating that there is one, or a very few copies of each catalase gene. The coding region of the Cat3 cDNA comprised 66% G + C, which led to a strong codon usage bias in this gene. This codon bias was also seen with the Cat2 transcripts, but not with those for Cat1. A high degree of similarity was found between the maize catalase nucleic acid and deduced amino-acid sequences and those of sweet potato and rat liver catalase.

Amino Acid Sequence↗

Structural proteins of mycobacteriophage I3: cloning, expression and sequence analysis of a gene encoding a 70-kDa structural protein.

The structural proteins of mycobacteriophage I3 have been analysed by sodium dodecyl sulfate-polyacrylamide-gel electrophoresis (SDS-PAGE), radioiodination and immunoblotting. Based on their abundance the 34- and 70-kDa bands appeared to represent the major structural proteins. Successful cloning and expression of the 70-kDa protein-encoding gene of phage I3 in Escherichia coli and its complete nucleotide sequence determination have been accomplished. A second (partial) open reading frame following the stop codon for the 70-kDa protein was also identified within the cloned fragment. The deduced amino-acid sequence of the 70-kDa protein and the codon usage patterns indicated the preponderance of codons, as predicted from the high G+C content of the genomic DNA of phage I3.

Amino Acid Sequence↗

A bacteriophage reagent for Salmonella: molecular studies on Felix 01.

Felix 01 (F01) is a bacteriophage originally isolated by Felix and Callow which lyses almost all Salmonella strains and has been widely used as a diagnostic test for this genus. Molecular information about this phage is entirely lacking. In the present study, the DNA of the phage was found to be a double-stranded linear molecule of about 80 kb. 11.5 kb has been sequenced and in this region A + T content is 60%. There are relatively few restriction endonuclease cleavage sites in the native genome and clones show this is due to their absence rather than modification. A restriction map of the genome has been constructed. The ends of the molecule cannot be ligated although they contain 5' phosphates. At least 60% of the genome must encode proteins. In the sequenced portion, many open reading frames exist and these are tightly packed together. These have been examined for homology to published proteins but only 1 to 17 shows similarity to known proteins. F01 is therefore the prototype of a new phage family. On the basis of restriction sites, codon usage and the distribution of nonsense codons in the unused reading frames, a strong case can be made for natural selection that reacts to mRNA structure and function.

Base Sequence↗

The complete mitochondrial DNA sequence of the horseshoe crab Limulus polyphemus.

We determined the complete 14,985-nt sequence of the mitochondrial DNA of the horseshoe crab Limulus polyphemus (Arthropoda: Xiphosura). This mtDNA encodes the 13 protein, 2 rRNA, and 22 tRNA genes typical for metazoans. The arrangement of these genes and about half of the sequence was reported previously; however, the sequence contained a large number of errors, which are corrected here. The two strands of Limulus mtDNA have significantly different nucleotide compositions. The strand encoding most mitochondrial proteins has 1. 25 times as many A's as T's and 2.33 times as many C's as G's. This nucleotide bias correlates with the biases in amino acid content and synonymous codon usage in proteins encoded by different strands and with the number of non-Watson-Crick base pairs in the stem regions of encoded tRNAs. The sizes of most mitochondrial protein genes in Limulus are either identical to or slightly smaller than those of their Drosophila counterparts. The usage of the initiation and termination codons in these genes seems to follow patterns that are conserved among most arthropod and some other metazoan mitochondrial genomes. The noncoding region of Limulus mtDNA contains a potential stem-loop structure, and we found a similar structure in the noncoding region of the published mtDNA of the prostriate tick Ixodes hexagonus. A simulation study was designed to evaluate the significance of these secondary structures; it revealed that they are statistically significant. No significant, comparable structure can be identified for the metastriate ticks Rhipicephalus sanguineus and Boophilus microplus. The latter two animals also share a mitochondrial gene rearrangement and an unusual structure of mt-tRNA(C) that is exactly the same association of changes as previously reported for a group of lizards. This suggests that the changes observed are not independent and that the stem-loop structure found in the noncoding regions of Limulus and Ixodes mtDNA may play the same role as that between trnN and trnC in vertebrates, i.e., the role of lagging strand origin of replication.

Animals↗

Amino acid translation program for full-length cDNA sequences with frameshift errors.

Here we present an amino acid translation program designed to suggest the position of experimental frameshift errors and predict amino acid sequences for full-length cDNA sequences having phred scores. Our program generates artificial insertions into artificial deletions from low-accuracy positions of the original sequence, thereby generating many candidate sequences. The validity of the most probable sequence (the likelihood that it represents the actual protein) is evaluated by using a score (V(a)) that is calculated in light of the Kozak consensus, preferred codon usage, and position of the initiation codon. To evaluate the software, we have used a database in which, out of 612 cDNA sequences, 524 (86%) carried 773 frameshift errors in the coding sequence. Our software detected and corrected 48% of the total frameshift errors in 62% of the total cDNA sequences with frameshift errors. The false positive rate of frameshift correction was 9%, and 91% of the suggested frameshifts were true.

Base Composition↗

Detection and evaluation of intron retention events in the human transcriptome.

Alternative splicing is a very frequent phenomenon in the human transcriptome. There are four major types of alternative splicing: exon skipping, alternative 3' splice site, alternative 5' splice site, and intron retention. Here we present a large-scale analysis of intron retention in a set of 21,106 known human genes. We observed that 14.8% of these genes showed evidence of at least one intron retention event. Most of the events are located within the untranslated regions (UTRs) of human transcripts. For those retained introns interrupting the coding region, the GC content, codon usage, and the frequency of stop codons suggest that these sequences are under selection for coding potential. Furthermore, 26% of the introns within the coding region participate in the coding of a protein domain. A comparison with mouse shows that at least 22% of all informative examples of retained introns in human are also present in the mouse transcriptome. We discuss that the data we present suggest that a significant fraction of the observed events is not spurious and might reflect biological significance. The analyses also allowed us to generate a reliable set of intron retention events that can be used for the identification of splicing regulatory elements.

Animals↗

Molecular evolutionary analysis of a histone gene repeating unit from Drosophila simulans.

A repeating unit of the histone gene cluster from Drosophila simulans containing the H1, H2A, H2B and H4 genes (the H3 gene region has already been analyzed) was cloned and analyzed. A nucleotide sequence of about 4.6 kbp was determined to study the nucleotide divergence and molecular evolution of the histone gene cluster. Comparison of the structure and nucleotide sequence with those of Drosophila melanogaster showed that the four histone genes were located at identical positions and in the same directions. The proportion of different nucleotide sites was 6.3% in total. The amino acid sequence of H1 was divergent, with a 5.1% difference. However, no amino acid change has been observed for the other three histone proteins. Analysis of the GC contents and the base substitution patterns in the two lineages, D. melanogaster and D. simulans, with a common ancestor showed the following. 1) A strong negative correlation was found between the GC content and the nucleotide divergence in the whole repeating unit. 2) The mode of molecular evolution previously found for the H3 gene was also observed for the whole repeating unit of histone genes; the nucleotide substitutions were stationary in the 3' and spacer regions, and there was a directional change of the codon usage to the AT-rich codons. 3) No distinct difference in the mode or pattern of molecular evolution was detected for the histone gene repeating unit in the D. melanogaster and D. simulans lineages. These results suggest that selectional pressure for the coding regions of histones, which eliminate A and T, is less effective in the D. melanogaster and D. simulans lineages than in the other GC-rich species.

Amino Acid Sequence↗

The complete nucleotide sequence of region 1 of the CFA/I fimbrial operon of human enterotoxigenic Escherichia coli.

The production of the plasmid-encoded fimbrial antigen CFA/I of enterotoxigenic Escherichia coli requires two DNA regions: CFA/I region 1 and CFA/I region 2. These two regions are separated by about 40 kb on the wildtype plasmid. CFA/I region 1 contains the structural genes, whereas CFA/I region 2 contains a positive regulator. The first two genes (cfaA and cfaB) and the cfaD' sequence of region 1 have already been described. Here the total nucleotide sequence of region 1 is presented. Two new genes in region 1 are described, named cfaC and cfaE. The GC content of the genes in region 1 is 33.6% which is substantially lower than normally found in E. coli genes (50%). The codon usage also differs from the standard codons used in E. coli.

Amino Acid Sequence↗