Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Trends in codon and amino acid usage in Thermotoga maritima.

The usage of synonymous codons and the frequencies of amino acids were investigated in the complete genome of the bacterium Thermotoga maritima using a multivariate statistical approach. The GC3 content of each gene was the most prominent source of variation of codon usage. Surprisingly the usage of UGU and UGC (synonymous triplets coding for Cys, the least frequent amino acid in this species) was detected as the second most prominent source of variation. However, this result is probably an artifact due to the very low frequency of Cys together with the nonbiased composition of this genome. The third trend was related to the preferential usage of a subset of codons among highly expressed genes, and these triplets are presumed to be translationally optimal. Concerning the amino acid usage, the hydropathy level of each protein (and therefore the frequency of charged residues) was the main trend, while the second factor was related to the frequency of usage of the smaller residues, suggesting that the cell economy strongly influences the architecture of the proteins. The third axis of the analysis discriminated the usage of Phe, Tyr, Trp (aromatic residues) plus Cys, Met, and His. These six residues have in common the property of being the preferential targets of reactive oxygen species, and therefore the anaerobic condition of T. maritima is an important factor for the amino acid frequencies. Finally, the Cys content of each protein was the fourth trend.

Amino Acids↗

msDNA of bacteria.

The msDNA-retron element represents the first prokaryotic member of the large and diverse retroelement family found in many eukaryotic genomes (Table II). This prokaryotic retroelement exists as a single copy element in the chromosome of two different bacterial groups: the common soil microbe M. xanthus and the enteric bacterium E. coli. It encodes an RT similar to the polymerases found in retroviruses, containing most of the strictly conserved amino acids found in all RTs. The RT is responsible for the production of an unusual extrachromosomal RNA-DNA molecule known as msDNA. Each composed of a short single strand of RNA and a short single strand of DNA, msDNAs vary considerably in their primary nucleotide sequences, but all share certain secondary structural features, including the unique 2',5' branch linkage that joins the 5' end of the DNA chain to the 2' position of an internal guanosine residue of the RNA strand. It is proposed that msDNA is synthesized by reverse transcription of a precursor RNA transcribed from a region of the retron containing the genes msr (encoding the RNA portion) and msd (encoding the DNA portion) and the ORF (encoding the RT). The precursor RNA transcript folds into a stable secondary structure that serves as both the primer and the template for the synthesis of msDNA. The msDNA-retron elements of E. coli are found in less than 10% of all strains observed, are heterogeneous in nature, and have an atypical aminoacid codon usage for this species, suggesting that this element was transmitted to E. coli by some other source. The presence of directly repeated 26-base-pair sequences flanking the junctions of the Ec67-retron of E. coli also suggests that it may be a mobile element. However, the msDNA-retrons of M. xanthus appear to be as old as other genes native to this species, based on codon-usage data for the RT genes and the fact that every strain of M. xanthus appears to have the same type of msDNA. If the msDNA-retron element originated with the myxobacteria, it would place the existence of retrons before the appearance of eukaryotic cells, suggesting that the bacterial element is perhaps the ancestral gene from which eukaryotic retroviruses and other retroelements evolved.(ABSTRACT TRUNCATED AT 400 WORDS)

Base Sequence↗

Synonymous codon choices in the extremely GC-poor genome of Plasmodium falciparum: compositional constraints and translational selection.

We have analyzed the patterns of synonymous codon preferences of the nuclear genes of Plasmodium falciparum, a unicellular parasite characterized by an extremely GC-poor genome. When all genes are considered, codon usage is strongly biased toward A and T in third codon positions, as expected, but multivariate statistical analysis detects a major trend among genes. At one end genes display codon choices determined mainly by the extreme genome composition of this parasite, and very probably their expression level is low. At the other end a few genes exhibit an increased relative usage of a particular subset of codons, many of which are C-ending. Since the majority of these few genes is putatively highly expressed, we postulate that the increased C-ending codons are translationally optimal. In conclusion, while codon usage of the majority of P. falciparum genes is determined mainly by compositional constraints, a small number of genes exhibit translational selection.

Animals↗

Gene "volatility" is most unlikely to reveal adaptation.

It has recently been claimed that adaptive molecular evolution can be detected within single genome sequences by use of gene "volatility" scores. However, the approach used was entirely based on the assumption that synonymous codon usage is normally shaped by selection for low volatility; this is most unlikely to be true. Furthermore, even if that assumption could be justified, the method would clearly lack power, detecting only genes where a very large number of nonsynonymous substitutions had occurred. Volatility scores are susceptible to other influences. The unusually high volatilities of the Mycobacterium tuberculosis and Plasmodium falciparum genes that were identified as putatively having undergone adaptive changes were largely the result of internally repetitive structures, in which unusual codon usage was caused by the mechanisms that generated this repetition rather than by adaptive changes.

Adaptation, Physiological↗

Genetic code 1990. Outlook.

The genetic code is evolving as shown by 9 departures from the universal code: 6 of them are in mitochondria and 3 are in nuclear codes. We propose that these changes are preceded by disappearance of a codon from coding sequences in mRNA of an organism or organelle. The function of the codon that disappears is taken by other, synonymous codons, so that there is no change in amino acid sequences of proteins. The deleted codon then reappears with a new function. Wobble pairing between anticodons and codons has evolved, starting with a single UNN anticodon pairing with 4 codons. Directional mutation pressure affects codon usage and may produce codon reassignments, especially of stop codons. Selenocysteine is coded by UGA, which is also a stop codon, and this anomaly is discussed. The outlook for discovery of more changes in the code is favorable, and open reading frames should be compared with actual sequential analyses of protein molecules in this search.

Anaerobiosis↗

Dietary arginine drives codon-dependent MHC class I translation and improves immunity in colon tumorigenesis and respiratory viral infection.

Amino acid levels fluctuate across diverse pathological conditions. Whether such amino acid modulations directly shape pathophysiology by regulating host gene expression remains unknown. We found that extracellular arginine restriction, observed in cancer and infection, represses specific arginine tRNAs-directly suppressing translation of major histocompatibility complex I (MHC class I) and antigen presentation. Arginine regulation of MHC class I was codon-usage dependent, as synonymous codon mutations prevented MHC class I modulation. Dietary arginine restriction impaired anti-viral immunity against influenza and SARS-CoV-2 and increased colon tumorigenesis. Conversely, increasing arginine availability via dietary supplementation or myeloid-specific arginase 1 deletion enhanced MHC class I protein levels, suppressed colon tumorigenesis, and improved viral infection outcomes. These disease-modulating effects were abolished in β2-microglobulin (B2m)-deficient mice. Thus, dietary modulation of a single amino acid critically influences codon-biased translation and MHC class I-mediated immunity to respiratory viral infections and cancer, revealing an unexpected mechanism and disease hazard for arginine deficiency and highlighting potential for amino acid-based translation modulation therapy.

Animals↗

Site-specific codon bias in bacteria.

Sequences of the gapA and ompA genes from 10 genera of enterobacteria have been analyzed. There is strong bias in codon usage, but different synonymous codons are preferred at different sites in the same gene. Site-specific preference for unfavored codons is not confined to the first 100 codons and is usually manifest between two codons utilizing the same tRNA. Statistical analyses, based on conclusions reached in an accompanying paper, show that the use of an unfavored codon at a given site in different genera is not due to common descent and must therefore be caused either by sequence-specific mutation or sequence-specific selection. Reasons are given for thinking that sequence-specific mutation cannot be responsible. We are unable to explain the preference between synonymous codons ending in C or T, but synonymous choice between A and G at third sites is largely explained by avoidance of AG-G (where the hyphen indicates the boundary between codons). We also observed that the preferred codon for proline in Enterobacter cloacea has changed from CCG to CCA.

Bacterial Outer Membrane Proteins↗

A Macintosh computer program for designing DNA sequences that code for specific peptides and proteins.

A computer program (PINCERS) is described for use in the design of synthetic genes and mixed-probe DNA sequences. A protein sequence is reverse translated with generation of synonymous codons at each position producing a degenerate sequence. In order to locate potential restriction enzyme sites, the degenerate sequence is searched with a library of restriction enzymes for sites that utilize any combination of synonymous codons. These sites are indicated in a map so that they may be incorporated into the synthetic gene sequence. The program allows the user to select the appropriate codon usage table for the organism of interest and then to set a threshold usage frequency below which codons are not generated. PINCERS may also be used to assist in planning the synthesis of mixed-probe DNA sequences for cross-hybridization experiments. It can identify regions of specified length with the protein sequence that have the least overall degeneracy, thereby minimizing the number of probes to be synthesized and, therefore, maximizing the concentration of a given probe sequence.

DNA↗

Codon bias variation in Staphylococcus aureus.

BACKGROUND: Staphylococcus aureus causes a multiplicity of human diseases acquired in community and healthcare settings alike around the globe. While most studies focus on coding changes to assess genome evolution and study genetic adaptation, interrogation of silent mutations in the form of synonymous codon usage bias is less well-studied. As such, understanding of patterns in codon bias at the gene and genome levels, and how codon bias impacts protein expression in S. aureus remains incomplete. METHODS: The codon bias of 2,565 protein encoding genes from NCTC 8325 was queried against all publicly available closed S. aureus genomes. Using public BioSample data, genomes were sorted by disease state, submitting institution, and collection site. Codon bias was assessed at the level of gene and genome using the codon adaptation index (CAI), calculated using 30S and 50S ribosomal genes. Gene set enrichment analysis was applied to determine associations between physiological functions, CAI gene scores, and interquartile ranges. CAI scores were also compared to an in vitro S. aureus proteomics database to correlate codon bias and protein expression. RESULTS: CAI scores varied within and between isolates at the gene and genome levels. Genes with ribosome-associated functions were most enriched among high CAI genes, and had low CAI interquartile ranges (IQR), suggesting selective pressure to maintain high expression of these genes across all S. aureus isolates. Genome sequences submitted by Aga Khan University Hospital, Nairobi, Kenya were most different from others. For the LAC USA 300 strain, CAI and protein expression were moderately positively correlated (cor&#x2009;=&#x2009;0.534, p&#x2009;<&#x2009;2.2e-16). CONCLUSIONS: Codon bias in S. aureus was shown to vary between gene, and to be a source of genetic variation between isolates; CAI and in vitro protein expression were positively correlated.

Staphylococcus aureus↗

Expression of enterovirus 70 capsid protein VP1 in Escherichia coli.

The VP1 gene of enterovirus 70 (EV70) possesses a large number of Escherichia coli low-usage codons (11.0%) and a bacterial ribosome binding site complementary sequence (RBSCS) 5'-UGUCUCCUUUUC-3' flanking the codon 139. Plasmids containing EV70 cDNA encoding the full-length VP1 failed to express in E. coli (BL21(DE3), Rosetta 2(DE3) or Rosetta (DE3)pLysS). High expression (>8% of total protein) of recombinant VP1 (rVP1m) in E. coli required engineering of the encoding cDNA (conserved modification of the native cDNA) by simultaneous substitution of a rare-codon cluster located between codons 103 and 132, and replacement of the RBSCS-TCCTTT sequence. The rare-codon frequencies of the cDNAs encoding VP1 non-overlapping terminal fragments N138 (1-138 aa) and C170 (141-310 aa) are similar (10.9 and 11.2%, respectively). However, in E. coli, high expression of recombinant C170 (rC170) required no modification of the native cDNA whereas high expression of recombinant N138 (rN138m) required minimal synonymous substitution of the above rare-codon cluster. The rare-codon cluster of EV70 VP1 gene has five least-usage arginine codons (AGG/AGA) and three tandem rare-codon pairs (AGGAGG, CUAAGG, and AGACUA). Our results suggest that the rare-codon cluster (its rare codon arrangement per se and/or its related mRNA secondary structure(s)) and the RBSCS in EV70 VP1 gene, not the rare-codon frequency, constitute the key elements that suppress its expression in E. coli.

Binding Sites↗

Effects of consecutive AGG codons on translation in Escherichia coli, demonstrated with a versatile codon test system.

A system for testing the effects of specific codons on gene expression is described. Tandem test and control genes are contained in a transcription unit for bacteriophage T7 RNA polymerase in a multicopy plasmid, and nearly identical test and control mRNAs are generated from the primary transcript by RNase III cleavages. Their coding sequences, derived from T7 gene 9, are translated efficiently and have few low-usage codons of Escherichia coli. The upstream test gene contains a site for insertion of test codons, and the downstream control gene has a 45-codon deletion that allows test and control mRNAs and proteins to be separated by gel electrophoresis. Codons can be inserted among identical flanking codons after codon 13, 223, or 307 in codon test vectors pCT1, pCT2, and pCT3, respectively, the third site being six codons from the termination codon. The insertion of two to five consecutive AGG (low-usage) arginine codons selectively reduced the production of full-length test protein to extents that depended on the number of AGG codons, the site of insertion, and the amount of test mRNA. Production of aberrant proteins was also stimulated at high levels of mRNA. The effects occurred primarily at the translational level and were not produced by CGU (high-usage) arginine codons. Our results are consistent with the idea that sufficiently high levels of the AGG mRNA can cause essentially all of the tRNA(AGG) in the cell to become sequestered in translating peptidyl-tRNA(AGG) -mRNA-ribosome complexes stalled at the first of two consecutive AGG codons and that the approach of an upstream translating ribosome stimulates a stalled ribosome of frameshift, hop, or terminate translation.

Arginine↗

Primary structure of the reaction center from Rhodopseudomonas sphaeroides.

The reaction center is a pigment-protein complex that mediates the initial photochemical steps of photosynthesis. The amino-terminal sequences of the L, M, and H subunits and the nucleotide and derived amino acid sequences of the L and M structural genes from Rhodopseudomonas sphaeroides have previously been determined. We report here the sequence of the H subunit, completing the primary structure determination of the reaction center from R. sphaeroides. The nucleotide sequence of the gene encoding the H subunit was determined by the dideoxy method after subcloning fragments into single-stranded M13 phage vectors. This information was used to derive the amino acid sequence of the corresponding polypeptide. The termini of the primary structure of the H subunit were established by means of the amino and carboxy terminal sequences of the polypeptide. The data showed that the H subunit is composed of 260 residues, corresponding to a molecular weight of 28,003. A molecular weight of 100,858 for the reaction center was calculated from the primary structures of the subunits and the cofactors. Examination of the genes encoding the reaction center shows that the codon usage is strongly biased towards codons ending in G and C. Hydropathy analysis of the H subunit sequence reveals one stretch of hydrophobic residues near the amino terminus; the L and M subunits contain five such stretches. From a comparison of the sequences of homologous proteins found in bacterial reaction centers and photosystem II of plants, an evolutionary tree was constructed. The analysis of evolutionary relationships showed that the L and M subunits of reaction centers and the D1 and D2 proteins of photosystem II are descended from a common ancestor, and that the rate of change in these proteins was much higher in the first billion years after the divergence of the reaction center and photosystem II than in the subsequent billion years represented by the divergence of the species containing these proteins.

Amino Acid Sequence↗

Spinach holo-acyl carrier protein: overproduction and phosphopantetheinylation in Escherichia coli BL21(DE3), in vitro acylation, and enzymatic desaturation of histidine-tagged isoform I.

Spinach ACP isoform I was overexpressed in Escherichia coli BL21(DE3) using a gene synthesized from codons associated with high-level expression in E. coli. The synthetic gene has extensive changes in codon usage (23 of 77 total codons) relative to that of the originally synthesized plant gene (P. D. Beremand et al., 1987, Arch. Biochem. Biophys. 256, 90-100). After expression of the new synthetic gene, purified ACP and ACP-His6 were obtained in yields of up to 70 mg L-1 of culture medium, compared to approximately 1-6 mg L-1 of purified ACP obtained from the gene composed of predicted spinach codons. In either shaken flask or fermentation culture, approximately 15% conversion to holo-ACP or holo-ACP-His6 was obtained regardless of the level of protein expression. However, coexpression of ACP-His6 with E. coli holo-ACP synthase in E. coli BL21(DE3) during pH- and dissolved O2-controlled fermentation routinely yielded greater than 95% conversion to holo-ACP-His6. Electrospray ionization mass spectrometric analysis of the purified recombinant ACPs revealed that the amino terminal Met was efficiently removed, but only if the bacterial cell lysates were prepared in the absence of EDTA. This observation is consistent with the inhibition of endogenous Met-aminopeptidase by removal of catalytically essential Co(II) and introduces the importance of considering the catalytic properties of host enzymes providing ad hoc posttranslational modification of recombinant proteins. Stearoyl-ACP-His6 was shown to be indistinguishable from stearoyl-ACP as a substrate for enzymatic acylation and desaturation. In combination, these studies provide a coordinated scheme to produce and characterize quantities of acyl-ACPs sufficient to support expanded biophysical and structural studies.

Acyl Carrier Protein↗

Isolation and characterization of a Ustilago maydis glyceraldehyde-3-phosphate dehydrogenase-encoding gene.

The complete nucleotide sequence of the glyceraldehyde-3-phosphate dehydrogenase gene from the corn smut fungus Ustilago maydis is reported. The gene encodes a 337-amino acid protein, parts of which show sequence identity to corresponding regions of GAPDH-encoding genes from other organisms. A single, putative 407-bp intron interrupts the tenth codon. Codon usage is highly biased for codons ending in cytosine.

Amino Acid Sequence↗

Enhanced readthrough of opal (UGA) stop codons and production of Mycoplasma pneumoniae P1 epitopes in Escherichia coli.

Expression of mycoplasma sequences in Escherichia coli is often hindered by an unusual mycoplasmal codon usage pattern: the UGA stop codon is utilized for tryptophan. This may result in the truncation of cloned proteins and may prevent the detection of products of many cloned genes. To circumvent this translation barrier, we have developed an expression system for the production of mycoplasma proteins in E. coli. The efficiency of an opal suppressor tRNA (trpT176) was augmented with other suppressor mutations (prfB3 or rrsB(SuUGA-delta C1054)) which influence termination events. System efficacy was analyzed by employing suppressor mutations in the expression of TGA-containing sequences from the P1 protein-encoding gene of Mycoplasma pneumoniae.

Adhesins, Bacterial↗

Molecular cloning of, and phylogenetic analysis of, an actin in Naegleria fowleri.

We cloned and sequenced an intronless actin gene from the amoebo-flagellate Naegleria fowleri, LEE strain, an opportunistic pathogen of man. Codon usage and third-position-codon nucleotide frequency were significantly different from Acanthamoeba, another amoeba genus which also includes opportunistic pathogens of man. Between the two amoebae, actin peptide sequences were 92.8% similar, while nucleotide sequences were only 70% similar. A phylogenetic reconstruction of actin amino acid sequences, using a distance method, placed Naegleria in a cluster with Plasmodium and Entamoeba.

Actins↗

Association of the phi nucleotide with codon bias, amino acid usage and expressivity: differences between Bacillus subtilis and Escherichia coli.

By measuring the non-randomness in Shine-Dalgarno regions it was recently shown that the compositional non-randomness peaks approximately 10 nucleotides upstream of the start codons. This position, termed the phi position, was furthermore shown to be associated with certain characteristics of the gene/protein and start codon usage. This raises the question whether codon usage in general is associated with the phi position. In this study, the connection between the phi nucleotide and general codon usage, both gene-wide and at the level of individual amino acids, was studied in Eschericia coli and Bacillus subtilis. E. coli but not B. subtilis shows a strong general association between the phi position and codon usage bias. In both species, the genes with higher expressivity show stronger conservation in the Shine-Dalgarno region compared to the genes with lower expressivity.

Amino Acids↗

Cloning and sequencing of a phospholipase C gene of Clostridium perfringens.

The gene encoding phospholipase C (alpha-toxin) of Clostridium perfringens was cloned into lambda gt10. The maximal size of the coding region was 1.4 kb and the minimum was 1.1 kb as determined by subcloning into the vector pBR322 and testing for activity. The nucleotide sequence of this region contained a single open reading frame of 1194 bp corresponding to a protein of Mr 45473 with a possible N-terminal signal sequence of 28 amino acids which when removed, would give a mature protein of Mr 42521. This is in good agreement with the reported size of 43 kDa. The coding region has a dG + dC content of 33.7%, and the codon usage displays a pronounced preference for codons with the lowest dG + dC content.

Amino Acid Sequence↗