Search PubMedSearch

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

The nuclear genomes of African and American trypanosomes are strikingly different.

We have investigated the compositional distributions of exons and their different codon positions, as well as the codon usage and amino-acid (aa) composition of the nuclear genomes of the African and American trypanosomes Trypanosoma brucei and T. cruzi. Very large differences between the two species were found in all the properties investigated. The most striking differences concern the compositional distributions of third codon positions and the extremely large nucleotide divergence of third codon position for homologous genes encoding proteins that are highly conserved in their aa sequences. Moreover, if coding sequences from each species are divided into two groups according to the GC levels in third codon positions, very different codon usages and aa compositions are found. This indicates a compositional compartmentalization in both genomes which had previously been detected in T. brucei (and T. equiperdum) by compositional fractionation.

Animals

Evolution of chromosome bands: molecular ecology of noncoding DNA.

Giemsa dark bands, G-bands, are a derived chromatin character that evolved along the chromosomes of early chordates. They are facultative heterochromatin reflecting acquisition of a late replication mechanism to repress tissue-specific genes. Subsequently, R-bands, the primitive chromatin state, became directionally GC rich as evidenced by Q-banding of mammalian and avian chromosomes. Contrary to predictions from the neutral mutation theory, noncoding DNA is positionally constrained along the banding pattern with short interspersed repeats in R-bands and long interspersed repeats in G-bands. Chromosomes seem dynamically stable: the banding pattern and gene arrangement along several human and murine autosomes has remained constant for 100 million years, whereas much of the noncoding DNA, especially retroposons, has changed. Several coding sequence attributes and probably mutation rates are determined more by where a gene lives than by what it does. R-band exons in homeotherms but not G-band exons have directionally acquired GC-rich wobble bases and the corresponding codon usage: CpG islands in mammals are specific to R-band exons, exons not facultatively heterochromatinized, and are independent of the tissue expression pattern of the gene. The dynamic organization of noncoding DNA suggests a feedback loop that could influence codon usage and stabilize the chromosome's chromatin pattern: DNA sequences determine affinities of----proteins that together form----a chromatin that modulates----rate constants for DNA modification that determine----DNA sequences. Theories of hierarchical selection and molecular ecology show how selection can act on Darwinian units of noncoding DNA at the genome level thus creating positionally constrained DNA and contributing minimal genetic load at the individual level.

Base Sequence

Genetics of lactobacilli: plasmids and gene expression.

This paper reviews the present knowledge of the structure and properties of small (< 5 kb) plasmids present in Lactobacillus spp. The data show that plasmids from Lactobacillus spp., like many plasmids from other Gram-positive bacteria, display a modular organization and replicate by a mechanism of rolling circle replication. Structurally, plasmids from lactobacilli are closely related to plasmids from other Gram-positive bacteria. They contain elements (plus- and minus origin of replication, element(s) for control of plasmid replication, mobilization function) showing extensive similarity to analogous elements in plasmids from these other organisms. It is believed that lactobacilli have acquired such elements by intra- and/or intergenic transfer mechanisms. The first part of the review is concluded with a description of plasmid vectors with a Lactobacillus replicon and integrative vectors, including data concerning their structural and segregational stability. In the second part of this review we describe the progress that has been made during the last few years in identifying and characterizing elements that control expression of genetic information in lactobacilli. Based on the sequence of eleven identified and twenty presumed promoters, some preliminary conclusions can be drawn regarding the structure of Lactobacillus promoters. A typical Lactobacillus promoter shows significant similarity to promoters from E. coli and B. subtilis. An analysis of published sequences of seventy genes indicates that the region encompassing the translation start codon AUG also shows extensive similarity to that of E. coli and B. subtilis. Codon usage of Lactobacillus genes is not random and shows interspecies as well as intraspecies heterogeneity. Interspecies differences may, in part, be explained by differences in G+C content of different lactobacilli. Differences in gene expression levels can, to a large extent, account for intraspecies differences of codon usage bias. Finally, we review the knowledge that has become available concerning protein secretion and heterologous gene expression in lactobacilli. This part is concluded with a compilation of data on the expression in Lactobacillus of heterologous genes under the control of their own promoter or under control of a Lactobacillus promoter.

Bacillus subtilis

Evidence that mutation patterns vary among Drosophila transposable elements.

In Drosophila melanogaster, codon usage in the open reading frames (ORFs) of transposable elements (TEs) differs greatly from that in other ORFs. In addition, while the ORFs from a single element are similar, there is considerable variation among elements. In the TE ORFs there are no indications of selection for the codons prevalent in the other D. melanogaster genes, but rather codon usage can be succinctly summarized in terms of the base composition at silent sites. We suggest that the particular silent site base composition of each TE is determined by an individual pattern of mutation. In many of the TEs there is an ORF encoding a protein with homology to reverse transcriptase; the amino acid sequences of these are quite divergent, and so it is possible that each of these incorporates certain mismatched bases at different frequencies during replication.

Animals

Analyses of frameshifting at UUU-pyrimidine sites.

Others have recently shown that the UUU phenylalanine codon is highly frameshift-prone in the 3'(rightward) direction at pyrimidine 3'contexts. Here, several approaches are used to analyze frameshifting at such sites. The four permutations of the UUU/C (phenylalanine) and CGG/U (arginine) codon pairs were examined because they vary greatly in their expected frameshifting tendencies. Furthermore, these synonymous sites allow direct tests of the idea that codon usage can control frameshifting. Frameshifting was measured for these dicodons embedded within each of two broader contexts: the Escherichia coli prfB (RF2 gene) programmed frameshift site and a 'normal' message site. The principal difference between these contexts is that the programmed frameshift contains a purine-rich sequence upstream of the slippery site that can base pair with the 3'end of 16 S rRNA (the anti-Shine-Dalgarno) to enhance frameshifting. In both contexts frameshift frequencies are highest if the slippery tRNAPhe is capable of stable base pairing in the shifted reading frame. This requirement is less stringent in the RF2 context, as if the Shine-Dalgarno interaction can help stabilize a quasi-stable rephased tRNA:message complex. It was previously shown that frameshifting in RF2 occurs more frequently if the codon 3'to the slippery site is read by a rare tRNA. Consistent with that earlier work, in the RF2 context frameshifting occurs substantially more frequently if the arginine codon is CGG, which is read by a rare tRNA. In contrast, in the 'normal' context frameshifting is only slightly greater at CGG than at CGU. It is suggested that the Shine-Dalgarno-like interaction elevates frameshifting specifically during the pause prior to translation of the second codon, which makes frameshifting exquisitely sensitive to the rate of translation of that codon. In both contexts frameshifting increases in a mutant strain that fails to modify tRNA base A37, which is 3'of the anticodon. Thus, those base modifications may limit frameshifting at UUU codons. Finally, statistical analyses show that UUU Ynn dicodons are extremely rare in E.coli genes that have highly biased codon usage.

Arginine

Computer-aided gene design.

A computer program, which runs on MS-DOS personal computers, is described that assists in the design of synthetic genes coding for proteins. The goal of the program is the design of a gene which (i) contains as many unique restriction sites as possible and (ii) uses a specific codon usage. The gene designed according to the criteria above is (i) suitable for 'modular mutagenesis' experiments and (ii) optimized for expression. The program 'reverse-translates' protein sequences into degenerated DNA sequences, generates a map of potential restriction sites and locates sequence positions where unique restriction sites can be accommodated. The nucleic acid sequence is then 'refined' according to a specific codon usage to remove any degeneration. Unique restriction sites, if potentially present, can be 'forced' into the degenerated nucleic acid sequence by using 'priority codes' assigned to different restriction sequences.

Amino Acid Sequence

Transgenic barley expressing a protein-engineered, thermostable (1,3-1,4)-beta-glucanase during germination.

The codon usage of a hybrid bacterial gene encoding a thermostable (1,3-1,4)-beta-glucanase was modified to match that of the barley (1,3-1,4)-beta-glucanase isoenzyme EII gene. Both the modified and unmodified bacterial genes were fused to a DNA segment encoding the barley high-pI alpha-amylase signal peptide downstream of the barley (1,3-1,4)-beta-glucanase isoenzyme EII gene promoter. When introduced into barley aleurone protoplasts, the bacterial gene with adapted codon usage directed synthesis of heat stable (1,3-1,4)-beta-glucanase, whereas activity of the heterologous enzyme was not detectable when protoplasts were transfected with the unmodified gene. In a different expression plasmid, the codon modified bacterial gene was cloned downstream of the barley high-pI alpha-amylase gene promoter and signal peptide coding region. This expression cassette was introduced into immature barley embryos together with plasmids carrying the bar and the uidA genes. Green, fertile plants were regenerated and approximately 75% of grains harvested from primary transformants synthesized thermostable (1,3-1,4)-beta-glucanase during germination. All three trans genes were detected in 17 progenies from a homozygous T1 plant.

Amino Acid Sequence

Codon preference in corynebacteria.

The codon usage (CU) of 34 genes from the closely related species, Brevibacterium lactofermentum and Corynebacterium glutamicum (BLCG), was analysed and compared with that of 23 genes from other Brevibacterium and Corynebacterium species. The G+C content of the BLCG genes ranged from 50 to 62%. A wider range was found in other corynebacterial genes (25-71%). The G+C contents of non-coding regions in glutamic acid bacteria are lower than those of the coding regions and both values are lower than the G+C content of ribosomal RNA (rRNA) sequences, suggesting an unusual biased mutation pressure. The CU and synonymous codon usage (SCU) analysis showed several common characteristics among the sequenced corynebacterial genes, consistent with the close relatedness of B. lactofermentum and C. glutamicum. A subset of 25 preferred codons were deduced from the presumably highly expressed genes and they encode most of the amino acid (aa) residues of the BLCG group. An analysis of the effective number of codons (Nc) was carried out in order to check the GC3s (G+C content at the silent third position of sense codons) dependence of the CU in corynebacteria. Nc values showed differences between the BLCG group and other corynebacterial sequences. A comparison of the most used codons for each aa showed a stronger similarity to Streptomyces than to Escherichia coli. The CU/SCU tables of corynebacteria are useful for identification of protein-coding regions, including start codons when they are uncertain, and for designing oligodeoxyribonucleotide probes from an aa sequence.

Base Sequence

Low codon bias and high rates of synonymous substitution in Drosophila hydei and D. melanogaster histone genes.

We have evaluated codon usage bias in Drosophila histone genes and have obtained the nucleotide sequence of a 5,161-bp D. hydei histone gene repeat unit. This repeat contains genes for all five histone proteins (H1, H2a, H2b, H3, and H4) and differs from the previously reported one by a second EcoRI site. These D. hydei repeats have been aligned to each other and to the 5.0-kb (i.e., long) and 4.8-kb (i.e., short) histone repeat types from D. melanogaster. In each species, base composition at synonymous sites is similar to the average genomic composition and approaches that in the small intergenic spacers of the histone gene repeats. Accumulation of synonymous changes at synonymous sites after the species diverged is quite high. Both of these features are consistent with the relatively low codon usage bias observed in these genes when compared with other Drosophila genes. Thus, the generalization that abundantly expressed genes in Drosophila have high codon bias and low rates of silent substitution does not hold for the histone genes.

Animals

msDNA of bacteria.

The msDNA-retron element represents the first prokaryotic member of the large and diverse retroelement family found in many eukaryotic genomes (Table II). This prokaryotic retroelement exists as a single copy element in the chromosome of two different bacterial groups: the common soil microbe M. xanthus and the enteric bacterium E. coli. It encodes an RT similar to the polymerases found in retroviruses, containing most of the strictly conserved amino acids found in all RTs. The RT is responsible for the production of an unusual extrachromosomal RNA-DNA molecule known as msDNA. Each composed of a short single strand of RNA and a short single strand of DNA, msDNAs vary considerably in their primary nucleotide sequences, but all share certain secondary structural features, including the unique 2',5' branch linkage that joins the 5' end of the DNA chain to the 2' position of an internal guanosine residue of the RNA strand. It is proposed that msDNA is synthesized by reverse transcription of a precursor RNA transcribed from a region of the retron containing the genes msr (encoding the RNA portion) and msd (encoding the DNA portion) and the ORF (encoding the RT). The precursor RNA transcript folds into a stable secondary structure that serves as both the primer and the template for the synthesis of msDNA. The msDNA-retron elements of E. coli are found in less than 10% of all strains observed, are heterogeneous in nature, and have an atypical aminoacid codon usage for this species, suggesting that this element was transmitted to E. coli by some other source. The presence of directly repeated 26-base-pair sequences flanking the junctions of the Ec67-retron of E. coli also suggests that it may be a mobile element. However, the msDNA-retrons of M. xanthus appear to be as old as other genes native to this species, based on codon-usage data for the RT genes and the fact that every strain of M. xanthus appears to have the same type of msDNA. If the msDNA-retron element originated with the myxobacteria, it would place the existence of retrons before the appearance of eukaryotic cells, suggesting that the bacterial element is perhaps the ancestral gene from which eukaryotic retroviruses and other retroelements evolved.(ABSTRACT TRUNCATED AT 400 WORDS)

Base Sequence

Genetic code 1990. Outlook.

The genetic code is evolving as shown by 9 departures from the universal code: 6 of them are in mitochondria and 3 are in nuclear codes. We propose that these changes are preceded by disappearance of a codon from coding sequences in mRNA of an organism or organelle. The function of the codon that disappears is taken by other, synonymous codons, so that there is no change in amino acid sequences of proteins. The deleted codon then reappears with a new function. Wobble pairing between anticodons and codons has evolved, starting with a single UNN anticodon pairing with 4 codons. Directional mutation pressure affects codon usage and may produce codon reassignments, especially of stop codons. Selenocysteine is coded by UGA, which is also a stop codon, and this anomaly is discussed. The outlook for discovery of more changes in the code is favorable, and open reading frames should be compared with actual sequential analyses of protein molecules in this search.

Anaerobiosis

Dietary arginine drives codon-dependent MHC class I translation and improves immunity in colon tumorigenesis and respiratory viral infection.

Amino acid levels fluctuate across diverse pathological conditions. Whether such amino acid modulations directly shape pathophysiology by regulating host gene expression remains unknown. We found that extracellular arginine restriction, observed in cancer and infection, represses specific arginine tRNAs-directly suppressing translation of major histocompatibility complex I (MHC class I) and antigen presentation. Arginine regulation of MHC class I was codon-usage dependent, as synonymous codon mutations prevented MHC class I modulation. Dietary arginine restriction impaired anti-viral immunity against influenza and SARS-CoV-2 and increased colon tumorigenesis. Conversely, increasing arginine availability via dietary supplementation or myeloid-specific arginase 1 deletion enhanced MHC class I protein levels, suppressed colon tumorigenesis, and improved viral infection outcomes. These disease-modulating effects were abolished in &#x3b2;2-microglobulin (B2m)-deficient mice. Thus, dietary modulation of a single amino acid critically influences codon-biased translation and MHC class I-mediated immunity to respiratory viral infections and cancer, revealing an unexpected mechanism and disease hazard for arginine deficiency and highlighting potential for amino acid-based translation modulation therapy.

Animals

Site-specific codon bias in bacteria.

Sequences of the gapA and ompA genes from 10 genera of enterobacteria have been analyzed. There is strong bias in codon usage, but different synonymous codons are preferred at different sites in the same gene. Site-specific preference for unfavored codons is not confined to the first 100 codons and is usually manifest between two codons utilizing the same tRNA. Statistical analyses, based on conclusions reached in an accompanying paper, show that the use of an unfavored codon at a given site in different genera is not due to common descent and must therefore be caused either by sequence-specific mutation or sequence-specific selection. Reasons are given for thinking that sequence-specific mutation cannot be responsible. We are unable to explain the preference between synonymous codons ending in C or T, but synonymous choice between A and G at third sites is largely explained by avoidance of AG-G (where the hyphen indicates the boundary between codons). We also observed that the preferred codon for proline in Enterobacter cloacea has changed from CCG to CCA.

Bacterial Outer Membrane Proteins

A Macintosh computer program for designing DNA sequences that code for specific peptides and proteins.

A computer program (PINCERS) is described for use in the design of synthetic genes and mixed-probe DNA sequences. A protein sequence is reverse translated with generation of synonymous codons at each position producing a degenerate sequence. In order to locate potential restriction enzyme sites, the degenerate sequence is searched with a library of restriction enzymes for sites that utilize any combination of synonymous codons. These sites are indicated in a map so that they may be incorporated into the synthetic gene sequence. The program allows the user to select the appropriate codon usage table for the organism of interest and then to set a threshold usage frequency below which codons are not generated. PINCERS may also be used to assist in planning the synthesis of mixed-probe DNA sequences for cross-hybridization experiments. It can identify regions of specified length with the protein sequence that have the least overall degeneracy, thereby minimizing the number of probes to be synthesized and, therefore, maximizing the concentration of a given probe sequence.

DNA

Codon bias variation in Staphylococcus aureus.

BACKGROUND: Staphylococcus aureus causes a multiplicity of human diseases acquired in community and healthcare settings alike around the globe. While most studies focus on coding changes to assess genome evolution and study genetic adaptation, interrogation of silent mutations in the form of synonymous codon usage bias is less well-studied. As such, understanding of patterns in codon bias at the gene and genome levels, and how codon bias impacts protein expression in S. aureus remains incomplete. METHODS: The codon bias of 2,565 protein encoding genes from NCTC 8325 was queried against all publicly available closed S. aureus genomes. Using public BioSample data, genomes were sorted by disease state, submitting institution, and collection site. Codon bias was assessed at the level of gene and genome using the codon adaptation index (CAI), calculated using 30S and 50S ribosomal genes. Gene set enrichment analysis was applied to determine associations between physiological functions, CAI gene scores, and interquartile ranges. CAI scores were also compared to an in vitro S. aureus proteomics database to correlate codon bias and protein expression. RESULTS: CAI scores varied within and between isolates at the gene and genome levels. Genes with ribosome-associated functions were most enriched among high CAI genes, and had low CAI interquartile ranges (IQR), suggesting selective pressure to maintain high expression of these genes across all S. aureus isolates. Genome sequences submitted by Aga Khan University Hospital, Nairobi, Kenya were most different from others. For the LAC USA 300 strain, CAI and protein expression were moderately positively correlated (cor&#x2009;=&#x2009;0.534, p&#x2009;<&#x2009;2.2e-16). CONCLUSIONS: Codon bias in S. aureus was shown to vary between gene, and to be a source of genetic variation between isolates; CAI and in vitro protein expression were positively correlated.

Staphylococcus aureus

Effects of consecutive AGG codons on translation in Escherichia coli, demonstrated with a versatile codon test system.

A system for testing the effects of specific codons on gene expression is described. Tandem test and control genes are contained in a transcription unit for bacteriophage T7 RNA polymerase in a multicopy plasmid, and nearly identical test and control mRNAs are generated from the primary transcript by RNase III cleavages. Their coding sequences, derived from T7 gene 9, are translated efficiently and have few low-usage codons of Escherichia coli. The upstream test gene contains a site for insertion of test codons, and the downstream control gene has a 45-codon deletion that allows test and control mRNAs and proteins to be separated by gel electrophoresis. Codons can be inserted among identical flanking codons after codon 13, 223, or 307 in codon test vectors pCT1, pCT2, and pCT3, respectively, the third site being six codons from the termination codon. The insertion of two to five consecutive AGG (low-usage) arginine codons selectively reduced the production of full-length test protein to extents that depended on the number of AGG codons, the site of insertion, and the amount of test mRNA. Production of aberrant proteins was also stimulated at high levels of mRNA. The effects occurred primarily at the translational level and were not produced by CGU (high-usage) arginine codons. Our results are consistent with the idea that sufficiently high levels of the AGG mRNA can cause essentially all of the tRNA(AGG) in the cell to become sequestered in translating peptidyl-tRNA(AGG) -mRNA-ribosome complexes stalled at the first of two consecutive AGG codons and that the approach of an upstream translating ribosome stimulates a stalled ribosome of frameshift, hop, or terminate translation.

Arginine

Primary structure of the reaction center from Rhodopseudomonas sphaeroides.

The reaction center is a pigment-protein complex that mediates the initial photochemical steps of photosynthesis. The amino-terminal sequences of the L, M, and H subunits and the nucleotide and derived amino acid sequences of the L and M structural genes from Rhodopseudomonas sphaeroides have previously been determined. We report here the sequence of the H subunit, completing the primary structure determination of the reaction center from R. sphaeroides. The nucleotide sequence of the gene encoding the H subunit was determined by the dideoxy method after subcloning fragments into single-stranded M13 phage vectors. This information was used to derive the amino acid sequence of the corresponding polypeptide. The termini of the primary structure of the H subunit were established by means of the amino and carboxy terminal sequences of the polypeptide. The data showed that the H subunit is composed of 260 residues, corresponding to a molecular weight of 28,003. A molecular weight of 100,858 for the reaction center was calculated from the primary structures of the subunits and the cofactors. Examination of the genes encoding the reaction center shows that the codon usage is strongly biased towards codons ending in G and C. Hydropathy analysis of the H subunit sequence reveals one stretch of hydrophobic residues near the amino terminus; the L and M subunits contain five such stretches. From a comparison of the sequences of homologous proteins found in bacterial reaction centers and photosystem II of plants, an evolutionary tree was constructed. The analysis of evolutionary relationships showed that the L and M subunits of reaction centers and the D1 and D2 proteins of photosystem II are descended from a common ancestor, and that the rate of change in these proteins was much higher in the first billion years after the divergence of the reaction center and photosystem II than in the subsequent billion years represented by the divergence of the species containing these proteins.

Amino Acid Sequence

Isolation and characterization of a Ustilago maydis glyceraldehyde-3-phosphate dehydrogenase-encoding gene.

The complete nucleotide sequence of the glyceraldehyde-3-phosphate dehydrogenase gene from the corn smut fungus Ustilago maydis is reported. The gene encodes a 337-amino acid protein, parts of which show sequence identity to corresponding regions of GAPDH-encoding genes from other organisms. A single, putative 407-bp intron interrupts the tenth codon. Codon usage is highly biased for codons ending in cytosine.

Amino Acid Sequence