Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Codon catalog usage is a genome strategy modulated for gene expressivity.

The nucleic acid sequence bank now contains 161 mRNAs, 43 new genes are added. One sequence, that of B. mori fibroin, is dropped due to uncertainty on the starting point for translation. Frequencies of all codons are given for each gene added and for each genome type in the total bank. A new series of correspondence analyses on codon use is presented, substantiating the genome hypothesis. Internal regulation of mRNA expression by different third base choices between quartet and duet codons is proposed for bacterial genes.

Amino Acid Sequence↗

The targeting of somatic hypermutation.

Somatic hypermutation does not occur randomly within immunoglobulin V genes but, rather, is preferentially targeted to certain nucleotide positions (hot spots) and away from others (cold spots). Cold spots often coincide with residues essential for V gene folding. Hotspots, which appear to be strategically located to favour affinity maturation, are most frequently located in the CDRs (particularly CDR1) though conserved hotspots are also found at the base of FR3. Hotspots are in part created by local DNA sequence and the strong biases of codon usage in V genes indicate that the genes have evolved such that somatic hypermutation is targeted to those parts of the V where it is likely to prove most useful. These features of mutational hotspots and biased codon usage are also evident in V genes of lower animals suggesting that diversification by strategic targeting of non-templated mutation may have evolved early in antigen receptor evolution.

Animals↗

Simultaneous horizontal gene transfer of a gene coding for ribosomal protein l27 and operational genes in Arthrobacter sp.

Phylogenetic analysis of bacterial L27 ribosomal proteins showed that, against taxonomy, the L27 protein from the Actinobacteria Arthrobacter sp. clusters with protein sequences from the Bacillus group. The L27 gene clusters in the Arthrobacter sp. genome with six genes responsible for creatinine and sarcosine degradation. Phylogenetic analyses of orthologue proteins encoded by three of these genes also showed a phylogenetic relationship with Bacillus species. Comparisons between the synonymous codon usage of the Arthrobacter sp. genes and those from complete genomes showed that Arthrobacter genes encoding the L27 ribosomal protein and the proteins responsible for the degradation of creatinine and sarcosine have a codon usage that is more similar to that of Bacillus species than that of Arthrobacter. We suggest that the Arthrobacter sp. genes encoding the L27 ribosomal protein and the proteins responsible for the degradation of creatinine and sarcosine were acquired simultaneously through horizontal gene transfer from an unknown Bacillus species.

Amino Acid Sequence↗

Genome variability and capsid structural constraints of hepatitis a virus.

The number of synonymous mutations per synonymous site (K(s)), the number of nonsynonymous mutations per nonsynonymous site (K(a)), and the codon usage statistic (N(c)) were calculated for several hepatitis A virus (HAV) isolates. While K(s) was similar to those of poliovirus (PV) and foot-and-mouth disease virus (FMDV), K(a) was 1 order of magnitude lower. The N(c) parameter provides information on codon usage bias and decreases when bias increases. The N(c) value in HAV was about 38, while in PV and FMDV, it was about 53. The emergence of 22 rare codons in front of 8 in PV and 7 in FMDV was detected. Most of the conserved rare codons of the P1 region were strategically located at the carboxy borders of beta barrels and alpha helices, their potential function being the assurance of proper folding of the capsid proteins through a decrease in the translation speed. This strategic location was not observed for amino acids encoded by the conserved rare codons of the 3D region. The percentage of bases with low pairing number values was higher in the latter region, suggesting a role of the conserved rare codons in the maintenance of RNA structure. Many of the rare codons in HAV are among the most frequent in humans, unlike in PV or in FMDV. This fact may be explained by the lack of cellular shutoff in HAV. One hypothesis is that HAV has evolved in order to avoid competition with its host for cellular tRNAs.

Amino Acid Sequence↗

Contrasting patterns of evolutionary divergence within the Acinetobacter calcoaceticus pca operon.

The six enzymes required for catabolism of protocatechuate to succinate and acetylCoA are encoded by the pca genes in the Gram-bacterium, Acinetobacter calcoaceticus. The clustered A. calcoaceticus cat genes encode an analogous set of enzymes associated with the metabolic dissimilation of catechol. The nucleotide (nt) sequences of pcaIJFB and pcaK, reported here, complete evidence showing that all of the pca structural genes are tightly grouped in the order pcaIJFBDKCHG within a single operon. The pcaIJF region is nearly identical in nt sequence to the A. calcoaceticus catIDJF region which exhibits a G+C content and a codon usage pattern exceptional for A. calcoaceticus. In contrast, pcaD, pcaC, pcaH and pcaG have diverged substantially from their evolutionary counterparts in the cat region; all of these divergent genes exhibit G+C contents and codon usage patterns that are typical for A. calcoaceticus. The pcaIJF and catIJF regions are known to exchange DNA sequence information, and this property may have contributed to their nt sequence conservation. The pcaK gene has no counterpart among known cat genes. The deduced amino-acid sequence of PcaK indicates that it may be a transmembrane protein associated with transport.

Acetyl Coenzyme A↗

On the informational content of overlapping genes in prokaryotic and eukaryotic viruses.

In genetic language a peculiar arrangement of biological information is provided by overlapping genes in which the same region of DNA can code for functionally unrelated messages. In this work, the informational content of overlapping genes belonging to prokaryotic and eukaryotic viruses was analyzed. Using information theory indices, we identified in the regions of overlap a first pattern, exhibiting a more uniform base composition and more severe constraints in base ordering with respect to the nonoverlapping regions. This pattern was found to be peculiar to coliphage, avian hepatitis B virus, human lentivirus, and plant luteovirus families. A second pattern, characterized by the occurrence of similar compositional constraints in both types of coding regions, was found to be limited to plant tymoviruses. At the level of codon usage, a low degree of correlation between overlapping and nonoverlapping coding regions characterized the first pattern, whereas a close link was found in tymoviruses, indicating a fine adaptation of the overlapping frame to the original codon choice of the virus. As a result of codon usage correlation analysis, deductions concerning the origin and evolution of several overlapping frames were also proposed. Comparison of amino acid composition revealed an increased frequency of amino acid residues with a high level of degeneracy (arginine, leucine, and serine) in the proteins encoded by overlapping genes; this peculiar feature of overlapping genes can be viewed as a way with which they may expand their coding ability and gain new, specialized functions.

Amino Acid Sequence↗

Expression of human lymphotoxin alpha in Aspergillus niger.

A gene-fusion expression strategy was applied for heterologous expression of human lymphotoxin alpha (LTalpha) in the Aspergillus niger AB1.13 protease-deficient strain. The LTalpha gene was fused with the A. niger glucoamylase GII-form as a carrier-gene, behind its transcription control and secretion signals. Special attention was paid to the influence of different codon usage on secretion of protein. In the case of human tumor necrosis factor alpha (TNFalpha) a dramatic change of secretion has been observed when human cDNA sequence was used instead of synthetic E. coli biased codons. In the case of LTalpha such a change of codon usage brought improvement at the RNA level, however, no increase in the quantity of secreted protein was observed, due to the proteolitic activity of the host organism. The estimated yield of secretion of LTalpha from A. niger into the soya medium was 50 pg l(-1) of culture.

Artificial Gene Fusion↗

Translation in Bacillus subtilis: roles and trends of initiation and termination, insights from a genome analysis.

We analysed the Bacillus subtilis protein coding sequences termini, and compared it to other genomes. The analysis focused on signals, com-positional biases of nucleotides, oligonucleotides, codons and amino acids and mRNA secondary structure. AUG is the preferred start codon in all genomes, independent of their G+C content, and seems to induce less stable mRNA structures. However, it is not conserved between homologous genes neither is it preferred in highly expressed genes. In B.subtilis the ribosome binding site is very strong. We found that downstream boxes do not seem to exist either in Escherichia coli or in B.subtilis. UAA stop codon usage is correlated with the G+C content and is strongly selected in highly expressed genes. We found less stable mRNA structures at both termini, which we related to mRNA-ribosome and mRNA-release-factor interactions. This pattern seems to impose a peculiar A-rich nucleotide and codon usage bias in these regions. Finally the analysis of all proteins from B.subtilis revealed a similar amino acid bias near both termini of proteins consisting of over-representation of hydrophilic residues. This bias near the stop codon is partially release-factor specific.

Algorithms↗

Characterizations of highly expressed genes of four fast-growing bacteria.

Predicted highly expressed (PHX) genes are characterized for the completely sequenced genomes of the four fast-growing bacteria Escherichia coli, Haemophilus influenzae, Vibrio cholerae, and Bacillus subtilis. Our approach to ascertaining gene expression levels relates to codon usage differences among certain gene classes: the collection of all genes (average gene), the ensemble of ribosomal protein genes, major translation/transcription processing factors, and genes for polypeptides of chaperone/degradation complexes. A gene is predicted highly expressed (PHX) if its codon frequencies are close to those of the ribosomal proteins, major translation/transcription processing factor, and chaperone/degradation standards but strongly deviant from the average gene codon frequencies. PHX genes identified by their codon usage frequencies among prokaryotic genomes commonly include those for ribosomal proteins, major transcription/translation processing factors (several occurring in multiple copies), and major chaperone/degradation proteins. Also PHX genes generally include those encoding enzymes of essential energy metabolism pathways of glycolysis, pyruvate oxidation, and respiration (aerobic and anaerobic), genes of fatty acid biosynthesis, and the principal genes of amino acid and nucleotide biosyntheses. Gene classes generally not PHX include most repair protein genes, virtually all vitamin biosynthesis genes, genes of two-component sensor systems, most regulatory genes, and most genes expressed in stationary phase or during starvation. Members of the set of PHX aminoacyl-tRNA synthetase genes contrast sharply between genomes. There are also subtle differences among the PHX energy metabolism genes between E. coli and B. subtilis, particularly with respect to genes of the tricarboxylic acid cycle. The good agreement of PHX genes of E. coli and B. subtilis with high protein abundances, as assessed by two-dimensional gel determination, is verified. Relationships of PHX genes with stoichiometry, multifunctionality, and operon structures are also examined. The spatial distribution of PHX genes within each genome reveals clusters and significantly long regions without PHX genes.

Amino Acyl-tRNA Synthetases↗

The nuclear genomes of African and American trypanosomes are strikingly different.

We have investigated the compositional distributions of exons and their different codon positions, as well as the codon usage and amino-acid (aa) composition of the nuclear genomes of the African and American trypanosomes Trypanosoma brucei and T. cruzi. Very large differences between the two species were found in all the properties investigated. The most striking differences concern the compositional distributions of third codon positions and the extremely large nucleotide divergence of third codon position for homologous genes encoding proteins that are highly conserved in their aa sequences. Moreover, if coding sequences from each species are divided into two groups according to the GC levels in third codon positions, very different codon usages and aa compositions are found. This indicates a compositional compartmentalization in both genomes which had previously been detected in T. brucei (and T. equiperdum) by compositional fractionation.

Animals↗

Evolution of chromosome bands: molecular ecology of noncoding DNA.

Giemsa dark bands, G-bands, are a derived chromatin character that evolved along the chromosomes of early chordates. They are facultative heterochromatin reflecting acquisition of a late replication mechanism to repress tissue-specific genes. Subsequently, R-bands, the primitive chromatin state, became directionally GC rich as evidenced by Q-banding of mammalian and avian chromosomes. Contrary to predictions from the neutral mutation theory, noncoding DNA is positionally constrained along the banding pattern with short interspersed repeats in R-bands and long interspersed repeats in G-bands. Chromosomes seem dynamically stable: the banding pattern and gene arrangement along several human and murine autosomes has remained constant for 100 million years, whereas much of the noncoding DNA, especially retroposons, has changed. Several coding sequence attributes and probably mutation rates are determined more by where a gene lives than by what it does. R-band exons in homeotherms but not G-band exons have directionally acquired GC-rich wobble bases and the corresponding codon usage: CpG islands in mammals are specific to R-band exons, exons not facultatively heterochromatinized, and are independent of the tissue expression pattern of the gene. The dynamic organization of noncoding DNA suggests a feedback loop that could influence codon usage and stabilize the chromosome's chromatin pattern: DNA sequences determine affinities of----proteins that together form----a chromatin that modulates----rate constants for DNA modification that determine----DNA sequences. Theories of hierarchical selection and molecular ecology show how selection can act on Darwinian units of noncoding DNA at the genome level thus creating positionally constrained DNA and contributing minimal genetic load at the individual level.

Base Sequence↗

Sequence survey of the genome of the opportunistic microsporidian pathogen, Vittaforma corneae.

The microsporidian Vittaforma corneae has been reported as a pathogen of the human stratum corneum, where it can cause keratitis, and is associated with systemic infections. In addition to this direct role as an infectious, etiologic agent of human disease, V. corneae has been used as a model organism for another microsporidian, Enterocytozoon bieneusi, a frequent and problematic pathogen of HIV-infected patients that, unlike V. corneae, is difficult to maintain and to study in vitro. Unfortunately, few molecular sequences are available for V. corneae. In this study, seventy-four genome survey sequences (GSS) were obtained from genomic DNA of spores of laboratory-cultured V. corneae. Approximately, 41 discontinuous kilobases of V. corneae were cloned and sequenced to generate these GSS. Putative identities were assigned to 44 of the V. corneae GSS based on BLASTX searches, representing 21 discrete proteins. Of these 21 deduced V. corneae proteins, only two had been reported previously from other microsporidia (until the recent report of the Encephalitozoon cuniculi genome). Two of the V. corneae proteins were of particular interest, reverse transcriptase and topoisomerase IV (parC). Since the existence of transposable elements in microsporidia is controversial, the presence of reverse transcriptase in V. corneae will contribute to resolution of this debate. The presence of topoisomerase IV was remarkable because this enzyme previously had been identified only from prokaryotes. The 74 GSS included 26.7 kilobases of unique sequences from which two statistics were generated: GC content and codon usage. The GC content of the unique GSS was 42%, lower than that of another microsporidian, E. cuniculi (48% for protein-encoding regions), and substantially higher than that predicted for a third microsporidian, Spraguea lophii (28%). A comparison using the Pearson correlation coefficient showed that codon usage in V. corneae was similar to that in the yeasts, Saccharomyces cerevisiae (r = 0.79) and Shizosaccharomyces pombe (r = 0.70), but was markedly dissimilar to E. cuniculi (r = 0.19).

Amino Acid Sequence↗

Increased expression and immunogenicity of sequence-modified human immunodeficiency virus type 1 gag gene.

A major challenge for the next generation of human immunodeficiency virus (HIV) vaccines is the induction of potent, broad, and durable cellular immune responses. The structural protein Gag is highly conserved among the HIV type 1 (HIV-1) gene products and is believed to be an important target for the host cell-mediated immune control of the virus during natural infection. Expression of Gag proteins for vaccines has been hampered by the fact that its expression is dependent on the HIV Rev protein and the Rev-responsive element, the latter located on the env transcript. Moreover, the HIV genome employs suboptimal codon usage, which further contributes to the low expression efficiency of viral proteins. In order to achieve high-level Rev-independent expression of the Gag protein, the sequences encoding HIV-1(SF2) p55(Gag) were modified extensively. First, the viral codons were changed to conform to the codon usage of highly expressed human genes, and second, the residual inhibitory sequences were removed. The resulting modified gag gene showed increases in p55(Gag) protein expression to levels that ranged from 322- to 966-fold greater than that for the native gene after transient expression of 293 cells. Additional constructs that contained the modified gag in combination with modified protease coding sequences were made, and these showed high-level Rev-independent expression of p55(Gag) and its cleavage products. Density gradient analysis and electron microscopy further demonstrated that the modified gag and gag protease genes efficiently expressed particles with the density and morphology expected for HIV virus-like particles. Mice immunized with DNA plasmids containing the modified gag showed Gag-specific antibody and CD8(+) cytotoxic T-lymphocyte (CTL) responses that were inducible at doses of input DNA 100-fold lower than those associated with plasmids containing the native gag gene. Most importantly, four of four rhesus monkeys that received two or three immunizations with modified gag plasmid DNA demonstrated substantial Gag-specific CTL responses. These results highlight the useful application of modified gag expression cassettes for increasing the potency of DNA and other gene delivery vaccine approaches against HIV.

AIDS Vaccines↗

Genetics of lactobacilli: plasmids and gene expression.

This paper reviews the present knowledge of the structure and properties of small (< 5 kb) plasmids present in Lactobacillus spp. The data show that plasmids from Lactobacillus spp., like many plasmids from other Gram-positive bacteria, display a modular organization and replicate by a mechanism of rolling circle replication. Structurally, plasmids from lactobacilli are closely related to plasmids from other Gram-positive bacteria. They contain elements (plus- and minus origin of replication, element(s) for control of plasmid replication, mobilization function) showing extensive similarity to analogous elements in plasmids from these other organisms. It is believed that lactobacilli have acquired such elements by intra- and/or intergenic transfer mechanisms. The first part of the review is concluded with a description of plasmid vectors with a Lactobacillus replicon and integrative vectors, including data concerning their structural and segregational stability. In the second part of this review we describe the progress that has been made during the last few years in identifying and characterizing elements that control expression of genetic information in lactobacilli. Based on the sequence of eleven identified and twenty presumed promoters, some preliminary conclusions can be drawn regarding the structure of Lactobacillus promoters. A typical Lactobacillus promoter shows significant similarity to promoters from E. coli and B. subtilis. An analysis of published sequences of seventy genes indicates that the region encompassing the translation start codon AUG also shows extensive similarity to that of E. coli and B. subtilis. Codon usage of Lactobacillus genes is not random and shows interspecies as well as intraspecies heterogeneity. Interspecies differences may, in part, be explained by differences in G+C content of different lactobacilli. Differences in gene expression levels can, to a large extent, account for intraspecies differences of codon usage bias. Finally, we review the knowledge that has become available concerning protein secretion and heterologous gene expression in lactobacilli. This part is concluded with a compilation of data on the expression in Lactobacillus of heterologous genes under the control of their own promoter or under control of a Lactobacillus promoter.

Bacillus subtilis↗

Evidence that mutation patterns vary among Drosophila transposable elements.

In Drosophila melanogaster, codon usage in the open reading frames (ORFs) of transposable elements (TEs) differs greatly from that in other ORFs. In addition, while the ORFs from a single element are similar, there is considerable variation among elements. In the TE ORFs there are no indications of selection for the codons prevalent in the other D. melanogaster genes, but rather codon usage can be succinctly summarized in terms of the base composition at silent sites. We suggest that the particular silent site base composition of each TE is determined by an individual pattern of mutation. In many of the TEs there is an ORF encoding a protein with homology to reverse transcriptase; the amino acid sequences of these are quite divergent, and so it is possible that each of these incorporates certain mismatched bases at different frequencies during replication.

Animals↗

Analyses of frameshifting at UUU-pyrimidine sites.

Others have recently shown that the UUU phenylalanine codon is highly frameshift-prone in the 3'(rightward) direction at pyrimidine 3'contexts. Here, several approaches are used to analyze frameshifting at such sites. The four permutations of the UUU/C (phenylalanine) and CGG/U (arginine) codon pairs were examined because they vary greatly in their expected frameshifting tendencies. Furthermore, these synonymous sites allow direct tests of the idea that codon usage can control frameshifting. Frameshifting was measured for these dicodons embedded within each of two broader contexts: the Escherichia coli prfB (RF2 gene) programmed frameshift site and a 'normal' message site. The principal difference between these contexts is that the programmed frameshift contains a purine-rich sequence upstream of the slippery site that can base pair with the 3'end of 16 S rRNA (the anti-Shine-Dalgarno) to enhance frameshifting. In both contexts frameshift frequencies are highest if the slippery tRNAPhe is capable of stable base pairing in the shifted reading frame. This requirement is less stringent in the RF2 context, as if the Shine-Dalgarno interaction can help stabilize a quasi-stable rephased tRNA:message complex. It was previously shown that frameshifting in RF2 occurs more frequently if the codon 3'to the slippery site is read by a rare tRNA. Consistent with that earlier work, in the RF2 context frameshifting occurs substantially more frequently if the arginine codon is CGG, which is read by a rare tRNA. In contrast, in the 'normal' context frameshifting is only slightly greater at CGG than at CGU. It is suggested that the Shine-Dalgarno-like interaction elevates frameshifting specifically during the pause prior to translation of the second codon, which makes frameshifting exquisitely sensitive to the rate of translation of that codon. In both contexts frameshifting increases in a mutant strain that fails to modify tRNA base A37, which is 3'of the anticodon. Thus, those base modifications may limit frameshifting at UUU codons. Finally, statistical analyses show that UUU Ynn dicodons are extremely rare in E.coli genes that have highly biased codon usage.

Arginine↗

ISSD Version 2.0: taxonomic range extended.

Two more organisms from different taxonomic groups were added to a new version of the Integrated Sequence-Structure Database (ISSD). ISSD serves as an integrated source of sequence and structure information for the analysis of correlations between mRNA synonymous codon usage and three-dimensional structure of the encoded proteins. ISSD now holds 88 non-homologous Escherichia coli proteins and 25 yeast Saccharomyces cerevisiae proteins in addition to the expanded set of mammalian proteins, which includes 166 proteins (107 in ISSD Version 1.0). Comparison of ISSD sequences with organism-specific codon usage data derived from CUTG database shows that it is a representative subset of the GenBank coding sequences data. Preliminary results of the statistical analysis confirm that sequence-structure correlations observed by us earlier are also present in the upgraded ISSD (Version 2.0), including bacterial and yeast proteins. The ISSD Version 2.0 release includes an improved Web-based data search and retrieval system and is accessible via URL http://www.protein.bio.msu.su/issd/. ISSD can be also accessed at ExPASy, URL http://www.expasy.ch/swissmod/swiss-model.htm l

Animals↗

The quality of merC, a module of the mer mosaic.

We examined a region of high variability in the mosaic mercury resistance (mer) operon of natural bacterial isolates from the primate intestinal microbiota. The region between the merP and merA genes of nine mer loci was sequenced and either the merC, the merF, or no gene was present. Two novel merC genes were identified. Overall nucleotide diversity, pi (per 100 sites), of the merC gene was greater (49.63) than adjacent merP (35.82) and merA (32.58) genes. However, the consequences of this variability for the predicted structure of the MerC protein are limited and putative functional elements (metal-binding ligands and transmembrane domains) are strongly conserved. Comparison of codon usage of the merTP, merC, and merA genes suggests that several merC genes are not coeval with their flanking sequences. Although evidence of homologous recombination within the very variable merC genes is not apparent, the flanking regions have higher homologies than merC, and recombination appears to be driving their overall sequence identities higher. The synonymous codon usage bias (EN(C)) values suggest greater variability in expression of the merC gene than in flanking genes in six different bacterial hosts. We propose a model for the evolution of MerC as a host-dependent, adventitious module of the mer operon.

Amino Acid Sequence↗