Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

DNA thermodynamic pressure: a potential contributor to genome evolution.

Codon usage bias is a feature of living organisms. The origin of this bias might be explained not only by external factors but also by the nature of the structure of deoxyribonucleic acid (DNA) itself. We have developed a point mutation simulation program of coding sequences, in which nucleotide replacement follows thermodynamic criteria. For this purpose we calculated the hydrogen bond-like and electrostatic energies of non-canonical base pairs in a 5 bp neighbourhood. Although the rate of non-canonical base pair formation is extremely low, such pairs occur with a preference towards a guanine (G) or cytosine (C) rather than an adenine (A) or thymine (T) replacement due to thermodynamic considerations. This feature, according to the simulation program, should result in an increase in the GC content of the genome over evolutionary time. In addition, codon bias towards a higher GC usage is also predicted. DNA sequence analysis of genes of the Trypanosomatidae lineage supported the hypothesis that DNA thermodynamic pressure is a driving force that impels increases in GC content and GC codon bias.

Algorithms↗

The genome and genes of Neurospora crassa.

Neurospora crassa is an organism with a 7-decade contribution to genetic research. in a genome of 42.9 Mb and just over 1000 map units, to date over 800 different genes have been identified by phenotype and/or map location, and 222 genes have been characterized by sequencing. Methods by which analysis of the genome has been carried out are discussed, including linkage, RFLP, and chromosome walking. Characterized centomeres, telomeres, the nucleolar organizer and the dispersed 5S rRNA genes are discussed. Analysis of the protein-encoding genes is undertaken, using new software for the querying of standard sequence databases. Gene analysis includes consensus sequences for transcription and RNA splicing and new insights into codon usage.

Chromosome Mapping↗

A comparison of the nucleotide sequences of eastern and western equine encephalomyelitis viruses with those of other alphaviruses and related RNA viruses.

The complete nucleotide sequence of a 1982 Florida strain of eastern equine encephalomyelitis (EEE) virus, and partial sequence of the nonstructural protein genes of western equine encephalomyelitis (WEE) virus, were determined. The EEE virus genome was 11,678 nucleotides in length, excluding the cap nucleotide and poly(A) tail, and the nucleotide composition was 28% A, 24% G, 25% C, and 23% U. The organization of both EEE and WEE virus genomes was like that of other alphaviruses and included a termination codon between the nsP3 and nsP4 genes. Codon usage for 10 of 20 amino acids was nonrandom in the EEE genome, and dinucleotide CpG-containing codons were underutilized in both genomes. The slight CpG deficiency was similar to that seen in other alphaviruses and plant viruses in the alphavirus-like group, but less than that of poliovirus and yellow fever virus. This slight deficiency may reflect adaptation for replication in both CpG-deficient vertebrates, as well as insects which do not have CpG-deficient genomes. Phylogenetic analyses using nonstructural protein amino acid sequences indicated that alphaviruses evolved from a common ancestor which existed a few thousand years ago. An intercontinental introduction of an ancestral virus from the Old to New World, or vice versa, probably resulted in two main extant groups: one includes New World (EEE and Venezuelan equine encephalitis) viruses, while the other includes Old World (Sindbis, Middelburg, O'nyong-nyong, Ross River, and Semliki Forest) viruses. The position of WEE virus in the phylogenetic trees indicated that, in addition to its capsid gene (C. S. Hahn et al. (1988) Proc. Natl. Acad. Sci. USA 85, 5997-6001), WEE virus acquired its nonstructural genes from an EEE-like ancestor during recombination.

Alphavirus↗

Codon modified human papillomavirus type 16 E7 DNA vaccine enhances cytotoxic T-lymphocyte induction and anti-tumour activity.

Polynucleotide immunisation with the E7 gene of human papillomavirus (HPV) type 16 induces only moderate levels of immune response, which may in part be due to limitation in E7 gene expression influenced by biased HPV codon usage. Here we compare for expression and immunogenicity polynucleotide expression plasmids encoding wild-type (pWE7) or synthetic codon optimised (pHE7) HPV16 E7 DNA. Cos-1 cells transfected with pHE7 expressed higher levels of E7 protein than similar cells transfected with pW7. C57BL/6 mice and F1 (C57x FVB) E7 transgenic mice immunised intradermally with E7 plasmids produced high levels of anti-E7 antibody. pHE7 induced a significantly stronger E7-specific cytotoxic T-lymphocyte response than pWE7 and 100% tumour protection in C57BL/6 mice, but neither vaccine induced CTL in partially E7 tolerant K14E7 transgenic mice. The data indicate that immunogenicity of an E7 polynucleotide vaccine can be enhanced by codon modification. However, this may be insufficient for priming E7 responses in animals with split tolerance to E7 as a consequence of expression of E7 in somatic cells.

Animals↗

Modulation of base-specific mutation and recombination rates enables functional adaptation within the context of the genetic code.

The persistence of life requires populations to adapt at a rate commensurate with the dynamics of their environment. Successful populations that inhabit highly variable environments have evolved mechanisms to increase the likelihood of successful adaptation. We introduce a 64 x 64 matrix to quantify base-specific mutation potential, analyzing four different replicative systems, error-prone PCR, mouse antibodies, a nematode, and Drosophila. Mutational tendencies are correlated with the structural evolution of proteins. In systems under strong selective pressure, mutational biases are shown to favor the adaptive search of space, either by base mutation or by recombination. Such adaptability is discussed within the context of the genetic code at the levels of replication and codon usage.

Adaptation, Biological↗

Mutation and selection on the anticodon of tRNA genes in vertebrate mitochondrial genomes.

The H-strand of vertebrate mitochondrial DNA is left single-stranded for hours during the slow DNA replication. This facilitates C-->U mutations on the H-strand (and consequently G-->A mutations on the L-strand) via spontaneous deamination which occurs much more frequently on single-stranded than on double-stranded DNA. For the 12 coding sequences (CDS) collinear with the L-strand, NNY synonymous codon families (where N stands for any of the four nucleotides and Y stands for either C or U) end mostly with C, and NNR and NNN codon families (where R stands for either A or G) end mostly with A. For the lone ND6 gene on the other strand, the codon bias is the opposite, with NNY codon families ending mostly with U and NNR and NNN codon families ending mostly with G. These patterns are consistent with the strand-specific mutation bias. The codon usage biased towards C-ending and A-ending in the 12 CDS sequences affects the codon-anticodon adaptation. The wobble site of the anticodon is always G for NNY codon families dominated by C-ending codons and U for NNR and NNN codon families dominated by A-ending codons. The only, but consistent, exception is the anticodon of tRNA-Met which consistently has a 5'-CAU-3' anticodon base-pairing with the AUG codon (the translation initiation codon) instead of the more frequent AUA. The observed CAU anticodon (matching AUG) would increase the rate of translation initiation but would reduce the rate of peptide elongation because most methionine codons are AUA, whereas the unobserved UAU anticodon (matching AUA) would increase the elongation rate at the cost of translation initiation rate. The consistent CAU anticodon in tRNA-Met suggests the importance of maximizing the rate of translation initiation.

Animals↗

Mutually symmetric and complementary triplets: differences in their use distinguish systematically between coding and non-coding genomic sequences.

The general property of asymmetry in word use in meaningful texts written in a variety of languages, motivates a quantification of the differences in the use of mutually symmetric triplets in genomic sequences. When this is done in the three reading frames, high values found for one of them are used as indication that the sequence is coding for a protein. Moreover, a similar quantification of the differences in the use of complementary triplets is introduced, again with predictive power of the coding character of a sequence. This method reflects the non-equivalence between sense and anti-sense strand of a coding segment. In both approaches, "linguistic asymmetry" in coding sequences is related to the form of the genetic code and to the bias in codon usage and amino acid use skews.

Algorithms↗

Codon discrimination due to presence of abundant non-cognate competitive tRNA.

It has been thought that preferential use of synonymous codons provides high efficiency and fidelity of protein synthesis through specific codon-anticodon interactions. In yeast genes, some codon boxes seem to prefer a codon which is unsuited for its cognate anticodon. Now, we propose that codon usage biases may arise due to presence of abundant non-cognate competitive tRNA capable of misreading a codon by C-U or G-U pairing in the middle position.

Codon↗

Cloning and nucleotide sequence of the aspartase gene of Pseudomonas fluorescens.

The aspartase gene (aspA) of Pseudomonas fluorescens was cloned and the nucleotide sequence of the 2,066-base-pair DNA fragment containing the aspA gene was determined. The amino acid sequence of the protein deduced from the nucleotide sequence was confirmed by N- and C-terminal sequence analysis of the purified enzyme protein. The deduced amino acid composition also fitted the previous amino acid analysis results well (Takagi et al. (1984) J. Biochem. 96, 545-552). These results indicate that aspartase of P. fluorescens consists of four identical subunits with a molecular weight of 50,859, composed of 472 amino acid residues. The coding sequence of the gene was preceded by a potential Shine-Dalgarno sequence and by a few promoter-like structures. Following the stop codon there was a structure which is reminiscent of the Escherichia coli rho-independent terminator. The G + C content of the coding sequence was found to be 62.3%. Inspection of the codon usage for the aspA gene revealed as high as 80.0% preference for G or C at the third codon position. The deduced amino acid sequence was 56.3% homologous with that of the enzyme of E. coli W (Takagi et al. (1985) Nucl. Acids Res. 13, 2063-2074). Cys-140 and Cys-430 of the E. coli enzyme, which had been assigned as functionally essential (Ida & Tokushige (1985) J. Biochem. 98, 793-797), were substituted by Ala-140 and Ala-431, respectively, in the P. fluorescens enzyme.

Amino Acid Sequence↗

Production of the Gram-positive Sarcina ventriculi pyruvate decarboxylase in Escherichia coli.

Sarcina ventriculi grows in a remarkable range of mesophilic environments from pH 2 to pH 10. During growth in acidic environments, where acetate is toxic, expression of pyruvate decarboxylase (PDC) serves to direct the flow of pyruvate into ethanol during fermentation. PDC is rare in bacteria and absent in animals, although it is widely distributed in the plant kingdom. The pdc gene from S. ventriculi is the first to be cloned and characterized from a Gram-positive bacterium. In Escherichia coli, the recombinant pdc gene from S. ventriculi was poorly expressed due to differences in codon usage that are typical of low-G+C organisms. Expression was improved by the addition of supplemental codon genes and this facilitated the 136-fold purification of the recombinant enzyme as a homo-tetramer of 58 kDa subunits. Unlike Zymomonas mobilis PDC, which exhibits Michaelis-Menten kinetics, S. ventriculi PDC is activated by pyruvate and exhibits sigmoidal kinetics similar to fungal and higher plant PDCs. Amino acid residues involved in the allosteric site for pyruvate in fungal PDCs were conserved in S. ventriculi PDC, consistent with a conservation of mechanism. Cluster analysis of deduced amino acid sequences confirmed that S. ventriculi PDC is quite distant from Z. mobilis PDC and plant PDCs. S. ventriculi PDC appears to have diverged very early from a common ancestor which included most fungal PDCs and eubacterial indole-3-pyruvate decarboxylases. These results suggest that the S. ventriculi pdc gene is quite ancient in origin, in contrast to the Z. mobilis pdc, which may have originated by horizontal transfer from higher plants.

Amino Acid Sequence↗

High level expression of a recombinant acid phytase gene in Pichia pastoris.

AIMS: To achieve high phytase yield with improved enzymatic activity in Pichia pastoris. METHODS AND RESULTS: The 1347-bp phytase gene of Aspergillus niger SK-57 was synthesized using a successive polymerase chain reaction and was altered by deleting intronic sequences, optimizing codon usage and replacing its original signal sequence with a synthetic signal peptide (designated MF4I) that is a codon-modified Saccharomyces cerevisiae mating factor alpha-prepro-leader sequence. The gene constructs containing wild type or modified phytase gene coding sequences under the control of the highly-inducible alcohol oxidase gene promoter with the MF4I- or wild type alpha-signal sequence were used to transform Pichia pastoris. The P. pastoris strain that expressed the modified phytase gene (phyA-sh) with MF4I sequence produced 6.1 g purified phytase per litre of culture fluid, with the phytase activity of 865 U ml(-1). The expressed phytase varied in size (64, 67, 87, 110 and 120 kDa), but could be deglycosylated to produce a homogeneous 64 kDa protein. The recombinant phytase had two pH optima (pH 2.5 and pH 5.5) and an optimum temperature of 60 degrees C. CONCLUSIONS: The P. pastoris strain with the genetically engineered phytase gene produced 6.1 g l(-1) of phytase or 865 U ml(-1) phytase activity, a 14.5-fold increase compared with the P. pastoris strain with the wild type phytase gene. SIGNIFICANCE AND IMPACT OF THE STUDY: The P. pastoris strain expressing the modified phytase gene with the MF4I signal peptide showed great potential as a commercial phytase production system.

6-Phytase↗

Molecular characterization of a 40-kDa outer membrane protein, FomA, of Fusobacterium periodonticum and comparison with Fusobacterium nucleatum.

The 40 kDa-outer membrane protein FomA of Fusobacterium periodonticum ATCC 33693 was found to exhibit heat modifiable properties, typical for a porin, and N-terminal sequencing indicated a close relationship to the porin FomA of Fusobacterium nucleatum. A polymerase chain reaction approach was therefore applied for sequencing the fomA gene of F. periodonticum, and nucleotide and deduced amino acid sequences were aligned and compared with the corresponding sequences of different strains of F. nucleatum. In all strains we found a common protein upstream of the fomA gene. The noncoding area upstream of the putative -35 region of the F. periodonticum fomA gene exhibited little sequence similarity with the F. nucleatum gene. The transcriptional unit of FomA, on the other hand, was very similar, with the similarities concentrated in domains that were interspersed with hypervariable regions. A topology model was made and compared with those made for F. nucleatum. This indicated that the great similarities reside in the membrane-spanning segments of the protein, while most cell surface exposed loops were hypervariable. The results strongly support the proposed model for FomA and also indicate that these taxa are related but on a lower level than the subspecies level. The codon usage of F. periodonticum is comparable to that of F. nucleatum, and the triplet AGA is the only codon used for arginine.

Amino Acid Sequence↗

Sequence analysis of the DNA encoding the Eco RI endonuclease and methylase.

The Eco RI endonuclease and methylase recognize the same hexanucleotide substrate sequence. We have determined the sequence of a fragment of DNA which encodes these enzymes using the chain-termination method of Sanger (Sanger, F., Nicklen, S., and Coulson, A. R. (1977) Proc. Natl. Acad. Sci. U. S. A. 74, 5463-5467). The amino acid sequences of both enzymes were derived from the DNA sequence. The coding regions selected include the only open translational frames of sufficient length to accommodate the enzymes. They coincide with previously established gene boundaries and orientation. The predicted amino acid sequences correlate well with analyses of the purified protein. Comparison of the nucleotide and protein sequences reveals no homology between the endonuclease and methylase which might provide insight into the origin of the restriction-modification system or the mechanism of common substrate recognition. Based on secondary structure predictions, the two enzymes also have grossly different molecular architecture. The base composition of the sequence is 65% A + T, and the codon usage is significantly different from that observed in several Escherichia coli chromosomal genes. In some cases, frequently selected codons are recognized by minor tRNA species. A spontaneous mutation in the endonuclease gene was isolated. Serine replaces arginine at residue 187. In crude extracts, Eco RI specific cleavage is approximately 0.3% wild type.

Amino Acid Sequence↗

Investigation on the causes of codon and amino acid usages variation between thermophilic Aquifex aeolicus and mesophilic Bacillus subtilis.

Base composition, codon usages and amino acid usages have been analyzed by taking 529 orthologous sequences of Aquifex aeolicus and Bacillus subtilis, having different optimal growth temperatures. These two bacteria do not have significant difference in overall GC composition, but GC(1+2) and GC3 levels were found to vary significantly. Significant increments in purine content and GC3 composition have been observed in the coding sequences of Aquifex aeolicus than its Bacillus subtilis counterparts. Correspondence analyses on codon and amino acid usages reveal that variation in base composition actually influences their codon and amino acid usages. Two selection pressures acting on the nucleotide level (GC3 and purine enrichment), causes variation in the amino acid usage differently in different protein secondary structures. Our results suggest that adaptation of amino acid usages in coil structure of Aquifex aeolicus proteins is under the control of both purine increment and GC3 composition, whereas the adaptation of the amino acids in the helical region of thermophilic bacteria is strongly influenced by the purine content. Evolutionary perspectives concerning the temperature adaptation of DNA and protein molecules of these two bacteria have been discussed on the basis of these results.

Amino Acids↗

Third position codon composition suggests two classes of genes within the Cauliflower mosaic virus genome.

The translation of viral mRNAs by host ribosomes is essential for infection. Hence, codon usage of virus genes may influence efficiency of infection. In addition, composition of nucleotides in the third position within codons of genes can reflect evolutionary relationships. In this study, third position codon composition was examined for the seven genes of eight Cauliflower mosaic virus isolates. Genes IV-VII had similar codon composition values and were termed Class 1 genes. Genes I-III possessed corresponding codon composition values and were termed Class 2 genes. The codon composition values of Class 1 and genes differed significantly. Neither Class 1 nor Class 2 genes had codon composition values identical to that of the host plant, Arabidopsis thaliana. However, Class 1 genes possessed codon composition values closer to those of the host than Class 2 genes. Examination of the genomes of three Rous sarcoma virus isolates indicated that codon composition values were similar for the gag, pol, and env genes but these genes differed significantly from the src genes. Since codon composition values for Rous sarcoma virus distinguished a "foreign" gene from the rest of the viral genome, it is possible that the Cauliflower mosaic virus genome is composed of genes from two different sources. Others have suggested that Cauliflower mosaic virus evolved in this manner and our data provide support for this hypothesis.

Arabidopsis↗

Strand-specific nucleotide composition bias in echinoderm and vertebrate mitochondrial genomes.

The gene organization of starfish mitochondrial DNA is identical with that of the sea urchin counterpart except for a reported inversion of an approximately 4.6-kb segment containing two structural genes for NADH dehydrogenase subunits 1 and 2 (ND 1 and ND 2). When the codon usage of each structural gene in starfish, sea urchin, and vertebrate mitochondrial DNAs is examined, it is striking that codons ending in T and G are preferentially used more for heavy strand-encoded genes, including starfish ND 1 and ND 2, than for light strand-encoded genes, including sea urchin ND 1 and ND 2. On the contrary, codons ending in A and C are preferentially used for the light strand-encoded genes rather than for the heavy strand-encoded ones. Moreover, G-U base pairs are more frequently found in the possible secondary structures of heavy strand-encoded tRNAs than in those of light strand-encoded tRNAs. These observations suggest the existence of a certain constraint operating on mitochondrial genomes from various animal phyla, which results in the accumulation of G and T on one strand, and A and C on the other.

Amino Acid Sequence↗

Cloning and characterization of xylanase A from the strain Bacillus sp. BP-7: comparison with alkaline pI-low molecular weight xylanases of family 11.

The xynA gene encoding a xylanase from the recently isolated Bacillus sp. strain BP-7 has been cloned and expressed in Escherichia coli. Recombinant xylanase A showed high activity on xylans from hardwoods and cereals, and exhibited maximum activity at pH 6 and 60 degrees C. The enzyme remained stable after incubation at 50 degrees C and pH 7 for 3 h, and it was strongly inhibited by Mn(2+), Fe(3+), Pb(2+), and Hg(2+). Analysis of xylanase A in zymograms showed an apparent molecular size of 24 kDa and a pI of above 9. The amino acid sequence of xylanase A, as deduced from xynA gene, shows homology to alkaline pI-low molecular weight xylanases of family 11 such as XynA from Bacillus subtilis. Analysis of codon usage in xynA from Bacillus sp. BP-7 shows that the G+C content at the first and second codon positions is notably different from the mean values found for glycosyl hydrolase genes from Bacillus subtilis.

Bacillus↗

The nucleotide sequence of the cloned tufA gene of Escherichia coli.

The 4 kb (8.5 % lambda units) EcoRI fragment harboring the tufA gene of Escherichia coli was cloned using plasmid pTUA1 (Shibuya et al., 1979) and its structure was analyzed. The nucleotide sequence of about 1500 base pairs, covering the C-terminal portion of elongation factor EF-G (fus gene), the intercistronic region between fus and tufA, the entire structural gene for tufA with the GUG initiation and UAA termination codons, and the 3' flanking region of tufA, was determined. Comparison of the tufA nucleotide sequence with the tufB sequence (An and Friesen, 1980) and the known amino acid sequence of EF-Tu (Arai et al., 1980) revealed that the products of genes tufA and tufB are identical except for one amino acid at the C-terminal, i.e., glycine for tufA and serine for tufB. Nucleotide differences between tufA and tufB were found at 13 positions. Among them, one in the initiation codon and the other one in the C-terminal amino acid codon had replacements at the first letter of the codons. The other eleven changes were in the third codon positions, which did not affect the amino acid coding. The pattern of codon usage in tufA and tufB is highly nonrandom, and remarkably similar to that in ribosomal protein genes, with the codons for the most abundant species of isoaccepting tRNAs being preferentially utilized (Post et al., 1979; Post and Nomura, 1980).

Bacterial Proteins↗