Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Coding in the noncoding DNA strand: A novel mechanism of gene evolution?

The question whether the noncoding DNA strand had or still has the capability for encoding functional polypeptides has been addressed in several articles. The theoretical background of the views advocating this idea arose from two groups of findings. One of them was based on various observations implying that the genetic code was adapted for double-strand coding. The other group of theories arose from the observation of gene-length overlapping open reading frames (O-ORFs) on the antisense DNA strand in a number of genes. In fact, the above theories, which I term selectionist, conceive a novel conception of gene evolution, proposing that new genes can be created by the utilization of antisense DNA strand. In contrast, neutralist theory claims that the O-ORFs are mere by-products of evolutionary processes acting to create special codon usage and base distribution patterns in the coding sequences.

Codon↗

Complete nucleotide sequences of the domestic cat (Felis catus) mitochondrial genome and a transposed mtDNA tandem repeat (Numt) in the nuclear genome.

The complete 17,009-bp mitochondrial genome of the domestic cat, Felis catus, has been sequenced and conforms largely to the typical organization of previously characterized mammalian mtDNAs. Codon usage and base composition also followed canonical vertebrate patterns, except for an unusual ATC (non-AUG) codon initiating the NADH dehydrogenase subunit 2 (ND2) gene. Two distinct repetitive motifs at opposite ends of the control region contribute to the relatively large size (1559 bp) of this carnivore mtDNA. Alignment of the feline mtDNA genome to a homologous 7946-bp nuclear mtDNA tandem repeat DNA sequence in the cat, Numt, indicates simple repeat motifs associated with insertion/deletion mutations. Overall DNA sequence divergence between Numt and cytoplasmic mtDNA sequence was only 5.1%. Substitutions predominate at the third codon position of homologous feline protein genes. Phylogenetic analysis of mitochondrial gene sequences confirms the recent transfer of the cytoplasmic mtDNA sequences to the domestic cat nucleus and recapitulates evolutionary relationships between mammal species.

Amino Acid Sequence↗

Genetic code and phylogenetic origin of oomycetous mitochondria.

We sequenced the 3'-terminal part of the COX3 gene encoding cytochrome c oxidase subunit 3 from mitochondria of Phytophthora parasitica (phylum Oomycota, kingdom Protoctista). Comparison of the sequence with known COX3 genes revealed that UGG is used as a tryptophan codon in contrast to UGA in the mitochondrial codes of most organisms other than green plants. A very high AT mutation pressure operates on the mitochondrial genome of Phytophthora, as revealed by codon usage and by A+T content of noncoding regions, which seems paradoxical because AT pressure causes tryptophan codon reassignment from UGG to UGA in mitochondria of most species. The genetic code and other data suggest that mitochondria of Oomycota share a direct common ancestor with mitochondria of plants and that mitochondria of the ancestor of Planta and Oomycota were acquired in a second endosymbiotic event, which occurred later than the acquisition of mitochondria by other eukaryotes.

Amino Acid Sequence↗

Characterization of nucleotidic sequences using maximum entropy techniques.

A statistical method for characterizing nucleotidic sequences based on maximum entropy techniques is presented. The method uses only codon usage tables and takes into account the length of sequences, and preserves the information contained in each codon by a punctual index. We present the methodological aspects of the analysis, showing an application relative to nucleotidic sequences of eukaryotes.

Animals↗

Codon bias and frequency-dependent selection on the hemagglutinin epitopes of influenza A virus.

Although the surface proteins of human influenza A virus evolve rapidly and continually produce antigenic variants, the internal viral genes acquire mutations very gradually. In this paper, we analyze the sequence evolution of three influenza A genes over the past two decades. We study codon usage as a discriminating signature of gene- and even residue-specific diversifying and purifying selection. Nonrandom codon choice can increase or decrease the effective local substitution rate. We demonstrate that the codons of hemagglutinin, particularly those in the antibody-combining regions, are significantly biased toward substitutional point mutations relative to the codons of other influenza virus genes. We discuss the evolutionary interpretation and implications of these biases for hemagglutinin's antigenic evolution. We also introduce information-theoretic methods that use sequence data to detect regions of recent positive selection and potential protein conformational changes.

Codon↗

Compositional biases and polyalanine runs in humans.

Human proteins containing polyalanine tracts tend to have runs of other amino acids and their open reading frames (ORFs) display a biased codon usage. Their alanine, glycine, proline, and histidine content strongly correlates with the GC content of the third codon base, suggesting that the compositional specificity of these proteins is dictated to a great extent by the evolution of their ORFs.

Codon↗

Characterization of the virE operon of the Agrobacterium Ti plasmid pTiA6.

The Agrobacterium tumefaciens Ti plasmid contains at least six transcriptional units (designated vir loci) which are essential for efficient crown gall tumorigenesis. Mutations in one of these loci, virE, result in a sharply attenuated virulence phenotype. In the present communication, we have analyzed the virE operon at the molecular level. This locus contains open reading frames coding for two hydrophilic proteins having molecular weights of approximately 7,000 daltons and 60,500 daltons. Using a maxicell strain of E. coli, we have visualized two proteins encoded by virE which correspond in size to these open reading frames. Analysis of codon usage of virE and seven other vir loci indicates that, in contrast to E. coli, all possible codons for a given amino acid are utilized at approximately the same frequency.

Amino Acid Sequence↗

Codon reading scheme in Mycoplasma pneumoniae revealed by the analysis of the complete set of tRNA genes.

The 33 genes encoding the complete set of tRNA species in Mycoplasma pneumoniae have been cloned and sequenced. They are organized into 5 clusters in addition to 9 single genes. No redundant gene was found, indicating that 33 tRNAs correspond to 32 different anticodons and decode all 62 codons used in this organism. There is only one single tRNA for each of the Ala, Leu, Pro, and Val family boxes. Therefore, a simplified decoding system resembling that recently described for Mycoplasma capricolum (1) has to also exist in M.pneumoniae. However, analysis of the anticodon set and codon usage revealed features characteristic of the latter: (i) there is no obvious preference toward AT rich synonymous codons, (ii) CGG codons are assigned for arginine and are translated by tRNA Arg(UCG), and (iii) CNN or GNN anticodons are encountered in the Ser, Thr, Arg, and Gly family boxes. We thus propose that this codon-anticodon recognition pattern has emerged in the 'M.pneumoniae cluster' under a genomic economization strategy but without the influence of AT pressure.

Amino Acid Sequence↗

Analysis and comparison of nucleotide sequences encoding the genes for [NiFe] and [NiFeSe] hydrogenases from Desulfovibrio gigas and Desulfovibrio baculatus.

The nucleotide sequences encoding the [NiFe] hydrogenase from Desulfovibrio gigas and the [NiFeSe] hydrogenase from Desulfovibrio baculatus (N.K. Menon, H.D. Peck, Jr., J. LeGall, and A.E. Przybyla, J. Bacteriol. 169:5401-5407, 1987; C. Li, H.D. Peck, Jr., J. LeGall, and A.E. Przybyla, DNA 6:539-551, 1987) were analyzed by the codon usage method of Staden and McLachlan. The reported reading frames were found to contain regions of low codon probability which are matched by more probable sequences in other frames. Renewed nucleotide sequencing showed the probable frames to be correct. The corrected sequences of the two small and large subunits share a significant degree of sequence homology. The small subunit, which contains 10 conserved cysteine residues, is likely to coordinate at least 2 iron-sulfur clusters, while the finding of a selenocysteine codon (TGA) near the 3' end of the [NiFeSe] large-subunit gene matched by a regular cysteine codon (TGC) in the [NiFe] large-subunit gene indicates the presence of some of the ligands to the active-site nickel in the large subunit.

Amino Acid Sequence↗

DNA thermodynamic pressure: a potential contributor to genome evolution.

Codon usage bias is a feature of living organisms. The origin of this bias might be explained not only by external factors but also by the nature of the structure of deoxyribonucleic acid (DNA) itself. We have developed a point mutation simulation program of coding sequences, in which nucleotide replacement follows thermodynamic criteria. For this purpose we calculated the hydrogen bond-like and electrostatic energies of non-canonical base pairs in a 5 bp neighbourhood. Although the rate of non-canonical base pair formation is extremely low, such pairs occur with a preference towards a guanine (G) or cytosine (C) rather than an adenine (A) or thymine (T) replacement due to thermodynamic considerations. This feature, according to the simulation program, should result in an increase in the GC content of the genome over evolutionary time. In addition, codon bias towards a higher GC usage is also predicted. DNA sequence analysis of genes of the Trypanosomatidae lineage supported the hypothesis that DNA thermodynamic pressure is a driving force that impels increases in GC content and GC codon bias.

Algorithms↗

The genome and genes of Neurospora crassa.

Neurospora crassa is an organism with a 7-decade contribution to genetic research. in a genome of 42.9 Mb and just over 1000 map units, to date over 800 different genes have been identified by phenotype and/or map location, and 222 genes have been characterized by sequencing. Methods by which analysis of the genome has been carried out are discussed, including linkage, RFLP, and chromosome walking. Characterized centomeres, telomeres, the nucleolar organizer and the dispersed 5S rRNA genes are discussed. Analysis of the protein-encoding genes is undertaken, using new software for the querying of standard sequence databases. Gene analysis includes consensus sequences for transcription and RNA splicing and new insights into codon usage.

Chromosome Mapping↗

A comparison of the nucleotide sequences of eastern and western equine encephalomyelitis viruses with those of other alphaviruses and related RNA viruses.

The complete nucleotide sequence of a 1982 Florida strain of eastern equine encephalomyelitis (EEE) virus, and partial sequence of the nonstructural protein genes of western equine encephalomyelitis (WEE) virus, were determined. The EEE virus genome was 11,678 nucleotides in length, excluding the cap nucleotide and poly(A) tail, and the nucleotide composition was 28% A, 24% G, 25% C, and 23% U. The organization of both EEE and WEE virus genomes was like that of other alphaviruses and included a termination codon between the nsP3 and nsP4 genes. Codon usage for 10 of 20 amino acids was nonrandom in the EEE genome, and dinucleotide CpG-containing codons were underutilized in both genomes. The slight CpG deficiency was similar to that seen in other alphaviruses and plant viruses in the alphavirus-like group, but less than that of poliovirus and yellow fever virus. This slight deficiency may reflect adaptation for replication in both CpG-deficient vertebrates, as well as insects which do not have CpG-deficient genomes. Phylogenetic analyses using nonstructural protein amino acid sequences indicated that alphaviruses evolved from a common ancestor which existed a few thousand years ago. An intercontinental introduction of an ancestral virus from the Old to New World, or vice versa, probably resulted in two main extant groups: one includes New World (EEE and Venezuelan equine encephalitis) viruses, while the other includes Old World (Sindbis, Middelburg, O'nyong-nyong, Ross River, and Semliki Forest) viruses. The position of WEE virus in the phylogenetic trees indicated that, in addition to its capsid gene (C. S. Hahn et al. (1988) Proc. Natl. Acad. Sci. USA 85, 5997-6001), WEE virus acquired its nonstructural genes from an EEE-like ancestor during recombination.

Alphavirus↗

Codon modified human papillomavirus type 16 E7 DNA vaccine enhances cytotoxic T-lymphocyte induction and anti-tumour activity.

Polynucleotide immunisation with the E7 gene of human papillomavirus (HPV) type 16 induces only moderate levels of immune response, which may in part be due to limitation in E7 gene expression influenced by biased HPV codon usage. Here we compare for expression and immunogenicity polynucleotide expression plasmids encoding wild-type (pWE7) or synthetic codon optimised (pHE7) HPV16 E7 DNA. Cos-1 cells transfected with pHE7 expressed higher levels of E7 protein than similar cells transfected with pW7. C57BL/6 mice and F1 (C57x FVB) E7 transgenic mice immunised intradermally with E7 plasmids produced high levels of anti-E7 antibody. pHE7 induced a significantly stronger E7-specific cytotoxic T-lymphocyte response than pWE7 and 100% tumour protection in C57BL/6 mice, but neither vaccine induced CTL in partially E7 tolerant K14E7 transgenic mice. The data indicate that immunogenicity of an E7 polynucleotide vaccine can be enhanced by codon modification. However, this may be insufficient for priming E7 responses in animals with split tolerance to E7 as a consequence of expression of E7 in somatic cells.

Animals↗

Mutually symmetric and complementary triplets: differences in their use distinguish systematically between coding and non-coding genomic sequences.

The general property of asymmetry in word use in meaningful texts written in a variety of languages, motivates a quantification of the differences in the use of mutually symmetric triplets in genomic sequences. When this is done in the three reading frames, high values found for one of them are used as indication that the sequence is coding for a protein. Moreover, a similar quantification of the differences in the use of complementary triplets is introduced, again with predictive power of the coding character of a sequence. This method reflects the non-equivalence between sense and anti-sense strand of a coding segment. In both approaches, "linguistic asymmetry" in coding sequences is related to the form of the genetic code and to the bias in codon usage and amino acid use skews.

Algorithms↗

Codon discrimination due to presence of abundant non-cognate competitive tRNA.

It has been thought that preferential use of synonymous codons provides high efficiency and fidelity of protein synthesis through specific codon-anticodon interactions. In yeast genes, some codon boxes seem to prefer a codon which is unsuited for its cognate anticodon. Now, we propose that codon usage biases may arise due to presence of abundant non-cognate competitive tRNA capable of misreading a codon by C-U or G-U pairing in the middle position.

Codon↗

Cloning and nucleotide sequence of the aspartase gene of Pseudomonas fluorescens.

The aspartase gene (aspA) of Pseudomonas fluorescens was cloned and the nucleotide sequence of the 2,066-base-pair DNA fragment containing the aspA gene was determined. The amino acid sequence of the protein deduced from the nucleotide sequence was confirmed by N- and C-terminal sequence analysis of the purified enzyme protein. The deduced amino acid composition also fitted the previous amino acid analysis results well (Takagi et al. (1984) J. Biochem. 96, 545-552). These results indicate that aspartase of P. fluorescens consists of four identical subunits with a molecular weight of 50,859, composed of 472 amino acid residues. The coding sequence of the gene was preceded by a potential Shine-Dalgarno sequence and by a few promoter-like structures. Following the stop codon there was a structure which is reminiscent of the Escherichia coli rho-independent terminator. The G + C content of the coding sequence was found to be 62.3%. Inspection of the codon usage for the aspA gene revealed as high as 80.0% preference for G or C at the third codon position. The deduced amino acid sequence was 56.3% homologous with that of the enzyme of E. coli W (Takagi et al. (1985) Nucl. Acids Res. 13, 2063-2074). Cys-140 and Cys-430 of the E. coli enzyme, which had been assigned as functionally essential (Ida & Tokushige (1985) J. Biochem. 98, 793-797), were substituted by Ala-140 and Ala-431, respectively, in the P. fluorescens enzyme.

Amino Acid Sequence↗

Production of the Gram-positive Sarcina ventriculi pyruvate decarboxylase in Escherichia coli.

Sarcina ventriculi grows in a remarkable range of mesophilic environments from pH 2 to pH 10. During growth in acidic environments, where acetate is toxic, expression of pyruvate decarboxylase (PDC) serves to direct the flow of pyruvate into ethanol during fermentation. PDC is rare in bacteria and absent in animals, although it is widely distributed in the plant kingdom. The pdc gene from S. ventriculi is the first to be cloned and characterized from a Gram-positive bacterium. In Escherichia coli, the recombinant pdc gene from S. ventriculi was poorly expressed due to differences in codon usage that are typical of low-G+C organisms. Expression was improved by the addition of supplemental codon genes and this facilitated the 136-fold purification of the recombinant enzyme as a homo-tetramer of 58 kDa subunits. Unlike Zymomonas mobilis PDC, which exhibits Michaelis-Menten kinetics, S. ventriculi PDC is activated by pyruvate and exhibits sigmoidal kinetics similar to fungal and higher plant PDCs. Amino acid residues involved in the allosteric site for pyruvate in fungal PDCs were conserved in S. ventriculi PDC, consistent with a conservation of mechanism. Cluster analysis of deduced amino acid sequences confirmed that S. ventriculi PDC is quite distant from Z. mobilis PDC and plant PDCs. S. ventriculi PDC appears to have diverged very early from a common ancestor which included most fungal PDCs and eubacterial indole-3-pyruvate decarboxylases. These results suggest that the S. ventriculi pdc gene is quite ancient in origin, in contrast to the Z. mobilis pdc, which may have originated by horizontal transfer from higher plants.

Amino Acid Sequence↗

Molecular characterization of a 40-kDa outer membrane protein, FomA, of Fusobacterium periodonticum and comparison with Fusobacterium nucleatum.

The 40 kDa-outer membrane protein FomA of Fusobacterium periodonticum ATCC 33693 was found to exhibit heat modifiable properties, typical for a porin, and N-terminal sequencing indicated a close relationship to the porin FomA of Fusobacterium nucleatum. A polymerase chain reaction approach was therefore applied for sequencing the fomA gene of F. periodonticum, and nucleotide and deduced amino acid sequences were aligned and compared with the corresponding sequences of different strains of F. nucleatum. In all strains we found a common protein upstream of the fomA gene. The noncoding area upstream of the putative -35 region of the F. periodonticum fomA gene exhibited little sequence similarity with the F. nucleatum gene. The transcriptional unit of FomA, on the other hand, was very similar, with the similarities concentrated in domains that were interspersed with hypervariable regions. A topology model was made and compared with those made for F. nucleatum. This indicated that the great similarities reside in the membrane-spanning segments of the protein, while most cell surface exposed loops were hypervariable. The results strongly support the proposed model for FomA and also indicate that these taxa are related but on a lower level than the subspecies level. The codon usage of F. periodonticum is comparable to that of F. nucleatum, and the triplet AGA is the only codon used for arginine.

Amino Acid Sequence↗