Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Characterization of nucleotidic sequences using maximum entropy techniques.

A statistical method for characterizing nucleotidic sequences based on maximum entropy techniques is presented. The method uses only codon usage tables and takes into account the length of sequences, and preserves the information contained in each codon by a punctual index. We present the methodological aspects of the analysis, showing an application relative to nucleotidic sequences of eukaryotes.

Animals↗

Codon bias and frequency-dependent selection on the hemagglutinin epitopes of influenza A virus.

Although the surface proteins of human influenza A virus evolve rapidly and continually produce antigenic variants, the internal viral genes acquire mutations very gradually. In this paper, we analyze the sequence evolution of three influenza A genes over the past two decades. We study codon usage as a discriminating signature of gene- and even residue-specific diversifying and purifying selection. Nonrandom codon choice can increase or decrease the effective local substitution rate. We demonstrate that the codons of hemagglutinin, particularly those in the antibody-combining regions, are significantly biased toward substitutional point mutations relative to the codons of other influenza virus genes. We discuss the evolutionary interpretation and implications of these biases for hemagglutinin's antigenic evolution. We also introduce information-theoretic methods that use sequence data to detect regions of recent positive selection and potential protein conformational changes.

Codon↗

Characterization of the virE operon of the Agrobacterium Ti plasmid pTiA6.

The Agrobacterium tumefaciens Ti plasmid contains at least six transcriptional units (designated vir loci) which are essential for efficient crown gall tumorigenesis. Mutations in one of these loci, virE, result in a sharply attenuated virulence phenotype. In the present communication, we have analyzed the virE operon at the molecular level. This locus contains open reading frames coding for two hydrophilic proteins having molecular weights of approximately 7,000 daltons and 60,500 daltons. Using a maxicell strain of E. coli, we have visualized two proteins encoded by virE which correspond in size to these open reading frames. Analysis of codon usage of virE and seven other vir loci indicates that, in contrast to E. coli, all possible codons for a given amino acid are utilized at approximately the same frequency.

Amino Acid Sequence↗

Codon reading scheme in Mycoplasma pneumoniae revealed by the analysis of the complete set of tRNA genes.

The 33 genes encoding the complete set of tRNA species in Mycoplasma pneumoniae have been cloned and sequenced. They are organized into 5 clusters in addition to 9 single genes. No redundant gene was found, indicating that 33 tRNAs correspond to 32 different anticodons and decode all 62 codons used in this organism. There is only one single tRNA for each of the Ala, Leu, Pro, and Val family boxes. Therefore, a simplified decoding system resembling that recently described for Mycoplasma capricolum (1) has to also exist in M.pneumoniae. However, analysis of the anticodon set and codon usage revealed features characteristic of the latter: (i) there is no obvious preference toward AT rich synonymous codons, (ii) CGG codons are assigned for arginine and are translated by tRNA Arg(UCG), and (iii) CNN or GNN anticodons are encountered in the Ser, Thr, Arg, and Gly family boxes. We thus propose that this codon-anticodon recognition pattern has emerged in the 'M.pneumoniae cluster' under a genomic economization strategy but without the influence of AT pressure.

Amino Acid Sequence↗

Analysis and comparison of nucleotide sequences encoding the genes for [NiFe] and [NiFeSe] hydrogenases from Desulfovibrio gigas and Desulfovibrio baculatus.

The nucleotide sequences encoding the [NiFe] hydrogenase from Desulfovibrio gigas and the [NiFeSe] hydrogenase from Desulfovibrio baculatus (N.K. Menon, H.D. Peck, Jr., J. LeGall, and A.E. Przybyla, J. Bacteriol. 169:5401-5407, 1987; C. Li, H.D. Peck, Jr., J. LeGall, and A.E. Przybyla, DNA 6:539-551, 1987) were analyzed by the codon usage method of Staden and McLachlan. The reported reading frames were found to contain regions of low codon probability which are matched by more probable sequences in other frames. Renewed nucleotide sequencing showed the probable frames to be correct. The corrected sequences of the two small and large subunits share a significant degree of sequence homology. The small subunit, which contains 10 conserved cysteine residues, is likely to coordinate at least 2 iron-sulfur clusters, while the finding of a selenocysteine codon (TGA) near the 3' end of the [NiFeSe] large-subunit gene matched by a regular cysteine codon (TGC) in the [NiFe] large-subunit gene indicates the presence of some of the ligands to the active-site nickel in the large subunit.

Amino Acid Sequence↗

DNA thermodynamic pressure: a potential contributor to genome evolution.

Codon usage bias is a feature of living organisms. The origin of this bias might be explained not only by external factors but also by the nature of the structure of deoxyribonucleic acid (DNA) itself. We have developed a point mutation simulation program of coding sequences, in which nucleotide replacement follows thermodynamic criteria. For this purpose we calculated the hydrogen bond-like and electrostatic energies of non-canonical base pairs in a 5 bp neighbourhood. Although the rate of non-canonical base pair formation is extremely low, such pairs occur with a preference towards a guanine (G) or cytosine (C) rather than an adenine (A) or thymine (T) replacement due to thermodynamic considerations. This feature, according to the simulation program, should result in an increase in the GC content of the genome over evolutionary time. In addition, codon bias towards a higher GC usage is also predicted. DNA sequence analysis of genes of the Trypanosomatidae lineage supported the hypothesis that DNA thermodynamic pressure is a driving force that impels increases in GC content and GC codon bias.

Algorithms↗

The genome and genes of Neurospora crassa.

Neurospora crassa is an organism with a 7-decade contribution to genetic research. in a genome of 42.9 Mb and just over 1000 map units, to date over 800 different genes have been identified by phenotype and/or map location, and 222 genes have been characterized by sequencing. Methods by which analysis of the genome has been carried out are discussed, including linkage, RFLP, and chromosome walking. Characterized centomeres, telomeres, the nucleolar organizer and the dispersed 5S rRNA genes are discussed. Analysis of the protein-encoding genes is undertaken, using new software for the querying of standard sequence databases. Gene analysis includes consensus sequences for transcription and RNA splicing and new insights into codon usage.

Chromosome Mapping↗

A comparison of the nucleotide sequences of eastern and western equine encephalomyelitis viruses with those of other alphaviruses and related RNA viruses.

The complete nucleotide sequence of a 1982 Florida strain of eastern equine encephalomyelitis (EEE) virus, and partial sequence of the nonstructural protein genes of western equine encephalomyelitis (WEE) virus, were determined. The EEE virus genome was 11,678 nucleotides in length, excluding the cap nucleotide and poly(A) tail, and the nucleotide composition was 28% A, 24% G, 25% C, and 23% U. The organization of both EEE and WEE virus genomes was like that of other alphaviruses and included a termination codon between the nsP3 and nsP4 genes. Codon usage for 10 of 20 amino acids was nonrandom in the EEE genome, and dinucleotide CpG-containing codons were underutilized in both genomes. The slight CpG deficiency was similar to that seen in other alphaviruses and plant viruses in the alphavirus-like group, but less than that of poliovirus and yellow fever virus. This slight deficiency may reflect adaptation for replication in both CpG-deficient vertebrates, as well as insects which do not have CpG-deficient genomes. Phylogenetic analyses using nonstructural protein amino acid sequences indicated that alphaviruses evolved from a common ancestor which existed a few thousand years ago. An intercontinental introduction of an ancestral virus from the Old to New World, or vice versa, probably resulted in two main extant groups: one includes New World (EEE and Venezuelan equine encephalitis) viruses, while the other includes Old World (Sindbis, Middelburg, O'nyong-nyong, Ross River, and Semliki Forest) viruses. The position of WEE virus in the phylogenetic trees indicated that, in addition to its capsid gene (C. S. Hahn et al. (1988) Proc. Natl. Acad. Sci. USA 85, 5997-6001), WEE virus acquired its nonstructural genes from an EEE-like ancestor during recombination.

Alphavirus↗

Codon modified human papillomavirus type 16 E7 DNA vaccine enhances cytotoxic T-lymphocyte induction and anti-tumour activity.

Polynucleotide immunisation with the E7 gene of human papillomavirus (HPV) type 16 induces only moderate levels of immune response, which may in part be due to limitation in E7 gene expression influenced by biased HPV codon usage. Here we compare for expression and immunogenicity polynucleotide expression plasmids encoding wild-type (pWE7) or synthetic codon optimised (pHE7) HPV16 E7 DNA. Cos-1 cells transfected with pHE7 expressed higher levels of E7 protein than similar cells transfected with pW7. C57BL/6 mice and F1 (C57x FVB) E7 transgenic mice immunised intradermally with E7 plasmids produced high levels of anti-E7 antibody. pHE7 induced a significantly stronger E7-specific cytotoxic T-lymphocyte response than pWE7 and 100% tumour protection in C57BL/6 mice, but neither vaccine induced CTL in partially E7 tolerant K14E7 transgenic mice. The data indicate that immunogenicity of an E7 polynucleotide vaccine can be enhanced by codon modification. However, this may be insufficient for priming E7 responses in animals with split tolerance to E7 as a consequence of expression of E7 in somatic cells.

Animals↗

Mutually symmetric and complementary triplets: differences in their use distinguish systematically between coding and non-coding genomic sequences.

The general property of asymmetry in word use in meaningful texts written in a variety of languages, motivates a quantification of the differences in the use of mutually symmetric triplets in genomic sequences. When this is done in the three reading frames, high values found for one of them are used as indication that the sequence is coding for a protein. Moreover, a similar quantification of the differences in the use of complementary triplets is introduced, again with predictive power of the coding character of a sequence. This method reflects the non-equivalence between sense and anti-sense strand of a coding segment. In both approaches, "linguistic asymmetry" in coding sequences is related to the form of the genetic code and to the bias in codon usage and amino acid use skews.

Algorithms↗

Codon discrimination due to presence of abundant non-cognate competitive tRNA.

It has been thought that preferential use of synonymous codons provides high efficiency and fidelity of protein synthesis through specific codon-anticodon interactions. In yeast genes, some codon boxes seem to prefer a codon which is unsuited for its cognate anticodon. Now, we propose that codon usage biases may arise due to presence of abundant non-cognate competitive tRNA capable of misreading a codon by C-U or G-U pairing in the middle position.

Codon↗

Cloning and nucleotide sequence of the aspartase gene of Pseudomonas fluorescens.

The aspartase gene (aspA) of Pseudomonas fluorescens was cloned and the nucleotide sequence of the 2,066-base-pair DNA fragment containing the aspA gene was determined. The amino acid sequence of the protein deduced from the nucleotide sequence was confirmed by N- and C-terminal sequence analysis of the purified enzyme protein. The deduced amino acid composition also fitted the previous amino acid analysis results well (Takagi et al. (1984) J. Biochem. 96, 545-552). These results indicate that aspartase of P. fluorescens consists of four identical subunits with a molecular weight of 50,859, composed of 472 amino acid residues. The coding sequence of the gene was preceded by a potential Shine-Dalgarno sequence and by a few promoter-like structures. Following the stop codon there was a structure which is reminiscent of the Escherichia coli rho-independent terminator. The G + C content of the coding sequence was found to be 62.3%. Inspection of the codon usage for the aspA gene revealed as high as 80.0% preference for G or C at the third codon position. The deduced amino acid sequence was 56.3% homologous with that of the enzyme of E. coli W (Takagi et al. (1985) Nucl. Acids Res. 13, 2063-2074). Cys-140 and Cys-430 of the E. coli enzyme, which had been assigned as functionally essential (Ida & Tokushige (1985) J. Biochem. 98, 793-797), were substituted by Ala-140 and Ala-431, respectively, in the P. fluorescens enzyme.

Amino Acid Sequence↗

Production of the Gram-positive Sarcina ventriculi pyruvate decarboxylase in Escherichia coli.

Sarcina ventriculi grows in a remarkable range of mesophilic environments from pH 2 to pH 10. During growth in acidic environments, where acetate is toxic, expression of pyruvate decarboxylase (PDC) serves to direct the flow of pyruvate into ethanol during fermentation. PDC is rare in bacteria and absent in animals, although it is widely distributed in the plant kingdom. The pdc gene from S. ventriculi is the first to be cloned and characterized from a Gram-positive bacterium. In Escherichia coli, the recombinant pdc gene from S. ventriculi was poorly expressed due to differences in codon usage that are typical of low-G+C organisms. Expression was improved by the addition of supplemental codon genes and this facilitated the 136-fold purification of the recombinant enzyme as a homo-tetramer of 58 kDa subunits. Unlike Zymomonas mobilis PDC, which exhibits Michaelis-Menten kinetics, S. ventriculi PDC is activated by pyruvate and exhibits sigmoidal kinetics similar to fungal and higher plant PDCs. Amino acid residues involved in the allosteric site for pyruvate in fungal PDCs were conserved in S. ventriculi PDC, consistent with a conservation of mechanism. Cluster analysis of deduced amino acid sequences confirmed that S. ventriculi PDC is quite distant from Z. mobilis PDC and plant PDCs. S. ventriculi PDC appears to have diverged very early from a common ancestor which included most fungal PDCs and eubacterial indole-3-pyruvate decarboxylases. These results suggest that the S. ventriculi pdc gene is quite ancient in origin, in contrast to the Z. mobilis pdc, which may have originated by horizontal transfer from higher plants.

Amino Acid Sequence↗

Molecular characterization of a 40-kDa outer membrane protein, FomA, of Fusobacterium periodonticum and comparison with Fusobacterium nucleatum.

The 40 kDa-outer membrane protein FomA of Fusobacterium periodonticum ATCC 33693 was found to exhibit heat modifiable properties, typical for a porin, and N-terminal sequencing indicated a close relationship to the porin FomA of Fusobacterium nucleatum. A polymerase chain reaction approach was therefore applied for sequencing the fomA gene of F. periodonticum, and nucleotide and deduced amino acid sequences were aligned and compared with the corresponding sequences of different strains of F. nucleatum. In all strains we found a common protein upstream of the fomA gene. The noncoding area upstream of the putative -35 region of the F. periodonticum fomA gene exhibited little sequence similarity with the F. nucleatum gene. The transcriptional unit of FomA, on the other hand, was very similar, with the similarities concentrated in domains that were interspersed with hypervariable regions. A topology model was made and compared with those made for F. nucleatum. This indicated that the great similarities reside in the membrane-spanning segments of the protein, while most cell surface exposed loops were hypervariable. The results strongly support the proposed model for FomA and also indicate that these taxa are related but on a lower level than the subspecies level. The codon usage of F. periodonticum is comparable to that of F. nucleatum, and the triplet AGA is the only codon used for arginine.

Amino Acid Sequence↗

Sequence analysis of the DNA encoding the Eco RI endonuclease and methylase.

The Eco RI endonuclease and methylase recognize the same hexanucleotide substrate sequence. We have determined the sequence of a fragment of DNA which encodes these enzymes using the chain-termination method of Sanger (Sanger, F., Nicklen, S., and Coulson, A. R. (1977) Proc. Natl. Acad. Sci. U. S. A. 74, 5463-5467). The amino acid sequences of both enzymes were derived from the DNA sequence. The coding regions selected include the only open translational frames of sufficient length to accommodate the enzymes. They coincide with previously established gene boundaries and orientation. The predicted amino acid sequences correlate well with analyses of the purified protein. Comparison of the nucleotide and protein sequences reveals no homology between the endonuclease and methylase which might provide insight into the origin of the restriction-modification system or the mechanism of common substrate recognition. Based on secondary structure predictions, the two enzymes also have grossly different molecular architecture. The base composition of the sequence is 65% A + T, and the codon usage is significantly different from that observed in several Escherichia coli chromosomal genes. In some cases, frequently selected codons are recognized by minor tRNA species. A spontaneous mutation in the endonuclease gene was isolated. Serine replaces arginine at residue 187. In crude extracts, Eco RI specific cleavage is approximately 0.3% wild type.

Amino Acid Sequence↗

Third position codon composition suggests two classes of genes within the Cauliflower mosaic virus genome.

The translation of viral mRNAs by host ribosomes is essential for infection. Hence, codon usage of virus genes may influence efficiency of infection. In addition, composition of nucleotides in the third position within codons of genes can reflect evolutionary relationships. In this study, third position codon composition was examined for the seven genes of eight Cauliflower mosaic virus isolates. Genes IV-VII had similar codon composition values and were termed Class 1 genes. Genes I-III possessed corresponding codon composition values and were termed Class 2 genes. The codon composition values of Class 1 and genes differed significantly. Neither Class 1 nor Class 2 genes had codon composition values identical to that of the host plant, Arabidopsis thaliana. However, Class 1 genes possessed codon composition values closer to those of the host than Class 2 genes. Examination of the genomes of three Rous sarcoma virus isolates indicated that codon composition values were similar for the gag, pol, and env genes but these genes differed significantly from the src genes. Since codon composition values for Rous sarcoma virus distinguished a "foreign" gene from the rest of the viral genome, it is possible that the Cauliflower mosaic virus genome is composed of genes from two different sources. Others have suggested that Cauliflower mosaic virus evolved in this manner and our data provide support for this hypothesis.

Arabidopsis↗

Strand-specific nucleotide composition bias in echinoderm and vertebrate mitochondrial genomes.

The gene organization of starfish mitochondrial DNA is identical with that of the sea urchin counterpart except for a reported inversion of an approximately 4.6-kb segment containing two structural genes for NADH dehydrogenase subunits 1 and 2 (ND 1 and ND 2). When the codon usage of each structural gene in starfish, sea urchin, and vertebrate mitochondrial DNAs is examined, it is striking that codons ending in T and G are preferentially used more for heavy strand-encoded genes, including starfish ND 1 and ND 2, than for light strand-encoded genes, including sea urchin ND 1 and ND 2. On the contrary, codons ending in A and C are preferentially used for the light strand-encoded genes rather than for the heavy strand-encoded ones. Moreover, G-U base pairs are more frequently found in the possible secondary structures of heavy strand-encoded tRNAs than in those of light strand-encoded tRNAs. These observations suggest the existence of a certain constraint operating on mitochondrial genomes from various animal phyla, which results in the accumulation of G and T on one strand, and A and C on the other.

Amino Acid Sequence↗

The nucleotide sequence of the cloned tufA gene of Escherichia coli.

The 4 kb (8.5 % lambda units) EcoRI fragment harboring the tufA gene of Escherichia coli was cloned using plasmid pTUA1 (Shibuya et al., 1979) and its structure was analyzed. The nucleotide sequence of about 1500 base pairs, covering the C-terminal portion of elongation factor EF-G (fus gene), the intercistronic region between fus and tufA, the entire structural gene for tufA with the GUG initiation and UAA termination codons, and the 3' flanking region of tufA, was determined. Comparison of the tufA nucleotide sequence with the tufB sequence (An and Friesen, 1980) and the known amino acid sequence of EF-Tu (Arai et al., 1980) revealed that the products of genes tufA and tufB are identical except for one amino acid at the C-terminal, i.e., glycine for tufA and serine for tufB. Nucleotide differences between tufA and tufB were found at 13 positions. Among them, one in the initiation codon and the other one in the C-terminal amino acid codon had replacements at the first letter of the codons. The other eleven changes were in the third codon positions, which did not affect the amino acid coding. The pattern of codon usage in tufA and tufB is highly nonrandom, and remarkably similar to that in ribosomal protein genes, with the codons for the most abundant species of isoaccepting tRNAs being preferentially utilized (Post et al., 1979; Post and Nomura, 1980).

Bacterial Proteins↗