Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Structural comparison of two nontandemly repeated yeast glyceraldehyde-3-phosphate dehydrogenase genes.

A hybrid plasmid (pgap63) was isolated which contains a second yeast glyceraldehyde-3-phosphate dehydrogenase structural gene. The complete nucleotide sequence of this gene was determined and compared with the primary structure of a yeast glyceraldehyde-3-phosphate dehydrogenase gene (pgap49) which was reported previously (Holland, J.P., and Holland, M.J. (1979) J. Biol. Chem. 254, 9839-9845). Based on the restriction endonuclease cleavage maps of the isolated segments of yeast DNA which contain these genes, the two genes are nontandemly duplicated. Greater than 94% of the nucleotides within the coding regions of these genes are homologous and the polypeptides encoded by the two structural genes differ by only 15 amino acid residues. Both genes have the same, highly biased, codon usage pattern and neither contains intervening sequences. Approximately 100 nucleotides adjacent to the ATG initiation codons and 130 nucleotides beyond the TAA termination codons are greater than 70% homologous. Structures within the flanking sequences of the genes which are potentially relevant to transcriptional and translational control are described. Several sequences (8 to 15 nucleotides in length) are repeated in both the 5' and 3' flanking sequences of the genes in a noninverted fashion. Finally, a rapid procedure for the isolation of spontaneous deletions within hybrid plasmid DNAs is described, as is the isolation of a structural gene deletion in pgap49.

Amino Acid Sequence↗

Messenger RNA release from ribosomes during 5'-translational blockage by consecutive low-usage arginine but not leucine codons in Escherichia coli.

In '5'-translational blockage', significantly reduced yields of proteins are synthesized in Escherichia coli when consecutive low-usage codons are inserted near translation starts of messages (with reduced or no effect when these same codons are inserted downstream). We tested the hypothesis that ribosomes encountering these low-usage codons near the translation start prematurely release the mRNA. RNA from polysome gradients was fractionated into pools of polysomes and monosomes and a ribosome-free pool. New hybridization probes, called 'molecular beacons', and standard slot blots were used to detect test messages containing either consecutive low-usage AGG (arginine) or synonymous high-usage CGU insertions near the 5' end. The results show an approximately twofold increase in the ratio of free to bound mRNA when the low-usage codons were present in the message compared with when high-usage codons were present. In contrast, there was no difference in the ratio of free to bound mRNA when consecutive low-usage CUA or high-usage CUG (leucine) codons were inserted or when the arginine codons were inserted near the 3' end. These data indicate that at least some mRNA is released from ribosomes during 5'-translational blockage by arginine but not leucine codons, and they support proposals that premature termination of translation can occur in some conditions in vivo in the absence of a stop codon.

Arginine↗

Why are translationally sub-optimal synonymous codons used in Escherichia coli?

Natural selection favors certain synonymous codons which aid translation in Escherichia coli, yet codons not favored by translational selection persist. We use the frequency distributions of synonymous polymorphisms to test three hypotheses for the existence of translationally sub-optimal codons: (1) selection is a relatively weak force, so there is a balance between mutation, selection, and drift; (2) at some sites there is no selection on codon usage, so some synonymous sites are unaffected by translational selection; and (3) translationally sub-optimal codons are favored by alternative selection pressures at certain synonymous sites. We find that when all the data is considered, model 1 is supported and both models 2 and 3 are rejected as sole explanations for the existence of translationally sub-optimal codons. However, we find evidence in favor of both models 2 and 3 when the data is partitioned between groups of amino acids and between regions of the genes. Thus, all three mechanisms appear to contribute to the existence of translationally sub-optimal codons in E. coli.

Codon↗

Random sequence analysis of genomic DNA of a hyperthermophile: Aquifex pyrophilus.

Aquifex pyrophilus is one of the hyperthermophilic bacteria that can grow at temperatures up to 95 degrees C. To obtain information about its genomic structure, random sequencing was performed on plasmid libraries containing 0.5-2 kb genomic DNA fragments of A. pyrophilus. Comparison of the obtained sequence tags with known proteins revealed that 123 tags showed strong similarity to previously identified proteins in the PIR or Genebank databases. These included three proteases, two amino acid racemases, and three enzymes utilizing oxygen as substrate. Although the GC ratio of the genome is about 40%, the codon usage of A. pyrophilus showed biased occurrence of G and C at the third position of codons, especially those for amino acids such as asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, lysine, and tyrosine. A higher ratio of positively charged amino acids in A. pyrophilus proteins as compared with proteins from mesophiles suggested that Aquifex proteins might contain increased ion-pair interaction that could help to maintain heat stability.

Amino Acid Sequence↗

Pseudogene in the genome of bacteriophage lambda?

We find a region in the non-coding part of bacteriophage lambda genome that codes for the conserved fold which repressors and other proteins use for specific DNA binding. The region is involved in a long open reading frame exceeding one kilobase and is read in the same frame as gene A in the opposite strand. The putative translation product of this open reading frame has a highly ordered secondary structure with a predominance of alpha helices, which is typical of repressors. In addition, codon usage in this frame suggests a protein-coding region. However, there is a TGA stop codon located between the putative gene start point and the region coding for the DNA binding fold. It thus appears that bacteriophage lambda had one more DNA binding protein, perhaps repressor, in the past that was inactivated by a mutation.

Amino Acid Sequence↗

Modification of GP63 genes from diverse species of Leishmania for expression of recombinant protein at high levels in Escherichia coli.

Toward the future development of a defined subunit vaccine against leishmaniasis is, high levels of recombinant GP63 for diverse species of Leishmania were produced in Escherichia coli. Several features of Leishmania GP63 genes were simultaneously modified with the polymerase chain reaction (PCR) using either cloned genes or total genomic DNA from Leishmania as template DNA for the PCR amplification reactions. The PCR products included only the coding region for the predicted mature form of GP63 that occurs on the surface of Leishmania, flanked by the appropriate translation signals and cloning sites for the production of recombinant GP63 as nonfusion protein in E. coli. When the codon usage in the GP63 gene was modified to reduce the guanine and cytosine content for the codons adjacent to the ATG initiation codon, rGP63 represented about 50% of total protein in E. coli. Mouse monoclonal antibodies raised against purified Leishmania major rGP63 had equivalent immunoblotting characteristics for native GP63 and recombinant GP63 with respect to linear determinants on GP63 expressed in diverse species of Leishmania. Human T cell lines and clones were derived from a patient infected with Leishmania braziliensis panamensis using rGP63 purified from an L. major GP63 expression clone as antigen.

Animals↗

Characterisation of the dihydrofolate reductase-thymidylate synthetase gene from human malaria parasites highly resistant to pyrimethamine.

To investigate the genetic basis of drug resistance in human malaria parasites, we have sequenced the entire dihydrofolate reductase thymidylate synthetase DHFR-TS bifunctional gene from the highly pyrimethamine-resistant K1 isolate of Plasmodium falciparum. The protein is predicted to consist of 607 amino acids (aa), (71,685 Da), with an N-terminal methionine encoded by the second start codon of the open reading frame. Compared to the sequence from drug-sensitive parasites, there are two nucleotide changes in the coding region which bring about a substitution of Arg for Cys at aa position 59 and Asn for Thr at aa position 108. Both changes occur in regions of the DHFR domain involved in inhibitor and cofactor binding and are hence strongly implicated in drug resistance. The gene is present as a single copy in both K1 and drug-sensitive FCR3 isolates, and is assigned to chromosome 4. Codon usage follows the pattern observed in that of malarial surface antigen genes, with the exception fo codons corresponding to Val and Pro. The Asn and Lys contents of the predicted protein are exceptionally high, these residues being particularly concentrated in the DHFR and junction domains.

Amino Acid Sequence↗

Delineation of coding areas in DNA sequences through assignment of codon probabilities.

Codon usage tables have been produced for E. coli, yeast, human, and mouse. The nonrandom employment of codons allows assignment of probability values to trinucleotides in any DNA sequence. These values represent the probability that a given trinucleotide is used as a codon in the organism from which the table is derived. For the graphical delineation of coding areas in DNA sequences, a probability is assigned to each trinucleotide equal to its frequency in the codon table. Averaging and smoothing procedures then greatly enhance the detectability of areas of high average codon probability and better represent the mean codon probability. These manipulations increase graphical clarity without altering the overall magnitude of probabilities. Averaging introduces an error of less than 0.5% between "raw" and smoothed data. This graphical delineation of coding sequences does not depend on the presence of punctuation, ribosomal binding sites, etc: moreover the delineation of introns and exons is also possible.

Base Sequence↗

Comparison of a vitellogenin gene between two distantly related rhabditid nematode species.

Three vitellogenin genes from the free-living nematode Caenorhabditis elegans have previously been characterized at the molecular level. In order to study evolutionary relationships within this poorly understood taxon, we have cloned a vitellogenin gene, CEW1-vit-6, from a distantly related species belonging to the same family as C. elegans. Screening of a genomic library with a probe to total poly(A+) RNA yielded three clones that hybridized more intensely than all others, and all three corresponded to a single gene homologous to C. elegans vit-6. Comparison of CEW1-vit-6 with Ce-vit-6 reveals both strong similarities and surprising differences. Life Ce-vit-6, the gene is about 5 kb long and contains four unusually small introns (38-41 nt), but only one interrupts the gene at the same location as a Ce-vit-6 intron. The promoter region contains five matches to Vitellogenin Promoter Element 1 (VPE1) and no matches to VPE2, both previously shown to be required for vit gene transcription in C. elegans. Codon usage is in general similar to that of the Ce-vit genes, but a few codon biases are quite different. Alignment of the CEW1-vit-6 protein with Ce-vit-6 and Ce-vit-2 products suggests the existence of two domains which have evolved at different rates. Sequence comparison shows that nematode vitellogenins are much more closely related to vertebrate than to insect vitellogenins.

Amino Acid Sequence↗

DNA and protein sequence homologies between the adhesins of Mycoplasma genitalium and Mycoplasma pneumoniae.

Mycoplasma genitalium and Mycoplasma pneumoniae are morphologically and serologically related pathogens that colonize the human host. Their successful parasitism appears to be dependent on the product, an adhesin protein, of a gene that is carried by each of these mycoplasmas. Here we describe the cloning and determine the sequence of the structural gene for the putative adhesin of M. genitalium and compare its sequence to the counterpart P1 gene of M. pneumoniae. Regions of homology that were consistent with the observed serological cross-reactivity between these adhesins were detected at both DNA and protein levels. However, the degree of homology between these two genes and their products was much higher than anticipated. Interestingly, the A + T content of the M. genitalium adhesin gene was calculated as 60.1%, which is substantially higher tham that of the P1 gene (46.5%). Comparisons of codon usage between the two organisms revealed that M. genitalium preferentially used A- and T-rich codons. A total of 65% of positions 3 and 56% of positions 1 in M. genitalium codons were either A or T, whereas M. pneumoniae utilized A or T for positions 3 and 1 at a frequency of 40 and 47%, respectively. The biased choice of the A- and T-rich codons in M. genitalium could also account for the preferential use of A- and T-rich codons in conservative amino acid substitutions found in the M. genitalium adhesin. These facts suggest that M. genitalium might have evolved independently of other human mycoplasma species, including M. pneumoniae.

Amino Acid Sequence↗

Characterization of a molluscum contagiosum virus homolog of the vaccinia virus p37K major envelope antigen.

We present the first nucleotide sequence data for molluscum contagiosum virus (MCV), an unclassified poxvirus. A 2,276-bp XhoI fragment from a near left-terminal fragment of MCV subtype I (MCVI) and a 1,920-bp XhoI fragment from the corresponding locus of MCV subtype II (MCVII) were sequenced and analyzed for open reading frames (ORFs). A large, complete ORF of 1,167 bp was present in both fragments. The putative polypeptide has a calculated molecular mass of 43 kDa (p43K protein) and was shown to have a high degree of homology to the vaccinia virus p37K major envelope antigen (40% amino acid identity and 22% conservative changes). The nucleotide content of the MCV fragments sequenced was 66% G or C. The codon usage within the gene for p43K reflected this high G + C content, with position 3 of codons being predominantly G or C (82 and 87% for MCVI and MCVII, respectively). The MCV p43K-encoding gene has motifs immediately upstream which are similar to those required for vaccinia virus late gene expression. The location and direction of transcription of the MCV p43K-encoding gene were equivalent to those of the vaccinia virus p37K gene, revealing similarity in genetic organization between MCV and vaccinia virus. Another, incomplete ORF was identified downstream of the p43K-encoding gene in both MCVI and MCVII. The sequence immediately upstream of this ORF overlapped the termination codon of the p43K-encoding gene and contained a motif which had homology to the derived consensus sequence for vaccinia virus early gene promoters.

Amino Acid Sequence↗

A synthetic E7 gene of human papillomavirus type 16 that yields enhanced expression of the protein in mammalian cells and is useful for DNA immunization studies.

A synthetic E7 gene of human papillomavirus (HPV) type 16 was generated that consists entirely of preferred human codons. Expression analysis of the synthetic E7 gene in human and animal cells showed levels of E7 protein 20- to 100-fold higher than those obtained with wild-type E7. Enhanced expression of E7 protein resulted from highly efficient translation, as well as increased stability of the E7 mRNA due to its codon optimization. Higher levels of E7 protein in cells transfected with synthetic E7 correlated with significant loss of cell viability in various human cell lines. In contrast, lower E7 protein expression driven by the wild-type gene resulted in a slight induction of cell proliferation. Furthermore, mice inoculated with plasmids expressing the synthetic E7 gene produced significantly higher levels of E7 antibodies than littermates injected with wild-type E7, suggesting that synthetic E7 may be useful for DNA immunization studies and the development of genetic vaccines against HPV-16. In view of these results, we hypothesize that HPVs may have retained a pattern of G + C content and codon usage distinct from that of their host cells in response to selective pressure. Thus, the nonhuman codon bias may have been conserved by HPVs to prevent compromising viability of the host cells by excessive viral early protein expression, as well as to evade the immune system.

Amino Acid Sequence↗

Coding in the noncoding DNA strand: A novel mechanism of gene evolution?

The question whether the noncoding DNA strand had or still has the capability for encoding functional polypeptides has been addressed in several articles. The theoretical background of the views advocating this idea arose from two groups of findings. One of them was based on various observations implying that the genetic code was adapted for double-strand coding. The other group of theories arose from the observation of gene-length overlapping open reading frames (O-ORFs) on the antisense DNA strand in a number of genes. In fact, the above theories, which I term selectionist, conceive a novel conception of gene evolution, proposing that new genes can be created by the utilization of antisense DNA strand. In contrast, neutralist theory claims that the O-ORFs are mere by-products of evolutionary processes acting to create special codon usage and base distribution patterns in the coding sequences.

Codon↗

Complete nucleotide sequences of the domestic cat (Felis catus) mitochondrial genome and a transposed mtDNA tandem repeat (Numt) in the nuclear genome.

The complete 17,009-bp mitochondrial genome of the domestic cat, Felis catus, has been sequenced and conforms largely to the typical organization of previously characterized mammalian mtDNAs. Codon usage and base composition also followed canonical vertebrate patterns, except for an unusual ATC (non-AUG) codon initiating the NADH dehydrogenase subunit 2 (ND2) gene. Two distinct repetitive motifs at opposite ends of the control region contribute to the relatively large size (1559 bp) of this carnivore mtDNA. Alignment of the feline mtDNA genome to a homologous 7946-bp nuclear mtDNA tandem repeat DNA sequence in the cat, Numt, indicates simple repeat motifs associated with insertion/deletion mutations. Overall DNA sequence divergence between Numt and cytoplasmic mtDNA sequence was only 5.1%. Substitutions predominate at the third codon position of homologous feline protein genes. Phylogenetic analysis of mitochondrial gene sequences confirms the recent transfer of the cytoplasmic mtDNA sequences to the domestic cat nucleus and recapitulates evolutionary relationships between mammal species.

Amino Acid Sequence↗

Genetic code and phylogenetic origin of oomycetous mitochondria.

We sequenced the 3'-terminal part of the COX3 gene encoding cytochrome c oxidase subunit 3 from mitochondria of Phytophthora parasitica (phylum Oomycota, kingdom Protoctista). Comparison of the sequence with known COX3 genes revealed that UGG is used as a tryptophan codon in contrast to UGA in the mitochondrial codes of most organisms other than green plants. A very high AT mutation pressure operates on the mitochondrial genome of Phytophthora, as revealed by codon usage and by A+T content of noncoding regions, which seems paradoxical because AT pressure causes tryptophan codon reassignment from UGG to UGA in mitochondria of most species. The genetic code and other data suggest that mitochondria of Oomycota share a direct common ancestor with mitochondria of plants and that mitochondria of the ancestor of Planta and Oomycota were acquired in a second endosymbiotic event, which occurred later than the acquisition of mitochondria by other eukaryotes.

Amino Acid Sequence↗

Characterization of nucleotidic sequences using maximum entropy techniques.

A statistical method for characterizing nucleotidic sequences based on maximum entropy techniques is presented. The method uses only codon usage tables and takes into account the length of sequences, and preserves the information contained in each codon by a punctual index. We present the methodological aspects of the analysis, showing an application relative to nucleotidic sequences of eukaryotes.

Animals↗

Codon bias and frequency-dependent selection on the hemagglutinin epitopes of influenza A virus.

Although the surface proteins of human influenza A virus evolve rapidly and continually produce antigenic variants, the internal viral genes acquire mutations very gradually. In this paper, we analyze the sequence evolution of three influenza A genes over the past two decades. We study codon usage as a discriminating signature of gene- and even residue-specific diversifying and purifying selection. Nonrandom codon choice can increase or decrease the effective local substitution rate. We demonstrate that the codons of hemagglutinin, particularly those in the antibody-combining regions, are significantly biased toward substitutional point mutations relative to the codons of other influenza virus genes. We discuss the evolutionary interpretation and implications of these biases for hemagglutinin's antigenic evolution. We also introduce information-theoretic methods that use sequence data to detect regions of recent positive selection and potential protein conformational changes.

Codon↗

Characterization of the virE operon of the Agrobacterium Ti plasmid pTiA6.

The Agrobacterium tumefaciens Ti plasmid contains at least six transcriptional units (designated vir loci) which are essential for efficient crown gall tumorigenesis. Mutations in one of these loci, virE, result in a sharply attenuated virulence phenotype. In the present communication, we have analyzed the virE operon at the molecular level. This locus contains open reading frames coding for two hydrophilic proteins having molecular weights of approximately 7,000 daltons and 60,500 daltons. Using a maxicell strain of E. coli, we have visualized two proteins encoded by virE which correspond in size to these open reading frames. Analysis of codon usage of virE and seven other vir loci indicates that, in contrast to E. coli, all possible codons for a given amino acid are utilized at approximately the same frequency.

Amino Acid Sequence↗