Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Cloning and sequencing of a Moraxella bovis pilin gene.

Moraxella bovis pili have been shown to play a major role in both infectivity and protective immunity of bovine infectious keratoconjunctivitis. Sonicated M. bovis DNA from the piliated strain EPP63 was inserted into the vector lambda gt11 with EcoRI linkers. Recombinant phage were screened with an oligonucleotide probe based on the amino-terminal portion of the DNA sequence of a Neisseria gonorrhoeae pilin gene. Two candidate phages produced a protein that comigrated with EPP63 beta pilin in sodium dodecyl sulfate-polyacrylamide gels and bound anti-pilus antisera. The 1.9-kilobase insert from one of these, lambda gt11M182, was subcloned in both orientations into pBR322, forming the plasmids pMxB7 and pMxB9, both of which produced beta pilin, as did pMxB12, a HindIII deletion derivative of pMxB7. In HB101(pMxB12), the M. bovis pilin protein was shown to be primarily localized in the inner membrane. The entire 939-base-pair insert of pMxB12 was sequenced, revealing a ribosome binding site just upstream of the coding region and an AT-rich region further upstream containing some potential RNA polymerase recognition sites. The translation of the sequence predicts a six-amino-acid leader sequence preceding the phenylalanine that begins the mature protein. Codon usage analysis of the M. bovis beta pilin gene revealed greater use of the CUA codon for leucine than usual for a well-expressed Escherichia coli gene. Comparisons of the M. bovis EPP63 beta pilin protein sequence with other pilin gene sequences are presented.

Amino Acid Sequence↗

Characterization of a bacteriocinogenic plasmid from Clostridium perfringens and molecular genetic analysis of the bacteriocin-encoding gene.

The bacteriocinogenic plasmid pIP404 from Clostridium perfringens was isolated and cloned in Escherichia coli, and its physical map was deduced. Expression of the bcn gene, encoding bacteriocin BCN5, is inducible by UV irradiation of C. perfringens and thus resembles the SOS-regulated bacteriocin genes of enteric bacteria. The location of bcn on pIP404 was established by a dot-blot procedure, using specific hybridization probes to analyze mRNA samples from induced and uninduced cultures. From the nucleotide sequence of its gene, the molecular weight of BCN5 was deduced to be 96,591, and a protein of this size was secreted by bacteriocin-producing cultures of C. perfringens. The primary structure of the protein suggests that it may function as an ionophore, since a hydrophobic domain, resembling those of the ionophoric colicins, is present at the COOH terminus. No bacteriocin activity could be detected in E. coli harboring plasmids bearing the bcn gene, even when the transcriptional and translational signals were replaced by those of lacZ. A possible explanation may be found in the unusual codon usage of the adenine-thymine-rich bcn gene, as this shows a preference for codons with a high adenine-plus-thymine content, especially in the wobble position. Many of the frequently used codons correspond to those recognized by minor tRNA species in E. coli. Consequently, bcn expression might be limited by tRNA availability in this bacterium.

Amino Acid Sequence↗

Identification and nucleotide sequence of the Leptospira biflexa serovar patoc trpE and trpG genes.

Leptospira biflexa is a representative of an evolutionarily distinct group of eubacteria. In order to better understand the genetic organization and gene regulatory mechanisms of this species, we have chosen to study the genes required for tryptophan biosynthesis in this bacterium. The nucleotide sequence of the region of the L. biflexa serovar patoc chromosome encoding the trpE and trpG genes has been determined. Four open reading frames (ORFs) were identified in this region, but only three ORFs were translated into proteins when the cloned genes were introduced into Escherichia coli. Analysis of the predicted amino acid sequences of the proteins encoded by the ORFs allowed us to identify the trpE and trpG genes of L. biflexa. Enzyme assays confirmed the identity of these two ORFs. Anthranilate synthase from L. biflexa was found to be subject to feedback inhibition by tryptophan. Codon usage analysis showed that there was a bias in L. biflexa towards the use of codons rich in A and T, as would be expected from its G + C content of 37%. Comparison of the amino acid sequences of the trpE gene product and the trpG gene product with corresponding gene products from other bacteria showed regions of highly conserved sequence.

Amino Acid Sequence↗

A comparison of group II introns of plastid tRNALysUUU genes encoding maturase protein.

All higher plant plastid genomes have six classes of tRNA genes containing introns. One of those is the tRNALysUUU gene, which encodes maturase protein. In the case of liverwort species from the genus Porella and mosses from the genus Plagiomnium, the maturase coding gene (matK) represents a truncated form of other plant matK genes: several subdomains of the reverse transcriptase-like domain and so-called domain X are not present in these ORFs. These ORFs probably represent pseudogenes of the matK gene. The analysis of codon usage within the matK gene revealed the presence of strong A/T pressure. The use of codons with the third letter being U or A varies from 71-93%. The comparison of maturase amino acid sequences at the family level shows a high identity between species. However, when liverwort and angiosperm maturase sequences are compared, the percentage of identity drops dramatically. The calculated values of the number of nucleotide substitutions vary considerably, even when liverwort species are compared pairwise. The phenetic tree of relationships between plant species on the basis of tRNALysUUU intron sequences concur with the generally accepted plant phylogeny.

Algorithms↗

Structure and evolution of bacterial adenylate cyclase: comparison between Escherichia coli and Erwinia chrysanthemi.

The cya genes, coding for adenylate cyclase, from Escherichia coli and Erwinia chrysanthemi B374 are compared after determination of a 3632 bp long nucleotide sequence of the hemC-cya region of E. chrysanthemi, encompassing the whole cya gene. In spite of a large divergence between the two organisms, especially visible in non coding regions, the amino acid sequence of the proteins are very similar, except at the very distal carboxyl end. Codon usage is different in the two organisms, and E. chrysanthemi tends to restrict translation to codons ending in G or C. Conservation of the translation initiation start region (including the poor ribosome binding site GGCG, and the TTG start codon), suggests that a specific protein synthesis process controls adenylate cyclase expression. Finally a palindromic unit, of primary sequence differing from the E. coli counterpart, borders the gene in E. chrysanthemi.

Adenylyl Cyclases↗

Structural comparison of two nontandemly repeated yeast glyceraldehyde-3-phosphate dehydrogenase genes.

A hybrid plasmid (pgap63) was isolated which contains a second yeast glyceraldehyde-3-phosphate dehydrogenase structural gene. The complete nucleotide sequence of this gene was determined and compared with the primary structure of a yeast glyceraldehyde-3-phosphate dehydrogenase gene (pgap49) which was reported previously (Holland, J.P., and Holland, M.J. (1979) J. Biol. Chem. 254, 9839-9845). Based on the restriction endonuclease cleavage maps of the isolated segments of yeast DNA which contain these genes, the two genes are nontandemly duplicated. Greater than 94% of the nucleotides within the coding regions of these genes are homologous and the polypeptides encoded by the two structural genes differ by only 15 amino acid residues. Both genes have the same, highly biased, codon usage pattern and neither contains intervening sequences. Approximately 100 nucleotides adjacent to the ATG initiation codons and 130 nucleotides beyond the TAA termination codons are greater than 70% homologous. Structures within the flanking sequences of the genes which are potentially relevant to transcriptional and translational control are described. Several sequences (8 to 15 nucleotides in length) are repeated in both the 5' and 3' flanking sequences of the genes in a noninverted fashion. Finally, a rapid procedure for the isolation of spontaneous deletions within hybrid plasmid DNAs is described, as is the isolation of a structural gene deletion in pgap49.

Amino Acid Sequence↗

Messenger RNA release from ribosomes during 5'-translational blockage by consecutive low-usage arginine but not leucine codons in Escherichia coli.

In '5'-translational blockage', significantly reduced yields of proteins are synthesized in Escherichia coli when consecutive low-usage codons are inserted near translation starts of messages (with reduced or no effect when these same codons are inserted downstream). We tested the hypothesis that ribosomes encountering these low-usage codons near the translation start prematurely release the mRNA. RNA from polysome gradients was fractionated into pools of polysomes and monosomes and a ribosome-free pool. New hybridization probes, called 'molecular beacons', and standard slot blots were used to detect test messages containing either consecutive low-usage AGG (arginine) or synonymous high-usage CGU insertions near the 5' end. The results show an approximately twofold increase in the ratio of free to bound mRNA when the low-usage codons were present in the message compared with when high-usage codons were present. In contrast, there was no difference in the ratio of free to bound mRNA when consecutive low-usage CUA or high-usage CUG (leucine) codons were inserted or when the arginine codons were inserted near the 3' end. These data indicate that at least some mRNA is released from ribosomes during 5'-translational blockage by arginine but not leucine codons, and they support proposals that premature termination of translation can occur in some conditions in vivo in the absence of a stop codon.

Arginine↗

Why are translationally sub-optimal synonymous codons used in Escherichia coli?

Natural selection favors certain synonymous codons which aid translation in Escherichia coli, yet codons not favored by translational selection persist. We use the frequency distributions of synonymous polymorphisms to test three hypotheses for the existence of translationally sub-optimal codons: (1) selection is a relatively weak force, so there is a balance between mutation, selection, and drift; (2) at some sites there is no selection on codon usage, so some synonymous sites are unaffected by translational selection; and (3) translationally sub-optimal codons are favored by alternative selection pressures at certain synonymous sites. We find that when all the data is considered, model 1 is supported and both models 2 and 3 are rejected as sole explanations for the existence of translationally sub-optimal codons. However, we find evidence in favor of both models 2 and 3 when the data is partitioned between groups of amino acids and between regions of the genes. Thus, all three mechanisms appear to contribute to the existence of translationally sub-optimal codons in E. coli.

Codon↗

Random sequence analysis of genomic DNA of a hyperthermophile: Aquifex pyrophilus.

Aquifex pyrophilus is one of the hyperthermophilic bacteria that can grow at temperatures up to 95 degrees C. To obtain information about its genomic structure, random sequencing was performed on plasmid libraries containing 0.5-2 kb genomic DNA fragments of A. pyrophilus. Comparison of the obtained sequence tags with known proteins revealed that 123 tags showed strong similarity to previously identified proteins in the PIR or Genebank databases. These included three proteases, two amino acid racemases, and three enzymes utilizing oxygen as substrate. Although the GC ratio of the genome is about 40%, the codon usage of A. pyrophilus showed biased occurrence of G and C at the third position of codons, especially those for amino acids such as asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, lysine, and tyrosine. A higher ratio of positively charged amino acids in A. pyrophilus proteins as compared with proteins from mesophiles suggested that Aquifex proteins might contain increased ion-pair interaction that could help to maintain heat stability.

Amino Acid Sequence↗

Pseudogene in the genome of bacteriophage lambda?

We find a region in the non-coding part of bacteriophage lambda genome that codes for the conserved fold which repressors and other proteins use for specific DNA binding. The region is involved in a long open reading frame exceeding one kilobase and is read in the same frame as gene A in the opposite strand. The putative translation product of this open reading frame has a highly ordered secondary structure with a predominance of alpha helices, which is typical of repressors. In addition, codon usage in this frame suggests a protein-coding region. However, there is a TGA stop codon located between the putative gene start point and the region coding for the DNA binding fold. It thus appears that bacteriophage lambda had one more DNA binding protein, perhaps repressor, in the past that was inactivated by a mutation.

Amino Acid Sequence↗

Modification of GP63 genes from diverse species of Leishmania for expression of recombinant protein at high levels in Escherichia coli.

Toward the future development of a defined subunit vaccine against leishmaniasis is, high levels of recombinant GP63 for diverse species of Leishmania were produced in Escherichia coli. Several features of Leishmania GP63 genes were simultaneously modified with the polymerase chain reaction (PCR) using either cloned genes or total genomic DNA from Leishmania as template DNA for the PCR amplification reactions. The PCR products included only the coding region for the predicted mature form of GP63 that occurs on the surface of Leishmania, flanked by the appropriate translation signals and cloning sites for the production of recombinant GP63 as nonfusion protein in E. coli. When the codon usage in the GP63 gene was modified to reduce the guanine and cytosine content for the codons adjacent to the ATG initiation codon, rGP63 represented about 50% of total protein in E. coli. Mouse monoclonal antibodies raised against purified Leishmania major rGP63 had equivalent immunoblotting characteristics for native GP63 and recombinant GP63 with respect to linear determinants on GP63 expressed in diverse species of Leishmania. Human T cell lines and clones were derived from a patient infected with Leishmania braziliensis panamensis using rGP63 purified from an L. major GP63 expression clone as antigen.

Animals↗

Characterisation of the dihydrofolate reductase-thymidylate synthetase gene from human malaria parasites highly resistant to pyrimethamine.

To investigate the genetic basis of drug resistance in human malaria parasites, we have sequenced the entire dihydrofolate reductase thymidylate synthetase DHFR-TS bifunctional gene from the highly pyrimethamine-resistant K1 isolate of Plasmodium falciparum. The protein is predicted to consist of 607 amino acids (aa), (71,685 Da), with an N-terminal methionine encoded by the second start codon of the open reading frame. Compared to the sequence from drug-sensitive parasites, there are two nucleotide changes in the coding region which bring about a substitution of Arg for Cys at aa position 59 and Asn for Thr at aa position 108. Both changes occur in regions of the DHFR domain involved in inhibitor and cofactor binding and are hence strongly implicated in drug resistance. The gene is present as a single copy in both K1 and drug-sensitive FCR3 isolates, and is assigned to chromosome 4. Codon usage follows the pattern observed in that of malarial surface antigen genes, with the exception fo codons corresponding to Val and Pro. The Asn and Lys contents of the predicted protein are exceptionally high, these residues being particularly concentrated in the DHFR and junction domains.

Amino Acid Sequence↗

Delineation of coding areas in DNA sequences through assignment of codon probabilities.

Codon usage tables have been produced for E. coli, yeast, human, and mouse. The nonrandom employment of codons allows assignment of probability values to trinucleotides in any DNA sequence. These values represent the probability that a given trinucleotide is used as a codon in the organism from which the table is derived. For the graphical delineation of coding areas in DNA sequences, a probability is assigned to each trinucleotide equal to its frequency in the codon table. Averaging and smoothing procedures then greatly enhance the detectability of areas of high average codon probability and better represent the mean codon probability. These manipulations increase graphical clarity without altering the overall magnitude of probabilities. Averaging introduces an error of less than 0.5% between "raw" and smoothed data. This graphical delineation of coding sequences does not depend on the presence of punctuation, ribosomal binding sites, etc: moreover the delineation of introns and exons is also possible.

Base Sequence↗

Comparison of a vitellogenin gene between two distantly related rhabditid nematode species.

Three vitellogenin genes from the free-living nematode Caenorhabditis elegans have previously been characterized at the molecular level. In order to study evolutionary relationships within this poorly understood taxon, we have cloned a vitellogenin gene, CEW1-vit-6, from a distantly related species belonging to the same family as C. elegans. Screening of a genomic library with a probe to total poly(A+) RNA yielded three clones that hybridized more intensely than all others, and all three corresponded to a single gene homologous to C. elegans vit-6. Comparison of CEW1-vit-6 with Ce-vit-6 reveals both strong similarities and surprising differences. Life Ce-vit-6, the gene is about 5 kb long and contains four unusually small introns (38-41 nt), but only one interrupts the gene at the same location as a Ce-vit-6 intron. The promoter region contains five matches to Vitellogenin Promoter Element 1 (VPE1) and no matches to VPE2, both previously shown to be required for vit gene transcription in C. elegans. Codon usage is in general similar to that of the Ce-vit genes, but a few codon biases are quite different. Alignment of the CEW1-vit-6 protein with Ce-vit-6 and Ce-vit-2 products suggests the existence of two domains which have evolved at different rates. Sequence comparison shows that nematode vitellogenins are much more closely related to vertebrate than to insect vitellogenins.

Amino Acid Sequence↗

DNA and protein sequence homologies between the adhesins of Mycoplasma genitalium and Mycoplasma pneumoniae.

Mycoplasma genitalium and Mycoplasma pneumoniae are morphologically and serologically related pathogens that colonize the human host. Their successful parasitism appears to be dependent on the product, an adhesin protein, of a gene that is carried by each of these mycoplasmas. Here we describe the cloning and determine the sequence of the structural gene for the putative adhesin of M. genitalium and compare its sequence to the counterpart P1 gene of M. pneumoniae. Regions of homology that were consistent with the observed serological cross-reactivity between these adhesins were detected at both DNA and protein levels. However, the degree of homology between these two genes and their products was much higher than anticipated. Interestingly, the A + T content of the M. genitalium adhesin gene was calculated as 60.1%, which is substantially higher tham that of the P1 gene (46.5%). Comparisons of codon usage between the two organisms revealed that M. genitalium preferentially used A- and T-rich codons. A total of 65% of positions 3 and 56% of positions 1 in M. genitalium codons were either A or T, whereas M. pneumoniae utilized A or T for positions 3 and 1 at a frequency of 40 and 47%, respectively. The biased choice of the A- and T-rich codons in M. genitalium could also account for the preferential use of A- and T-rich codons in conservative amino acid substitutions found in the M. genitalium adhesin. These facts suggest that M. genitalium might have evolved independently of other human mycoplasma species, including M. pneumoniae.

Amino Acid Sequence↗

Characterization of a molluscum contagiosum virus homolog of the vaccinia virus p37K major envelope antigen.

We present the first nucleotide sequence data for molluscum contagiosum virus (MCV), an unclassified poxvirus. A 2,276-bp XhoI fragment from a near left-terminal fragment of MCV subtype I (MCVI) and a 1,920-bp XhoI fragment from the corresponding locus of MCV subtype II (MCVII) were sequenced and analyzed for open reading frames (ORFs). A large, complete ORF of 1,167 bp was present in both fragments. The putative polypeptide has a calculated molecular mass of 43 kDa (p43K protein) and was shown to have a high degree of homology to the vaccinia virus p37K major envelope antigen (40% amino acid identity and 22% conservative changes). The nucleotide content of the MCV fragments sequenced was 66% G or C. The codon usage within the gene for p43K reflected this high G + C content, with position 3 of codons being predominantly G or C (82 and 87% for MCVI and MCVII, respectively). The MCV p43K-encoding gene has motifs immediately upstream which are similar to those required for vaccinia virus late gene expression. The location and direction of transcription of the MCV p43K-encoding gene were equivalent to those of the vaccinia virus p37K gene, revealing similarity in genetic organization between MCV and vaccinia virus. Another, incomplete ORF was identified downstream of the p43K-encoding gene in both MCVI and MCVII. The sequence immediately upstream of this ORF overlapped the termination codon of the p43K-encoding gene and contained a motif which had homology to the derived consensus sequence for vaccinia virus early gene promoters.

Amino Acid Sequence↗

A synthetic E7 gene of human papillomavirus type 16 that yields enhanced expression of the protein in mammalian cells and is useful for DNA immunization studies.

A synthetic E7 gene of human papillomavirus (HPV) type 16 was generated that consists entirely of preferred human codons. Expression analysis of the synthetic E7 gene in human and animal cells showed levels of E7 protein 20- to 100-fold higher than those obtained with wild-type E7. Enhanced expression of E7 protein resulted from highly efficient translation, as well as increased stability of the E7 mRNA due to its codon optimization. Higher levels of E7 protein in cells transfected with synthetic E7 correlated with significant loss of cell viability in various human cell lines. In contrast, lower E7 protein expression driven by the wild-type gene resulted in a slight induction of cell proliferation. Furthermore, mice inoculated with plasmids expressing the synthetic E7 gene produced significantly higher levels of E7 antibodies than littermates injected with wild-type E7, suggesting that synthetic E7 may be useful for DNA immunization studies and the development of genetic vaccines against HPV-16. In view of these results, we hypothesize that HPVs may have retained a pattern of G + C content and codon usage distinct from that of their host cells in response to selective pressure. Thus, the nonhuman codon bias may have been conserved by HPVs to prevent compromising viability of the host cells by excessive viral early protein expression, as well as to evade the immune system.

Amino Acid Sequence↗

Coding in the noncoding DNA strand: A novel mechanism of gene evolution?

The question whether the noncoding DNA strand had or still has the capability for encoding functional polypeptides has been addressed in several articles. The theoretical background of the views advocating this idea arose from two groups of findings. One of them was based on various observations implying that the genetic code was adapted for double-strand coding. The other group of theories arose from the observation of gene-length overlapping open reading frames (O-ORFs) on the antisense DNA strand in a number of genes. In fact, the above theories, which I term selectionist, conceive a novel conception of gene evolution, proposing that new genes can be created by the utilization of antisense DNA strand. In contrast, neutralist theory claims that the O-ORFs are mere by-products of evolutionary processes acting to create special codon usage and base distribution patterns in the coding sequences.

Codon↗