Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Specific use of start codons and cellular localization of splice variants of human phosphodiesterase 9A gene.

BACKGROUND: Phosphodiesterases are an important protein family that catalyse the hydrolysis of cyclic nucleotide monophosphates (cAMP and cGMP), second intracellular messengers responsible for transducing a variety of extra-cellular signals. A number of different splice variants have been observed for the human phosphodiesterase 9A gene, a cGMP-specific high-affinity PDE. These mRNAs differ in the use of specific combinations of exons located at the 5' end of the gene while the 3' half, that codes for the catalytic domain of the protein, always has the same combination of exons. It was observed that to deduce the protein sequence with the catalytic domain from all the variants, at least two ATG start codons have to be used. Alternatively some variants code for shorter non-functional polypeptides. RESULTS: In the present study, we expressed different splice variants of PDE9A in HeLa and Cos-1 cells with EGFP fluorescent protein in phase with the catalytic domain sequence in order to test the different start codon usage in each splice variant. It was found that at least two ATG start codons may be used and that the open reading frame that includes the catalytic domain may be translated. In addition the proteins produced from some of the splice variants are targeted to membrane ruffles and cellular vesicles while other variants appear to be cytoplasmic. A hypothesis about the functional meaning of these results is discussed. CONCLUSION: Our data suggest the utilization of two different start codons to produce a variety of different PDE9A proteins, allowing specific subcellular location of PDE9A splice variants.

3',5'-Cyclic-AMP Phosphodiesterases↗

Relating physicochemical properties of amino acids to variable nucleotide substitution patterns among sites.

Markov-process models of codon substitution were implemented that account for features of DNA sequence evolution (such as transition/transversion bias and codon usage bias) as well as heterogeneity of amino acid substitution pattern over sites. The codon (amino acid) sites are assumed to come from several classes (such as secondary structure categories), among which the rate of amino acid substitution and the effect of amino acid chemical properties vary. Parameters are estimated by the maximum likelihood method, which accounts for the phylogenetic relationship among species and corrects for multiple hits at the same site. The likelihood ratio test is used to compare models. Mitochondrial cytochrome b genes of 28 primate species are analyzed. The site-heterogeneity models provide much better fit to previous homogeneous models.

Amino Acid Substitution↗

[Characteristics of the context shift in the frequency of synonymic codons in Escherichia coli].

We have demonstrated, that coding regions of E. coli DNA exhibit the non-random shifts of codon usage frequency depending both on the type of the 3' nucleotide adjacent to the codon and the degree of gene expression. Analysis of primary data--statistics of tetranucleotide occurrences--was performed by the techniques of contingency tables. The results of the investigation allowed us to suggest that the phenomenon observed is connected with the influence of the 3'-adjacent nucleotide on the level of missens-errors and another types of inaccuracy during translation. The specific advantages of such mutations in the third position of the codon are based on the adaptation of the codon to the 3' context in order to increase the efficiency of translation.

Base Sequence↗

Organization and structure of Volvox alpha-tubulin genes.

Southern analysis of Volvox genomic DNA revealed two genes homologous to Chlamydomonas reinhardtii alpha-tubulin cDNA. Restriction fragment length polymorphism analysis indicated that the two genes are not genetically linked. Clones representing one of the alpha-tubulin genes have been isolated from a genomic library of Volvox carteri f. nagariensis. A 3153 bp BamHI fragment containing the entire alpha-tubulin gene (1802 bp) plus 707 bp of the 5'- and 644 bp of the 3'-untranslated regions has been sequenced, revealing the following features: (1) the derived alpha-tubulin primary structure of 451 amino acids is highly conserved, differing in two residues from the alpha 1- and in two additional residues from the alpha 2-tubulin of C. reinhardtii; (2) in comparison to the C. reinhardtii genes, the Volvox alpha-tubulin gene contains a third intron; positions of the other two introns are precisely conserved; (3) codon usages are biased towards G or C, and against A, in the third position; 19 codons are absent from the alpha-tubulin coding sequence, and 5 of these are not used in any of 7 compiled Volvox genes; (4) transcription begins with an A, 30 bp downstream of the putative TATA box; upstream of the TATA box is a 14 bp sequence similar to consensus sequences found in all 4 C. reinhardtii tubulin genes and believed to regulate promoter function.

Amino Acid Sequence↗

Global mRNA stability is not associated with levels of gene expression in Drosophila melanogaster but shows a negative correlation with codon bias.

A multitude of factors contribute to the regulation of gene expression in living cells. The relationship between codon usage bias and gene expression has been extensively studied, and it has been shown that codon bias may have adaptive significance in many unicellular and multicellular organisms. Given the central role of mRNA in post-transcriptional regulation, we hypothesize that mRNA stability is another important factor associated either with positive or negative regulation of gene expression. We have conducted genome-wide studies of the association between gene expression (measured as transcript abundance in public EST databases), mRNA stability, codon bias, GC content, and gene length in Drosophila melanogaster. To remove potential bias of gene length inherently present in EST libraries, gene expression is measured as normalized transcript abundance. It is demonstrated that codon bias and GC content in second codon position are positively associated with transcript abundance. Gene length is negatively associated with transcript abundance. The stability of thermodynamically predicted mRNA secondary structures is not associated with transcript abundance, but there is a negative correlation between mRNA stability and codon bias. This finding does not support the hypothesis that codon bias has evolved as an indirect consequence of selection favoring thermodynamically stable mRNA molecules.

Animals↗

Mutational analysis of the Streptomyces scabies esterase signal peptide.

Ten site-directed mutations affecting the predicted 39-amino-acid signal peptide of the Streptomyces scabies esterase were used to examine start-codon usage and esterase secretion in S. lividans. The first of two in-frame AUG codons was preferred for translation initiation. Removal of 2 of the 4 positively charged amino acids at the amino terminus of the signal peptide reduced esterase expression more than 100-fold; however, deletion of all 4 charged residues reduced expression by only 2- to 5-fold. Deletion of 4 or 8 amino acids from the hydrophobic core of the signal peptide reduced esterase production more than 200-fold, and a signal peptide processing site deletion completely disrupted esterase expression. For all constructs in which a mutation in the signal sequence decreased esterase production, esterase mRNA levels were also reduced, suggesting that a defect in secretion or processing affected esterase transcript abundance.

Amino Acid Sequence↗

Transcriptional and translational regulation of the expression of the l(2)gl tumor suppressor gene of Drosophila melanogaster.

By structural, biochemical and molecular genetic analyses, we have investigated the different mechanisms that control the expression of the lethal(2) giant larvae gene, a tumor suppressor gene of Drosophila melanogaster. Transcription of the l(2)gl gene is controlled by two highly identical promoters that result from the duplication of the 2.8 kb proximal portion of the gene. These two repeats are 96% homologous. Reverse genetic analysis has shown that each promoter can drive gene expression. In addition to the promoters, both repeats express two or three exons according to the pattern of splicing. The most distal exon in the second repeat is required because it contains the ATG initiating codon at the beginning of the open reading frame. The 3' untranslated region appears to contain motifs that specifically destabilize the transcript. Deletion of this region results in the formation of more stable mRNAs. The l(2)gl gene is characterized by an unusual codon usage that may reflect an enhanced translation efficiency by moderating the strength of pairing between codons and anticodons and may therefore increase the expressivity of this gene. Analysis of the spatio-temporal expression of the l(2)gl transcripts and proteins has shown that transcripts and proteins are produced ubiquitously during early embryogenesis, at a time when expression of the gene is required for preventing tumorigenesis. In the second half of embryogenesis, l(2)gl expression becomes restricted to tissues that do not show any phenotypic alteration in mutant animals. The l(2)gl protein exhibits two distinct intracellular localizations. It is preferentially found free in the cytoplasm but can become associated with the inner face of the plasma membrane where it is restricted to domains facing contiguous cells. In particular, the l(2)gl protein is absent from the basal and apical domains of the plasma membrane. The aim of the current research is directed towards understanding the functional relevance of the l(2)gl protein binding to the plasma membrane and its role in the control of cell proliferation and differentiation.

Animals↗

Three stages in the evolution of the genetic code.

A diversification of the genetic code based on the number of codons available for the proteinous amino acids is established. Three groups of amino acids during evolution of the code are distinguished. On the basis of their chemical complexity those amino acids emerging later in a translation process are derived. Codon number and chemical complexity indicate that His, Phe, Tyr, Cys and either Lys or Asn were introduced in the second stage, whereas the number of codons alone gives evidence that Trp and Met were introduced in the third stage. The amino acids of stage 1 use purine-rich codons, while all the amino acids introduced in the second stage, in contrast, use pyrimidines in the third position of their codons. A low abundance of pyrimidines during early translation is derived. This assumption is supported by experiments on non-enzymatic replication and interactions of hairpin loops with a complementary strand. A back extrapolation concludes a high purine content of the first nucleic acids, which gradually decreased during their evolution. Amino acids independently available from prebiotic synthesis were thus correlated to purine-rich codons. Implications on the prebiotic replication are discussed also in the light of recent codon usage data.

Amino Acids↗

Cloning and sequencing of a Moraxella bovis pilin gene.

Moraxella bovis pili have been shown to play a major role in both infectivity and protective immunity of bovine infectious keratoconjunctivitis. Sonicated M. bovis DNA from the piliated strain EPP63 was inserted into the vector lambda gt11 with EcoRI linkers. Recombinant phage were screened with an oligonucleotide probe based on the amino-terminal portion of the DNA sequence of a Neisseria gonorrhoeae pilin gene. Two candidate phages produced a protein that comigrated with EPP63 beta pilin in sodium dodecyl sulfate-polyacrylamide gels and bound anti-pilus antisera. The 1.9-kilobase insert from one of these, lambda gt11M182, was subcloned in both orientations into pBR322, forming the plasmids pMxB7 and pMxB9, both of which produced beta pilin, as did pMxB12, a HindIII deletion derivative of pMxB7. In HB101(pMxB12), the M. bovis pilin protein was shown to be primarily localized in the inner membrane. The entire 939-base-pair insert of pMxB12 was sequenced, revealing a ribosome binding site just upstream of the coding region and an AT-rich region further upstream containing some potential RNA polymerase recognition sites. The translation of the sequence predicts a six-amino-acid leader sequence preceding the phenylalanine that begins the mature protein. Codon usage analysis of the M. bovis beta pilin gene revealed greater use of the CUA codon for leucine than usual for a well-expressed Escherichia coli gene. Comparisons of the M. bovis EPP63 beta pilin protein sequence with other pilin gene sequences are presented.

Amino Acid Sequence↗

Characterization of a bacteriocinogenic plasmid from Clostridium perfringens and molecular genetic analysis of the bacteriocin-encoding gene.

The bacteriocinogenic plasmid pIP404 from Clostridium perfringens was isolated and cloned in Escherichia coli, and its physical map was deduced. Expression of the bcn gene, encoding bacteriocin BCN5, is inducible by UV irradiation of C. perfringens and thus resembles the SOS-regulated bacteriocin genes of enteric bacteria. The location of bcn on pIP404 was established by a dot-blot procedure, using specific hybridization probes to analyze mRNA samples from induced and uninduced cultures. From the nucleotide sequence of its gene, the molecular weight of BCN5 was deduced to be 96,591, and a protein of this size was secreted by bacteriocin-producing cultures of C. perfringens. The primary structure of the protein suggests that it may function as an ionophore, since a hydrophobic domain, resembling those of the ionophoric colicins, is present at the COOH terminus. No bacteriocin activity could be detected in E. coli harboring plasmids bearing the bcn gene, even when the transcriptional and translational signals were replaced by those of lacZ. A possible explanation may be found in the unusual codon usage of the adenine-thymine-rich bcn gene, as this shows a preference for codons with a high adenine-plus-thymine content, especially in the wobble position. Many of the frequently used codons correspond to those recognized by minor tRNA species in E. coli. Consequently, bcn expression might be limited by tRNA availability in this bacterium.

Amino Acid Sequence↗

Identification and nucleotide sequence of the Leptospira biflexa serovar patoc trpE and trpG genes.

Leptospira biflexa is a representative of an evolutionarily distinct group of eubacteria. In order to better understand the genetic organization and gene regulatory mechanisms of this species, we have chosen to study the genes required for tryptophan biosynthesis in this bacterium. The nucleotide sequence of the region of the L. biflexa serovar patoc chromosome encoding the trpE and trpG genes has been determined. Four open reading frames (ORFs) were identified in this region, but only three ORFs were translated into proteins when the cloned genes were introduced into Escherichia coli. Analysis of the predicted amino acid sequences of the proteins encoded by the ORFs allowed us to identify the trpE and trpG genes of L. biflexa. Enzyme assays confirmed the identity of these two ORFs. Anthranilate synthase from L. biflexa was found to be subject to feedback inhibition by tryptophan. Codon usage analysis showed that there was a bias in L. biflexa towards the use of codons rich in A and T, as would be expected from its G + C content of 37%. Comparison of the amino acid sequences of the trpE gene product and the trpG gene product with corresponding gene products from other bacteria showed regions of highly conserved sequence.

Amino Acid Sequence↗

Archaeology and evolution of transfer RNA genes in the Escherichia coli genome.

Transfer RNA genes tend to be presented in multiple copies in the genomes of most organisms, from bacteria to eukaryotes. The evolution and genomic structure of tRNA genes has been a somewhat neglected area of molecular evolution. Escherichia coli, the first phylogenetic species for which more than two different strains have been sequenced, provides an invaluable framework to study the evolution of tRNA genes. In this work, a detailed analysis of the tRNA structure of the genomes of Escherichia coli strains K12, CFT073, and O157:H7, Shigella flexneri 2a 301, and Salmonella typhimurium LT2 was carried out. A phylogenetic analysis of these organisms was completed, and an archaeological map depicting the main events in the evolution of tRNA genes was drawn. It is shown that duplications, deletions, and horizontal gene transfers are the main factors driving tRNA evolution in these genomes. On average, 0.64 tRNA insertions/duplications occur every million years (Myr) per genome per lineage, while deletions occur at the slower rate of 0.30 per million years per genome per lineage. This work provides a first genomic glance at the problem of tRNA evolution as a repetitive process, and the relationship of this mechanism to genome evolution and codon usage is discussed.

Codon↗

A comparison of group II introns of plastid tRNALysUUU genes encoding maturase protein.

All higher plant plastid genomes have six classes of tRNA genes containing introns. One of those is the tRNALysUUU gene, which encodes maturase protein. In the case of liverwort species from the genus Porella and mosses from the genus Plagiomnium, the maturase coding gene (matK) represents a truncated form of other plant matK genes: several subdomains of the reverse transcriptase-like domain and so-called domain X are not present in these ORFs. These ORFs probably represent pseudogenes of the matK gene. The analysis of codon usage within the matK gene revealed the presence of strong A/T pressure. The use of codons with the third letter being U or A varies from 71-93%. The comparison of maturase amino acid sequences at the family level shows a high identity between species. However, when liverwort and angiosperm maturase sequences are compared, the percentage of identity drops dramatically. The calculated values of the number of nucleotide substitutions vary considerably, even when liverwort species are compared pairwise. The phenetic tree of relationships between plant species on the basis of tRNALysUUU intron sequences concur with the generally accepted plant phylogeny.

Algorithms↗

Structure and evolution of bacterial adenylate cyclase: comparison between Escherichia coli and Erwinia chrysanthemi.

The cya genes, coding for adenylate cyclase, from Escherichia coli and Erwinia chrysanthemi B374 are compared after determination of a 3632 bp long nucleotide sequence of the hemC-cya region of E. chrysanthemi, encompassing the whole cya gene. In spite of a large divergence between the two organisms, especially visible in non coding regions, the amino acid sequence of the proteins are very similar, except at the very distal carboxyl end. Codon usage is different in the two organisms, and E. chrysanthemi tends to restrict translation to codons ending in G or C. Conservation of the translation initiation start region (including the poor ribosome binding site GGCG, and the TTG start codon), suggests that a specific protein synthesis process controls adenylate cyclase expression. Finally a palindromic unit, of primary sequence differing from the E. coli counterpart, borders the gene in E. chrysanthemi.

Adenylyl Cyclases↗

Structural comparison of two nontandemly repeated yeast glyceraldehyde-3-phosphate dehydrogenase genes.

A hybrid plasmid (pgap63) was isolated which contains a second yeast glyceraldehyde-3-phosphate dehydrogenase structural gene. The complete nucleotide sequence of this gene was determined and compared with the primary structure of a yeast glyceraldehyde-3-phosphate dehydrogenase gene (pgap49) which was reported previously (Holland, J.P., and Holland, M.J. (1979) J. Biol. Chem. 254, 9839-9845). Based on the restriction endonuclease cleavage maps of the isolated segments of yeast DNA which contain these genes, the two genes are nontandemly duplicated. Greater than 94% of the nucleotides within the coding regions of these genes are homologous and the polypeptides encoded by the two structural genes differ by only 15 amino acid residues. Both genes have the same, highly biased, codon usage pattern and neither contains intervening sequences. Approximately 100 nucleotides adjacent to the ATG initiation codons and 130 nucleotides beyond the TAA termination codons are greater than 70% homologous. Structures within the flanking sequences of the genes which are potentially relevant to transcriptional and translational control are described. Several sequences (8 to 15 nucleotides in length) are repeated in both the 5' and 3' flanking sequences of the genes in a noninverted fashion. Finally, a rapid procedure for the isolation of spontaneous deletions within hybrid plasmid DNAs is described, as is the isolation of a structural gene deletion in pgap49.

Amino Acid Sequence↗

Messenger RNA release from ribosomes during 5'-translational blockage by consecutive low-usage arginine but not leucine codons in Escherichia coli.

In '5'-translational blockage', significantly reduced yields of proteins are synthesized in Escherichia coli when consecutive low-usage codons are inserted near translation starts of messages (with reduced or no effect when these same codons are inserted downstream). We tested the hypothesis that ribosomes encountering these low-usage codons near the translation start prematurely release the mRNA. RNA from polysome gradients was fractionated into pools of polysomes and monosomes and a ribosome-free pool. New hybridization probes, called 'molecular beacons', and standard slot blots were used to detect test messages containing either consecutive low-usage AGG (arginine) or synonymous high-usage CGU insertions near the 5' end. The results show an approximately twofold increase in the ratio of free to bound mRNA when the low-usage codons were present in the message compared with when high-usage codons were present. In contrast, there was no difference in the ratio of free to bound mRNA when consecutive low-usage CUA or high-usage CUG (leucine) codons were inserted or when the arginine codons were inserted near the 3' end. These data indicate that at least some mRNA is released from ribosomes during 5'-translational blockage by arginine but not leucine codons, and they support proposals that premature termination of translation can occur in some conditions in vivo in the absence of a stop codon.

Arginine↗

Why are translationally sub-optimal synonymous codons used in Escherichia coli?

Natural selection favors certain synonymous codons which aid translation in Escherichia coli, yet codons not favored by translational selection persist. We use the frequency distributions of synonymous polymorphisms to test three hypotheses for the existence of translationally sub-optimal codons: (1) selection is a relatively weak force, so there is a balance between mutation, selection, and drift; (2) at some sites there is no selection on codon usage, so some synonymous sites are unaffected by translational selection; and (3) translationally sub-optimal codons are favored by alternative selection pressures at certain synonymous sites. We find that when all the data is considered, model 1 is supported and both models 2 and 3 are rejected as sole explanations for the existence of translationally sub-optimal codons. However, we find evidence in favor of both models 2 and 3 when the data is partitioned between groups of amino acids and between regions of the genes. Thus, all three mechanisms appear to contribute to the existence of translationally sub-optimal codons in E. coli.

Codon↗

Random sequence analysis of genomic DNA of a hyperthermophile: Aquifex pyrophilus.

Aquifex pyrophilus is one of the hyperthermophilic bacteria that can grow at temperatures up to 95 degrees C. To obtain information about its genomic structure, random sequencing was performed on plasmid libraries containing 0.5-2 kb genomic DNA fragments of A. pyrophilus. Comparison of the obtained sequence tags with known proteins revealed that 123 tags showed strong similarity to previously identified proteins in the PIR or Genebank databases. These included three proteases, two amino acid racemases, and three enzymes utilizing oxygen as substrate. Although the GC ratio of the genome is about 40%, the codon usage of A. pyrophilus showed biased occurrence of G and C at the third position of codons, especially those for amino acids such as asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, lysine, and tyrosine. A higher ratio of positively charged amino acids in A. pyrophilus proteins as compared with proteins from mesophiles suggested that Aquifex proteins might contain increased ion-pair interaction that could help to maintain heat stability.

Amino Acid Sequence↗