Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Mammalian mitochondrial DNA evolution: a comparison of the cytochrome b and cytochrome c oxidase II genes.

The evolution of two mitochondrial genes, cytochrome b and cytochrome c oxidase subunit II, was examined in several eutherian mammal orders, with special emphasis on the orders Artiodactyla and Rodentia. When analyzed using both maximum parsimony, with either equal or unequal character weighting, and neighbor joining, neither gene performed with a high degree of consistency in terms of the phylogenetic hypotheses supported. The phylogenetic inconsistencies observed for both these genes may be the result of several factors including differences in the rate of nucleotide substitution among particular lineages (especially between orders), base composition bias, transition/transversion bias, differences in codon usage, and different constraints and levels of homoplasy associated with first, second, and third codon positions. We discuss the implications of these findings for the molecular systematics of mammals, especially as they relate to recent hypotheses concerning the polyphyly of the order Rodentia, relationships among the Artiodactyla, and various interordinal relationships.

Animals↗

In-phase implies large likelihood for independent codon model: distinguishing coding from non-coding sequences.

It is proven that under the independent codon model, the likelihood of a DNA coding sequence read according to the correct frame is asymptotically larger than that read with an incorrect frame. Based on this proposition, a single set of probabilities of the codon usage is enough for discriminating the six frames of coding sequences under the independent codon model. The direct coding sequence of Escherichia coli genome is taken as an example to examine the codon independency by using the mutual information and chi2 analysis. The contrast between the coding frame and the two offset frames is evident. A self-learning approach for generating training set is proposed to estimate probability parameters.

Codon↗

Molecular characterization of cDNA encoding for adenylate kinase of rice (Oryza sativa L.).

Two types of genes (Adk-a, and Adk-b) encoding for adenylate kinase (AK, EC 2.7.4.3.) were isolated from the cDNA library constructed from poly(A)+ RNA of rice (Oryza sativa L.). Two cDNAs were heterogeneous at 5' and 3' ends of non-coding sequences and had possible polyadenylation signals. One of the genes, Adk-a, had 1154 bp sequences encoding 241 amino acid residues, while the other type, Adk-b, contained 1085 bp sequences encoding for 243 amino acid residues. Homology between Adk-a and Adk-b was 73.7% in nucleotide sequences, and 90.8% in amino acid level. Two genes showed about 53% homology to bovine mitochondrial adenylate kinase (AK2) at nucleotide and amino acid levels. Concerning the codon usage of rice AK genes, T was abundant at the third position of a codon in the reading frames. In order to examine the enzyme activity of the protein encoded by the rice cDNA, Adk-a was cloned into an expression vector, pUC119, which was introduced into Escherichia coli strain CV2, a temperature-sensitive mutant of adenylate kinase. We found that the transformant carrying the rice Adk-a gene in the sense orientation recovered cell growth at non-permissive high temperature (42 degrees C) and expressed enzyme activities higher than the untransformed CV2 and the transformant possessing Adk-a cDNA in the antisense orientation. These observations suggest that rice Adk-a codes a biologically active enzyme. Furthermore, sucrose was found to regulate the transcription of AK genes in rice cell cultures. Organ related accumulation of mRNA in whole plants was also found.

Adenylate Kinase↗

Molecular evolution of ependymin and the phylogenetic resolution of early divergences among euteleost fishes.

The rate and pattern of DNA evolution of ependymin, a single-copy gene coding for a highly expressed glycoprotein in the brain matrix of teleost fishes, is characterized and its phylogenetic utility for fish systematics is assessed. DNA sequences were determined from catfish, electric fish, and characiforms and compared with published ependymin sequences from cyprinids, salmon, pike, and herring. Among these groups, ependymin amino acid sequences were highly divergent (up to 60% sequence difference), but had surprisingly similar hydropathy profiles and invariant glycosylation sites, suggesting that functional properties of the proteins are conserved. Comparison of base composition at third codon positions and introns revealed AT-rich introns and GC-rich third codon positions, suggesting that the biased codon usage observed might not be due to mutational bias. Phylogenetic information content of third codon positions was surprisingly high and sufficient to recover the most basal nodes of the tree, in spite of the observation that pairwise distances (at third codon positions) were well above the presumed saturation level. This finding can be explained by the high proportion of phylogenetically informative nonsynonymous changes at third codon positions among these highly divergent proteins. Ependymin DNA sequences have established the first molecular evidence for the monophyly of a group containing salmonids and esociforms. In addition, ependymin suggests a sister group relationship of electric fish (Gymnotiformes) and Characiformes, constituting a significant departure from currently accepted classifications. However, relationships among characiform lineages were not completely resolved by ependymin sequences in spite of seemingly appropriate levels of variation among taxa and considerably low levels of homoplasy in the data (consistency index = 0.7). If the diversification of Characiformes took place in an "explosive" manner, over a relatively short period of time this pattern should also be observed using other phylogenetic markers. Poor conservation of ependymin's primary structure hinders the design of efficient primers for PCR that could be used in wide-ranging fish systematic studies. However, alternative methods like PCR amplification from cDNA used here should provide promising comparative sequence data for the resolution of phylogenetic relationships among other basal lineages of teleost fishes.

Amino Acid Sequence↗

Comparison and cross-species expression of the acetyl-CoA synthetase genes of the Ascomycete fungi, Aspergillus nidulans and Neurospora crassa.

The genes encoding the acetate-inducible enzyme acetyl-coenzyme A synthetase from Neurospora crassa and Aspergillus nidulans (acu-5 and facA, respectively) have been cloned and their sequences compared. The predicted amino acid sequence of the Aspergillus enzyme has 670 amino acid residues and that of the Neurospora enzyme either 626 or 606 residues, depending upon which of the two possible initiation codons is used. The amino acid sequences following the second alternative AUG show 86% homology between the two species; the extended N-terminal sequences show no homology. The Neurospora protein is characterized by the appearance of the S(T)PXX sequence motif where the amino acid homologies break down. The codon usage is biased in both genes, with a marked deficiency, especially in Neurospora, of codons with A in the third position. The facA transcribed sequence contains six introns: one in the long leader sequence, one in the 5' coding sequence not homologous with acu-5, and four within the sequence that is largely similar to that of acu-5. Only one intron, corresponding in size and position to the furthest downstream of the facA introns, is found in acu-5. The evolution of introns during the divergence of these two Ascomycete fungi is discussed. Each of the two genes has been transferred by transformation into the other species. Each species is evidently able to splice out the other's introns. Most transformants have normal acetate-induction of acetyl-CoA synthetase, implying that the two genes respond to transcriptional control signals common to both species, in spite of the striking divergence of their 5' ends.

Acetate-CoA Ligase↗

Cryptic plasmid of Neisseria gonorrhoeae: complete nucleotide sequence and genetic organization.

The naturally occurring cryptic plasmid pJD1 of Neisseria gonorrhoeae is 4,207 base pairs long and is found in about 96% of gonococcal strains. The total probable coding capacity of pJD1 was determined from the complete nucleotide sequence by using computational probes to identify open reading frames with similar codon usage and by screening for the presence of ribosomal binding sites before the start codons. Candidates for promoters and terminators were also found in the sequence. Based on these findings, we propose a model for the genetic organization of the plasmid. The model predicts two transcriptional units, each composed of five compactly spaced genes. A promoter of one of the transcripts was shown to function in Escherichia coli, and the products of three of the five genes in this operon were identified in minicell expression experiments. Of these, the cppA gene encoded a 9-kilodalton protein, and the cppB and cppC genes both coded for 24-kilodalton proteins. No expression of the other transcriptional unit was detected, but two genes in this operon were expressed in minicells when transcribed from an E. coli promoter. The experimental data were consistent with the model.

Amino Acid Sequence↗

Identification, sequence analysis, and expression of a Corynebacterium glutamicum gene cluster encoding the three glycolytic enzymes glyceraldehyde-3-phosphate dehydrogenase, 3-phosphoglycerate kinase, and triosephosphate isomerase.

To investigate a possible chromosomal clustering of glycolytic enzyme genes in Corynebacterium glutamicum, a 6.4-kb DNA fragment located 5' adjacent to the structural phosphoenolpyruvate carboxylase (PEPCx) gene ppc was isolated. Sequence analysis of the ppc-proximal part of this fragment identified a cluster of three glycolytic genes, namely, the glyceraldehyde-3-phosphate dehydrogenase (GAPDH) gene gap, the 3-phosphoglycerate kinase (PGK) gene pgk, and the triosephosphate isomerase (TPI) gene tpi. The four genes are organized in the order gap-pgk-tpi-ppc and are separated by 215 bp (gap and pgk), 78 bp (pgk and tpi), and 185 bp (tpi and ppc). The predicted gene product of gap consists of 336 amino acids (M(r) of 36,204), that of pgk consists of 403 amino acids (M(r) of 42,654), and that of tpi consists of 259 amino acids (M(r) of 27,198). The amino acid sequences of the three enzymes show up to 62% (GAPDH), 48% (PGK), and 44% (TPI) identity in comparison with respective enzymes from other organisms. The gap, pgk, tpi, and ppc genes were cloned into the C. glutamicum-Escherichia coli shuttle vector pEK0 and introduced into C. glutamicum. Relative to the wild type, the recombinant strains showed up to 20-fold-higher specific activities of the respective enzymes. On the basis of codon usage analysis of gap, pgk, tpi, and previously sequenced genes from C. glutamicum, a codon preference profile for this organism which differs significantly from those of E. coli and Bacillus subtilis is presented.

Amino Acid Sequence↗

The two beta-tubulin genes of Chlamydomonas reinhardtii code for identical proteins.

The two beta-tubulin genes of the unicellular green alga Chlamydomonas reinhardtii are expressed coordinately after deflagellation and produce two transcripts of 2.1 and 2.0 kilobases. Full-length cDNA clones corresponding to the transcript of each gene were isolated. DNA sequences were obtained from the cDNA clones and from cloned tubulin gene fragments. Both genes contained 1,332 base pairs of coding sequence, with only 19 nucleotide differences between the genes. Because all the differences occurred at the third base position of a codon and did not change the predicted amino acid sequence, we concluded that both beta-tubulin genes code for the same protein of 443 amino acids. The predicted amino acid sequence is 89 and 72% homologous with beta-tubulins from chicken and yeast cells, respectively. Each gene had three intervening sequences, which occurred at identical positions. Although the first two intervening sequences were not conserved between the two genes, the nucleotide sequence of the third intervening sequence was 89% conserved between the genes. The codon usage in the tubulin genes of C. reinhardtii was very biased: only 37 different codons were used. Striking differences occurred between the codons used in these nuclear genes and C. reinhardtii chloroplast genes.

Amino Acid Sequence↗

Relating physicochemical properties of amino acids to variable nucleotide substitution patterns among sites.

Markov-process models of codon substitution were implemented that account for features of DNA sequence evolution (such as transition/transversion bias and codon usage bias) as well as heterogeneity of amino acid substitution pattern over sites. The codon (amino acid) sites are assumed to come from several classes (such as secondary structure categories), among which the rate of amino acid substitution and the effect of amino acid chemical properties vary. Parameters are estimated by the maximum likelihood method, which accounts for the phylogenetic relationship among species and corrects for multiple hits at the same site. The likelihood ratio test is used to compare models. Mitochondrial cytochrome b genes of 28 primate species are analyzed. The site-heterogeneity models provide much better fit to previous homogeneous models.

Amino Acid Substitution↗

[Characteristics of the context shift in the frequency of synonymic codons in Escherichia coli].

We have demonstrated, that coding regions of E. coli DNA exhibit the non-random shifts of codon usage frequency depending both on the type of the 3' nucleotide adjacent to the codon and the degree of gene expression. Analysis of primary data--statistics of tetranucleotide occurrences--was performed by the techniques of contingency tables. The results of the investigation allowed us to suggest that the phenomenon observed is connected with the influence of the 3'-adjacent nucleotide on the level of missens-errors and another types of inaccuracy during translation. The specific advantages of such mutations in the third position of the codon are based on the adaptation of the codon to the 3' context in order to increase the efficiency of translation.

Base Sequence↗

Organization and structure of Volvox alpha-tubulin genes.

Southern analysis of Volvox genomic DNA revealed two genes homologous to Chlamydomonas reinhardtii alpha-tubulin cDNA. Restriction fragment length polymorphism analysis indicated that the two genes are not genetically linked. Clones representing one of the alpha-tubulin genes have been isolated from a genomic library of Volvox carteri f. nagariensis. A 3153 bp BamHI fragment containing the entire alpha-tubulin gene (1802 bp) plus 707 bp of the 5'- and 644 bp of the 3'-untranslated regions has been sequenced, revealing the following features: (1) the derived alpha-tubulin primary structure of 451 amino acids is highly conserved, differing in two residues from the alpha 1- and in two additional residues from the alpha 2-tubulin of C. reinhardtii; (2) in comparison to the C. reinhardtii genes, the Volvox alpha-tubulin gene contains a third intron; positions of the other two introns are precisely conserved; (3) codon usages are biased towards G or C, and against A, in the third position; 19 codons are absent from the alpha-tubulin coding sequence, and 5 of these are not used in any of 7 compiled Volvox genes; (4) transcription begins with an A, 30 bp downstream of the putative TATA box; upstream of the TATA box is a 14 bp sequence similar to consensus sequences found in all 4 C. reinhardtii tubulin genes and believed to regulate promoter function.

Amino Acid Sequence↗

Mutational analysis of the Streptomyces scabies esterase signal peptide.

Ten site-directed mutations affecting the predicted 39-amino-acid signal peptide of the Streptomyces scabies esterase were used to examine start-codon usage and esterase secretion in S. lividans. The first of two in-frame AUG codons was preferred for translation initiation. Removal of 2 of the 4 positively charged amino acids at the amino terminus of the signal peptide reduced esterase expression more than 100-fold; however, deletion of all 4 charged residues reduced expression by only 2- to 5-fold. Deletion of 4 or 8 amino acids from the hydrophobic core of the signal peptide reduced esterase production more than 200-fold, and a signal peptide processing site deletion completely disrupted esterase expression. For all constructs in which a mutation in the signal sequence decreased esterase production, esterase mRNA levels were also reduced, suggesting that a defect in secretion or processing affected esterase transcript abundance.

Amino Acid Sequence↗

Transcriptional and translational regulation of the expression of the l(2)gl tumor suppressor gene of Drosophila melanogaster.

By structural, biochemical and molecular genetic analyses, we have investigated the different mechanisms that control the expression of the lethal(2) giant larvae gene, a tumor suppressor gene of Drosophila melanogaster. Transcription of the l(2)gl gene is controlled by two highly identical promoters that result from the duplication of the 2.8 kb proximal portion of the gene. These two repeats are 96% homologous. Reverse genetic analysis has shown that each promoter can drive gene expression. In addition to the promoters, both repeats express two or three exons according to the pattern of splicing. The most distal exon in the second repeat is required because it contains the ATG initiating codon at the beginning of the open reading frame. The 3' untranslated region appears to contain motifs that specifically destabilize the transcript. Deletion of this region results in the formation of more stable mRNAs. The l(2)gl gene is characterized by an unusual codon usage that may reflect an enhanced translation efficiency by moderating the strength of pairing between codons and anticodons and may therefore increase the expressivity of this gene. Analysis of the spatio-temporal expression of the l(2)gl transcripts and proteins has shown that transcripts and proteins are produced ubiquitously during early embryogenesis, at a time when expression of the gene is required for preventing tumorigenesis. In the second half of embryogenesis, l(2)gl expression becomes restricted to tissues that do not show any phenotypic alteration in mutant animals. The l(2)gl protein exhibits two distinct intracellular localizations. It is preferentially found free in the cytoplasm but can become associated with the inner face of the plasma membrane where it is restricted to domains facing contiguous cells. In particular, the l(2)gl protein is absent from the basal and apical domains of the plasma membrane. The aim of the current research is directed towards understanding the functional relevance of the l(2)gl protein binding to the plasma membrane and its role in the control of cell proliferation and differentiation.

Animals↗

Three stages in the evolution of the genetic code.

A diversification of the genetic code based on the number of codons available for the proteinous amino acids is established. Three groups of amino acids during evolution of the code are distinguished. On the basis of their chemical complexity those amino acids emerging later in a translation process are derived. Codon number and chemical complexity indicate that His, Phe, Tyr, Cys and either Lys or Asn were introduced in the second stage, whereas the number of codons alone gives evidence that Trp and Met were introduced in the third stage. The amino acids of stage 1 use purine-rich codons, while all the amino acids introduced in the second stage, in contrast, use pyrimidines in the third position of their codons. A low abundance of pyrimidines during early translation is derived. This assumption is supported by experiments on non-enzymatic replication and interactions of hairpin loops with a complementary strand. A back extrapolation concludes a high purine content of the first nucleic acids, which gradually decreased during their evolution. Amino acids independently available from prebiotic synthesis were thus correlated to purine-rich codons. Implications on the prebiotic replication are discussed also in the light of recent codon usage data.

Amino Acids↗

Cloning and sequencing of a Moraxella bovis pilin gene.

Moraxella bovis pili have been shown to play a major role in both infectivity and protective immunity of bovine infectious keratoconjunctivitis. Sonicated M. bovis DNA from the piliated strain EPP63 was inserted into the vector lambda gt11 with EcoRI linkers. Recombinant phage were screened with an oligonucleotide probe based on the amino-terminal portion of the DNA sequence of a Neisseria gonorrhoeae pilin gene. Two candidate phages produced a protein that comigrated with EPP63 beta pilin in sodium dodecyl sulfate-polyacrylamide gels and bound anti-pilus antisera. The 1.9-kilobase insert from one of these, lambda gt11M182, was subcloned in both orientations into pBR322, forming the plasmids pMxB7 and pMxB9, both of which produced beta pilin, as did pMxB12, a HindIII deletion derivative of pMxB7. In HB101(pMxB12), the M. bovis pilin protein was shown to be primarily localized in the inner membrane. The entire 939-base-pair insert of pMxB12 was sequenced, revealing a ribosome binding site just upstream of the coding region and an AT-rich region further upstream containing some potential RNA polymerase recognition sites. The translation of the sequence predicts a six-amino-acid leader sequence preceding the phenylalanine that begins the mature protein. Codon usage analysis of the M. bovis beta pilin gene revealed greater use of the CUA codon for leucine than usual for a well-expressed Escherichia coli gene. Comparisons of the M. bovis EPP63 beta pilin protein sequence with other pilin gene sequences are presented.

Amino Acid Sequence↗

Characterization of a bacteriocinogenic plasmid from Clostridium perfringens and molecular genetic analysis of the bacteriocin-encoding gene.

The bacteriocinogenic plasmid pIP404 from Clostridium perfringens was isolated and cloned in Escherichia coli, and its physical map was deduced. Expression of the bcn gene, encoding bacteriocin BCN5, is inducible by UV irradiation of C. perfringens and thus resembles the SOS-regulated bacteriocin genes of enteric bacteria. The location of bcn on pIP404 was established by a dot-blot procedure, using specific hybridization probes to analyze mRNA samples from induced and uninduced cultures. From the nucleotide sequence of its gene, the molecular weight of BCN5 was deduced to be 96,591, and a protein of this size was secreted by bacteriocin-producing cultures of C. perfringens. The primary structure of the protein suggests that it may function as an ionophore, since a hydrophobic domain, resembling those of the ionophoric colicins, is present at the COOH terminus. No bacteriocin activity could be detected in E. coli harboring plasmids bearing the bcn gene, even when the transcriptional and translational signals were replaced by those of lacZ. A possible explanation may be found in the unusual codon usage of the adenine-thymine-rich bcn gene, as this shows a preference for codons with a high adenine-plus-thymine content, especially in the wobble position. Many of the frequently used codons correspond to those recognized by minor tRNA species in E. coli. Consequently, bcn expression might be limited by tRNA availability in this bacterium.

Amino Acid Sequence↗

Identification and nucleotide sequence of the Leptospira biflexa serovar patoc trpE and trpG genes.

Leptospira biflexa is a representative of an evolutionarily distinct group of eubacteria. In order to better understand the genetic organization and gene regulatory mechanisms of this species, we have chosen to study the genes required for tryptophan biosynthesis in this bacterium. The nucleotide sequence of the region of the L. biflexa serovar patoc chromosome encoding the trpE and trpG genes has been determined. Four open reading frames (ORFs) were identified in this region, but only three ORFs were translated into proteins when the cloned genes were introduced into Escherichia coli. Analysis of the predicted amino acid sequences of the proteins encoded by the ORFs allowed us to identify the trpE and trpG genes of L. biflexa. Enzyme assays confirmed the identity of these two ORFs. Anthranilate synthase from L. biflexa was found to be subject to feedback inhibition by tryptophan. Codon usage analysis showed that there was a bias in L. biflexa towards the use of codons rich in A and T, as would be expected from its G + C content of 37%. Comparison of the amino acid sequences of the trpE gene product and the trpG gene product with corresponding gene products from other bacteria showed regions of highly conserved sequence.

Amino Acid Sequence↗

Structure and evolution of bacterial adenylate cyclase: comparison between Escherichia coli and Erwinia chrysanthemi.

The cya genes, coding for adenylate cyclase, from Escherichia coli and Erwinia chrysanthemi B374 are compared after determination of a 3632 bp long nucleotide sequence of the hemC-cya region of E. chrysanthemi, encompassing the whole cya gene. In spite of a large divergence between the two organisms, especially visible in non coding regions, the amino acid sequence of the proteins are very similar, except at the very distal carboxyl end. Codon usage is different in the two organisms, and E. chrysanthemi tends to restrict translation to codons ending in G or C. Conservation of the translation initiation start region (including the poor ribosome binding site GGCG, and the TTG start codon), suggests that a specific protein synthesis process controls adenylate cyclase expression. Finally a palindromic unit, of primary sequence differing from the E. coli counterpart, borders the gene in E. chrysanthemi.

Adenylyl Cyclases↗