Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Rare codons in E. coli and S. typhimurium signal sequences.

Codon usage has been examined in the signal sequences of 27 genes encoding proteins which possess leader peptides, and are inner-membrane located or exported. The results have been compared with codon usage in the corresponding coding sequences of most of the mature proteins. A bias is observed in the usage of rare codons for two of the three hydrophobic amino acids for which there are rare codons. Since hydrophobic residues are predominant in leader peptides, we suggest that a resulting concentration of rare codons in the signal sequence may play a role (or have played a role in the evolutionary past) in the secretion process by delaying translation.

Base Sequence

DNA Translator and Aligner: HyperCard utilities to aid phylogenetic analysis of molecules.

DNA Translator and Aligner are molecular phylogenetics HyperCard stacks for Macintosh computers. They manipulate sequence data to provide graphical gene mapping, conversions, translations and manual multiple-sequence alignment editing. DNA Translator is able to convert documented GenBank or EMBL documented sequences into linearized, rescalable gene maps whose gene sequences are extractable by clicking on the corresponding map button or by selection from a scrolling list. Provided gene maps, complete with extractable sequences, consist of nine metazoan, one yeast, and one ciliate mitochondrial DNAs and three green plant chloroplast DNAs. Single or multiple sequences can be manipulated to aid in phylogenetic analysis. Sequences can be translated between nucleic acids and proteins in either direction with flexible support of alternate genetic codes and ambiguous nucleotide symbols. Multiple aligned sequence output from diverse sources can be converted to Nexus, Hennig86 or PHYLIP format for subsequent phylogenetic analysis. Input or output alignments can be examined with Aligner, a convenient accessory stack included in the DNA Translator package. Aligner is an editor for the manual alignment of up to 100 sequences that toggles between display of matched characters and normal unmatched sequences. DNA Translator also generates graphic displays of amino acid coding and codon usage frequency relative to all other, or only synonymous, codons for approximately 70 select organism-organelle combinations. Codon usage data is compatible with spreadsheet or UWGCG formats for incorporation of additional molecules of interest. The complete package is available via anonymous ftp and is free for non-commercial uses.

Amino Acid Sequence

"Silent" sites in Drosophila genes are not neutral: evidence of selection among synonymous codons.

The patterns of synonymous codon usage in 91 Drosophila melanogaster genes have been examined. Codon usage varies strikingly among genes. This variation is associated with differences in G+C content at silent sites, but (unlike the situation in mammalian genes) these differences are not correlated with variation in intron base composition and so are not easily explicable in terms of mutational biases. Instead, those genes with high G+C content at silent sites, resulting from a strong "preference" for a particular subset of the codons that are mostly C-ending, appear to be the more highly expressed genes. This suggests that G+C content is reduced in sequences where selective constraints are weaker, as indeed seen in a pseudogene. These and other data discussed are consistent with the effects of translational selection among synonymous codons, as seen in unicellular organisms. The existence of selective constraints on silent substitutions, which may vary in strength among genes, has implications for the use of silent molecular clocks.

Animals

Molecular characterization of the tdc operon of Escherichia coli K-12.

The nucleotide sequence of a 2-kilobase DNA fragment of the tdc region of Escherichia coli K-12, previously cloned in this laboratory, revealed two open reading frames, tdcC and ORFX, downstream from the tdcB gene (formerly designated tdc) encoding biodegradative threonine dehydratase. A 24-base-pair sequence separated tdcC from the dehydratase coding region, and an untranslated region of 60 nucleotides, which contains a recognizable -10 consensus sequence, was found between tdcC and ORFX. The deduced amino acid sequence of tdcC showed it to be a large hydrophobic polypeptide of 431 amino acid residues, whereas ORFX coded for a small 135-residue polypeptide lacking glutamine and tryptophan. A computer-assisted sequence analysis revealed no similarity among the tdcB, tdcC, and ORFX polypeptides, and a search of the GenBank database failed to detect similarity with any other known proteins. The tdc genes and ORFX showed similar codon usage and, in analogy with other bacterial genes, showed codon usage typical for genes expressed at an intermediate level. Transcriptional analysis with S1 nuclease indicated two distinct transcription start sites upstream of the tdcB gene in regions previously identified as promoterlike elements P1 and P2. Interestingly, expression of tdcB and tdcC, but not ORFX, was contingent upon the presence of P1. These results taken together tend to suggest that the biodegradative threonine dehydratase is the second gene in a polycistronic transcription unit constituting a novel operon (tdcABC) in E. coli implicated in anaerobic threonine metabolism.

Amino Acid Sequence

Organization and expression of algal (Chlamydomonas reinhardtii) mitochondrial DNA.

The mitochondrial genome of Chlamydomonas reinhardtii, a unicellular green alga, is a linear 15.8 kilobase pair (kbp) molecule. In gene arrangement and mode of expression, as well as in size, it differs radically from the large (200-2400 kbp) mitochondrial genomes of higher plants. Heterologous hybridization experiments and nucleotide sequence analysis have revealed that C. reinhardtii mitochondrial DNA (mtDNA) is a compactly organized genome specifying at least eight proteins, a minimum of three transfer RNAs, and large subunit (LS) and small subunit (SS) ribosomal RNAs. Both strands of the mtDNA encode genetic information, with genes organized into perhaps a single transcriptional unit on each strand. Stable transcripts have been identified by Northern hybridization analysis, and transcript termini have been mapped by primer extension and S1 nuclease protection experiments. The results suggest that mature RNAs, which virtually saturate the genome, are generated by precise endonucleolytic cleavage of long precursors, with specific motifs (both primary sequence and secondary structure) implicated as processing signals. Codon usage in C. reinhardtii mitochondria is highly biased, with eight codons entirely absent from all protein-coding genes; however, even though codon usage is restricted, it appears that C. reinhardtii mtDNA cannot encode the minimum number of tRNAs needed to support mitochondrial protein synthesis. The most striking feature of C. reinhardtii mtDNA is the division of SS and LS rRNA genes into a number of separate subgenic coding segments ('modules') that are interspersed with one another and with protein-coding and tRNA genes. We have identified abundant small RNAs, transcribed from these modules, that approximate to the latter in size. This indicates that splicing of rRNA 'pieces' does not occur in this system. Rather, the mature rRNAs apparently exist and function as non-covalent complexes of small RNAs (four in SS rRNA, at least eight in LS rRNA), held together by intermolecular base pairing. These complexes contain all the conserved elements of the minimal secondary structures that define the functional core of conventional LS and SS rRNAs.

Base Sequence

The Vitreoscilla hemoglobin gene: molecular cloning, nucleotide sequence and genetic expression in Escherichia coli.

Vitreoscilla hemoglobin is involved in oxygen metabolism of this bacterium, possibly in an unusual role for a microbe. We have isolated the Vitreoscilla hemoglobin structural gene from a pUC19 genomic library using mixed oligodeoxy-nucleotide probes based on the reported amino acid sequence of the protein. The gene is expressed in Escherichia coli from its natural promoter as a major cellular protein. The nucleotide sequence, which is in complete agreement with the known amino acid sequence of the protein, suggests the existence of promoter and ribosome binding sites with a high degree of homology to consensus E. coli upstream sequences. In the case of at least some amino acids, a codon usage bias can be detected which is different from the biased codon usage pattern in E. coli. The downstream sequence exhibits homology with the 3' end sequences of several plant leghemoglobin genes. E. coli cells expressing the gene contain greater than fivefold more heme than controls.

Amino Acid Sequence

Structure and expression of the gene encoding the periplasmic arylsulfatase of Chlamydomonas reinhardtii.

Chlamydomonas reinhardtii produces a periplasmic arylsulfatase in response to sulfur deprivation. We have isolated and sequenced arylsulfatase cDNAs from a lambda gt11 expression library. The amino acid sequence of the protein, as deduced from the nucleotide sequence, has features characteristic of secreted proteins, including a signal sequence and putative glycosylation sites. The gene has a broad codon usage with seven codons, all having A residues in the third position, not previously observed in C. reinhardtii genes. Arylsulfatase transcription is tightly regulated by sulfur availability. The approximately 2.7 kb arylsulfatase transcript is very susceptible to degradation, disappearing in less than an hour after sulfur starved cells are administered either sulfate or alpha-amanitin. The accumulation of the arylsulfatase transcript is also suppressed by the addition of cycloheximide. Transcription initiation from the arylsulfatase gene occurs approximately 100 bp upstream of the initiation codon, in a region that is 5' to a 43 bp imperfect inverted repeat. Preceding the transcription start site are sequences similar to those present in promoter regions of other genes from C. reinhardtii.

Amino Acid Sequence

Suppression of the negative effect of minor arginine codons on gene expression; preferential usage of minor codons within the first 25 codons of the Escherichia coli genes.

AGA and AGG codons for arginine are the least used codons in Escherichia coli, which are encoded by a rare tRNA, the product of the dnaY gene. We examined the positions of arginine residues encoded by AGA/AGG codons in 678 E. coli proteins. It was found that AGA/AGG codons appear much more frequently within the first 25 codons. This tendency becomes more significant in those proteins containing only one AGA or AGG codon. Other minor codons such as CUA, UCA, AGU, ACA, GGA, CCC and AUA are also found to be preferentially used within the first 25 codons. The effects of the AGG codon on gene expression were examined by inserting one to five AGG codons after the 10th codon from the initiation codon of the lacZ gene. The production of beta-galactosidase decreased as more AGG codons were inserted. With five AGG codons, the production of beta-galactosidase (Gal-AGG5) completely ceased after a mid-log phase of cell growth. After 22 hr induction of the lacZ gene, the overall production of Gal-AGG5 was 11% of the control production (no insertion of arginine codons). When five CGU codons, the major arginine codon were inserted instead of AGG, the production of beta-galactosidase (Gal-CGU5) continued even after stationary phase and the overall production was 66% of the control. The negative effect of the AGG codons on the Gal-AGG5 production was found to be dependent upon the distance between the site of the AGG codons and the initiation codon. As the distance was increased by inserting extra sequences between the two codons, the production of Gal-AGG5 increased almost linearly up to 8 fold. From these results, we propose that the position of the minor codons in an mRNA plays an important role in the regulation of gene expression possibly by modulating the stability of the initiation complex for protein synthesis.

Amino Acid Sequence

Primary structure of the tolC gene that codes for an outer membrane protein of Escherichia coli K12.

We present the nucleotide sequence of the tolC gene of Escherichia coli K12, and the amino acid sequence of the TolC protein (an outer membrane protein) as deduced from it. The mature TolC protein comprises 467 amino acid residues, and, as previously reported (1), a signal sequence of 22 amino acid residues is attached to the N-terminus. The C-terminus of the gene is followed by a stem-loop structure (8 base pair stem, 4 base loop) which may be a rho-independent termination signal. The codon usage of the gene is nonrandom; the major isoaccepting species of tRNA are preferentially utilised, or, among synonomous codons recognized by the same tRNA, those codons are used which can interact better with the anticodon (2,3). In contrast to the codon usage for other outer membrane proteins of E. coli (4) the rare arginine codons AGA and AGG are used once and twice respectively.

Bacterial Outer Membrane Proteins

[Role of the code redundancy in determining cotranslational protein folding].

It has been demonstrated earlier in our laboratory that rare codon clusters can determine the boundaries of the polypeptide chain fragments of the same secondary structure type during the co-translational protein folding. According to this data, co-translational protein folding can occur under condition of a correlation between the frequency of codon choice in mRNAs and the relative abundance of their isoaccepting tRNAs. The alterations in the spectrum and concentrations of the isoaccepting tRNAs in different cells were demonstrated by many authors. The existence of a mechanism of the coordinate regulation of the levels (activities) of the isoaccepting tRNAs, corresponding aminoacyl-tRNA synthetases and mRNAs predominantly translated at a given moment of time can be suggested. Such a mechanism can ensure the needed accuracy of the protein folding process. Analysis of gene sequences of various pro- and eukaryotic organisms carried out in the present work revealed that the codon usage frequency spectra of simultaneously synthesized proteins are similar. The relative appearance of the most rare and frequent codons in investigated gene sequences displays a high degree of conservatism. It has also been found that structural-homologous proteins from different organisms (cytochromes c, myoglobins) have very similar codon frequency distribution profiles. This property retains despite the significant variations in the codon usage spectra in the investigated gene sequences. The data obtained indicate that the codon distribution in mRNAs whose diversity is mainly conditioned by the genetic code redundance is a program that determines translational rates of different mRNA parts thus controlling the spatial folding of the synthesized peptide chain.

Animals

Comparison of the nucleoside sequence of trpA and sequences immediately beyond the trp operon of Klebsiella aerogenes. Salmonella typhimurium and Escherichia coli.

The nucleotide sequence of trpA of Klebsiella aerogenes is presented and compared with the trpA sequences of Salmonella typhimurium and Escherichia coli. The majority of the approximately 200 differences between each pair of trpA's are single nucleotide pair changes that do not alter the amino acid sequence. Codon usage conforms to the general patterns revealed by examination of other prokaryotic gene sequences. However, codon usage in K. aerogenes trpA reflects the high G+C content of the genome of this organism. The DNA sequences just beyond trpA, the presumed transcription termination region, are also compared for the three species. Perusal of these sequences indicates that the secondary structure of the transcript segment just beyond trpA has been preserved, while the primary sequence has diverged appreciably.

Amino Acid Sequence

Complete nucleotide sequence and genetic organization of the bacteriocinogenic plasmid, pIP404, from Clostridium perfringens.

The complete nucleotide sequence of the bacteriocinogenic plasmid, pIP404, from Clostridium perfringens has been determined. The plasmid genome comprises 10,207 bp and has a dA + dT content of 75%. Functions have been tentatively assigned to 6 of the 10 open reading frames and an origin-like region of repeated sequence identified. The codon usage of this extremely dA + dT rich plasmid is highly unusual and displays a pronounced preference for codons with the lowest dG + dC content. Only one of the genes from pIP404 was expressed at a significant level in Escherichia coli, suggesting that the atypical codon usage could represent a major obstacle to heterologous gene expression.

Base Sequence

Preferential use of A- and U-rich codons for Mycoplasma capricolum ribosomal proteins S8 and L6.

The nucleotide sequence of the 1.3 kilobase-pair DNA segment, which contains the genes for ribosomal proteins S8 and L6, and a part of L18 of Mycoplasma capricolum, has been determined and compared with the corresponding sequence in Escherichia coli (Cerretti et al., Nucl. Acids Res. 11, 2599, 1983). Identities of the predicted amino acid sequences of S8 and L6 between the two organisms are 54% and 42%, respectively. The A + T content of the M. capricolum genes is 71%, which is much higher than that of E. coli (49%). Comparisons of codon usage between the two organisms have revealed that M. capricolum preferentially uses A- and U-rich codons. More than 90% of the codon third positions and 57% of the first positions in M. capricolum is either A or U, whereas E. coli uses A or U for the third and the first positions at a frequency of 51% and 36%, respectively. The biased choice of the A- and U-rich codons in this organism has been also observed in the codon replacements for conservative amino acid substitutions between M. capricolum and E. coli. These facts suggest that the codon usage of M. capricolum is strongly influenced by the high A + T content of the genome.

Adenine

The sequence of the chloroplast atpB gene and its flanking regions in Chlamydomonas reinhardtii.

The chloroplast (cp)-encoded CF1 ATPase beta-subunit gene (atpB) of Chlamydomonas reinhardtii and its flanking regions have been sequenced. The derived amino acid (aa) sequence is highly homologous to that of the beta-subunit gene in Escherichia coli, bovine heart mitochondria, and higher plant cp. In contrast to all other cp genomes, the CF1 epsilon subunit gene (atpE) does not lie at the 3' end of the atpB gene but maps to a position 92 kb away in the other single-copy region. Northern blots confirm that the beta subunit is not encoded as part of a dicistronic message as it is in higher plants. The region just upstream from the atpB gene in C. reinhardtii contains two small open reading frames (ORFs) and not the gene for the large subunit of ribulose-1,5-bisphosphate carboxylase/oxygenase as is found in cp genomes of higher plants. No transcripts for either ORF were detected, but the codon usage in these ORFs as well as in the atpB gene follows the unique pattern of codon usage previously seen in other cp genes in C. reinhardtii.

Amino Acid Sequence

The ribosomal protein S8 from Thermus thermophilus VK1. Sequencing of the gene, overexpression of the protein in Escherichia coli and interaction with rRNA.

The gene of the ribosomal protein S8 from Thermus thermophilus VK1 has been isolated from a genomic library by hybridization of an oligonucleotide coding for the N-terminal amino acid sequence of the protein, amplified by PCR and sequenced. Nucleotide sequence reveals an open reading frame coding for a protein of 138 amino acid residues (M(r) 15,839). The codon usage shows that 94% of the codons possess G or C in the third position, and agrees with the preferential usage of codons of high G+C content in the bacteria of the genus Thermus. The amino acid sequence of the protein shows 48% identity with the protein from Escherichia coli. Ribosomal protein S8 from T. thermophilus has been expressed in E. coli under the control of the T7 promoter and purified to homogeneity by heat treatment of the extract followed by cation-exchange chromatography. Conditions were defined in which T. thermophilus protein S8 binds specifically an homologous 16S rRNA fragment containing the putative S8 binding site with an apparent association constant of 5 x 10(7) M-1. The overexpressed protein binds the rRNA with the same affinity as that extracted from T. thermophilus, indicating that the thermophilic protein is correctly folded in E. coli. The specificity of this binding is dependent on the ionic strength. The protein S8 from T. thermophilus recognizes the E. coli rRNA binding sites as efficiently as the S8 protein from E. coli. This result agrees with sequence comparisons of the S8 binding site on the small subunit rRNA from E. coli and from T. thermophilus, showing strong similarities in the regions involved in the interaction. It suggests that the structural features responsible for the recognition are conserved in the mesophilic and thermophilic eubacteria, despite structural peculiarities in the thermophilic partners conferring thermostability.

Amino Acid Sequence

[Use of codons in plant lectins].

Codon usage in the coding region of mature lectins has been examined for 11 plant species (8 leguminoseae, 1 euphorbiaceae, 2 gramineae). The different legume lectins exhibit nearly the same codon usage pattern whereas the choice for the silent position of codons is non-random.

Codon

Ornithine decarboxylase and trypanothione reductase genes in Leishmania braziliensis guyanensis.

Ornithine decarboxylase and trypanothione reductase are the key enzymes in polyamine and trypanothione metabolism in kinetoplastids. Using a heterologous Trypanosoma brucei brucei probe for ornithine decarboxylase and a mixed synthetic probe of 29 oligonucleotides for trypanothione reductase, we have detected the putative genes for these enzymes by Southern blot hybridization using genomic DNA of Leishmania braziliensis guyanensis MHOM/SR/80/CUMC 1. The trypanothione reductase probe was constructed both from the conserved codon usage of the redox active site for other flavin oxidoreductases over a wide evolutionary scale, and the preferred codon usage for other genes in species of Leishmania.

Amino Acid Sequence

Contextual constraints on codon pair usage: structural and biological implications.

Complementary DNA sequence data of 278 protein coding genes from prokaryotic systems have been analysed at the level of near neighbour codon pairs. Our analysis points out that constraints exist even at the level of near neighbour codon pairs. These constraints are in addition to those which arise due to relative levels of tRNA. Codon pairs, which in the data base have different occurrence values from their expected values, neither have common secondary structure nor do have better stabilization due to high base stacking. Our study points out that there are strong interaction between constituent codons in these codon pairs. These strongly interacting codon pairs, we suggest, are involved in the formation of three dimensional structural elements of cDNA/mRNA and interact with ribosome and thus modulate translation.

Base Sequence