Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

DNA Translator and Aligner: HyperCard utilities to aid phylogenetic analysis of molecules.

DNA Translator and Aligner are molecular phylogenetics HyperCard stacks for Macintosh computers. They manipulate sequence data to provide graphical gene mapping, conversions, translations and manual multiple-sequence alignment editing. DNA Translator is able to convert documented GenBank or EMBL documented sequences into linearized, rescalable gene maps whose gene sequences are extractable by clicking on the corresponding map button or by selection from a scrolling list. Provided gene maps, complete with extractable sequences, consist of nine metazoan, one yeast, and one ciliate mitochondrial DNAs and three green plant chloroplast DNAs. Single or multiple sequences can be manipulated to aid in phylogenetic analysis. Sequences can be translated between nucleic acids and proteins in either direction with flexible support of alternate genetic codes and ambiguous nucleotide symbols. Multiple aligned sequence output from diverse sources can be converted to Nexus, Hennig86 or PHYLIP format for subsequent phylogenetic analysis. Input or output alignments can be examined with Aligner, a convenient accessory stack included in the DNA Translator package. Aligner is an editor for the manual alignment of up to 100 sequences that toggles between display of matched characters and normal unmatched sequences. DNA Translator also generates graphic displays of amino acid coding and codon usage frequency relative to all other, or only synonymous, codons for approximately 70 select organism-organelle combinations. Codon usage data is compatible with spreadsheet or UWGCG formats for incorporation of additional molecules of interest. The complete package is available via anonymous ftp and is free for non-commercial uses.

Amino Acid Sequence↗

"Silent" sites in Drosophila genes are not neutral: evidence of selection among synonymous codons.

The patterns of synonymous codon usage in 91 Drosophila melanogaster genes have been examined. Codon usage varies strikingly among genes. This variation is associated with differences in G+C content at silent sites, but (unlike the situation in mammalian genes) these differences are not correlated with variation in intron base composition and so are not easily explicable in terms of mutational biases. Instead, those genes with high G+C content at silent sites, resulting from a strong "preference" for a particular subset of the codons that are mostly C-ending, appear to be the more highly expressed genes. This suggests that G+C content is reduced in sequences where selective constraints are weaker, as indeed seen in a pseudogene. These and other data discussed are consistent with the effects of translational selection among synonymous codons, as seen in unicellular organisms. The existence of selective constraints on silent substitutions, which may vary in strength among genes, has implications for the use of silent molecular clocks.

Animals↗

Cloning and hemolysin-mediated secretory expression of a codon-optimized synthetic human interleukin-6 gene in Escherichia coli.

Previously, we constructed human interleukin-6 (hIL-6)-secreting Escherichia coli and Salmonella typhimurium strains by fusion of the hIL-6 cDNA to the HlyA(s) secretional signal, utilizing the hemolysin export apparatus for extracellular delivery of a bioactive hIL-6-hemolysin (hIL-6-HlyA(s)) fusion protein. Molecular analysis of the secretion process revealed that low secretion levels were due to inefficient gene expression. To adapt the codon usage in hIL-6 cDNA to the E. coli codon bias, a synthetic hIL-6Ec gene variant was constructed from 20 overlapping oligonucleotides, yielding a 561-bp fragment, which comprises the complete hIL-6 cDNA sequence. Genetic fusion of the hIL-6Ec gene with the hlyA(s) secretional signal as an integral part of the hemolysin operon resulted in 3-fold higher hIL-6-HlyA(s) secretion levels in E. coli, compared to a strain expressing the original hIL-6-hlyA(s) fusion gene. An increase in the electrophoretic mobility of secreted hIL-6-HlyA(s) in non-reducing SDS-PAGE, similar to that found for recombinant mature hIL-6, and the absence of such a mobility shift in the intracellular hIL-6-HlyA(s) protein fraction indicated that in hIL-6-HlyA(s) most probably correct intramolecular disulfide bond formation occurred during the secretion step. To confirm the disulfide bond formation, hIL-6-HlyA(s) was purified by a single-step immunoaffinity chromatography from culture supernatant in yields of 18 microg/L culture supernatant with purity in the range of 60%. These results demonstrate that codon usage has an impact on the hemolysin-mediated secretion of hIL-6 and, furthermore, provide evidence that the hemolysin system enables secretory delivery of disulfide-bridged proteins.

Amino Acid Sequence↗

Molecular characterization of the tdc operon of Escherichia coli K-12.

The nucleotide sequence of a 2-kilobase DNA fragment of the tdc region of Escherichia coli K-12, previously cloned in this laboratory, revealed two open reading frames, tdcC and ORFX, downstream from the tdcB gene (formerly designated tdc) encoding biodegradative threonine dehydratase. A 24-base-pair sequence separated tdcC from the dehydratase coding region, and an untranslated region of 60 nucleotides, which contains a recognizable -10 consensus sequence, was found between tdcC and ORFX. The deduced amino acid sequence of tdcC showed it to be a large hydrophobic polypeptide of 431 amino acid residues, whereas ORFX coded for a small 135-residue polypeptide lacking glutamine and tryptophan. A computer-assisted sequence analysis revealed no similarity among the tdcB, tdcC, and ORFX polypeptides, and a search of the GenBank database failed to detect similarity with any other known proteins. The tdc genes and ORFX showed similar codon usage and, in analogy with other bacterial genes, showed codon usage typical for genes expressed at an intermediate level. Transcriptional analysis with S1 nuclease indicated two distinct transcription start sites upstream of the tdcB gene in regions previously identified as promoterlike elements P1 and P2. Interestingly, expression of tdcB and tdcC, but not ORFX, was contingent upon the presence of P1. These results taken together tend to suggest that the biodegradative threonine dehydratase is the second gene in a polycistronic transcription unit constituting a novel operon (tdcABC) in E. coli implicated in anaerobic threonine metabolism.

Amino Acid Sequence↗

Causal analysis of CpG suppression in the Mycoplasma genome.

Some bacterial genomes are known to have low CpG dinucleotide frequencies. While their causes are not clearly understood, the frequency of CpG is suppressed significantly in the genome of Mycoplasma genitalium, but not in that of Mycoplasma pneumoniae. We compared orthologous gene pairs of the two closely related species to analyze CpG substitution patterns between these two genomes. We also divided genome sequences into three regions: protein-coding, noncoding, and RNA-coding, and obtained the CpG frequencies for each region for each organism. It was found that the observed/expected ratio of CpG dinucleotides is low in both the protein-coding and noncoding regions; while that ratio is in the normal range in the RNA-coding region. Our results indicate that CpG suppression of the Mycoplasma genome is not caused by (1) biased usage amino acid; (2) biased usage of synonymous codon; or (3) methylation effects by the CpG methyltransferase in the genomes of their hosts. Instead, we consider it likely that a certain global pressure, such as genome-wide pressure for the advantages of DNA stability or replication, has the effect of decreasing CpG over the entire genome, which, in turn, resulted in the biased codon usage.

Base Composition↗

Organization and expression of algal (Chlamydomonas reinhardtii) mitochondrial DNA.

The mitochondrial genome of Chlamydomonas reinhardtii, a unicellular green alga, is a linear 15.8 kilobase pair (kbp) molecule. In gene arrangement and mode of expression, as well as in size, it differs radically from the large (200-2400 kbp) mitochondrial genomes of higher plants. Heterologous hybridization experiments and nucleotide sequence analysis have revealed that C. reinhardtii mitochondrial DNA (mtDNA) is a compactly organized genome specifying at least eight proteins, a minimum of three transfer RNAs, and large subunit (LS) and small subunit (SS) ribosomal RNAs. Both strands of the mtDNA encode genetic information, with genes organized into perhaps a single transcriptional unit on each strand. Stable transcripts have been identified by Northern hybridization analysis, and transcript termini have been mapped by primer extension and S1 nuclease protection experiments. The results suggest that mature RNAs, which virtually saturate the genome, are generated by precise endonucleolytic cleavage of long precursors, with specific motifs (both primary sequence and secondary structure) implicated as processing signals. Codon usage in C. reinhardtii mitochondria is highly biased, with eight codons entirely absent from all protein-coding genes; however, even though codon usage is restricted, it appears that C. reinhardtii mtDNA cannot encode the minimum number of tRNAs needed to support mitochondrial protein synthesis. The most striking feature of C. reinhardtii mtDNA is the division of SS and LS rRNA genes into a number of separate subgenic coding segments ('modules') that are interspersed with one another and with protein-coding and tRNA genes. We have identified abundant small RNAs, transcribed from these modules, that approximate to the latter in size. This indicates that splicing of rRNA 'pieces' does not occur in this system. Rather, the mature rRNAs apparently exist and function as non-covalent complexes of small RNAs (four in SS rRNA, at least eight in LS rRNA), held together by intermolecular base pairing. These complexes contain all the conserved elements of the minimal secondary structures that define the functional core of conventional LS and SS rRNAs.

Base Sequence↗

Optimizing heterologous expression in dictyostelium: importance of 5' codon adaptation.

Expression of heterologous proteins in Dictyostelium discoideum presents unique research opportunities, such as the functional analysis of complex human glycoproteins after random mutagenesis. In one study, human chorionic gonadotropin (hCG) and human follicle stimulating hormone were expressed in Dictyostelium. During the course of these experiments, we also investigated the role of codon usage and of the DNA sequence upstream of the ATG start codon. The Dictyostelium genome has a higher AT content than the human, resulting in a different codon preference. The hCG-beta gene contains three clusters with infrequently used codons that were changed to codons that are preferred by Dictyostelium. The results reported here show that optimizing the first 5-17 codons of the hCG gene contributes to 4- to 5-fold increased expression levels, but that further optimization has no significant effect. These observations suggest that optimal codon usage contributes to ribosome stabilization, but does not play an important role during the elongation phase of translation. Furthermore, adapting the 5'-sequence of the hCG gene to the Dictyostelium 'Kozak'-like sequence increased expression levels approximately 1.5-fold. Thus, using both codon optimization and 'Kozak' adaptation, a 6- to 8-fold increase in expression levels could be obtained for hCG.

Amino Acid Sequence↗

The Vitreoscilla hemoglobin gene: molecular cloning, nucleotide sequence and genetic expression in Escherichia coli.

Vitreoscilla hemoglobin is involved in oxygen metabolism of this bacterium, possibly in an unusual role for a microbe. We have isolated the Vitreoscilla hemoglobin structural gene from a pUC19 genomic library using mixed oligodeoxy-nucleotide probes based on the reported amino acid sequence of the protein. The gene is expressed in Escherichia coli from its natural promoter as a major cellular protein. The nucleotide sequence, which is in complete agreement with the known amino acid sequence of the protein, suggests the existence of promoter and ribosome binding sites with a high degree of homology to consensus E. coli upstream sequences. In the case of at least some amino acids, a codon usage bias can be detected which is different from the biased codon usage pattern in E. coli. The downstream sequence exhibits homology with the 3' end sequences of several plant leghemoglobin genes. E. coli cells expressing the gene contain greater than fivefold more heme than controls.

Amino Acid Sequence↗

Molecular cloning and sequence analysis of the gene of the molybdenum-containing aldehyde oxido-reductase of Desulfovibrio gigas. The deduced amino acid sequence shows similarity to xanthine dehydrogenase.

In this report, we describe the isolation of a 4020-bp genomic PstI fragment of Desulfovibrio gigas harboring the aldehyde oxido-reductase gene. The aldehyde oxido-reductase gene spans 2718 bp of genomic DNA and codes for a protein with 906 residues. The protein sequence shows an average 52% (+/- 1.5%) similarity to xanthine dehydrogenase from different organisms. The codon usage of the aldehyde oxidoreductase is almost identical to a calculated codon usage of the Desulfovibrio bacteria.

Aldehyde Oxidoreductases↗

[Analysis of apolipoprotein gene family in codon space--non-random selection of nucleotide changes in evolution].

The choice of nucleotide changes in DNA evolution can be either selectively neutral or biased. To study how apolipoprotein gene selects the nucleotide substitutions in the course of evolution, a codon space is constructed in which its DNA sequence can be mapped as a matrix of nucleotide frequencies in three codon positions. Accordingly, a number of methods that measure the nonrandomness of nucleotide distribution in codon space are developed based on maximum entropy techniques to define the nature of nucleotide change selection in evolution. By these methods, we demonstrated that the nucleotide composition in 1st and 3rd codon position of apolipoprotein genes is highly nonrandom, which appears to be a result of non-neutral selection of codon positions by adenosine and thymidine. In addition, this paper is also concerned in the divergence of synonymos codon usage and its correlation to taxonomic distances among species. As a result, a codon usage clock was reported in apolipoprotein A-I. Our studies suggest that non-random selection of nucleotide changes in codon space may represent an evolutionary characteristics of apolipoprotein genes.

Animals↗

A DnaB intein in Rhodothermus marinus: indication of recent intein homing across remotely related organisms.

A dnaB gene encoding a homologue of the Escherichia coli DNA helicase DnaB was cloned and sequenced in the thermophilic eubacterium Rhodothermus marinus, predicting a DnaB protein that harbors an intein. This DnaB intein is 428 amino acid residues long, has several putative intein sequence motifs (including two putative endonuclease motifs), and is capable of protein splicing when produced in E. coli cells. The R. marinus DnaB intein is a close homologue of a DnaB intein in the cyanobacterium Synechocystis sp. strain PCC6803. The two inteins are positioned identically in their respective DnaB proteins. They also share a 54% sequence identity (74% sequence similarity) that is markedly higher than the 37% sequence identity shared by the extein sequences of the two DnaB proteins. Horizontal intein transfer (homing) is therefore invoked to relate these two DnaB inteins. The codon usage of R. marinus DnaB intein coding sequence differs markedly from the codon usages of its flanking extein coding sequences and other genes in the same genome, suggesting more recent acquisition of the DnaB intein in this organism.

Amino Acid Sequence↗

How mitochondria redefine the code.

Annotated, complete DNA sequences are available for 213 mitochondrial genomes from 132 species. These provide an extensive sample of evolutionary adjustment of codon usage and meaning spanning the history of this organelle. Because most known coding changes are mitochondrial, such data bear on the general mechanism of codon reassignment. Coding changes have been attributed variously to loss of codons due to changes in directional mutation affecting the genome GC content (Osawa and Jukes 1988), to pressure to reduce the number of mitochondrial tRNAs to minimize the genome size (Anderson and Kurland 1991), and to the existence of transitional coding mechanisms in which translation is ambiguous (Schultz and Yarus 1994a). We find that a succession of such steps explains existing reassignments well. In particular, (1) Genomic variation in the prevalence of a codon's third-position nucleotide predicts relative mitochondrial codon usage well, though GC content does not. This is because A and T, and G and C, are uncorrelated in mitochondrial genomes. (2) Codons predicted to reach zero usage (disappear) do so more often than expected by chance, and codons that do disappear are disproportionately likely to be reassigned. However, codons predicted to disappear are not significantly more likely to be reassigned. Therefore, low codon frequencies can be related to codon reassignment, but appear to be neither necessary nor sufficient for reassignment. (3) Changes in the genetic code are not more likely to accompany smaller numbers of tRNA genes and are not more frequent in smaller genomes. Thus, mitochondrial codons are not reassigned during demonstrable selection for decreased genome size. Instead, the data suggest that both codon disappearance and codon reassignment depend on at least one other event. This mitochondrial event (leading to reassignment) occurs more frequently when a codon has disappeared, and produces only a small subset of possible reassignments. We suggest that coding ambiguity, the extension of a tRNA's decoding capacity beyond its original set of codons, is the second event. Ambiguity can act alone but often acts in concert with codon disappearance, which promotes codon reassignment.

Base Composition↗

Structure and expression of the gene encoding the periplasmic arylsulfatase of Chlamydomonas reinhardtii.

Chlamydomonas reinhardtii produces a periplasmic arylsulfatase in response to sulfur deprivation. We have isolated and sequenced arylsulfatase cDNAs from a lambda gt11 expression library. The amino acid sequence of the protein, as deduced from the nucleotide sequence, has features characteristic of secreted proteins, including a signal sequence and putative glycosylation sites. The gene has a broad codon usage with seven codons, all having A residues in the third position, not previously observed in C. reinhardtii genes. Arylsulfatase transcription is tightly regulated by sulfur availability. The approximately 2.7 kb arylsulfatase transcript is very susceptible to degradation, disappearing in less than an hour after sulfur starved cells are administered either sulfate or alpha-amanitin. The accumulation of the arylsulfatase transcript is also suppressed by the addition of cycloheximide. Transcription initiation from the arylsulfatase gene occurs approximately 100 bp upstream of the initiation codon, in a region that is 5' to a 43 bp imperfect inverted repeat. Preceding the transcription start site are sequences similar to those present in promoter regions of other genes from C. reinhardtii.

Amino Acid Sequence↗

Development of Polymorphic EST Markers Suitable for Genetic Linkage Mapping of Catfish.

: Expressed sequence tag (EST) markers are important for gene mapping and for marker-assisted selection (MAS). To develop EST markers for use in catfish gene mapping, 100 randomly picked complementary DNAs from the channel catfish (Ictalurus punctatus) pituitary library were sequenced. The EST sequences were used to design primers to amplify channel catfish and blue catfish (I. furcatus) genomic DNAs. Polymerase chain reaction products of the ESTs were analyzed to determine length polymorphism between the channel catfish and blue catfish. Eleven polymorphic EST markers were identified. Five of the 11 EST markers were from known genes and the other six were from unidentified ESTs. Seven ESTs were found to be associated with microsatellite sequences. Analysis of channel catfish gene sequences indicated highly biased codon usage, with 16 codons being preferably used. These codons were more preferably used in highly expressed ribosomal protein genes and in highly expressed pituitary hormone genes. G/C-rich codons are less used in channel catfish than those in other vertebrates suggesting AT-richness of the channel catfish genome.

Journal Article↗

Catalyzing bacterial speciation: correlating lateral transfer with genetic headroom.

Unlike crown eukaryotic species, microbial species are created by continual processes of gene loss and acquisition promoted by horizontal genetic transfer. The amounts of foreign DNA in bacterial genomes, and the rate at which this is acquired, are consistent with gene transfer as the primary catalyst for microbial differentiation. However, the rate of successful gene transfer varies among bacterial lineages. The heterogeneity in foreign DNA content is directly correlated with amount of genetic headroom intrinsic to a bacterial species. Genetic headroom reflects the amount of potentially dispensable information--reflected in codon usage bias and codon context bias--that can be transiently sacrificed to allow experimentation with functions introduced by gene transfer. In this way, genetic headroom offers a potential metric for assessing the propensity of a lineage to speciate.

Bacteria↗

Molecular population genetics and evolution of a prion-like protein in Saccharomyces cerevisiae.

The prion-like behavior of Sup35p, the eRF3 homolog in the yeast Saccharomyces cerevisiae, mediates the activity of the cytoplasmic nonsense suppressor known as [PSI(+)]. Sup35p is divided into three regions of distinct function. The N-terminal and middle (M) regions are required for the induction and propagation of [PSI(+)] but are not necessary for translation termination or cell viability. The C-terminal region encompasses the termination function. The existence of the N-terminal region in SUP35 homologs of other fungi has led some to suggest that this region has an adaptive function separate from translation termination. To examine this hypothesis, we sequenced portions of SUP35 in 21 strains of S. cerevisiae, including 13 clinical isolates. We analyzed nucleotide polymorphism within this species and compared it to sequence divergence from a sister species, S. paradoxus. The N domain of Sup35p is highly conserved in amino acid sequence and is highly biased in codon usage toward preferred codons. Amino acid changes are under weak purifying selection based on a quantitative analysis of polymorphism and divergence. We also conclude that the clinical strains of S. cerevisiae are not recently derived and that outcrossing between strains in S. cerevisiae may be relatively rare in nature.

Amino Acid Sequence↗

Suppression of the negative effect of minor arginine codons on gene expression; preferential usage of minor codons within the first 25 codons of the Escherichia coli genes.

AGA and AGG codons for arginine are the least used codons in Escherichia coli, which are encoded by a rare tRNA, the product of the dnaY gene. We examined the positions of arginine residues encoded by AGA/AGG codons in 678 E. coli proteins. It was found that AGA/AGG codons appear much more frequently within the first 25 codons. This tendency becomes more significant in those proteins containing only one AGA or AGG codon. Other minor codons such as CUA, UCA, AGU, ACA, GGA, CCC and AUA are also found to be preferentially used within the first 25 codons. The effects of the AGG codon on gene expression were examined by inserting one to five AGG codons after the 10th codon from the initiation codon of the lacZ gene. The production of beta-galactosidase decreased as more AGG codons were inserted. With five AGG codons, the production of beta-galactosidase (Gal-AGG5) completely ceased after a mid-log phase of cell growth. After 22 hr induction of the lacZ gene, the overall production of Gal-AGG5 was 11% of the control production (no insertion of arginine codons). When five CGU codons, the major arginine codon were inserted instead of AGG, the production of beta-galactosidase (Gal-CGU5) continued even after stationary phase and the overall production was 66% of the control. The negative effect of the AGG codons on the Gal-AGG5 production was found to be dependent upon the distance between the site of the AGG codons and the initiation codon. As the distance was increased by inserting extra sequences between the two codons, the production of Gal-AGG5 increased almost linearly up to 8 fold. From these results, we propose that the position of the minor codons in an mRNA plays an important role in the regulation of gene expression possibly by modulating the stability of the initiation complex for protein synthesis.

Amino Acid Sequence↗

The complete mitochondrial genome of Tupaia belangeri and the phylogenetic affiliation of scandentia to other eutherian orders.

The complete mitochondrial genome of Tupaia belangeri, a representative of the eutherian order Scandentia, was determined and compared with full-length mitochondrial sequences of other eutherian orders described to date. The complete mitochondrial genome is 16, 754 nt in length, with no obvious deviation from the general organization of the mammalian mitochondrial genome. Thus, features such as start codon usage, incomplete stop codons, and overlapping coding regions, as well as the presence of tandem repeats in the control region, are within the range of mammalian mitochondrial (mt) DNA variation. To address the question of a possible close phylogenetic relationship between primates and Tupaia, the evolutionary affinities among primates, Tupaia and bats as representatives of the Archonta superorder, ferungulates, guinea pigs, armadillos, rats, mice, and hedgehogs were examined on the basis of the complete mitochondrial DNA sequences. The opossum sequence was used as an outgroup. The trees, estimated from 12 concatenated genes encoded on the mitochondrial H-strand, add further molecular evidence against an Archonta monophyly. With the new data described in this paper, most of both the mitochondrial and the nuclear data point away from Scandentia as the closest extant relatives to primates. Instead, the complete mitochondrial data support a clustering of Scandentia with Lagomorpha connecting to the branch leading to ferungulates. This closer phylogenetic relationship of Tupaia to rabbits than to primates first received support from several analyses of nuclear and partial mitochondrial DNA data sets. Given that short sequences are of limited use in determining deep mammalian relationships, the partial mitochondrial data available to date support this hypothesis only tentatively. Our complete mitochondrial genome data therefore add considerably more evidence in support of this hypothesis.

Animals↗