Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Chromosomal localization of the human hexabrachion (tenascin) gene and evidence for recent reduplication within the gene.

Using analysis of rodent-human somatic cell hybrids as well as in situ hybridization of hexabrachion cDNA probes to normal human metaphase chromosomes, we have localized the human hexabrachion gene to chromosome 9, bands q32-q34. We also put forward the hypothesis that there has been a recent reduplication of a small segment of the human hexabrachion gene. We support this hypothesis by comparison of codon usage in this segment of the gene to codon usage in the remainder of the gene. This hypothesis is also supported by comparison of the sequence of human hexabrachion to that of the chicken hexabrachion. In addition, the latter comparison shows that the reduplication most likely occurred after the divergence of mammalian and avian species.

Amino Acid Sequence↗

Synonymous codon preferences in bacteriophage T4: a distinctive use of transfer RNAs from T4 and from its host Escherichia coli.

Codon usage data of bacteriophage T4 genes were compiled and synonymous codon preferences were investigated in comparison with tRNA availabilities in an infected cell. Since the genome of T4 is highly AT rich and its codon usage pattern is significantly different from that of its host Escherichia coli, certain codons of T4 genes need to be translated by appropriate host transfer RNAs present in minor amounts. To avoid this predicament, T4 phage seems to direct the synthesis of its own tRNA molecules and these phage tRNAs are suggested to supplement the host tRNA population with isoacceptors that are normally present in minor amounts. A positive correlation was found in that the frequency of E. coli optimal codons in T4 genes increases as the number of protein monomers per phage particle increases. A negative correlation was also found between the number of protein monomers per phage and the frequency of "T4 optimal codons", which are defined as those codons that are efficiently recognized by T4 tRNAs. From these observations it was proposed that tRNAs from the host are predominantly used for translation of highly expressed T4 genes while tRNAs from T4 tend to be used for translation of weakly expressed T4 genes. This distinctive tRNA-usage in T4 may be an optimization of translational efficiency, and an adjustment of T4-encoded tRNAs to the synonymous codon preferences, which are largely influenced by the high genomic AT-content, would have occurred during evolution.

Bacteriophage T4↗

[Regularities of the nucleotide sequence at the 5'-end of the codon in Escherichia coli genes].

The frequencies of occurrence of nucleotides at the 5' side of codons have been determined in highly and weakly expressed genes from E. coli. Significant constraints on the nucleotide 5' to some codons were found in highly expressed genes. Certain rules of synonymous codon usage depending on the amino acid 3' of the codon were established. E. g., codon possessing quanosine in the third position (NNG) are preferred over NNA if the next amino acid is lysine (P less than 10(-5)). On the other hand, rules of synonymous codon usage in relation to 5' flanking nucleotide were found. For example, when coding for aspartic acid, GAC codon is preferred over GAU (P less than 0.001) if uridine is 5' to codon and on the contrary GAU is favoured (P less than 0.0001) if quanosine is at the 5' side of aspartic acid codon. These rules can be used in the chemical synthesis of genes designed for expression in E. coli.

Base Sequence↗

The contributions of replication orientation, gene direction, and signal sequences to base-composition asymmetries in bacterial genomes.

Asymmetries in base composition between the leading and the lagging strands have been observed previously in many prokaryotic genomes. Since a majority of genes is encoded on the leading strand in these genomes, previous analyses have not been able to determine the relative contribution to the base composition skews of replication processes and transcriptional and/or translational forces. Using qualitative graphical presentations and quantitative statistical analyses (analysis of variance), we have found that a significant proportion of the GC and AT skews can be attributed to replication orientation, i.e., the sequence of a gene is influenced by whether it is encoded on the leading or lagging strand. This effect of replication orientation on skews is independent of, and can be opposite in sign to, the effects of transcriptional or translational processes, such as selection for codon usage, amino acid preferences, expression levels (inferred from codon adaptation index), or potential short signal sequences (e.g., chi sequences). Mutational differences between the leading and the lagging strands are the most likely explanation for a significant proportion of the base composition skew in these bacterial genomes. The finding that base composition skews due to replication orientation are independent of those due to selection for function of the encoded protein may complicate the interpretation of phylogenetic relationships, conserved positions in nucleotide or amino acid sequence alignments, and codon usage patterns.

Analysis of Variance↗

Detecting anomalous gene clusters and pathogenicity islands in diverse bacterial genomes.

A gene in a genome is defined as putative alien (pA) if its codon usage difference from the average gene exceeds a high threshold and codon usage differences from ribosomal protein genes, chaperone genes and protein-synthesis-processing factors are also high. pA gene clusters in bacterial genomes are relevant for detecting genomic islands (GIs), including pathogenicity islands (PAIs). Four other analyses appropriate to this task are G+C genome variation (the standard method); genomic signature divergences (dinucleotide bias); extremes of codon bias; and anomalies of amino acid usage. For example, the cagA domain of Helicobacter pylori is highly deviant in its genome signature and codon bias from the rest of the genome. Using these methods we can detect two potential PAIs in the Neisseria meningitidis genome, which contain hemagglutinin and/or hemolysin-related genes. Additionally, G+C variation and genome signature differences of the Mycobacterium tuberculosis genome indicate two pA gene clusters.

Bacteria↗

Rare codons in E. coli and S. typhimurium signal sequences.

Codon usage has been examined in the signal sequences of 27 genes encoding proteins which possess leader peptides, and are inner-membrane located or exported. The results have been compared with codon usage in the corresponding coding sequences of most of the mature proteins. A bias is observed in the usage of rare codons for two of the three hydrophobic amino acids for which there are rare codons. Since hydrophobic residues are predominant in leader peptides, we suggest that a resulting concentration of rare codons in the signal sequence may play a role (or have played a role in the evolutionary past) in the secretion process by delaying translation.

Base Sequence↗

DNA Translator and Aligner: HyperCard utilities to aid phylogenetic analysis of molecules.

DNA Translator and Aligner are molecular phylogenetics HyperCard stacks for Macintosh computers. They manipulate sequence data to provide graphical gene mapping, conversions, translations and manual multiple-sequence alignment editing. DNA Translator is able to convert documented GenBank or EMBL documented sequences into linearized, rescalable gene maps whose gene sequences are extractable by clicking on the corresponding map button or by selection from a scrolling list. Provided gene maps, complete with extractable sequences, consist of nine metazoan, one yeast, and one ciliate mitochondrial DNAs and three green plant chloroplast DNAs. Single or multiple sequences can be manipulated to aid in phylogenetic analysis. Sequences can be translated between nucleic acids and proteins in either direction with flexible support of alternate genetic codes and ambiguous nucleotide symbols. Multiple aligned sequence output from diverse sources can be converted to Nexus, Hennig86 or PHYLIP format for subsequent phylogenetic analysis. Input or output alignments can be examined with Aligner, a convenient accessory stack included in the DNA Translator package. Aligner is an editor for the manual alignment of up to 100 sequences that toggles between display of matched characters and normal unmatched sequences. DNA Translator also generates graphic displays of amino acid coding and codon usage frequency relative to all other, or only synonymous, codons for approximately 70 select organism-organelle combinations. Codon usage data is compatible with spreadsheet or UWGCG formats for incorporation of additional molecules of interest. The complete package is available via anonymous ftp and is free for non-commercial uses.

Amino Acid Sequence↗

"Silent" sites in Drosophila genes are not neutral: evidence of selection among synonymous codons.

The patterns of synonymous codon usage in 91 Drosophila melanogaster genes have been examined. Codon usage varies strikingly among genes. This variation is associated with differences in G+C content at silent sites, but (unlike the situation in mammalian genes) these differences are not correlated with variation in intron base composition and so are not easily explicable in terms of mutational biases. Instead, those genes with high G+C content at silent sites, resulting from a strong "preference" for a particular subset of the codons that are mostly C-ending, appear to be the more highly expressed genes. This suggests that G+C content is reduced in sequences where selective constraints are weaker, as indeed seen in a pseudogene. These and other data discussed are consistent with the effects of translational selection among synonymous codons, as seen in unicellular organisms. The existence of selective constraints on silent substitutions, which may vary in strength among genes, has implications for the use of silent molecular clocks.

Animals↗

Cloning and hemolysin-mediated secretory expression of a codon-optimized synthetic human interleukin-6 gene in Escherichia coli.

Previously, we constructed human interleukin-6 (hIL-6)-secreting Escherichia coli and Salmonella typhimurium strains by fusion of the hIL-6 cDNA to the HlyA(s) secretional signal, utilizing the hemolysin export apparatus for extracellular delivery of a bioactive hIL-6-hemolysin (hIL-6-HlyA(s)) fusion protein. Molecular analysis of the secretion process revealed that low secretion levels were due to inefficient gene expression. To adapt the codon usage in hIL-6 cDNA to the E. coli codon bias, a synthetic hIL-6Ec gene variant was constructed from 20 overlapping oligonucleotides, yielding a 561-bp fragment, which comprises the complete hIL-6 cDNA sequence. Genetic fusion of the hIL-6Ec gene with the hlyA(s) secretional signal as an integral part of the hemolysin operon resulted in 3-fold higher hIL-6-HlyA(s) secretion levels in E. coli, compared to a strain expressing the original hIL-6-hlyA(s) fusion gene. An increase in the electrophoretic mobility of secreted hIL-6-HlyA(s) in non-reducing SDS-PAGE, similar to that found for recombinant mature hIL-6, and the absence of such a mobility shift in the intracellular hIL-6-HlyA(s) protein fraction indicated that in hIL-6-HlyA(s) most probably correct intramolecular disulfide bond formation occurred during the secretion step. To confirm the disulfide bond formation, hIL-6-HlyA(s) was purified by a single-step immunoaffinity chromatography from culture supernatant in yields of 18 microg/L culture supernatant with purity in the range of 60%. These results demonstrate that codon usage has an impact on the hemolysin-mediated secretion of hIL-6 and, furthermore, provide evidence that the hemolysin system enables secretory delivery of disulfide-bridged proteins.

Amino Acid Sequence↗

Molecular characterization of the tdc operon of Escherichia coli K-12.

The nucleotide sequence of a 2-kilobase DNA fragment of the tdc region of Escherichia coli K-12, previously cloned in this laboratory, revealed two open reading frames, tdcC and ORFX, downstream from the tdcB gene (formerly designated tdc) encoding biodegradative threonine dehydratase. A 24-base-pair sequence separated tdcC from the dehydratase coding region, and an untranslated region of 60 nucleotides, which contains a recognizable -10 consensus sequence, was found between tdcC and ORFX. The deduced amino acid sequence of tdcC showed it to be a large hydrophobic polypeptide of 431 amino acid residues, whereas ORFX coded for a small 135-residue polypeptide lacking glutamine and tryptophan. A computer-assisted sequence analysis revealed no similarity among the tdcB, tdcC, and ORFX polypeptides, and a search of the GenBank database failed to detect similarity with any other known proteins. The tdc genes and ORFX showed similar codon usage and, in analogy with other bacterial genes, showed codon usage typical for genes expressed at an intermediate level. Transcriptional analysis with S1 nuclease indicated two distinct transcription start sites upstream of the tdcB gene in regions previously identified as promoterlike elements P1 and P2. Interestingly, expression of tdcB and tdcC, but not ORFX, was contingent upon the presence of P1. These results taken together tend to suggest that the biodegradative threonine dehydratase is the second gene in a polycistronic transcription unit constituting a novel operon (tdcABC) in E. coli implicated in anaerobic threonine metabolism.

Amino Acid Sequence↗

Causal analysis of CpG suppression in the Mycoplasma genome.

Some bacterial genomes are known to have low CpG dinucleotide frequencies. While their causes are not clearly understood, the frequency of CpG is suppressed significantly in the genome of Mycoplasma genitalium, but not in that of Mycoplasma pneumoniae. We compared orthologous gene pairs of the two closely related species to analyze CpG substitution patterns between these two genomes. We also divided genome sequences into three regions: protein-coding, noncoding, and RNA-coding, and obtained the CpG frequencies for each region for each organism. It was found that the observed/expected ratio of CpG dinucleotides is low in both the protein-coding and noncoding regions; while that ratio is in the normal range in the RNA-coding region. Our results indicate that CpG suppression of the Mycoplasma genome is not caused by (1) biased usage amino acid; (2) biased usage of synonymous codon; or (3) methylation effects by the CpG methyltransferase in the genomes of their hosts. Instead, we consider it likely that a certain global pressure, such as genome-wide pressure for the advantages of DNA stability or replication, has the effect of decreasing CpG over the entire genome, which, in turn, resulted in the biased codon usage.

Base Composition↗

Organization and expression of algal (Chlamydomonas reinhardtii) mitochondrial DNA.

The mitochondrial genome of Chlamydomonas reinhardtii, a unicellular green alga, is a linear 15.8 kilobase pair (kbp) molecule. In gene arrangement and mode of expression, as well as in size, it differs radically from the large (200-2400 kbp) mitochondrial genomes of higher plants. Heterologous hybridization experiments and nucleotide sequence analysis have revealed that C. reinhardtii mitochondrial DNA (mtDNA) is a compactly organized genome specifying at least eight proteins, a minimum of three transfer RNAs, and large subunit (LS) and small subunit (SS) ribosomal RNAs. Both strands of the mtDNA encode genetic information, with genes organized into perhaps a single transcriptional unit on each strand. Stable transcripts have been identified by Northern hybridization analysis, and transcript termini have been mapped by primer extension and S1 nuclease protection experiments. The results suggest that mature RNAs, which virtually saturate the genome, are generated by precise endonucleolytic cleavage of long precursors, with specific motifs (both primary sequence and secondary structure) implicated as processing signals. Codon usage in C. reinhardtii mitochondria is highly biased, with eight codons entirely absent from all protein-coding genes; however, even though codon usage is restricted, it appears that C. reinhardtii mtDNA cannot encode the minimum number of tRNAs needed to support mitochondrial protein synthesis. The most striking feature of C. reinhardtii mtDNA is the division of SS and LS rRNA genes into a number of separate subgenic coding segments ('modules') that are interspersed with one another and with protein-coding and tRNA genes. We have identified abundant small RNAs, transcribed from these modules, that approximate to the latter in size. This indicates that splicing of rRNA 'pieces' does not occur in this system. Rather, the mature rRNAs apparently exist and function as non-covalent complexes of small RNAs (four in SS rRNA, at least eight in LS rRNA), held together by intermolecular base pairing. These complexes contain all the conserved elements of the minimal secondary structures that define the functional core of conventional LS and SS rRNAs.

Base Sequence↗

Optimizing heterologous expression in dictyostelium: importance of 5' codon adaptation.

Expression of heterologous proteins in Dictyostelium discoideum presents unique research opportunities, such as the functional analysis of complex human glycoproteins after random mutagenesis. In one study, human chorionic gonadotropin (hCG) and human follicle stimulating hormone were expressed in Dictyostelium. During the course of these experiments, we also investigated the role of codon usage and of the DNA sequence upstream of the ATG start codon. The Dictyostelium genome has a higher AT content than the human, resulting in a different codon preference. The hCG-beta gene contains three clusters with infrequently used codons that were changed to codons that are preferred by Dictyostelium. The results reported here show that optimizing the first 5-17 codons of the hCG gene contributes to 4- to 5-fold increased expression levels, but that further optimization has no significant effect. These observations suggest that optimal codon usage contributes to ribosome stabilization, but does not play an important role during the elongation phase of translation. Furthermore, adapting the 5'-sequence of the hCG gene to the Dictyostelium 'Kozak'-like sequence increased expression levels approximately 1.5-fold. Thus, using both codon optimization and 'Kozak' adaptation, a 6- to 8-fold increase in expression levels could be obtained for hCG.

Amino Acid Sequence↗

The Vitreoscilla hemoglobin gene: molecular cloning, nucleotide sequence and genetic expression in Escherichia coli.

Vitreoscilla hemoglobin is involved in oxygen metabolism of this bacterium, possibly in an unusual role for a microbe. We have isolated the Vitreoscilla hemoglobin structural gene from a pUC19 genomic library using mixed oligodeoxy-nucleotide probes based on the reported amino acid sequence of the protein. The gene is expressed in Escherichia coli from its natural promoter as a major cellular protein. The nucleotide sequence, which is in complete agreement with the known amino acid sequence of the protein, suggests the existence of promoter and ribosome binding sites with a high degree of homology to consensus E. coli upstream sequences. In the case of at least some amino acids, a codon usage bias can be detected which is different from the biased codon usage pattern in E. coli. The downstream sequence exhibits homology with the 3' end sequences of several plant leghemoglobin genes. E. coli cells expressing the gene contain greater than fivefold more heme than controls.

Amino Acid Sequence↗

Molecular cloning and sequence analysis of the gene of the molybdenum-containing aldehyde oxido-reductase of Desulfovibrio gigas. The deduced amino acid sequence shows similarity to xanthine dehydrogenase.

In this report, we describe the isolation of a 4020-bp genomic PstI fragment of Desulfovibrio gigas harboring the aldehyde oxido-reductase gene. The aldehyde oxido-reductase gene spans 2718 bp of genomic DNA and codes for a protein with 906 residues. The protein sequence shows an average 52% (+/- 1.5%) similarity to xanthine dehydrogenase from different organisms. The codon usage of the aldehyde oxidoreductase is almost identical to a calculated codon usage of the Desulfovibrio bacteria.

Aldehyde Oxidoreductases↗

[Analysis of apolipoprotein gene family in codon space--non-random selection of nucleotide changes in evolution].

The choice of nucleotide changes in DNA evolution can be either selectively neutral or biased. To study how apolipoprotein gene selects the nucleotide substitutions in the course of evolution, a codon space is constructed in which its DNA sequence can be mapped as a matrix of nucleotide frequencies in three codon positions. Accordingly, a number of methods that measure the nonrandomness of nucleotide distribution in codon space are developed based on maximum entropy techniques to define the nature of nucleotide change selection in evolution. By these methods, we demonstrated that the nucleotide composition in 1st and 3rd codon position of apolipoprotein genes is highly nonrandom, which appears to be a result of non-neutral selection of codon positions by adenosine and thymidine. In addition, this paper is also concerned in the divergence of synonymos codon usage and its correlation to taxonomic distances among species. As a result, a codon usage clock was reported in apolipoprotein A-I. Our studies suggest that non-random selection of nucleotide changes in codon space may represent an evolutionary characteristics of apolipoprotein genes.

Animals↗

A DnaB intein in Rhodothermus marinus: indication of recent intein homing across remotely related organisms.

A dnaB gene encoding a homologue of the Escherichia coli DNA helicase DnaB was cloned and sequenced in the thermophilic eubacterium Rhodothermus marinus, predicting a DnaB protein that harbors an intein. This DnaB intein is 428 amino acid residues long, has several putative intein sequence motifs (including two putative endonuclease motifs), and is capable of protein splicing when produced in E. coli cells. The R. marinus DnaB intein is a close homologue of a DnaB intein in the cyanobacterium Synechocystis sp. strain PCC6803. The two inteins are positioned identically in their respective DnaB proteins. They also share a 54% sequence identity (74% sequence similarity) that is markedly higher than the 37% sequence identity shared by the extein sequences of the two DnaB proteins. Horizontal intein transfer (homing) is therefore invoked to relate these two DnaB inteins. The codon usage of R. marinus DnaB intein coding sequence differs markedly from the codon usages of its flanking extein coding sequences and other genes in the same genome, suggesting more recent acquisition of the DnaB intein in this organism.

Amino Acid Sequence↗

How mitochondria redefine the code.

Annotated, complete DNA sequences are available for 213 mitochondrial genomes from 132 species. These provide an extensive sample of evolutionary adjustment of codon usage and meaning spanning the history of this organelle. Because most known coding changes are mitochondrial, such data bear on the general mechanism of codon reassignment. Coding changes have been attributed variously to loss of codons due to changes in directional mutation affecting the genome GC content (Osawa and Jukes 1988), to pressure to reduce the number of mitochondrial tRNAs to minimize the genome size (Anderson and Kurland 1991), and to the existence of transitional coding mechanisms in which translation is ambiguous (Schultz and Yarus 1994a). We find that a succession of such steps explains existing reassignments well. In particular, (1) Genomic variation in the prevalence of a codon's third-position nucleotide predicts relative mitochondrial codon usage well, though GC content does not. This is because A and T, and G and C, are uncorrelated in mitochondrial genomes. (2) Codons predicted to reach zero usage (disappear) do so more often than expected by chance, and codons that do disappear are disproportionately likely to be reassigned. However, codons predicted to disappear are not significantly more likely to be reassigned. Therefore, low codon frequencies can be related to codon reassignment, but appear to be neither necessary nor sufficient for reassignment. (3) Changes in the genetic code are not more likely to accompany smaller numbers of tRNA genes and are not more frequent in smaller genomes. Thus, mitochondrial codons are not reassigned during demonstrable selection for decreased genome size. Instead, the data suggest that both codon disappearance and codon reassignment depend on at least one other event. This mitochondrial event (leading to reassignment) occurs more frequently when a codon has disappeared, and produces only a small subset of possible reassignments. We suggest that coding ambiguity, the extension of a tRNA's decoding capacity beyond its original set of codons, is the second event. Ambiguity can act alone but often acts in concert with codon disappearance, which promotes codon reassignment.

Base Composition↗