Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

Transcription of the triose-phosphate-isomerase gene of Schizosaccharomyces pombe initiates from a start point different from that in Saccharomyces cerevisiae.

Gene tpi, encoding the glycolytic enzyme triose phosphate isomerase (TPI) from the fission yeast Schizosaccharomyces pombe was cloned by complementation of a Saccharomyces cerevisiae tpil mutant. Nucleotide sequence analysis of the cloned gene revealed a single open reading frame (ORF) encoding a protein 59% homologous to S. cerevisiae TPI. The gene has a very high codon usage bias. Messenger RNA synthesis initiates at two points located 38 and 44 nucleotides downstream from a TATA box promoter sequence. In S. cerevisiae, transcription of this S. pombe gene initiates about 26 nucleotides downstream from the S. pombe start points. This observation indicates that the two yeasts have diverged in the mechanism which determines the 5' end of the messenger RNA relative to the TATA box. It appears that in some respects the transcription initiation mechanism of S. pombe more closely resembles that of higher eukaryotes than does the S. cerevisiae mechanism.

Amino Acid Sequence↗

Whole genome analysis reveals a high incidence of non-optimal codons in secretory signal sequences of Escherichia coli.

Translational pausing may occur due to a number of mechanisms, including the presence of non-optimal codons, and it is thought to play a role in the folding of specific polypeptide domains during translation and in the facilitation of signal peptide recognition during sec-dependent protein targeting. In this whole genome analysis of Escherichia coli we have found that non-optimal codons in the signal peptide-encoding sequences of secretory genes are overrepresented relative to the "mature" portions of these genes; this is in addition to their overrepresentation in the 5'-regions of genes encoding non-secretory proteins. We also find increased non-optimal codon usage at the 3' ends of most E. coli genes, in both non-secretory and secretory sequences. Whereas presumptive translational pausing at the 5' and 3' ends of E. coli messenger RNAs may clearly have a general role in translation, we suggest that it also has a specific role in sec-dependent protein export, possibly in facilitating signal peptide recognition. This finding may have important implications for our understanding of how the majority of non-cytoplasmic proteins are targeted, a process that is essential to all biological cells.

Codon↗

Physical map of alkaliphilic Bacillus firmus OF4 and detection of a large endogenous plasmid.

Extremely alkaliphilic Bacillus firmus OF4 is among the best characterized of this group of alkaliphiles. Together with alkaliphilic Bacillus C-125 and numerous non-alkaliphilic Bacillus species whose chromosomes and gene organizations are currently being studied in detail, work on B. firmus OF4 offers the opportunity to discern whether there are features of chromosome and gene organization that are associated with alkaliphily. A physical map of the B. firmus OF4 is consistent with a circular chromosome of approximately 4 Mb, with an extrachromosomal element of 110 kb also detected. The previously identified cadmium-resistance locus and transposition functions in B. firmus OF4 were localized to the extrachromosomal element, whose genes exhibit a slightly different pattern of codon usage from chromosomal genes. No clustering of genes thus far identified with roles in alkaliphily has been found. Direct repeat sequences (DRS) were previously reported upstream of a gene encoding a Na+/H+ antiporter that has a role in pH homeostasis. In the current analyses, these sequences were found to be present in multiple copies on the chromosome, most of which are present in one 920-kb fragment. Such sequences might play a role in DNA rearrangements that allow amplification of important genes in this region.

Bacillus↗

Rates and patterns of molecular evolution in inbred and outbred Arabidopsis.

The evolution of self-fertilization is associated with a large reduction in the effective rate of recombination and a corresponding decline in effective population size. If many spontaneous mutations are slightly deleterious, this shift in the breeding system is expected to lead to a reduced efficacy of natural selection and genome-wide changes in the rates of molecular evolution. Here, we investigate the effects of the breeding system on molecular evolution in the highly self-fertilizing plant Arabidopsis thaliana by comparing its coding and noncoding genomic regions with those of its close outcrossing relative, the self-incompatible A. lyrata. More distantly related species in the Brassicaceae are used as outgroups to polarize the substitutions along each lineage. In contrast to expectations, no significant difference in the rates of protein evolution is observed between selfing and outcrossing Arabidopsis species. Similarly, no consistent overall difference in codon bias is observed between the species, although for low-biased genes A. lyrata shows significantly higher major codon usage. There is also evidence of intron size evolution in A. thaliana, which has consistently smaller introns than its outcrossing congener, potentially reflecting directional selection on intron size. The results are discussed in the context of heterogeneity in selection coefficients across loci and the effects of life history and population structure on rates of molecular evolution. Using estimates of substitution rates in coding regions and approximate estimates of divergence and generation times, the genomic deleterious mutation rate (U) for amino acid substitutions in Arabidopsis is estimated to be approximately 0.2-0.6 per generation.

Arabidopsis↗

Rapid evolution of the plastid translational apparatus in a nonphotosynthetic plant: loss or accelerated sequence evolution of tRNA and ribosomal protein genes.

The vestigial plastid genome of Epifagus virginiana (beechdrops), a nonphotosynthetic parasitic flowering plant, is functional but lacks six ribosomal protein and 13 tRNA genes found in the chloroplast DNAs of photosynthetic flowering plants. Import of nuclear gene products is hypothesized to compensate for many of these losses. Codon usage and amino acid usage patterns in Epifagus plastic genes have not been affected by the tRNA gene losses, though a small shift in the base composition of the whole genome (toward A+T-richness) is apparent. The ribosomal protein and tRNA genes that remain have had a high rate of molecular evolution, perhaps due to relaxation of constraints on the translational apparatus. Despite the compactness and extensive gene loss, one translational gene (infA, encoding initiation factor 1) that is a pseudogene in tobacco has been maintained intact in Epifagus.

Amino Acid Sequence↗

Detection of genes with atypical nucleotide sequence in microbial genomes.

Along the gene, nucleotides in various codon positions tend to exert a slight but observable influence on the nucleotide choice at neighboring positions. Such context biases are different in different organisms and can be used as genomic signatures. In this paper, we will focus specifically on the dinucleotide composed of a third codon position nucleotide and its succeeding first position nucleotide. Using the 16 possible dinucleotide combinations, we calculate how well individual genes conform to the observed mean dinucleotide frequencies of an entire genome, forming a distance measure for each gene. It is found that genes from different genomes can be separated with a high degree of accuracy, according to these distance values. In particular, we address the problem of recent horizontal gene transfer, and how imported genes may be evaluated by their poor assimilation to the host's context biases. By concentrating on the third- and succeeding first position nucleotides, we eliminate most spurious contributions from codon usage and amino-acid requirements, focusing mainly on mutational effects. Since imported genes are expected to converge only gradually to genomic signatures, it is possible to question whether a gene present in only one of two closely related organisms has been imported into one organism or deleted in the other. Striking correlations between the proposed distance measure and poor homology are observed when Escherichia coli genes are compared to Salmonella typhi, indicating that sets of outlier genes in E. coli may contain a high number of genes that have been imported into E. coli, and not deleted in S. typhi.

Bacteria↗

Genomic heterogeneity of background substitutional patterns in Drosophila melanogaster.

Mutation is the underlying force that provides the variation upon which evolutionary forces can act. It is important to understand how mutation rates vary within genomes and how the probabilities of fixation of new mutations vary as well. If substitutional processes across the genome are heterogeneous, then examining patterns of coding sequence evolution without taking these underlying variations into account may be misleading. Here we present the first rigorous test of substitution rate heterogeneity in the Drosophila melanogaster genome using almost 1500 nonfunctional fragments of the transposable element DNAREP1_DM. Not only do our analyses suggest that substitutional patterns in heterochromatic and euchromatic sequences are different, but also they provide support in favor of a recombination-associated substitutional bias toward G and C in this species. The magnitude of this bias is entirely sufficient to explain recombination-associated patterns of codon usage on the autosomes of the D. melanogaster genome. We also document a bias toward lower GC content in the pattern of small insertions and deletions (indels). In addition, the GC content of noncoding DNA in Drosophila is higher than would be predicted on the basis of the pattern of nucleotide substitutions and small indels. However, we argue that the fast turnover of noncoding sequences in Drosophila makes it difficult to assess the importance of the GC biases in nucleotide substitutions and small indels in shaping the base composition of noncoding sequences.

Animals↗

Stability of structures of the epsilon subunit and terminator of thermophilic ATPase.

F1-type ATPase is the central enzyme for ATP synthesis in most organisms. Because of the extreme reconstitutability of thermophilic ATPase (TF1) and diversity of the minor subunits of F1 type ATPase, an operon coding for TF1 was isolated from DNA of thermophilic bacterium PS3, and its terminal region containing the epsilon subunit (TF1 epsilon) and terminator was sequenced. The primary structure of the epsilon subunit (Mr = 14 333) was deduced from the nucleotide sequence (396 base-pairs) and amino-acid sequence of its amino terminus. The conclusions drawn from the results are as follows. Homologies: TF1 epsilon shows only 6% homology with the epsilon subunits of eight species reported, but 50% homology with Escherichia coli epsilon and 41% with chloroplast. The residues having a tendency to form reverse turns (Gly, Pro and Tyr) and His are relatively well conserved. Unlike some F1 epsilon types TF1 epsilon has no ATPase inhibitor activity and is not homologous with ATPase inhibitor. TF1 epsilon is essential to connect F1 to F0, like the b subunit, and is weakly homologous with the b subunit of F0F1. The cause of 3 beta: 1 epsilon subunit stoichiometry: The ribosome binding sequence of TF1 epsilon is TAGGN7, which is incomplete compared with that of TF1 beta. The codon usage for TF1 epsilon is similar to that for TF1 epsilon. The cause of stability of TF1 epsilon and its gene: There are 18 ionic groups at the putative reverse turns and the N- and C-termini of TF1 epsilon, but only 10 ionic groups in the corresponding sites of E. coli epsilon subunit. These ionic groups enhance the external polarity of TF1 epsilon and may intensify subunit-subunit interaction. There is a terminator at the 3' end of the TF1 epsilon gene, which is stabilized by a long (13 base-pairs) stem.

Bacteria↗

The nucleotide sequence of POX18, a gene encoding a small oleate-inducible peroxisomal protein from Candida tropicalis.

We report the molecular cloning and nucleotide sequence of the nuclear gene, POX18, encoding an oleate-inducible peroxisomal protein from the yeast Candida tropicalis. POX18 has a single open reading frame of 381 nucleotides (nt), which encodes a protein of 127 amino acids. The predicted Mr of this protein is 13,792. Codon usage in the expression of POX18 is non-random, and shows a pattern similar to that used for other peroxisomal genes from C. tropicalis and highly expressed genes from Saccharomyces cerevisiae. Northern analysis of total RNA from oleate-grown cells determined that POX18 mRNA is approximately 750 nt in length. The POX18 gene was expressed in vitro, which resulted in a single translation product that co-migrated in denaturing polyacrylamide gels with an abundant peroxisomal protein (apparent mass of 16 kDa) and was immunoprecipitated by an antiserum against peroxisomal protein.

Amino Acid Sequence↗

The complete mitochondrial genomes of the sea lily Gymnocrinus richeri and the feather star Phanogenia gracilis: signature nucleotide bias and unique nad4L gene rearrangement within crinoids.

Complete DNA sequences have been determined for the mitochondrial genomes of the crinoids Phanogenia gracilis (15892 bp) and Gymnocrinus richeri (15966 bp). The mitochondrial genetic map of the stalkless feather star P. gracilis is identical to that of the comatulid feather star Florometra serratissima (Scouras, A., Smith, M.J., 2001. Mol. Biol. Evol. 18, 61-73). The mitochondrial gene order of the stalked crinoid G. richeri differs from that of F. serratissima and P. gracilis by the transposition of the nad4L protein gene. The G. richeri nad4L mitochondrial map position is unique among metazoa and is likely a derived feature in this stalked crinoid. Nucleotide compositional analyses of protein genes encoded on the major sense strand confirm earlier conclusions regarding a crinoid-distinctive T over C bias. All three crinoids exhibit high T levels in third codon positions, whereas other echinoderm classes favor A or C in the third codon position. The nucleotide bias is reflected in the relative synonymous codon usage patterns of crinoids versus other echinoderms. We suggest that the nucleotide bias of crinoids, in comparison to other echinoderms, indicates that a physical inversion of the origin of replication has occurred in the crinoid lineage. Evolutionary rate tests support the use of the cytochrome b (cob) gene in molecular phylogenetic analyses of echinoderms. A consensus echinoderm tree was generated based on cytochrome b nucleotide alignments that placed the asteroids as a sister group to a clade containing the ophiuroids and the (echinoids+holothuroids) with the crinoids basal to the rest of the echinoderm classes: [Crinoid,(Asteroid,(Ophiuroid,(Echinoid,Holothuroid)))].

Animals↗

SGP-1: prediction and validation of homologous genes based on sequence alignments.

Conventional methods of gene prediction rely on the recognition of DNA-sequence signals, the coding potential or the comparison of a genomic sequence with a cDNA, EST, or protein database. Reasons for limited accuracy in many circumstances are species-specific training and the incompleteness of reference databases. Lately, comparative genome analysis has attracted increasing attention. Several analysis tools that are based on human/mouse comparisons are already available. Here, we present a program for the prediction of protein-coding genes, termed SGP-1 (Syntenic Gene Prediction), which is based on the similarity of homologous genomic sequences. In contrast to most existing tools, the accuracy of depends little on species-specific properties such as codon usage or the nucleotide distribution. may therefore be applied to nonstandard model organisms in vertebrates as well as in plants, without the need for extensive parameter training. In addition to predicting genes in large-scale genomic sequences, the program may be useful to validate gene structure annotations from databases. To this end, SGP-1 output also contains comparisons between predicted and annotated gene structures in HTML format. The program can be accessed via a Web server at http://soft.ice.mpg.de/sgp-1. The source code, written in ANSI C, is available on request from the authors.

Algorithms↗

Molecular cloning and sequencing of the gene encoding the fimbrial subunit protein of Bacteroides gingivalis.

The gene encoding the fimbrial subunit protein of Bacteroides gingivalis 381, fimbrilin, has been cloned and sequenced. The gene was present as a single copy on the bacterial chromosome, and the codon usage in the gene conformed closely to that expected for an abundant protein. The predicted size of the mature protein was 35,924 daltons, and the secretory form may have had a 10-amino-acid, hydrophilic leader sequence similar to the leader sequences of the MePhe fimbriae family. The protein sequence had no marked similarity to known fimbrial sequences, and no homologous sequences could be found in other black-pigmented Bacteroides species, suggesting that fimbrillin represents a class of fimbrial subunit protein of limited distribution.

Amino Acid Sequence↗

Structure of the Schizosaccharomyces pombe cytochrome c gene.

The cytochrome c gene of the fission yeast Schizosaccharomyces pombe has been cloned by using the Saccharomyces cerevisiae iso-1-cytochrome c gene as a molecular hybridization probe. The DNA sequence and the 5' termini of the mRNA transcripts of the gene have been determined. The DNA sequence has confirmed, with two exceptions, the previously determined protein sequence. The nonrandom distribution of silent third base differences which was observed between the two cytochrome c genes of S. cerevisiae does not extend to the S. pombe cytochrome c gene, suggesting that there are no constraints other than protein function and codon usage which have acted to conserve the cytochrome DNA sequences of the two yeasts. Introduction of the S. pombe cytochrome c gene on a yeast plasmid into a S. cerevisiae mutant which lacked functional cytochrome c transformed that recipient strain for the ability to grow on a nonfermentable carbon source. This implies that the S. pombe cytochrome c gene has all the regulatory signals which are required for its expression in S. cerevisiae, and that none of the amino acid differences between the cytochrome c proteins of the two yeasts has a drastic effect on the function of the protein in vivo.

Ascomycota↗

The nucleotide sequence of the essential cell-division gene ftsZ of Escherichia coli.

The nucleotide sequence of a 1.8-kb fragment of Escherichia coli DNA containing the essential cell division gene ftsZ is reported. The FtsZ protein has an Mr of 40294 and has 23% charged residues with a calculated isoelectric point of 4.9. The codon usage of the ftsZ gene reflects that of a highly expressed gene. Also located on this DNA fragment is the 3' end of the ftsA gene and the 5' end of the envA gene. These designations were confirmed by locating Tn5 insertions within the ends of these genes that inactivate each of these genes. A potential promoter for ftsZ overlapped the 3' end of the ftsA gene. A Tn5 insertion was located within the 3' end of the ftsA and within this potential promoter. No transcription terminators were evident between ftsA and ftsZ or between ftsZ and envA.

Amino Acid Sequence↗

Structure and organization of hip, an operon that affects lethality due to inhibition of peptidoglycan or DNA synthesis.

High-frequency persistence to the lethal effects of inhibition of either DNA or peptidoglycan synthesis, the Hip phenotype, results from mutations at the hip locus of Escherichia coli K-12. The nucleotide sequence of DNA fragments which complement these mutations revealed an operon consisting of a possible regulatory region, including sequences with modest homology to an E. coli promoter, and two open reading frames which are translated both in vitro and in vivo. The stop codon of a 264-bp open reading frame, hipB, and the start codon of a 1,320-bp open reading frame, hipA, share an adenine residue. Assays of promoter strength, the location of the probable promoter with respect to the start of transcription, and codon usage all indicate that hipB and hipA are weakly expressed genes. The activity of the promoter is impaired by an adjacent downstream sequence which includes the coding region of hipB. The impairment is partially relieved by insertion of a premature translation termination signal within the coding region of hipB, suggesting involvement of the HipB protein in the regulation of this promoter. The arrangement of hipB and hipA within the operon and the toxicity of hipA for strains defective in or lacking hipB suggest an important interaction between the products of these genes.

Amino Acid Sequence↗

Laterally transferred elements and high pressure adaptation in Photobacterium profundum strains.

BACKGROUND: Oceans cover approximately 70% of the Earth's surface with an average depth of 3800 m and a pressure of 38 MPa, thus a large part of the biosphere is occupied by high pressure environments. Piezophilic (pressure-loving) organisms are adapted to deep-sea life and grow optimally at pressures higher than 0.1 MPa. To better understand high pressure adaptation from a genomic point of view three different Photobacterium profundum strains were compared. Using the sequenced piezophile P. profundum strain SS9 as a reference, microarray technology was used to identify the genomic regions missing in two other strains: a pressure adapted strain (named DSJ4) and a pressure-sensitive strain (named 3TCK). Finally, the transcriptome of SS9 grown under different pressure (28 MPa; 45 MPa) and temperature (4 degrees C; 16 degrees C) conditions was analyzed taking into consideration the differentially expressed genes belonging to the flexible gene pool. RESULTS: These studies indicated the presence of a large flexible gene pool in SS9 characterized by various horizontally acquired elements. This was verified by extensive analysis of GC content, codon usage and genomic signature of the SS9 genome. 171 open reading frames (ORFs) were found to be specifically absent or highly divergent in the piezosensitive strain, but present in the two piezophilic strains. Among these genes, six were found to also be up-regulated by high pressure. CONCLUSION: These data provide information on horizontal gene flow in the deep sea, provide additional details of P. profundum genome expression patterns and suggest genes which could perform critical functions for abyssal survival, including perhaps high pressure growth.

Atmospheric Pressure↗

Nucleotide sequence of the Escherichia coli gap gene. Different evolutionary behavior of the NAD+-binding domain and of the catalytic domain of D-glyceraldehyde-3-phosphate dehydrogenase.

A 1523-base-pair DNA fragment, spanning the gap gene from Escherichia coli, has been sequenced. It contains an open-reading frame whose length (330 amino acids) is in agreement with D-glyceraldehyde-3-phosphate dehydrogenase (GAPDH) molecular mass. This coding sequence is preceded by a Shine-Dalgarno complementary sequence and by two overlapping promoter-like structures. The codon usage within gap is consistent with that expected for a gene which is strongly expressed. The amino acid sequence of the E. coli GAPDH, deduced from the DNA sequence, contains all the amino acids postulated to play a functional role in GAPDH. Comparison of the E. coli enzyme with enzymes from other species reveals different evolutionary behaviour of the NAD+-binding domain and of the catalytic domain of GAPDH. The E. coli enzyme is found to be more similar to eucaryotic enzymes than to enzymes from thermophilic bacteria. This observation is discussed in terms of adaptation to growth at high temperature.

Amino Acid Sequence↗

Cloning of the blasticidin S deaminase gene (BSD) from Aspergillus terreus and its use as a selectable marker for Schizosaccharomyces pombe and Pyricularia oryzae.

Aspergillus terreus produces a unique enzyme, blasticidin S deaminase, which catalyzes the deamination of blasticidin S (BS), and in consequence confers high resistance to the antibiotic. A cDNA clone derived from the structural gene for BS deaminase (BSD) was isolated by transforming Escherichia coli with an Aspergillus cDNA expression library and directly selecting for the ability to grow in the presence of the antibiotic. The complete nucleotide sequence of BSD was determined and proved to contain an open reading frame of 393 bp, encoding a polypeptide of 130 amino acids. Comparison of its nucleotide sequence with that of bsr, the BS deaminase gene isolated from Bacillus cereus, indicated no homology and a large difference in codon usage. The activity of BSD expressed in E. coli was easily quantified by an assay based on spectrophotometric recording. The BSD gene was placed in a shuttle vector for Schizosaccharomyces pombe, downstream of the SV40 early region promoter, and this allowed direct selection with BS at high frequency, following transformation into the yeast. The BSD gene was also employed as a selectable marker for Pyricularia oryzae, which could not be transformed to BS resistance by bsr. These result promise that the BSD gene will be useful as a new dominant selectable marker for eukaryotes.

Amino Acid Sequence↗