Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Analysis of the primary structure and promoter function of a pyruvate decarboxylase gene (PDC1) from Saccharomyces cerevisiae.

The PDC1 gene of Saccharomyces cerevisiae, encoding pyruvate decarboxylase was sequenced. The gene contains an open reading frame of 1647 base pairs. The codon usage shows the same strong bias as found for some other glycolytic enzymes. Transcription starts mainly at -30 and terminates 100 base pairs downstream of the termination codon. In some strains a second termination site, 46 base pairs upstream of the stop codon was observed. The function of the promoter region was analyzed by fusion to the bacterial structural gene encoding beta-lactamase (bla). On multicopy plasmid or integrated in the genome, the expression of the bla gene showed the regulation of the authentic PDC1 gene.

Amino Acid Sequence↗

The complete mitochondrial genomes of the sea lily Gymnocrinus richeri and the feather star Phanogenia gracilis: signature nucleotide bias and unique nad4L gene rearrangement within crinoids.

Complete DNA sequences have been determined for the mitochondrial genomes of the crinoids Phanogenia gracilis (15892 bp) and Gymnocrinus richeri (15966 bp). The mitochondrial genetic map of the stalkless feather star P. gracilis is identical to that of the comatulid feather star Florometra serratissima (Scouras, A., Smith, M.J., 2001. Mol. Biol. Evol. 18, 61-73). The mitochondrial gene order of the stalked crinoid G. richeri differs from that of F. serratissima and P. gracilis by the transposition of the nad4L protein gene. The G. richeri nad4L mitochondrial map position is unique among metazoa and is likely a derived feature in this stalked crinoid. Nucleotide compositional analyses of protein genes encoded on the major sense strand confirm earlier conclusions regarding a crinoid-distinctive T over C bias. All three crinoids exhibit high T levels in third codon positions, whereas other echinoderm classes favor A or C in the third codon position. The nucleotide bias is reflected in the relative synonymous codon usage patterns of crinoids versus other echinoderms. We suggest that the nucleotide bias of crinoids, in comparison to other echinoderms, indicates that a physical inversion of the origin of replication has occurred in the crinoid lineage. Evolutionary rate tests support the use of the cytochrome b (cob) gene in molecular phylogenetic analyses of echinoderms. A consensus echinoderm tree was generated based on cytochrome b nucleotide alignments that placed the asteroids as a sister group to a clade containing the ophiuroids and the (echinoids+holothuroids) with the crinoids basal to the rest of the echinoderm classes: [Crinoid,(Asteroid,(Ophiuroid,(Echinoid,Holothuroid)))].

Animals↗

Synonymous codon usage in adenoviruses: influence of mutation, selection and protein hydropathy.

Trends in synonymous codon usage in adenoviruses have been examined through the multivariate statistical analysis on the annotated protein-coding regions of 22 adenoviral species, for which complete genome sequences are available. One of the major determinants of such trends is the G+C content at third codon positions of the genes, the average value of which varied from one viral genome to other depending on the overall mutational bias of the species. G3S and C3S interacted synergistically along the first principal axis of correspondence analysis on the Relative Synonymous Codon Usage of adenoviral genes, but antagonistically along the second principal axis. The intra-genomic variation in codon usage pattern in adenoviruses is generally influenced by asymmetrical mutational bias in two DNA strands. Other major determinants of the trends are the natural selection, putatively operative at the level of translation and quite interestingly, hydropathy of the encoded proteins. The trends in codon usage, though characterized by distinct virus-specific mutational bias, do not exhibit any sign of host-specificity. Significant variations are observed in synonymous codon choice in structural and nonstructural genes of adenoviruses.

Adenoviridae↗

Mutational and selective pressures on codon and amino acid usage in Buchnera, endosymbiotic bacteria of aphids.

We have explored compositional variation at synonymous (codon usage) and nonsynonymous (amino acid usage) positions in three complete genomes of Buchnera, endosymbiotic bacteria of aphids, and also in their orthologs in Escherichia coli, a close free-living relative. We sought to discriminate genes of variable expression levels in order to weigh the relative contributions of mutational bias and selection in the genomic changes following symbiosis. We identified clear strand asymmetries, distribution biases (putative high-expression genes were found more often on the leading strand), and a residual slight codon bias within each strand. Amino acid usage was strongly biased in putative high-expression genes, characterized by avoidance of aromatic amino acids, but above all by greater conservation and resistance to AT enrichment. Despite the almost complete loss of codon bias and heavy mutational pressure, selective forces are still strong at nonsynonymous sites of a fraction of the genome. However, Buchnera from Baizongia pistaciae appears to have suffered a stronger symbiotic syndrome than the two other species.

Amino Acids↗

Spectinomycin operon of Micrococcus luteus: evolutionary implications of organization and novel codon usage.

The complete DNA sequence of the Micrococcus luteus spectinomycin (spc) operon and its adjacent regions has been determined. The sequence has revealed the presence of genes that are homologous to those of the Escherichia coli ribosomal and related proteins, L14, L24, L5, S8, L6, L18, S5, L30, L15, and secretion protein Y (sec Y), and the gene for adenylate kinase (adk). The gene arrangement in the spc operon is essentially the same as that of E. coli except for the absence in the M. luteus spc operon of the genes for S14 and X protein that exist in the E. coli spc operon. SecY and adk seem to be composed of another operon (adk operon) with at least an open reading frame. The deduced amino acid sequences for these ribosomal proteins are well conserved among the two species (40-65% identity). Reflecting the high genomic guanine and cytosine (GC) content of M. luteus (74%), the codon usage of the genes is extremely biased toward use of G and C, about 94% of the codon third positions being G or C. Seven codons, AUA, AAA, AGA, UUA, GUA, CUA, and CAA, all of which have A at the codon third positions, are completely absent in the M. luteus genes examined. Out of 11 genes in the M. luteus spc and adk operons, 5 (10) use GUG (UGA) and 6 (1) use AUG (UAA) as an initiation (termination) codon.

Amino Acid Sequence↗

Translational selection shapes codon usage in the GC-rich genome of Chlamydomonas reinhardtii.

In unicellular species codon usage is determined by mutational biases and natural selection. Among prokaryotes, the influence of these factors is different if the genome is skewed towards AT or GC, since in AT-rich organisms translational selection is absent. On the other hand, in AT-rich unicellular eukaryotes the two factors are present. In order to understand if GC-rich genomes display a similar behavior, the case of Chlamydomonas reinhardtii was studied. Since we found that translational selection strongly influences codon usage in this species, we conclude that there is not a common pattern among unicellular organisms.

AT Rich Sequence↗

Patterns of context-dependent codon biases.

The association of codon context and codon usage was studied in seven bacteria as well as Schizosaccharomyces pombe and Encephalitozoon cuniculi. The association is strongest in magnitude closest to the codons of interest but there is apparently no rule about which of the two contexts is generally strongest associated to codon usage. In all bacterial species and in the intron-rich Sch. pombe it was furthermore observed from plots of chi2 versus N that the wobble positions of codons in the proximity cause regular peaks both upstream and downstream. This observation is discussed in relation to a possible effect of mutational pressure on the association of codon usage and codon context. Absence of peaks corresponding to the wobble positions in the intron-poor En. cuniculi, and presence in Sch. pombe, may indicate that the role of introns in the context-dependent codon bias is negligible.

Animals↗

Codon usage and gene expression level in Dictyostelium discoideum: highly expressed genes do 'prefer' optimal codons.

Codon usage patterns in the slime mould Dictyostelium discoideum have been re-examined (a total of 58 genes have been analysed). Considering the extreme A + T-richness of this genome (G + C = 22%), there is a surprising degree of codon usage variation among genes. For example, G + C content at silent sites varies from less than 10% to greater than 30%. It was previously suggested [Warrick, H.M. and Spudich, J.A. (1988) Nucleic Acids Res. 16: 6617-6635] that highly expressed genes contain fewer 'optimal' codons than genes expressed at lower levels. However, it appears that the optimal codons were misidentified. Multivariate statistical analysis shows that the greatest variation among genes is in relative usage of a particular subset of codons (about one per amino acid), many of which are C-ending. We have identified these as optimal codons, since (i) their frequency is positively correlated with gene expression level, and (ii) there is a strong mutation bias in this genome towards A and T nucleotides. Thus, codon usage in D. discoideum can be explained by a balance between the forces of mutational bias and translational selection.

Codon↗

A new measure to study phylogenetic relations in the brown algal order Ectocarpales: the "codon impact parameter".

We analyse forty-seven chloroplast genes of the large subunit of RuBisCO, from the algal order Ectocarpales, sourced from GenBank. Codon-usage weighted by the nucleotide base-bias defines our score called the codon-impact-parameter. This score is used to obtain phylogenetic relations amongst the 47 Ectocarpales. We compare our classification with the ones done earlier.

Base Composition↗

Rare codons in E. coli and S. typhimurium signal sequences.

Codon usage has been examined in the signal sequences of 27 genes encoding proteins which possess leader peptides, and are inner-membrane located or exported. The results have been compared with codon usage in the corresponding coding sequences of most of the mature proteins. A bias is observed in the usage of rare codons for two of the three hydrophobic amino acids for which there are rare codons. Since hydrophobic residues are predominant in leader peptides, we suggest that a resulting concentration of rare codons in the signal sequence may play a role (or have played a role in the evolutionary past) in the secretion process by delaying translation.

Base Sequence↗

GC-biased segregation of noncoding polymorphisms in Drosophila.

The study of base composition evolution in Drosophila has been achieved mostly through the analysis of coding sequences. Third codon position GC content, however, is influenced by both neutral forces (e.g., mutation bias) and natural selection for codon usage optimization. In this article, large data sets of noncoding DNA sequence polymorphism in D. melanogaster and D. simulans were gathered from public databases to try to disentangle these two factors-noncoding sequences are not affected by selection for codon usage. Allele frequency analyses revealed an asymmetric pattern of AT vs. GC noncoding polymorphisms: AT --> GC mutations are less numerous, and tend to segregate at a higher frequency, than GC --> AT ones, especially at GC-rich loci. This is indicative of nonstationary evolution of base composition and/or of GC-biased allele transmission. Fitting population genetics models to the allele frequency spectra confirmed this result and favored the hypothesis of a biased transmission. These results, together with previous reports, suggest that GC-biased gene conversion has influenced base composition evolution in Drosophila and explain the correlation between intron and exon GC content.

Animals↗

Evolutionary aspects of trypanosomes: analysis of genes.

The genes for four glycolytic enzymes of Trypanosoma brucei have been analyzed. The proteins encoded by these genes show 38-57% identity with their counterparts in other organisms, whether pro- or eukaryotic. These data are consistent with a phylogenetic tree in which trypanosomes diverged very early from the main branch of the eukaryotic lineage. No definite conclusion can be drawn yet about the evolutionary origin of glycosomes, the microbodies of trypanosomes which contain most enzymes of the glycolytic pathway. A bias could be observed in the codon usage of the glycolytic genes and genes for other housekeeping proteins, indicating that trypanosomes may have selected a nucleotide sequence that enables efficient translation. However, the genes for variant surface glycoproteins (VSGs) do not show such a bias. This lack of preference for special codons is explained by the high evolutionary rate that could be observed for VSG genes.

Animals↗

Structure of the Escherichia coli K12 regulatory gene tyrR. Nucleotide sequence and sites of initiation of transcription and translation.

The nucleotide sequence of 1964 base pairs of the Escherichia coli K12 chromosome containing the autogenously regulated regulatory gene tyrR has been determined. The site of initiation of transcription of tyrR has been mapped by primer-extension analysis, and the initiation codon has been identified by site-specific deletion mutagenesis. The nucleotide sequence predicts a subunit molecular weight of 53,099 for the TyrR protein. Codon usage in the tyrR structural gene shows a bias toward those synonymic codons which are used rarely in efficiently expressed E. coli genes. The nucleotide sequence of a 22-base pair region adjacent to the promoter and distal to the structural gene exhibits considerable identity with corresponding regions of other genes regulated by tyrR. It is proposed that this is a site for repression by the TyrR protein.

Amino Acid Sequence↗

Identification and nucleotide sequence of the Leptospira biflexa serovar patoc trpE and trpG genes.

Leptospira biflexa is a representative of an evolutionarily distinct group of eubacteria. In order to better understand the genetic organization and gene regulatory mechanisms of this species, we have chosen to study the genes required for tryptophan biosynthesis in this bacterium. The nucleotide sequence of the region of the L. biflexa serovar patoc chromosome encoding the trpE and trpG genes has been determined. Four open reading frames (ORFs) were identified in this region, but only three ORFs were translated into proteins when the cloned genes were introduced into Escherichia coli. Analysis of the predicted amino acid sequences of the proteins encoded by the ORFs allowed us to identify the trpE and trpG genes of L. biflexa. Enzyme assays confirmed the identity of these two ORFs. Anthranilate synthase from L. biflexa was found to be subject to feedback inhibition by tryptophan. Codon usage analysis showed that there was a bias in L. biflexa towards the use of codons rich in A and T, as would be expected from its G + C content of 37%. Comparison of the amino acid sequences of the trpE gene product and the trpG gene product with corresponding gene products from other bacteria showed regions of highly conserved sequence.

Amino Acid Sequence↗

Amplification and molecular cloning of the IMP dehydrogenase gene of Leishmania donovani.

A mutant (MPA100) strain of Leishmania donovania was generated from a wild type (D1700) population by virtue of its ability to survive the selective pressure of gradually increasing concentrations of mycophenolic acid (MPA), an inhibitor of IMP dehydrogenase (IMPDH) activity. Comparative growth experiments revealed that the MPA100 strain was 100-fold more resistant to MPA toxicity and cross-resistant to ribavarin, another inhibitor of IMPDH. A direct comparison of IMPDH levels in D1700 and MPA100 cells showed that the latter expressed at least 20-fold higher enzyme activity. In order to evaluate the mechanism by which MPA100 cells overexpressed IMPDH, the leishmanial gene encoding IMPDH was isolated from a genomic library in EMBL3 by cross-hybridization to a mouse IMPDH cDNA, and a 2.3-kilobase EcoRV-PstI fragment was subcloned into a Bluescript vector and sequenced. The EcoRV-PstI fragment contained an open reading frame of 514 amino acids that encompassed the entire leishmanial IMPDH coding sequence. The predicted amino acid sequence showed a 52.5% identity with that of the corresponding human IMPDH. The codon usage of the leishmanial IMPDH gene reflected a strong bias toward codons containing either G or C in the wobble position. The EcoRV-PstI fragment hybridized to a 3.0-kilobase mRNA that was expressed at 10-20-fold greater levels in the MPA100 cells. Using the EcoRV-PstI fragment as a probe, the increased amount of IMPDH activity and IMPDH mRNA in the MPA100 cells could be attributed to an approximately 10-20-fold amplification of the leishmanial IMPDH gene.

Amino Acid Sequence↗

Nucleotide sequence of the gene encoding the nitrogenase iron protein (nifH) of Azospirillum brasilense and identification of a region controlling nifH transcription.

The DNA sequence was determined for the Azospirillum brasilense nifH gene and part of the nifD gene. The nifH gene is 885 bp long and encodes 293 amino acid residues. The region upstream of the nifH open reading frame contains a putative promoter whose sequence shows perfect homology with promoters of other diazotrophic bacteria and two putative upstream activator sequences. Experiments with the promoter-probe vector pAF300 showed that this region promotes transcription in response to the nitrogen and oxygen availability of the cell. The amino acid sequence was deduced from the DNA nucleotide sequence of nifH; the polypeptide contains the four cysteine residues highly conserved among other nifH products and an arginine residue at position 101 which could be the site of the modification occurring during the "switch-off" of nitrogenase. The codon usage appears to be very biased reflecting the high G + C content of the Azospirillum nifH gene. In a comparison of the amino acid sequence with the other 18 known nifH gene products, the A. brasilense nifH product showed the highest level of homology with fast-growing Rhizobia suggesting interesting evolutionary implications.

Amino Acid Sequence↗

The ambush hypothesis: hidden stop codons prevent off-frame gene reading.

Coding sequences lack stop codons, but many stops appear off-frame. Off-frame stops (stops in -1 and +1 shifted reading frames, termed hidden stops) terminate frame-shifted translation, potentially decreasing energy, and resource waste on nonfunctional proteins. Benefits may include reduced waste elimination costs and avoidance of potentially cytotoxic frame-shifted products. Our "ambush" hypothesis suggests that hidden stops are sometimes selected for. Codons of many amino acids can contribute to hidden stops, depending on the synonymous position state and adjacent codons. In vertebrate mitochondria, 31.75% of all amino acid combinations can form hidden stops. Codons with more potential to form hidden stops have greater usage frequency and bias in their favor among synonymous codons. Among primates, predicted mitochondrial rRNA secondary structure stability correlates negatively with the number of hidden stops in the mitochondrial genome. The taxonomic distribution of genetic codes suggests that +1 frameshifts might be more frequent than -1 frameshifts. This is confirmed by analyses of primate mitochondrial genomes: species with unstable rRNAs have more +1 stops, but the correlation is weak for -1 stops. High hidden stop density seems to be an adaptation in species with slippage prone ribosomes (unstable rRNAs). Hidden stops may thus compensate for reduced efficiency of some parts of the biosynthetic machinery. Some experimental data confirm our hypothesis: gene expression increases with the experimentally manipulated number of stops in the promoter region of a gene, suggesting biotechnological applications.

Animals↗

Codon usage domains over bacterial chromosomes.

The geography of codon bias distributions over prokaryotic genomes and its impact upon chromosomal organization are analyzed. To this aim, we introduce a clustering method based on information theory, specifically designed to cluster genes according to their codon usage and apply it to the coding sequences of Escherichia coli and Bacillus subtilis. One of the clusters identified in each of the organisms is found to be related to expression levels, as expected, but other groups feature an over-representation of genes belonging to different functional groups, namely horizontally transferred genes, motility, and intermediary metabolism. Furthermore, we show that genes with a similar bias tend to be close to each other on the chromosome and organized in coherent domains, more extended than operons, demonstrating a role of translation in structuring bacterial chromosomes. It is argued that a sizeable contribution to this effect comes from the dynamical compartimentalization induced by the recycling of tRNAs, leading to gene expression rates dependent on their genomic and expression context.

Amino Acids↗