Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Selection conflicts, gene expression, and codon usage trends in yeast.

Synonymous codon usage in yeast appears to be influenced by natural selection on gene expression, as well as regional variation in compositional bias. Because of the large number of potential targets of selection (i.e., most of the codons in the genome) and presumed small selection coefficients, codon usage is an excellent model for studying factors that limit the effectiveness of selection. We use factor analysis to identify major trends in codon usage for 5836 genes in Saccharomyces cerevisiae. The primary factor is strongly correlated with gene expression, consistent with the model that a subset of codons allows for more efficient translation. The secondary factor is very strongly correlated with third codon position GC content and probably reflects regional variation in compositional bias. We find that preferred codon usage decreases in the face of three potential limitations on the effectiveness of selection: reduced recombination rate, increased gene length, and reduced intergenic spacing. All three patterns are consistent with the Hill-Robertson effect (reduced effectiveness of selection among linked targets). A reduction in gene expression in closely spaced genes may also reflect selection conflicts due to antagonistic pleiotropy.

Codon↗

Homologous nucleotide sequences at the 5' termini of messenger RNAs synthesized from the yeast enolase and glyceraldehyde-3-phosphate dehydrogenase gene families. The primary structure of a third yeast glyceraldehyde-3-phosphate dehydrogenase gene.

Genomic DNA containing a third yeast glyceraldehyde-3-phosphate dehydrogenase structural gene has been isolated on a bacterial plasmid designated pgap11. The complete nucleotide sequence of this structural gene was determined. The gene contains no intervening sequences, codon usage is highly biased, and the nucleotide sequence of the coding portion of this gene is 90% homologous to the other two glyceraldehyde-3-phosphate dehydrogenase genes (Holland, J. P., and Holland, M. J. (1980) J. Biol. Chem. 255, 2596-2605). Based on the extent of nucleotide sequence divergence among the three glyceraldehyde-3-phosphate dehydrogenase genes, it is likely that they arose as a consequence of two duplication events and the gene contained on the hybrid plasmid designated pgap11 is a product of the first duplication event. All three structural genes share extensive nucleotide sequence homology in the 5'-noncoding regions adjacent to the three respective translational initiation codons. The gene contained on pgap11 is not homologous to the others downstream from the respective translational termination codon, however. The 5' termini of messenger RNAs synthesized from the three glyceraldehyde-3-phosphate dehydrogenase and two yeast enolase genes have been mapped to sites ranging from 36 to 82 nucleotides upstream from the respective translational initiation codons. In each case the 5' terminus of the mRNA maps to a region of strong nucleotide sequence homology which is shared by all five structural genes. These latter data confirm that all five structural genes are expressed during vegetative cell growth and further support the hypothesis that a portion of the 5'-noncoding flanking region of the yeast glyceraldehyde-3-phosphate dehydrogenase and enolase genes evolved from a common precursor sequence.

Base Sequence↗

Molecular cloning, primary structure and disruption of the structural gene of aldolase from Saccharomyces cerevisiae.

A yeast cDNA genetic library in a bacteriophage expression vector was screened using an antiserum reacting with fructose 1,6-bisphosphate aldolase from Saccharomyces cerevisiae. Radio-labelled probes of selected immunopositive clones were used for screening of a yeast genomic library. From the genomic clones a yeast/Escherichia coli shuttle plasmid was constructed containing on a 1990-base-pair fragment the entire structural gene FBA1 coding for yeast aldolase. The primary structure of the FBA1 gene was determined. An open reading frame comprises 1077 base pairs coding for a protein of 359 amino acids with a predicted molecular mass of 39,608 Da. As observed for other strongly expressed yeast genes, codon usage is extremely biased. The 810 base pairs at the 5' end and the 90 base pairs at the 3' end of the coding region of the cloned FBA1 gene are sufficient for normal expression and show characteristic elements present in the noncoding sequences of other yeast genes. Aldolase is the major protein in yeast cells transformed with a high-copy-number plasmid containing the FBA1 gene. The aldolase gene was disrupted by insertion of the yeast URA3 gene into the coding region of one FBA1 allele in a homozygous diploid ura3 strain. The haploid offsprings with the defective aldolase allele fba1::URA3 lack aldolase enzymatic activity and fail to grow in media containing as a carbon source metabolites of only one side of the aldolase reaction.

Amino Acid Sequence↗

Mammalian mitochondrial DNA evolution: a comparison of the cytochrome b and cytochrome c oxidase II genes.

The evolution of two mitochondrial genes, cytochrome b and cytochrome c oxidase subunit II, was examined in several eutherian mammal orders, with special emphasis on the orders Artiodactyla and Rodentia. When analyzed using both maximum parsimony, with either equal or unequal character weighting, and neighbor joining, neither gene performed with a high degree of consistency in terms of the phylogenetic hypotheses supported. The phylogenetic inconsistencies observed for both these genes may be the result of several factors including differences in the rate of nucleotide substitution among particular lineages (especially between orders), base composition bias, transition/transversion bias, differences in codon usage, and different constraints and levels of homoplasy associated with first, second, and third codon positions. We discuss the implications of these findings for the molecular systematics of mammals, especially as they relate to recent hypotheses concerning the polyphyly of the order Rodentia, relationships among the Artiodactyla, and various interordinal relationships.

Animals↗

Sequence analysis of the DdPYR5-6 gene coding for UMP synthase in Dictyostelium discoideum and comparison with orotate phosphoribosyl transferases and OMP decarboxylases.

A Dictyostelium discoideum DNA fragment that complements the ura3 and the ura5 mutants of Saccharomyces cerevisiae has been sequenced. It contains an open reading frame of 478 codons capable of encoding a polypeptide of molecular weight 52475. This gene, named DdPYR5-6, encodes a bifunctional protein composed of the orotate phosphoribosyl transferase (OPRTase) and the orotidine-5'-phosphate decarboxylase (OMPdecase) domains described for UMP synthase in mammals. The existence of separate domains for the two activities was suspected because deletion of the N-terminal coding segment of the gene eliminated the ura5 but not the ura3 complementing activity. We have now confirmed that the two parts of the open reading frame share homology with known OPRTase and OMPdecase sequences. Several blocks of sequence are conserved among OPRTase from bacteria, fungi and slime mold and one of them corresponds to the consensus sequence for phosphoribosylbinding sites. The OMPdecase domain shows extensive similarity with the yeast and Neurospora crassa enzymes, suggesting that they have evolved from an ancestral gene which was fused to the OPRTase gene in D. discoideum. It is less related to the bacterial enzyme but all these sequences present conserved blocks of homology which could identify the active site. The codon usage is strongly biased in a manner similar to that found for other D. discoideum genes. The flanking DNA contains homopolymers of A and T and alternating sequences that are characteristic of the gene organization in D. discoideum.

Amino Acid Sequence↗

Relative rates of nucleotide substitution in frogs.

Accurate estimation of relative mutation rates of mitochondrial DNA (mtDNA) and single-copy nuclear DNA (scnDNA) within lineages contributes to a general understanding of molecular evolutionary processes and facilitates making demographic inferences from population genetic data. The rate of divergence at synonymous sites ( K(s)) may be used as a surrogate for mutation rate. Such data are available for few organisms and no amphibians. Relative to mammals and birds, amphibian mtDNA is thought to evolve slowly, and the K(s) ratio of mtDNA to scnDNA would be expected to be low as well. Relative K(s) was estimated from a mitochondrial gene, ND2, and a nuclear gene, c-myc, using both "approximate" and likelihood methods. Three lineages of congeneric frogs were studied and this ratio was found to be approximately 16, the highest of previously reported ratios. No evidence of a low K(s) in the nuclear gene was found: c-myc codon usage was not biased, the K(s) was double the intron divergence rate, and the absolute K(s) was similar to estimates obtained here for other genes from other frog species. A high K(s) in mitochondrial vs. nuclear genes was unexpected in light of previous reports of a slow rate of mtDNA evolution in amphibians. These results highlight the need for further investigation of the effects of life history on mutation rates.

Animals↗

Contextual constraints on synonymous codon choice.

We have studied the statistical constraints on synonymous codon choice to evaluate various proposals regarding the origin of the bias in synonymous codon usage observed by Fiers et al. (1975), Air et al. (1976), Grantham et al. (1980) and others. We have determined the statistical dependence of the degenerate third base on either of its nearest neighbors in mitochondrial, prokaryotic, and eukaryotic coding sequences. We noted an increasing dependence of the third base on its nearest neighbors in moving from mitochondria to prokaryotes to eukaryotes. A statistical model assuming random equiprobable selection of synonymous codons was found grossly adequate for the mitochondria, but totally inadequate for prokaryotes and eukaryotes. A model assuming selection of synonymous codons reflecting a genomic strategy, i.e. the genome hypothesis of Grantham et al. (1980), gave a good approximation of the mitochondrial sequences. A statistical model which exactly maintains codon frequency, but allows the position of corresponding synonymous codons to vary was only grossly adequate for prokaryotes and totally inadequate for eukaryotes. The results of these simulations are consistent with the measures on experimental sequences and suggest that a "frequency constraint" model such as that of Grantham et al. (1980) may be an adequate explanation of the codon usage in mitochondria. However, in addition to this frequency constraint, there may be constraints on synonymous codon choice in prokaryotes due to codon context. Furthermore, any proposal to explain codon usage in eukaryotes must involve a constraint on the context of a codon in the sequence.

Amino Acid Sequence↗

Sequence of one alpha- and two beta-tubulin genes of Tetrahymena pyriformis. Structural and functional relationships with other eukaryotic tubulin genes.

Macronuclear DNA of the ciliate Tetrahymena pyriformis contains only one size class of fragments coding for alpha-tubulin, alpha TT. We have isolated alpha TT from a partial plasmid library, using Chlamydomonas reinhardtii alpha-tubulin gene as a probe. This gene as well as the two beta-tubulin genes, beta TT1 and beta TT2, have been sequenced. None of these genes contains introns and all use TGA as the stop codon. In the coding region of the two beta-tubulin genes, there are several TAA and TAG stop codons that probably code for glutamine. The codon usage is very biased. Regions flanking the tubulin coding sequences are A + T-rich (75%) and quite different among themselves. In these regions there are several putative transcription-regulatory sequences. Nuclear transcripts begin and terminate at multiple sites. The beta-tubulin proteins differ only in two amino acid residues. Primary structure of Tetrahymena tubulins as well as their hydropathy indexes show a high degree of homology with tubulins from other organisms. Two-dimensional electrophoretic analysis of the ciliary tubulins shows the presence of eight alpha-tubulins and four beta-tubulins. The alpha-tubulins migrate faster than the beta-tubulins, in contrast with what happens with brain tubulins. We suggest that there are several alpha- and beta-tubulin isoforms and the migratory inversion observed may be due to post-translational modifications.

Amino Acid Sequence↗

Actin in the oomycetous fungus Phytophthora infestans is the product of several genes.

Actin (ACT) in Phytophthora infestans is encoded by at least two genes, in contrast to unicellular and other filamentous fungi where there is a single gene. These genes (designated actA and actB) have been isolated from a genomic library of P. infestans. The complete nucleotide sequence of both genes has been determined. Unlike the actin-encoding genes (act) of other filamentous fungi, no introns are obvious in the coding region, a feature shared with the act genes of certain protists. Northern blotting and primer extension studies of the mRNA show that actA and actB are actively transcribed in mycelium, sporangia and germinating cysts but only at a low level in the case of actB. Both genes display bias in their codon usage. This is more extreme in actA. The deduced ACTB protein is strikingly similar to that of the Phytophthora megasperma actin and is more diverged from other actins than ACTA.

Actins↗

Actin-encoding cDNAs and gene expression during the intermolt cycle of the Bermuda land crab Gecarcinus lateralis.

Two actin-encoding cDNAs (act1 and act2) from Gecarcinus lateralis have been sequenced or partially sequenced and the corresponding proteins deduced. The act1 cDNA has a complete ORF; the act2 cDNA lacks most of the 5' end of the coding region. The nucleotide (nt) sequences of both clones are very similar to act sequences of many organisms, the most closely related being from another arthropod, the silkmoth Bombyx mori. The proteins Act1 and Act2 are more similar to vertebrate cytoplasmic actin isoforms (beta-actins) than to vertebrate muscle actins (alpha-actins); they are also more similar to animal actins than to those of fungi or plants. Codon usage is strongly biased toward C or G in the third position. The deduced number of amino acid (aa) residues and calculated Mr for Act1 are 376 aa and 41.94 kDa, respectively. The deduced aa sequence of Act1 is very similar to those of muscle actins of B. mori and Drosophila melanogaster. Southern blots indicated seven to eleven act genes in the crab genome. Northern blots probed with a segment from the 3' UTR of act1 showed a single band of approx. 1.6 kb in poly(A)+ mRNAs from epidermis, limb bud or claw muscle and in total RNAs from ovary and gill, and two bands of approx. 1.6 and 1.8 kb in total RNA from midgut gland. Western blots of one-dimensional gels of proteins from the four layers of the exoskeleton, epidermis, limb buds and claw muscle were probed with a monoclonal Ab against chicken gizzard actin; tissue- and stage-specific changes in actin content were observed. The presence of several isoforms, and differences in their number and occurrence at various stages of the intermolt cycle, were detected on Western blots of two-dimensional gels.

Actins↗

Location and sequence analysis of a 2-hydroxy-6-oxo-6-phenylhexa-2,4-dienoate hydrolase-encoding gene (bpdF) of the biphenyl/polychlorinated biphenyl degradation pathway in Rhodococcus sp. M5.

The 2-hydroxy-6-oxo-6-phenylhexa-2,4-dienoate (HOPD) hydrolase-encoding gene (bpdF) in the biphenyl (BP)/polychlorinated biphenyl (PCB)-degrading bacterium, Rhodococcus sp. M5 (M5), was found to be located within a 4.5-kb HindIII-BamHI genomic DNA that was 5.4 kb downstream from the bpdC1C2BADE gene cluster. The deduced amino acid (aa) sequence of bpdF revealed that the hydrolase contains 297 aa (32679 Da) that was verified by expression in the Escherichia coli T7 RNA polymerase/promoter system. Unlike previously known HOPD hydrolases, the aa sequence of BpdF appears unique. Interestingly, all HOPD hydrolases and related proteins from the phenol and toluene/xylene degradation pathways, were found to have a bias in the codon usage in the catalytic Ser within the conserved VGNS(M/F)GG motif.

Amino Acid Sequence↗

Composition strand asymmetries in prokaryotic genomes: mutational bias and biased gene orientation.

Most prokaryotic genomes display strand compositional asymmetries, but the reasons for these biases remain unclear. When the distribution of gene orientation is biased, as it often is, this may induce a bias in composition, as codon frequencies are not identical. We show here that this effect can be estimated and removed, and that the residual base skews are the highest at third base codon positions and lower at first and second positions. This strongly suggests that compositional asymmetries result from 1) a replication-related mutational bias that is filtered through selective pressure and/or from 2) an uneven distribution of gene orientation. In most cases, the mutational bias alters the codon usage and amino acid frequencies of the leading and the lagging strand. However, these features are not ubiquitous amongst prokaryotes, and the biological reasons for them remain to be found.

Bacillus subtilis↗

Complete mitochondrial DNA sequence and amino acid analysis of the cytochrome C oxidase subunit I (COI) from Aedes aegypti.

The complete sequence of the yellow fever mosquito, Aedes aegypti, mitochondrial cytochrome c oxidase subunit 1 gene has been identified. The nucleotide sequence codes for a 512 amino acid peptide. The AeCOI sequence is A + T rich (68.6%) and the codon usage is highly biased toward a preference for A- or T-ending triplets. The A. aegypti COI peptide shows high homology, up to 93% identity, with several other insect sequences and a phylogenetic analysis indicates that the A. aegypti sequence is closely related to two other mosquito species, Anopheles gambiae and A. quadrimaculatus. Comparisons of the nucleotide sequence for four A. aegypti laboratory strains revealed single nucleotide polymorphisms, with 25 nucleotide sites showing SNPs between strains. All SNPs occurred as synonymous transitions such that the peptide sequence is conserved among A. aegypti strains. RT-PCR analysis showed that COI is expressed at similar levels in all developmental stages and tissues.

AT Rich Sequence↗

Isolation and characterization of a Neurospora crassa ribosomal protein gene homologous to CYH2 of yeast.

We have isolated and characterized a Neurospora crassa gene homologous to the yeast CYH2 gene encoding L29, a cycloheximide sensitivity-conferring protein of the cytoplasmic ribosome. The cloned Neurospora gene was isolated by cross-hybridization to CYH2. It was sequenced from both cDNA and genomic clones. The coding region is interrupted by seven intervening sequences. Its deduced amino acid sequence shows 70% homology to that of yeast ribosomal protein L29 and 60% homology to that of mammalian ribosomal protein L27', suggesting that the protein has an important role in ribosomal function. The pattern of codon usage is highly biased, consistent with high translation efficiency. There is a single copy of this gene in N. crassa, and R. Metzenberg and coworkers have mapped its genetic location to the vicinity of the cyh-2 locus.

Animals↗

Cloning, sequencing, and expression of the mig gene of Mycobacterium avium, which codes for a secreted macrophage-induced protein.

Mycobacterium avium is an intracellular pathogen that has evolved to be a frequent cause of disseminated infection in immunocompromised patients. Although these bacilli are readily phagocytized, they are able to survive and even multiply within human macrophages. The process whereby mycobacteria circumvent the lytic functions of the macrophages is currently not well understood, but this is a key aspect in the pathogenicity of all pathogenic mycobacteria. Previously, we identified a gene in M. avium, designated mig (for macrophage-induced gene), the expression of which is induced when the bacilli grow in human macrophages (G. Plum and J. E. Clark-Curtiss, Infect. Immun. 62:476-483, 1994). In the present study we show that (i) the nucleotide sequence of the mig gene has an open reading frame of 295 amino acids with a strong bias for mycobacterial codon usage, (ii) the mig gene also codes for a putative signal peptide of 19 amino acid residues, (iii) mig is induced by acidity to be expressed as an early-secreted 30-kDa protein, and (iv) the Mig protein exhibits an AMP-binding domain signature. However, beyond this motif which is common to enzymes that activate a large variety of substrates, no homologies to known sequences are found. We also show that (v) Mycobacterium smegmatis strains expressing the Mig protein have a limited advantage for survival in macrophages. These findings may be concordant with a role of the mig gene in the virulence of M. avium.

Amino Acid Sequence↗

Molecular cloning of cDNA and analysis of protein secondary structure of Candida albicans enolase, an abundant, immunodominant glycolytic enzyme.

We isolated and sequenced a clone for Candida albicans enolase from a C. albicans cDNA library by using molecular genetic techniques. The 1.4-kbp cDNA encoded one long open reading frame of 440 amino acids which was 87 and 75% similar to predicted enolases of Saccharomyces cerevisiae and enolases from other organisms, respectively. The cDNA included the entire coding region and predicted a protein of molecular weight 47,178. The codon usage was highly biased and similar to that found for the highly expressed EF-1 alpha proteins of C. albicans. Northern (RNA) blot analysis showed that the enolase cDNA hybridized to an abundant C. albicans mRNA of 1.5 kb present in both yeast and hyphal growth forms. The polypeptide product of the cloned cDNA, which was purified as a recombinant protein fused to glutathione S-transferase, had enolase enzymatic activity and inhibited radioimmunoprecipitation of a single C. albicans protein of molecular weight 47,000. Analysis of the predicted C. albicans enolase showed strong conservation in regions of alpha helices, beta sheets, and beta turns, as determined by comparison with the crystal structure of apo-enolase A of S. cerevisiae. The lack of cysteine residues and a two-amino-acid insertion in the main domain differentiated C. albicans enolase from S. cerevisiae enolase. Immunofluorescence of whole C. albicans cells by using a mouse antiserum generated against the purified fusion protein showed that enolase is not located on the surface of C. albicans. Recombinant C. albicans enolase will be useful in understanding the pathogenesis and host immune response in disseminated candidiasis, since enolase is an immunodominant antigen which circulates during disseminated infections.

Amino Acid Sequence↗

Cloning, sequencing, and expression of the Zymomonas mobilis phosphoglycerate mutase gene (pgm) in Escherichia coli.

Phosphoglycerate mutase is an essential glycolytic enzyme for Zymomonas mobilis, catalyzing the reversible interconversion of 3-phosphoglycerate and 2-phosphoglycerate. The pgm gene encoding this enzyme was cloned on a 5.2-kbp DNA fragment and expressed in Escherichia coli. Recombinants were identified by using antibodies directed against purified Z. mobilis phosphoglycerate mutase. The pgm gene contains a canonical ribosome-binding site, a biased pattern of codon usage, a long upstream untranslated region, and four promoters which share sequence homology. Interestingly, adhA and a D-specific 2-hydroxyacid dehydrogenase were found on the same DNA fragment and appear to form a cluster of genes which function in central metabolism. The translated sequence for Z. mobilis pgm was in full agreement with the 40 N-terminal amino acid residues determined by protein sequencing. The primary structure of the translated sequence is highly conserved (52 to 60% identity with other phosphoglycerate mutases) and also shares extensive homology with bisphosphoglycerate mutases (51 to 59% identity). Since Southern blots indicated the presence of only a single copy of pgm in the Z. mobilis chromosome, it is likely that the cloned pgm gene functions to provide both activities. Z. mobilis phosphoglycerate mutase is unusual in that it lacks the flexible tail and lysines at the carboxy terminus which are present in the enzyme isolated from all other organisms examined.

2,3-Diphosphoglycerate↗

The genomic pattern of tDNA operon expression in E. coli.

In fast-growing microorganisms, a tRNA concentration profile enriched in major isoacceptors selects for the biased usage of cognate codons. This optimizes translational rate for the least mass invested in the translational apparatus. Such translational streamlining is thought to be growth-regulated, but its genetic basis is poorly understood. First, we found in reanalysis of the E. coli tRNA profile that the degree to which it is translationally streamlined is nearly invariant with growth rate. Then, using least squares multiple regression, we partitioned tRNA isoacceptor pools to predicted tDNA operons from the E. coli K12 genome. Co-expression of tDNAs in operons explains the tRNA profile significantly better than tDNA gene dosage alone. Also, operon expression increases significantly with proximity to the origin of replication, oriC, at all growth rates. Genome location explains about 15% of expression variation in a form, at a given growth rate, that is consistent with replication-dependent gene concentration effects. Yet the change in the tRNA profile with growth rate is less than would be expected from such effects. We estimated per-copy expression rates for all tDNA operons that were consistent with independent estimates for rDNA operons. We also found that tDNA operon location, and the location dependence of expression, were significantly different in the leading and lagging strands. The operonic organization and genomic location of tDNA operons are significant factors influencing their expression. Nonrandom patterns of location and strandedness shown by tDNA operons in E. coli suggest that their genomic architecture may be under selection to satisfy physiological demand for tRNA expression at high growth rates.

Journal Article↗