Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Synonymous substitutions in the Xdh gene of Drosophila: heterogeneous distribution along the coding region.

The Xdh (rosy) region of Drosophila subobscura has been sequenced and compared to the homologous region of D. pseudoobscura and D. melanogaster. Estimates of the numbers of synonymous substitutions per site (Ks) confirm that Xdh has a high synonymous substitution rate. The distributions of both nonsynonymous and synonymous substitutions along the coding region were found to be heterogeneous. Also, no relationship has been detected between Ks estimates and codon usage bias along the gene, in contrast with the generally observed relationship among genes. This heterogeneous distribution of synonymous substitutions along the Xdh gene, which is expression-level independent, could be explained by a differential selection pressure on synonymous sites along the coding region acting on mRNA secondary structure. The synonymous rate in the Xdh coding region is lower in the D. subobscura than in the D. pseudoobscura lineage, whereas the reverse is true for the Adh gene.

Animals↗

Transcription of the triose-phosphate-isomerase gene of Schizosaccharomyces pombe initiates from a start point different from that in Saccharomyces cerevisiae.

Gene tpi, encoding the glycolytic enzyme triose phosphate isomerase (TPI) from the fission yeast Schizosaccharomyces pombe was cloned by complementation of a Saccharomyces cerevisiae tpil mutant. Nucleotide sequence analysis of the cloned gene revealed a single open reading frame (ORF) encoding a protein 59% homologous to S. cerevisiae TPI. The gene has a very high codon usage bias. Messenger RNA synthesis initiates at two points located 38 and 44 nucleotides downstream from a TATA box promoter sequence. In S. cerevisiae, transcription of this S. pombe gene initiates about 26 nucleotides downstream from the S. pombe start points. This observation indicates that the two yeasts have diverged in the mechanism which determines the 5' end of the messenger RNA relative to the TATA box. It appears that in some respects the transcription initiation mechanism of S. pombe more closely resembles that of higher eukaryotes than does the S. cerevisiae mechanism.

Amino Acid Sequence↗

Translational selection and yeast proteome evolution.

The primary structures of peptides may be adapted for efficient synthesis as well as proper function. Here, the Saccharomyces cerevisiae genome sequence, DNA microarray expression data, tRNA gene numbers, and functional categorizations of proteins are employed to determine whether the amino acid composition of peptides reflects natural selection to optimize the speed and accuracy of translation. Strong relationships between synonymous codon usage bias and estimates of transcript abundance suggest that DNA array data serve as adequate predictors of translation rates. Amino acid usage also shows striking relationships with expression levels. Stronger correlations between tRNA concentrations and amino acid abundances among highly expressed proteins than among less abundant proteins support adaptation of both tRNA abundances and amino acid usage to enhance the speed and accuracy of protein synthesis. Natural selection for efficient synthesis appears to also favor shorter proteins as a function of their expression levels. Comparisons restricted to proteins within functional classes are employed to control for differences in amino acid composition and protein size that reflect differences in the functional requirements of proteins expressed at different levels.

Adaptation, Physiological↗

Comparative analysis of genes encoding methyl coenzyme M reductase in methanogenic bacteria.

The sequence of the gene cluster encoding the methyl coenzyme M reductase (MCR) in Methanococcus voltae was determined. It contains five open reading frames (ORF), three of which encode the known enzyme subunits. Putative ribosome binding sites were found in front of all ORFs. They differ in their degrees of complementarity to the 3' end of the 16 S rRNA, which is discussed in terms of different translation efficiencies of the respective genes. The codon usage bias is different in the subunit encoding genes compared with the two other ORFs in the cluster and two other known genes of Mc. voltae. This is interpreted in terms of increased translational accuracy of the highly expressed MCR subunit genes. The derived polypeptide sequences encoded by the five ORFs of the MCR cluster were compared to those of the respective genes in Methanobacterium thermoautotrophicum Marburg and Methanosarcina barkeri. Conserved regions were detected in the enzyme subunits, which are candidates for factor binding domains. Conserved hydrophobic sequences found in the alpha and beta subunits are discussed with respect to the membrane association of the enzyme.

Amino Acid Sequence↗

The putative acetyl-CoA synthetase gene of Cryptosporidium parvum and a new conserved protein motif in acetyl-CoA synthetases.

We determined the nucleotide (nt) sequence of the putative gene encoding acetyl-coenzyme A synthetase (ACS) from the parasitic protozoan Cryptosporidium parvum. The gene is single copy, located on a chromosome of approximately 1.08 mb, and has no introns. The gene is characterized by low codon usage bias and encodes a 694-amino acid (aa) protein with a predicted molecular size of 78 kDa, similar to other ACSs from different prokaryotic and eukaryotic species. Comparison of multiple protein alignments of ACSs revealed a new conserved sequence motif PKT(R/V/L)SGK(I/V/T)(T/M/V/K)R(R/N) near the C-terminus, which may be a signature for ACSs. This motif shares significant homology with sequences from other members of the AMP-binding family, has secondary structure similar to the purine-binding motif of ATP- and GTP-ases, and may play a role in the enzymatic activity of proteins from the AMP-binding family.

Acetate-CoA Ligase↗

The complete maternal and paternal mitochondrial genomes of the Mediterranean mussel Mytilus galloprovincialis: implications for the doubly uniparental inheritance mode of mtDNA.

The maternal (F) and paternal (M) mitochondrial genomes of the mussel Mytilus galloprovincialis have diverged by about 20% in nucleotide sequence but retained identical gene content and gene arrangement and similar nucleotide composition and codon usage bias. Both lack the ATPase8 subunit gene, have two tRNAs for methionine and a longer open-reading frame for cox3 than seen in other mollusks. Between the F and M genomes, tRNAs are most conserved followed by rRNAs and protein-coding genes, even though the degree of divergence varies considerably among the latter. Divergence at nad3 is exceptionally low most likely because this gene includes the origin of transcription of the lagging strand (O(L)). Noncoding regions are the least conserved with the notable exception of the central domain of the main control region and a segment of another noncoding region immediately following nad3. The amino acid divergence (14%) of the two genomes is smaller than in two other pairs of conspecific genomes that are available in GenBank, that of the clam Venerupis philippinarum (34%) and of the fresh water mussel Inversidens japanensis (50%), suggesting that doubly uniparental inheritance of mtDNA emerged at different times in the three species or that there has been a relatively recent replacement of the male genome by the female in the Mytilus line. The latter hypothesis is supported from phylogenetic and population studies of Mytilidae. That the M genome contains a full complement of genes with no premature termination codons argues against it being a selfish element that rides with the sperm. It is shorter than the F by 118 bp, which apparently cannot account for the postulated replicative advantage of this genome over the F in male gonads. The high similarity of the two genomes explains why the F genome may assume the role of the M genome, but it does not exclude the possibility that for this to happen some M-specific sequences must be transferred on to the F genome by means of recombination. If such sequences exist they would most likely be located in noncoding regions.

Animals↗

Design and expression of a synthetic phyC gene encoding the neutral phytase in Pichia pastoris.

The 1074-bp phyCs gene (optimized phyC gene) encoding neutral phytase was designed and synthesized according to the methylotrophic yeast Pichia pastoris codon usage bias without altering the protein sequence. The expression vector, pP9K-phyCs, was linearized and transformed in P. pastoris. The yield of total extracellular phytase activity was 17.6 U/ml induced in Buffered Methanol-complex Medium (BMMY) and 18.5 U/ml in Wheat Bran Extract Induction (WBEI) medium at the flask scale, respectively, improving over 90 folds compared with the wild-type isolate. Purified enzyme showed temperature optimum of 70 degrees and pH optimum of 7.5. The enzyme activity retained 97% of the relative activity after incubation at 80 degrees for 5 min. Because of the heavy glycosylation the expressed phytase had a molecular size of approximately 51 kDa. After deglycosylation by endoglycosylase H (EndoH(f)), the enzyme had an apparent molecular size of 42 kDa. Its property and thermostability was affected by the glycosylation.

6-Phytase↗

PF-IND: probability algorithm and software for separation of plant and fungal sequences.

The separation of plant and fungal sequences in EST pools by bioinformatic methods is difficult because of sequence similarities between plants and fungi, lack of enough sequence information, and the short length of the isolated fragments. An algorithm and software that utilize the differences in codon usage bias to discriminate between plant and fungal sequences are described. The software (PF-IND) includes five pairs of fungi and their host plants that can be used to analyze a large number of related species. Analysis of a sequence provides an arbitrary value that defines the likelihood that a sequence will be a fungal or a plant gene. The software can distinguish between homologous fungal and plant genes and it helps identify the correct reading frame of unknown expressed sequence tags (ESTs) for which BLAST analyses do not provide clear information. Short sequences of 100-150 bp can be analyzed with high confidence. PF-IND analysis of 100 sequences derived from fungal infected plants identified the origin of 94 sequences. Only 66 sequences were identified by a BLASTX analysis of the same 100 ESTs. Overall, PF-IND is a novel bioinformatic tool aimed at assisting the research of fungus-plant interactions.

Algorithms↗

Analyses of frameshifting at UUU-pyrimidine sites.

Others have recently shown that the UUU phenylalanine codon is highly frameshift-prone in the 3'(rightward) direction at pyrimidine 3'contexts. Here, several approaches are used to analyze frameshifting at such sites. The four permutations of the UUU/C (phenylalanine) and CGG/U (arginine) codon pairs were examined because they vary greatly in their expected frameshifting tendencies. Furthermore, these synonymous sites allow direct tests of the idea that codon usage can control frameshifting. Frameshifting was measured for these dicodons embedded within each of two broader contexts: the Escherichia coli prfB (RF2 gene) programmed frameshift site and a 'normal' message site. The principal difference between these contexts is that the programmed frameshift contains a purine-rich sequence upstream of the slippery site that can base pair with the 3'end of 16 S rRNA (the anti-Shine-Dalgarno) to enhance frameshifting. In both contexts frameshift frequencies are highest if the slippery tRNAPhe is capable of stable base pairing in the shifted reading frame. This requirement is less stringent in the RF2 context, as if the Shine-Dalgarno interaction can help stabilize a quasi-stable rephased tRNA:message complex. It was previously shown that frameshifting in RF2 occurs more frequently if the codon 3'to the slippery site is read by a rare tRNA. Consistent with that earlier work, in the RF2 context frameshifting occurs substantially more frequently if the arginine codon is CGG, which is read by a rare tRNA. In contrast, in the 'normal' context frameshifting is only slightly greater at CGG than at CGU. It is suggested that the Shine-Dalgarno-like interaction elevates frameshifting specifically during the pause prior to translation of the second codon, which makes frameshifting exquisitely sensitive to the rate of translation of that codon. In both contexts frameshifting increases in a mutant strain that fails to modify tRNA base A37, which is 3'of the anticodon. Thus, those base modifications may limit frameshifting at UUU codons. Finally, statistical analyses show that UUU Ynn dicodons are extremely rare in E.coli genes that have highly biased codon usage.

Arginine↗

Ribosomal protein L9 is the product of GRC5, a homolog of the putative tumor suppressor QM in S. cerevisiae.

Genes encoding members of the highly conserved QM family have been identified in eukaryotic organisms from yeast to man. Results of previous studies have suggested roles for QM in control of cell growth and proliferation, perhaps as a tumor suppressor, and in energy metabolism. We identified recessive lethal alleles of the Saccharomyces cerevisiae QM homolog GRC5 that increased GCN4 expression when present in multiple copies. These alleles encode truncated forms of the yeast QM protein Grc5p. Using a functional epitope-tagged GRC5 allele, we localized Grc5p to a 60S fraction that contained the large ribosomal subunit. Two-dimensional gel analysis of highly purified yeast ribosomes indicated that Grc5p corresponds to 60S ribosomal protein L9. This identification is consistent with the predicted physical characteristics of eukaryotic QM proteins, the highly biased codon usage of GRC5, and the presence of putative Rap1p-binding sites in the 5' sequences of the yeast GRC5 gene.

Alleles↗

The neutral theory is dead. Long live the neutral theory.

The neutral theory of molecular evolution has been instrumental in organizing our thinking about the nature of evolutionary forces shaping variation at the DNA level. More importantly, it has provided empiricists with a strong set of testable predictions and hence, a useful null hypothesis against which to test for the presence of selection. Evidence indicates that the neutral theory cannot explain key features of protein evolution nor patterns of biased codon usage in certain species. Whereas we now have a reasonable model of selection acting on synonymous changes in Drosophila, protein evolution remains poorly understood. Despite limitations in the applicability of the neutral theory, it is likely to remain an integral part of the quest to understand molecular evolution.

Animals↗

Sequence of a 7.8 kb segment on the left arm of yeast chromosome XI reveals four open reading frames, including the CAP1 gene, an intron-containing gene and a gene encoding a homolog to the mammalian UOG-1 gene.

We report here the DNA sequence of a segment of chromosome XI of Saccharomyces cerevisiae extending over 7.8 kb. The segment contains four long open reading frames, YKL150, YKL153, YKL155 and YKL156, YKL155 corresponds to the CAP1 gene. YKL153 contains an intron and shows an extremely biased codon usage suggestive of a highly expressed protein. YKL156 is a homolog to UOG-1, an open reading frame associated with the cDNA clone of the mammalian growth/differentiation factor 1. YKL150 reveals common motifs to both the RNA polymerase II elongation factor of Drosophila melanogaster and to the yeast PPR2 gene product.

Actin Capping Proteins↗

Two distinct yeast proteins are related to the mammalian ribosomal polypeptide L7.

The RLP7 gene of Saccharomyces cerevisiae was cloned, sequenced and localized to the right arm of chromosome XIV, close to the centromere. It encodes a predicted polypeptide (RLP7p) of 322 amino acids, with a calculated molecular mass of 36 kDa and an isoelectric point of 9.6. Putative open reading frames very similar to RLP7 are present in two other yeasts, Kluyveromyces lactis and Candida utilis. The RLP7p gene product has significant sequence similarity to the S. cerevisiae YL8 polypeptide of the large ribosomal subunit (Mizuta et al., 1992), itself homologous to the L7 subunit of mammalian ribosomes. However, RLP7p and YL8 do not functionally replace each other, since an rlp7-delta::HIS3 strain is completely inviable. Judging from its predicted mass, isoelectric point and amino acid sequence, RLP7p does not correspond to any ribosomal component biochemically identified so far in S. cerevisiae, and also differs from all known ribosomal proteins by the low codon usage bias of its gene.

Amino Acid Sequence↗

Two-dimensional protein map of Saccharomyces cerevisiae: construction of a gene-protein index.

This publication marks the beginning of the construction of a gene-protein index that relates proteins which are resolved on the two-dimensional protein map of Saccharomyces cerevisiae with their corresponding genes. We report the identification of 36 novel polypeptide spots on the yeast protein map. They correspond to the products of 26 genes. Together with the polypeptide spots previously identified, this raises to 41 the number of genes whose products have been identified on the protein map. The proteins identified here are concerned with four major areas of yeast cellular physiology: carbon metabolism, heat shock, amino acid biosynthesis and purine biosynthesis. Given the molecular weight and isoelectric point of the identified proteins, and the codon-usage bias of the corresponding genes, it can be estimated that 25 to 35% of all the soluble yeast proteins are detectable under the labelling and running gel conditions used in this study.

Amino Acid Sequence↗

Heterologous gene expression in a membrane-protein-specific system.

We have constructed an expression system for heterologous proteins which uses the molecular machinery responsible for the high level production of bacteriorhodopsin in Halobacterium salinarum. Cloning vectors were assembled that fused sequences of the bacterio-opsin gene (bop) to coding sequences of heterologous genes and generated DNA fragments with cloning sites that permitted transfer of fused genes into H. salinarum expression vectors. Gene fusions include: (i) carboxyl-terminal-tagged bacterio-opsin; (ii) a carboxyl-terminal fusion with the catalytic subunit of the Escherichia coli aspartate transcarbamylase; (iii) the human muscarinic receptor, subtype M1; (iv) the human serotonin receptor, type 5HT2c; and (v) the yeast alpha mating factor receptor, Ste2. Characterization of the expression of these fusions revealed that the bop gene coding region contains previously undescribed molecular determinants which are critical for high level expression. For example, introduction of immunogenic and purification tag sequences into the C-terminal coding region significantly decreased bop gene mRNA and protein accumulation. The bacteriorhodopsin-aspartate transcarbamylase fusion protein was expressed at 7 mg per liter of culture, demonstrating that E. coli codon usage bias did not limit the system's potential for high level expression. The work presented describes initial efforts in the development of a novel heterologous protein expression system, which may have unique advantages for producing multiple milligram quantities of membrane-associated proteins.

Amino Acid Sequence↗

Structure, evolution and expression of the mitochondrial ADP/ATP translocator gene from Chlamydomonas reinhardtii.

The first AUG in the Chlamydomonas reinhardtii ADP/ATP translocator (CRANT) mRNA initiates an open reading frame (ORF) which is very similar (51-79% amino acid identity) to other ANT proteins. In contrast to higher plants, no evidence for a long amino-terminal extension was obtained. The 5' non-transcribed region of the single-copy CRANT gene contains sequence motifs present in other C. reinhardtii nuclear genes. Four introns, whose positions are not conserved in other ANT genes, interrupt the protein coding region. A short heat shock specifically reduces CRANT mRNA levels. CRANT mRNA levels were unaffected by a mutation in photosynthesis. In a dark/light regime CRANT mRNA levels are high in the dark phase and low in the early light phase. Data on translation initiation sites, splice junctions and the codon preferences of C. reinhardtii nuclear genes were compiled. With the exception of two rare codons, ACA and GGA, the CRANT gene exhibits the biased codon usage of C. reinhardtii nuclear genes that are highly expressed during normal vegetative growth.

Amino Acid Sequence↗

Nucleotide sequences encoding and promoting expression of three antibiotic resistance genes indigenous to Streptomyces.

Promoter-probe plasmid vectors were used to isolate putative promoter-containing DNA fragments of three Streptomyces antibiotic resistance genes, the rRNA methylase (tsr) gene of S. azureus, the aminoglycoside phosphotransferase (aph) gene of S. fradiae, and the viomycin phosphotransferase (vph) gene of S. vinaceus. DNA sequence analysis was carried out for all three of the fragments and for the protein-coding regions of the tsr and vph genes. No sequences resembling typical E. coli promoters or Bacillus vegetatively-expressed promoters were identified. Furthermore, none of the three DNA fragments found to be transcriptionally active in Streptomyces could initiate transcription when introduced into E. coli. An extremely biased codon usage pattern that reflects the high G + C composition of Streptomyces DNA was observed for the protein-coding regions of the tsr and vph genes, and of the previously sequenced aph gene. This pattern enabled delineation of the protein-coding region and identification of the coding strand of the genes.

Base Sequence↗

The maize plastid psbB-psbF-petB-petD gene cluster: spliced and unspliced petB and petD RNAs encode alternative products.

The chloroplast psbB, psbF, petB, and petD genes are cotranscribed and give rise to many overlapping RNAs. The mechanism and significance of this mode of expression are of interest, particularly because the accumulation of the psb and pet gene products respond differently to both light and, in C4 species such as maize, developmental signals. We present an analysis of the maize psbB, psbF, petB, and petD genes and intergenic regions. The genes are organized similarly in maize (a C4 species) and in several C3 species. Functional class II-like introns interrupt the 5' ends of petB and petD. Both spliced and unspliced RNAs accumulate; these encode alternative forms of the petB and petD proteins, differing at their N-termini. Promoter-like elements between psbF and petB, and biased codon usage suggest that the differential regulation of the psb and pet genes might be achieved at both the transcriptional and translational levels.

Amino Acid Sequence↗