Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Regional base composition variation along yeast chromosome III: evolution of chromosome primary structure.

The recent determination of the complete sequence of chromosome III from the yeast Saccharomyces cerevisiae allows, for the first time, the investigation of the long range primary structure of a eukaryotic chromosome. We have found that, against a background G+C level of about 35%, there are two regions (one in each chromosome arm) in which G+C values rise to over 50%. This effect is seen in silent sites within genes, but not in noncoding intergenic sequences. The variation in G+C content is not related to differential selection of synonymous codons, and probably reflects mutational biases. That the intergenic regions do not exhibit the same phenomenon is particularly interesting, and suggests that they are under substantial constraint. The yeast chromosome may be a model of the structure of the human genome, since there is evidence that it is also a mosaic of long regions of different base compositions, reflected in wide variation of G+C content at silent sites among genes. Two possible causes of this regional effect, replication timing, and recombination frequency, are discussed.

Animals↗

An Integrated Sequence-Structure Database incorporating matching mRNA sequence, amino acid sequence and protein three-dimensional structure data.

We have constructed a non-homologous database, termed the Integrated Sequence-Structure Database (ISSD) which comprises the coding sequences of genes, amino acid sequences of the corresponding proteins, their secondary structure and straight phi,psi angles assignments, and polypeptide backbone coordinates. Each protein entry in the database holds the alignment of nucleotide sequence, amino acid sequence and the PDB three-dimensional structure data. The nucleotide and amino acid sequences for each entry are selected on the basis of exact matches of the source organism and cell environment. The current version 1.0 of ISSD is available on the WWW at http://www.protein.bio.msu.su/issd/ and includes 107 non-homologous mammalian proteins, of which 80 are human proteins. The database has been used by us for the analysis of synonymous codon usage patterns in mRNA sequences showing their correlation with the three-dimensional structure features in the encoded proteins. Possible ISSD applications include optimisation of protein expression, improvement of the protein structure prediction accuracy, and analysis of evolutionary aspects of the nucleotide sequence-protein structure relationship.

Algorithms↗

ISSD Version 2.0: taxonomic range extended.

Two more organisms from different taxonomic groups were added to a new version of the Integrated Sequence-Structure Database (ISSD). ISSD serves as an integrated source of sequence and structure information for the analysis of correlations between mRNA synonymous codon usage and three-dimensional structure of the encoded proteins. ISSD now holds 88 non-homologous Escherichia coli proteins and 25 yeast Saccharomyces cerevisiae proteins in addition to the expanded set of mammalian proteins, which includes 166 proteins (107 in ISSD Version 1.0). Comparison of ISSD sequences with organism-specific codon usage data derived from CUTG database shows that it is a representative subset of the GenBank coding sequences data. Preliminary results of the statistical analysis confirm that sequence-structure correlations observed by us earlier are also present in the upgraded ISSD (Version 2.0), including bacterial and yeast proteins. The ISSD Version 2.0 release includes an improved Web-based data search and retrieval system and is accessible via URL http://www.protein.bio.msu.su/issd/. ISSD can be also accessed at ExPASy, URL http://www.expasy.ch/swissmod/swiss-model.htm l

Animals↗

Secondary structure of MS2 phage RNA and bias in code word usage.

Based on the secondary structural model of MS2 RNA, it is shown that, in base-pairing regions of the RNA, there is a bias in the use of synonymous codons which favours C and/or G over U and/or A in the third codon positions, and that in non-pairing regions, there is an opposite bias which favours U and/or A over C and/or G. This nature is interpreted as a result of selective constraint which stabilises the secondary structure of the single-stranded RNA genome of the MS2 phage.

Base Sequence↗

Thermoadaptation trait revealed by the genome sequence of thermophilic Geobacillus kaustophilus.

We present herein the first complete genome sequence of a thermophilic Bacillus-related species, Geobacillus kaustophilus HTA426, which is composed of a 3.54 Mb chromosome and a 47.9 kb plasmid, along with a comparative analysis with five other mesophilic bacillar genomes. Upon orthologous grouping of the six bacillar sequenced genomes, it was found that 1257 common orthologous groups composed of 1308 genes (37%) are shared by all the bacilli, whereas 839 genes (24%) in the G.kaustophilus genome were found to be unique to that species. We were able to find the first prokaryotic sperm protamine P1 homolog, polyamine synthase, polyamine ABC transporter and RNA methylase in the 839 unique genes; these may contribute to thermophily by stabilizing the nucleic acids. Contrasting results were obtained from the principal component analysis (PCA) of the amino acid composition and synonymous codon usage for highlighting the thermophilic signature of the G.kaustophilus genome. Only in the PCA of the amino acid composition were the Bacillus-related species located near, but were distinguishable from, the borderline distinguishing thermophiles from mesophiles on the second principal axis. Further analysis revealed some asymmetric amino acid substitutions between the thermophiles and the mesophiles, which are possibly associated with the thermoadaptation of the organism.

Adaptation, Physiological↗

New Onto-Tools: Promoter-Express, nsSNPCounter and Onto-Translate.

The Onto-Tools suite is composed of an annotation database and eight complementary, web-accessible data mining tools: Onto-Express, Onto-Compare, Onto-Design, Onto-Translate, Onto-Miner, Pathway-Express, Promoter-Express and nsSNPCounter. Promoter-Express is a new tool added to the Onto-Tools ensemble that facilitates the identification of transcription factor binding sites active in specific conditions. nsSNPCounter is another new tool that allows computation and analysis of synonymous and non-synonymous codon substitutions for studying evolutionary rates of protein coding genes. Onto-Translate has also been enhanced to expand its scope and accuracy by fully utilizing the capabilities of the Onto-Tools database. Currently, Onto-Translate allows arbitrary mappings between 28 types of IDs for 53 organisms. Onto-Tools are freely available at http://vortex.cs.wayne.edu/Projects.html.

Binding Sites↗

Orotate phosphoribosyltransferase from Thermus thermophilus: overexpression in Escherichia coli, purification and characterization.

Orotate phosphoribosyltransferase (OPRTase, EC2.4.2.10) plays a role in de novo synthesis of pyrimidine nucleotide and transfers orotate to 5-phosphoribosyl-1-pyrophosphate (PRPP) to form orotidine-5'-monophosphate (OMP). To obtain heat-stable OPRTase and to elucidate the mechanism of heat stability, this enzyme from Thermus thermophilus was expressed in Escherichia coli and purified. The pyrE gene of T. thermophilus which encodes OPRTase, contains an open reading frame of 549 base pairs with 69% G+C content. Since this gene expressed itself inefficiently in E. coli, the 5' and 3' ends of the coding regions were replaced with synonymous codons which contain more A+T and corresponds to major codons for E. coli. Introduction of the modified gene fragments into a plasmid having a tac promoter resulted in production of a polypeptide of molecular weight (M(r)) 20,000 in the presence of isopropyl-beta-D-thiogalactopyranoside (IPTG) in E. coli. This protein represented as much as 16% of the bacterial total protein and showed the OPRTase activity. Three purification steps, consisting of heat treatment at 65 degrees C, 40% ammonium sulfate fractionation, and KCl gradient elution from DEAE-Sephadex A-50, resulted in highly purified single polypeptide. The optimum activity of the purified OPRTase was observed at 150 mM KCl, pH 9.0, 75-80 degrees C, and in the presence of 100 microM PRPP. The activation energy of this enzyme reaction was 20.3 kJ/mol. The Km of this enzyme for orotate as a substrate was 75 microM and the maximum specific activity was 300 units/mg protein under the optimum conditions. The purified OPRTase was stable for 20 min at 85 degrees C.

Amino Acid Sequence↗

Elevated rates of nonsynonymous substitution in island birds.

Slightly deleterious mutations are expected to fix at relatively higher rates in small populations than in large populations. Support for this prediction of the nearly-neutral theory of molecular evolution comes from many cases in which lineages inferred to differ in long-term average population size have different rates of nonsynonymous substitution. However, in most of these cases, the lineages differ in many other ways as well, leaving open the possibility that some factor other than population size might have caused the difference in substitution rates. We compared synonymous and nonsynonymous substitutions in the mitochondrial cyt b and ND2 genes of nine closely related island and mainland lineages of ducks and doves. We assumed that island taxa had smaller average population sizes than those of their mainland sister taxa for most of the time since they were established. In all nine cases, more nonsynonymous substitutions occurred on the island branch, but synonymous substitutions showed no significant bias. As in previous comparisons of this kind, the lineages with smaller populations might differ in other respects that tend to increase rates of nonsynonymous substitution, but here such differences are expected to be slight owing to the relatively recent origins of the island taxa. An examination of changes to apparently "preferred" and "unpreferred" synonymous codons revealed no consistent difference between island and mainland lineages.

Amino Acid Sequence↗

A comparative mitogenomic analysis of the potential adaptive value of Arctic charr mtDNA introgression in brook charr populations (Salvelinus fontinalis Mitchill).

Wild brook charr populations (Salvelinus fontinalis) completely introgressed with the mitochondrial genome (mtDNA) of arctic charr (Salvelinus alpinus) are found in several lakes of northeastern Québec, Canada. Mitochondrial respiratory enzymes of these populations are thus encoded by their own nuclear DNA and by arctic charr mtDNA. In the present study we performed a comparative sequence analysis of the whole mitochondrial genome of both brook and arctic charr to identify the distribution of mutational differences across these two genomes. This analysis revealed 47 amino acid replacements, 45 of which were confined to subunits of the NADH dehydrogenase complex (Complex I), one in the cox3 gene (Complex IV), and one in the atp8 gene (Complex V). A cladistic approach performed with brook charr, arctic charr, and two other salmonid fishes (rainbow trout [Oncorhynchus mykiss] and Atlantic salmon [Salmo salar]) revealed that only five amino acid replacements were specific to the charr comparison and not shared with the other two salmonids. In addition, five amino acid substitutions localized in the nad2 and nad5 genes denoted negative scores according to the functional properties of amino acids and, therefore, could possibly have an impact on the structure and functional properties of these mitochondrial peptides. The comparison of both brook and arctic charr mtDNA with that of rainbow trout also revealed a relatively constant mutation rate for each specific gene among species, whereas the rate was quite different among genes. This pattern held for both synonymous and nonsynonymous nucleotide positions. These results, therefore, support the hypothesis of selective constraints acting on synonymous codon usage.

Amino Acid Substitution↗

Determinants of substitution rates in mammalian genes: expression pattern affects selection intensity but not mutation rate.

To determine whether gene expression patterns affect mutation rates and/or selection intensity in mammalian genes, we studied the relationships between substitution rates and tissue distribution of gene expression. For this purpose, we analyzed 2,400 human/rodent and 834 mouse/rat orthologous genes, and we measured (using expressed sequence tag data) their expression patterns in 19 tissues from three development states. We show that substitution rates at nonsynonymous sites are strongly negatively correlated with tissue distribution breadth: almost threefold lower in ubiquitous than in tissue-specific genes. Nonsynonymous substitution rates also vary considerably according to the tissues: the average rate is twofold lower in brain-, muscle-, retina- and neuron-specific genes than in lymphocyte-, lung-, and liver-specific genes. Interestingly, 5' and 3' untranslated regions (UTRs) show exactly the same trend. These results demonstrate that the expression pattern is an essential factor in determining the selective pressure on functional sites in both coding and noncoding regions. Conversely, silent substitution rates do not vary with expression pattern, even in ubiquitously expressed genes. This latter result thus suggests that synonymous codon usage is not constrained by selection in mammals. Furthermore, this result also indicates that there is no reduction of mutation rates in genes expressed in the germ line, contrary to what had been hypothesized based on the fact that transcribed DNA is more efficiently repaired than nontranscribed DNA.

Animals↗

The complete mitochondrial DNA sequence of the horseshoe crab Limulus polyphemus.

We determined the complete 14,985-nt sequence of the mitochondrial DNA of the horseshoe crab Limulus polyphemus (Arthropoda: Xiphosura). This mtDNA encodes the 13 protein, 2 rRNA, and 22 tRNA genes typical for metazoans. The arrangement of these genes and about half of the sequence was reported previously; however, the sequence contained a large number of errors, which are corrected here. The two strands of Limulus mtDNA have significantly different nucleotide compositions. The strand encoding most mitochondrial proteins has 1. 25 times as many A's as T's and 2.33 times as many C's as G's. This nucleotide bias correlates with the biases in amino acid content and synonymous codon usage in proteins encoded by different strands and with the number of non-Watson-Crick base pairs in the stem regions of encoded tRNAs. The sizes of most mitochondrial protein genes in Limulus are either identical to or slightly smaller than those of their Drosophila counterparts. The usage of the initiation and termination codons in these genes seems to follow patterns that are conserved among most arthropod and some other metazoan mitochondrial genomes. The noncoding region of Limulus mtDNA contains a potential stem-loop structure, and we found a similar structure in the noncoding region of the published mtDNA of the prostriate tick Ixodes hexagonus. A simulation study was designed to evaluate the significance of these secondary structures; it revealed that they are statistically significant. No significant, comparable structure can be identified for the metastriate ticks Rhipicephalus sanguineus and Boophilus microplus. The latter two animals also share a mitochondrial gene rearrangement and an unusual structure of mt-tRNA(C) that is exactly the same association of changes as previously reported for a group of lizards. This suggests that the changes observed are not independent and that the stem-loop structure found in the noncoding regions of Limulus and Ixodes mtDNA may play the same role as that between trnN and trnC in vertebrates, i.e., the role of lagging strand origin of replication.

Animals↗

Isolation and molecular phylogenetic analysis of actin-coding regions from Emiliania huxleyi, a Prymnesiophyte alga, by reverse transcriptase and PCR methods.

Reverse transcriptase and polymerase chain reaction methods were used to amplify and clone actin cDNAs from the chlorophylls a + C-containing unicellular alga, Emiliania huxleyi (Prymnesiophyta). Actins in E. huxleyi are defined by a gene family containing at least six distinct coding regions that were derived from relatively recent gene duplications. Five of the coding regions (types 1, 2, and 4-6) varied only among synonymous codons. A nonsynonomous change in a sixth coding region (type 3 actin) produced a serine-to-phenylalanine replacement. The G + C composition of third positions in E. huxleyi actin genes is 98%, which contrasts with the mean value of 50% G + C content for first and second positions. Distance-matrix and parsimony analyses of actin genes identified the prymnesiophytes as a photosynthetic lineage that is not already related to other eukaryotic algal groups.

Actins↗

Molecular drift of the bride of sevenless (boss) gene in Drosophila.

DNA sequences were determined for three to five alleles of the bride-of-sevenless (boss) gene in each of four species of Drosophila. The product of boss is a transmembrane receptor for a ligand coded by the sevenless gene that triggers differentiation of the R7 photoreceptor cell in the compound eye. Population parameters affecting the rate and pattern of molecular evolution of boss were estimated from the multinomial configurations of nucleotide polymorphisms of synonymous codons. The time of divergence between D. melanogaster and D. simulans was estimated as approximately 1 Myr, that between D. teissieri and D. yakuba as approximately 0.75 Myr, and that between the two pairs of sibling species as approximately 2 Myr. (The boss genes themselves have estimated divergence times approximately 50% greater than the species divergence times.) The effective size of the species was estimated as approximately 5 x 10(6), and the average mutation rate was estimated as 1-2 x 10(-9)/nucleotide/generation. The ratio of amino acid polymorphisms within species to fixed differences between species suggests that approximately 25% of all possible single-step amino acid replacements in the boss gene product may be selectively neutral or nearly neutral. The data also imply that random genetic drift has been responsible for virtually all of the observed differences in the portion of the boss gene analyzed among the four species.

Alcohol Dehydrogenase↗

Reduced natural selection associated with low recombination in Drosophila melanogaster.

Synonymous codons are not used equally in many organisms, and the extent of codon bias varies among loci. Earlier studies have suggested that more highly expressed loci in Drosophila melanogaster are more biased, consistent with findings from several prokaryotes and unicellular eukaryotes that codon bias is partly due to natural selection for translational efficiency. We link this model of varying selection intensity to the population-genetics prediction that the effectiveness of natural selection is decreased under reduced recombination. In analyses of 385 D. melanogaster loci, we find that codon bias is reduced in regions of low recombination (i.e., near centromeres and telomeres and on the fourth chromosome). The effect does not appear to be a linear function of recombination rate; rather, it seems limited to regions with the very lowest levels of recombination. The large majority of the genome apparently experiences recombination at a sufficiently high rate for effective natural selection against suboptimal codons. These findings support models of the Hill-Robertson effect and genetic hitchhiking and are largely consistent with multiple reports of low levels of DNA sequence variation in regions of low recombination.

Animals↗

Sequence of the ebgA gene of Escherichia coli: comparison with the lacZ gene.

We have sequenced the ebgA (evolved beta-galactosidase) gene of Escherichia coli K12. The sequence shows 50% nucleotide identity with the E. coli lacZ gene, demonstrating that the two genes are related by descent from a common ancestral gene. Comparison of the two sequences suggests that the ebgA gene has recently been under selection. A significant excess of identical, rather than synonymous, codons used to encode identical amino acids at the same positions in the aligned sequences implies that some form of selection is operating directly at the DNA level. This selection is independent of, and in addition to, selection based on codon usage or on function of the gene products.

Amino Acid Sequence↗

DNA and the neutral theory.

The neutral theory claims that the great majority of evolutionary changes at the molecular (DNA) level are caused not by Darwinian selection but by random fixation of selectively neutral or nearly neutral mutants. The theory also asserts that the majority of protein and DNA polymorphisms are selectively neutral and that they are maintained in the species by mutational input balanced by random extinction. In conjunction with diffusion models (the stochastic theory) of gene frequencies in finite populations, it treats these phenomena in quantitative terms based on actual observations. Although the theory has been strongly criticized by the 'selectionists', supporting evidence has accumulated over the years. Particularly, the recent outburst of DNA sequence data lends strong support to the theory both with respect to evolutionary base substitutions and DNA polymorphism, including rapid evolutionary base substitutions in pseudogenes. In addition, the observed pattern of synonymous codon choice can now be readily explained in the framework of this theory. I review these recent findings in the light of the neutral theory.

Animals↗

DNA sequence evolution: the sounds of silence.

Silent sites (positions that can undergo synonymous substitutions) in protein-coding genes can illuminate two evolutionary processes. First, despite being silent, they may be subject to natural selection. Among eukaryotes this is exemplified by yeast, where synonymous codon usage patterns are shaped by selection for particular codons that are more efficiently and/or accurately translated by the most abundant tRNAs; codon usage across the genome, and the abundance of different tRNA species, are highly co-adapted. Second, in the absence of selection, silent sites reveal underlying mutational patterns. Codon usage varies enormously among human genes, and yet silent sites do not appear to be influenced by natural selection, suggesting that mutation patterns vary among regions of the genome. At first, the yeast and human genomes were thought to reflect a dichotomy between unicellular and multicellular organisms. However, it now appears that natural selection shapes codon usage in some multicellular species (e.g. Drosophila and Caenorhabditis), and that regional variations in mutation biases occur in yeast. Silent sites (in serine codons) also provide evidence for mutational events changing adjacent nucleotides simultaneously.

Animals↗

Evolutionary relationships within a subgroup of HERV-K-related human endogenous retroviruses.

The prototype endogenous retrovirus HERV-K10 was identified in the human genome by its homology to the exogenous mouse mammary tumour virus. By analysis of a short 244 bp segment of the reverse transcriptase (RT) gene of other HERV-K10-like sequences, it has become clear that these elements represent an extended family consisting of multiple groups (the HML-1 to HML-6 subgroups). Some of these elements are transcriptionally active and contain an intact open reading frame for the RT protein, raising the possibility that this family is still expanding through retrotransposition. To better define the relationship of these endogenous retroviruses, we identified ten new members of the HML-2 subgroup. PCR was used to amplify reverse-transcribed RNA of a 595 bp region of the RT gene in a variety of human cell samples, including normal and leukaemic bone marrow and peripheral blood, placenta cells and a transformed T cell line. We provide an extensive phylogenetic analysis of the relationships for this cluster of HERV-K-related endogenous retroviral elements. Nucleotide diversity values for nonsynonymous versus synonymous codon positions indicate that moderately strong selection is or was operating on these retroviral RT gene segments. The evolution of this class of endogenous retroelements is discussed.

Base Sequence↗