Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Selectionism and neutralism in molecular evolution.

Charles Darwin proposed that evolution occurs primarily by natural selection, but this view has been controversial from the beginning. Two of the major opposing views have been mutationism and neutralism. Early molecular studies suggested that most amino acid substitutions in proteins are neutral or nearly neutral and the functional change of proteins occurs by a few key amino acid substitutions. This suggestion generated an intense controversy over selectionism and neutralism. This controversy is partially caused by Kimura's definition of neutrality, which was too strict (|2Ns|< or =1). If we define neutral mutations as the mutations that do not change the function of gene products appreciably, many controversies disappear because slightly deleterious and slightly advantageous mutations are engulfed by neutral mutations. The ratio of the rate of nonsynonymous nucleotide substitution to that of synonymous substitution is a useful quantity to study positive Darwinian selection operating at highly variable genetic loci, but it does not necessarily detect adaptively important codons. Previously, multigene families were thought to evolve following the model of concerted evolution, but new evidence indicates that most of them evolve by a birth-and-death process of duplicate genes. It is now clear that most phenotypic characters or genetic systems such as the adaptive immune system in vertebrates are controlled by the interaction of a number of multigene families, which are often evolutionarily related and are subject to birth-and-death evolution. Therefore, it is important to study the mechanisms of gene family interaction for understanding phenotypic evolution. Because gene duplication occurs more or less at random, phenotypic evolution contains some fortuitous elements, though the environmental factors also play an important role. The randomness of phenotypic evolution is qualitatively different from allele frequency changes by random genetic drift. However, there is some similarity between phenotypic and molecular evolution with respect to functional or environmental constraints and evolutionary rate. It appears that mutation (including gene duplication and other DNA changes) is the driving force of evolution at both the genic and the phenotypic levels.

Animals↗

New methods for estimating the numbers of synonymous and nonsynonymous substitutions.

New methods for estimating the numbers of synonymous and nonsynonymous substitutions per site were developed. The methods are unweighted pathway methods based on Kimura's two-parameter model. Computer simulations were conducted to evaluate the accuracies of the new methods, Nei and Gojobori's (NG) method, Miyata and Yasunaga's (MY) method, Li, Wu, and Luo's (LWL) method, and Pamilo, Bianchi, and Li's (PBL) method. The following results were obtained: (1) The NG, MY, and LWL methods give overestimates of the number of synonymous substitutions and underestimates of the number of nonsynonymous substitutions. The major cause for the biased estimation is that these three methods underestimate the number of synonymous sites and overestimate the number of nonsynonymous sites. (2) The PBL method gives better estimates of the numbers of synonymous and nonsynonymous substitutions than those obtained by the NG, MY, and LWL methods. (3) The new methods also give better estimates of the numbers of synonymous and nonsynonymous substitutions than those obtained by the NG, MY, and LWL methods. In addition, estimates of the numbers of synonymous and nonsynonymous sites obtained by the new methods are reasonably accurate. (4) In some cases, the new methods and the PBL method give biased estimates of substitution numbers. However, from the number of nucleotide substitutions at the third position of codons, we can examine whether estimates obtained by the new methods are good or not, whereas we cannot make an examination of estimates obtained by the PBL method. (5) When there are strong transition/transversion and nucleotide-frequency biases like mitochondrial genes, all of the above methods give biased estimates of substitution numbers. In such cases, Kondo et al.'s method is recommended to be used for estimating the number of synonymous substitutions, although their method cannot estimate the number of nonsynonymous substitutions and is time-consuming. These results, particularly result (1), call for reexaminations of some genes. This is because evolutionary pictures of genes have often been discussed on the basis of results obtained by the NG, MY, and LWL methods, which are favorable for the neutral theory of molecular evolution.

Animals↗

Specific features of immunoglobulin VH genes of the Antarctic teleost Trematomus bernacchii.

The somatic recombination of different germline-encoded gene segments constitutes a principal source of antibody diversity. In order to investigate the diversity in recombined gene segments encoding the immunoglobulin heavy chain of the Antarctic teleost Trematomus bernacchii, a VH library was constructed by 5'-RACE (rapid amplification of cDNA ends) using RNA isolated from the spleen of an individual specimen. Analysis of cDNA sequences of 45 rearranged VH/D/JH segments revealed specific features, such as: high number of repeats, up to 8 bp long, and palindromic sequences, especially in CDRs (complementary determining regions); occurrence of the RGYW consensus, known as mutational hot spot, higher than in other species. Sixty-four percent of single base substitutions was found within this motif. In addition, the usage of serine codons showed a clear bias for AGY in CDRs, particularly in CDR2, and for TCN in FRs (framework regions). In CDRs, the frequency of non-synonymous changes was higher than that of synonymous changes. Diversity generated by insertions/deletions occurred more often than in other species; inserted bases were often repeats of adjacent bases. In particular the CDR2 showed the highest length variability as compared to other species. Alignment of VH sequences indicated that also the gene conversion mechanism may contribute to generating diversity. These data indicate a CDR mutability higher than in other species and provide some insights into the hypermutational events that may also occur in teleosts.

Animals↗

Extent of gene duplication in the genomes of Drosophila, nematode, and yeast.

We conducted a detailed analysis of duplicate genes in three complete genomes: yeast, Drosophila, and Caenorhabditis elegans. For two proteins belonging to the same family we used the criteria: (1) their similarity is > or =I (I = 30% if L > or = 150 a.a. and I = 0.01n + 4.8L(-0.32(1 + exp(-L/1000))) if L < 150 a.a., where n = 6 and L is the length of the alignable region), and (2) the length of the alignable region between the two sequences is > or = 80% of the longer protein. We found it very important to delete isoforms (caused by alternative splicing), same genes with different names, and proteins derived from repetitive elements. We estimated that there were 530, 674, and 1,219 protein families in yeast, Drosophila, and C. elegans, respectively, so, as expected, yeast has the smallest number of duplicate genes. However, for the duplicate pairs with the number of substitutions per synonymous site (K(S)) < 0.01, Drosophila has only seven pairs, whereas yeast has 58 pairs and nematode has 153 pairs. After considering the possible effects of codon usage bias and gene conversion, these numbers became 6, 55, and 147, respectively. Thus, Drosophila appears to have much fewer young duplicate genes than do yeast and nematode. The larger numbers of duplicate pairs with K(S) < 0.01 in yeast and C. elegans were probably largely caused by block duplications. At any rate, it is clear that the genome of Drosophila melanogaster has undergone few gene duplications in the recent past and has much fewer gene families than C. elegans.

Animals↗

A new nuclear gene for insect phylogenetics: dopa decarboxylase is informative of relationships within Heliothinae (Lepidoptera: Noctuidae).

The lack of a readily accessible roster of nuclear genes informative at various taxonomic levels is a bottleneck for molecular systematics. In this report, we describe the first phylogenetic application of the sequence that encodes the enzyme dopa decarboxylase (DDC). For 14 test species within the noctuid moth subfamily Heliothinae that represent the previously best-supported groupings, a 690-bp fragment of DDC resolved relationships that are largely concordant with prior evidence from elongation factor-1 alpha (EF-1 alpha), morphology, and allozymes. Although both synonymous and nonsynonymous changes occur in DDC substantially more rapidly than they do in EF-1 alpha, DDC divergences within Heliothinae are below saturation at all codon positions. Analysis of DDC and EF-1 alpha in combination resulted in increased bootstrap support for several groupings. As a first estimate of previously unresolved relationships, DDC sequences were analyzed from 16 additional heliothines, for a total of 30 heliothine species plus outgroups. Previous relationships based on DDC were generally stable with increased taxon sampling, although a two- to eightfold downweighting of codon position 3 was required for complete concordance with the 14-species result. The weighted strict consensus trees were largely resolved and were congruent with most although not all previous hypotheses based on either morphology or EF-1 alpha. The proposed phylogeny suggests that the major agricultural pest heliothines belong to a single clade, characterized by polyphagy and associated life history traits, within this largely host-specific moth subfamily. DDC holds much promise for phylogenetic analysis of Tertiary-age animal groups.

Animals↗

Evidence for genetic drift in endosymbionts (Buchnera): analyses of protein-coding genes.

Buchnera, the bacterial endosymbionts of aphids, undergo severe population bottlenecks during maternal transmission through their hosts. Previous studies suggest an increased effect of drift within these strictly asexual, small populations, resulting in an increased fixation of slightly deleterious mutations. This study further explores sequence evolution in Buchnera using three approaches. First, patterns of codon usage were compared across several homologous Escherichia coli and Buchnera loci, in order to test the prediction that selection for the use of optimal codons is less effective in small populations. A chi 2-based measure of codon bias was developed to adjust for the overall A + T richness of silent positions in the endosymbionts. In contrast to E. coli homologues, adaptive codon bias across Buchnera loci is markedly low, and patterns of codon usage lack a strong relationship with gene expression level. These data suggest that codon usage in Buchnera has been shaped largely by mutational pressure and drift rather than by selection for translational efficiency. One exception to the overall lack of bias is groEL, which is known to be constitutively overexpressed in Buchnera and other endosymbionts. Second, relative-rate tests show elevated rates of sequence evolution of numerous protein-coding loci across Buchnera, compared to E. coli. Finally, consistently higher ratios of nonsynonymous to synonymous substitutions in Buchnera loci relative to the enteric bacteria strongly suggest the accumulation of nonsynonymous substitutions in endosymbiont lineages. Combined, these results suggest a decreased effectiveness of purifying selection in purging endosymbiont populations of slightly deleterious mutations, particularly those affecting codon usage and amino acid identity.

Animals↗

Concerted evolution at a multicopy locus in the protozoan parasite Theileria parva: extreme divergence of potential protein-coding sequences.

Concerted evolution of multicopy gene families in vertebrates is recognized as an important force in the generation of biological novelty but has not been documented for the multicopy genes of protozoa. A multicopy locus, Tpr, which consists of tandemly arrayed open reading frames (ORFs) containing several repeated elements has been described for Theileria parva. Herein we show that probes derived from the 5'/N-terminal ends of ORFs in the genomic DNAs of T. parva Uganda (1,108 codons) and Boleni (699 codons) hybridized with multicopy sequences in homologous DNA but did not detect similar sequences in the DNA of 14 heterologous T. parva stocks and clones. The probe sequences were, however, protein coding according to predictive algorithms and codon usage. The 3'/C-terminal ends of the Uganda and Boleni ORFs exhibited 75% similarity and identity, respectively, to the previously identified Tpr1 and Tpr2 repetitive elements of T. parva Muguga. Tpr1-homologous sequences were detected in two additional species of Theileria. Eight different Tpr1-homologous transcripts were present in piroplasm mRNA from a single T. parva Muguga-infected animal. The Tpr1 and Tpr2 amino acid sequences contained six predicted membrane-associated segments. The ratio of synonymous to nonsynonymous substitutions indicates that Tpr1 evolves like protein-encoding DNA. The previously determined nucleotide sequence of the gene encoding the p67 antigen is completely identical in T. parva Muguga, Boleni, and Uganda, including the third base in codons. The data suggest that concerted evolution can lead to the radical divergence of coding sequences and that this can be a mechanism for the generation of novel genes.

Amino Acid Sequence↗

Comparative genomics of the Mill family: a rapidly evolving MHC class I gene family.

Mill (MHC class I-like located near the leukocyte receptor complex) is a novel family of class I genes identified in mice that is most closely related to the human MICA/B family. In the present study, we isolated Mill cDNA from rats and carried out a comparative genomic analysis. Rats have two Mill genes orthologous to mouse Mill1 and Mill2 near the leukocyte receptor complex, with expression patterns similar to those of their mouse counterparts. Interspecies sequence comparison indicates that Mill is one of the most rapidly evolving class I gene families and that non-synonymous substitutions occur more frequently than synonymous substitutions in its alpha 1 domain, implicating the involvement of Mill in immune defenses. Interestingly, the alpha 2 domain of rat Mill2 contains a premature stop codon in many inbred strains, indicating that Mill2 is not essential for survival. A computer search of the database identified a horse Mill-like expressed sequence tag, indicating that Mill emerged before the radiation of mammals. Hence, the failure to find Mill in human indicates strongly that it was lost from the human lineage. Our present work provides convincing evidence that Mill is akin to the MICA/B family, yet constitutes a distinct gene family.

Amino Acid Sequence↗

Synonymous substitution rates in Drosophila: mitochondrial versus nuclear genes.

Synonymous substitution rates in mitochondrial and nuclear genes of Drosophila were compared. To make accurate comparisons, we considered the following: (1) relative synonymous rates, which do not require divergence time estimates, should be used; (2) methods estimating divergence should take into account base composition; (3) only very closely related species should be used to avoid effects of saturation; (4) the heterogeneity of rates should be examined. We modified the methods estimating synonymous substitution numbers to account for base composition bias. By using these methods, we found that mitochondrial genes have 1.7-3.4 times higher synonymous substitution rates than the fastest nuclear genes or 4.5-9.0 times higher rates than the average nuclear genes. The average rate of synonymous transversions was 2.7 (estimated from the melanogaster species subgroup) or 2.9 (estimated from the obscura group) times higher in mitochondrial genes than in nuclear genes. Synonymous transversions in mitochondrial genes occurred at an approximately equivalent rate to those in the fastest nuclear genes. This last result is not consistent with the hypothesis that the difference in turnover rates between mitochondrial and nuclear genomes is the major factor determining higher synonymous substitution rates in mtDNA. We conclude that the difference in synonymous substitution rates is due to a combination of two factors: a higher transitional mutation rate in mtDNA and constraints on nuclear genes due to selection for codon usage.

Animals↗

Host-symbiont conflicts: positive selection on an outer membrane protein of parasitic but not mutualistic Rickettsiaceae.

The Rickettsiaceae is a family of intracellular bacterial symbionts that includes both vertically transmitted parasites that spread by manipulating the reproduction of their host (Wolbachia in arthropods) and horizontally transmitted parasites (represented by Cowdria ruminantium), and mutualists (Wolbachia pipientis in nematode worms). We have investigated the nature of natural selection acting on an outer membrane protein, the wsp gene in Wolbachia and its homologue map1 in Cowdria, thought likely to be involved in host-parasite interactions in these bacteria. The ratio of nonsynonymous to synonymous substitution rates (d(N)/d(S)) at individual amino acid sites or at lineages within the gene's phylogeny was estimated using maximum likelihood models of codon substitution. The first hypothesis we tested was that this protein is under positive selection in the parasitic but not in the mutualistic Rickettsiaceae. This hypothesis was supported as positive selection and was detected in Cowdria and arthropod Wolbachia sequence evolution but not in the evolution of Wolbachia sequences from nematodes. Furthermore, this selection was concentrated outside the transmembrane region of the protein and, therefore, in the regions of the protein that may interact with the host. The second hypothesis tested was that positive selection would be stronger in the strains of arthropod Wolbachia that distort the host sex ratio than in those that induce cytoplasmic incompatibility. However, we found no support for this hypothesis. In conclusion, our results are consistent with the hypothesis that antagonistic coevolution causes faster evolution of surface protein sequences in parasites than in mutualists. Confirmation of this conclusion awaits the replication of these results both in additional genes and across more bacterial taxa. The regions of the wsp and map1 genes we identified as likely to be involved in host-parasite arms races should be examined in future studies of parasite virulence and host immune responses, and during the design of vaccines.

Amino Acid Sequence↗

Bayes empirical bayes inference of amino acid sites under positive selection.

Codon-based substitution models have been widely used to identify amino acid sites under positive selection in comparative analysis of protein-coding DNA sequences. The nonsynonymous-synonymous substitution rate ratio (d(N)/d(S), denoted omega) is used as a measure of selective pressure at the protein level, with omega > 1 indicating positive selection. Statistical distributions are used to model the variation in omega among sites, allowing a subset of sites to have omega > 1 while the rest of the sequence may be under purifying selection with omega < 1. An empirical Bayes (EB) approach is then used to calculate posterior probabilities that a site comes from the site class with omega > 1. Current implementations, however, use the naive EB (NEB) approach and fail to account for sampling errors in maximum likelihood estimates of model parameters, such as the proportions and omega ratios for the site classes. In small data sets lacking information, this approach may lead to unreliable posterior probability calculations. In this paper, we develop a Bayes empirical Bayes (BEB) approach to the problem, which assigns a prior to the model parameters and integrates over their uncertainties. We compare the new and old methods on real and simulated data sets. The results suggest that in small data sets the new BEB method does not generate false positives as did the old NEB approach, while in large data sets it retains the good power of the NEB approach for inferring positively selected sites.

Alleles↗

Evaluation of methods for determination of a reconstructed history of gene sequence evolution.

With whole-genome sequences being completed at an increasing rate, it is important to develop and assess tools to analyze them. Following annotation of the protein content of a genome, one can compare sequences with previously characterized homologous genes to detect novel functions within specific proteins in the evolution of the newly sequenced genome. One common statistical method to detect such changes is to compare the ratios of nonsynonymous (K(a)) to synonymous (K(s)) nucleotide substitution rates. Here, the effects of several parameters that can influence this calculation (sequence reconstruction method, phylogenetic tree branch length weighting, GC content, and codon bias) are examined. Also, two new alternative measures of adaptive evolution, the point accepted mutations (PAM)/neutral evolutionary distance (NED) ratio and the sequence space assessment (SSA) statistic are presented. All of these methods are compared using two sequence families: the recent divergence of leptin orthologs in primates, and the more ancient divergence of the deoxyribonucleoside kinase family. The examination of these and other measures to detect changes of gene function along branches of a phylogenetic tree will become increasingly important in the postgenomic era.

Algorithms↗

The phylogenetic utility of the codon-degeneracy model.

The codon-degeneracy model (CDM) predicts relative frequencies of substitution for any set of homologous protein-coding DNA sequences based on patterns of nucleotide degeneracy, codon composition, and the assumption of selective neutrality. However, at present, the CDM is reliant on outside estimates of transition bias. A new method by which the power of the CDM can be used to find a synonymous transition bias that is optimal for any given phylogenetic tree topology is presented. An example is illustrated that utilizes optimized transition biases to generate CDM GF-scores for every possible phylogenetic tree for pocket gophers of the genus Orthogeomys. The resulting distribution of CDM GF-scores is compared and contrasted with the results of maximum parsimony and maximum likelihood methods. Although convergence on a single tree topology by the CDM and another method indicates greater support for that particular tree, the value of CDM GF-score as the sole optimality criterion for phylogeny reconstruction remains to be determined. It is clear, however, that the a priori estimation of an optimum transition bias from codon composition has a direct application to differentiating between alternative trees.

Animals↗

Phylogeny and intraspecific variability of holoparasitic Orobanche (Orobanchaceae) inferred from plastid rbcL sequences.

The rbcL sequences of 106 specimens representing 28 species of the four recognized sections of Orobanche were analyzed and compared. Most sequences represent pseudogenes with premature stop codons. This study confirms that the American lineage (sects. Gymnocaulis and Myzorrhiza) contains potentially functional rbcL-copies with intact open reading frames and low rates of non-synonymous substitutions. For the first time, this is also shown for a member of the Eurasian lineage, O. coerulescens of sect. Orobanche, while all other investigated species of sects. Orobanche and Trionychon contain pseudogenes with distorted reading frames and significantly higher rates of non-synonymous substitutions. Phylogenetic analyses of the rbcL sequences give equivocal results concerning the monophyly of Orobanche, and the American lineage might be more closely related to Boschniakia and Cistanche than to the other sections of Orobanche. Additionally, species of sect. Trionychon phylogenetically nest in sect. Orobanche. This is in concordance with results from other plastid markers (rps2 and matK), but in disagreement with other molecular (nuclear ITS), morphological, and karyological data. This might indicate that the ancestor of sect. Trionychon has captured the plastid genome, or parts of it, of a member of sect. Orobanche. Apart from the phylogenetically problematic position of sect. Trionychon, the phylogenetic relationships within sect. Orobanche are similar to those inferred from nuclear ITS data and are close to the traditional groupings traditionally recognized based on morphology. The intraspecific variation of rbcL is low and is neither correlated with intraspecific morphological variability nor with host range. Ancestral character reconstruction using parsimony suggests that the ancestor of O. sect. Orobanche had a narrow host range.

Base Sequence↗

Heterogeneity in regional GC content and differential usage of codons and amino acids in GC-poor and GC-rich regions of the genome of Apis mellifera.

The honeybee (Apis mellifera) has a genome with a wide variation in GC content showing 2 clear modal GC values, in some ways reminiscent of an isochore-like structure. To gain insight into causes and consequences of this pattern, we used a comparative approach to study the genome-wide alignment of primarily coding sequence of A. mellifera with Drosophila melanogaster and Anopheles gambiae. The latter 2 species show a higher average GC content than A. mellifera and no indications of bimodality, suggesting that the GC-poor mode is a derived condition in honeybee. In A. mellifera, synonymous sites of genes generally adopt the GC content of the region in which they reside. A large proportion of genes in GC-poor regions have not been assigned to the honeybee assembly because of the low sequence complexity of their genome neighborhood. The synonymous substitution rate between A. mellifera and the other species is very close to saturation, but analyses of nonsynonymous substitutions as well as amino acid substitutions indicate that the GC-poor regions are not evolving faster than the GC-rich regions. We describe the codon usage and amino acid usage and show that they are remarkably heterogeneous within the honeybee genome between the 2 different GC regions. Specifically, the genes located in GC-poor regions show a much larger deviation in both codon usage bias and amino acid usage from the Dipterans than the genes located in the GC-rich regions.

Amino Acids↗

CHOP: visualization of 'wobbling' and isolation of highly conserved regions from aligned DNA sequences.

The web software CHOP was developed to visualize the 'wobbling' in the third codon position of aligned DNA sequences. The simple features of this tool allow users to easily find regions suspected of containing coding sequences (CDSs). The program also allows visualization of the nucleotide diversity between two genomic or gene sequences by graphically plotting the percentage identity between the two sequences. CHOP can also isolate highly conserved regions within both CDSs and non-CDSs. Highly conserved regions within CDSs include the regions with lower rates of synonymous substitution in which nucleotide sequences are expected to be under strong selective pressure. CHOP is available at http://bunsei2.med.u-tokai.ac.jp:8080/~ohtsuka/cds_finding.html.

Animals↗

The S7 gene and VP7 protein are highly conserved among temporally and geographically distinct American isolates of epizootic hemorrhagic disease virus.

Complete sequences of genome segment 7 (S7) from six isolates of epizootic hemorrhagic disease virus serotype 1 (EHDV-1) and 37 isolates of serotype 2 (EHDV-2) were determined. These isolates were made between 1978 and 2001 from the southeast, mid-Atlantic, Midwest and intermountain United States. Analysis of the S7 sequence similarities showed 98.1% identity among the EHDV-1 isolates and 91.0% identity among the EHDV-2 isolates. Comparison of the deduced amino acid similarities showed an even greater degree of similarity among the isolates (100% among the EHDV-1 isolates and 98.9% identity among the EHDV-2 isolates). There was only 75.8% identity between the EHDV-1 and EHDV-2 isolates at the nucleic acid level; however, there was 93.7% identity between the two groups at the amino acid level. The ratio of non-synonymous to synonymous nucleotide indicates a strong selection for silent substitutions. There was no evidence for reassortment between EHDV-1 and EHDV-2 isolates. The high degree of conservation of S7 gene codons and the VP7 protein, suggests that little variation is allowed in preserving the function of this protein. The high degree of conservation also validates the use of diagnostic tests for EHDV based on S7 and VP7.

Capsid Proteins↗

Translation conditional models for protein coding sequences.

A coding sequence is defined as a DNA sequence coding the primary structure of a protein (a polypeptide). Such a sequence must satisfy a specific constraint, which consists in coding a functional protein. As the genetic code is degenerated, there exists, for a given polypeptide, a set of synonymous sequences which would code the same polypeptide. Translation conditional models are being defined on such sets. The aim of this paper is to give a common formalism. Besides the codon bias model, a few other conditional models will be defined. Statistical estimators and comparison methods will be briefly presented. These models can be used for gene classification, or to find out, in a real sequence, remarkable features. An example will be presented on Escherichia coli genes.

Bacterial Proteins↗