Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Synonymous and nonsynonymous substitutions in genes from Gramineae: intragenic correlations.

In this work, we have investigated the relationships between synonymous and nonsynonymous rates and base composition in coding sequences from Gramineae to analyze the factors underlying the variation in substitutional rates. We have shown that in these genes the rates of nucleotide divergence, both synonymous and nonsynonymous, are, to some extent, dependent on each other and on the base composition. In the first place, the variation in nonsynonymous rate is related to the GC level at the second codon position (the higher the GC(2) level, the higher the amino acid replacement rate). The correlation is especially strong with T(2), the coefficients being significant in the three data sets analyzed. This correlation between nonsynonymous rate and base composition at the second codon position is also detectable at the intragenic level, which implies that the factors that tend to increase the intergenic variance in nonsynonymous rates also affect the intragenic variance. On the other hand, we have shown that the synonymous rate is strongly correlated with the GC(3) level. This correlation is observed both across genes and at the intragenic level. Similarly, the nonsynonymous rate is also affected at the intragenic level by GC(3) level, like the silent rate. In fact, synonymous and nonsynonymous rates exhibit a parallel behavior in relation to GC(3) level, indicating that the intragenic patterns of both silent and amino acid divergence rates are influenced in a similar way by the intragenic variation of GC(3). This result, taken together with the fact that the number of genes displaying intragenic correlation coefficients between synonymous and nonsynonymous rates is not very high, but higher than random expectation (in the three data sets analyzed), strongly suggests that the processes of silent and amino acid replacement divergence are, at least in part, driven by common evolutionary forces in genes from Gramineae.

Base Composition↗

Substitution rates in hepatitis delta virus.

Substitution rates were estimated for the coding and noncoding regions of the hepatitis delta virus (HDV). The estimated rates of synonymous substitution in HDV were lower than the rates of substitution at non-synonymous sites and in the noncoding region. HDV has lower synonymous substitution rates than the hepatitis C virus, though both are RNA viruses. The relatively low rate of synonymous substitution in HDV may be due to a strong preference of G and C nucleotides at third codon positions. Variation in substitution rate among HDV lineages may be correlated with the clinical development of the HDV-induced hepatitis. The phylogenetic tree inferred for 24 HDV strains reveals similarities between lineages isolated from the same geographic region.

Genome, Viral↗

Allelic variants of the human MHC class I chain-related B gene (MICB).

The human major histocompatibility complex (MHC) is located within a 4 megabase segment on chromosome 6p21.3. Recently, a highly divergent MHC class I chain-related gene family, MIC was identified within the class I region. The MICA and MICB genes in this family have unique patterns of tissue expression. The MICA gene is highly polymorphic, with more than 20 alleles identified to date. To elucidate the extent of MICB allelic variations, we sequenced exons 2 (alpha 1), 3 (alpha 2), 4 (alpha 3), and 5 (transmembrane) as well as introns 2 and 4 of this gene in 46 HLA homozygous B-cell lines. We report the identification of eleven alleles based on seven non-synonymous, two synonymous, and four intronic nucleotide variations. Interestingly, one allele has a nonsense mutation resulting in a premature termination codon in the alpha 2 domain. Thus, MICB appears to have fewer alleles than MICA, not unlike the allelic ratio between the HLA-C and -B loci. A preliminary linkage analysis of the MICB alleles with those of the closely located MICA and HLA-B genes revealed no conspicuous linkage disequilibrium between them, implying the presence of a potential recombination hotspot between the MICB and MICA genes.

Alleles↗

Overestimated frequency of a possible emphysema-susceptibility allele when microsomal epoxide hydrolase is genotyped by the conventional polymerase chain reaction-based method.

A recent association study suggested that the His113 variant of microsomal epoxide hydrolase (mEPHX) may confer a risk for development of emphysema, presumably by increasing susceptibility to smoking injury. Before considering a possible role of this enzyme in pulmonary disease, we attempted to characterize the genetic polymorphism further. The Tyr/His113 polymorphism within exon 3 of mEPHX was initially examined in 62 healthy individuals by conventional methods involving polymerase chain reaction (PCR)-based determination of a restriction fragment length polymorphism (RFLP). Genomic nucleotide sequences, including the polymorphic site and the downstream primer sequence, were further analyzed in 95 unrelated, healthy Japanese volunteers by single-stranded conformation polymorphism (SSCP) analysis and direct sequencing. Genotyping by the first method (PCR-RFLP) revealed that the allelic distribution in our test population apparently deviated from Hardy-Weinberg equilibrium. Sequence analysis showed that a synonymous nucleotide substitution, AAG to AAA (Lys119), was located just within the published primer site. The AAA at codon 119 was present only in alleles with Tyr113, and its frequency reached 0.31 in our panel of 190 Japanese alleles. This substitution potentially hampered PCR amplification because of the nucleotide mismatch, with the result that the frequency of the Tyr113 variation was underestimated. The frequency of His113, a possible emphysema susceptibility allele of the mEPHX gene, was thus overestimated when human DNA samples were genotyped in the conventional way. Depending on the population(s) tested, this anomaly could represent a pitfall for PCR-based association studies.

Alleles↗

Tax & rex: overlapping genes of the Deltaretrovirus group.

Bovine leukemia virus and human T-cell leukemia viruses I and II, members of the Deltaretrovirus group, have two regulatory genes, tax and rex, that are coded in overlapping reading frames. We found that sequence variations in the rex gene of each virus result in amino acid differences significantly more often than variations in the tax gene. For all three viruses the highest ratio of non-synonymous to synonymous changes was found in the rex gene. In the overlapping regions of tax and rex, the second codon position of Rex corresponds to the third codon position of Tax. Nucleotide C was present in all genes of the three viruses at the highest frequency and this bias was most pronounced in the rex gene. More specifically we found that the C bias and nucleotide variation is greatest at the second codon position of Rex and the third codon position of Tax in the area of tax/rex overlap. Changes in the second codon position of Rex always resulted in amino acid change whereas changes in the third codon position of Tax resulted in amino acid changes less than a third of the time. Analysis of the amino acid frequencies in both proteins shows that there is a disproportionately large percentage of the amino acids alanine, proline, serine and threonine (the four amino acids whose second codon position is C) in Rex. These findings led us to hypothesize that the Rex protein can withstand more amino acid changes than can the Tax protein suggesting that the Tax protein experiences higher evolutionary constraints and is the more conserved of the two proteins.

Amino Acid Sequence↗

Dramatically elevated rate of mitochondrial substitution in lice (Insecta: Phthiraptera).

Few estimates of relative substitution rates, and the underlying mutation rates, exist between mitochondrial and nuclear genes in insects. Previous estimates for insects indicate a 2-9 times faster substitution rate in mitochondrial genes relative to nuclear genes. Here we use novel methods for estimating relative rates of substitution, which incorporate multiple substitutions, and apply these methods to a group of insects (lice, Order: Phthiraptera). First, we use a modification of copath analysis (branch length regression) to construct independent comparisons of rates, consisting of each branch in a phylogenetic tree. The branch length comparisons use maximum likelihood models to correct for multiple substitution. In addition, we estimate codon-specific rates under maximum likelihood for the different genes and compare these values. Estimates of the relative synonymous substitution rates between a mitochondrial (COI) and nuclear (EF-1alpha) gene in lice indicate a relative rate of several 100 to 1. This rapid relative mitochondrial rate (>100 times) is at least an order of magnitude faster than previous estimates for any group of organisms. Comparisons using the same methods for another group of insects (aphids) reveals that this extreme relative rate estimate is not simply attributable to the methods we used, because estimates from aphids are substantially lower. Taxon sampling affects the relative rate estimate, with comparisons involving more closely related taxa resulting in a higher estimate. Relative rate estimates also increase with model complexity, indicating that methods accounting for more multiple substitution estimate higher relative rates.

Animals↗

Effects of gene expression on molecular evolution in Arabidopsis thaliana and Arabidopsis lyrata.

We analyzed the complete genome sequence of Arabidopsis thaliana and sequence data from 83 genes in the outcrossing A. lyrata, to better understand the role of gene expression on the strength of natural selection on synonymous and replacement sites in Arabidopsis. From data on tRNA gene abundance, we find a good concordance between codon preferences and the relative abundance of isoaccepting tRNAs in the complete A. thaliana genome, consistent with models of translational selection. Both EST-based and new quantitative measures of gene expression (MPSS) suggest that codon preferences derived from information on tRNA abundance are more strongly associated with gene expression than those obtained from multivariate analysis, which provides further support for the hypothesis that codon bias in Arabidopsis is under selection mediated by tRNA abundance. Consistent with previous results, analysis of protein evolution reveals a significant correlation between gene expression level and amino acid substitution rate. Analysis by MPSS estimates of gene expression suggests that this effect is primarily the result of a correlation between the number of tissues in which a gene is expressed and the rate of amino acid substitution, which indicates that the degree of tissue specialization may be an important determinant of the rate of protein evolution in Arabidopsis.

Arabidopsis↗

LISTA, LISTA-HOP and LISTA-HON: a comprehensive compilation of protein encoding sequences and its associated homology databases from the yeast Saccharomyces.

We continued our effort to make a comprehensive database (LISTA) for the yeast Saccharomyces cerevisiae. In this database each sequence has been attributed a single genetic name. In the case of duplicated sequences a simple method has been applied to distinguish between sequences of one and the same gene from non-allelic sequences of duplicated genes. If necessary, synonyms are given in the case of allelic duplicated sequences. Thus sequences can be found either by the name or by synonyms given in LISTA. Each entry contains the genetic name, the mnemonic from the EMBL data bank, the codon bias, reference of the publication of the sequence, Chromosomal location as far as known, Swissprot and EMBL accession numbers. To obtain more information on the included sequences, each entry has been screened against non-redundant nucleotide and protein data bank collections resulting in LISTA-HON and LISTA-HOP. The LISTA data base can be linked to the associated data sets or to nucleotide and protein banks by the Sequence Retrieval System (SRS).

Base Sequence↗

LISTA, LISTA-HOP and LISTA-HON: a comprehensive compilation of protein encoding sequences and its associated homology databases from the yeast Saccharomyces.

We continued our effort to make a comprehensive database (LISTA) for the yeast Saccharomyces cerevisiae. As in previous editions the genetic names are consistently associated to each sequence with a known and confirmed ORF. If necessary, synonyms are given in the case of allelic duplicated sequences. Although the first publication of a sequence gives-according to our rules-the genetic name of a gene, in some instances more commonly used names are given to avoid nomenclature problems and the use of ancient designations which are no longer used. In these cases the old designation is given as synonym. Thus sequences can be found either by the name or by synonyms given in LISTA. Each entry contains the genetic name, the mnemonic from the EMBL data bank, the codon bias, reference of the publication of the sequence, Chromosomal location as far as known, SWISSPROT and EMBL accession numbers. New entries will also contain the name from the systematic sequencing efforts. Since the release of LISTA4.1 we update the database continuously. To obtain more information on the included sequences, each entry has been screened against non-redundant nucleotide and protein data bank collections resulting in LISTA-HON and LISTA-HOP. This release includes reports from full Smith and Watermann peptide-level searches against a non-redundant protein sequence database. The LISTA data base can be linked to the associated data sets or to nucleotide and protein banks by the Sequence Retrieval System (SRS). The database is available by FTP and on World Wide Web.

Amino Acid Sequence↗

Identification of new endogenous retroviral sequences belonging to the HERV-W family in human cancer cells.

A human endogenous retroviral family (HERV-W) has recently been identified on chromosome 7 that contains a single complete open reading frame putatively encoding an envelope protein. We have identified fifteen HERV-W families on human genomic DNA using a monochromosomal panel in our previous study. In order to identify additional families, we examined genomic DNA derived from cancer cell lines (A549, AZ521, OVCAR3, RT4) using the PCR approach. Five env fragments of a HERV-W family were newly identified and analyzed. They showed a high degree of nucleotide sequence similarity (93-97%) to that of the other HERV-W families. Translation of the env fragments showed no frameshift and termination codon by deletion/insertion or point mutation in clones A549-2, AZ521-3, and OVCAR-3. The ratio of synonymous to nonsynonymous substitutions indicated that negative selective pressure is acting on A549-2, AZ521-3, and OVCAR-3 sequences. These env gene sequences could be associated with an active provirus in human cancer cells (A549, AZ521, OVCAR3).

Amino Acid Sequence↗

Selection profiles in RNA viruses reflect the characteristics of viruses more than individual proteins.

Proteins that are exposed on the surface of a virus are frequently subject to strong selection to escape from neutralizing antibodies. To investigate whether surface-exposed (SE) and non-exposed (NE) proteins encoded by RNA viruses exhibit different patterns of evolution under selection, we analyzed 244 protein-coding genes from 28 species of RNA viruses representing 15 taxonomic families. First, we show that gene-wide rates of non-synonymous (dN) and synonymous (dS) substitutions do not differentiate between SE and NE proteins. To incorporate variation in substitution rates among codon sites, we inferred the posterior distribution over a fixed grid of dN and dS rates for each alignment. This 'evolutionary fingerprint' provides a common framework for comparing the selection profiles of non-homologous genes. Next, we computed the Wasserstein distance for every pair of fingerprints, which is analogous to amount of work required to reshape one distribution to another. After compensating for differences in genetic variation among alignments, we found a small but significant difference between the fingerprints of SE and NE proteins (PERMANOVA, P&#x2009;=&#x2009;0.03). However, we observed larger and more significant effects of whether the virus is enveloped (P&#x2009;<&#x2009;10-5) and the interaction between these factors (P=6.9&#xd7;10-4). The latter effects were driven by high levels of purifying selection in capsid proteins of Picornaviruses. Furthermore, greater amounts of variation in fingerprints were explained by significant differences among virus families and modes of transmission (P&#x2009;<&#x2009;10-5). These results imply the pattern of selection on a virus protein is shaped more by characteristics of the virus than the protein itself.

RNA Viruses↗

Three novel single nucleotide polymorphisms in UGT1A9.

Three novel single nucleotide polymorphisms (SNPs) were found in the UDP-glucuronosyltransferase (UGT) 1A9 gene from 97 Japanese subjects (47 cancer patients and 50 cardiovascular disease patients). The detected SNPs were as follows: 1) SNP, MPJ6_U1A006; GENE NAME, UGT1A9; ACCESSION NUMBER, AF297093; LENGTH, 25 bases; 5'-AATTCTCTTAGGG/TTTCTCAGATGCC-3'. 2) SNP, MPJ6_U1A007; GENE NAME, UGT1A9; ACCESSION NUMBER, AF297093; LENGTH, 25 bases; 5'-TGTTACGGAGTAT/GGATCTCTACAGC-3'. 3) SNP, MPJ6_U1A031; GENE NAME, UGT1A9; ACCESSION NUMBER, AF297093; LENGTH, 25 bases; 5'-ACTCATTCTCAGG/AGGGCATGAGGTG-3'. All three SNPs were located in exon 1 with frequencies of 0.036 for MPJ6_U1A006, and 0.005 for MPJ6_U1A007 and MPJ6_U1A031. SNP MPJ6_U1A007 (726T>G) results in formation of a termination codon TAG (Y242X). The other two SNPs, MPJ6_U1A006 (588G>T) and MPJ6_U1A031 (153G>A), result in synonymous changes (G196G and R51R, respectively).

Journal Article↗

Natural selection and the frequency distributions of "silent" DNA polymorphism in Drosophila.

In Escherichia coli, Saccharomyces cerevisiae, and Drosophila melanogaster, codon bias may be maintained by a balance among mutation pressure, genetic drift, and natural selection favoring translationally superior codons. Under such an evolutionary model, silent mutations fall into two fitness categories: preferred mutations that increase codon bias and unpreferred changes in the opposite direction. This prediction can be tested by comparing the frequency spectra of synonymous changes segregating within populations; natural selection will elevate the frequencies of advantageous mutations relative to that of deleterious changes. The frequency distributions of preferred and unpreferred mutations differ in the predicted direction among 99 alleles of two D. pseudoobscura genes and five alleles of eight D. simulans genes. This result confirms the existence of fitness classes of silent mutations. Maximum likelihood estimates suggest that selection intensity at silent sites is, on average, very weak in both D. pseudoobscura and D. simulans (magnitude of NS approximately 1). Inference of evolutionary processes from within-species sequence variation is often hindered by the assumption of a stationary frequency distribution. This assumption can be avoided when identifying the action of selection and tested when estimating selection intensity.

Animals↗

Are residues in a protein folding nucleus evolutionarily conserved?

Protein is the working molecule of the cell, and evolution is the hallmark of life. It is important to understand how protein folding and evolution influence each other. Several studies correlating experimental measurement of residue participation in folding nucleus and sequence conservation have reached different conclusions. These studies are based on assessment of sequence conservation at folding nucleus sites using entropy or relative entropy measurement derived from multiple sequence alignment. Here we report analysis of conservation of folding nucleus using an evolutionary model alternative to entropy-based approaches. We employ a continuous time Markov model of codon substitution to distinguish mutation fixed by evolution and mutation fixed by chance. This model takes into account bias in codon frequency, bias-favoring transition over transversion, as well as explicit phylogenetic information. We measure selection pressure using the ratio omega of synonymous versus non-synonymous substitution at individual residue site. The omega-values are estimated using the PAML method, a maximum-likelihood estimator. Our results show that there is little correlation between the extent of kinetic participation in protein folding nucleus as measured by experimental phi-value and selection pressure as measured by omega-value. In addition, two randomization tests failed to show that folding nucleus residues are significantly more conserved than the whole protein, or the median omega value of all residues in the protein. These results suggest that at the level of codon substitution, there is no indication that folding nucleus residues are significantly more conserved than other residues. We further reconstruct candidate ancestral residues of the folding nucleus and suggest possible test tube mutation studies for testing folding behavior of ancient folding nucleus.

Amino Acid Sequence↗

Codon usage bias and base composition of nuclear genes in Drosophila.

The nuclear genes of Drosophila evolve at various rates. This variation seems to correlate with codon-usage bias. In order to elucidate the determining factors of the various evolutionary rates and codon-usage bias in the Drosophila nuclear genome, we compared patterns of codon-usage bias with base compositions of exons and introns. Our results clearly show the existence of selective constraints at the translational level for synonymous (silent) sites and, on the other hand, the neutrality or near neutrality of long stretches of nucleotide sequence within noncoding regions. These features were found for comparisons among nuclear genes in a particular species (Drosophila melanogaster, Drosophila pseudoobscura and Drosophila virilis) as well as in a particular gene (alcohol dehydrogenase) among different species in the genus Drosophila. The patterns of evolution of synonymous sites in Drosophila are more similar to those in the prokaryotes than they are to those in mammals. If a difference in the level of expression of each gene is a main reason for the difference in the degree of selective constraint, the evolution of synonymous sites of Drosophila genes would be sensitive to the level of expression among genes and would change as the level of expression becomes altered in different species. Our analysis verifies these predictions and also identifies additional selective constraints at the translational level in Drosophila.

Animals↗

Modifying the sequence of an immunoglobulin V-gene alters the resulting pattern of hypermutation.

Affinity maturation of antibodies requires localized hypermutation and antigen selection. Hypermutation is particularly active in certain regions (notably the CDRs of light and heavy chains) due to the local accumulation of hot spots. We have now analyzed the role of individual nucleotides in the origin of hot spots and show that mutability is largely defined by the nucleotide sequence. We compared the mutability profile of wild-type and modified kappa transgenes that contain silent mutations in the CDR1 segment. We found a new hot spot created at the third base of Ser-31 when its wild-type AGT codon was substituted by AGC. Two major hot spots associated with this AGC vanished when Ser-31 was encoded by the synonymous TCA. In addition to these, which were the most prominent changes, there were compensatory alterations in mutability of residues not directly related to the introduced silent mutations, so that the average hypermutation remained constant. Thus, mutations arising early in the immune response, even silent ones, could affect the mutability of critical residues and alter the pattern of affinity maturation. When analyzing hybridomas, we detected such alterations, but they seemed to better correlate with changes in average rather than local mutation rates. Overall, this paper shows how evolution could have optimized the mutability of individual residues to minimize deleterious mutations. Thus, the optimal strategy for affinity maturation may involve the incorporation of multiple point mutations before antigen selection of the relevant cells.

Animals↗

TAPI polymorphisms in several human ethnic groups: characteristics, evolution, and genotyping strategies.

Genetic variations in the locus encoding the transporter associated with antigen processing, subunit 1 (TAP1), were systematically studied using samples from Caucasians, Africans, Brazilians, and compared with data from chimpanzees. PCR-amplified genomic sequences corresponding to the 11 exons were analyzed by single-strand conformation polymorphism (SSCP) and sequencing. Six nonsynonymous and 2 synonymous single nucleotide polymorphisms (SNPs) were found to be common in one ethnic group or another, and they involved codons 254 (Gly-GGC/Gly-GGT) in exon 3, 333 (Ile-ATC/Val-GTC) in exon 4, 370 (Ala-GCT/Val-GTT) in exon 5, 458 (Val-GTG/Leu-TTG) in exon 6, 518 (Val-GTC/Ile-ATC) in exon 7, 637 (Asp-GAC/Gly-GGC), 648 (Arg-CGA/Gln-CAA) and 661 (Pro-CCG/Pro-CCA) in exon 10. At each SNP site the sequence listed first was predominant in all ethnic groups. Several SNPs segregated on the same chromosome regardless of populations and species. Together, the SNPs produced 5 major human TAP1 alleles, 4 of which matched the officially recognized alleles *0101, *02011, *0301, and *0401; the 5th allele differed from each of those by at least 4 SNPs. Overall, TAP1*0101 was the predominant allele in all ethnic groups, with frequencies ranging from 0.667 in Zambians to 0.808 in US Caucasians. The TAP1*0401 frequency showed the greatest difference between Africans (0.221-0.254) and Caucasians (0.033), with Brazilians (0.058) fitting in the middle. Consistent with earlier work based on Caucasians and gorillas, *0101 appeared to be the newest human TAP1 allele, suggesting a dramatic spread of *0101 into all human populations examined. Characterization of TAP1 polymorphisms allowed the design of a PCR-based genotyping scheme that targeted 7 SNP sites and required 2 separate genotyping techniques.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

Adaptive evolution of the insulin gene in caviomorph rodents.

Insulin is a conservative molecule among mammals, maintaining both its structure and function. Rodents that belong to the Suborder Hystricognathi represent an exception, having a very divergent molecule with unusual physiological properties. In this work, we analyzed the evolutionary pattern of the insulin gene in caviomorph rodents (South American hystricomorph rodents). We found that these rodents have higher rates of nonsynonymous:synonymous substitutions (d(N)/d(S)) than nonhystricomorph rodents and that values are heterogeneous inside the group. We estimated codons under positive selection, specifically the second binding site (A13 and B17) and others related with hexamerization (B18, B20, and B22). In the monomer structure, all selected sites formed a single patch around the second binding site. In the hexamer structure, these amino acids were grouped into three major patches. In this structure, contacts between B chains involved all selected sites (except B18), and between faces in the center of the molecule, all contacts were among selected sites. While there is no clear hypothesis regarding the cause of this drastic change, experimental evidence does show that this group of rodents has some peculiarities in growth function, and, whether coincidental or not, these changes appeared together with important changes in life-history traits.

Adaptation, Physiological↗