Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Large number of polymorphic nucleotides and a termination codon in the env gene of the endogenous human retrovirus ERV3.

The terminal portion of the pol gene and the entire env gene of the human endogenous retrovirus ERV3 was screened for polymorphic nucleotides. For this purpose fragments amplified from the desired regions of ERV3 were subjected to single strand conformational analysis (SSCP analysis). Using this approach, we detected 13 polymorphic nucleotides, namely four in the pol gene and nine in the env gene. Three of the nucleotide substitutions were synonymous (not affecting the amino acid code). One of the non-synonymous nucleotide substitutions changed an arginine codon to a termination codon. The alleles at the different polymorphic sites could be arranged into five ERV3 haplotypes, two of which were new. To evaluate the possible significance of the termination codon, which precludes expression of a putative immunoregulatory factor, we examined samples of DNA from patients with multiple sclerosis, a demyelinating disease of presumed autoimmune etiology. We did not find an association between the ERV3 allele with the termination codon and this disease. Perhaps the presence of a stop codon combined with the high number of non-synonymous nucleotide substitutions in the reading frame of the env gene reflects absence of selective constraints during evolution. Obviously, our findings contradict the assumption that the reading frame of the ERV3 env gene has been conserved throughout evolution.

Codon, Terminator↗

Positive selection at sites of multiple amino acid replacements since rat-mouse divergence.

New alleles become fixed owing to random drift of nearly neutral mutations or to positive selection of substantially advantageous mutations. After decades of debate, the fraction of fixations driven by selection remains uncertain. Within 9,390 genes, we analysed 28,196 codons at which rat and mouse differ from each other at two nucleotide sites and 1,982 codons with three differences. At codons where rat-mouse divergence involved two non-synonymous substitutions, both of them occurred in the same lineage, either rat or mouse, in 64% of cases; however, independent substitutions would occur in the same lineage with a probability of only 50%. All three non-synonymous substitutions occurred in the same lineage for 46% of codons, instead of the 25% expected. Furthermore, comparison of 12 pairs of prokaryotic genomes also shows clumping of multiple non-synonymous substitutions in the same lineage. This pattern cannot be explained by correlated mutation or episodes of relaxed negative selection, but instead indicates that positive selection acts at many sites of rapid, successive amino acid replacement.

Alleles↗

Novel alleles of the chemokine-receptor gene CCR5.

The CCR5 gene encodes a cell-surface chemokine-receptor molecule that serves as a coreceptor for macrophage-tropic strains of HIV-1. Mutations in this gene may alter expression or function of the protein product, thereby altering chemokine binding/signaling or HIV-1 infection of cells that normally express CCR5 protein. Indeed, homozygotes for a 32-bp deletion allele of CCR5 (CCR5-delta 32), which causes a frameshift at amino acid 185, are relatively resistant to HIV-1 infection. Here we report the identification of 16 additional mutations in the coding region of the CCR5 gene, all but 3 of which are codon altering or "nonsynonymous." Most mutations were rare (found only once or twice in the sample); five were detected exclusively among African Americans, whereas eight were observed only in Caucasians. The mutations included 11 codon-altering nonsynonymous variants, one trinucleotide deletion, one chain-termination mutant, and three synonymous mutations. The high predominance of codon-altering alleles among CCR5 mutants (14/17 [81%], including CCR5-delta 32) is consistent with an adaptive accumulation of function-altering alleles for this gene, perhaps as a consequence of historic selective pressures.

Alleles↗

Evidence for positive selection on a sexual reproduction gene in the diatom genus Thalassiosira (Bacillariophyta).

Single likelihood ancestor counting (SLAC), fixed effects likelihood (FEL), and several random effects likelihood (REL) methods were utilized to identify positively and negatively selected sites in sexually induced gene 1 (Sig1) of four different Thalassiosira species. The SLAC analysis did not find any sites affected by positive selection but suggested 13 sites influenced by negative selection. The SLAC approach may be too conservative because of low sequence divergence. The FEL and REL analyses revealed over 60 negatively selected sites and two positively selected sites that were unique to each method. The REL method may not be able to reliably identify individual sites under selection when applied to short sequences with low divergence. Instead, we proposed a new alignment-wide test for adaptive evolution based on codon models with variation in synonymous and nonsynonymous substitution rates among sites and found evidence for diversifying evolution without relying on site-by-site testing. The performance of the FEL and REL approaches was evaluated by subjecting the tests to a type I error rate simulation analysis, using the specific characteristics of the Sig1 data set. Simulation results indicated that the FEL test had reasonable Type I errors, while REL might have been too liberal, suggesting that the two positively selected sites identified by FEL (codons 94 and 174) are not likely to be false positives. The evolution of these codon sites, one of which is located in functional domain II, appears to be associated with divergence among the three major Thalassiosira lineages.

Animals↗

Natural selection and algorithmic design of mRNA.

Messenger RNA (mRNA) sequences serve as templates for proteins according to the triplet code, in which each of the 4(3) = 64 different codons (sequences of three consecutive nucleotide bases) in RNA either terminate transcription or map to one of the 20 different amino acids (or residues) which build up proteins. Because there are more codons than residues, there is inherent redundancy in the coding. Certain residues (e.g., tryptophan) have only a single corresponding codon, while other residues (e.g., arginine) have as many as six corresponding codons. This freedom implies that the number of possible RNA sequences coding for a given protein grows exponentially in the length of the protein. Thus nature has wide latitude to select among mRNA sequences which are informationally equivalent, but structurally and energetically divergent. In this paper, we explore how nature takes advantage of this freedom and how to algorithmically design structures more energetically favorable than have been built through natural selection. In particular: (1) Natural Selection--we perform the first large-scale computational experiment comparing the stability of mRNA sequences from a variety of organisms to random synonymous sequences which respect the codon preferences of the organism. This experiment was conducted on over 27,000 sequences from 34 microbial species with 36 genomic structures. We provide evidence that in all genomic structures highly stable sequences are disproportionately abundant, and in 19 of 36 cases highly unstable sequences are disproportionately abundant. This suggests that the stability of mRNA sequences is subject to natural selection. (2) Artificial Selection--motivated by these biological results, we examine the algorithmic problem of designing the most stable and unstable mRNA sequences which code for a target protein. We give a polynomial-time dynamic programming solution to the most stable sequence problem (MSSP), which is asymptotically no more complex than secondary structure prediction. We show that the corresponding least stable sequence problem (LSSP) is NP-complete, and develop two heuristics for the construction of such sequences. We have implemented these algorithms, and present experimental results placing the high/low stability sequences in context with both wildtype and random encodings. Our implementation has already been applied to the design of RNA "code-words" creating little or no secondary structure in RNA computing (Brenneman and Condon, 2001; Marathe et al., 2001), and we anticipate a variety of other applications of this work to sequence design problems (Skiena, 2001).

Algorithms↗

New methods for detecting positive selection at single amino acid sites.

Inferring positive selection at single amino acid sites is of particular importance for studying evolutionary mechanisms of a protein. For this purpose, Suzuki and Gojobori (1999) developed a method (SG method) for comparing the rates of synonymous and nonsynonymous substitutions at each codon site in a protein-coding nucleotide sequence, using ancestral codons at interior nodes of the phylogenetic tree as inferred by the maximum parsimony method. In the SG method, however, selective neutrality of nucleotide substitutions cannot be tested at codon sites, where only termination codons are inferred at any interior node or the number of equally parsimonious inferences of ancestral codons at all interior nodes exceeds 10,000. Here I present a modified SG method which is free from these problems. Specifically, I use the distance-based Bayesian method for inferring the single most likely ancestral codon from 61 sense codons at each interior node. In the computer simulation and real data analysis, the modified SG method showed a higher overall efficiency of detecting positive selection than the original SG method, particularly at highly polymorphic codon sites. These results indicate that the modified SG method is useful for inferring positive selection at codon sites where neutrality cannot be tested by the original SG method. I also discuss that the p-distance is preferable to the number of synonymous substitutions for inferring the phylogenetic tree in the SG method, and present a maximum likelihood method for detecting positive selection at single amino acid sites, which produced reasonable results in the real data analysis.

Amino Acids↗

A new algorithm for analysis of within-host HIV-1 evolution.

A new algorithm for inferring the evolution of within-host viral sequences is presented. A sequential-linking approach is developed so that a longitudinal phylogenetic tree can be reconstructed from sequential molecular data that are obtained at different time points from the same host. The algorithm employs a codon-based model, which uses a Markov process to describe substitutions between codons, to calculate nonsynonymous and synonymous substitution rates and to distinguish positive selection and neutral evolution. The algorithm is applied to a data set of the V3 region of the HIV-1 envelope genes sequenced at different years after the infection of a single patient. The results suggest that this algorithm may provide a more realistic description of viral evolution than traditional evolutionary models, because it accounts for both neutral and adaptive evolution, and reconstructs a longitudinal phylogenetic tree that describes the dynamic process of viral evolution.

Algorithms↗

The molecular clock revisited: the rate of synonymous vs. replacement change in Drosophila.

Rates of synonymous and nonsynonymous substitution were investigated for 24 genes in three Drosophila species, D. pseudoobscura, D. subobscura, and D. melanogaster. D. pseudoobscura and D. subobscura, two distantly related members of the obscura clade, differ on average by 0.29 synonymous nucleotide substitutions per site. D. melanogaster differs from the two obscura species by an average of 0.81 synonymous substitutions per site. Using a method developed by Gillespie, we investigated the variance to mean ratio, or Index of Dispersion, R, of substitutions along the three species' branches to test the fundamental prediction of the neutral theory of molecular evolution, E(R) = 1. For nonsynonymous substitutions, the average R, Ra is 1.6, which is not significantly different from the neutral theory prediction. Only 5 of the 24 genes had significantly large Ra valves, and 12 of the genes had Ra estimates of less than one. In contrast, the Index of Dispersion for synonymous substitutions was significantly large for 12 of the 24 genes, with an average of R(s) = 4.4, also statistically significant. These findings contrast with results for mammals, which showed overdispersion of nonsynonymous substitutions, but not of synonymous substitutions. Weak selection acting to maintain codon bias in Drosophila, but not in mammals, may be important in explaining the high variance in the rate of synonymous substitutions in this group of organisms.

Animals↗

Synonymous and nonsynonymous substitution rates in diatoms: a comparison between chloroplast and nuclear genes.

Rates of synonymous and nonsynonymous nucleotide substitutions and codon usage bias (ENC) were estimated for a number of nuclear and chloroplast genes in a sample of centric and pennate diatoms. The results suggest that DNA evolution has taken place, on an average, at a slower rate in the chloroplast genes than in the nuclear genes: a rate variation pattern similar to that observed in land plants. Synonymous substitution rates in the chloroplast genes show a negative association with the degree of codon usage bias, suggesting that genes with a higher degree of codon usage bias have evolved at a slower rate. While this relationship has been shown in both prokaryotes and multicellular eukaryotes, it has not been demonstrated before in diatoms.

Cell Nucleus↗

Interspecific sequence comparison of the muscle-myosin heavy-chain genes from Drosophila hydei and Drosophila melanogaster.

The muscle-myosin heavy-chain (mMHC) gene of Drosophila hydei has been sequenced completely (size 23.3 kb). The sequence comparison with the D. melanogaster mMHC gene revealed that the exon-intron pattern is identical. The protein coding regions show a high degree of conservation (97%). The alternatively spliced exons (3a-b, 7a-d, 9a-c, 11a-e, and 15a-b) display more variations in the number of nonsynonymous and synonymous substitutions than the common exons (2, 4, 5, 6, 8, 10, 12, 13, 14, 16, 17, and 19). The base composition at synonymous sites of fourfold degenerate codons (third position) is not biased in the alternative exons. In the common exons there exists a bias for C and against A. These findings imply that the alternative exons of the Drosophila mMHC gene evolve at a different, in several cases higher, rate than the common ones. The 5' splice junctions and 5' and 3' untranslated regions show a high level of similarity, indicating a functional constraint on these sequences. The intron regions vary considerably in length within one species, but the corresponding introns are very similar in length between the two species and all contain stretches of sequence similarity. A particular example is the first intron, which contains multiple regions of similarity. In the conserved regions of intron 12 (head-tail border) sequences were found which have the potential to direct another smaller mMHC transcript.

Amino Acid Sequence↗

Origin and evolutionary pathways of the H1 hemagglutinin gene of avian, swine and human influenza viruses: cocirculation of two distinct lineages of swine virus.

The nucleotide sequences of the HA1 domain of the H1 hemagglutinin genes of A/duck/Hong Kong/36/76, A/duck/Hong Kong/196/77, A/sw/North Ireland/38, A/sw/Cambridge/39 and A/Yamagata/120/86 viruses were determined, and their evolutionary relationships were compared with those of previously sequenced hemagglutinin (H1) genes from avian, swine and human influenza viruses. A pairwise comparison of the nucleotide sequences revealed that the genes can be segregated into three groups, the avian, swine and human virus groups. With the exception of two swine strains isolated in the 1930s, a high degree of nucleotide sequence homology exists within the group. Two phylogenetic trees constructed from the substitutions at the synonymous site and the third codon position showed that the H1 hemagglutinin genes can be divided into three host-specific lineages. Examination of 21 hemagglutinin genes from the human and swine viruses revealed that two distinct lineages are present in the swine population. The swine strains, sw/North Ireland/38 and sw/Cambridge/39, are clearly on the human lineage, suggesting that they originate from a human A/WSN/33-like variant. However, the classic swine strain, sw/Iowa/15/30, and the contemporary human viruses are not direct descendants of the 1918 human pandemic strain, but did diverge from a common ancestral virus around 1905. Furthermore, previous to this the above mammalian viruses diverged from the lineage containing the avian viruses at about 1880.

Amino Acid Sequence↗

A simple method for estimating evolutionary rates of base substitutions through comparative studies of nucleotide sequences.

Some simple formulae were obtained which enable us to estimate evolutionary distances in terms of the number of nucleotide substitutions (and, also, the evolutionary rates when the divergence times are known). In comparing a pair of nucleotide sequences, we distinguish two types of differences; if homologous sites are occupied by different nucleotide bases but both are purines or both pyrimidines, the difference is called type I (or "transition" type), while, if one of the two is a purine and the other is a pyrimidine, the difference is called type II (or "transversion" type). Letting P and Q be respectively the fractions of nucleotide sites showing type I and type II differences between two sequences compared, then the evolutionary distance per site is K = -(1/2) ln [(1-2P-Q) square root of 1-2Q]. The evolutionary rate per year is then given by k = K/(2T), where T is the time since the divergence of the two sequences. If only the third codon positions are compared, the synonymous component of the evolutionary base substitutions per site is estimated by K'S = -(1/2) ln (1-2P-Q). Also, formulae for standard errors were obtained. Some examples were worked out using reported globin sequences to show that synonymous substitutions occur at much higher rates than amino acid-altering substitutions in evolution.

Animals↗

No receptor-binding domain adaptation detected in within-host H5N1 surveillance of 4,559 US dairy outbreak sequences.

BACKGROUND: The 2024-2026 US H5N1 clade 2.3.4.4b dairy cattle outbreak has been characterised primarily through consensus-level phylogenetics. Whether mammalian-adaptation variants are emerging at sub-consensus frequencies within infected hosts, particularly at the haemagglutinin receptor-binding domain (RBD), remains unknown because no systematic within-host variant analysis of the public sequencing corpus has been performed. METHODS: We conducted a pre-registered, corpus-wide intrahost single-nucleotide variant (iSNV) analysis of all publicly available H5N1 cattle, feline-spillover, and retail-milk sequences on the NCBI Sequence Read Archive (4559 samples across 7 BioProjects). A dual-caller concordance pipeline (iVar + LoFreq) with empirically determined allele frequency (AF) threshold (3%, set via four-criterion validation including synthetic spike-in controls) was applied to an 11-site Tier 1 mammalian-adaptation panel spanning the polymerase complex, haemagglutinin RBD, and accessory proteins. Within-host nucleotide diversity was compared across host categories. RESULTS: The HA RBD sites Q226L and G228S (H3 numbering) showed zero detections across >4300 adequately sequenced samples at all AF thresholds tested (1-5%), despite the pipeline detecting other non-synonymous variants at these exact codon positions (upper 95% CI for prevalence: 0.08%). Seven of eleven adaptation sites carried statistically significant iSNV signals after Bonferroni correction (corrected α = 0.00417), though all at low prevalence (≤2.95%). Genotype stratification showed that most polymerase-site detections reflected genotype structure rather than within-host emergence: the apparent PB2 631 L→M "reversion" was largely the ancestral avian state of the D1.1 genotype (20 of 23 detections), which never acquired the 631L mammalian adaptation, with only two genuine sub-consensus events in the B3.13 background, while consensus-level PB2 701N was a fixed feature of the D1.1 genotype (10 of 14 detections) rather than independent sub-consensus emergence. Cattle exhibited significantly higher within-host nucleotide diversity than feline-spillover samples (π = 1.59 × 10-4 vs 6.11 × 10-5; Kruskal-Wallis p = 6.6 × 10-15), a finding that persisted after depth-matching (p = 4.6 × 10-5); this may reflect prolonged mammary-gland infection, though sampling differences and host biology cannot be excluded. CONCLUSIONS: We did not detect HA receptor-switching adaptation (the acquisition of human-type α2,6 receptor binding via Q226L/G228S) at any tested allele frequency in the US dairy H5N1 outbreak. Sub-consensus mammalian-adaptation signals exist at polymerase-complex sites but at low prevalence, are genotype-structured rather than independently recurrent, and require functional characterisation before informing risk assessment.

Dairy cattle↗

SARS-CoV genome polymorphism: a bioinformatics study.

A dataset of 103 SARS-CoV isolates (101 human patients and 2 palm civets) was investigated on different aspects of genome polymorphism and isolate classification. The number and the distribution of single nucleotide variations (SNVs) and insertions and deletions, with respect to a "profile", were determined and discussed ("profile" being a sequence containing the most represented letter per position). Distribution of substitution categories per codon positions, as well as synonymous and non-synonymous substitutions in coding regions of annotated isolates, was determined, along with amino acid (a.a.) property changes. Similar analysis was performed for the spike (S) protein in all the isolates (55 of them being predicted for the first time). The ratio Ka/Ks confirmed that the S gene was subjected to the Darwinian selection during virus transmission from animals to humans. Isolates from the dataset were classified according to genome polymorphism and genotypes. Genome polymorphism yields to two groups, one with a small number of SNVs and another with a large number of SNVs, with up to four subgroups with respect to insertions and deletions. We identified three basic nine-locus genotypes: TTTT/TTCGG, CGCC/TTCAT, and TGCC/TTCGT, with four subgenotypes. Both classifications proposed are in accordance with the new insights into possible epidemiological spread, both in space and time.

Amino Acid Sequence↗

PDA: a pipeline to explore and estimate polymorphism in large DNA databases.

Polymorphism studies are one of the main research areas of this genomic era. To date, however, no available web server or software package has been designed to automate the process of exploring and estimating nucleotide polymorphism in large DNA databases. Here, we introduce a novel software, PDA, Pipeline Diversity Analysis, that automatically can (i) search for polymorphic sequences in large databases, and (ii) estimate their genetic diversity. PDA is a collection of modules, mainly written in Perl, which works sequentially as follows: unaligned sequence retrieved from a DNA database are automatically classified by organism and gene, and aligned using the ClustalW algorithm. Sequence sets are regrouped depending on their similarity scores. Main diversity parameters, including polymorphism, synonymous and non-synonymous substitutions, linkage disequilibrium and codon bias are estimated both for the full length of the sequences and for specific functional regions. Program output includes a database with all sequences and estimations, and HTML pages with summary statistics, the performed alignments and a histogram maker tool. PDA is an essential tool to explore polymorphism in large DNA databases for sequences from different genes, populations or species. It has already been successfully applied to create a secondary database. PDA is available on the web at http://pda.uab.es/.

Databases, Nucleic Acid↗

MamPol: a database of nucleotide polymorphism in the Mammalia class.

Multi-locus and multi-species nucleotide diversity studies would benefit enormously from a public database encompassing high-quality haplotypic sequences with their associated genetic diversity measures. MamPol, 'Mammalia Polymorphism Database', is a website containing all the well-annotated polymorphic sequences available in GenBank for the Mammalia class grouped by name of organism and gene. Diversity measures of single nucleotide polymorphisms are provided for each set of haplotypic homologous sequences, including polymorphism at synonymous and non-synonymous sites, linkage disequilibrium and codon bias. Data gathering, calculation of diversity measures and daily updates are automatically performed using PDA software. The MamPol website includes several interfaces for browsing the contents of the database and making customizable comparative searches of different species or taxonomic groups. It also contains a set of tools for simple re-analysis of the available data and a statistics section that is updated daily and summarizes the contents of the database. MamPol is available at http://mampol.uab.es/ and can be downloaded via FTP.

Animals↗

Genotype-phenotype correlation in gemistocytic astrocytomas.

OBJECTIVE: Gemistocytic astrocytomas often behave aggressively and carry the least favorable prognosis among diffuse astrocytomas. The frequency of p53 mutations has been reported to be significantly higher in the gemistocytic variant as compared with other astrocytomas. METHODS: Between 1985 and 1998, we selected 25 tumor samples from among 201 samples from patients with gemistocytic astrocytomas operated on at the Mayo Clinic. Exons 5 to 8 of the p53 gene were sequenced using an automated deoxyribonucleic acid sequencer. Morphometric characterization of individual gemistocytes was performed using an image analysis program. RESULTS: Of 25 tissue samples analyzed, 16 were found to carry a p53 missense mutation (three in exon 5, three in exon 6, one in exon 7, and nine in exon 8), and one sequence variant was synonymous. Mutations were clustered at codons 151 (2 of 17 mutations), 193 (3 of 17 mutations), and 273 (5 of 17 mutations) of the p53 gene. Patients whose tumors carried a p53 mutation were significantly younger than other patients, and their tumors tended to accumulate more p53 protein than those of other patients. Phenotype analysis of gemistocytes revealed that the sizes of tumor cell nuclei and of entire tumor cells in the same tissue area were positively correlated. Smaller tumor cell nuclei tended to be less circular or more atypical. In addition, more atypical gemistocytes were found in tumors lacking a wild-type p53 allele as well as in tissue from patients whose postoperative survival was shorter. CONCLUSION: Our data confirm that the frequency of p53 mutations is significantly higher (approximately twofold) in gemistocytic astrocytomas as compared with other astrocytoma subtypes. Whether the high frequency of p53 mutations contributes to the more aggressive behavior of gemistocytic astrocytomas, however, remains unclear.

Adult↗