Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Sequence evolution in bacterial endosymbionts having extreme base compositions.

A major limitation on ability to reconstruct bacterial evolution is the lack of dated ancestors that might be used to evaluate and calibrate molecular clocks. Vertically transmitted symbionts that have cospeciated with animal hosts offer a firm basis for calibrating sequence evolution in bacteria, since fossils of the hosts can be used to date divergence events. Sequences for a functionally diverse set of genes have been obtained for bacterial endosymbionts (Buchnera) from two pairs of aphid host species, each pair diverging 50-70 MYA. Using these dates and estimated numbers of Buchnera generations per year, we calculated rates of base substitution for neutral and selected sites of protein-coding genes and overall rates for rRNA genes. Buchnera shows homogeneity among loci with regard to synonymous rate. The Buchnera synonymous rate is about twice that for low-codon-bias genes of Escherichia coli-Salmonella typhimurium on an absolute timescale, and fourfold higher on a generational timescale. Nonsynonymous substitutions show a greater rate disparity in favor of Buchnera, a result consistent with a genomewide decrease in selection efficiency in Buchnera. Ratios of synonymous to nonsynonymous substitutions differ for the two pairs of Buchnera, indicating that selection efficiency varies among lineages. Like numerous other intracellular bacteria, such as Rickettsia and Wolbachia, Buchnera has accumulated amino acids with codons rich in A or T. Phylogenetic reconstruction of amino acid replacements indicates that replacements yielding increased A + T predominated early in the evolution of Buchnera, with the trend slowing or stopping during the last 50 Myr. This suggests that base composition in Buchnera has approached a limit enforced by selective constraint acting on protein function.

AT Rich Sequence↗

Specific compositional patterns of synonymous positions in homologous mammalian genes.

All 69 homologous coding sequences that are currently available in four mammalian orders were aligned and the synonymous (ie., third) positions of quartet (fourfold degenerate) codons were divided into three classes (that will be called conserved, intermediate, and variable), according to whether they show no change, one change, and more than one change, respectively. The three classes were analyzed in their compositional patterns. In the majority of GC-rich genes, the three classes of positions (but especially conserved positions) exhibited significantly different base compositions compared to expectations based on a "random" substitution process from the "ancestral" (consensus) sequence to the present-day (actual) sequences. Significant differences were rare in GC-poor genes. An analysis of the present results indicates that natural selection plays a role in the synonymous nucleotide substitution process, especially in GC-rich genes which represent the vast majority of mammalian genes.

Animals↗

Genetic polymorphism and natural selection in the malaria parasite Plasmodium falciparum.

We have studied the genetic polymorphism at 10 Plasmodium falciparum loci that are considered potential targets for specific antimalarial vaccines. The polymorphism is unevenly distributed among the loci; loci encoding proteins expressed on the surface of the sporozoite or the merozoite (AMA-1, CSP, LSA-1, MSP-1, MSP-2, and MSP-3) are more polymorphic than those expressed during the sexual stages or inside the parasite (EBA-175, Pfs25, PF48/45, and RAP-1). Comparison of synonymous and nonsynonymous substitutions indicates that natural selection may account for the polymorphism observed at seven of the 10 loci studied. This inference depends on the assumption that synonymous substitutions are neutral, which we test by analyzing codon bias and G+C content in a set of 92 gene loci. We find evidence for an overall trend towards increasing A+T richness, but no evidence for mutation bias. Although the neutrality of synonymous substitutions is not definitely established, this trend towards an A+T rich genome cannot explain the accumulation of substitutions at least in the case of four genes (AMA-1, CSP, LSA-1, and PF48/45) because the Gleft and right arrow C transversions are more frequent than expected. Moreover, the Tajima test manifests positive natural selection for the MSP-1 and, less strongly, MSP-3 polymorphisms; the McDonald-Kreitman test manifests natural selection at LSA-1 and PF48/45. We conclude that there is definite evidence for positive natural selection in the genes encoding AMA-1, CSP, LSA-1, MSP-1, and Pfs48/45. For four other loci, EBA-175, MSP-2, MSP-3, and RAP-1, the evidence is limited. No evidence for natural selection is found for Pfs25.

Animals↗

Choice of base at silent codon site 3 is not selectively neutral in eucaryotic structural genes: it maintains excess short runs of weak and strong hydrogen bonding bases.

On the average in the coding sequences of 30 eucaryotic structural genes the weak hydrogen bonding, W, (A or T) or strong hydrogen bonding, S, (C or G) base in codon site 3 was chosen to be unlike its neighbors on both sides up to two sites away. This preference produced the nonrandom excess of runs W and S of length one and two and the deficit of long runs observed earlier (Blaisdell 1982). The neighbors in the different codon, 3' to codon site 3, were as important in determining the choice as were the neighbors 5' in the same codon. Every amino acid except methionine and tryptophan, of least frequent occurrence, permits choice of W or S. The persistence of this preference could explain the observation that the rate of substitution of codon site 3 in functional genes is considerably less than in synonymous pseudo genes.

Amino Acids↗

Evidence for positive selection in foot-and-mouth disease virus capsid genes from field isolates.

The nature of selection on capsid genes of foot-and-mouth disease virus (FMDV) was characterized by examining the ratio of nonsynonymous to synonymous substitutions in 11 data sets of sequences obtained from six different serotypes of FMDV. Using a method of analysis that assigns each codon position to one of a number of estimated values of nonsynonymous to synonymous ratio, significant evidence of positive selection was identified in 5 data sets, operating at 1-7% of codon positions. Evidence of positive selection was identified in complete capsid sequences of serotypes A and C and in VP1 sequences of serotypes SAT 1 and 2. Sequences of serotype SAT-2 recovered from a persistently infected African buffalo also revealed evidence for positive selection. Locations of codons under positive selection coincide closely with those of antigenic sites previously identified with the use of monoclonal antibody escape mutants. The vast majority of codons are under mild to strong purifying selection. However, these results suggest that arising antigenic variants benefit from a selective advantage in their interaction with the immune system, either during the course of an infection or in transmission to individuals with previous exposure to antigen. Analysis of amino acid usage at sites under positive selection indicates that this selective advantage can be conferred by amino acid substitutions that share physicochemically similar properties.

Amino Acid Sequence↗

Substrate recognition of type III secretion machines--testing the RNA signal hypothesis.

Secretion by the type III pathway of Gram-negative microbes transports polypeptides into the extracellular medium or into the cytoplasm of host cells during infection. In pathogenic Yersinia spp., type III machines recognize 14 different Yop protein substrates via discrete signals genetically encoded in 7-15 codons at the 5' portion of yop genes. Although the signals necessary and sufficient for substrate recognition of Yop proteins have been mapped, a clear mechanism on how proteins are recognized by the machinery and then initiated into the transport pathway has not yet emerged. As synonymous substitutions, mutations that alter mRNA sequence but not codon specificity, affect the function of some secretion signals, recent work with several different microbes tested the hypothesis of an RNA-encoded secretion signal for polypeptides that travel the type III pathway. This review summarizes experimental observations and mechanistic models for substrate recognition in this field.

Amino Acid Sequence↗

The codon-degeneracy model of molecular evolution.

Mitochondrial genetic codons can be categorized by four patterns of nucleotide-site degeneracy based on varying combinations of twofold- or nondegenerate sites at first codon positions and twofold- or fourfold-degenerate sites at third codon positions. Herein, a model of molecular evolution is introduced that uses these patterns to calculate expected substitution frequencies for each codon position and substitution type relative to overall number of synonymous or nonsynonymous substitutions. Regions of the pocket gopher cytochrome oxidase subunit I (COI) and cytochrome b (cyt-b) genes are analyzed using this model. Chi-square distributions are used to produce relative goodness-of-fit (GF) scores for measuring the difference between substitution frequencies predicted by the codon-degeneracy model (CDM), and frequencies inferred using a well-supported phylogenetic tree of closely related species. The GF scores for expected and observed synonymous (GF(syn) = 0.429, p = 0.807) and nonsynonymous (GF(ns) = 2.309, p = 0.679) substitution frequencies resulted in a failure to reject the CDM as a null hypothesis for the molecular evolution of COI and cyt-b in pocket gophers. Alternative tree topologies and calculations of transition bias for these data result in higher GF scores.

Animals↗

Evidence of diversifying selection in human papillomavirus type 16 E6 but not E7 oncogenes.

Human papillomavirus type 16 is a common sexually transmitted pathogen capable of giving rise to cervical intraepithelial neoplasia and invasive carcinoma through the expression and activity of two adjacent oncogenes: E6 and E7. Naturally occurring amino acid variation is commonly observed in the E6 protein but to a much lesser extent in E7. In order to investigate the evolutionary mechanisms involved in the generation and maintenance of this variation, we examine 42 distinct E6-E7 haplotypes using codon-based genealogical techniques. These techniques involve estimation of the ratio of nonsynonymous to synonymous substitutions (dn/ds) and allow testing for directional (positive) natural selection. Positive selection was detected for four codon sites within the E6 oncogene but not in any E7 codons. The amino acid compositions and locations of selected sites are described. Possible sources of natural selection including antiviral immune pressure and polymorphism of host cellular proteins are discussed.

Amino Acid Sequence↗

Evolution of a new nonclassical MHC class I locus in two Old World primate species.

HLA-G is a nonclassical major histocompatibility complex (MHC) class I molecule that is expressed only in the human placenta, suggesting that it plays an important role at the fetal-maternal interface. In rhesus monkeys, which have similar placentation to humans, the HLA-G orthologue is a pseudogene. However, rhesus monkeys express a novel placental MHC class I molecule, Mamu-AG, which has HLA-G-like characteristics. Phylogenetic analysis of AG alleles in two Old World primate species, the baboon and the rhesus macaque, revealed limited diversity characteristic of a nonclassical MHC class I locus. Gene trees constructed using classical and nonclassical primate MHC class I alleles demonstrated that the AG locus was most closely related to the classical A locus. Interestingly, gene tree analyses suggested that the AG alleles were most closely related to a subset of A alleles which are the products of an ancestral interlocus recombination event between the A and B loci. Calculation of the rates of synonymous and nonsynonymous substitution at the AG locus revealed that positive selection was not acting on the codons encoding the peptide binding region. In exon 4, however, the rate of nonsynonymous substitution was significantly lower than the rate of synonymous substitution, suggesting that negative selection was acting on these codons.

Amino Acid Sequence↗

Possible identity of transcription and translation signals in early vital systems.

The distribution of codons was analysed in three classes of eukaryote proteins having widely different evolutionary rates: 78 histones, 40 tubulins, and seven fibrinogens. In this set of genes, (i) it was confirmed that codons which are components of known transcription signals, like ATA, are used infrequently when a synonym is available, particularly in the more constrained proteins, and (ii) it was observed that the three codons which have an iso-accepting transfer with anticodon UAA, UAG or UGA are also suppressed. Then, the distribution of UAA, UAG and UGA trimers was studied in 498 tDNAs and 198 rDNAs. It was found that these trimers are weakly but significantly suppressed in tDNAs and to a lesser extent in rDNAs. It was advanced that the present suppression of ATA, which codes for Methionine in several mitochondria, and of the TAA, TAG and TGA trimers in tDNAs, might be an indication that at the very early stages of the evolution of translation and transcription the signals for initiation and termination were shared by the two processes.

Amino Acid Sequence↗

Arrangement and nucleotide sequence of the gene (fus) encoding elongation factor G (EF-G) from the hyperthermophilic bacterium Aquifex pyrophilus: phylogenetic depth of hyperthermophilic bacteria inferred from analysis of the EF-G/fus sequences.

The gene fus (for EF-G) of the hyperthermophilic bacterium Aquifex pyrophilus was cloned and sequenced. Unlike the other bacteria, which display the streptomycin-operon arrangement of EF genes (5'-rps12-rps7-fus-tuf-3'), the Aquifex fus gene (700 codons) is not preceded by the two small ribosomal subunit genes although it is still followed by a tuf gene (for EF-Tu). The opposite strand upstream from the EF-G coding locus revealed an open reading frame (ORF) encoding a polypeptide having 52.5% identity with an E. coli protein (the pdxJ gene product) involved in pyridoxine condensation. The Aquifex EF-G was aligned with available homologs representative of Deinococci, high G+C Gram positives, Proteobacteria, cyanobacteria, and several Archaea. Outgroup-rooted phylogenies were constructed from both the amino acid and the DNA sequences using first and second codon positions in the alignments except sites containing synonymous changes. Both datasets and alternative tree-making methods gave a consistent topology, with Aquifex and Thermotoga maritima (a hyperthermophile) as the first and the second deepest offshoots, respectively. However, the robustness of the inferred phylogenies is not impressive. The branching of Aquifex more deeply than Thermotoga and the branching of Thermotoga more deeply than the other taxa examined are given at bootstrap values between 65 and 70% in the fus-based phylogenies, while the EF-G(2)-based phylogenies do not provide a statistically significant level of support (< or = 50% bootstrap confirmation) for the emergence of Thermotoga between Aquifex and the successive offshoot (Thermus genus). At present, therefore, the placement of Aquifex at the root of the bacterial tree, albeit reproducible, can be asserted only with reservation, while the emergence of Thermotoga between the Aquificales and the Deinococci remains (statistically) indeterminate.

Amino Acid Sequence↗

Evidence for the adaptive evolution of the carbon fixation gene rbcL during diversification in temperature tolerance of a clade of hot spring cyanobacteria.

Determining the molecular basis of enzyme adaptation is central to understanding the evolution of environmental tolerance but is complicated by the fact that not all amino acid differences between ecologically divergent taxa are adaptive. Analysing patterns of nucleotide sequence evolution can potentially guide the investigation of protein adaptation by identifying candidate codon sites on which diversifying selection has been operating. Here, I test whether there is evidence for molecular adaptation of the carbon fixation gene rbcL for a clade of hot spring cyanobacteria in the genus Synechococcus that has diverged in thermotolerance. Amino acid replacements during Synechococcus radiation have resulted in an increase in the number of hydrophobic residues in the RbcLs of more thermotolerant strains. A similar increase in hydrophobicity has been observed for many thermostable proteins. Maximum likelihood models which allow for heterogeneity among codon sites in the ratio of nonsynonymous to synonymous nucleotide substitutions estimated a class of amino acid sites as a target of positive selection. Depending on the model, a single amino acid site that interacts with a flexible element involved in the opening and closing of the active site was estimated with either low or moderate support to be a member of this class. Site-directed mutagenesis approaches are being explored in order to directly test its adaptive significance.

Adaptation, Biological↗

Polymorphisms in nucleotide excision repair genes, smoking and breast cancer in African Americans and whites: a population-based case-control study.

Polymorphisms exist in several genes involved in nucleotide excision repair (NER), the principal pathway for removal of smoking-induced DNA damage. An epidemiologic study was conducted to determine whether these polymorphisms modify the association between smoking and breast cancer. DNA samples and exposure histories were analyzed as part of a large population-based case-control study of breast cancer in North Carolina. The study population included 2311 cases (894 African Americans, 1417 whites) and 2022 controls (788 African Americans, 1234 whites). Odds ratios (ORs) were calculated for breast cancer and smoking, and for breast cancer and nine non-synonymous coding polymorphisms in six NER genes (XPD codons 312 and 751, RAD23B codon 249, XPG codon 1104, XPC codon 939, XPF codons 415 and 662, and ERCC6 codons 1213 and 1230). Modification of ORs for smoking by single and combined NER genotypes was investigated. In this study population, smoking was more strongly associated with breast cancer in African American women compared with white women. Among African American women, the association of breast cancer and smoking was strongest among women with specific combinations of NER genotypes. Evidence for multiplicative interaction was found between combined NER genotypes and smoking dose (likelihood ratio test P = 0.06), duration (P = 0.09), time since cessation (P = 0.02), age at initiation (P = 0.04) and former smoking (P = 0.03). No interactions were observed in white women. Therefore, polymorphisms in NER genes may modify the relationship between breast cancer and smoking. These results are consistent with previous evidence of exposure-specific p53 mutations in breast tumors from current and former smokers, suggesting that smoking may play a role in breast cancer etiology.

Adult↗

The cobalamin (coenzyme B12) biosynthetic genes of Escherichia coli.

The enteric bacterium Escherichia coli synthesizes cobalamin (coenzyme B12) only when provided with the complex intermediate cobinamide. Three cobalamin biosynthetic genes have been cloned from Escherichia coli K-12, and their nucleotide sequences have been determined. The three genes form an operon (cob) under the control of several promoters and are induced by cobinamide, a precursor of cobalamin. The cob operon of E. coli comprises the cobU gene, encoding the bifunctional cobinamide kinase-guanylyltransferase; the cobS gene, encoding cobalamin synthetase; and the cobT gene, encoding dimethylbenzimidazole phosphoribosyltransferase. The physiological roles of these sequences were verified by the isolation of Tn10 insertion mutations in the cobS and cobT genes. All genes were named after their Salmonella typhimurium homologs and are located at the corresponding positions on the E. coli genetic map. Although the nucleotide sequences of the Salmonella cob genes and the E. coli cob genes are homologous, they are too divergent to have been derived from an operon present in their most recent common ancestor. On the basis of comparisons of G+C content, codon usage bias, dinucleotide frequencies, and patterns of synonymous and nonsynonymous substitutions, we conclude that the cob operon was introduced into the Salmonella genome from an exogenous source. The cob operon of E. coli may be related to cobalamin synthetic genes now found among non-Salmonella enteric bacteria.

Bacterial Proteins↗

Reevaluation of amino acid variability of the human immunodeficiency virus type 1 gp120 envelope glycoprotein and prediction of new discontinuous epitopes.

To elucidate the evolutionary mechanisms of the human immunodeficiency virus type 1 gp120 envelope glycoprotein at the single-site level, the degree of amino acid variation and the numbers of synonymous and nonsynonymous substitutions were examined in 186 nucleotide sequences for gp120 (subtype B). Analyses of amino acid variabilities showed that the level of variability was very different from site to site in both conserved (C1 to C5) and variable (V1 to V5) regions previously assigned. To examine the relative importance of positive and negative selection for each amino acid position, the numbers of synonymous and nonsynonymous substitutions that occurred at each codon position were estimated by taking phylogenetic relationships into account. Among the 414 codon positions examined, we identified 33 positions where nonsynonymous substitutions were significantly predominant. These positions where positive selection may be operating, which we call putative positive selection (PS) sites, were found not only in the variable loops but also in the conserved regions (C1 to C4). In particular, we found seven PS sites at the surface positions of the alpha-helix (positions 335 to 347 in the C3 region) in the opposite face for CD4 binding. Furthermore, two PS sites in the C2 region and four PS sites in the C4 region were detected in the same face of the protein. The PS sites found in the C2, C3, and C4 regions were separated in the amino acid sequence but close together in the three-dimensional structure. This observation suggests the existence of discontinuous epitopes in the protein's surface including this alpha-helix, although the antigenicity of this area has not been reported yet.

Amino Acid Sequence↗

A likelihood approach for comparing synonymous and nonsynonymous nucleotide substitution rates, with application to the chloroplast genome.

A model of DNA sequence evolution applicable to coding regions is presented. This represents the first evolutionary model that accounts for dependencies among nucleotides within a codon. The model uses the codon, as opposed to the nucleotide, as the unit of evolution, and is parameterized in terms of synonymous and nonsynonymous nucleotide substitution rates. One of the model's advantages over those used in methods for estimating synonymous and nonsynonymous substitution rates is that it completely corrects for multiple hits at a codon, rather than taking a parsimony approach and considering only pathways of minimum change between homologous codons. Likelihood-ratio versions of the relative-rate test are constructed and applied to data from the complete chloroplast DNA sequences of Oryza sativa, Nicotiana tabacum, and Marchantia polymorpha. Results of these tests confirm previous findings that substitution rates in the chloroplast genome are subject to both lineage-specific and locus-specific effects. Additionally, the new tests suggest tha the rate heterogeneity is due primarily to differences in nonsynonymous substitution rates. Simulations help confirm previous suggestions that silent sites are saturated, leaving no evidence of heterogeneity in synonymous substitution rates.

Chloroplasts↗

Evidence for translational selection in codon usage in Echinococcus spp.

We analysed the intragenomic variation in codon usage in Echinococcus spp. by correspondence analysis. This approach detected a trend among genes which was correlated with expression levels. Among the (presumed) highly expressed sequences we found an increased usage of a subset of codons, almost all of them G- or C- ending. Since an increase in these bases at the synonymous sites is against the mutational bias (these genomes are slightly A+-T- rich), we conclude that codon usage in Echinococcus is the result of an equilibrium between compositional pressure and selection, the latter acting at the level of translation, mainly on highly expressed genes. This is the first report where translational selection for codon usage is detected among Platyhelminthes.

Animals↗