Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Not so different after all: a comparison of methods for detecting amino acid sites under selection.

We consider three approaches for estimating the rates of nonsynonymous and synonymous changes at each site in a sequence alignment in order to identify sites under positive or negative selection: (1) a suite of fast likelihood-based "counting methods" that employ either a single most likely ancestral reconstruction, weighting across all possible ancestral reconstructions, or sampling from ancestral reconstructions; (2) a random effects likelihood (REL) approach, which models variation in nonsynonymous and synonymous rates across sites according to a predefined distribution, with the selection pressure at an individual site inferred using an empirical Bayes approach; and (3) a fixed effects likelihood (FEL) method that directly estimates nonsynonymous and synonymous substitution rates at each site. All three methods incorporate flexible models of nucleotide substitution bias and variation in both nonsynonymous and synonymous substitution rates across sites, facilitating the comparison between the methods. We demonstrate that the results obtained using these approaches show broad agreement in levels of Type I and Type II error and in estimates of substitution rates. Counting methods are well suited for large alignments, for which there is high power to detect positive and negative selection, but appear to underestimate the substitution rate. A REL approach, which is more computationally intensive than counting methods, has higher power than counting methods to detect selection in data sets of intermediate size but may suffer from higher rates of false positives for small data sets. A FEL approach appears to capture the pattern of rate variation better than counting methods or random effects models, does not suffer from as many false positives as random effects models for data sets comprising few sequences, and can be efficiently parallelized. Our results suggest that previously reported differences between results obtained by counting methods and random effects models arise due to a combination of the conservative nature of counting-based methods, the failure of current random effects models to allow for variation in synonymous substitution rates, and the naive application of random effects models to extremely sparse data sets. We demonstrate our methods on sequence data from the human immunodeficiency virus type 1 env and pol genes and simulated alignments.

Amino Acids↗

Origin and evolution of overlapping genes in the family Microviridae.

The possibility of creating novel genes from pre-existing sequences, known as overprinting, is a widespread phenomenon in small viruses. Here, the origin and evolution of gene overlap in the bacteriophages belonging to the family Microviridae have been investigated. The distinction between ancestral and derived frames was carried out by comparing the patterns of codon usage in overlapping and non-overlapping genes. By this approach, a gradual increase in complexity of the phage genome--from an ancestral state lacking gene overlap to a derived state with a high density of genetic information--was inferred. Genes encoding less-essential proteins, yet playing a role in phage growth and diffusion, were predicted to be novel genes that originated by overprinting. Evaluation of the rates of synonymous and non-synonymous substitution yielded evidence for overlapping genes under positive selection in one frame and purifying selection in the alternative frame.

Coliphages↗

Receptor-like genes in the major resistance locus of lettuce are subject to divergent selection.

Disease resistance genes in plants are often found in complex multigene families. The largest known cluster of disease resistance specificities in lettuce contains the RGC2 family of genes. We compared the sequences of nine full-length genomic copies of RGC2 representing the diversity in the cluster to determine the structure of genes within this family and to examine the evolution of its members. The transcribed regions range from at least 7.0 to 13.1 kb, and the cDNAs contain deduced open reading frames of approximately 5. 5 kb. The predicted RGC2 proteins contain a nucleotide binding site and irregular leucine-rich repeats (LRRs) that are characteristic of resistance genes cloned from other species. Unique features of the RGC2 gene products include a bipartite LRR region with >40 repeats. At least eight members of this family are transcribed. The level of sequence diversity between family members varied in different regions of the gene. The ratio of nonsynonymous (Ka) to synonymous (Ks) nucleotide substitutions was lowest in the region encoding the nucleotide binding site, which is the presumed effector domain of the protein. The LRR-encoding region showed an alternating pattern of conservation and hypervariability. This alternating pattern of variation was also found in all comparisons within families of resistance genes cloned from other species. The Ka /Ks ratios indicate that diversifying selection has resulted in increased variation at these codons. The patterns of variation support the predicted structure of LRR regions with solvent-exposed hypervariable residues that are potentially involved in binding pathogen-derived ligands.

Amino Acid Sequence↗

Amino acid and nucleotide recurrence in aligned sequences: synonymous substitution patterns in association with global and local base compositions.

The tendency for repetitiveness of nucleotides in DNA sequences has been reported for a variety of organisms. We show that the tendency for repetitive use of amino acids is widespread and is observed even for segments conserved between human and Drosophila melanogaster at the level of >50% amino acid identity. This indicates that repetitiveness influences not only the weakly constrained segments but also those sequence segments conserved among phyla. Not only glutamine (Q) but also many of the 20 amino acids show a comparable level of repetitiveness. Repetitiveness in bases at codon position 3 is stronger for human than for D.melanogaster, whereas local repetitiveness in intron sequences is similar between the two organisms. While genes for immune system-specific proteins, but not ancient human genes (i.e. human homologs of Escherichia coli genes), have repetitiveness at codon bases 1 and 2, repetitiveness at codon base 3 for these groups is similar, suggesting that the human genome has at least two mechanisms generating local repetitiveness. Neither amino acid nor nucleotide repetitiveness is observed beyond the exon boundary, denying the possibility that such repetitiveness could mainly stem from natural selection on mRNA or protein sequences. Analyses of mammalian sequence alignments show that while the 'between gene' GC content heterogeneity, which is linked to 'isochores', is a principal factor associated with the bias in substitution patterns in human, 'within gene' heterogeneity in nucleotide composition is also associated with such bias on a more local scale. The relationship amongst the various types of repetitiveness is discussed.

Amino Acid Sequence↗

Mutation in the human gene for 3 beta-hydroxysteroid dehydrogenase type II leading to male pseudohermaphroditism without salt loss.

A 5-year-old XY pseudohermaphrodite was found to have a defect of steroid biosynthesis consistent with a partial deficiency of the enzyme 3 beta-hydroxysteroid dehydrogenase (3 beta-HSD). Circulating concentrations of delta 5 steroids and delta 5 urinary steroid metabolites were elevated and remained elevated after orchidectomy. There was no evidence of salt loss, plasma renin being within normal limits, and no detectable glucocorticoid abnormality. The coding sequences of the genes for 3 beta-HSD types I and II were amplified by PCR and screened for mutations by denaturing gradient gel electrophoresis (DGGE) and manual and automatic DNA sequencing. A mutation in the gene for 3 beta-HSD type II was observed at codon 173 (CTA-->CGA), leading in the affected patient to a homozygous substitution in which the leucine at residue 173 was altered to an arginine (L173R). The propositus's 2-year-old XX sister was also homozygous for L173R and showed the biochemical characteristics of partial 3 beta-HSD deficiency without clinical symptoms or signs. The mutation segregated as an autosomal recessive. Three related heterozygous adult females showed evidence of a small over-production of delta 5 steroids and steroid metabolites and a variable reduction in ovarian function. Concentrations of delta 5 steroids and steroid metabolites in the heterozygous father of the propositus were within the normal range. These data are discussed in relation to the endocrine causes of pseudohermaphroditism and hirsutism. Evidence for tight linkage between the genes for 3 beta-HSD types I and II was obtained using a microsatellite polymorphism in the third intron of the gene for 3 beta-HSD type II and synonymous and non-synonymous mutations and polymorphisms in the gene for 3 beta-HSD type I. The latter polymorphisms were located 88 bp apart at the 3' end of the type I coding sequence and could be physically resolved as haplotypes using DGGE. The application of DGGE to the analysis of mutations in members of a multigene family is discussed.

3-Hydroxysteroid Dehydrogenases↗

Sequence variation and gene duplication at MHC DQB loci of baiji (Lipotes vexillifer), a Chinese river dolphin.

The major histocompatibility complex (MHC) is a fundamental part of the vertebrate immune system, and the high variability in many MHC genes is thought to play an important role in the recognition of parasites. Baiji (Lipotes vexillifer) is one of the most endangered species in the world. Its wild population has declined to fewer than 100 individuals and has a very high risk of becoming extinct in the near future. In this study we present a first step in the molecular characterization of a DQB-like locus of baiji by nucleotide sequence analysis of the polymorphic exon 2 segments. In the examined 172 bp sequences from a group of 18 incidentally captured or stranded individuals, 48 variable sites were determined and 43 alleles were identified, many of which were represented by only one clone. Three to seven alleles were found in each individual, suggesting gene duplications. No deletion, insertion, or exceptional stop codon was detected, suggesting these alleles function in vivo. Phylogenetic reconstruction using neighbor joining grouped the 43 alleles into two distinct lineages, differing by seven nucleotides and four amino acids. Substitutions of amino acids tend to be clustered around sites postulated to be responsible for selective peptide recognition. In the peptide-binding region (PBR) of the DQB locus, the average number of nonsynonymous substitutions per site is greater than that of synonymous substitutions per site (0.1962 versus 0.0256, respectively). Nucleotide and amino acid sequences both showed a relatively high level of similarity (nucleotides 90.6%; amino acids 80.6%) to those of beluga whale (Delphinapterus leucas) and narwhal (Monodon monoceros). The high level of baiji MHC polymorphism revealed in the present study has not been reported in other cetaceans and could be a consequence of the small baiji population adapting to freshwater with a relatively high level of pathogens.

Amino Acid Sequence↗

Eight novel single nucleotide polymorphisms in ABCG2/BCRP in Japanese cancer patients administered irinotacan.

Eight novel single nucleotide polymorphisms (SNPs) were found in the gene encoding the ATP-binding cassette transporter, ABCG2/BCRP, from 60 Japanese individuals administered the anti-cancer drug irinotecan. The detected SNPs were as follows: 1) SNP, MPJ6_AG2005 (IVS2-93T>C); Gene Name, ABCG2; Accession Number, NT_006204; 2) SNP, MPJ6_AG2007 (IVS3+71_72 insT); Gene Name, ABCG2; Accession Number, NT_006204; 3) SNP, MPJ6_AG2012 (IVS6-204C>T); Gene Name, ABCG2; Accession Number, NT_006204; 4) SNP, MPJ6_AG2015 (at nucleotide 1098G>A (exon 9) from the A of the translation initiation codon); Gene Name, ABCG2; Accession Number, NT_006204; 5) SNP, MPJ6_AG2017 (1291T>C (exon 11)); Gene Name, ABCG2; Accession Number, NT_006204; 6) SNP, MPJ6_AG2019 (IVS11-135G>A); Gene Name, ABCG2; Accession Number, NT_006204; 7) SNP, MPJ6_AG2020 (1465T>C (exon 12)); Gene Name, ABCG2; Accession Number, NT_006204; 8) SNP, MPJ6_AG2023 (IVS13+65T>G); Gene Name, ABCG2; Accession Number, NT_006204.MPJ6_AG2015 was a synonymous SNP (E366E). MPJ6_AG2017 and MPJ6_AG2020 resulted in amino acid alterations, F431L and F489L, respectively.

Journal Article↗

Human type I hair keratin pseudogene phihHaA has functional orthologs in the chimpanzee and gorilla: evidence for recent inactivation of the human gene after the Pan-Homo divergence.

In addition to nine functional genes, the human type I hair keratin gene cluster contains a pseudogene, phihHaA (KRTHAP1), which is thought to have been inactivated by a single base-pair substitution that introduced a premature TGA termination codon into exon 4. Large-scale genotyping of human, chimpanzee, and gorilla DNAs revealed the homozygous presence of the phihHaA nonsense mutation in humans of different ethnic backgrounds, but its absence in the functional orthologous chimpanzee (cHaA) and gorilla (gHaA) genes. Expression analyses of the encoded cHaA and gHaA hair keratins served to highlight dramatic differences between the hair keratin phenotypes of contemporary humans and the great apes. The relative numbers of synonymous and non-synonymous substitutions in the phihHaA and cHaA genes, as inferred by using the gHaA gene as an outgroup, suggest that the human hHaA gene was inactivated only recently, viz., less than 240,000 years ago. This implies that the hair keratin phenotype of hominids prior to this date, and after the Pan-Homo divergence some 5.5 million years ago, could have been identical to that of the great apes. In addition, the homozygous presence of the phihHaA exon 4 nonsense mutation in some of the earliest branching lineages among extant human populations lends strong support to the "single African origin" hypothesis of modern humans.

Amino Acid Sequence↗

Evolutionary history of the Asr gene family.

The Asr gene family is widespread in higher plants. Most Asr genes are up-regulated under different environmental stress conditions and during fruit ripening. ASR proteins are localized in the nucleus and their likely function is transcriptional regulation. In cultivated tomato, we identified a novel fourth family member, named Asr4, which maps close to its sibling genes Asr1-Asr2-Asr3 and displays an unshared region coding for a domain containing a 13-amino acid repeat. In this work we were able to expand our previous analysis for Asr2 and investigated the coding regions of the four known Asr paralogous genes in seven tomato species from different geographic locations. In addition, we performed a phylogenetic analysis on ASR proteins. The first conclusion drawn from this work is that tomato ASR proteins cluster together in the tree. This observation can be explained by a scenario of concerted evolution or birth and death of genes. Secondly, our study showed that Asr1 is highly conserved at both replacement and synonymous sites within the genus Lycopersicon. ASR1 protein sequence conservation might be associated with its multiple functions in different tissues while the low rate of synonymous substitutions suggests that silent variation in Asr1 is selectively constrained, which is probably related to its high expression levels. Finally, we found that Asr1 activation under water stress is not conserved between Lycopersicon species.

Amino Acid Sequence↗

A comprehensive compilation of 1001 nucleotide sequences coding for proteins from the yeast Saccharomyces cerevisiae (= ListA2)

The amount of nucleotide sequence data is increasing exponentially. We therefore continued our effort to make a comprehensive database for the yeast Saccharomyces cerevisiae. In this database (ListA2) we have compiled 1001 protein coding sequences from this organism. Each sequence has been attributed a single genetic name and in the case of allelic duplicated sequences, synonyms are given, if necessary. For the nomenclature we have introduced a standard principle for naming gene sequences based on priority rules. We have also applied a simple method to distinguish duplicated sequences of one and the same gene from non-allelic sequences of duplicated genes. By using these principles we have sorted out a lot of confusion in the literature and databanks. Along with the genetic name, the mnemonic from the EMBL databank, the codon bias, reference of the publication of the sequence and the EMBL accession numbers are included for each entry. The database is available on request.

Alleles↗

LISTA, a comprehensive compilation of nucleotide sequences encoding proteins from the yeast Saccharomyces.

The amount of nucleotide sequence data is increasing exponentially. We therefore made an effort to make a comprehensive database (LISTA) for the yeast Saccharomyces cerevisiae. Each sequence has been attributed a single genetic name and in the case of allelic duplicated sequences, synonyms are given, if necessary. For the nomenclature we have introduced a standard principle for naming gene sequences based on priority rules. We have also applied a simple method to distinguish duplicated sequences of one and the same gene from non-allelic sequences of duplicated genes. By using these principles we have sorted out a lot of confusion in the literature and databanks. Along with the genetic name, the mnemonic from the EMBL databank, the codon bias, reference of the publication of the sequence and the EMBL accession numbers are included in each entry.

Base Sequence↗

Rapp-Hodgkin ectodermal dysplasia syndrome: the clinical and molecular overlap with Hay-Wells syndrome.

We report on the clinical and molecular abnormalities in a 7-month-old girl and her mother with an ectodermal dysplasia disorder that most closely resembles Rapp-Hodgkin syndrome (RHS). At birth, the child had bilateral cleft palate, a narrow pinched nose, small chin, and hypoplastic nipples, and suffered from respiratory distress, feeding difficulties, and poor weight gain, although developmental progress was normal. Her mother had a cleft palate, sparse hair, high forehead, dental anomalies, a narrow nose, dysplastic nails, and reduced sweating. Sequencing of the p63 gene in genomic DNA from both individuals revealed a heterozygous frameshift mutation, 1721delC, in exon 14. This mutation has not been described previously and is the seventh report of a pathogenic p63 gene mutation in RHS. The frameshift results in changes to the tail of p63 with the addition of 90 missense amino acids downstream and a delayed termination codon that extends the protein by 21 amino acids. This mutation is predicted to disrupt the normal repressive function of the transactivation inhibitory domain leading to gain-of-function for at least two isoforms of the p63 transcription factor. The expanding p63 mutation database demonstrates that there is considerable overlap between the molecular pathology of RHS and Hay-Wells syndrome, with identical mutations in some cases, and that these two disorders may in fact be synonymous.

Adult↗

Physicochemical evolution and molecular adaptation of the cetacean and artiodactyl cytochrome b proteins.

Cetaceans have most likely experienced metabolic shifts since evolutionarily diverging from their terrestrial ancestors, shifts that may be reflected in the proteins such as cytochrome b that are responsible for metabolic efficiency. However, accepted statistical methods for detecting molecular adaptation are largely biased against even moderately conservative proteins because the primary criterion involves a comparison of nonsynonymous and synonymous substitution rates (dN/dS); they do not allow for the possibility that adaptation may come in the form of very few amino acid changes. We apply the MM01 model to the possible molecular adaptation of cytochrome b among cetaceans because it does not rely on a dN/dS ratio, instead evaluating positive selection in terms of the amino acid properties that comprise protein phenotypes that selection at the molecular level may act upon. We also apply the codon-degeneracy model (CDM), which focuses on evaluating overall patterns of nucleotide substitution in terms of base exchange, codon position, and synonymy to estimate the overall effect of selection. Using these relatively new models, we characterize the molecular adaptation that has occurred in the cetacean cytochrome b protein by comparing revealed amino acid replacement patterns to those found among artiodactyls, the modern terrestrial mammals found to be most closely related to cetaceans. Our findings suggest that several regions of the cetacean cytochrome b protein have experienced molecular adaptation. Also, these adaptations are spatially associated with domain structure, protein function, and the structure and function of the cytochrome bc(1) complex and its constituents. We also have found a general correlation between the results of the analytical software programs TreeSAAP (which implements the MM01 model) and CDM (which implements the codon-degeneracy model).

Adaptation, Physiological↗

Cloning and chromosomal mapping of URA3 genes of Pichia farinosa and P. sorbitophila encoding orotidine-5'-phosphate decarboxylase.

The PfURA3 gene, which encodes orotidine-5'-phosphate decarboxylase, of osmotolerant yeast Pichia farinosa NFRI 3,621, was cloned by complementation of the ura3 mutation of Saccharomyces cerevisiae. The nucleotide sequence of the PfURA3 gene and its deduced amino acid sequence indicated that the gene encodes a protein (PfUra3p) of 267 amino acids. Pulsed-field gel electrophoresis and subsequent Southern blot analysis showed that the genome of P. farinosa NFRI 3621 consisted of seven chromosomes, each approximately 1.1-2.2 Mb in size (11.8 Mb in total) and that PfURA3 was located on chromosome V. Pichia sorbitophila is considered as a synonym of P. farinosa. The genome of P. sorbitophila IFO10021 may consist of 12 chromosomes, each approximately 1.2-2.2 Mb in size. P. sorbitophila has two copies of URA3 genes, termed PsURA3 and PsURA30, which were located on chromosome VIII and III, respectively. The difference between PfURA3 and PsURA3 was only two amino acid substitutions, whereas that between PsURA3 and PsURA30 was six amino acid substitutions and the deletion of the C-terminal amino acid by a stop codon insertion. The sequences of PfURA3, PsURA3 and PsURA30 have been deposited in the DDBJ data library under Accession Nos AB071417, AB109042 and AB109043, respectively.

Amino Acid Sequence↗

Sequence analysis of Potato leafroll virus isolates reveals genetic stability, major evolutionary events and differential selection pressure between overlapping reading frame products.

In order to investigate the genetic diversity of Potato leafroll virus (PLRV), seven new complete genomic sequences of isolates collected worldwide were compared with the five sequences available in GenBank. Then, a restricted polymorphic region of the genome was chosen to further analyse new sequences. The sequences of PLRV open reading frames (ORFs) 3 and 4 were also compared with those of two other poleroviruses and the non-synonymous to synonymous substitution ratio distribution was analysed in overlapping and non-overlapping regions of the genome using maximum-likelihood models. Results confirmed that PLRV sequences from around the world are very closely related and showed that the region encoding protein P0 allowed the detection of three groups of isolates. When compared to other poleroviruses, PLRV was the most conserved in both ORFs 3 and 4. However, the results suggest that important events, such as deletion, mutation at a stop codon and intraspecific homologous recombination events, have occurred during the evolution of PLRV. Finally, it was shown that the translation products of ORFs 0 and 3 are significantly more conserved than those of the overlapping ORFs 1 and 4, respectively. All together, the results allow the proposal of new hypotheses to explain the apparent genetic stability of PLRV and its evolution.

Biological Evolution↗

Hepatitis C in human immunodeficiency virus-coinfected patients: increased variability in the hypervariable envelope coding domain.

Patients coinfected with the hepatitis C virus (HCV) and the human immunodeficiency virus (HIV) were studied with regard to nucleotide sequence variability in the E2/NS1 first hypervariable region of the HCV genome. The nucleotide variability within individual patients was compared to patients infected only with HCV. The proportion of predicted synonymous and nonsynonymous amino acid changes, and the relationship to putative high-antigenicity sites, were evaluated in the hypervariable envelope domain. Ninety-one clones from 10 patients with HCV/HIV coinfection were sequenced, following polymerase chain reaction (PCR) amplification of the hypervariable region. The control HCV group included 53 clones from 7 patients. Sequence analysis encompassed the region coding for amino acids 384 to 414. Consensus sequences from each patient were used as the internal standard for nonsynonymous amino acid codon variability. Cumulative proportional comparison at each amino acid site revealed increased variability in HCV RNA from patients with HCV/HIV coinfection versus HCV alone (P < .05). The greatest variability was observed at amino acids 386, 397, 400, 402, 405, 407, and 414, with >l0 percent clonal variation at these sites. Jameson-Wolf plots were used to predict putative high-antigenicity domains. Nonsynonymous clonal variation resulted in alteration of putative antigenic sites within the hypervariable region. All clones had at least one high-probability site. Clones with unique predicted antigenic domains were observed more frequently in HIV/HCV coinfected patients, and, independent of viral titer, were consistent with increased sequence variability. These data suggest an accumulation of envelope variants in the HCV/HIV coinfected patients, which could be related to ineffective viral clearance, and may help explain prior reports of interferon (IFN) resistance in this patient group.

Adult↗

Association of a G2014A transition in exon 8 of the estrogen receptor-alpha gene with postmenopausal osteoporosis.

We report the association of a newly identified synonymous G2014A single nucleotide polymorphism (SNP) which does not alter the amino acid sequence in exon 8 of the estrogen receptor-alpha (ERalpha) gene with osteoporosis in Thai postmenopausal women. Subjects consisted of 228 postmenopausal women aged more than 55 years divided into two groups--with vertebral or femoral osteoporosis (n = 106) or without osteoporosis (n = 122)--according to bone mineral density (BMD) criteria. The exon 8 G2014A SNP, which is 6 nucleotides upstream from the end of the stop codon, was identified by PCR-RFLP. Data are expressed as the mean and 95% CI. The allele frequency of the G2014A polymorphism was 26.4% in osteoporotic subjects and was significantly higher than that in non-osteoporotic women (15.2%) (p<0.05). By stepwise logistic regression analysis, it was found that the G2014A polymorphism was related to the presence of osteoporosis (odds ratio 2.7 per A allele, 95% CI 1.49-4.76) independently of body weight (odds ratio 0.93 per kg, 95% CI 0.89-0.96) and years since menopause (odds ratio 1.12 per year, 95% CI 1.08-1.19). In a multiple linear regression model, L2-L4 BMD of osteoporotic subjects was associated with body weight (p<0.05), endogenous estradiol levels (p<0.05) and the G2014A genotype (p<0.001), while it was related only to body weight (p<0.05) and estradiol levels in non-osteoporotic women (p<0.05). We conclude that a G2014A SNP in exon 8 of ERalpha is associated with the presence and severity of postmenopausal osteoporosis. Linkage disequilibrium between this polymorphism and the 3'-untranslated region of the ERalpha gene which may participate in the regulation of ERalpha gene expression remains to be determined.

Aged↗