Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

HIV-1 strains specific for Dutch injecting drug users in heterosexually infected individuals in The Netherlands.

OBJECTIVE: To study the molecular epidemiology of HIV-1 subtype B amongst heterosexually infected individuals in The Netherlands. DESIGN: The study population comprised 54 individuals infected by subtype B viruses through heterosexual contacts. Serum samples were collected between 1988 and 1996. METHODS: Sequences of the gp120 V3 region were obtained from serum samples and analysed by using the signature pattern and phylogenetic methods. RESULTS: In 22 (41%) out of 54 subtype B sequences from heterosexually infected individuals, the synonymous nucleotide substitution in the second glycine codon at the tip of the V3 loop (the GGC pattern), previously identified as specific for Dutch injecting drug users (IDU), was found. The other previously described IDU sequence patterns were observed significantly more often among GGC- than among non-GGC-containing sequences. In addition, we identified another amino-acid change specific for the GGC sequences. In the phylogenetic and principal coordinate analyses, the GGC sequences from heterosexually infected individuals clustered separately from the non-GGC sequences and together with the IDU consensus sequence. Both the nonsynonymous and particularly the synonymous distances amongst the GGC sequences were significantly lower than amongst the non-GGC sequences. CONCLUSIONS: Our data provide evidence for a common origin of the viruses in Dutch IDU and the GGC viruses in heterosexuals. We suggest that a considerable proportion of the viruses in heterosexually infected individuals in The Netherlands may have originated from Dutch IDU.

Adult↗

Evolution of base composition and codon usage bias in the genus Flavivirus.

The extent to which base composition and codon usage vary among RNA viruses, and the possible causes of this bias, is undetermined in most cases. A maximum-likelihood statistical method was used to test whether base composition and codon usage bias covary with arthropod association in the genus Flavivirus, a major source of disease in humans and animals. Flaviviruses are transmitted by mosquitoes, by ticks, or directly between vertebrate hosts. Those viruses associated with ticks were found to have a significantly lower G+C content than non-vector-borne flaviviruses and this difference was present throughout the genome at all amino acids and codon positions. In contrast, mosquito-borne viruses had an intermediate G+C content which was not significantly different from those of the other two groups. In addition, biases in dinucleotide and codon usage that were independent of base composition were detected in all flaviviruses, but these did not covary with arthropod association. However, the overall effect of these biases was slight, suggesting only weak selection at synonymous sites. A preliminary analysis of base composition, codon usage, and vector specificity in other RNA virus families also revealed a possible association between base composition and vector specificity, although with biases different from those seen in the Flavivirus genus.

Amino Acids↗

The case for an error minimizing standard genetic code.

Since discovering the pattern by which amino acids are assigned to codons within the standard genetic code, investigators have explored the idea that natural selection placed biochemically similar amino acids near to one another in coding space so as to minimize the impact of mutations and/or mistranslations. The analytical evidence to support this theory has grown in sophistication and strength over the years, and counterclaims questioning its plausibility and quantitative support have yet to transcend some significant weaknesses in their approach. These weaknesses are illustrated here by means of a simple simulation model for adaptive genetic code evolution. There remain ill explored facets of the 'error minimizing' code hypothesis, however, including the mechanism and pathway by which an adaptive pattern of codon assignments emerged, the extent to which natural selection created synonym redundancy, its role in shaping the amino acid and nucleotide languages, and even the correct interpretation of the adaptive codon assignment pattern: these represent fertile areas for future research.

Amino Acids↗

A simple model based on mutation and selection explains trends in codon and amino-acid usage and GC composition within and across genomes.

BACKGROUND: Correlations between genome composition (in terms of GC content) and usage of particular codons and amino acids have been widely reported, but poorly explained. We show here that a simple model of processes acting at the nucleotide level explains codon usage across a large sample of species (311 bacteria, 28 archaea and 257 eukaryotes). The model quantitatively predicts responses (slope and intercept of the regression line on genome GC content) of individual codons and amino acids to genome composition. RESULTS: Codons respond to genome composition on the basis of their GC content relative to their synonyms (explaining 71-87% of the variance in response among the different codons, depending on measure). Amino-acid responses are determined by the mean GC content of their codons (explaining 71-79% of the variance). Similar trends hold for genes within a genome. Position-dependent selection for error minimization explains why individual bases respond differently to directional mutation pressure. CONCLUSIONS: Our model suggests that GC content drives codon usage (rather than the converse). It unifies a large body of empirical evidence concerning relationships between GC content and amino-acid or codon usage in disparate systems. The relationship between GC content and codon and amino-acid usage is ahistorical; it is replicated independently in the three domains of living organisms, reinforcing the idea that genes and genomes at mutation/selection equilibrium reproduce a unique relationship between nucleic acid and protein composition. Thus, the model may be useful in predicting amino-acid or nucleotide sequences in poorly characterized taxa.

Amino Acids↗

Hill-Robertson interference is a minor determinant of variations in codon bias across Drosophila melanogaster and Caenorhabditis elegans genomes.

According to population genetics models, genomic regions with lower crossing-over rates are expected to experience less effective selection because of Hill-Robertson interference (HRi). The effect of genetic linkage is thought to be particularly important for a selection of weak intensity such as selection affecting codon usage. Consistent with this model, codon bias correlates positively with recombination rate in Drosophila melanogaster and Caenorhabditis elegans. However, in these species, the G+C content of both noncoding DNA and synonymous sites correlates positively with recombination, which suggests that mutation patterns and recombination are associated. To remove this effect of mutation patterns on codon bias, we used the synonymous sites of lowly expressed genes that are expected to be effectively neutral sites. We measured the differences between codon biases of highly expressed genes and their lowly expressed neighbors. In D. melanogaster we find that HRi weakly reduces selection on codon usage of genes located in regions of very low recombination; but these genes only comprise 4% of the total. In C. elegans we do not find any evidence for the effect of recombination on selection for codon bias. Computer simulations indicate that HRi poorly enhances codon bias if the local recombination rate is greater than the mutation rate. This prediction of the model is consistent with our data and with the current estimate of the mutation rate in D. melanogaster. The case of C. elegans, which is highly self-fertilizing, is discussed. Our results suggest that HRi is a minor determinant of variations in codon bias across the genome.

Animals↗

Genetic variation in clinical varicella-zoster virus isolates collected in Ireland between 2002 and 2003.

Analysis of genetic variation in 16 varicella-zoster virus (VZV) isolates selected at random and circulating in the Irish population between March 2002 and February 2003 was carried out. A 919 bp fragment of the glycoprotein E gene (open reading frame 68) encompassing codon 150, at which a non-synonymous mutation defines the escape mutant VZV-MSP, and including two other epitope regions e1 and c1, was sequenced. No new single nucleotide polymorphisms (SNPs) were detected, indicating stability of these epitopes in clinical isolates of VZV. However, when four informative polymorphic markers consisting of defined regions from genes 1, 21, 50, and 54 were sequenced 14 variable nucleotide positions were identified. Phylogenetic analysis showed the presence of three highly supported clades A, B, and C circulating in the Irish population. Approximately one third (6/16; 37.5%) of the Irish VZV isolates in this study belonged to genotype C, 4/16 (25%) to genotype A, and 4/16 (25%) to genotype B. A smaller number 2/16 (12.5%) belonged to genotype J1. This indicates remarkable heterogeneity in the Irish population given the small sample size. No evidence was found to suggest any of the 16 isolates was a recombinant. These findings have implications for the model of geographic isolation of VZV clades to certain regions as the circulating Irish VZV population appears to comprise approximately equal numbers of each of the main genotypes. This data is inconsistent with a model of strict geographical separation of VZV genotypes and suggests that VZV diversity is more pronounced in certain areas than had been thought previously.

Adult↗

Evidence for horizontal transfer from Streptococcus to Escherichia coli of the kfiD gene encoding the K5-specific UDP-glucose dehydrogenase.

Capsular polysaccharides are important virulence factors both in Gram-positive and Gram-negative bacteria. A similar cluster organization of the genes involved in the synthesis of bacterial exopolysaccharides has been postulated in both cases, suggesting that these clusters evolved by module assembly. Horizontal gene transfer has been postulated to explain the polymorphism found in these cellular polymers. The cap1 K and cap3A genes coding for the pneumococcal type 1 and type 3 UDP-glucose dehydrogenases, respectively, have been compared with other UDP-sugar dehydrogenases. We have observed that the evolutionary distance between Cap1K and Cap3A is approximately equal to that found between Cap1K (or Cap3A) and other UDP-GlcDH of families evolutionarily distant like KfiD, the dehydrogenase from Escherichia coli K5. On the basis of comparisons of G + C content, patterns of synonymous and nonsynonymous substitutions, dinucleotide frequencies, and codon usage bias, we conclude that the kfiD gene has been introduced into E. coli from an exogenous source, probably from a streptococcal species.

Bacterial Capsules↗

Relationships between genomic base content and distribution of mass in coded proteins.

The aim of this research was to examine the possible significance of genome/protein relationships in terms of effects on distribution of mass, especially in proteins. Amino acid residues in proteins have side-chains and polypeptide segments. We use "SCM" (side-chain mass), "MCM" (main-chain mass), and "deltaM" (SCM-MCM) as the deviation from "mass balance." Total MCM of the 61 amino acids in the standard code, 3412, equals total SCM: they form a mass balanced set (mean deltaM = 0). Of 14 natural variants of the code, seven have slightly positive mean deltaM values and seven have slightly negative values. Codes with the standard amino acids assigned randomly to the 20 codon sets of the standard code have about one chance in 3,300 of producing a mass balanced set. In natural proteins, as %A + T increases, the proportion of the mass in the side-chains also increases, by about half the amount calculated for standard genes with various AT/GC ratios, partly due to selection of codons with greater variability in composition at synonymous sites. For 203 representative species (including organelles), the total protein mass is distributed approximately equally between SCM and MCM (overall mean deltaM/amino acid residue, -0.06). The attainment of some overall macromolecular mass balance may have been a criterion for selecting the codon/amino acid pairs. When both structural and dynamic requirements are considered, a genetic code based on hydrophobicity and mass balance as key properties seems likely.

AT Rich Sequence↗

A sliding window-based method to detect selective constraints in protein-coding genes and its application to RNA viruses.

Here we present a new sliding window-based method specially designed to detect selective constraints in specific regions of a multiple protein-coding sequence alignment. In contrast to previous window-based procedures, our method is based on a nonarbitrary statistical approach to find the appropriate codon-window size to test deviations of synonymous (d(S)) and nonsynonymous (d(N)) nucleotide substitutions from the expectation. The probabilities of d(N) and d(S) are obtained from simulated data and used to detect significant deviations of d(N) and d(S) in a specific window region of the real sequence alignment. The nonsynonymous-to-synonymous rate ratio (w = d(N)/d(S)) was used to highlight selective constraints in any window wherein d(S) or d(N) was significantly different from the expectation. In these significant windows, w and its variance [V(w)] were calculated and used to test the neutral hypothesis. Computer simulations showed that the method is accurate even for highly divergent sequences. The main advantages of the new method are that it (i) uses a statistically appropriate window size to detect different selective patterns, (ii) is computationally less intensive than maximum likelihood methods, and (iii) detects saturation of synonymous sites, which can give deviations from neutrality. Hence, it allows the analysis of highly divergent sequences and the test of different alternative hypothesis as well. The application of the method to different human immunodeficiency virus type 1 and to foot-and-mouth disease virus genes confirms the action of positive selection on previously described regions as well as on new regions.

Base Sequence↗

Detection of single nucleotide polymorphisms in 24 kDa dimeric alpha-amylase inhibitors from cultivated wheat and its diploid putative progenitors.

Seventeen new genes encoding 24 kDa family dimeric alpha-amylase inhibitors had been characterized from cultivated wheat and its diploid putative progenitors. And the different alpha-amylase inhibitors in this family, which were determined by coding regions single nucleotide polymorphisms (cSNPs) of their genes, were investigated. The amino acid sequences of 24 kDa alpha-amylase inhibitors shared very high coherence (91.2%). It indicated that the dimeric alpha-amylase inhibitors in the 24 kDa family were derived from common ancestral genes by phylogenetic analysis. Eight alpha-amylase inhibitor genes were characterized from one hexaploid wheat variety, and clustered into four subgroups, indicating that the 24 kDa dimeric alpha-amylase inhibitors in cultivated wheat were encoded by multi-gene. Forty-five cSNPs, including 35 transitions and 10 transversions, were found, and resulted in a total of ten amino acid changes. The cSNPs at the first site of a codon cause much more nonsynonymous (92.9%) than synonymous mutations, while nonsynonymous and synonymous mutations were almost equal when the cSNPs were at the third site. It was observed that there was Ile105 instead of Val105 at the active region Val104-Val105-Asp106-Ala107 of the alpha-amylase inhibitor by cSNPs in some inhibitors from Aegilops speltoides, diploid and hexaploid wheats.

Amino Acid Sequence↗

Single strand conformational polymorphism analysis of human CD1 genes in different ethnic groups.

CD1 molecules are able to present unusual antigens, lipids or glycolipids from mycobacterium cell walls to T lymphocytes. Previous studies have suggested that polymorphism of these genes is very limited, in contrast with classical major histocompatibility complex (MHC) antigen-presenting molecules. Our aim was to study possible allelic variations of exons 2 and 3, encoding for the alpha1 and alpha2 domains, respectively, of human CD1A, -B, -C and -D genes. We analyzed genomic samples of unrelated, healthy individuals from different ethnic background: 70 Caucasians from Europe, 33 Black Africans (13 from Tanzania and 20 Zulus), 19 Caucasians from the Sahara and 44 Asian individuals. We have found CD1A to be a biallelic locus with a common allele which was present in the majority of the individuals studied. The second allele differed from the common one by a single-point mutation, resulting in a change of Cys to Trp at position 52 in the alpha1 domain. This second allele was found in heterozygosis in 7 out of 70 Caucasians from Europe (allelic frequencies P=0.95 and q=0.05). In the Chinese population, we found the second allele present in heterozygosis in 19 from the 44 individuals studied, and we also found 6 homozygous individuals for the second allele (allelic frequencies P=0.64 and q=0.35). In addition, we detected a synonymous mutation (C to T transition) in codon 34 of CD1C exon 2 in 4 out of 20 Zulus and in 2 of the 13 Blacks from Tanzania.

Africa↗

Long-term reinfection of the human genome by endogenous retroviruses.

Endogenous retrovirus (ERV) families are derived from their exogenous counterparts by means of a process of germ-line infection and proliferation within the host genome. Several families in the human and mouse genomes now consist of many hundreds of elements and, although several candidates have been proposed, the mechanism behind this proliferation has remained uncertain. To investigate this mechanism, we reconstructed the ratio of nonsynonymous to synonymous changes and the acquisition of stop codons during the evolution of the human ERV family HERV-K(HML2). We show that all genes, including the env gene, which is necessary only for movement between cells, have been under continuous purifying selection. This finding strongly suggests that the proliferation of this family has been almost entirely due to germ-line reinfection, rather than retrotransposition in cis or complementation in trans, and that an infectious pool of endogenous retroviruses has persisted within the primate lineage throughout the past 30 million years. Because many elements within this pool would have been unfixed, it is possible that the HERV-K(HML2) family still contains infectious elements at present, despite their apparent absence in the human genome sequence. Analysis of the env gene of eight other HERV families indicated that reinfection is likely to be the most common mechanism by which endogenous retroviruses proliferate in their hosts.

Endogenous Retroviruses↗

The rate with which spontaneous mutation alters the electrophoretic mobility of polypeptides.

Studies of a Japanese population, involving a total of 539,170 locus tests distributed over 36 polypeptides, yielded three presumptive spontaneous mutations altering the electrophoretic mobility of the polypeptide. This corresponds to a mutation rate of 0.6 X 10(-5) per locus per generation. The a priori probability that undetected discrepancies between legal and biological parentage might in our test system result in an apparent electrophoretic mutation in this population is calculated to be only 0.3 X 10(-7) per locus per generation. Since electrophoresis only detects about half of the amino acid substitutions due to mutations of nucleotides, the corrected rate for mutations causing amino acid substitutions in polypeptides is 1.2 X 10(-5) per locus per generation. With allowance for synonymous mutations and those resulting in stop codons, the total mutation rate for nucleotide changes in the exons encoding a polypeptide becomes approximately equal to 1.8 X 10(-5) per locus per generation. When the present observations are combined with all of the other available data concerning mutation resulting in electrophoretic variants, the electrophoretic rate drops to 0.3 X 10(-5) per locus per generation, the total locus rate drops to roughly 1.0 X 10(-5), and the nucleotide rate drops to 1 X 10(-8). Even with this lower estimate, given approximately equal to 2 X 10(9) nucleotides in the haploid genome and an average of 10(3) exon nucleotides per polypeptide encoded, the implication, if these exon rates can be generalized, is of approximately equal to 20 nucleotide mutations per gamete per generation. This estimate of the frequency of "point" mutations does not include small duplications, rearrangements, or deletions resulting from unequal crossing-over, transcription errors, etc.

Child↗

Silent mutation in the V3 region characteristic of HIV type 1 env subtype B strains from injecting drug users in the former Soviet Union.

New independent states of the former Soviet Union are facing a rapidly growing epidemic of HIV-1 among injecting drug users (IDUs). This epidemic is caused by three HIV-1 populations, one belonging to HIV-1 subtype A (IDU-A), another to subtype B (IDU-B), and the third being a recombinant of the IDU-A and IDU-B viruses (IDU-A/B, gagA/envB). Each of these populations is characterized by a high level of genetic homogeneity. We identified a unique synonymous nucleotide substitution in the first isoleucine codon at the IHIGPGR motif (ATT), which was observed in the env subtype B V3 sequences derived from IDUs in Russia and the Ukraine. This substitution was observed in none of 179 sequences obtained from IDUs in western Europe, northern America, and Asia. Molecular epidemiological analysis of HIV-1 strains based on this sequence pattern could be useful for tracing the origin and spread of the IDU-B viruses to other countries and risk groups.

Base Sequence↗

Evidence for HIV type 1 strains of U.S. intravenous drug users as founders of AIDS epidemic among intravenous drug users in northern Europe.

To establish an epidemiological link between HIV-1 epidemics in U.S. and European homosexual men and intravenous drug users (IVDUs) we analyzed the HIV-1 gp120 V3 sequences in both risk groups. Signature pattern analysis revealed that the V3 sequences of viruses from IVDUs in Northern Europe are distinguishable from those of homosexual men on the basis of one amino acid and two synonymous nucleotide substitutions, which the most conserved was a synonymous nucleotide substitution in the second glycine codon at the tip of the gp120 V3 loop (GGC). This substitution was seen in 17 of 20 (85%) viruses of IVDUs in Northern Europe, in none of 41 homosexual men in either Europe or the United States, and in 5 of 11 (45%) U.S. IVDUs sequences analyzed. Subsequent phylogenetic and multivariate principal coordinate (PCOORD) analyses showed that 16 of 20 (80%) of the Northern European IVDU sequences clustered together with the 5 U.S. IVDU sequences carrying the GGC substitution and away from the sequences of homosexual men from either Europe or the United States. Taken together with the higher level of heterogeneity of U.S. IVDU sequences compared to the Dutch IVDU sequences taken at the same time, these data present suggestive evidence for a U.S. instead of a European origin of the AIDS epidemic among Northern European IVDUs.

Acquired Immunodeficiency Syndrome↗

A family of genes clustered at the Triplo-lethal locus of Drosophila melanogaster has an unusual evolutionary history and significant synteny with Anopheles gambiae.

Within the unique Triplo-lethal region (Tpl) of the Drosophila melanogaster genome we have found a cluster of 20 genes encoding a novel family of proteins. This family is also present in the Anopheles gambiae genome and displays remarkable synteny and sequence conservation with the Drosophila cluster. The family is also present in the sequenced genome of D. pseudoobscura, and homologs have been found in Aedes aegypti mosquitoes and in four other insect orders, but it is not present in the sequenced genome of any noninsect species. Phylogenetic analysis suggests that the cluster evolved prior to the divergence of Drosophila and Anopheles (250 MYA) and has been highly conserved since. The ratio of synonymous to nonsynonymous substitutions and the high codon bias suggest that there has been selection on this family both for expression level and function. We hypothesize that this gene family is Tpl, name it the Osiris family, and consider possible functions. We also predict that this family of proteins, due to the unique dosage sensitivity and the lack of homologs in noninsect species, would be a good target for genetic engineering or novel insecticides.

Amino Acid Sequence↗

Evolution of MHC class II beta chain-encoding genes in the Lake Tana barbel species flock (Barbus intermedius complex).

Major histocompatibility complex (MHC) class II protein polymorphism is maintained in allelic lineages which evolve in a trans-specific manner, passing from one species to descendant species. Selection pressure on peptide binding residues should be greatest during speciation, when organisms move into new environments and their MHC molecules encounter new pathogens. The isolation of MHC genes from teleost fishes, the most diverse group of vertebrates, has created possibilities for testing this hypothesis. The large barbels of Lake Tana have undergone an adaptive radiation within the last 5 million years, producing 14 morphotypes which inhabit different ecological niches within the lake. We studied the variability in class II beta chain-encoding genes of four of these morphotypes using polymerase chain reaction amplification and DNA sequencing. The sequences obtained were orthologous to four of the known class II genes from the common carp, from which barbels diverged approximately 32 million years ago. When subjected to phylogenetic analysis, the 48 sequences clustered into groups which represent allelic lineages. A comparison of nonsynonymous and synonymous substitutions between the peptide binding region codons and non-peptide binding region codons of these sequences revealed that they are under strong selective pressure.

Amino Acid Sequence↗