Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “nucleotide polymorphism patterns”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Rhizobial 16S rRNA and dnaK genes: mosaicism and the uncertain phylogenetic placement of Rhizobium galegae.

The phylogenetic relatedness among 12 agriculturally important species in the order Rhizobiales was estimated by comparative 16S rRNA and dnaK sequence analyses. Two groups of related species were identified by neighbor-joining and maximum-parsimony analysis. One group consisted of Mesorhizobium loti and Mesorhizobium ciceri, and the other group consisted of Agrobacterium rhizogenes, Rhizobium tropici, Rhizobium etli, and Rhizobium leguminosarum. Although bootstrap support for the placement of the remaining six species varied, A. tumefaciens, Agrobacterium rubi, and Agrobacterium vitis were consistently associated in the same subcluster. The three other species included Rhizobium galegae, Sinorhizobium meliloti, and Brucella ovis. Among these, the placement of R. galegae was the least consistent, in that it was placed flanking the A. rhizogenes-Rhizobium cluster in the dnaK nucleotide sequence trees, while it was placed with the other three Agrobacterium species in the 16S rRNA and the DnaK amino acid trees. In an effort to explain the inconsistent placement of R. galegae, we examined polymorphic site distribution patterns among the various species. Localized runs of nucleotide sequence similarity were evident between R. galegae and certain other species, suggesting that the R. galegae genes are chimeric. These results provide a tenable explanation for the weak statistical support often associated with the phylogenetic placement of R. galegae, and they also illustrate a potential pitfall in the use of partial sequences for species identification.

Alleles↗

CLUSTAG: hierarchical clustering and graph methods for selecting tag SNPs.

UNLABELLED: Cluster and set-cover algorithms are developed to obtain a set of tag single nucleotide polymorphisms (SNPs) that can represent all the known SNPs in a chromosomal region, subject to the constraint that all SNPs must have a squared correlation R2>C with at least one tag SNP, where C is specified by the user. AVAILABILITY: http://hkumath.hku.hk/web/link/CLUSTAG/CLUSTAG.html CONTACT: mng@maths.hku.hk.

Algorithms↗

Regulation of intracellular localization of human MTH1, OGG1, and MYH proteins for repair of oxidative DNA damage.

In mammalian cells, more than one genome has to be maintained throughout the entire life of the cell, one in the nucleus and the other in mitochondria. It seems likely that the genomes in mitochondria are highly exposed to reactive oxygen species (ROS) as a result of their respiratory function. Human MTH1 (hMTH1) protein hydrolyzes oxidized purine nucleoside triphosphates, such as 8-oxo-dGTP, 8-oxo-dATP, and 2-hydroxy (OH)-dATP, thus suggesting that these oxidized nucleotides are deleterious for cells. Here, we report that a single-nucleotide polymorphism (SNP) in the human MTH1 gene alters splicing patterns of hMTH1 transcripts, and that a novel hMTH1 polypeptide with an additional mitochondrial targeting signal is produced from the altered hMTH1 mRNAs; thus, intracellular location of hMTH1 is likely to be affected by a SNP. These observations strongly suggest that errors caused by oxidized nucleotides in mitochondria have to be avoided in order to maintain the mitochondrial genome, as well as the nuclear genome, in human cells. Based on these observations, we further characterized expression and intracellular localization of 8-oxoG DNA glycosylase (hOGG1) and 2-OH-A/adenine DNA glycosylase (hMYH) in human cells. These two enzymes initiate base excision repair reactions for oxidized bases in DNA generated by direct oxidation of DNA or by incorporation of oxidized nucleotides. We describe the detection of the authentic hOGG1 and hMYH proteins in mitochondria, as well as nuclei in human cells, and how their intracellular localization is regulated by alternative splicing of each transcript.

Adaptor Proteins, Signal Transducing↗

Sequence alterations in the reduced folate carrier are observed in osteosarcoma tumor samples.

High-dose methotrexate is a standard component of therapy for high-grade osteosarcoma. Its effectiveness may be limited by intrinsic and acquired resistance. Decreased reduced folate carrier (RFC) expression has been shown in approximately half of osteosarcomas at diagnosis. Mutations and polymorphisms in the RFC gene have been reported in various cell lines. The purpose of this study was to investigate sequence alterations in the RFC gene in osteosarcoma tumor samples. The entire coding region of the RFC gene in samples from 162 osteosarcoma patients was screened by DNA single-stranded conformational polymorphism, followed by direct sequencing of any region with altered mobility. A previously identified polymorphism at cDNA position number 174 of RFC exon 2 was observed. Sixty-one samples (37.6%) were heterozygous with both A/G at this position (His(27)/Arg(27)), 52 samples (32.2%) were homozygous with G (Arg(27)), and 49 samples (30.2%) were homozygous with A (His(27)). Fifteen (9.2%) samples were identified with other RFC sequence variants in exon 2, none of which have been reported. The sequence variants in exon 2 included a G to A substitution at cDNA position 231, a G to A substitution at cDNA position 155, a C to T substitution at cDNA position 114, and a T to C substitution at cDNA position 104, resulting in a serine to asparagine substitution at amino acid 46, a glutamate to lysine substitution at amino acid 21, an alanine to valine substitution at amino acid 7, and a serine to proline substitution at amino acid 4, respectively. A deletion of A at cDNA position 126 resulting in a frameshift was also observed. Some of these variants were observed in multiple samples. Eight samples had altered single-stranded conformational polymorphism patterns in exon 3 that were associated with nucleotide changes that altered the amino acid sequence. All of these RFC sequence variants appeared to be heterozygous. Heterozygous C/T and homozygous C also were observed at RFC cDNA position 790 in exon 3, which does not alter the amino acid coding sequence. This study shows that RFC sequence alterations are frequent in samples from osteosarcoma patients. Additional studies are under way to determine the clinical significance of these sequence alterations and their effect on methotrexate transport and resistance.

Amino Acid Sequence↗

Genotyping of single nucleotide polymorphism using model-based clustering.

MOTIVATION: Single nucleotide polymorphisms have been investigated as biological markers and the representative high-throughput genotyping method is a combination of the Invader assay and a statistical clustering method. A typical statistical clustering method is the k-means method, but it often fails because of the lack of flexibility. An alternative fast and reliable method is therefore desirable. RESULTS: This paper proposes a model-based clustering method using a normal mixture model and a well-conceived penalized likelihood. The proposed method can judge unclear genotypings to be re-examined and also work well even when the number of clusters is unknown. Some results are illustrated and then satisfactory genotypings are shown. Even when the conventional maximum likelihood method and the typical k-means clustering method failed, the proposed method succeeded.

Algorithms↗

Choosing SNPs using feature selection.

A major challenge for genomewide disease association studies is the high cost of genotyping large number of single nucleotide polymorphisms (SNPs). The correlations between SNPs, however, make it possible to select a parsimonious set of informative SNPs, known as "tagging" SNPs, able to capture most variation in a population. Considerable research interest has recently focused on the development of methods for finding such SNPs. In this paper, we present an efficient method for finding tagging SNPs. The method does not involve computation-intensive search for SNP subsets but discards redundant SNPs using a feature selection algorithm. In contrast to most existing methods, the method presented here does not limit itself to using only correlations between SNPs in local groups. By using correlations that occur across different chromosomal regions, the method can reduce the number of globally redundant SNPs. Experimental results show that the number of tagging SNPs selected by our method is smaller than by using block-based methods. Supplementary website: http://htsnp.stanford.edu/FSFS/.

Algorithms↗

A comparative study of machine-learning methods to predict the effects of single nucleotide polymorphisms on protein function.

MOTIVATION: The large volume of single nucleotide polymorphism data now available motivates the development of methods for distinguishing neutral changes from those which have real biological effects. Here, two different machine-learning methods, decision trees and support vector machines (SVMs), are applied for the first time to this problem. In common with most other methods, only non-synonymous changes in protein coding regions of the genome are considered. RESULTS: In detailed cross-validation analysis, both learning methods are shown to compete well with existing methods, and to out-perform them in some key tests. SVMs show better generalization performance, but decision trees have the advantage of generating interpretable rules with robust estimates of prediction confidence. It is shown that the inclusion of protein structure information produces more accurate methods, in agreement with other recent studies, and the effect of using predicted rather than actual structure is evaluated. AVAILABILITY: Software is available on request from the authors.

Algorithms↗

Choosing SNPs using feature selection.

A major challenge for genomewide disease association studies is the high cost of genotyping large number of single nucleotide polymorphisms (SNP). The correlations between SNPs, however, make it possible to select a parsimonious set of informative SNPs, known as "tagging" SNPs, able to capture most variation in a population. Considerable research interest has recently focused on the development of methods for finding such SNPs. In this paper, we present an efficient method for finding tagging SNPs. The method does not involve computation-intensive search for SNP subsets but discards redundant SNPs using a feature selection algorithm. In contrast to most existing methods, the method presented here does not limit itself to using only correlations between SNPs in local groups. By using correlations that occur across different chromosomal regions, the method can reduce the number of globally redundant SNPs. Experimental results show that the number of tagging SNPs selected by our method is smaller than by using block-based methods.

Artificial Intelligence↗

MACGT: multi-dimensional automated clustering genotyping tool for analysis of microarray-based mini-sequencing data.

SUMMARY: Multi-dimensional Automated Clustering Genotyping Tool (MACGT) is a Java application that clusters complex multi-dimensional vector data derived from single nucleotide polymorphism (SNP) genotyping experiments using mini-sequencing based microarray chemistries such as arrayed primer extension (APEX). Spot intensity output files from microarray experiments across multiple samples are imported into MACGT. The datasets can include four channels of intensity data for each spot, replica spots for each SNP probe and multiple probe types (APEX and allele-specific APEX probes) on both DNA strands for each SNP. MACGT automatically clusters these multi-dimensionality datasets for each SNP across multiple samples. Incorporation of additional array datasets from known samples that have previously validated SNP genotype calls allows unknown samples to be automatically assigned a genotype based on the clustering, along with numerical measures of confidence for each genotype call. Calling accuracy by MACGT exceeds 98% when applied to genotyping data from APEX microarrays, and can be increased to >99.5% by applying thresholds to the confidence measures.

Algorithms↗

Performance in the Wisconsin Card Sorting Test and the C957T polymorphism of the DRD2 gene in healthy volunteers.

INTRODUCTION: Previous studies have associated a decreased striatal D2 dopamine receptor (DRD2) binding with impaired performance in cognitive tasks. In vivo studies have found a lower DRD2 binding associated with the CC genotype of the C957T single nucleotide polymorphism (SNP) of the DRD2 gene. OBJECTIVE: The aim of this study was to investigate the relationship between executive functions and the C957T DRD2 SNP. We hypothesized that the CC genotype would be associated with a poorer executive functioning. METHODS: Our sample consisted of 83 healthy volunteers (28 males and 55 females; mean age 25.2, SD 1.7 years). To assess executive functions, the Wisconsin Card Sorting Test was used, considering the variables perseverative errors, perseverative responses, and number of categories achieved. The genotype distribution was 13 CC, 41 CT, and 29 TT, satisfying Hardy-Weinberg equilibrium. RESULTS: Carriers of the CC genotype, compared with carriers of the CT/TT genotypes, achieved significantly fewer categories (5.00 vs. 5.81; p = 0.004), made a greater number of perseverative errors (13.46 vs. 8.39; p = 0.018), and had a greater number of perseverative responses (14.92 vs. 8.94; p = 0.014). CONCLUSIONS: Our results support the hypothesis that the C957T DRD2 SNP may influence cognitive performance through its repercussions on central dopaminergic function.

Adult↗

Automatic scoring and quality assessment using accuracy bounds for FP-TDI SNP genotyping data.

BACKGROUND: Human diversity, namely single nucleotide polymorphisms (SNPs), is becoming a focus of biomedical research. Despite the binary nature of SNP determination, the majority of genotyping assay data need a critical evaluation for genotype calling. We applied statistical models to improve the automated analysis of 2-dimensional SNP data. METHODS: We derived several quantities in the framework of Gaussian mixture models that provide figures of merit to objectively measure the data quality. The accuracy of individual observations is scored as the probability of belonging to a certain genotype cluster, while the assay quality is measured by the overlap between the genotype clusters. RESULTS: The approach was extensively tested with a dataset of 438 nonredundant SNP assays comprising >150,000 datapoints. The performance of our automatic scoring method was compared with manual assignments. The agreement for the overall assay quality is remarkably good, and individual observations were scored differently by man and machine in 2.6% of cases, when applying stringent probability threshold values. CONCLUSION: Our definition of bounds for the accuracy for complete assays in terms of misclassification probabilities goes beyond other proposed analysis methods. We expect the scoring method to minimise human intervention and provide a more objective error estimate in genotype calling.

Algorithms↗

Coverage and characteristics of the Affymetrix GeneChip Human Mapping 100K SNP set.

Improvements in technology have made it possible to conduct genome-wide association mapping at costs within reach of academic investigators, and experiments are currently being conducted with a variety of high-throughput platforms. To provide an appropriate context for interpreting results of such studies, we summarize here results of an investigation of one of the first of these technologies to be publicly available, the Affymetrix GeneChip Human Mapping 100K set of single nucleotide polymorphisms (SNPs). In a systematic analysis of the pattern and distribution of SNPs in the Mapping 100K set, we find that SNPs in this set are undersampled from coding regions (both nonsynonymous and synonymous) and oversampled from regions outside genes, relative to SNPs in the overall HapMap database. In addition, we utilize a novel multilocus linkage disequilibrium (LD) coefficient based on information content (analogous to the information content scores commonly used for linkage mapping) that is equivalent to the familiar measure r2 in the special case of two loci. Using this approach, we are able to summarize for any subset of markers, such as the Affymetrix Mapping 100K set, the information available for association mapping in that subset, relative to the information available in the full set of markers included in the HapMap, and highlight circumstances in which this multilocus measure of LD provides substantial additional insight about the haplotype structure in a region over pairwise measures of LD.

Chromosome Mapping↗

BNTagger: improved tagging SNP selection using Bayesian networks.

Genetic variation analysis holds much promise as a basis for disease-gene association. However, due to the tremendous number of candidate single nucleotide polymorphisms (SNPs), there is a clear need to expedite genotyping by selecting and considering only a subset of all SNPs. This process is known as tagging SNP selection. Several methods for tagging SNP selection have been proposed, and have shown promising results. However, most of them rely on strong assumptions such as prior block-partitioning, bi-allelic SNPs, or a fixed number or location of tagging SNPs. We introduce BNTagger, a new method for tagging SNP selection, based on conditional independence among SNPs. Using the formalism of Bayesian networks (BNs), our system aims to select a subset of independent and highly predictive SNPs. Similar to previous prediction-based methods, we aim to maximize the prediction accuracy of tagging SNPs, but unlike them, we neither fix the number nor the location of predictive tagging SNPs, nor require SNPs to be bi-allelic. In addition, for newly-genotyped samples, BNTagger directly uses genotype data as input, while producing as output haplotype data of all SNPs. Using three public data sets, we compare the prediction performance of our method to that of three state-of-the-art tagging SNP selection methods. The results demonstrate that our method consistently improves upon previous methods in terms of prediction accuracy. Moreover, our method retains its good performance even when a very small number of tagging SNPs are used.

Algorithms↗

Single nucleotide polymorphisms in exons of the apo(a) kringles IV types 6 to 10 domain affect Lp(a) plasma concentrations and have different patterns in Africans and Caucasians.

Lipoprotein(a) [Lp(a)] is a complex of apolipoprotein(a) [apo(a)] and low-density lipoprotein which is associated with atherothrombotic disease. Lp(a) plasma levels are controlled to a large extent by the apo(a) gene locus. Known polymorphisms in the apo(a) gene, including the kringle (K) IV-2 variable number of tandem repeats, explain only part of the large interindividual variability and do not explain the differences in Lp(a) concentrations between major human ethnic groups. Here we performed screening for single nucleotide polymorphisms (SNPs) in exons and flanking intron sequences of the apo(a) K IV types 6, 8, 9 and 10 which represent 1.3 kb of coding sequence in two African (Khoi San, Black South Africans) and one Caucasian (Tyroleans) populations and investigated whether they affect Lp(a) levels. Together, 768 alleles were analyzed. We identified 14 SNPs, including 11 non-synonymous SNPs (eight of which involved conserved residues), one splice site and two synonymous base changes. No sequence variants common to Africans and Caucasians were found. Several of the newly identified SNPs showed significant effects on Lp(a) plasma concentrations. The substitutions S37F in K IV-6 and G17R in K IV-8 were associated with Lp(a) levels significantly below average in Africans. In contrast, the R18W substitution in K IV-9, which occurred with a frequency of 8% in Khoi San, resulted in a significantly increased Lp(a) concentration. Together, our data suggest that several SNPs in the coding sequence of apo(a) affect Lp(a) levels. This indicates that many SNPs may have subtle effects on the gene product.

Alleles↗

N-methyl-D-aspartate receptor NR1 subunit gene (GRIN1) in schizophrenia: TDT and case-control analyses.

The N-methyl-d-aspartate glutamate receptors (NMDAR) act in the CNS as regulators of the release of neurotransmitters such as dopamine, noradrenaline, acetylcholine, and GABA. It has been suggested that a weakened glutamatergic tone increases the risk of sensory overload and of exaggerated responses in the monoaminergic system, which is consistent with the symptomatology of schizophrenia. We studied two silent polymorphisms in GRIN1. GRIN1/1 is a G/C substitution localized on the 5' untranslated region; GRIN1/10 is an A/G substitution localized in exon 6 of GRIN1. Minor allele frequencies in our sample were calculated to be 0.05 and 0.2 respectively. We genotyped 86 nuclear families and 91 ethnically matched case-control pairs. Both samples were collected from the Toronto area. We tested the hypothesis that GRIN1 polymorphisms were associated with schizophrenia using the transmission disequilibrium test (TDT) and comparing allele frequencies between cases and controls. The results are as follows: GRIN1/1: chi(2) = 2.19, P = 0.14; GRIN1/10: chi(2) = 1.5, P = 0.22. For the case-control sample: GRIN1/1: chi(2) = 0.013, P = 0.908; GRIN1/10: chi(2) = 0.544, P = 0.461. No significant results were obtained. Haplotype analyses showed a borderline significant result for the 2,1 haplotype (chi(2) = 3.86, P-value = 0.049). An analysis of variance (ANOVA) to evaluate the association between genetic makeup and age at onset was performed, with no significant results: GRIN1/1, F[df = 2] = 0.42, P-value = 0.659; GRIN1/10, F[df = 2] = 0.16, P-value = 0.853. We are currently collecting additional samples to increase the power of the analyses.

Age of Onset↗

Randomly distributed crossovers may generate block-like patterns of linkage disequilibrium: an act of genetic drift.

There is considerable interest in identifying and characterizing block-like patterns of linkage disequilibrium (LD; haplotype blocks) in the human genome as these may facilitate the identification of complex disease genes via genome-wide association studies. Although recombination hot-spots have been suggested as the primary mechanism to explain the block-like pattern of LD, other forces, such as genetic drift, may also be important. To this end, we have studied the effect of various recombination models on patterns of LD by using extensive simulations. As expected, haplotype blocks were observed under a model allowing recombination hot-spots. However, we also observed similar block-like patterns in the models where recombination crossovers are randomly and uniformly distributed, and we demonstrate that these blocks are generated by genetic drift. We caution that genetic drift may be an alternative mechanism (in addition to recombination hot-spots) that can lead to block-like patterns of LD. Our findings highlight the necessity of characterizing haplotype blocks in world-wide populations.

Computer Simulation↗

Population-based and family-based association studies of an (AC)n dinucleotide repeat in alpha-7 nicotinic receptor subunit gene and schizophrenia.

The human alpha-7 neuronal nicotinic receptor subunit (CHRNA7) gene, located at chromosome 15q13.2, represents a strong candidate gene for schizophrenia. We have examined an (AC)n dinucleotide repeat in intron 2 of the CHRNA7 gene, which was previously shown to be strongly linked with schizophrenia, using both population-based and family-based association studies. In the population-based study, no significant differences between the genotype and allele frequency distributions in schizophrenia patients and control subjects were observed after correction for multiple testing, although a nominally significant association between the most common allele and schizophrenia was observed (P = 0.023, uncorrected for multiple testing). In the family-based study, there is no significant over-transmission (Transmitted/Non-transmitted: 61/50) of the same allele in 160 family trios. Overall, our results do not support a major role for the (AC)n dinucleotide repeat in schizophrenia susceptibility in Han Chinese. Further large-scale genetic studies based on a set of single nucleotide polymorphisms (SNPs) that fully characterize the linkage disequilibrium patterns at the CHRNA7 gene are necessary to determine the relevance of this gene as a risk factor for schizophrenia susceptibility.

Adult↗

Patterns of linkage disequilibrium in the human genome.

Particular alleles at neighbouring loci tend to be co-inherited. For tightly linked loci, this might lead to associations between alleles in the population a property known as linkage disequilibrium (LD). LD has recently become the focus of intense study in the hope that it might facilitate the mapping of complex disease loci through whole-genome association studies. This approach depends crucially on the patterns of LD in the human genome. In this review, we draw on empirical studies in humans and Drosophila, as well as simulation studies, to assess the current state of knowledge about patterns of LD, and consider the implications for the use of LD as a mapping tool.

Animals↗