Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “SNP analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Population genetic analysis of ascertained SNP data.

The large single nucleotide polymorphism (SNP) typing projects have provided an invaluable data resource for human population geneticists. Almost all of the available SNP loci, however, have been identified through a SNP discovery protocol that will influence the allelic distributions in the sampled loci. Standard methods for population genetic analysis based on the available SNP data will, therefore, be biased. This paper discusses the effect of this ascertainment bias on allelic distributions and on methods for quantifying linkage disequilibrium and estimating demographic parameters. Several recently developed methods for correcting for the ascertainment bias will also be discussed.

Bias↗

Genotools SNP manager: a new software for automated high-throughput MALDI-TOF mass spectrometry SNP genotyping.

Analysis of single nucleotide polymorphisms (SNPs) is a rapidly growing field of research that provides insights into the most common type of differences between individual genomes. The resulting information has a strong impact in the fields of pharmacogenomics, drug development, forensic medicine, and diagnostics of specific disease markers. The technique of matrix-assisted laser desorption/ionization time of flight mass spectrometry (MALDI-TOF MS) has been shown to be a highly suitable tool for the analysis of DNA. It supplies a very versatile method for addressing a high-throughput SNP genotyping approach. Here, we present the Bruker genotools SNP MANAGER, a new software tool suitable for highly automated MALDI-TOF MS SNP genotyping. The genotools SNP MANAGER administers the sample preparation data, calculates masses of allele-specific primer extension products, performs genotyping analysis, and displays the results. In the current study, we have used the genotools SNP MANAGER to perform an automated duplex SNP analysis of two biallelic markers from the promoter of the gene encoding the inflammatory mediator interleukin-6.

DNA↗

Analysis of SNP-expression association matrices.

High throughput expression profiling and genotyping technologies provide the means to study the genetic determinants of population variation in gene expression variation. In this paper we present a general statistical framework for the simultaneous analysis of gene expression data and SNP genotype data measured for the same cohort. The framework consists of methods to associate transcripts with SNPs affecting their expression, algorithms to detect subsets of transcripts that share significantly many associations with a subset of SNPs, and methods to visualize the identified relations. We apply our framework to SNP-expression data collected from 50 breast cancer patients. Our results demonstrate an overabundance of transcript-SNP associations in this data, and pinpoint SNPs that are potential master regulators of transcription. We also identify several statistically significant transcript-subsets with common putative regulators that fall into well-defined functional categories.

Algorithms↗

Analysis of SNP-expression association matrices.

High throughput expression profiling and genotyping technologies provide the means to study the genetic determinants of population variation in gene expression variation. In this paper we present a general statistical framework for the simultaneous analysis of gene expression data and SNP genotype data measured for the same cohort. The framework consists of methods to associate transcripts with SNPs affecting their expression, algorithms to detect subsets of transcripts that share significantly many associations with a subset of SNPs, and methods to visualize the identified relations. We apply our framework to SNP-expression data collected from 49 breast cancer patients. Our results demonstrate an overabundance of transcript-SNP associations in this data, and pinpoint SNPs that are potential master regulators of transcription. We also identify several statistically significant transcript-subsets with common putative regulators that fall into well-defined functional categories.

Algorithms↗

A comprehensive SNP-based genetic analysis of inbred mouse strains.

Dense genetic maps of mammalian genomes facilitate a variety of biological studies including the mapping of polygenic traits, positional cloning of monogenic traits, mapping of quantitative or qualitative trait loci, marker association, allelic imbalance, speed congenic construction, and evolutionary or phylogenetic comparison. In particular, single nucleotide polymorphisms (SNPs) have proved useful because of their abundance and compatibility with multiple high-throughput technology platforms. SNP genotyping is especially suited for the genetic analysis of model organisms such as the mouse because biallelic markers remain fully informative when used to characterize crosses between inbred strains. Here we report the mapping and genotyping of 673 SNPs (including 519 novel SNPs) in 55 of the most commonly used mouse strains. These data have allowed us to construct a phylogenetic tree that correlates and expands known genealogical relationships and clarifies the origin of strains previously having an uncertain ancestry. All 55 inbred strains are distinguishable genetically using this SNP panel. Our data reveal an uneven SNP distribution consistent with a mosaic pattern of inheritance and provide some insight into the changing dynamics of the physical architecture of the genome. Furthermore, these data represent a valuable resource for the selection of markers and the design of experiments that require the genetic distinction of any pair of mouse inbred strains such as the generation of congenic mice, positional cloning, and the mapping of quantitative or qualitative trait loci.

Animals↗

Rapid genetic mapping of ESTs using SNP pyrosequencing and indel analysis.

We describe an effective systematic approach to genetic mapping of cDNA clones, including those obtained from EST sequencing. The EST of interest is first partially sequenced from the 3'-end. PCR primers which bracket the 3'-UTR segment of the cDNA are designed. The corresponding gene segment is amplified from the parents of the mapping population, using primers equipped with 3'- and 5'-extensions to facilitate direct sequencing of PCR products. Comparison of the sequences obtained from the mapping parents frequently reveals single nucleotide polymorphisms or insertion / deletion polymorphisms, which can then be genotyped in a mapping population. The genotyping of SNPs is performed by pyrosequencing, a sequencing-by-synthesis method that has been used successfully in SNP diagnostics. SNP analysis of up to 96 samples, a number required to produce meaningful genetic segregation data, can be rapidly accomplished in parallel. The parental genotype of three loci, stearoyl-ACP desaturase, nucleoside-diphosphate kinase and sucrose synthetase-1 were determined by conventional sequencing, and the polymorphism so identified were scored by the pyrosequencing of 94 individuals of a maize recombinant-inbred population. These loci were successfully placed onto chromosomes 3, 7 and 9 respectively. This method is generally applicable to most plant species, which show sufficient sequence diversity in the 3'-UTR region of genes.

3' Untranslated Regions↗

SNP and haplotype analysis of a novel tryptophan hydroxylase isoform (TPH2) gene provide evidence for association with major depression.

Tryptophan hydroxylase (TPH), being the rate-limiting enzyme in the biosynthesis of serotonin plays a major role as candidate gene in several psychiatric disorders. Recently, a second TPH isoform (TPH2) was identified in mice, which was exclusively present in the brain. In a previous post-mortem study of our own group, we could demonstrate that TPH2 is also expressed in the human brain, but not in peripheral tissues. This is the first report of an association study between polymorphisms in the TPH2 gene and major depression (MD). We performed single-nucleotide polymorphism (SNP), haplotype and linkage disequlibrium studies on 300 depressed patients and 265 healthy controls with 10 SNPs in the TPH2 gene. Significant association was detected between one SNP (P=0.0012, global P=0.0051) and MD. Haplotype analysis produced additional support for association (P<0.0001, global P=0.0001). Our findings provide evidence for an involvement of genetic variants of the TPH2 gene in the pathogenesis of MD and might be a hint on the repeatedly discussed duality of the serotonergic system. These results may open up new research strategies for the analysis of the observed disturbances in the serotonergic system in patients suffering from several other psychiatric disorders.

Adult↗

Genome wide in silico SNP-tumor association analysis.

BACKGROUND: Carcinogenesis occurs, at least in part, due to the accumulation of mutations in critical genes that control the mechanisms of cell proliferation, differentiation and death. Publicly accessible databases contain millions of expressed sequence tag (EST) and single nucleotide polymorphism (SNP) records, which have the potential to assist in the identification of SNPs overrepresented in tumor tissue. METHODS: An in silico SNP-tumor association study was performed utilizing tissue library and SNP information available in NCBI's dbEST (release 092002) and dbSNP (build 106). RESULTS: A total of 4865 SNPs were identified which were present at higher allele frequencies in tumor compared to normal tissues. A subset of 327 (6.7%) SNPs induce amino acid changes to the protein coding sequences. This approach identified several SNPs which have been previously associated with carcinogenesis, as well as a number of SNPs that now warrant further investigation CONCLUSIONS: This novel in silico approach can assist in prioritization of genes and SNPs in the effort to elucidate the genetic mechanisms underlying the development of cancer.

Databases, Genetic↗

Detection and functional analysis of an SNP in the promoter of the human ferritin H gene that modulates the gene expression.

The H ferritin promoter spans approximately 150 bp, upstream of the transcription start and is composed by two cis-elements in position -132 (A box) and -62 (B-box), respectively. The A box is recognized by the transcription factor Sp1, and the B-box by a protein complex called Bbf, which includes the CAAT binding factor NF-Y. In this study we performed a functional analysis of an H ferritin promoter allele carrying a G to T substitution adjacent to the Bbf binding site, in position -69. In vitro studies with reporter constructs revealed a significantly reduced transcriptional activity of this allele compared to that of the w.t. promoter that was mirrored by a decrease in Bbf binding. In vivo, this variant genotype is accompanied by a reduced amount of the H mRNA in peripheral blood lymphocytes.

Alleles↗

SNP identification and analysis in part of intron 2 of goat MSTN gene and variation within and among species.

Part of intron 2 of the myostatin (MSTN) gene of 140 goats from 24 populations and 38 sheep from 8 breeds were sequenced, and similar sequences of different species from Gene bank were also obtained to study MSTN diversity within and among species. The results indicated that there were seven polymorphic sites in the sequenced region of goat, which have not been separated by recombination (or recurrent mutation), presented complete linkage disequilibrium, and could be sorted into three haplotypes. There was no polymorphic site in the sequenced region of sheep. The haplotype diversity, nucleotide diversity, and average number of single nucleotide polymorphism (SNP) differences of goats from the South group are higher than those of North group, and the corresponding value of the Foreign group is also higher than that of Chinese. The genetic differentiation (0.7558) between the Foreign and Chinese group is significant. There are two main haplotypes of the MSTN intron 2 in the goat, which may represent two ancestral types, in support of the theory that domestic goats in the world mainly originated from two ancestors based on morphology, history, archaeology, and molecular markers. The sequence differences of the MSTN intron 2 among species are greater than those within species.

Animals↗

Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.

The vast amount of protein sequence data now available, together with accumulating experimental knowledge of protein function, enables modeling of protein sequence and function evolution. The PANTHER database was designed to model evolutionary sequence-function relationships on a large scale. There are a number of applications for these data, and we have implemented web services that address three of them. The first is a protein classification service. Proteins can be classified, using only their amino acid sequences, to evolutionary groups at both the family and subfamily levels. Specific subfamilies, and often families, are further classified when possible according to their functions, including molecular function and the biological processes and pathways they participate in. The second application, then, is an expression data analysis service, where functional classification information can help find biological patterns in the data obtained from genome-wide experiments. The third application is a coding single-nucleotide polymorphism scoring service. In this case, information about evolutionarily related proteins is used to assess the likelihood of a deleterious effect on protein function arising from a single substitution at a specific amino acid position in the protein. All three web services are available at http://www.pantherdb.org/tools.

Amino Acid Substitution↗

SNPSplicer: systematic analysis of SNP-dependent splicing in genotyped cDNAs.

Functional annotation of SNPs (as generated by HapMap (http://www.hapmap.org) for instance) is a major challenge. SNPs that lead to single amino acid substitutions, stop codons, or frameshift mutations can be readily interpreted, but these represent only a fraction of known SNPs. Many SNPs are located in sequences of splicing relevance-the canonical splice site consensus sequences, exonic and intronic splice enhancers or silencers (exonic splice enhancer [ESE], intronic splice enhancer [ISE], exonic splicing silencer [ESS], and intronic splicing silencer [ISS]), and others. We propose using sets of matching DNA and complementary DNA (cDNA) as a screening method to investigate the potential splice effects of SNPs in RT-PCR experiments with tissue material from genotyped sources. We have developed a software solution (SNPSplicer; http://www.ikmb.uni-kiel.de/snpsplicer) that aids in the rapid interpretation of such screening experiments. The utility of the approach is illustrated for SNPs affecting the donor splice sites (rs2076530:A>G, rs3816989:G>A) leading to the use of a cryptic splice site and exon skipping, respectively, and an exonic splice enhancer SNP (rs2274987:C/T), leading to inclusion of a new exon. We anticipate that this methodology may help in the functional annotation of SNPs in a more high-throughput fashion.

Alternative Splicing↗

ArrayFusion: a web application for multi-dimensional analysis of CGH, SNP and microarray data.

UNLABELLED: ArrayFusion annotates conventional CGH results and various types of microarray data from a range of platforms (cDNA, expression, exon, SNP, array-CGH and ChIP-on-chip) and converts them into standard formats which can be visualized in genome browsers (Affymetrix Integrated Genome Browser and GBrowse in the HapMap Project). Converted files can then be imported simultaneously into a single genome browser to benefit a collective interpretation between different array results. ArrayFusion therefore provides a new type of tool facilitating the integration of CGH and array results to provide new experimental directions. AVAILABILITY: http://microarray.ym.edu.tw/tools/arrayfusion

Algorithms↗

ALOHOMORA: a tool for linkage analysis using 10K SNP array data.

SUMMARY: ALOHOMORA is a software tool designed to facilitate genome-wide linkage studies performed with high-density single nucleotide polymorphism (SNP) marker panels such as the Affymetrix GeneChip(R) Human Mapping 10K Array. Genotype data are converted into appropriate formats for a number of common linkage programs and subjected to standard quality control routines before linkage runs are started. ALOHOMORA is written in Perl and may be used to perform state-of-the-art linkage scans in small and large families with any genetic model. Options for using different genetic maps or ethnicity-specific allele frequencies are implemented. Graphic outputs of whole-genome multipoint LOD score values are provided for the entire dataset as well as for individual families. AVAILABILITY: ALOHOMORA is available free of charge for non-commercial research institutions. For more details, see http://gmc.mdc-berlin.de/alohomora/

Algorithms↗

Direct analysis of unphased SNP genotype data in population-based association studies via Bayesian partition modelling of haplotypes.

We describe a novel method for assessing the strength of disease association with single nucleotide polymorphisms (SNPs) in a candidate gene or small candidate region, and for estimating the corresponding haplotype relative risks of disease, using unphased genotype data directly. We begin by estimating the relative frequencies of haplotypes consistent with observed SNP genotypes. Under the Bayesian partition model, we specify cluster centres from this set of consistent SNP haplotypes. The remaining haplotypes are then assigned to the cluster with the "nearest" centre, where distance is defined in terms of SNP allele matches. Within a logistic regression modelling framework, each haplotype within a cluster is assigned the same disease risk, reducing the number of parameters required. Uncertainty in phase assignment is addressed by considering all possible haplotype configurations consistent with each unphased genotype, weighted in the logistic regression likelihood by their probabilities, calculated according to the estimated relative haplotype frequencies. We develop a Markov chain Monte Carlo algorithm to sample over the space of haplotype clusters and corresponding disease risks, allowing for covariates that might include environmental risk factors or polygenic effects. Application of the algorithm to SNP genotype data in an 890-kb region flanking the CYP2D6 gene illustrates that we can identify clusters of haplotypes with similar risk of poor drug metaboliser (PDM) phenotype, and can distinguish PDM cases carrying different high-risk variants. Further, the results of a detailed simulation study suggest that we can identify positive evidence of association for moderate relative disease risks with a sample of 1,000 cases and 1,000 controls.

Algorithms↗

SNP identification, haplotype analysis, and parental origin of mutations in TSC2.

Inactivating mutations in the TSC2 gene, consisting of 41coding exons in 40 kb on 16p13, cause the hamartoma syndrome tuberous sclerosis. During TSC2 mutational analysis we identified ten SNPs that occur within or close to exon boundaries at minor allele frequencies greater than 5%. We determined the haplotypes for six of these SNPs and the microsatellite marker kg8 in the 3' region of TSC2 in a set of 40 parent-child trios. The most common haplotypes accounted for 53%, 11%, 6%, and 5% of chromosomes. Thirty-eight TSC2 mutation-bearing haplotypes had a similar distribution, indicating that there was no haplotype that predisposed to mutation in this region of TSC2. Family analysis was possible in 12 sporadic cases, and indicated that the mother was the parent of origin in 7 cases (3 point mutations, 2 small deletions, 2 large deletions), while the father was in 5 cases (2 point mutations, 3 small deletions). We conclude that TSC2 mutations occur at substantial frequency on both the maternally and paternally derived TSC2 alleles, in contrast to many other genetic diseases including NF1. The observations have implications for genetic counseling in TSC.

Adult↗