Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “nucleotide polymorphism patterns”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Patterns of linkage disequilibrium in the human genome.

Particular alleles at neighbouring loci tend to be co-inherited. For tightly linked loci, this might lead to associations between alleles in the population a property known as linkage disequilibrium (LD). LD has recently become the focus of intense study in the hope that it might facilitate the mapping of complex disease loci through whole-genome association studies. This approach depends crucially on the patterns of LD in the human genome. In this review, we draw on empirical studies in humans and Drosophila, as well as simulation studies, to assess the current state of knowledge about patterns of LD, and consider the implications for the use of LD as a mapping tool.

Animals↗

Detecting single-feature polymorphisms using oligonucleotide arrays and robustified projection pursuit.

MOTIVATION: Genomic DNA was hybridized to oligonucleotide microarrays to identify single-feature polymorphisms (SFP) for Arabidopsis, which has a genome size of approximately 130 Mb. However, that method does not work well for organisms such as barley, with a much larger 5200 Mb genome. In the present study, we demonstrate SFP detection using a small number of replicate datasets and complex RNA as a surrogate for barley DNA. To identify single probes defining SFPs in the data, we developed a method using robustified projection pursuit (RPP). This method first evaluates, for each probe set, the overall differentiation of signal intensities between two genotypes and then measures the contribution of the individual probes within the probe set to the overall differentiation. RESULTS: RNA from whole seedlings with and without dehydration stress provided 'present' calls for approximately 75% of probe sets. Using triplicated data, among the 5% of 'present' probe sets identified as most likely to contain at least one SFP probe, at least 80% are correctly predicted. This was determined by direct sequencing of PCR amplicons derived from barley genomic DNA. Using a 5 percentile cutoff, we defined 2007 SFP probes contained in 1684 probe sets by combining three parental genotype comparisons: Steptoe versus Morex, Morex versus Barke and Oregon Wolfe Barley Dominant versus Recessive. AVAILABILITY: The algorithm is available upon request from the corresponding author. CONTACT: xinping.cui@ucr.edu SUPPLEMENTARY INFORMATION: http://faculty.ucr.edu/~xpcui.

Algorithms↗

An efficient comprehensive search algorithm for tagSNP selection using linkage disequilibrium criteria.

MOTIVATION: Selecting SNP markers for genome-wide association studies is an important and challenging task. The goal is to minimize the number of markers selected for genotyping in a particular platform and therefore reduce genotyping cost while simultaneously maximizing the information content provided by selected markers. RESULTS: We devised an improved algorithm for tagSNP selection using the pairwise r(2) criterion. We first break down large marker sets into disjoint pieces, where more exhaustive searches can replace the greedy algorithm for tagSNP selection. These exhaustive searches lead to smaller tagSNP sets being generated. In addition, our method evaluates multiple solutions that are equivalent according to the linkage disequilibrium criteria to accommodate additional constraints. Its performance was assessed using HapMap data. AVAILABILITY: A computer program named FESTA has been developed based on this algorithm. The program is freely available and can be downloaded at http://www.sph.umich.edu/csg/qin/FESTA/

Algorithms↗

A self-tuning method for one-chip SNP identification.

Current methods for interpreting oligonucleotide-based SNP-detection microarrays, SNP chips, are based on statistics and require extensive parameter tuning as well as extremely high-resolution images of the chip being processed. We present a method, based on a simple data-classification technique called nearest-neighbors that, on haploid organisms, produces results comparable to the published results of the leading statistical methods and requires very little in the way of parameter tuning. Furthermore, it can interpret SNP chips using lower-resolution scanners of the type more typically used in current microarray experiments. Along with our algorithm, we present the results of a SNP-detection experiment where, when independently applying this algorithm to six identical SARS SNP chips, we correctly identify all 24 SNPs in a particular strain of the SARS virus, with between 6 and 13 false positives across the six experiments.

Algorithms↗

Gene mapping and marker clustering using Shannon's mutual information.

Finding the causal genetic regions underlying complex traits is one of the main aims in human genetics. In the context of complex diseases, which are believed to be controlled by multiple contributing loci of largely unknown effect and position, it is especially important to develop general yet sensitive methods for gene mapping. We discuss the use of Shannon's information theory for population-based gene mapping of discrete and quantitative traits and for marker clustering. Various measures of mutual information were employed in order to develop a comprehensive framework for gene mapping analyses. An algorithm aimed at finding so-called relevance chains of causal markers is proposed. Moreover, entropy measures are used in conjunction with multidimensional scaling to visualize clusters of genetic markers. The relevance chain algorithm successfully detected the two causal regions in a simulated scenario. The approach has also been applied to a published clinical study on autoimmune (Graves') disease. Results were consistent with those of standard statistical methods, but identified an additional locus of interest in the promotor region of the associated gene CTLA4. The developed software is freely available at http://www.Int.ei.tum.de/download/InfoGeneMap/.

Algorithms↗

The fine-scale structure of recombination rate variation in the human genome.

The nature and scale of recombination rate variation are largely unknown for most species. In humans, pedigree analysis has documented variation at the chromosomal level, and sperm studies have identified specific hotspots in which crossing-over events cluster. To address whether this picture is representative of the genome as a whole, we have developed and validated a method for estimating recombination rates from patterns of genetic variation. From extensive single-nucleotide polymorphism surveys in European and African populations, we find evidence for extreme local rate variation spanning four orders in magnitude, in which 50% of all recombination events take place in less than 10% of the sequence. We demonstrate that recombination hotspots are a ubiquitous feature of the human genome, occurring on average every 200 kilobases or less, but recombination occurs preferentially outside genes.

Base Composition↗

Analysis of SNP-expression association matrices.

High throughput expression profiling and genotyping technologies provide the means to study the genetic determinants of population variation in gene expression variation. In this paper we present a general statistical framework for the simultaneous analysis of gene expression data and SNP genotype data measured for the same cohort. The framework consists of methods to associate transcripts with SNPs affecting their expression, algorithms to detect subsets of transcripts that share significantly many associations with a subset of SNPs, and methods to visualize the identified relations. We apply our framework to SNP-expression data collected from 50 breast cancer patients. Our results demonstrate an overabundance of transcript-SNP associations in this data, and pinpoint SNPs that are potential master regulators of transcription. We also identify several statistically significant transcript-subsets with common putative regulators that fall into well-defined functional categories.

Algorithms↗

Haplotype-based association studies of IGFBP1 and IGFBP3 with prostate and breast cancer risk: the multiethnic cohort.

Collective evidence suggests that the insulin-like growth factor (IGF) system plays a role in prostate and breast cancer risk. IGF-binding proteins (IGFBP) are the principal regulatory molecules that modulate IGF-I bioavailability in the circulation and tissues. To examine whether inherited differences in the IGFBP1 and IGFBP3 genes influence prostate and breast cancer susceptibility, we conducted two large population-based association studies of African Americans, Native Hawaiians, Japanese Americans, Latinos, and Whites. To thoroughly assess the genetic variation across the two loci, we (a) sequenced the IGFBP1 and IGFBP3 exons in 95 aggressive prostate and 95 advanced breast cancer cases to ensure that we had identified all common missense variants and (b) characterized the linkage disequilibrium patterns and common haplotypes by genotyping 36 single nucleotide polymorphisms (SNP) spanning 71 kb across the loci ( approximately 20 kb upstream and approximately 40 kb downstream, respectively) in a panel of 349 control subjects of the five racial/ethnic groups. No new missense SNPs were found. We identified three regions of strong linkage disequilibrium and selected a subset of 23 tagging SNPs that could accurately predict both the common IGFBP1 and IGFBP3 haplotypes and the remaining 13 SNPs. We tested the association between IGFBP1 and IGFBP3 genotypes and haplotypes for their associations with prostate and breast cancer risk in two large case-control studies nested within the Multiethnic Cohort [prostate cases/controls = 2,320/2,290; breast cases (largely postmenopausal)/controls = 1,615/1,962]. We observed no strong associations between IGFBP1 and IGFBP3 genotypes or haplotypes with either prostate or breast cancer risk. Our results suggest that common genetic variation in the IGFBP1 and IGFBP3 genes do not substantially influence prostate and breast cancer susceptibility.

Black or African American↗

GPNN: power studies and applications of a neural network method for detecting gene-gene interactions in studies of human disease.

BACKGROUND: The identification and characterization of genes that influence the risk of common, complex multifactorial disease primarily through interactions with other genes and environmental factors remains a statistical and computational challenge in genetic epidemiology. We have previously introduced a genetic programming optimized neural network (GPNN) as a method for optimizing the architecture of a neural network to improve the identification of gene combinations associated with disease risk. The goal of this study was to evaluate the power of GPNN for identifying high-order gene-gene interactions. We were also interested in applying GPNN to a real data analysis in Parkinson's disease. RESULTS: We show that GPNN has high power to detect even relatively small genetic effects (2-3% heritability) in simulated data models involving two and three locus interactions. The limits of detection were reached under conditions with very small heritability (<1%) or when interactions involved more than three loci. We tested GPNN on a real dataset comprised of Parkinson's disease cases and controls and found a two locus interaction between the DLST gene and sex. CONCLUSION: These results indicate that GPNN may be a useful pattern recognition approach for detecting gene-gene and gene-environment interactions.

Algorithms↗

Multilocus analysis of SNP and metabolic data within a given pathway.

BACKGROUND: Complex traits, which are under the influence of multiple and possibly interacting genes, have become a subject of new statistical methodological research. One of the greatest challenges facing human geneticists is the identification and characterization of susceptibility genes for common multifactorial diseases and their association to different quantitative phenotypic traits. RESULTS: Two types of data from the same metabolic pathway were used in the analysis: categorical measurements of 18 SNPs; and quantitative measurements of plasma levels of several steroids and their precursors. Using the combinatorial partitioning method we tested various thresholds for each metabolic trait and each individual SNP locus. One SNP in CYP19, 3UTR, two SNPs in CYP1B1 (R48G and A119S) and one in CYP1A1 (T461N) were significantly differently distributed between the high and low level metabolic groups. The leave one out cross validation method showed that 6 SNPs in concert make 65% correct prediction of phenotype. Further we used pattern recognition, computing the p-value by Monte Carlo simulation to identify sets of SNPs and physiological characteristics such as age and weight that contribute to a given metabolic level. Since the SNPs detected by both methods reside either in the same gene (CYP1B1) or in 3 different genes in immediate vicinity on chromosome 15 (CYP19, CYP11 and CYP1A1) we investigated the possibility that they form intragenic and intergenic haplotypes, which may jointly account for a higher activity in the pathway. We identified such haplotypes associated with metabolic levels. CONCLUSION: The methods reported here may enable to study multiple low-penetrance genetic factors that together determine various quantitative phenotypic traits. Our preliminary data suggest that several genes coding for proteins involved in a common pathway, that happen to be located on common chromosomal areas and may form intragenic haplotypes, together account for a higher activity of the whole pathway.

Aged↗

A tool for selecting SNPs for association studies based on observed linkage disequilibrium patterns.

The design of genetic association studies using single-nucleotide polymorphisms (SNPs) requires the selection of subsets of the variants providing high statistical power at a reasonable cost. SNPs must be selected to maximize the probability that a causative mutation is in linkage disequilibrium (LD) with at least one marker genotyped in the study. The HapMap project performed a genome-wide survey of genetic variation with about a million SNPs typed in four populations, providing a rich resource to inform the design of association studies. A number of strategies have been proposed for the selection of SNPs based on observed LD, including construction of metric LD maps and the selection of haplotype tagging SNPs. Power calculations are important at the study design stage to ensure successful results. Integrating these methods and annotations can be challenging: the algorithms required to implement these methods are complex to deploy, and all the necessary data and annotations are deposited in disparate databases. Here, we present the SNPbrowser Software, a freely available tool to assist in the LD-based selection of markers for association studies. This stand-alone application provides fast query capabilities and swift visualization of SNPs, gene annotations, power, haplotype blocks, and LD map coordinates. Wizards implement several common SNP selection workflows including the selection of optimal subsets of SNPs (e.g. tagging SNPs). Selected SNPs are screened for their conversion potential to either TaqMan SNP Genotyping Assays or the SNPlex Genotyping System, two commercially available genotyping platforms, expediting the set-up of genetic studies with an increased probability of success.

Computational Biology↗

Linkage disequilibrium patterns across a recombination gradient in African Drosophila melanogaster.

Previous multilocus surveys of nucleotide polymorphism have documented a genome-wide excess of intralocus linkage disequilibrium (LD) in Drosophila melanogaster and D. simulans relative to expectations based on estimated mutation and recombination rates and observed levels of diversity. These studies examined patterns of variation from predominantly non-African populations that are thought to have recently expanded their ranges from central Africa. Here, we analyze polymorphism data from a Zimbabwean population of D. melanogaster, which is likely to be closer to the standard population model assumptions of a large population with constant size. Unlike previous studies, we find that levels of LD are roughly compatible with expectations based on estimated rates of crossing over. Further, a detailed examination of genes in different recombination environments suggests that markers near the telomere of the X chromosome show considerably less linkage disequilibrium than predicted by rates of crossing over, suggesting appreciable levels of exchange due to gene conversion. Assuming that these populations are near mutation-drift equilibrium, our results are most consistent with a model that posits heterogeneity in levels of exchange due to gene conversion across the X chromosome, with gene conversion being a minor determinant of LD levels in regions of high crossing over. Alternatively, if levels of exchange due to gene conversion are not negligible in regions of high crossing over, our results suggest a marked departure from mutation-drift equilibrium (i.e., toward an excess of LD) in this Zimbabwean population. Our results also have implications for the dynamics of weakly selected mutations in regions of reduced crossing over.

Animals↗

Inferring the fitness effects of DNA mutations from polymorphism and divergence data: statistical power to detect directional selection under stationarity and free recombination.

The fitness effects of classes of DNA mutations can be inferred from patterns of nucleotide variation. A number of studies have attributed differences in levels of polymorphism and divergence between silent and replacement mutations to the action of natural selection. Here, I investigate the statistical power to detect directional selection through contrasts of DNA variation among functional categories of mutations. A variety of statistical approaches are applied to DNA data simulated under Sawyer and Hartl's Poisson random field model. Under assumptions of free recombination and stationarity, comparisons that include both the frequency distributions of mutations segregating within populations and the numbers of mutations fixed between populations have substantial power to detect even very weak selection. Frequency distribution and divergence tests are applied to silent and replacement mutations among five alleles of each of eight Drosophila simulans genes. Putatively "preferred" silent mutations segregate at higher frequencies and are more often fixed between species than "unpreferred" silent changes, suggesting fitness differences among synonymous codons. Amino acid changes tend to be either rare polymorphisms or fixed differences, consistent with a combination of deleterious and adaptive protein evolution. In these data, a substantial fraction of both silent and replacement DNA mutations appear to affect fitness.

Adaptation, Biological↗

Characterization and functional investigation of single nucleotide polymorphisms (SNPs) in the human TLR5 gene.

Toll-like receptors recognize pathogen-associated molecular patterns (PAMPs) and TLR5 is the pathogen recognition receptor (PRR) for bacterial flagellin. Patients carrying a R392 stop polymorphism display an inflammatory phenotype and increased susceptibility to pneumonia caused by the flagellated bacteria Legionella pneumophila. While this suggests that TLR5 mutations may be clinically relevant, functional data are not available for the majority of the other TLR5 polymorphisms. We have characterized all known single nucleotide polymorphisms (SNPs) of TLR5 for their functional relevance upon stimulation in transiently transfected CHO-K1 cells. Among the 13 missense SNPs of TLR5 reported in the human genetic databases, three SNPs (c.1174C>T, p.R392X; c.2081A>G, p.D694G; and c.2464C>T, p.L822F) were found to be functionally relevant in transiently transfected CHO-K1 cells. The prevalences of these functionally relevant SNPs in our investigation were 11.9 %, 0 %, and 0 %, in healthy donors. The p.D694G and p.L822F SNPs are of low frequency in the Caucasian population though further investigations of the common p.R392X variant alone or of functional relevant TLR5 SNPs in combination with other TLR SNPs will elucidate their possible role on disease susceptibility in humans and may facilitate clinical diagnosis.

Alleles↗

High-throughput single-strand conformation polymorphism analysis by automated capillary electrophoresis: robust multiplex analysis and pattern-based identification of allelic variants.

Genetic diagnosis of an inherited disease or cancer often involves analysis for unknown point mutations in several genes; therefore, rapid and automated techniques that can process a large number of samples are needed. We describe a method for high-throughput single-strand conformation polymorphism (SSCP) analysis using automated capillary electrophoresis. The operating temperature of a commercially available capillary electrophoresis instrument (ABI PRISM 310) was expanded by installation of a cheap in-house designed cooling system, thereby allowing us to perform automated SSCP analysis at 14-45 degrees C. We have used the method for detection of point mutations associated with the inherited cardiac disorders long QT syndrome (LQTS) and hypertrophic cardiomyopathy (HCM). The sensitivity of the method was 100% when 34 different point mutations were analyzed, including two previously unpublished LQTS-associated mutations (F157C in KVLQT1 and G572R in HERG), as well as eight novel normal variants in HERG and MYH7. The analyzed polymerase chain reaction (PCR) fragments ranged in size from 166 to 1,223 bp. Seventeen different sequence contexts were analyzed. Three different electrophoresis temperatures were used to obtain 100% sensitivity. Two mutants could not be detected at temperatures greater than 20 degrees C. The method has a high resolution and good reproducibility and is very robust, making multiplex SSCP analysis and pattern-based identification of known allelic variants as single nucleotide polymorphisms (SNPs) possible. These possibilities, combined with automation and short analysis time, make the method suitable for high-throughput tasks, such as genetic screening.

Alleles↗

Genetic diversity and evolution of Mycoplasma capricolum subsp. capripneumoniae strains from eastern Africa assessed by 16S rDNA sequence analysis.

Mycoplasma capricolum subsp. capripneumoniae (M. capripneumoniae), the causal agent of contagious caprine pleuropneumonia (CCPP), is a member of the so-called Mycoplasma mycoides cluster. These mycoplasmas have two rRNA operons in which intraspecific variations have been demonstrated. The sequences of the 16S rRNA genes of both operons from 13 field strains of M. capripneumoniae from three neighbouring African countries (Kenya, Ethiopia, and Tanzania) were determined. Four new and unique polymorphism patterns reflecting the intraspecific variations were found. Two of these patterns included length differences between the rrnA and rrnB operons. The length difference in one of the patterns was caused by a two-nucleotide insert (TG) in the rrnB operon and the length difference in the other pattern was due to a three-nucleotide deletion, also in the rrnB operon. Another pattern was characterised by a polymorphic position caused by a mutation that is known to cause streptomycin resistance in other bacterial species. The strain with this pattern was also found to be resistant to streptomycin. Streptomycin resistant clones were selected from four M. capripneumoniae strains to further investigate the correlation of this mutation to streptomycin resistance. Mutations in the 16S rRNA genes had occurred in two of these strains. The fourth pattern included a new polymorphism in position 1059. The results show that polymorphisms in M. capripneumoniae strains can be used as epidemiological markers for CCPP in smaller geographical areas and to study the molecular evolution of this species.

Animals↗

Genetic variations in the WFS1 gene in Japanese with type 2 diabetes and bipolar disorder.

Diabetic and psychiatric symptoms often appear in patients with Wolfram syndrome, and obligate carriers of WFS1 have increased prevalence of type 2 diabetes and are more likely to require hospitalization for psychiatric illness including bipolar disorder. To identify the polymorphisms in Japanese, we examined a region of approximately 50 kb covering the entire WFS1 gene, and evaluated the patterns of linkage disequilibrium. We found a total of 42 variations including 8 novel coding single nucleotide polymorphisms (A6T, A134A, N159N, T170T, E237K, R383C, V412L, and V503G), 14 novel non-coding polymorphisms, and 2 linkage disequilibrium blocks. We also performed association studies in patients with type 2 diabetes mellitus and patients with bipolar disorder. The haplotype comprising R456 and H611 was most associated with type 2 diabetes (p = 0.013) and the haplotype comprising g. -15503C/T and g. 16226G/A was most associated with bipolar disorder (p = 0.006), but neither reached significant difference after multiple adjustment. These genetic variations and linkage disequilibrium patterns in WFS1 in Japanese should be useful in further investigation of genetic diversities of WFS1 and various related disorders.

Aged↗

Molecular scanning of insulin-responsive glucose transporter (GLUT4) gene in NIDDM subjects.

We investigated the prevalence of mutations in the gene encoding the major insulin-responsive facilitative glucose transporter (GLUT4) in patients with non-insulin-dependent diabetes mellitus (NIDDM). All 11 exons of the GLUT4 gene from 30 British white subjects with NIDDM were amplified using the polymerase chain reaction and screened for nucleotide sequence variation using the single-stranded conformation polymorphism (SSCP) method. No variation between the study subjects was detected in exons 1-3, 4b-8, and 10. Variant SSCP patterns were detected in exons 4a and 9. SSCP variation in exon 4a was revealed by direct nucleotide sequencing to be due to a common silent polymorphism (AAC----AAT at Asn130). One NIDDM patient demonstrated a variant SSCP pattern in exon 9. This was caused by a point mutation (GTC----ATC) at codon 383, which leads to the conservative substitution of isoleucine for valine in the putative fifth extracellular loop of the transporter. Allele-specific oligonucleotide hybridization was used to examine the frequency of this mutation in 240 Welsh white subjects (160 with NIDDM and 80 controls). The Val----Ile383 mutation was found in the heterozygous state in two diabetic subjects and no control subjects. We conclude that mutations of the GLUT4 coding sequence are very uncommon in this population of subjects with typical NIDDM. Determining whether the Ile383 GLUT4 variant present in 3 diabetic subjects contributes in any way to their disease will require further study.

Alleles↗