Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Association mapping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Association mapping and fine mapping with TreeLD.

SUMMARY: The program package TreeLD implements a unified approach to association mapping and fine mapping of complex trait loci and a novel approach to visualizing association data, based on an inferred ancestry of the sample. Fundamentally, the TreeLD approach is based on the idea that the evidence for association at a particular position is contained in the ancestral tree relating the sampled chromosomes at that position. TreeLD provides an easy-to-use interface and can be applied to case-control, TDT trio and quantitative trait data.

Algorithms↗

Family-based association mapping provides evidence for a gene for reading disability on chromosome 15q.

Family-based association mapping was used to follow up reports of linkage between reading disability (RD) and a genomic region on chromosome 15q. Using a two-stage approach, we ascertained 101 (stage 1) and 77 (stage 2) parent-proband trios, in which RD was characterized rigorously. In stage 1, a set of eight microsatellite markers spanning the region of putative linkage was used and a highly significant association was detected between RD and a three-marker haplotype (D15S994/D15S214/D15S146: P and empirical P < 0.001). A significant association with the same three-marker haplotype was also observed in the second-stage sample (P = 0.009, empirical P = 0.006). Our data therefore provide strong evidence for one or more genes contributing to RD being located in the vicinity of the region including D15S146 and D15S994. In addition, our results provide support for association analysis being a useful method to map susceptibility loci for complex disorders.

Chromosome Mapping↗

Quantitative-trait homozygosity and association mapping and empirical genomewide significance in large, complex pedigrees: fasting serum-insulin level in the Hutterites.

We present methods for linkage and association mapping of quantitative traits for a founder population with a large, known genealogy. We detect linkage to quantitative-trait loci (QTLs) through a multipoint homozygosity-mapping method. We propose two association methods, one of which is single point and uses a general two-allele model and the other of which is multipoint and uses homozygosity by descent for a particular allele. In all three methods, we make extensive use of the pedigree and genotype information, while keeping the computations simple and efficient. To assess significance, we have developed a permutation-based test that takes into account the covariance structure due to relatedness of individuals and can be used to determine empirical genomewide and locus-specific P values. In the case of multivariate-normally distributed trait data, the permutation-based test is asymptotically exact. The test is broadly applicable to a variety of mapping methods that fall within the class of linear statistical models (e.g., variance-component methods), under the assumption of random ascertainment with respect to the phenotype. For obtaining genomewide P values, our proposed method is appropriate when positions of markers are independent of the observed linkage signal, under the null hypothesis. We apply our methods to a genome screen for fasting insulin level in the Hutterites. We detect significant genomewide linkage on chromosome 19 and suggestive evidence of QTLs on chromosomes 1 and 16.

Alleles↗

A population-based latent variable approach for association mapping of quantitative trait loci.

A population-based latent variable approach is proposed for association mapping of quantitative trait loci (QTL), using multiple closely linked genetic markers within a small candidate region in the genome. By incorporating QTL as latent variables into a penetrance model, the QTL are flexible to characterize either alleles at putative trait loci or potential risk haplotypes/sub-haplotypes of the markers. Under a general likelihood framework, we develop an EM-based algorithm to estimate genetic effects of the QTL and haplotype frequencies of the QTL and markers jointly. Closed form solutions derived in the maximization step of the EM procedure for updating the joint haplotype frequencies of QTL and markers can effectively reduce the computational intensity. Various association measures between QTL and markers can then be derived from the haplotype frequencies of markers and used to infer QTL positions. The likelihood ratio statistic also provides a joint test for association between a quantitative trait and marker genotypes without requiring adjustment for the multiple testing. Extensive simulation studies are performed to evaluate the approach.

Algorithms↗

Association mapping of complex diseases in linked regions: estimation of genetic effects and feasibility of testing rare variants.

Association mapping in linked regions is a current major approach for the identification of genes for complex diseases. Loci contributing to linkage, even with small values of sibling recurrence risk (lambda(s)), may be equivalent to substantial underlying genetic effects for association studies. For disease alleles with a frequency as low as 1%, highly reliable association studies (80% power for significance level alpha=10(-6)) require only 277, 781, and 1289 families or cases and controls for loci detected with lambda(s) of 1.5, 1.1, and 1.05, respectively, under a multiplicative genetic model. Under alternative models, provided epistatic effects are minor, larger achievable sample sizes will provide sufficient power to map almost any disease gene that may have initially contributed to linkage.

Bias↗

Genetic association mapping based on discordant sib pairs: the discordant-alleles test.

Family-based tests of association provide the opportunity to test for an association between a disease and a genetic marker. Such tests avoid false-positive results produced by population stratification, so that evidence for association may be interpreted as evidence for linkage or causation. Several methods that use family-based controls have been proposed, including the haplotype relative risk, the transmission-disequilibrium test, and affected family-based controls. However, because these methods require genotypes on affected individuals and their parents, they are not ideally suited to the study of late-onset diseases. In this paper, we develop several family-based tests of association that use discordant sib pairs (DSPs) in which one sib is affected with a disease and the other sib is not. These tests are based on statistics that compare counts of alleles or genotypes or that test for symmetry in tables of alleles or genotypes. We describe the use of a permutation framework to assess the significance of these statistics. These DSP-based tests provide the same general advantages as parent-offspring trio-based tests, while being applicable to essentially any disease; they may also be tailored to particular hypotheses regarding the genetic model. We compare the statistical properties of our DSP-based tests by computer simulation and illustrate their use with an application to Alzheimer disease and the apolipoprotein E polymorphism. Our results suggest that the discordant-alleles test, which compares the numbers of nonmatching alleles in DSPs, is the most powerful of the tests we considered, for a wide class of disease models and marker types. Finally, we discuss advantages and disadvantages of the DSP design for genetic association mapping.

Alleles↗

Whole genome association mapping by incompatibilities and local perfect phylogenies.

BACKGROUND: With current technology, vast amounts of data can be cheaply and efficiently produced in association studies, and to prevent data analysis to become the bottleneck of studies, fast and efficient analysis methods that scale to such data set sizes must be developed. RESULTS: We present a fast method for accurate localisation of disease causing variants in high density case-control association mapping experiments with large numbers of cases and controls. The method searches for significant clustering of case chromosomes in the "perfect" phylogenetic tree defined by the largest region around each marker that is compatible with a single phylogenetic tree. This perfect phylogenetic tree is treated as a decision tree for determining disease status, and scored by its accuracy as a decision tree. The rationale for this is that the perfect phylogeny near a disease affecting mutation should provide more information about the affected/unaffected classification than random trees. If regions of compatibility contain few markers, due to e.g. large marker spacing, the algorithm can allow the inclusion of incompatibility markers in order to enlarge the regions prior to estimating their phylogeny. Haplotype data and phased genotype data can be analysed. The power and efficiency of the method is investigated on 1) simulated genotype data under different models of disease determination 2) artificial data sets created from the HapMap ressource, and 3) data sets used for testing of other methods in order to compare with these. Our method has the same accuracy as single marker association (SMA) in the simplest case of a single disease causing mutation and a constant recombination rate. However, when it comes to more complex scenarios of mutation heterogeneity and more complex haplotype structure such as found in the HapMap data our method outperforms SMA as well as other fast, data mining approaches such as HapMiner and Haplotype Pattern Mining (HPM) despite being significantly faster. For unphased genotype data, an initial step of estimating the phase only slightly decreases the power of the method. The method was also found to accurately localise the known susceptibility variants in an empirical data set--the DeltaF508 mutation for cystic fibrosis--where the susceptibility variant is already known--and to find significant signals for association between the CYP2D6 gene and poor drug metabolism, although for this dataset the highest association score is about 60 kb from the CYP2D6 gene. CONCLUSION: Our method has been implemented in the Blossoc (BLOck aSSOCiation) software. Using Blossoc, genome wide chip-based surveys of 3 million SNPs in 1000 cases and 1000 controls can be analysed in less than two CPU hours.

Chromosome Mapping↗

Bayesian association mapping for quantitative traits in a mixture of two populations.

We introduce a novel Bayesian approach to estimate and account for population structure simultaneously with association mapping of multiple quantitative trait loci. The method is designed for an analysis of unrelated individuals from a mixture of two populations (no admixture), where the individual population memberships are unknown. In our approach, the population structure is estimated and accounted for by using data on additional "grouping" markers which are assumed to be in Hardy-Weinberg equilibrium within the populations but have different allele frequencies between the populations. We use Bayesian hierarchical modeling and Markov chain Monte Carlo estimation, where we allow both population stratification and genetic heterogeneity. In our model the number of quantitative trait loci and their positions are treated as random variables, and we obtain their posterior distributions. Here we select the candidate and the grouping markers based on results from a preliminary SOLAR analysis.

Bayes Theorem↗

Linkage and association mapping of the LRP5 locus on chromosome 11q13 in type 1 diabetes.

Linkage of chromosome 11q13 to type 1 diabetes (T1D) was first reported from genome scans (Davies et al. 1994; Hashimoto et al. 1994) resulting in P <2.2 x 10(-5) (Luo et al. 1996) and designated IDDM4 ( insulin dependent diabetes mellitus 4). Association mapping under the linkage peak using 12 polymorphic microsatellite markers suggested some evidence of association with a two-marker haplotype, D11S1917*03-H0570POLYA*02, which was under-transmitted to affected siblings and over-transmitted to unaffected siblings ( P=1.5 x 10(-6)) (Nakagawa et al. 1998). Others have reported evidence for T1D association of the microsatellite marker D11S987, which is approximately 100 kb proximal to D11S1917 (Eckenrode et al. 2000). We have sequenced a 400-kb interval surrounding these loci and identified four genes, including the low-density lipoprotein receptor related protein (LRP5) gene, which has been considered as a functional candidate gene for T1D (Hey et al. 1998; Twells et al. 2001). Consequently, we have developed a comprehensive SNP map of the LRP5 gene region, and identified 95 SNPs encompassing 269 kb of genomic DNA, characterised the LD in the region and haplotypes (Twells et al. 2003). Here, we present our refined linkage curve of the IDDM4 region, comprising 32 microsatellite markers and 12 SNPs, providing a peak MLS=2.58, P=5 x 10(-4), at LRP5 g.17646G>T. The disease association data, largely focused in the LRP5 region with 1,106 T1D families, provided no further evidence for disease association at LRP5 or at D11S987. A second dataset, comprising 1,569 families from Finland, failed to replicate our previous findings at LRP5. The continued search for the variants of the putative IDDM4 locus will greatly benefit from the future development of a haplotype map of the genome.

Chromosome Mapping↗

TEAM: a tool for the integration of expression, and linkage and association maps.

The identification of genes primarily responsible for complex genetic disorders is a daunting task. Despite the assignment of many susceptibility loci, there has only been limited success in identifying disease genes based solely on positional information from genome-wide screens. The incorporation of several complementary strategies in a single integrated approach should facilitate and further enhance the efficacy of this search for genes. To permit the integration of linkage, association and expression data, together with functional annotations, we have developed a Java-based software tool: TEAM (tool for the integration of expression, and linkage and association maps). TEAM includes a genome viewer, capable of overlaying karyobands, genes, markers, linkage graphs, association data, gene expression levels and functional annotations in one composite view. Data management, analysis and filtering functionality was implemented and extended with links to the Ensembl, Unigene and Gene Ontology databases to facilitate gene annotation. Filtering functionality can help prevent the exclusion of poorly annotated, but differentially expressed, genes that reside in candidate regions that show linkage or association. Here we demonstrate the program's functionality in our study on coeliac disease (OMIM 212750), a multifactorial gluten-sensitive enteropathy. We performed a combined data analysis of a genome-wide linkage screen in 82 Dutch families with affected siblings and the microarray expression profiles of 18,110 cDNAs in 22 intestinal biopsies.

Celiac Disease↗

Association mapping of segregating sites in the early trypsin gene and susceptibility to dengue-2 virus in the mosquito Aedes aegypti.

Evidence suggests that midgut trypsins in Aedes aegypti condition the mosquito's ability to become infected with the dengue-2 flavivirus (DEN2). The activity of early trypsin protein peaks approximately 3 h after blood feeding and then drops within a few hours. We use association mapping to test the hypothesis that segregating sites in early trypsin condition midgut susceptibility to DEN2 virus. A total of 1642 females from throughout Mexico and the southern US were fed an artificial blood meal containing DEN2. After 2 weeks, mosquito heads and midguts were tested for DEN2. Mosquitoes with an infected head were classified as susceptible, those without a midgut infection had an infection barrier, and those with an infected gut but no head infection had an escape barrier. The early trypsin gene was amplified in two overlapping pieces from each mosquito and analyzed for single strand conformation polymorphisms (SSCPs). Unique SSCP genotypes were sequenced and 90 segregating sites were found. The dataset was divided into the four geographic regions within which Ae. aegypti is panmictic in Mexico. Heterogeneity chi2 analyses between alleles or genotypes and infection phenotypes demonstrated significant associations but allelic and genotypic effects were inconsistent among geographic regions. No consistent associations were found between segregating sites in early trypsin and susceptibility to DEN2 in Ae. aegypti in Mexico.

Aedes↗

A mixed-model approach to association mapping using pedigree information with an illustration of resistance to Phytophthora infestans in potato.

Association or linkage disequilibrium (LD)-based mapping strategies are receiving increased attention for the identification of quantitative trait loci (QTL) in plants as an alternative to more traditional, purely linkage-based approaches. An attractive property of association approaches is that they do not require specially designed crosses between inbred parents, but can be applied to collections of genotypes with arbitrary and often unknown relationships between the genotypes. A less obvious additional attractive property is that association approaches offer possibilities for QTL identification in crops with hard to model segregation patterns. The availability of candidate genes and targeted marker systems facilitates association approaches, as will appropriate methods of analysis. We propose an association mapping approach based on mixed models with attention to the incorporation of the relationships between genotypes, whether induced by pedigree, population substructure, or otherwise. Furthermore, we emphasize the need to pay attention to the environmental features of the data as well, i.e., adequate representation of the relations among multiple observations on the same genotypes. We illustrate our modeling approach using 25 years of Dutch national variety list data on late blight resistance in the genetically complex crop of potato. As markers, we used nucleotide binding-site markers, a specific type of marker that targets resistance or resistance-analog genes. To assess the consistency of QTL identified by our mixed-model approach, a second independent data set was analyzed. Two markers were identified that are potentially useful in selection for late blight resistance in potato.

Chromosome Mapping↗

Recent history of artificial outcrossing facilitates whole-genome association mapping in elite inbred crop varieties.

Genomewide association studies depend on the extent of linkage disequilibrium (LD), the number and distribution of markers, and the underlying structure in populations under study. Outbreeding species generally exhibit limited LD, and consequently, a very large number of markers are required for effective whole-genome association genetic scans. In contrast, several of the world's major food crops are self-fertilizing inbreeding species with narrow genetic bases and theoretically extensive LD. Together these are predicted to result in a combination of low resolution and a high frequency of spurious associations in LD-based studies. However, inbred elite plant varieties represent a unique human-induced pseudo-outbreeding population that has been subjected to strong selection for advantageous alleles. By assaying 1,524 genomewide SNPs we demonstrate that, after accounting for population substructure, the level of LD exhibited in elite northwest European barley, a typical inbred cereal crop, can be effectively exploited to map traits by using whole-genome association scans with several hundred to thousands of biallelic SNPs.

Crosses, Genetic↗

Association mapping: where we've been, where we're going.

This paper provides a review of recent work in the area of marker-phenotype association studies, specifically as used for localizing--or mapping--genes affecting a trait of interest. We describe the basis of association mapping and discuss a number of the commonly used techniques. We have also included references to various papers that have evaluated the use of these methods.

Analysis of Variance↗

An empirical comparison of case-control and trio based study designs in high throughput association mapping.

Motivated by high throughput genotyping technology, our aim in this study was to experimentally compare the power and accuracy of case-control and family trio based approaches for haplotype based, large scale, association gene mapping. We compared trio based and case-control study designs in different disease models, and partitioned the performance differences into separate components: those from the sample ascertainment, the effective sample size, and the haplotyping approaches. For systematic and controlled tests, we simulated a rapidly expanding and relatively young isolated population. The experiments were also replicated with real asthma data. We used computationally efficient methods that scale up to large amounts of both markers and individuals. Mapping is based on a haplotype association test for haplotypes of 1-10 markers. For population based haplotype reconstruction, we use HaploRec, and compare it to both a simple trio based inference and true haplotypes. Firstly and surprisingly, statistically inferred population based haplotypes can be equally powerful as true haplotypes. Secondly, as expected, the effective sample size has a clear effect on both gene detection power and mapping accuracy. Thirdly, the sample ascertainment method does not have much effect on mapping accuracy. Finally, an interesting side result is that the simple haplotype association test clearly outperformed exhaustive allelic transmission disequilibrium tests. The results suggest that the case-control design is a powerful alternative to the more laborious family based ascertainment approach, especially for large datasets, and wherever population stratification can be controlled.

Algorithms↗

Multilocus association mapping using variable-length Markov chains.

I propose a new method for association-based gene mapping that makes powerful use of multilocus data, is computationally efficient, and is straightforward to apply over large genomic regions. The approach is based on the fitting of variable-length Markov chain models, which automatically adapt to the degree of linkage disequilibrium (LD) between markers to create a parsimonious model for the LD structure. Edges of the fitted graph are tested for association with trait status. This approach can be thought of as haplotype testing with sophisticated windowing that accounts for extent of LD to reduce degrees of freedom and number of tests while maximizing information. I present analyses of two published data sets that show that this approach can have better power than single-marker tests or sliding-window haplotypic tests.

Chromosome Mapping↗

Are moment bounds on the recombination fraction between a marker and a disease locus too good to be true? Allelic association mapping revisited for simple genetic diseases in the Finnish population.

In the past several years, allelic association has helped map a number of rare genetic diseases in the human genome. A commonly used upper bound on the recombination fraction between the disease gene and an associated marker is known to be biased downward, so there is the possibility that an investigator could be misled. This upper bound is based on a moment equation that can be derived within the context of a Poisson branching process, so its performance can be compared with a recently proposed likelihood bound. We show that the confidence level of the moment upper bound is much lower than expected, while the confidence level of the likelihood bound is in line with expectation. The effects of mutation at either the marker or disease locus on the upper bounds are also investigated. Results indicate that mutation is not an important force for typical mutation rates, unless the recombination fraction between the marker and disease locus is very small or the disease allele is very rare in the general population. Finally, the impact of sample size on the likelihood bound is investigated. The results are illustrated with data on 10 simple genetic diseased in the Finnish population.

Alleles↗

Linkage and association mapping of a chromosome 1q21-q24 type 2 diabetes susceptibility locus in northern European Caucasians.

We have identified a region on chromosome 1q21-q24 that was significantly linked to type 2 diabetes in multiplex families of Northern European ancestry and also in Pima Indians, Amish families, and families from France and England. We sought to narrow and map this locus using a combination of linkage and association approaches by typing microsatellite markers at 1.2 and 0.5 cM densities, respectively, over a region of 37 cM (23.5 Mb). We tested linkage by parametric and nonparametric approaches and association using both case-control and family-based methods. In the 40 multiplex families that provided the previous evidence for linkage, the highest parametric, recessive logarithm of odds (LOD) score was 5.29 at marker D1S484 (168.5 cM, 157.5 Mb) without heterogeneity. Nonparametric linkage (NPL) statistics (P = 0.00009), SimWalk2 Statistic A (P = 0.0002), and sib-pair analyses (maximum likelihood score = 6.07) all mapped to the same location. The one LOD CI was narrowed to 156.8-158.9 Mb. Under recessive, two-point linkage analysis, adjacent markers D1S2675 (171.5 cM, 158.9 Mb) and D1S1679 (172 cM, 159.1 Mb) showed LOD scores >3.0. Nonparametric analyses revealed a second linkage peak at 180 cM near marker D1S1158 (163.3 Mb, NPL score 3.88, P = 0.0001), which was also supported by case-control (marker D1S194, 178 cM, 162.1 Mb; P = 0.003) and family-based (marker ATA38A05, 179 cM, 162.5 Mb; P = 0.002) association studies. We propose that the replicated linkage findings actually encompass at least two closely spaced regions, with a second susceptibility region located telomeric at 162.5-164.7 Mb.

Case-Control Studies↗