Search PubMed⌕ Search

PubMed · 15588316

Screening large-scale association study data: exploiting interactions using random forests.

Abstract

BACKGROUND: Genome-wide association studies for complex diseases will produce genotypes on hundreds of thousands of single nucleotide polymorphisms (SNPs). A logical first approach to dealing with massive numbers of SNPs is to use some test to screen the SNPs, retaining only those that meet some criterion for further study. For example, SNPs can be ranked by p-value, and those with the lowest p-values retained. When SNPs have large interaction effects but small marginal effects in a population, they are unlikely to be retained when univariate tests are used for screening. However, model-based screens that pre-specify interactions are impractical for data sets with thousands of SNPs. Random forest analysis is an alternative method that produces a single measure of importance for each predictor variable that takes into account interactions among variables without requiring model specification. Interactions increase the importance for the individual interacting variables, making them more likely to be given high importance relative to other variables. We test the performance of random forests as a screening procedure to identify small numbers of risk-associated SNPs from among large numbers of unassociated SNPs using complex disease models with up to 32 loci, incorporating both genetic heterogeneity and multi-locus interaction. RESULTS: Keeping other factors constant, if risk SNPs interact, the random forest importance measure significantly outperforms the Fisher Exact test as a screening tool. As the number of interacting SNPs increases, the improvement in performance of random forest analysis relative to Fisher Exact test for screening also increases. Random forests perform similarly to the univariate Fisher Exact test as a screening tool when SNPs in the analysis do not interact. CONCLUSIONS: In the context of large-scale genetic association studies where unknown interactions exist among true risk-associated SNPs or SNPs and environmental covariates, screening SNPs using random forest analyses can significantly reduce the number of SNPs that need to be retained for further study compared to standard univariate screening methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kathryn L Lunetta, L Brooke Hayward, Jonathan Segal, Paul Van Eerdewegh. 2004-12-10. Screening large-scale association study data: exploiting interactions using random forests.. https://doi.org/10.1186/1471-2156-5-32

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Control of confounding through secondary samples.

The control of confounding is essential in many statistical problems, especially in those that attempt to estimate exposure effects. In some cases, in addition to the 'primary' sample, there is another 'secondary' sample which, though having no direct information about the exposure effect, contains information about the confounding factors. The purpose of this article is to study the influence of the secondary sample on likelihood inference for the exposure effect. In particular, we investigate the interplay between the efficiency improvement and the possible bias introduced by the secondary sample as a function of the degree of confounding in the primary sample and the sizes of the primary and secondary samples. In the case of weak confounding, the secondary sample can only little improve estimation of the exposure effect, whereas with strong confounding the secondary sample can be much more useful. On the other hand, it will be more important to consider possible biasing effects in the latter case. For illustration, we use a formal example of a generalized linear model and a real example with sparse data from a case-control study of the association between gastric cancer and HM-CAP/Band 120.

Case-Control Studies↗

On estimation of the variance in Cochran-Armitage trend tests for genetic association using case-control studies.

The Cochran-Armitage trend test has been used in case-control studies for testing genetic association. As the variance of the test statistic is a function of unknown parameters, e.g. disease prevalence and allele frequency, it must be estimated. The usual estimator combining data for cases and controls assumes they follow the same distribution under the null hypothesis. Under the alternative hypothesis, however, the cases and controls follow different distributions. Thus, the power of the trend tests may be affected by the variance estimator used. In particular, the usual method combining both cases and controls is not an asymptotically unbiased estimator of the null variance when the alternative is true. Two different estimates of the null variance are available which are consistent under both the null and alternative hypotheses. In this paper, we examine sample size and small sample power performance of trend tests, which are optimal for three common genetic models as well as a robust trend test based on the three estimates of the variance and provide guidelines for choosing an appropriate test.

Case-Control Studies↗

Multiple hypothesis testing strategies for genetic case-control association studies.

The genetic case-control association study of unrelated subjects is a leading method to identify single nucleotide polymorphisms (SNPs) and SNP haplotypes that modulate the risk of complex diseases. Association studies often genotype several SNPs in a number of candidate genes; we propose a two-stage approach to address the inherent statistical multiple comparisons problem. In the first stage, each gene's association with disease is summarized by a single p-value that controls a familywise error rate. In the second stage, summary p-values are adjusted for multiplicity using a false discovery rate (FDR) controlling procedure. For the first stage, we consider marginal and joint tests of SNPs and haplotypes within genes, and we construct an omnibus test that combines SNP and haplotype analysis. Simulation studies show that when disease susceptibility is conferred by a SNP, and all common SNPs in a gene are genotyped, marginal analysis of SNPs using the Simes test has similar or higher power than marginal or joint haplotype analysis. Conversely, haplotype analysis can be more powerful when disease susceptibility is conferred by a haplotype. The omnibus test tracks the more powerful of the two approaches, which is generally unknown. Multiple testing balances the desire for statistical power against the implicit costs of false positive results, which up to now appear to be common in the literature.

Case-Control Studies↗