Search PubMed⌕ Search

Biomedical subjects

Brooke Hayward

Publications and source records attributed to Brooke Hayward.

4 recordsLinked to original sources

Identifying SNPs predictive of phenotype using random forests.

There has been a great interest and a few successes in the identification of complex disease susceptibility genes in recent years. Association studies, where a large number of single-nucleotide polymorphisms (SNPs) are typed in a sample of cases and controls to determine which genes are associated with a specific disease, provide a powerful approach for complex disease gene mapping. Genes of interest in those studies may contain large numbers of SNPs that classical statistical methods cannot handle simultaneously without requiring prohibitively large sample sizes. By contrast, high-dimensional nonparametric methods thrive on large numbers of predictors. This work explores the application of one such method, random forests, to the problem of identifying SNPs predictive of the phenotype in the case-control study design. A random forest is a collection of classification trees grown on bootstrap samples of observations, using a random subset of predictors to define the best split at each node. The observations left out of the bootstrap samples are used to estimate prediction error. The importance of a predictor is quantified by the increase in misclassification occurring when the values of the predictor are randomly permuted. We extend the concept of importance to pairs of predictors, to capture joint effects, and we explore the behavior of importance measures over a range of two-locus disease models in the presence of a varying number of SNPs unassociated with the phenotype. We illustrate the application of random forests with a data set of asthma cases and unaffected controls genotyped at 42 SNPs in ADAM33, a previously identified asthma susceptibility gene. SNPs and SNP pairs highly associated with asthma tend to have the highest importance index value, but predictive importance and association do not always coincide.

Case-Control Studies↗

Prognostic significance of DCC and p27Kip1 in colorectal cancer.

The progression of colorectal cancer is a multistage process associated with specific molecular alterations. The stepwise accumulation of these multiple genetic mutations progressively results in the acquisition of neoplastic cell behavior. The genetic abnormalities associated with the expression of metastatic phenotype, therefore, may be of prognostic significance in the clinical treatment of colorectal cancer patients. In this study, the immunohistochemical expression of the deleted in colorectal cancer gene (DCC) and p27Kip1 was assessed in 168 paraffin-embedded, formalin-fixed tumors of patients with stage II and III colorectal cancer. Kaplan-Meier survival curves and log-rank statistics were used to analyze survival times after curative primary tumor resection, and Cox proportional hazards models were used to adjust the assessment of demographic and clinical covariates. Loss of DCC or p27Kip1 expression had no influence on survival in patients with stage II or III colorectal cancer. The 5-year survival rates of DCC-positive and DCC-negative tumors were 51.8% and 35.7% (P=0.40), respectively. The 5-year survival rate of patients with p27Kip1-positive tumors was 47.9%, whereas the rate for patients with p27Kip1-negative tumors was 38.8% (P=0.68). After adjustment for all evaluated variables, neither DCC or p27Kip1 was found to be a predictor of survival (risk ratio for DCC, 0.98; 95% confidence interval, 0.66-1.56; P=0.92; risk ratio for p27Kip1, 0.87; 95% confidence interval, 0.58-1.29; P=0.49). The present study demonstrated that the expression of neither DCC nor p27Kip1 was predictive in poor survival outcome in patients with stage II or III colorectal cancer.

Biomarkers, Tumor↗

Mapping complex traits using Random Forests.

Random Forest is a prediction technique based on growing trees on bootstrap samples of data, in conjunction with a random selection of explanatory variables to define the best split at each node. In the case of a quantitative outcome, the tree predictor takes on a numerical value. We applied Random Forest to the first replicate of the Genetic Analysis Workshop 13 simulated data set, with the sibling pairs as our units of analysis and identity by descent (IBD) at selected loci as our explanatory variables. With the knowledge of the true model, we performed two sets of analyses on three phenotypes: HDL, triglycerides, and glucose. The goal was to approach the mapping of complex traits from a multivariate perspective. The first set of analyses mimics a candidate gene approach with a high proportion of true genes among the predictors while the second set represents a genome scan analysis using microsatellite markers. Random Forest was able to identify a few of the major genes influencing the phenotypes, such as baseline HDL and triglycerides, but failed to identify the major genes regulating baseline glucose levels.

Chromosome Mapping↗

Association of the ADAM33 gene with asthma and bronchial hyperresponsiveness.

Asthma is a common respiratory disorder characterized by recurrent episodes of coughing, wheezing and breathlessness. Although environmental factors such as allergen exposure are risk factors in the development of asthma, both twin and family studies point to a strong genetic component. To date, linkage studies have identified more than a dozen genomic regions linked to asthma. In this study, we performed a genome-wide scan on 460 Caucasian families and identified a locus on chromosome 20p13 that was linked to asthma (log(10) of the likelihood ratio (LOD), 2.94) and bronchial hyperresponsiveness (LOD, 3.93). A survey of 135 polymorphisms in 23 genes identified the ADAM33 gene as being significantly associated with asthma using case-control, transmission disequilibrium and haplotype analyses (P = 0.04 0.000003). ADAM proteins are membrane-anchored metalloproteases with diverse functions, which include the shedding of cell-surface proteins such as cytokines and cytokine receptors. The identification and characterization of ADAM33, a putative asthma susceptibility gene identified by positional cloning in an outbred population, should provide insights into the pathogenesis and natural history of this common disease.

ADAM Proteins↗