Search PubMed⌕ Search

Biomedical subjects

Bill C White

Publications and source records attributed to Bill C White.

5 recordsLinked to original sources

A flexible computational framework for detecting, characterizing, and interpreting statistical patterns of epistasis in genetic studies of human disease susceptibility.

Detecting, characterizing, and interpreting gene-gene interactions or epistasis in studies of human disease susceptibility is both a mathematical and a computational challenge. To address this problem, we have previously developed a multifactor dimensionality reduction (MDR) method for collapsing high-dimensional genetic data into a single dimension (i.e. constructive induction) thus permitting interactions to be detected in relatively small sample sizes. In this paper, we describe a comprehensive and flexible framework for detecting and interpreting gene-gene interactions that utilizes advances in information theory for selecting interesting single-nucleotide polymorphisms (SNPs), MDR for constructive induction, machine learning methods for classification, and finally graphical models for interpretation. We illustrate the usefulness of this strategy using artificial datasets simulated from several different two-locus and three-locus epistasis models. We show that the accuracy, sensitivity, specificity, and precision of a naïve Bayes classifier are significantly improved when SNPs are selected based on their information gain (i.e. class entropy removed) and reduced to a single attribute using MDR. We then apply this strategy to detecting, characterizing, and interpreting epistatic models in a genetic study (n = 500) of atrial fibrillation and show that both classification and model interpretation are significantly improved.

Atrial Fibrillation↗

Multilocus analysis of hypertension: a hierarchical approach.

While hypertension is a complex disease with a well-documented genetic component, genetic studies often fail to replicate findings. One possibility for such inconsistency is that the underlying genetics of hypertension is not based on single genes of major effect, but on interactions among genes. To test this hypothesis, we studied both single locus and multilocus effects, using a case-control design of subjects from Ghana. Thirteen polymorphisms in eight candidate genes were studied. Each candidate gene has been shown to play a physiological role in blood pressure regulation and affects one of four pathways that modulate blood pressure: vasoconstriction (angiotensinogen, angiotensin converting enzyme - ACE, angiotensin II receptor), nitric oxide (NO) dependent and NO independent vasodilation pathways and sodium balance (G protein-coupled receptor kinase, GRK4). We evaluated single site allelic and genotypic associations, multilocus genotype equilibrium and multilocus genotype associations, using multifactor dimensionality reduction (MDR). For MDR, we performed systematic reanalysis of the data to address the role of various physiological pathways. We found no significant single site associations, but the hypertensive class deviated significantly from genotype equilibrium in more than 25% of all multilocus comparisons (2,162 of 8,178), whereas the normotensive class rarely did (11 of 8,178). The MDR analysis identified a two-locus model including ACE and GRK4 that successfully predicted blood pressure phenotype 70.5% of the time. Thus, our data indicate epistatic interactions play a major role in hypertension susceptibility. Our data also support a model where multiple pathways need to be affected in order to predispose to hypertension.

Alleles↗

Integrated analysis of genetic, genomic and proteomic data.

The rapid expansion of methods for measuring biological data ranging from DNA sequence variations to mRNA expression and protein abundance presents the opportunity to utilize multiple types of information jointly in the study of human health and disease. Organisms are complex systems that integrate inputs at myriad levels to arrive at an observable phenotype. Therefore, it is essential that questions concerning the etiology of phenotypes as complex as common human diseases take the systemic nature of biology into account, and integrate the information provided by each data type in a manner analogous to the operation of the body itself. While limited in scope, the initial forays into the joint analysis of multiple data types have yielded interesting results that would not have been reached had only one type of data been considered. These early successes, along with the aforementioned theoretical appeal of data integration, provide impetus for the development of methods for the parallel, high-throughput analysis of multiple data types. The idea that the integrated analysis of multiple data types will improve the identification of biomarkers of clinical endpoints, such as disease susceptibility, is presented as a working hypothesis.

Animals↗

Proteomic patterns of tumour subsets in non-small-cell lung cancer.

BACKGROUND: Proteomics-based approaches complement the genome initiatives and may be the next step in attempts to understand the biology of cancer. We used matrix-assisted laser desorption/ionisation mass spectrometry directly from 1-mm regions of single frozen tissue sections for profiling of protein expression from surgically resected tissues to classify lung tumours. METHODS: Proteomic spectra were obtained and aligned from 79 lung tumours and 14 normal lung tissues. We built a class-prediction model with the proteomic patterns in a training cohort of 42 lung tumours and eight normal lung samples, and assessed their statistical significance. We then applied this model to a blinded test cohort, including 37 lung tumours and six normal lung samples, to estimate the misclassification rate. FINDINGS: We obtained more than 1600 protein peaks from histologically selected 1 mm diameter regions of single frozen sections from each tissue. Class-prediction models based on differentially expressed peaks enabled us to perfectly classify lung cancer histologies, distinguish primary tumours from metastases to the lung from other sites, and classify nodal involvement with 85% accuracy in the training cohort. This model nearly perfectly classified samples in the independent blinded test cohort. We also obtained a proteomic pattern comprised of 15 distinct mass spectrometry peaks that distinguished between patients with resected non-small-cell lung cancer who had poor prognosis (median survival 6 months, n=25) and those who had good prognosis (median survival 33 months, n=41, p<0.0001). INTERPRETATION: Proteomic patterns obtained directly from small amounts of fresh frozen lung-tumour tissue could be used to accurately classify and predict histological groups as well as nodal involvement and survival in resected non-small-cell lung cancer.

Biomarkers, Tumor↗

Optimization of neural network architecture using genetic programming improves detection and modeling of gene-gene interactions in studies of human diseases.

BACKGROUND: Appropriate definition of neural network architecture prior to data analysis is crucial for successful data mining. This can be challenging when the underlying model of the data is unknown. The goal of this study was to determine whether optimizing neural network architecture using genetic programming as a machine learning strategy would improve the ability of neural networks to model and detect nonlinear interactions among genes in studies of common human diseases. RESULTS: Using simulated data, we show that a genetic programming optimized neural network approach is able to model gene-gene interactions as well as a traditional back propagation neural network. Furthermore, the genetic programming optimized neural network is better than the traditional back propagation neural network approach in terms of predictive ability and power to detect gene-gene interactions when non-functional polymorphisms are present. CONCLUSION: This study suggests that a machine learning strategy for optimizing neural network architecture may be preferable to traditional trial-and-error approaches for the identification and characterization of gene-gene interactions in common, complex human diseases.

Algorithms↗