Search PubMed⌕ Search

Biomedical subjects

James D Iglehart

Publications and source records attributed to James D Iglehart.

3 recordsLinked to original sources

Recursive SVM feature selection and sample classification for mass-spectrometry and microarray data.

BACKGROUND: Like microarray-based investigations, high-throughput proteomics techniques require machine learning algorithms to identify biomarkers that are informative for biological classification problems. Feature selection and classification algorithms need to be robust to noise and outliers in the data. RESULTS: We developed a recursive support vector machine (R-SVM) algorithm to select important genes/biomarkers for the classification of noisy data. We compared its performance to a similar, state-of-the-art method (SVM recursive feature elimination or SVM-RFE), paying special attention to the ability of recovering the true informative genes/biomarkers and the robustness to outliers in the data. Simulation experiments show that a 5%- approximately 20% improvement over SVM-RFE can be achieved regard to these properties. The SVM-based methods are also compared with a conventional univariate method and their respective strengths and weaknesses are discussed. R-SVM was applied to two sets of SELDI-TOF-MS proteomics data, one from a human breast cancer study and the other from a study on rat liver cirrhosis. Important biomarkers found by the algorithm were validated by follow-up biological experiments. CONCLUSION: The proposed R-SVM method is suitable for analyzing noisy high-throughput proteomics and microarray data and it outperforms SVM-RFE in the robustness to noise and in the ability to recover informative features. The multivariate SVM-based method outperforms the univariate method in the classification performance, but univariate methods can reveal more of the differentially expressed features especially when there are correlations between the features.

Algorithms↗

Loss of heterozygosity and its correlation with expression profiles in subclasses of invasive breast cancers.

Gene expression array profiles identify subclasses of breast cancers with different clinical outcomes and different molecular features. The present study attempted to correlate genomic alterations (loss of heterozygosity; LOH) with subclasses of breast cancers having distinct gene expression signatures. Hierarchical clustering of expression array data from 89 invasive breast cancers identified four major expression subclasses. Thirty-four of these cases representative of the four subclasses were microdissected and allelotyped using genome-wide single nucleotide polymorphism detection arrays (Affymetrix, Inc.). LOH was determined by comparing tumor and normal single nucleotide polymorphism allelotypes. A newly developed statistical tool was used to determine the chromosomal regions of frequent LOH. We found that breast cancers were highly heterogeneous, with the proportion of LOH ranging widely from 0.3% to >60% of heterozygous markers. The most common sites of LOH were on 17p, 17q, 16q, 11q, and 14q, sites reported in previous LOH studies. Signature LOH events were discovered in certain expression subclasses. Unique regions of LOH on 5q and 4p marked a subclass of breast cancers with "basal-like" expression profiles, distinct from other subclasses. LOH on 1p and 16q occurred preferentially in a subclass of estrogen receptor-positive breast cancers. Finding unique LOH patterns in different groups of breast cancer, in part defined by expression signatures, adds confidence to newer schemes of molecular classification. Furthermore, exclusive association between biological subclasses and restricted LOH events provides rationale to search for targeted genes.

Breast Neoplasms↗