Search PubMed⌕ Search

Biomedical subjects

Daniel Q Naiman

Publications and source records attributed to Daniel Q Naiman.

10 recordsLinked to original sources

Random data set generation to support microarray analysis.

As microarray analyses become increasingly routine, involving the simultaneous investigation of huge numbers of genes, researchers can easily search for and uncover what appear to be promising patterns in their data. In such circumstances tools are needed to help decide the extent to which these patterns are meaningful or can be explained by chance alone. The purpose of this chapter is to describe examples of the use of microarray analysis for inferential purposes and how validation of inference is addressed by Monte-Carlo techniques, which essentially amounts to investigation of statistical methods on synthetic or random data sets.

Animals↗

Estimates of exposures to perchlorate from consumption of human milk, dairy milk, and water, and comparison to current reference dose.

To develop an enforceable drinking water standard from a health-based reference dose, sources of exposure and relevant exposure factors across the U.S. population must be considered. Human exposures, expressed as an estimated daily exposure, can be used to evaluate the health protectiveness of a range of potential regulatory values, thus providing a scientific foundation on which decisions can be based. Recent evidence points to detectable levels of perchlorate in milk and other foods. The purpose of this article is to estimate human exposure to perchlorate from ingestion of drinking water, human milk, and dairy milk. Drinking-water exposure was based on a range of possible regulatory values, derived from the recently established reference dose. Exposure to perchlorate from the consumption of milk was based on exploratory Food and Drug Administration dairy milk data, and on additional published perchlorate concentrations in dairy and human milk samples. This effort is exploratory in nature due to the limited data available at this time. However, it is anticipated that these exposure estimates and comparison with the current reference dose will stimulate dialogue and research that will advance the risk assessment for perchlorate.

Animals↗

Cortical reconstruction using implicit surface evolution: accuracy and precision analysis.

Two different studies were conducted to assess the accuracy and precision of an algorithm developed for automatic reconstruction of the cerebral cortex from T1-weighted magnetic resonance (MR) brain images. Repeated scans of three different brains were used to quantify the precision of the algorithm, and manually selected landmarks on different sulcal regions throughout the cortex were used to analyze the accuracy of the three reconstructed surfaces: inner, central, and pial. We conclude that the algorithm can find these surfaces in a robust fashion and with subvoxel accuracy, typically with an accuracy of one third of a voxel, although this varies with brain region and cortical geometry. Parameters were adjusted on the basis of this analysis in order to improve the algorithm's overall performance.

Algorithms↗

Robust prostate cancer marker genes emerge from direct integration of inter-study microarray data.

MOTIVATION: DNA microarray data analysis has been used previously to identify marker genes which discriminate cancer from normal samples. However, due to the limited sample size of each study, there are few common markers among different studies of the same cancer. With the rapid accumulation of microarray data, it is of great interest to integrate inter-study microarray data to increase sample size, which could lead to the discovery of more reliable markers. RESULTS: We present a novel, simple method of integrating different microarray datasets to identify marker genes and apply the method to prostate cancer datasets. In this study, by applying a new statistical method, referred to as the top-scoring pair (TSP) classifier, we have identified a pair of robust marker genes (HPN and STAT6) by integrating microarray datasets from three different prostate cancer studies. Cross-platform validation shows that the TSP classifier built from the marker gene pair, which simply compares relative expression values, achieves high accuracy, sensitivity and specificity on independent datasets generated using various array platforms. Our findings suggest a new model for the discovery of marker genes from accumulated microarray data and demonstrate how the great wealth of microarray data can be exploited to increase the power of statistical analysis. CONTACT: leixu@jhu.edu.

Algorithms↗

Simple decision rules for classifying human cancers from gene expression profiles.

MOTIVATION: Various studies have shown that cancer tissue samples can be successfully detected and classified by their gene expression patterns using machine learning approaches. One of the challenges in applying these techniques for classifying gene expression data is to extract accurate, readily interpretable rules providing biological insight as to how classification is performed. Current methods generate classifiers that are accurate but difficult to interpret. This is the trade-off between credibility and comprehensibility of the classifiers. Here, we introduce a new classifier in order to address these problems. It is referred to as k-TSP (k-Top Scoring Pairs) and is based on the concept of 'relative expression reversals'. This method generates simple and accurate decision rules that only involve a small number of gene-to-gene expression comparisons, thereby facilitating follow-up studies. RESULTS: In this study, we have compared our approach to other machine learning techniques for class prediction in 19 binary and multi-class gene expression datasets involving human cancers. The k-TSP classifier performs as efficiently as Prediction Analysis of Microarray and support vector machine, and outperforms other learning methods (decision trees, k-nearest neighbour and naïve Bayes). Our approach is easy to interpret as the classifier involves only a small number of informative genes. For these reasons, we consider the k-TSP method to be a useful tool for cancer classification from microarray gene expression data. AVAILABILITY: The software and datasets are available at http://www.ccbm.jhu.edu CONTACT: actan@jhu.edu.

Algorithms↗

p-Value simulation for affected sib pair multiple testing.

A standard approach to calculation of critical values for affected sib pair multiple testing is based on: (a) fully informative markers, (b) Haldane map function assumptions leading to a Markov chain model for inheritance vectors, (c) central limit approximation to averages of sampled inheritance vectors leading to an Ornstein-Uhlenbeck process approximation, and (d) simple approximations to the maximum of such a process. Under these assumptions, assuming equispaced or close to equispaced markers, if the sample size is large, an approximation is available that is easy to calculate and performs well. However, for small sample sizes, a large number of markers, and for small p-values, there is good reason to be cautious about the use of the Gaussian approximation. We develop an algorithm for calculation of multiple testing p-values based on the standard Markov chain model, avoiding the use of Gaussian (large sample) approximation. We illustrate the use of this algorithm by demonstrating some inadequacies of the Gaussian approximation.

Algorithms↗

Classifying gene expression profiles from pairwise mRNA comparisons.

We present a new approach to molecular classification based on mRNA comparisons. Our method, referred to as the top-scoring pair(s) (TSP) classifier, is motivated by current technical and practical limitations in using gene expression microarray data for class prediction, for example to detect disease, identify tumors or predict treatment response. Accurate statistical inference from such data is difficult due to the small number of observations, typically tens, relative to the large number of genes, typically thousands. Moreover, conventional methods from machine learning lead to decisions which are usually very difficult to interpret in simple or biologically meaningful terms. In contrast, the TSP classifier provides decision rules which i) involve very few genes and only relative expression values (e.g., comparing the mRNA counts within a single pair of genes); ii) are both accurate and transparent; and iii) provide specific hypotheses for follow-up studies. In particular, the TSP classifier achieves prediction rates with standard cancer data that are as high as those of previous studies which use considerably more genes and complex procedures. Finally, the TSP classifier is parameter-free, thus avoiding the type of over-fitting and inflated estimates of performance that result when all aspects of learning a predictor are not properly cross-validated.

Journal Article↗

Importance sampling method of correction for multiple testing in affected sib-pair linkage analysis.

Using the Genetic Analysis Workshop 13 simulated data set, we compared the technique of importance sampling to several other methods designed to adjust p-values for multiple testing: the Bonferroni correction, the method proposed by Feingold et al., and naïve Monte Carlo simulation. We performed affected sib-pair linkage analysis for each of the 100 replicates for each of five binary traits and adjusted the derived p-values using each of the correction methods. The type I error rates for each correction method and the ability of each of the methods to detect loci known to influence trait values were compared. All of the methods considered were conservative with respect to type I error, especially the Bonferroni method. The ability of these methods to detect trait loci was also low. However, this may be partially due to a limitation inherent in our binary trait definitions.

Computer Simulation↗

New genes involved in cancer identified by retroviral tagging.

Retroviral insertional mutagenesis in BXH2 and AKXD mice induces a high incidence of myeloid leukemia and B- and T-cell lymphoma, respectively. The retroviral integration sites (RISs) in these tumors thus provide powerful genetic tags for the discovery of genes involved in cancer. Here we report the first large-scale use of retroviral tagging for cancer gene discovery in the post-genome era. Using high throughput inverse PCR, we cloned and analyzed the sequences of 884 RISs from a tumor panel composed primarily of B-cell lymphomas. We then compared these sequences, and another 415 RIS sequences previously cloned from BXH2 myeloid leukemias and from a few AKXD lymphomas, against the recently assembled mouse genome sequence. These studies identified 152 loci that are targets of retroviral integration in more than one tumor (common retroviral integration sites, CISs) and therefore likely to encode a cancer gene. Thirty-six CISs encode genes that are known or predicted to be genes involved in human cancer or their homologs, whereas others encode candidate genes that have not yet been examined for a role in human cancer. Our studies demonstrate the power of retroviral tagging for cancer gene discovery in the post-genome era and indicate a largely unrecognized complexity in mouse and presumably human cancer.

Animals↗

A comprehensive method for genome scans.

In applications involving the use of genome scans the problem of correcting for multiple testing figures prominently. A frequently used approach is the Bonferroni adjustment, but this is known to be often severely conservative. As an alternative we use the method of importance sampling to accurately and efficiently obtain required exceedance probabilities. This method is comprehensive in the sense that it has application to exceedance probabilities for other classes of test statistics, such as those for linkage disequilibrium or Hardy-Weinberg equilibrium at multiple loci. We illustrate the importance sampling technique by focusing on affected sib pair tests done at a large number of fully informative markers. We demonstrate how our approach can be used to obtain exceedance probabilities for arbitrary marker spacings, and we compare our approach with that of Feingold et al. [1993], which uses the method of large deviations and does not provide the means for adjusting for unequal marker spacing.

Genetic Linkage↗