Search PubMed⌕ Search

PubMed · 10535635

Feature selection with limited datasets.

Abstract

Computer-aided diagnosis has the potential of increasing diagnostic accuracy by providing a second reading to radiologists. In many computerized schemes, numerous features can be extracted to describe suspect image regions. A subset of these features is then employed in a data classifier to determine whether the suspect region is abnormal or normal. Different subsets of features will, in general, result in different classification performances. A feature selection method is often used to determine an "optimal" subset of features to use with a particular classifier. A classifier performance measure (such as the area under the receiver operating characteristic curve) must be incorporated into this feature selection process. With limited datasets, however, there is a distribution in the classifier performance measure for a given classifier and subset of features. In this paper, we investigate the variation in the selected subset of "optimal" features as compared with the true optimal subset of features caused by this distribution of classifier performance. We consider examples in which the probability that the optimal subset of features is selected can be analytically computed. We show the dependence of this probability on the dataset sample size, the total number of features from which to select, the number of features selected, and the performance of the true optimal subset. Once a subset of features has been selected, the parameters of the data classifier must be determined. We show that, with limited datasets and/or a large number of features from which to choose, bias is introduced if the classifier parameters are determined using the same data that were employed to select the "optimal" subset of features.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

M A Kupinski, M L Giger. 1999. Feature selection with limited datasets.. https://doi.org/10.1118/1.598821

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Assessment of blinding in pharmacotherapy and noninvasive neuromodulation randomized controlled trials for neuropathic pain in adults.

In randomized controlled trials (RCTs), study participants and research personnel are often blinded to minimize biases related to knowing treatment allocation. To determine if blinding was effective, participants may be asked which treatment they believe they received ("treatment guess"). This descriptive review characterized blinding assessment (BA) reporting in pharmacotherapy and neuromodulation neuropathic pain RCTs. Of 288 papers, 36 (12.5%) reported a BA. One paper reported the results of 2 studies, so in total 37 studies with a BA were assessed. Of these, 19 were crossover, 17 parallel, and 1 partial crossover in design. All 37 studies assessed participant blinding, and 10 also assessed investigator blinding. Approximately 27% included an "unsure" answer option for treatment guess, and 38% asked the reason for the guess. There were no clear patterns in BA reporting across time nor based on treatment type. Seventeen trials provided sufficient data to calculate Bang Blinding Index (BI) to determine blinding success. Participants remained blinded (BI = 0 &#xb1; 0.2) in 10/17 placebo and 10/17 treatment arms, 6 placebo and 5 treatment arms had a BI > 0.2 suggesting possible unblinding, whereas 1 placebo and 2 treatment arms had a BI < -0.2 suggesting misinformed guessing. Overall, we found that BAs are done in a minority of published neuropathic pain trials and with variable methodology. Given the importance of minimizing risk of bias because of treatment unblinding, future studies should consider including BAs, and further consensus building is necessary to determine if and how BAs should be conducted and interpreted in analgesic clinical trials.

Bias↗

A residuals-based transition model for longitudinal analysis with estimation in the presence of missing data.

We propose a transition model for analysing data from complex longitudinal studies. Because missing values are practically unavoidable in large longitudinal studies, we also present a two-stage imputation method for handling general patterns of missing values on both the outcome and the covariates by combining multiple imputation with stochastic regression imputation. Our model is a time-varying auto-regression on the past innovations (residuals), and it can be used in cases where general dynamics must be taken into account, and where the model selection is important. The entire estimation process was carried out using available procedures in statistical packages such as SAS and S-PLUS. To illustrate the viability of the proposed model and the two-stage imputation method, we analyse data collected in an epidemiological study that focused on various factors relating to childhood growth. Finally, we present a simulation study to investigate the behaviour of our two-stage imputation procedure.

Bias↗

HIV viral dynamic models with dropouts and missing covariates.

In recent years HIV viral dynamic models have received great attention in AIDS studies. Often, subjects in these studies may drop out for various reasons such as drug intolerance or drug resistance, and covariates may also contain missing data. Statistical analyses ignoring informative dropouts and missing covariates may lead to misleading results. We consider appropriate methods for HIV viral dynamic models with informative dropouts and missing covariates and evaluate these methods via simulations. A real data set is analysed, and the results show that the initial viral decay rate, which may reflect the efficacy of the anti-HIV treatment, may be over-estimated if dropout patients are ignored. We also find that the current or immediate previous viral load values may be most predictive for patients' dropout. These results may be important for HIV/AIDS studies.

Bias↗