Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Model selection for a medical diagnostic decision support system: a breast cancer detection case.

There are a number of different quantitative models that can be used in a medical diagnostic decision support system (MDSS) including parametric methods (linear discriminant analysis or logistic regression), non-parametric models (K nearest neighbor, or kernel density) and several neural network models. The complexity of the diagnostic task is thought to be one of the prime determinants of model selection. Unfortunately, there is no theory available to guide model selection. Practitioners are left to either choose a favorite model or to test a small subset using cross validation methods. This paper illustrates the use of a self-organizing map (SOM) to guide model selection for a breast cancer MDSS. The topological ordering properties of the SOM are used to define targets for an ideal accuracy level similar to a Bayes optimal level. These targets can then be used in model selection, variable reduction, parameter determination, and to assess the adequacy of the clinical measurement system. These ideas are applied to a successful model selection for a real-world breast cancer database. Diagnostic accuracy results are reported for individual models, for ensembles of neural networks, and for stacked predictors.

Algorithms↗

Model selection for covariance structures analysis in nursing research.

Covariance structures analysis is often used in nursing research to appraise statistical models reflecting complex human health processes. The model selection approach in covariance structures analysis is designed to select the "best" model from a specified set of theoretically defensible, competing alternatives, all of which are viewed as approximations. Model selection criteria explicitly incorporate both model misfit in the population and sampling error to evaluate the set of models. The result is that interpretability of model parameters and goodness-of-fit are enhanced simultaneously. Relative merits of the model selection approach are identified in light of technical concerns, parsimony, and use of scientific theory in nursing.

Analysis of Variance↗

Likelihood-based diagnostics for influential individuals in non-linear mixed effects model selection.

PURPOSE: Data from single individuals, or a small group of subjects may influence non-linear mixed effects model selection. Diagnostics routinely applied in model building may identify such individuals, but these methods are not specifically designed for that purpose and are, therefore, not optimal. We describe two likelihood-based diagnostics for identifying individuals that can influence the choice between two competing models. METHODS: One method is based on a jackknife of the raw data on the individual level and refitting the model to each new data set. The second method is a calculation which utilises the contribution each individual make to the objective function values under each of the two models. The two methods were applied to model selection during analysis of a real data set. RESULTS: The agreement between the methods was high. Individuals for whom there was a discrepancy between the methods tended to be those for which neither of the contending models described the data appropriately. Both methods identified individuals that influenced the model selection. CONCLUSIONS: Two objective, specific and quantitative methods for identifying influential individuals in nonlinear mixed effects model selection have been presented. One of the methods doesn't require additional model fitting and is therefore particularly attractive.

Age Factors↗

Maintenance of genetic variation with a frequency-dependent selection model as compared to the overdominant model.

A frequency-dependent selection model proposed by Huang, Singh and Kojima (1971) was found to be more effective at maintaining genetic variation in a finite population than the overdominant model. The fourth moment parameter of the distribution of unfixed states showed that there was a more platykurtic distribution for the frequency-dependent model. This agreed well with the expected gene frequency change found for an infinite population.

Analysis of Variance↗

Discriminant analysis using the unweighted sum of binary variables: a comparison of model selection methods.

Many clinical decision-making rules are equivalent to linear discriminant functions that involve the unweighted sum of binary variables (SBV). We briefly consider the geometry of this restriction and then propose a number of methods for forward stepwise selection of SBV models. Using a simulation study, we compare the performance of these methods under a wide range of plausible conditions and show that no single method is uniformly superior for selecting models of a fixed size. Factors of general importance in relative method performance are the ratio of sample size to the number of candidate variables and the class-conditional moment structure of the data. We conclude by offering some practical strategies for SBV model construction.

Computer Simulation↗

A worst-case optimal parameter selection model of cancer chemotherapy.

An optimal parameter selection model of cancer chemotherapy in which two system parameters are unknown is formulated as a worst-case optimal parameter selection model. The model assumes that the unknown parameters lie within a known set. The system constraints must be satisfied over this entire set, and the objective function minimized in the worst case. The continuous dependence of the objective function and the system constraints upon the unknown parameters can be removed, making a numerical solution tractable. For the data considered it is proven that a cure is impossible no matter what the values of the unknown parameters in the parameter set. The optimal policy is shown to be relatively low dose intensity for the majority of the treatment, with the remaining drug delivered towards the end of the treatment interval.

Humans↗

How carcinogens (or telomere dysfunction) induce genetic instability: associated-selection model.

Carcinogens induce carcinogen-specific genetic instability (defects in DNA repair). According to the 'direct-selection' model, defects in DNA repair per se provide an immediate growth advantage. According to the 'associated-selection' model, carcinogens merely select for cells with adaptive mutations. Like any mutations, adaptive mutations occur predominantly in genetically unstable cells. The 'associated-selection' model predicts that carcinogen-driven selection minimizes cytotoxic but maximizes mutagenic effects of carcinogens. A purely mutagenic (neither cytotoxic, nor cytostatic) environment will favor effective DNA repair, whereas any growth-limiting conditions (telomerase deficiency, anticancer drugs) will select for genetically unstable cells. Genetic instability is a postmark of selective pressure rather than a hallmark of cancer per se. Once selected, genetic instability facilitates the development of resistance to any other growth-limiting conditions. As an example, a putative link between prior exposure to carcinogens and the ability to develop a telomerase-independent growth is discussed.

Carcinogens↗

Model selection in non-nested hidden Markov models for ion channel gating.

An important task in the application of Markov models to the analysis of ion channel data is the determination of the correct gating scheme of the ion channel under investigation. Some prior knowledge from other experiments can reduce significantly the number of possible models. If these models are standard statistical procedures nested like likelihood ratio testing, provide reliable selection methods. In the case of non-nested models, information criteria like AIC, BIC, etc., are used. However, it is not known if any of these criteria provide a reliable selection method and which is the best one in the context of ion channel gating. We provide an alternative approach to model selection in the case of non-nested models with an equal number of open and closed states. The models to choose from are embedded in a properly defined general model. Therefore, we circumvent the problems of model selection in the non-nested case and can apply model selection procedures for nested models.

Animals↗

Tests and model selection for the general growth curve model.

The model considered here is a generalized multivariate analysis of variance model useful especially for many types of growth curve problems including biological growth and technology substitution. It is defined as Yp x N = Xp x m tau m x r Ar x N + epsilon p x N, where tau is unknown, and X and A are known design matrices of ranks m less than p and r less than N, respectively. Furthermore, the columns of epsilon are independent p-variate normal with mean vector 0 and common covariance matrix sigma. In general, p is the number of time (or spatial) points observed on each of the N cases, (m - 1) is the degree of polynomial in time, and r is the number of groups. The main focus of this paper is the selection of models for the general growth curve model with regard to the covariance matrix sigma. Likelihood ratio tests and selection procedures based on sample reuse and predictions are proposed. Special emphasis is on the serial covariance structure for sigma, which has been shown to be quite important in the prediction of biological data and technology substitution data. One-population and K-population problems are considered. Some of the results are illustrated with two sets of biological data.

Animals↗

Model selection for extended quasi-likelihood models in small samples.

We develop a small sample criterion (AICc) for the selection of extended quasi-likelihood models. In contrast to the Akaike information criterion (AIC). AICc provides a more nearly unbiased estimator for the expected Kullback-Leibler information. Consequently, it often selects better models than AIC in small samples. For the logistic regression model, Monte Carlo results show that AICc outperforms AIC, Pregibon's (1979, Data Analytic Methods for Generalized Linear Models. Ph.D. thesis. University of Toronto) Cp*, and the Cp selection criteria of Hosmer et al. (1989, Biometrics 45, 1265-1270). Two examples are presented.

Age Factors↗

Best harmony, unified RPCL and automated model selection for unsupervised and supervised learning on Gaussian mixtures, three-layer nets and ME-RBF-SVM models.

After introducing the fundamentals of BYY system and harmony learning, which has been developed in past several years as a unified statistical framework for parameter learning, regularization and model selection, we systematically discuss this BYY harmony learning on systems with discrete inner-representations. First, we shown that one special case leads to unsupervised learning on Gaussian mixture. We show how harmony learning not only leads us to the EM algorithm for maximum likelihood (ML) learning and the corresponding extended KMEAN algorithms for Mahalanobis clustering with criteria for selecting the number of Gaussians or clusters, but also provides us two new regularization techniques and a unified scheme that includes the previous rival penalized competitive learning (RPCL) as well as its various variants and extensions that performs model selection automatically during parameter learning. Moreover, as a by-product, we also get a new approach for determining a set of 'supporting vectors' for Parzen window density estimation. Second, we shown that other special cases lead to three typical supervised learning models with several new results. On three layer net, we get (i) a new regularized ML learning, (ii) a new criterion for selecting the number of hidden units, and (iii) a family of EM-like algorithms that combines harmony learning with new techniques of regularization. On the original and alternative models of mixture-of-expert (ME) as well as radial basis function (RBF) nets, we get not only a new type of criteria for selecting the number of experts or basis functions but also a new type of the EM-like algorithms that combines regularization techniques and RPCL learning for parameter learning with either least complexity nature on the original ME model or automated model selection on the alternative ME model and RBF nets. Moreover, all the results for the alternative ME model are also applied to other two popular nonparametric statistical approaches, namely kernel regression and supporting vector machine. Particularly, not only we get an easily implemented approach for determining the smoothing parameter in kernel regression, but also we get an alternative approach for deciding the set of supporting vectors in supporting vector machine.

Algorithms↗

The application of sample selection models to outcomes research: the case of evaluating the effects of antidepressant therapy on resource utilization.

Non-randomized studies of treatment effects have come under criticism because of their failure to control for potential biases introduced by unobserved variables correlated with treatment selection and outcomes. This paper describes the basic concepts of sample selection models--a technique used widely in the economics evaluation literature for nearly two decades--and discusses the potential role of these models in outcomes research. In addition, it presents a case study of the application of the sample selection modelling approach to evaluation of the effects of antidepressant therapies on medical expenditures for physician services. This case study presents empirical comparisons of alternative model specifications and discusses practical issues in evaluation of sample selection models. We demonstrate that, in this particular case, sample selection models yield very different conclusions regarding treatment effects than traditional ordinary least squares regression.

Antidepressive Agents↗

A general asymptotic property of two-locus selection models.

It is shown that any two-locus, two-allele model of selection with constant fitnesses has at least one polymorphic equilibrium for which the linkage association measure, D, is arbitrarily close to zero for large enough recombination, R. As R----+/- infinity, D----0 in such a way that the product l = RD----a non-zero finite constant. There may be 1, 3, or 5 distinct asymptotic equilibria, depending upon fitness parameters.

Alleles↗

Bayes factor of model selection validates FLMP.

The fuzzy logical model of perception (FLMP; Massaro, 1998) has been extremely successful at describing performance across a wide range of ecological domains as well as for a broad spectrum of individuals. An important issue is whether this descriptive ability is theoretically informative or whether it simply reflects the model's ability to describe a wider range of possible outcomes. Previous tests and contrasts of this model with others have been adjudicated on the basis of both a root mean square deviation (RMSD) for goodness-of-fit and an observed RMSD relative to a benchmark RMSD if the model was indeed correct. We extend the model evaluation by another technique called Bayes factor (Kass & Raftery, 1995; Myung & Pitt, 1997). The FLMP maintains its significant descriptive advantage with this new criterion. In a series of simulations, the RMSD also accurately recovers the correct model under actual experimental conditions. When additional variability was added to the results, the models continued to be recoverable. In addition to its descriptive accuracy, RMSD should not be ignored in model testing because it can be justified theoretically and provides a direct and meaningful index of goodness-of-fit. We also make the case for the necessity of free parameters in model testing. Finally, using Newton's law of universal gravitation as an analogy, we argue that it might not be valid to expect a model's fit to be invariant across the whole range of possible parameter values for the model. We advocate that model selection should be analogous to perceptual judgment, which is characterized by the optimal use of multiple sources of information (e.g., the FLMP). Conclusions about models should be based on several selection criteria.

Bayes Theorem↗

Cooperative selection of movements: the optimal selection model.

How one selects a movement when faced with alternative ways of doing a task is a central problem in human motor control. Moving the fingertip a short distance can be achieved with any of an infinite number of combinations of knuckle, wrist, elbow, shoulder, and hip movements. The question therefore arises: how is a unique combination chosen? In our model, choice is achieved by consideration of the similarity between the task requirements and the optimal biomechanical performance of each limb segment. Two variants of the model account for the movements that are selected when subjects freely oscillate the fingertip and when they tap against an obstacle. An important feature of both is that the impulse of collision with an obstacle (as in drumming with the hand or tapping with the finger) is assumed to be controlled in part by aiming for a point beyond the surface being struck. Thus, a force-related control variable may be represented and controlled spatially.

Arm↗

Bayesian model selection: analysis of a survival model with a surviving fraction.

We describe a methodology for model comparison in a Bayesian framework as applied to survival with a surviving fraction. This is illustrated using a case study of a randomized and controlled clinical trial investigating time until recurrence of depression. Posterior distributions are simulated using Metropolis-within-Gibbs Markov chain methods. Models reflecting the effects of covariates on the log odds of being in the surviving fraction, the log of the hazard rate, as well as both and neither are compared. Bayes factors for comparing the models are obtained by using the bridge sampling method of calculating normalizing constants.

Algorithms↗

Computing minimum description length for robust linear regression model selection.

A minimum description length (MDL) and stochastic complexity approach for model selection in robust linear regression is studied in this paper. Computational aspects and implementation of this approach to practical problems are the focuses of the study. Particularly, we provide both algorithms and a package of S language programs for computing the stochastic complexity and proceeding with the associated model selection. A simulation study is then presented for illustration and comparing the MDL approach with the commonly used AIC and BIC methods. Finally, an application is given to a physiological study of triathlon athletes.

Algorithms↗

Counting probability distributions: differential geometry and model selection.

A central problem in science is deciding among competing explanations of data containing random errors. We argue that assessing the "complexity" of explanations is essential to a theoretically well-founded model selection procedure. We formulate model complexity in terms of the geometry of the space of probability distributions. Geometric complexity provides a clear intuitive understanding of several extant notions of model complexity. This approach allows us to reconceptualize the model selection problem as one of counting explanations that lie close to the "truth." We demonstrate the usefulness of the approach by applying it to the recovery of models in psychophysics.

Models, Theoretical↗