Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Bayes factor of model selection validates FLMP.

The fuzzy logical model of perception (FLMP; Massaro, 1998) has been extremely successful at describing performance across a wide range of ecological domains as well as for a broad spectrum of individuals. An important issue is whether this descriptive ability is theoretically informative or whether it simply reflects the model's ability to describe a wider range of possible outcomes. Previous tests and contrasts of this model with others have been adjudicated on the basis of both a root mean square deviation (RMSD) for goodness-of-fit and an observed RMSD relative to a benchmark RMSD if the model was indeed correct. We extend the model evaluation by another technique called Bayes factor (Kass & Raftery, 1995; Myung & Pitt, 1997). The FLMP maintains its significant descriptive advantage with this new criterion. In a series of simulations, the RMSD also accurately recovers the correct model under actual experimental conditions. When additional variability was added to the results, the models continued to be recoverable. In addition to its descriptive accuracy, RMSD should not be ignored in model testing because it can be justified theoretically and provides a direct and meaningful index of goodness-of-fit. We also make the case for the necessity of free parameters in model testing. Finally, using Newton's law of universal gravitation as an analogy, we argue that it might not be valid to expect a model's fit to be invariant across the whole range of possible parameter values for the model. We advocate that model selection should be analogous to perceptual judgment, which is characterized by the optimal use of multiple sources of information (e.g., the FLMP). Conclusions about models should be based on several selection criteria.

Bayes Theorem↗

Cooperative selection of movements: the optimal selection model.

How one selects a movement when faced with alternative ways of doing a task is a central problem in human motor control. Moving the fingertip a short distance can be achieved with any of an infinite number of combinations of knuckle, wrist, elbow, shoulder, and hip movements. The question therefore arises: how is a unique combination chosen? In our model, choice is achieved by consideration of the similarity between the task requirements and the optimal biomechanical performance of each limb segment. Two variants of the model account for the movements that are selected when subjects freely oscillate the fingertip and when they tap against an obstacle. An important feature of both is that the impulse of collision with an obstacle (as in drumming with the hand or tapping with the finger) is assumed to be controlled in part by aiming for a point beyond the surface being struck. Thus, a force-related control variable may be represented and controlled spatially.

Arm↗

Model selection and parameter estimation for ion channel recordings with an application to the K+ outward-rectifier in barley leaf.

We present a statistical method, and its accompanying algorithms, for the selection of a mathematical model of the gating mechanism of an ion channel and for the estimation of the parameters of this model. The method assumes a hidden Markov model that incorporates filtering, colored noise and state-dependent white excess noise for the recorded data. The model selection and parameter estimation are performed via a Bayesian approach using Markov chain Monte Carlo. The method is illustrated by its application to single-channel recordings of the K(+) outward-rectifier in barley leaf.

Algorithms↗

Bayesian model selection: analysis of a survival model with a surviving fraction.

We describe a methodology for model comparison in a Bayesian framework as applied to survival with a surviving fraction. This is illustrated using a case study of a randomized and controlled clinical trial investigating time until recurrence of depression. Posterior distributions are simulated using Metropolis-within-Gibbs Markov chain methods. Models reflecting the effects of covariates on the log odds of being in the surviving fraction, the log of the hazard rate, as well as both and neither are compared. Bayes factors for comparing the models are obtained by using the bridge sampling method of calculating normalizing constants.

Algorithms↗

Computing minimum description length for robust linear regression model selection.

A minimum description length (MDL) and stochastic complexity approach for model selection in robust linear regression is studied in this paper. Computational aspects and implementation of this approach to practical problems are the focuses of the study. Particularly, we provide both algorithms and a package of S language programs for computing the stochastic complexity and proceeding with the associated model selection. A simulation study is then presented for illustration and comparing the MDL approach with the commonly used AIC and BIC methods. Finally, an application is given to a physiological study of triathlon athletes.

Algorithms↗

Generalized additive selection models for the analysis of studies with potentially nonignorable missing outcome data.

Rotnitzky, Robins, and Scharfstein (1998, Journal of the American Statistical Association 93, 1321-1339) developed a methodology for conducting sensitivity analysis of studies in which longitudinal outcome data are subject to potentially nonignorable missingness. In their approach, they specify a class of fully parametric selection models, indexed by a non- or weakly identified selection bias function that indicates the degree to which missingness depends on potentially unobservable outcomes. Estimation of the parameters of interest proceeds by varying the selection bias function over a range considered plausible by subject-matter experts. In this article, we focus on cross-sectional, univariate outcome data and extend their approach to a class of semiparametric selection models, using generalized additive restrictions. We propose a backfitting algorithm to estimate the parameters of the generalized additive selection model. For estimation of the mean outcome, we propose three types of estimating functions: simple inverse weighted, doubly robust, and orthogonal. We present the results of a data analysis and a simulation study.

Acquired Immunodeficiency Syndrome↗

Counting probability distributions: differential geometry and model selection.

A central problem in science is deciding among competing explanations of data containing random errors. We argue that assessing the "complexity" of explanations is essential to a theoretically well-founded model selection procedure. We formulate model complexity in terms of the geometry of the space of probability distributions. Geometric complexity provides a clear intuitive understanding of several extant notions of model complexity. This approach allows us to reconceptualize the model selection problem as one of counting explanations that lie close to the "truth." We demonstrate the usefulness of the approach by applying it to the recovery of models in psychophysics.

Models, Theoretical↗

The multifocused faculty selection model: a design for hiring the best faculty.

The multifocused faculty selection model is designed to assist the faculty search committee in two ways. First, this model organizes the stages for obtaining information about the applicant's abilities related to faculty performance. Second, the model provides a consistent method for faculty selection as the membership of the department's search committee changes. The applicant selection techniques used in this model include descriptive interviewing, computer assisted data, and the teaching demonstration.

Faculty, Nursing↗

General kin selection models for genetic evolution of sib altruism in diploid and haplodiploid species.

A population genetic approach is presented for general analysis and comparison of kin selection models of sib and half-sib altruism. Nine models are described, each assuming a particular mode of inheritance, number of female inseminations, and Mendelian dominance of the altruist gene. In each model, the selective effects of altruism are described in terms of two general fitness functions, A(beta) and S(beta), giving respectively the expected fitness of an altruist and a nonaltruist as a function of the fraction of altruists beta in a given sibship. For each model, exact conditions are reported for stability at altruist and nonaltruist fixation. Under the Table 3 axions, the stability conditions may then be partially ordered on the basis of implications holding between pairs of conditions. The partial orderings are compared with predictions of the kin selection theory of Hamilton.

Biological Evolution↗

Reverse engineering galactose regulation in yeast through model selection.

We examine the application of statistical model selection methods to reverse-engineering the control of galactose utilization in yeast from DNA microarray experiment data. In these experiments, relationships among gene expression values are revealed through modifications of galactose sugar level and genetic perturbations through knockouts. For each gene variable, we select predictors using a variety of methods, taking into account the variance in each measurement. These methods include maximization of log-likelihood with Cp, AIC, and BIC penalties, bootstrap and cross-validation error estimation, and coefficient shrinkage via the Lasso.

Journal Article↗

A chance-selection model for cell differentiation.

A chance-selection model is proposed to explain cell differentiation. It is based on the general idea that stochasticity at the molecular level generates diversity in cell types whereas cell interactions impose a characteristic order on the developing embryo. In this model, gene expression depends on stochastic molecular interactions between transcriptional regulators and DNA. Random diffusion of these regulators along DNA causes differential gene expression in differentiating cells.The role of phosphorylation or dephosphorylation of transcriptional regulators, triggered by cell interactions, is to control their random diffusion and to stabilize stochastic gene expression in differentiated cells. This model is based on well documented phenomena: random diffusion of DNA binding molecules along DNA, phosphorylation or dephosphorylation of transcription factors by protein kinases or phosphatases and control of DNA binding of transcription factors through this latter process. The different explanatory powers of deterministic and stochastic models are discussed.

Journal Article↗

Selection models and pattern-mixture models for incomplete data with covariates.

Most models for incomplete data are formulated within the selection model framework. This paper studies similarities and differences of modeling incomplete data within both selection and pattern-mixture settings. The focus is on missing at random mechanisms and on categorical data. Point and interval estimation is discussed. A comparison of both approaches is done on side effects in a psychiatric study.

Biometry↗

Evolutionarily stable strategies in food selection models with fitness sets.

Most current models for optimal food selection apply to ecological and behavioural optimization. In this paper optimal food selection theory is extended to apply to evolutionary optimization. A general evolutionary model for optimal food selection must incorporate the concept of fitness sets--or that variables, changing as a result of natural selection in evolutionary time, cannot, in general, vary independently of each other. A "Charnov type" optimal food selection model with a fitness set is investigated, and evolutionarily stable strategy (ESS) solutions of the evolutionary variables (i.e., the efficiencies of using available food types) are found. From this analysis it follows that the relative frequency of various food types in the environment may, under specified conditions, influence the evolutionarily optimal diet. Secondly, the analysis demonstrates that a food type not in the optimal diet may, in evolutionary time, be added to this by becoming more abundant. Thirdly, it follows from the analysis that the ecological result of MacArthur and Pianka, that food types are worth eating even if there is competition for them, is not generally applicable when referring to an evolutionary time scale. Finally, it is pointed out that for the diet to be an ESS, it is necessary that the consumer's density is stable and that the consumer's population dynamics are subjected to some density-dependent factor.

Animals↗

Accounting for uncertainty in the tree topology has little effect on the decision-theoretic approach to model selection in phylogeny estimation.

Currently available methods for model selection used in phylogenetic analysis are based on an initial fixed-tree topology. Once a model is picked based on this topology, a rigorous search of the tree space is run under that model to find the maximum-likelihood estimate of the tree (topology and branch lengths) and the maximum-likelihood estimates of the model parameters. In this paper, we propose two extensions to the decision-theoretic (DT) approach that relax the fixed-topology restriction. We also relax the fixed-topology restriction for the Bayesian information criterion (BIC) and the Akaike information criterion (AIC) methods. We compare the performance of the different methods (the relaxed, restricted, and the likelihood-ratio test [LRT]) using simulated data. This comparison is done by evaluating the relative complexity of the models resulting from each method and by comparing the performance of the chosen models in estimating the true tree. We also compare the methods relative to one another by measuring the closeness of the estimated trees corresponding to the different chosen models under these methods. We show that varying the topology does not have a major impact on model choice. We also show that the outcome of the two proposed extensions is identical and is comparable to that of the BIC, Extended-BIC, and DT. Hence, using the simpler methods in choosing a model for analyzing the data is more computationally feasible, with results comparable to the more computationally intensive methods. Another outcome of this study is that earlier conclusions about the DT approach are reinforced. That is, LRT, Extended-AIC, and AIC result in more complicated models that do not contribute to the performance of the phylogenetic inference, yet cause a significant increase in the time required for data analysis.

Computational Biology↗

Molecular diagnosis. Classification, model selection and performance evaluation.

OBJECTIVES: We discuss supervised classification techniques applied to medical diagnosis based on gene expression profiles. Our focus lies on strategies of adaptive model selection to avoid overfitting in high-dimensional spaces. METHODS: We introduce likelihood-based methods, classification trees, support vector machines and regularized binary regression. For regularization by dimension reduction, we describe feature selection methods: feature filtering, feature shrinkage and wrapper approaches. In small sample-size situations efficient methods of data re-use are needed to assess the predictive power of a model. We discuss two issues in using cross-validation: the difference between in-loop and out-of-loop feature selection, and estimating model parameters in nested-loop cross-validation. RESULTS: Gene selection does not reduce the dimensionality of the model. Tuning parameters enable adaptive model selection. The feature selection bias is a common pitfall in performance evaluation. Model selection and performance evaluation can be combined by nested-loop cross-validation. CONCLUSIONS: Classification of microarrays is prone to overfitting. A rigorous and unbiased assessment of the predictive power of the model is a must.

Gene Expression Profiling↗

Pedigree analysis package (PAP) vs. MORGAN: model selection and hypothesis testing on a large pedigree.

The MORGAN package of programs is compared to a commonly used package, PAP, with respect to model selection in segregation analysis of a quantitative trait. MORGAN uses Monte Carlo Markov chain (MCMC) methods to estimate the likelihood, whereas both versions of PAP used employ an approximation to the likelihood for the mixed model. Comparisons are done by using results obtained from simulated data. All simulations were done on the same 232-member pedigree using data generated under each of several variations of models, which included different combinations of environmental, polygenic, and major gene components. PAP, version 4.0, and MORGAN gave similar results with respect to model selection for the majority of situations, suggesting that MCMC methods provide a computationally tractable approach for analysis of more complex models that cannot be analyzed by more direct computational methods. PAP, version 3.0, gave somewhat more disparate results compared with either PAP version 4.0 or MORGAN. Both MORGAN and the two versions of PAP confirmed that the major gene component is much easier to detect in the presence of some dominance. All three packages frequently falsely accepted the polygenic model when there was high residual heritability.

Computer Simulation↗

Evaluation of candidate gene effects for beef backfat via Bayesian model selection.

Candidate gene approaches provide tools for exploring and localizing causative genes affecting quantitative traits and the underlying variation may be better understood by determining the relative magnitudes of effects of their polymorphisms. Diacyglycerol O-acyltransferase 1 (DGAT1), fatty acid binding protein (heart) 3 (FABP3), growth hormone 1 (GH1), leptin (LEP) and thyroglobulin (TG) have been previously identified as genes contributing to genetic control of subcutaneous fat thickness (SFT) in beef cattle. In the present research, Bayesian model selection was used to evaluate effects of these five candidate genes by comparing competing non-nested models and treating candidate gene effects as either random or fixed. The analyses were implemented in SAS to simplify the programming and computation. Phenotypic data were gathered from a F(2) population of Wagyu x Limousin cattle. The five candidate genes had significant but varied effects on SFT in this population. Bayesian model selection identified the DGAT1 model as the one with the greatest model probability, whether candidate gene effects were considered random or fixed, and DGAT1 had the greatest additive effect on SFT. The SAS codes developed in the study are freely available and can be downloaded at: http://www.ansci.wsu.edu/programs/.

Adiposity↗

An overview of the posters presented at Watermatex 2000. III: Model selection and calibration/optimal experimental design.

This paper presents an overview of the posters presented in sessions 7 and 8 of the Watermatex 2000 conference. These posters present two aspects of modelling biological processes--model selection and calibration. Special attention is given to the papers on OED (Optimal Experimental Design), which is a method of optimising the data collection for model selection and calibration. The presence of these presentations at the conference highlights the continuing significance of modelling and stresses the requirement of improvements in modelling techniques. The papers provide some contribution to this end.

Calibration↗