Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Model Selection Based on Minimum Description Length.

We introduce the minimum description length (MDL) principle, a general principle for inductive inference based on the idea that regularities (laws) underlying data can always be used to compress data. We introduce the fundamental concept of MDL, called the stochastic complexity, and we show how it can be used for model selection. We briefly compare MDL-based model selection to other approaches and we informally explain why we may expect MDL to give good results in practical applications. Copyright 2000 Academic Press.

Journal Article↗

Selection models and pattern-mixture models to analyse longitudinal quality of life data subject to drop-out.

Longitudinally observed quality of life data with large amounts of drop-out are analysed. First we used the selection modelling framework, frequently used with incomplete studies. An alternative method consists of using pattern-mixture models. These are also straightforward to implement, but result in a different set of parameters for the measurement and drop-out mechanisms. Since selection models and pattern-mixture models are based upon different factorizations of the joint distribution of measurement and drop-out mechanisms, comparing both models concerning, for example, treatment effect, is a useful form of a sensitivity analysis.

Aged↗

A study of early stopping and model selection applied to the papermaking industry.

This paper addresses the issues of neural network model development and maintenance in the context of a complex task taken from the papermaking industry. In particular, it describes a comparison study of early stopping techniques and model selection, both to optimise neural network models for generalisation performance. The results presented here show that early stopping via use of a Bayesian model evidence measure is a viable way of optimising performance while also making maximum use of all the data. In addition, they show that ten-fold cross-validation performs well as a model selector and as an estimator of prediction accuracy. These results are important in that they show how neural network models may be optimally trained and selected for highly complex industrial tasks where the data are noisy and limited in number.

Algorithms↗

A selection model for motion processing in area MT of primates.

A computational model for motion processing in area MT is presented that is based on the observed response properties of cortical neurons and is consistent with the visual perception of partially occluded and transparent moving stimuli. In contrast to models of motion processing that assume spatial continuity and fail to compute the correct velocity for these visual stimuli, our model produces a distributed segmentation of the image into disjoint patches that represent distinct objects moving with common velocities. A key element in the model is the selection of regions of the visual field where the velocity estimates are most reliable. The processing units in the motion model that perform the selection have nonclassical receptive fields similar to those observed in area MT (Allman et al., 1985). The psychophysical responses of the model to coherently moving random dots and transparent plaid gratings are similar to those observed in primates.

Animals↗

Use of fuzzy set theory to extend Dhawan's journal selection model: ranking the biomedical informatics serials.

OBJECTIVE: Experts disagree on the parameters to use to identify the "best" serials within a scientific field. The author set out to develop an extension to Dhawan's journal selection model for ranking serials in any scientific field. METHODS: Comparison of three different instantiations of Dhawan's model were used to rank thirty-four biomedical informatics serials. RESULTS: The first instantiation of Dhawan's model identified seven serials and divided them into two groups. The second instantiation of Dhawan's model identified twelve serials and separated them into two groups. Using fuzzy set theory the new extended model produced a rank ordered list of the top twelve biomedical informatics serials. CONCLUSIONS: Use of fuzzy set theory to assign set membership and combine data in Dhawan's journal selection model allows one to: (1) eliminate the need to determine arbitrary cutoff points for inclusion of serials within each of Dhawan's evaluation criteria categories, (2) combine data from disparate sources, and (3) obtain a rank-ordered list of the biomedical informatics serials rather than simply identifying a set of the "top" serials. Such a ranked list provides librarians and researchers alike with the information necessary to help them make their biomedical informatics serial selection decisions based on objective, quantifiable data.

Abstracting and Indexing↗

Model selection in neural networks.

In this article, we examine how model selection in neural networks can be guided by statistical procedures such as hypothesis tests, information criteria and cross validation. The application of these methods in neural network models is discussed, paying attention especially to the identification problems encountered. We then propose five specification strategies based on different statistical procedures and compare them in a simulation study. As the results of the study are promising, it is suggested that a statistical analysis should become an integral part of neural network modeling.

Journal Article↗

A numerical solution to the equilibria of the two-locus two-allele selection model.

Examination of the equilibria of the standard two-locus two-allele selection model leads to the construction of a polynomial with coefficients derived from selective values in the genotypic fitness matrix. This polynomial can be partially factored algebraically and numerical techniques are available to extract the roots of the remainder. Each root provides a possible value of the disequilibrium coefficient and the gametic frequencies at equilibrium, and these can be readily checked for stability.

Alleles↗

Influence of arterial vs. venous sampling site on nicotine tolerance model selection and parameter estimation.

In this modeling study we utilize previously published nicotine pharmacokinetic (PK) and pharmacodynamic (PD, heart rate) data to investigate the influence of PK sampling site (venous vs. arterial) on the selection of a specific PD tolerance model and estimation of its parameters. We describe a general model for tolerance which includes as special cases feedback (TF), and kinetic based tolerance (TK) models. A TK model has arterial plasma drug concentrations (Ca) driving (hypothetical) effect (Ce) and antagonist (Cm) site concentrations, which drive a non-feedback effect (Enf): tolerance depends on the relative rate of equilibration of Ce and Cm with Ca. The TF model adds feedback which makes tolerance depend on Enf, not just on drug kinetics for nicotine. The arterial-sampling-analysis (PKPDa) has Ca driving Ce and Cm. The venous-sampling-analysis (PKPDv) does the same but estimates Ca from venous data by means of deconvolution. A TF model (with Cm = Ce) was always selected in the PKPDa. According to this model tolerance developed rapidly with a median half-life of 6.6 min, and median decrease of effect due to tolerance of 31%. Different variants of the TF or TK models were selected in the PKPDv. Parameter estimates for PKPDv show higher variability, and, for the TF model, lower rate and extent of tolerance development and threefold increase in EC50. The study shows that (i) TF models are more appropriate than TK models to describe nicotine effect data, (ii) venous sampling may lead to incorrect model selection and inaccurate and imprecise parameter estimation in respect to arterial sampling, and (iii) arterial sampling should be preferred for accurate (non-steady-state) PD modeling.

Arteries↗

Bayesian methods for quantitative trait loci mapping based on model selection: approximate analysis using the Bayesian information criterion.

We describe an approximate method for the analysis of quantitative trait loci (QTL) based on model selection from multiple regression models with trait values regressed on marker genotypes, using a modification of the easily calculated Bayesian information criterion to estimate the posterior probability of models with various subsets of markers as variables. The BIC-delta criterion, with the parameter delta increasing the penalty for additional variables in a model, is further modified to incorporate prior information, and missing values are handled by multiple imputation. Marginal probabilities for model sizes are calculated, and the posterior probability of nonzero model size is interpreted as the posterior probability of existence of a QTL linked to one or more markers. The method is demonstrated on analysis of associations between wood density and markers on two linkage groups in Pinus radiata. Selection bias, which is the bias that results from using the same data to both select the variables in a model and estimate the coefficients, is shown to be a problem for commonly used non-Bayesian methods for QTL mapping, which do not average over alternative possible models that are consistent with the data.

Alleles↗

Homology modeling using parametric alignment ensemble generation with consensus and energy-based model selection.

The accuracy of a homology model based on the structure of a distant relative or other topologically equivalent protein is primarily limited by the quality of the alignment. Here we describe a systematic approach for sequence-to-structure alignment, called 'K*Sync', in which alignments are generated by dynamic programming using a scoring function that combines information on many protein features, including a novel measure of how obligate a sequence region is to the protein fold. By systematically varying the weights on the different features that contribute to the alignment score, we generate very large ensembles of diverse alignments, each optimal under a particular constellation of weights. We investigate a variety of approaches to select the best models from the ensemble, including consensus of the alignments, a hydrophobic burial measure, low- and high-resolution energy functions, and combinations of these evaluation methods. The effect on model quality and selection resulting from loop modeling and backbone optimization is also studied. The performance of the method on a benchmark set is reported and shows the approach to be effective at both generating and selecting accurate alignments. The method serves as the foundation of the homology modeling module in the Robetta server.

Amino Acid Sequence↗

Model selection in spatio-temporal electromagnetic source analysis.

Several methods [model selection procedures (MSPs)] to determine the number of sources in electroencephalogram (EEG) and magnetoencphalogram (MEG) data have previously been investigated in an instantaneous analysis. In this paper, these MSPs are extended to a spatio-temporal analysis if possible. It is seen that the residual variance (RV) tends to overestimate the number of sources. The Akaike information criterion (AIC) and the Wald test on amplitudes (WA) and the Wald test on locations (WL) have the highest probabilities of selecting the correct number of sources. The WA has the advantage that it offers the opportunity to test which source is active at which time sample.

Algorithms↗

Maintenance of handedness polymorphism in humans: a frequency-dependent selection model.

Frequency-dependent selection is an important process in the maintenance of genetic variation in fitness. In humans, it has been proposed that the polymorphism of handedness is maintained by negative frequency-dependent selection, through a strategic advantage of left-handers in fighting interactions. Using simple mathematical models, we explore: (1) whether it is possible to predict the range of left-handedness frequencies observed in human populations by the frequency and the violence of fighting interactions; (2) the consequences of the sex differences in the probability of transmission of hand preference to offspring. We show that a wide range of values of the frequency of left-handers can be obtained with realistic changes of the parameters values. Our models reinforce the idea that negative frequency-dependence may have played a role in maintaining left-handedness in human populations, and provide further support for the importance of fighting interactions in the evolution of hand preference. Moreover, they suggest an explanation for the occurrence of left-handedness among women in this context, namely an indirect selective advantage through their male offspring.

Aggression↗

Model selection for quantitative trait locus analysis in polyploids.

Over the years, substantial gains have been made in locating regions of agricultural genomes associated with characteristics, diseases, and agroeconomic traits. These gains have relied heavily on the ability to statistically estimate the association between DNA markers and regions of a genome (quantitative trait loci or QTL) related to a particular trait. The majority of these advances have focused on diploid species, even though many important agricultural crops are, in fact, polyploid. The purpose of our work is to initiate an algorithmic approach for model selection and QTL detection in polyploid species. This approach involves the construction of all possible chromosomal configurations (models) that may result in a gamete, model reduction based on estimation of marker dosage from progeny data, and lastly model selection. While simplified for initial explanation, our approach has demonstrated itself to be extendible to many breeding schemes and less restricted settings.

Agriculture↗

Dependence of in vitro-in vivo correlation analysis acceptability on model selections.

The objectives of this study were to assess the influence of pharmacokinetic and dissolution model selection on the success of in vitro-in vivo correlation (IVIVC) analysis of fast-, medium-, and slow-dissolving metoprolol tartrate immediate-release tablet formulations. Several different compartmental models were fit to the fast formulation plasma data. Three candidate models with the best fits were ranked 1, 2, and 3 and used to predict AUC and Cmax of the medium and slow formulations. Acceptability of each model to predict the medium and slow formulations was determined using +/- 20% as the limit for acceptable relative prediction error. When the best dissolution models were used, models 1 and 2 each failed to adequately predict Cmax for slow formulation (-26.8% error for model 1 and -20.4% error for model 2). However, the less appropriate model 3 adequately predicted Cmax for slow formulation (-15.1% error). The selection of the dissolution model also determined the outcome of IVIVC analysis, again with a less appropriate model resulting in successful prediction. When the Weibull function was used to characterize dissolution, model 2 failed to adequately predict Cmax for slow formulation (-20.4% error); however, model 2 adequately predicted Cmax for slow formulation when dissolution was characterized using the poorer fitting first-order model (-14.4% error). These results indicate that the success or failure of external validation of these metoprolol tartrate tablets was dependent on the pharmacokinetic and the dissolution models employed. Considering the role of subjectivity in identifying pharmacokinetic and dissolution models, these findings suggest a need to develop objective criteria to identify models a priori to IVIVC analysis.

Biological Availability↗

Model selection and the estimation of odds ratios in the presence of extraneous factors.

This paper deals with model selection and the estimation of odds ratios from cross-classified frequencies in the presence of extraneous factors. The odds ratio is estimated in different ways dependent on whether the extraneous factor is modelled as an effect modifier, a confounder, or neither. Routinely this choice is based on statistical tests of null hypotheses. By contrast, we propose selection of the model which is estimated to maximize the accuracy of the estimator of the odds ratio, on average. We demonstrate how a non-parametric bootstrap method can be used to carry out the selection, and illustrate the methodology using an example on use of oral contraceptives and myocardial infarction.

Adult↗

The maintenance of single-locus polymorphism. I. Numerical studies of a viability selection model.

The ability of viability selection to maintain single-locus polymorphism is investigated with two models in which the population is bombarded with a series of mutations with random fitnesses. In the first model, the population is allowed to reach equilibrium before mutation resumes; in the second the iterations and mutation occur simultaneously. Monte Carlo simulations of these models show that viability selection is easily able to maintain stable 6- or 7-allele polymorphisms and that monomorphisms and diallelic polymorphisms are uncommon. The question of how monomorphisms arise is also discussed.

Alleles↗

A simple clustering technique to improve QSAR model selection and predictivity: application to a receptor independent 4D-QSAR analysis of cyclic urea derived inhibitors of HIV-1 protease.

A training set of 50 tetrahydropyrimidine-2-one based inhibitors of HIV-1 protease, for which the -log K(i) values were measured, was used to construct receptor independent 4D-QSAR models. A novel clustering technique was employed to facilitate and improve model selection as well as test set predictions. Following the manifold model theory, five unique models were chosen by the clustering algorithm (q(2) = 0.81-0.84). The models were used to map the atom type morphology of the inhibitor binding site of HIV-1 protease as well as to predict the potencies (-log K(i)) of 10 test set compounds. The rank-difference correlation coefficient was used to evaluate the quality of the test set predictions, which was improved from 0.39 to 0.68 when the clustering technique was applied. The set of five models, collectively, identify the important binding characteristics of the HIV protease receptor site. This study demonstrates that the selected simple clustering technique provides a discrete algorithm for model selection, as well as improving the quality of test set, or unknown, compound prediction as determined by the rank-difference correlation coefficient.

Algorithms↗