Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Statistics versus statistical science in the regulatory process.

This paper reviews the established practice of providing evidence to regulatory authorities about the claimed properties (such as efficacy and safety) of new pharmaceutical products. The established conventions and procedures are contrasted with scientific concepts and principles. The following issues are discussed: (a) recruitment of subjects and its connection to treatment heterogeneity; (b) the measurement process and the handling of missing data; (c) data transformation and the use of generalized linear models; (d) model selection and model checking; (e) the 'cult of the single trial' and the use of prior information; and (f) hypothesis testing and the P-value culture.

Area Under Curve↗

A model for selecting assessment methods for evaluating medical students in African medical schools.

Introduction of more effective and standardized assessment methods for testing students' performance in Africa's medical institutions has been hampered by severe financial and personnel shortages. Nevertheless, some African institutions have recognized the problem and are now revising their medical curricula, and, therefore, their assessment methods. These institutions, and those yet to come, need guidance on selecting assessment methods so as to adopt models that can be sustained locally. The authors provide a model for selecting assessment methods for testing medical students' performance in African medical institutions. The model systematically evaluates factors that influence implementation of an assessment method. Six commonly used methods (the essay examinations, short-answer questions, multiple-choice questions, patient-based clinical examination, problem-based oral examination [POE], and objective structured clinical examination) are evaluated by scoring and weighting against performance, cost, suitability, and safety factors. In the model, the highest score identifies the most appropriate method. Selection of an assessment method is illustrated using two institutional models, one depicting an ideal situation in which the objective structured clinical examination was preferred, and a second depicting the typical African scenario in which the essay and short-answer-question examinations were best. The POE method received the highest score and could be recommended as the most appropriate for Africa's medical institutions, but POE assessments require changing the medical curricula to a problem-based learning approach. The authors' model is easy to understand and promotes change in the medical curriculum and method of student assessment.

Africa↗

Improving the quality of protein structure models by selecting from alignment alternatives.

BACKGROUND: In the area of protein structure prediction, recently a lot of effort has gone into the development of Model Quality Assessment Programs (MQAPs). MQAPs distinguish high quality protein structure models from inferior models. Here, we propose a new method to use an MQAP to improve the quality of models. With a given target sequence and template structure, we construct a number of different alignments and corresponding models for the sequence. The quality of these models is scored with an MQAP and used to choose the most promising model. An SVM-based selection scheme is suggested for combining MQAP partial potentials, in order to optimize for improved model selection. RESULTS: The approach has been tested on a representative set of proteins. The ability of the method to improve models was validated by comparing the MQAP-selected structures to the native structures with the model quality evaluation program TM-score. Using the SVM-based model selection, a significant increase in model quality is obtained (as shown with a Wilcoxon signed rank test yielding p-values below 10(-15)). The average increase in TMscore is 0.016, the maximum observed increase in TM-score is 0.29. CONCLUSION: In template-based protein structure prediction alignment is known to be a bottleneck limiting the overall model quality. Here we show that a combination of systematic alignment variation and modern model scoring functions can significantly improve the quality of alignment-based models.

Computer Simulation↗

Coinfection and superinfection in RNA virus populations: a selection-mutation model.

In this paper, we present a general selection-mutation model of evolution on a one-dimensional continuous fitness space. The formulation of our model includes both the classical diffusion approach to mutation process as well as an alternative approach based on an integral operator with a mutation kernel. We show that both approaches produce fundamentally equivalent results. To illustrate the suitability of our model, we focus its analytical study into its application to recent experimental studies of in vitro viral evolution. More specifically, these experiments were designed to test previous theoretical predictions regarding the effects of multiple infection dynamics (i.e., coinfection and superinfection) on the virulence of evolving viral populations. The results of these experiments, however, did not match with previous theory. By contrast, the model we present here helps to understand the underlying viral dynamics on these experiments and makes new testable predictions about the role of parameters such the time between successive infections and the growth rates of resident and invading populations.

Evolution, Molecular↗

Disequilibrium in two-locus mutation-selection balance models.

Equilibrium behavior of two-locus mutation-selection balance models is analyzed using perturbation techniques. The classical result of Haldane for one locus is shown to carry over to two loci, if fitnesses are replaced by marginal fitnesses. If the fitness of the double heterozygote is smaller than would be produced by a multiplicative model, as in additive or quantitative fitness models, the disequilibrium is negative--an excess of gametes with one rare allele. In this case the disequilibrium can be as large as one-half its maximum value possible, if the recombination rate is small, not greater than the strength of selection. If the fitness of the double heterozygote is larger than would be produced by a multiplicative model, the disequilibrium is positive, and is very small relative to its maximum value possible, even if the recombination rate is zero.

Heterozygote↗

Convergent and parallel evolution: a model illustrating selection, phylogeny and phenetic similarity.

A model is proposed which considers the structural relationships of body characteristics and their role in a concept wherein the phylogenetic relationship of the organisms under study is interpreted as constituting degrees of convergent or parallel evolution. The model also accounts for the relationships between selection pressures, phylogeny, and phenetic expression. The phenomena of convergent and parallel evolution are based upon the observations of similar characteristics, the geometric concept, magnitude and similarity of selection pressures, and the phylogenetic relationship of the groups in question.

Animals↗

T cell repertoire formation displays characteristics of qualitative models of thymic selection.

The use of T cell receptor elements varies between mouse strains, reflecting a balance between positive and negative selection. The presence of H-2E biases V alpha and V beta usage through major histocompatibility class II isotype preferences of V elements, and mammary tumor virus-dependent, negative selection. Quantitative models of thymic selection predict that negative selection equates to 'excess' positive selection, whereas qualitative models suggest that positive and negative selection are opposing forces. This report attempts to distinguish between the models by assessing whether, at the level of the T cell repertoire, positive and negative selection have quantitative or qualitative characteristics. The data show that the effect of bearing V alpha and V beta regions which are both preferentially (or negatively) selected in the presence of H-2E is additive or synergistic, whilst positive stimuli counteract negative ones. The data thus provide support for qualitative models of thymic selection.

Animals↗

Comparison among some models of orientation selectivity.

Several models exist for explaining primary visual cortex (V1) orientation tuning. The modified feedforward model (MFM) and the recurrent model (RM) are major examples. We have implemented these two models, at the same level of detail, alongside a few newer variations, and thoroughly compared their receptive-field structures. We found that antiphase inhibition in the MFM enhances both spatial phase information and orientation tuning, producing well-tuned simple cells. This remains true for a newer version of the MFM that incorporates untuned complex-cell inhibition. In contrast, when the recurrent connections in the RM are strong enough to produce typical V1 orientation tuning, they also eliminate spatial phase information, making the cells complex. Introducing phase specificity into the connections of the RM (as done in an original version of the RM) can make the cells phase sensitive, but the cells show an incorrect 90 degrees peak shift of orientation tuning under opposite contrast signs. An inhibition-dominant version of the RM can generate well-tuned cells across the simple-complex spectrum, but it predicts that the net effect of cortical interactions is to suppress feedforward excitation across all orientations in simple cells. Finally, adding antiphase inhibition used in the MFM into the RM produces a most general model. We call this new model the modified recurrent model (MRM) and show that this model can also produce well-tuned cells throughout the simple-complex spectrum. Unlike the inhibition-dominant RM, the MRM is consistent with data from cat V1, suggesting that the net effect of cortical interactions is to boost simple cell responses at the preferred orientation. These results suggest that the MFM is well suited for explaining orientation tuning in simple cells, whereas the standard RM is for complex cells. The assignment of the RM to complex cells also avoids conflicts between the RM and the experiments of cortical inactivation (done on simple cells) and the spatial-frequency dependency of orientation tuning (found in simple cells). Because orientation-tuned V1 cells show a continuum of simple- to complex-cell behavior, the MRM provides the best description of V1 data.

Animals↗

Statistical limitations in functional neuroimaging. I. Non-inferential methods and statistical models.

Functional neuroimaging (FNI) provides experimental access to the intact living brain making it possible to study higher cognitive functions in humans. In this review and in a companion paper in this issue, we discuss some common methods used to analyse FNI data. The emphasis in both papers is on assumptions and limitations of the methods reviewed. There are several methods available to analyse FNI data indicating that none is optimal for all purposes. In order to make optimal use of the methods available it is important to know the limits of applicability. For the interpretation of FNI results it is also important to take into account the assumptions, approximations and inherent limitations of the methods used. This paper gives a brief overview over some non-inferential descriptive methods and common statistical models used in FNI. Issues relating to the complex problem of model selection are discussed. In general, proper model selection is a necessary prerequisite for the validity of the subsequent statistical inference. The non-inferential section describes methods that, combined with inspection of parameter estimates and other simple measures, can aid in the process of model selection and verification of assumptions. The section on statistical models covers approaches to global normalization and some aspects of univariate, multivariate, and Bayesian models. Finally, approaches to functional connectivity and effective connectivity are discussed. In the companion paper we review issues related to signal detection and statistical inference.

Bayes Theorem↗

Pairwise multiple comparisons: a model comparison approach versus stepwise procedures.

Researchers in the behavioural sciences have been presented with a host of pairwise multiple comparison procedures that attempt to obtain an optimal combination of Type I error control, power, and ease of application. However, these procedures share one important limitation: intransitive decisions. Moreover, they can be characterized as a piecemeal approach to the problem rather than a holistic approach. Dayton has recently proposed a new approach to pairwise multiple comparisons testing that eliminates intransitivity through a model selection procedure. The present study compared the model selection approach (and a protected version) with three powerful and easy-to-use stepwise multiple comparison procedures in terms of the proportion of times that the procedure identified the true pattern of differences among a set of means across several one-way layouts. The protected version of the model selection approach selected the true model a significantly greater proportion of times than the stepwise procedures and, in most cases, was not affected by variance heterogeneity and non-normality.

Humans↗

Comparative performance of Bayesian and AIC-based measures of phylogenetic model uncertainty.

Reversible-jump Markov chain Monte Carlo (RJ-MCMC) is a technique for simultaneously evaluating multiple related (but not necessarily nested) statistical models that has recently been applied to the problem of phylogenetic model selection. Here we use a simulation approach to assess the performance of this method and compare it to Akaike weights, a measure of model uncertainty that is based on the Akaike information criterion. Under conditions where the assumptions of the candidate models matched the generating conditions, both Bayesian and AIC-based methods perform well. The 95% credible interval contained the generating model close to 95% of the time. However, the size of the credible interval differed with the Bayesian credible set containing approximately 25% to 50% fewer models than an AIC-based credible interval. The posterior probability was a better indicator of the correct model than the Akaike weight when all assumptions were met but both measures performed similarly when some model assumptions were violated. Models in the Bayesian posterior distribution were also more similar to the generating model in their number of parameters and were less biased in their complexity. In contrast, Akaike-weighted models were more distant from the generating model and biased towards slightly greater complexity. The AIC-based credible interval appeared to be more robust to the violation of the rate homogeneity assumption. Both AIC and Bayesian approaches suggest that substantial uncertainty can accompany the choice of model for phylogenetic analyses, suggesting that alternative candidate models should be examined in analysis of phylogenetic data. [AIC; Akaike weights; Bayesian phylogenetics; model averaging; model selection; model uncertainty; posterior probability; reversible jump.].

Bayes Theorem↗

Model weights and the foundations of multimodel inference.

Statistical thinking in wildlife biology and ecology has been profoundly influenced by the introduction of AIC (Akaike's information criterion) as a tool for model selection and as a basis for model averaging. In this paper, we advocate the Bayesian paradigm as a broader framework for multimodel inference, one in which model averaging and model selection are naturally linked, and in which the performance of AIC-based tools is naturally evaluated. Prior model weights implicitly associated with the use of AIC are seen to highly favor complex models: in some cases, all but the most highly parameterized models in the model set are virtually ignored a priori. We suggest the usefulness of the weighted BIC (Bayesian information criterion) as a computationally simple alternative to AIC, based on explicit selection of prior model probabilities rather than acceptance of default priors associated with AIC. We note, however, that both procedures are only approximate to the use of exact Bayes factors. We discuss and illustrate technical difficulties associated with Bayes factors, and suggest approaches to avoiding these difficulties in the context of model selection for a logistic regression. Our example highlights the predisposition of AIC weighting to favor complex models and suggests a need for caution in using the BIC for computing approximate posterior model weights.

Animals↗

Potassium channels: a computer prediction of structure and selectivity.

Model structures for the pore of the potassium channels Shaker and ROMK1 are predicted. The models arise from computer simulations and suggest reasons for the striking selectivity of these channels for K+ and the blocking of ROMK1 by internal Mg2+. The modelled structure of the Shaker pore is supported by mutagenesis data. The mutagenesis experiments indicate the side chains responsible for binding to blocking agents [tetraethylammonium (TEA) and charybdotoxin (CTX)] and the model has these side chains suitably oriented for binding. An aromatic K+ binding site part way down the pore is also predicted by the Shaker pore model.

Amino Acid Sequence↗

BYY harmony learning, structural RPCL, and topological self-organizing on mixture models.

The Bayesian Ying-Yang (BYY) harmony learning acts as a general statistical learning framework, featured by not only new regularization techniques for parameter learning but also a new mechanism that implements model selection either automatically during parameter learning or via a new class of model selection criteria used after parameter learning. In this paper, further advances on BYY harmony learning by considering modular inner representations are presented in three parts. One consists of results on unsupervisedmixture models, ranging from Gaussian mixture based Mean Square Error (MSE) clustering, elliptic clustering, subspace clustering to NonGaussian mixture based clustering not only with each cluster represented via either Bernoulli-Gaussian mixtures or independent real factor models, but also with independent component analysis implicitly made on each cluster. The second consists of results on supervised mixture-of-experts (ME) models, including Gaussian ME, Radial Basis Function nets, and Kernel regressions. The third consists of two strategies for extending the above structural mixtures into self-organized topological maps. All these advances are introduced with details on three issues, namely, (a) adaptive learning algorithms, especially elliptic, subspace, and structural rival penalized competitive learning algorithms, with model selection made automatically during learning; (b) model selection criteria for being used after parameter learning, and (c) how these learning algorithms and criteria are obtained from typical special cases of BYY harmony learning.

Bayes Theorem↗

Polymorphism in an inbreeding population under models involving underdominance.

Models of selection favoring homozygotes over the heterozygotes and involving frequency-dependency in their competitive abilities were simulated in order to determine the conditions for maintaining stable polymorphism at a diallelic locus in large inbreeding populations. With heavier inbreeding, frequency-dependency could increasingly override the effects of underdominance in both pure stand and the competing ability components of fitness in terms of yielding stable nontrivial equilibriums. The significance of such selection models is discussed for the retention of variability in inbreeding populations with a minimum of segregational load and higher overall stability in contrast to the overdominance models.

Genes, Dominant↗

Marker pair selection for mapping quantitative trait loci.

Mapping of quantitative trait loci (QTL) for backcross and F(2) populations may be set up as a multiple linear regression problem, where marker types are the regressor variables. It has been shown previously that flanking markers absorb all information on isolated QTL. Therefore, selection of pairs of markers flanking QTL is useful as a direct approach to QTL detection. Alternatively, selected pairs of flanking markers can be used as cofactors in composite interval mapping (CIM). Overfitting is a serious problem, especially if the number of regressor variables is large. We suggest a procedure denoted as marker pair selection (MPS) that uses model selection criteria for multiple linear regression. Markers enter the model in pairs, which reduces the number of models to be considered, thus alleviating the problem of overfitting and increasing the chances of detecting QTL. MPS entails an exhaustive search per chromosome to maximize the chance of finding the best-fitting models. A simulation study is conducted to study the merits of different model selection criteria for MPS. On the basis of our results, we recommend the Schwarz Bayesian criterion (SBC) for use in practice.

Chromosome Mapping↗

The robustness of Hamilton's rule with inbreeding and dominance: kin selection and fixation probabilities under partial sib mating.

Assessing the validity of Hamilton's rule when there is both inbreeding and dominance remains difficult. In this article, we provide a general method based on the direct fitness formalism to address this question. We then apply it to the question of the evolution of altruism among diploid full sibs and among haplodiploid sisters under inbreeding resulting from partial sib mating. In both cases, we find that the allele coding for altruism always increases in frequency if a condition of the form rb>c holds, where r depends on the rate of sib mating alpha but not on the frequency of the allele, its phenotypic effects, or the dominance of these effects. In both examples, we derive expressions for the probability of fixation of an allele coding for altruism; comparing these expressions with simulation results allows us to test various approximations often made in kin selection models (weak selection, large population size, large fecundity). Increasing alpha increases the probability of fixation of recessive altruistic alleles (h<1/2), while it can increase or decrease the probability of fixation of dominant altruistic alleles (h>1/2).

Altruism↗

Classifiability-based omnivariate decision trees.

Top-down induction of decision trees is a simple and powerful method of pattern classification. In a decision tree, each node partitions the available patterns into two or more sets. New nodes are created to handle each of the resulting partitions and the process continues. A node is considered terminal if it satisfies some stopping criteria (for example, purity, i.e., all patterns at the node are from a single class). Decision trees may be univariate, linear multivariate, or nonlinear multivariate depending on whether a single attribute, a linear function of all the attributes, or a nonlinear function of all the attributes is used for the partitioning at each node of the decision tree. Though nonlinear multivariate decision trees are the most powerful, they are more susceptible to the risks of overfitting. In this paper, we propose to perform model selection at each decision node to build omnivariate decision trees. The model selection is done using a novel classifiability measure that captures the possible sources of misclassification with relative ease and is able to accurately reflect the complexity of the subproblem at each node. The proposed approach is fast and does not suffer from as high a computational burden as that incurred by typical model selection algorithms. Empirical results over 26 data sets indicate that our approach is faster and achieves better classification accuracy compared to statistical model select algorithms.

Algorithms↗