Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Evolutionary model selection with a genetic algorithm: a case study using stem RNA.

The choice of a probabilistic model to describe sequence evolution can and should be justified. Underfitting the data through the use of overly simplistic models may miss out on interesting phenomena and lead to incorrect inferences. Overfitting the data with models that are too complex may ascribe biological meaning to statistical artifacts and result in falsely significant findings. We describe a likelihood-based approach for evolutionary model selection. The procedure employs a genetic algorithm (GA) to quickly explore a combinatorially large set of all possible time-reversible Markov models with a fixed number of substitution rates. When applied to stem RNA data subject to well-understood evolutionary forces, the models found by the GA 1) capture the expected overall rate patterns a priori; 2) fit the data better than the best available models based on a priori assumptions, suggesting subtle substitution patterns not previously recognized; 3) cannot be rejected in favor of the general reversible model, implying that the evolution of stem RNA sequences can be explained well with only a few substitution rate parameters; and 4) perform well on simulated data, both in terms of goodness of fit and the ability to estimate evolutionary rates. We also investigate the utility of several distance measures for comparing and contrasting inferred evolutionary models. Using widely available small computer clusters, our approach allows, for the first time, to evaluate the performance of existing RNA evolutionary models by comparing them with a large pool of candidate models and to validate common modeling assumptions. In addition, the new method provides the foundation for rigorous selection and comparison of substitution models for other types of sequence data.

Algorithms↗

Mutual selection model for weighted networks.

For most networks, the connection between two nodes is the result of their mutual affinity and attachment. In this paper, we propose a mutual selection model to characterize the weighted networks. By introducing a general mechanism of mutual selection, the model can produce power-law distributions of degree, weight, and strength, as confirmed in many real networks. Moreover, we also obtained the nontrivial clustering coefficient C, degree assortativity coefficient r, and degree-strength correlation depending on a single parameter m. These results are supported by present empirical evidence. Studying the degree-dependent average clustering coefficient C(k) and the degree-dependent average nearest neighbors' degree k(nn)(k) also provide us with a better description of the hierarchies and organizational architecture of weighted networks.

Journal Article↗

Selecting models of nucleotide substitution: an application to human immunodeficiency virus 1 (HIV-1).

The blind use of models of nucleotide substitution in evolutionary analyses is a common practice in the viral community. Typically, a simple model of evolution like the Kimura two-parameter model is used for estimating genetic distances and phylogenies, either because other authors have used it or because it is the default in various phylogenetic packages. Using two statistical approaches to model fitting, hierarchical likelihood ratio tests and the Akaike information criterion, we show that different viral data sets are better explained by different models of evolution. We demonstrate our results with the analysis of HIV-1 sequences from a hierarchy of samples; sequences within individuals, individuals within subtypes, and subtypes within groups. We also examine results for three different gene regions: gag, pol, and env. The Kimura two-parameter model was not selected as the best-fit model for any of these data sets, despite its widespread use in phylogenetic analyses of HIV-1 sequences. Furthermore, the model complexity increased with increasing sequence divergence. Finally, the molecular-clock hypothesis was rejected in most of the data sets analyzed, throwing into question clock-based estimates of divergence times for HIV-1. The importance of models in evolutionary analyses and their repercussions on the derived conclusions are discussed.

Databases, Factual↗

Error thresholds in a mutation-selection model with Hopfield-type fitness.

The deterministic limit of a Hopfield-type mutation-selection model in the sequence space approach is investigated. Genotypes are identified with two-letter sequences. Mutation is modelled as a Markov process, fitness functions are of Hopfield type, where the fitness of a sequence is determined by the Hamming distances to a number of predefined patterns. Using a maximum principle for the population mean fitness in equilibrium, the error threshold phenomenon is studied for quadratic Hopfield-type fitness functions with small numbers of patterns. Different from previous investigations of the Hopfield model, the system shows error threshold behaviour not for all fitness functions, but only for certain parameter values.

Algorithms↗

Stability analysis of the partial selfing selection model.

We undertake a detailed study of the one-locus two-allele partial selfing selection model. We show that a polymorphic equilibrium can exist only in the cases of overdominance and underdominance and only for a certain range of selfing rates. Furthermore, when it exists, we show that the polymorphic equilibrium is unique. The local stability of the polymorphic equilibrium is investigated and exact analytical conditions are presented. We also carry out an analysis of local stability of the fixation states and then conclude that only overdominance can maintain polymorphism in the population. When the linear local analysis is inconclusive, a quadratic analysis is performed. For some sets of selective values, we demonstrate global convergence. Finally, we compare and discuss results under the partial selfing model and the random mating model.

Alleles↗

Selection in complex genetic systems. VI. Equilibrium properties of two locus selection models with partial selfing.

The results of a combined analytical and numerical study of two locus selection models with partial selfing indicate that several commonly held opinions about the effects of partial self-fertilization do not hold in general. For example, the heterozygosity of a population may actually increase as the selfing rate is increased. Similarly, selection strong enough to guarantee a two locus polymorphism with complete selfing does not necessarily guarantee a two locus polymorphism with intermediate amounts of self-fertilization. The results presented here and a brief review of previously existing results indicate that the predictions of population genetic models based on the assumption of random mating will not be greatly altered by a small amount of self-fertilization, unless the loci involved are tightly linked. On the other hand, the results presented indicate that a very small amount of outcrossing may lead to marked differences from the expectation based on complete self-fertilization.

Biometry↗

Two-view multibody structure-and-motion with outliers through model selection.

Multibody structure-and-motion (MSaM) is the problem to establish the multiple-view geometry of several views of a 3D scene taken at different times, where the scene consists of multiple rigid objects moving relative to each other. We examine the case of two views. The setting is the following: Given are a set of corresponding image points in two images, which originate from an unknown number of moving scene objects, each giving rise to a motion model. Furthermore, the measurement noise is unknown, and there are a number of gross errors, which are outliers to all models. The task is to find an optimal set of motion models for the measurements. It is solved through Monte-Carlo sampling, careful statistical analysis of the sampled set of motion models, and simultaneous selection of multiple motion models to best explain the measurements. The framework is not restricted to any particular model selection mechanism because it is developed from a Bayesian viewpoint: Different model selection criteria are seen as different priors for the set of moving objects, which allow one to bias the selection procedure for different purposes.

Algorithms↗

Comparing the performance of two indices for spatial model selection: application to two mortality data.

The statistical analysis of spatially correlated data has become an important scientific research topic lately. The analysis of the mortality or morbidity rates observed at different areas may help to decide if people living in certain locations are considered at higher risk than others. Once the statistical model for the data of interest has been chosen, further effort can be devoted to identifying the areas under higher risks. Many scientists, including statisticians, have tried the conditional autoregressive (CAR) model to describe the spatial autocorrelation among the observed data. This model has greater smoothing effect than the exchangeable models, such as the Poisson gamma model for spatial data. This paper focuses on comparing the two types of models using the index LG, the ratio of local to global variability. Two applications, Taiwan asthma mortality and Scotland lip cancer, are considered and the use of LG is illustrated. The estimated values for both data sets are small, implying a Poisson gamma model may be favoured over the CAR model. We discuss the implications for the two applications respectively. To evaluate the performance of the index LG, we also compute the Bayes factor, a Bayesian model selection criterion, to see which model is preferred for the two applications and simulation data. To derive the value of LG, we estimate its posterior mode based on samples derived from the BUGS program, while for Bayes factor we use the double Laplace-Metropolis method, Schwarz criterion, and a modified harmonic mean for approximations. The results of LG and Bayes factor are consistent. We conclude that LG is fairly accurate as an index for selection between Poisson gamma and CAR model. When easy and fast computation is of concern, we recommend using LG as the first and less costly index.

Asthma↗

Model selection and optimal sampling in high-throughput experimentation.

The practical difficulties encountered in analyzing the kinetics of new reactions are considered from the viewpoint of the capabilities of state-of-the-art high-throughput systems. There are three problems. The first problem is that of model selection, i.e., choosing the correct reaction rate law. The second problem is how to obtain good estimates of the reaction parameters using only a small number of samples once a kinetic model is selected. The third problem is how to perform both functions using just one small set of measurements. To solve the first problem, we present an optimal sampling protocol to choose the correct kinetic model for a given reaction, based on T-optimal design. This protocol is then tested for the case of second-order and pseudo-first-order reactions using both experiments and computer simulations. To solve the second problem, we derive the information function for second-order reactions and use this function to find the optimal sampling points for estimating the kinetic constants. The third problem is further complicated by the fact that the optimal measurement times for determining the correct kinetic model differ from those needed to obtain good estimates of the kinetic constants. To solve this problem, we propose a Pareto optimal approach that can be tuned to give the set of best possible solutions for the two criteria. One important advantage of this approach is that it enables the integration of a priori knowledge into the workflow.

Journal Article↗

A multi-dimensional coalescent process applied to multi-allelic selection models and migration models.

For a sample of two genes from a population divided into an arbitrary number of allele classes, a general mathematical framework is developed to address the expectation and variance of the time of the most recent common ancestor. Depending on the meaning of allele classes and the manner in which genes can change among them, this framework can be applied to a diversity of population genetic models. By adoption of the infinite sites model, the effect on heterozygosity is modelled for balancing selection among allele classes, mutation between allele classes, migration among populations, and gene conversion between loci. Most results are described for a continuous time approximation to a discrete generation model. It is also shown how the discrete generation model can be used to study the hitch-hiking effect of favorable mutations.

Alleles↗

A selection model to estimate the interaction between food particles and the post-canine teeth in human mastication.

Food comminution during chewing is the composite result of selection and breakage. In the selection process, every food particle has a chance of being placed between the antagonistic post-canine teeth and being subjected to subsequent breakage. The selection chance, being the ratio between the number of selected and offered particles, has been mathematically described as a function of the number of particles offered, in terms of the number of breakage sites available on the teeth and particle affinity, i.e. the fraction of breakage sites occupied by one particle. The assumption has been made that particles are successively selected during a jaw-closing phase and that the selection chance of subsequent particles having the opportunity to occupy a breakage site proportionally decreases with the unoccupied fraction of the breakage sites left. The number of selected particles of a single size then asymptotically approaches the total number of breakage sites available for that size, when the number of particles offered increases. The critical particle number, derived from the measure of particle affinity, indicates the number of particles by which the breakage sites become saturated. The selection model for single particle sizes has been successfully applied to describe one-chew experiments, using various numbers and sizes of particles made of a silicone-rubber. After pseudo-chewing movements the subjects were unexpectedly instructed to carry out a real chew on particles (half-cubes). Undamaged, hence non-selected half-cubes could afterwards be distinguished from broken particles. The model has been extended to a particle mixture to describe the selection of particles of a certain size while other particles of different sizes are present. If a two-way competition between smaller and larger particles is assumed, the model predicts that the ratios of the selection chances between different particle sizes do not depend upon the numbers of the particles in the mixture.

Food↗

A continuous selective model for an X-linked locus.

Neglecting age-structure, but taking into account matings with differential fertility in Mendelian reproduction, a continuous selective model is formulated for a single X-linked locus with an arbitrary number of alleles. Without restricting the mating system, differential equations are derived for the genotypic and allelic frequencies. Assuming random mating, no selection, and constant fertilities and mortalities, these differential equations are solved explicitly. For this case, in contrast to the corresponding phenomenon in the usual model with discrete, non-overlapping generation, the difference between the frequencies of any allele in males and females approaches zero without oscillation.

Age Factors↗

Selection models for repeated measurements with non-random dropout: an illustration of sensitivity.

The outcome-based selection model of Diggle and Kenward for repeated measurements with non-random dropout is applied to a very simple example concerning the occurrence of mastitis in dairy cows, in which the occurrence of mastitis can be modelled as a dropout process. It is shown through sensitivity analysis how the conclusions concerning the dropout mechanism depend crucially on untestable distributional assumptions. This example is exceptional in that from a simple plot of the data two outlying observations can be identified that are the source of the apparent evidence for non-random dropout and also provide an explanation of the behaviour of the sensitivity analysis. It is concluded that a plausible model for the data does not require the assumption of non-random dropout.

Animals↗

Oligogenic model selection using the Bayesian Information Criterion: linkage analysis of the P300 Cz event-related brain potential.

The traditional likelihood-based approach to hypothesis testing may not be an optimal strategy for evaluating oligogenic models of inheritance. Under oligogenic inheritance the number of possible multilocus models can become very large; there may be several competing linkage models having similar likelihoods; and comparisons among non-nested models can be required to determine if a given multilocus model provides a significantly better fit to observed phenotypic variation than an alternative model. We propose an efficient Bayesian approach to oligogenic model selection that makes use of existing model likelihoods, and show how model uncertainty can be incorporated into parameter estimation.

Alcoholism↗

Exploratory Bayesian model selection for serial genetics data.

Characterizing the process by which molecular and cellular level changes occur over time will have broad implications for clinical decision making and help further our knowledge of disease etiology across many complex diseases. However, this presents an analytic challenge due to the large number of potentially relevant biomarkers and the complex, uncharacterized relationships among them. We propose an exploratory Bayesian model selection procedure that searches for model simplicity through independence testing of multiple discrete biomarkers measured over time. Bayes factor calculations are used to identify and compare models that are best supported by the data. For large model spaces, i.e., a large number of multi-leveled biomarkers, we propose a Markov chain Monte Carlo (MCMC) stochastic search algorithm for finding promising models. We apply our procedure to explore the extent to which HIV-1 genetic changes occur independently over time.

Algorithms↗

Pruning and model-selecting algorithms in the RBF frameworks constructed by support vector learning.

This paper presents the pruning and model-selecting algorithms to the support vector learning for sample classification and function regression. When constructing RBF network by support vector learning we occasionally obtain redundant support vectors which do not significantly affect the final classification and function approximation results. The pruning algorithms primarily based on the sensitivity measure and the penalty term. The kernel function parameters and the position of each support vector are updated in order to have minimal increase in error, and this makes the structure of SVM network more flexible. We illustrate this approach with synthetic data simulation and face detection problem in order to demonstrate the pruning effectiveness.

Algorithms↗

A loss function approach to model selection in nonlinear principal components.

The nonlinear transformation of the input variables that characterises the first nonlinear principal component is modelled as a linear sum of radially-symmetric kernel functions. It is shown that the parameters of the variance maximising transformation may be obtained through the minimisation of a loss function measuring departure from homogeneity. An alternating least squares algorithm is given. This is used as the basis of a cross-validation routine for model selection.

Journal Article↗

Note on "Comparison of model selection for regression" by Vladimir Cherkassky and Yunqian Ma.

While Cherkassky and Ma (2003) raise some interesting issues in comparing techniques for model selection, their article appears to be written largely in protest of comparisons made in our book, Elements of Statistical Learning (2001). Cherkassky and Ma feel that we falsely represented the structural risk minimization (SRM) method, which they defend strongly here. In a two-page section of our book (pp. 212-213), we made an honest attempt to compare the SRM method with two related techniques, Aikaike information criterion (AIC) and Bayesian information criterion (BIC). Apparently, we did not apply SRM in the optimal way. We are also accused of using contrived examples, designed to make SRM look bad. Alas, we did introduce some careless errors in our original simulation--errors that were corrected in the second and subsequent printings. Some of these errors were pointed out to us by Cherkassky and Ma (we supplied them with our source code), and as a result we replaced the assessment "SRM performs poorly overall" with a more moderate "the performance of SRM is mixed" (p. 212).

Models, Neurological↗