Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Bayesian phylogenetic model selection using reversible jump Markov chain Monte Carlo.

A common problem in molecular phylogenetics is choosing a model of DNA substitution that does a good job of explaining the DNA sequence alignment without introducing superfluous parameters. A number of methods have been used to choose among a small set of candidate substitution models, such as the likelihood ratio test, the Akaike Information Criterion (AIC), the Bayesian Information Criterion (BIC), and Bayes factors. Current implementations of any of these criteria suffer from the limitation that only a small set of models are examined, or that the test does not allow easy comparison of non-nested models. In this article, we expand the pool of candidate substitution models to include all possible time-reversible models. This set includes seven models that have already been described. We show how Bayes factors can be calculated for these models using reversible jump Markov chain Monte Carlo, and apply the method to 16 DNA sequence alignments. For each data set, we compare the model with the best Bayes factor to the best models chosen using AIC and BIC. We find that the best model under any of these criteria is not necessarily the most complicated one; models with an intermediate number of substitution types typically do best. Moreover, almost all of the models that are chosen as best do not constrain a transition rate to be the same as a transversion rate, suggesting that it is the transition/transversion rate bias that plays the largest role in determining which models are selected. Importantly, the reversible jump Markov chain Monte Carlo algorithm described here allows estimation of phylogeny (and other phylogenetic model parameters) to be performed while accounting for uncertainty in the model of DNA substitution.

Algorithms↗

On model selection for standard curve in assay development.

This paper discusses the selection of an appropriate statistical model for representing standard curve in assay development. This is an important issue in assay validation because the accuracy and reliability of the assay result depend on the selected standard curve. In this study, we propose a selection procedure, which is based on the R2 and the mean squared error of the estimation sample, to determine the "best" model. An example concerning an assay validation study is used to illustrate the application of the proposed procedure to discriminate among the five commonly used statistical models.

Calibration↗

Statistical inference and model selection for the 1861 Hagelloch measles epidemic.

A stochastic epidemic model is proposed which incorporates heterogeneity in the spread of a disease through a population. In particular, three factors are considered: the spatial location of an individual's home and the household and school class to which the individual belongs. The model is applied to an extremely informative measles data set and the model is compared with nested models, which incorporate some, but not all, of the aforementioned factors. A reversible jump Markov chain Monte Carlo algorithm is then introduced which assists in selecting the most appropriate model to fit the data.

Adolescent↗

Distinguishing the hitchhiking and background selection models.

A simple method to distinguish hitchhiking and background selection is proposed. It is based on the observation that these models make different predictions about the average level of nucleotide diversity in regions of low recombination. The method is applied to data from Drosophila melanogaster and two highly selfing tomato species.

Animals↗

Effects of metronidazole, tetracycline, and bismuth-metronidazole-tetracycline triple therapy in the Helicobacter pylori SS1 mouse model after 1 day of dosing: development of an H. pylori lead selection model.

We evaluated the effect of optimized doses and dosing schedules of metronidazole, tetracycline, and bismuth-metronidazole-tetracycline (BMT) triple therapy with only 1 day of dosing on Helicobacter pylori SS1 titers in a mouse model. A reduction of bacterial titers was observable with 22.5 and 112.5 mg of metronidazole per kg of body weight (as well as BMT) given twice daily and four times daily (QID). Two hundred milligrams of tetracycline per kilogram, given QID, resulted in only a slight reduction of H. pylori titers in the stomach. We argue that optimization of doses based on antimicrobial drug levels in the animal and shortened (1 or 2 days) drug administration can be used to facilitate early evaluation of putative anti-H. pylori drug candidates in lieu of using human doses and extended schedules (7 to 14 days), as can be deduced from the results seen with these antimicrobial agents.

Animals↗

Model Comparisons and Model Selections Based on Generalization Criterion Methodology.

The purpose of this article is to formalize the generalization criterion method for model comparison. The method has the potential to provide powerful comparisons of complex and nonnested models that may also differ in terms of numbers of parameters. The generalization criterion differs from the better known cross-validation criterion in the following critical procedure. Although both employ a calibration stage to estimate parameters, cross-validation employs a replication sample from the same design for the validation stage, whereas generalization employs a new design for the critical stage. Two examples of the generalization criterion method are presented that demonstrate its usefulness for selecting a model based on sound scientific principles out of a set that also contains models lacking sound scientific principles that are either overly complex or oversimplified. The main advantage of the generalization criterion is its reliance on extrapolations to new conditions. After all, accurate a priori predictions to new conditions are the hallmark of a good scientific theory. Copyright 2000 Academic Press.

Journal Article↗

Model selection for incomplete and design-based samples.

The Akaike information criterion, AIC, is one of the most frequently used methods to select one or a few good, optimal regression models from a set of candidate models. In case the sample is incomplete, the naive use of this criterion on the so-called complete cases can lead to the selection of poor or inappropriate models. A similar problem occurs when a sample based on a design with unequal selection probabilities, is treated as a simple random sample. In this paper, we consider a modification of AIC, based on reweighing the sample in analogy with the weighted Horvitz-Thompson estimates. It is shown that this weighted AIC-criterion provides better model choices for both incomplete and design-based samples. The use of the weighted AIC-criterion is illustrated on data from the Belgian Health Interview Survey, which motivated this research. Simulations show its performance in a variety of settings.

Adult↗

Advanced hemostatic dressing development program: animal model selection criteria and results of a study of nine hemostatic dressings in a model of severe large venous hemorrhage and hepatic injury in Swine.

BACKGROUND: An advanced hemostatic dressing is needed to augment current methods for the control of life-threatening hemorrhage. A systematic approach to the study of dressings is described. We studied the effects of nine hemostatic dressings on blood loss using a model of severe venous hemorrhage and hepatic injury in swine. METHODS: Swine were treated using one of nine hemostatic dressings. Dressings used the following primary active ingredients: microfibrillar collagen, oxidized cellulose, thrombin, fibrinogen, propyl gallate, aluminum sulfate, and fully acetylated poly-N-acetyl glucosamine. Standardized liver injuries were induced, dressings were applied, and resuscitation was initiated. Blood loss, hemostasis, and 60-minute survival were quantified. RESULTS: The American Red Cross hemostatic dressing (fibrinogen and thrombin) reduced (p < 0.01) posttreatment blood loss (366 mL; 95% confidence interval, 175-762 mL) and increased (p < 0.05) the percentage of animals in which hemostasis was attained (73%), compared with gauze controls (2,973 mL; 95% confidence interval, 1,414-6,102 mL and 0%, respectively). No other dressing was effective. The number of vessels lacerated was positively related to pretreatment blood loss and negatively related to hemostasis. CONCLUSION: The hemorrhage model allowed differentiation among topical hemostatic agents for severe hemorrhage. The American Red Cross hemostatic dressing was effective and warrants further development.

Animals↗

The effect of smoking on health using a sequential self-selection model.

We estimate a structural model of individual smoking behaviour emphasizing the role of individual risk belief on smoking choices. Our model consists of five equations: two selection equations for initiation and cessation decisions, and three switching outcome regressions for nonsmokers, ex-smokers, and current smokers. The presence of significant self-selectivity implies that the health effects of smoking based on sample proportions do not correctly indicate the true risk of cigarette smoking. Further, our evidence suggests that the self-selection in the cessation decision, but not in the initiation decision, is consistent with economic rationality. We estimate the model by full information maximum likelihood (FIML) with starting values from heteroskedasticity corrected Heckman-Lee two-step method using newly released Health and Retirement Study (HRS) data.

Adult↗

HIV-1 viral fitness estimation using exchangeable on subsets priors and prior model selection.

The phenotype-genotype problem is a fundamental problem of biology where an organism's genotype (genetic information) predicts its phenotype (observable characteristic). Viral fitness, defined as the reproductive capacity of a virus compared to a standard, is a continuous phenotype. We construct models to predict viral fitness as a function of mutation away from the standard wildtype virus. Data of this nature are difficult to analyse because there are potentially many more parameters than observations. We treat this issue as a regression problem using a prior with both a shrinkage component and a variable selection component. The key to practical implementation of the model is the prior specification for the regression coefficients. We use results from the scientific literature to construct several informative exchangeable within subsets priors (ESP). We use prior model selection (PMS) to select among our priors. Two novel graphics present results from five models each with 71 predictors.

Codon↗

Polynomial neural network for linear and non-linear model selection in quantitative-structure activity relationship studies on the internet.

This article presents a self-organising multilayered iterative algorithm that provides linear and non-linear polynomial regression models thus allowing the user to control the number and the power of the terms in the models. The accuracy of the algorithm is compared to the partial least squares (PLS) algorithm using fourteen data sets in quantitative-structure activity relationship studies. The calculated data show that the proposed method is able to select simple models characterized by a high prediction ability and thus provides a considerable interest in quantitative-structure activity relationship studies. The software is developed using client-server protocol (Java and C++ languages) and is available for world-wide users on the Web site of the authors.

Internet↗

Comparison of Bayesian model averaging and stepwise methods for model selection in logistic regression.

Logistic regression is the standard method for assessing predictors of diseases. In logistic regression analyses, a stepwise strategy is often adopted to choose a subset of variables. Inference about the predictors is then made based on the chosen model constructed of only those variables retained in that model. This method subsequently ignores both the variables not selected by the procedure, and the uncertainty due to the variable selection procedure. This limitation may be addressed by adopting a Bayesian model averaging approach, which selects a number of all possible such models, and uses the posterior probabilities of these models to perform all inferences and predictions. This study compares the Bayesian model averaging approach with the stepwise procedures for selection of predictor variables in logistic regression using simulated data sets and the Framingham Heart Study data. The results show that in most cases Bayesian model averaging selects the correct model and out-performs stepwise approaches at predicting an event of interest.

Age Factors↗

[A selection model for predicting poultry meat productivity].

A model for estimating growth intensity parametres in poultry in early ontogenesis has been worked out. The growth intensity parametres were studied; high efficiency of using the model for prognosis of meat productivity in poultry at the age of 49-56 days are shown. The application techniques of the model offered for increasing poultry growth energy were examined.

Animals↗

Evaluation of decay times in coupled spaces: Bayesian decay model selection.

This paper applies Bayesian probability theory to determination of the decay times in coupled spaces. A previous paper [N. Xiang and P. M. Goggans, J. Acoust. Soc. Am. 110, 1415-1424 (2001)] discussed determination of the decay times in coupled spaces from Schroeder's decay functions using Bayesian parameter estimation. To this end, the previous paper described the extension of an existing decay model [N. Xiang, I. Acoust. Soc. Am. 98, 2112-2121 (1995)] to incorporate one or more decay modes for use with Bayesian inference. Bayesian decay time estimation will obtain reasonable results only when it employs an appropriate decay model with the correct number of decay modes. However, in architectural acoustics practice, the number of decay modes may not be known when evaluating Schroeder's decay functions. The present paper continues the endeavor of the previous paper to apply Bayesian probability inference for comparison and selection of an appropriate decay model based upon measured data. Following a summary of Bayesian model comparison and selection, it discusses selection of a decay model in terms of experimentally measured Schroeder's decay functions. The present paper, along with the Bayesian decay time estimation described previously, suggests that Bayesian probability inference presents a suitable approach to the evaluation of decay times in coupled spaces.

Journal Article↗

Gametophytic selection in Arabidopsis thaliana supports the selective model of intron length reduction.

Why do highly expressed genes have small introns? This is an important issue, not least because it provides a testing ground to compare selectionist and neutralist models of genome evolution. Some argue that small introns are selectively favoured to reduce the costs of transcription. Alternatively, large introns might permit complex regulation, not needed for highly expressed genes. This "genome design" hypothesis evokes a regionalized model of control of expression and hence can explain why intron size covaries with intergene distance, a feature also consistent with the hypothesis that highly expressed genes cluster in genomic regions with high deletion rates. As some genes are expressed in the haploid stage and hence subject to especially strong purifying selection, the evolution of genes in Arabidopsis provides a novel testing ground to discriminate between these possibilities. Importantly, controlling for expression level, genes that are expressed in pollen have shorter introns than genes that are expressed in the sporophyte. That genes flanking pollen-expressed genes have average-sized introns and intergene distances argues against regional mutational biases and genomic design. These observations thus support the view that selection for efficiency contributes to the reduction in intron length and provide the first report of a molecular signature of strong gametophytic selection.

Arabidopsis↗