Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Use of Akaike information criteria for model selection and inference. An application to assess prevention of gastrointestinal parasitism and respiratory mortality of Guinean goats in Kolda, Senegal.

A field experiment was carried out in Kolda (southern Senegal) from July 1986 to July 1988. Its goals were to: (1) describe the patterns of mortality of female Guinean goats by age, season and year; (2) assess preventive measures against respiratory diseases and gastrointestinal parasitism in reducing mortality; and (3) estimate the overall impact of these measures on survival to 1 year of age. Preventive measures for respiratory disease included vaccination against peste des petits ruminants (PPR) and pneumonic pasteurellosis (Pasteurella multocida types A and D). Control of gastrointestinal parasites was by deworming does with morantel (7.5mg kg(-1), three times during the rainy season). The effects of vaccines and deworming were tested in a randomised factorial field experiment with villages being the experimental units. A total of 19 villages, 113 goat herds and 1,458 goats were included in the study. Generalised linear models of survival for five cohorts of goats (defined by five different birth seasons) used a binomial assumption for the response distribution and a complementary log-log link. Explanatory variables included age, season, year, vaccination, deworming and their interactions. A complex a priori model was built on the basis of previous epidemiological knowledge; a purposely selected set of simpler models was compared to this full model by the Akaike information criterion (AIC) and derived statistics. Inference on 1-year survival and treatment effects accounted for model-selection uncertainty. It was carried out with a bootstrap procedure and used information from the whole set of selected models. Large variations in mortality by year and season were observed but no regular seasonal pattern was apparent. Mortality probabilities of kids in dewormed groups decreased quickly after birth, but remained elevated up to 9 months of age in the non-dewormed groups. Deworming lowered the risk of mortality. Vaccination alone was not protective (except during an observed outbreak of PPR).

Animals↗

High nucleotide sequence variation in a region of low recombination in Drosophila simulans is consistent with the background selection model.

We surveyed nucleotide sequence variation at glucose dehydrogenase (Gld), in a region of low recombination on chromosome 3R, from a population sample of Drosophila simulans. The levels of nucleotide variation were surprisingly high. There was no departure from the expectation of a neutral model for the level of polymorphism, indicating no evidence of a selective sweep in this region. There was a significant deficiency of singleton polymorphisms according to the Fu and Li test, although Tajima and Hudson, Kreitman, and Aguade (HKA) tests do not provide evidence of a significant elevation of variation due to balancing selection. Genetic map data for the D. simulans third chromosome were used to calculate expected values of pi for Gld under a current model of background selection, varying the values for the parameter sh (selection coefficient against deleterious mutations). We show that the recombinational landscape of D. simulans is sufficiently different from that of D. melanogaster that we expect higher variation under the background selection model, even when effective population sizes are assumed to be equal. The data for Gld were tested against the predictions using computer simulations of the distribution of the number of segregating sites conditioned on pi. Background selection alone can explain our observations as long as sh is larger than 0.005 and species-level effective population size is assumed to be several-fold larger than in D. melanogaster. Alternatively, the deleterious mutation rate may be smaller in D. simulans, or balancing selection may be acting nearby, thereby reducing the effect of background selection.

Animals↗

Quantifying epidemiologic risk factors using non-parametric regression: model selection remains the greatest challenge.

Logistic regression is widely used to estimate relative risks (odds ratios) from case-control studies, but when the study exposure is continuous, standard parametric models may not accurately characterize the exposure-response curve. Semi-parametric generalized linear models provide a useful extension. In these models, the exposure of interest is modelled flexibly using a regression spline or a smoothing spline, while other variables are modelled using conventional methods. When coupled with a model-selection procedure based on minimizing a cross-validation score, this approach provides a non-parametric, objective, and reproducible method to characterize the exposure-response curve by one or several models with a favourable bias-variance trade-off. We applied this approach to case-control data to estimate the dose-response relationship between alcohol consumption and risk of oral cancer among African Americans. We did not find a uniquely 'best' model, but results using linear, cubic, and smoothing splines were consistent: there does not appear to be a risk-free threshold for alcohol consumption vis-à-vis the development of oral cancer. This finding was not apparent using a standard step-function model. In our analysis, the cross-validation curve had a global minimum and also a local minimum. In general, the phenomenon of multiple local minima makes it more difficult to interpret the results, and may present a computational roadblock to non-parametric generalized additive models of multiple continuous exposures. Nonetheless, the semi-parametric approach appears to be a practical advance.

Black or African American↗

A model selection algorithm for a posteriori probability estimation with neural networks.

This paper proposes a novel algorithm to jointly determine the structure and the parameters of a posteriori probability model based on neural networks (NNs). It makes use of well-known ideas of pruning, splitting, and merging neural components and takes advantage of the probabilistic interpretation of these components. The algorithm, so called a posteriori probability model selection (PPMS), is applied to an NN architecture called the generalized softmax perceptron (GSP) whose outputs can be understood as probabilities although results shown can be extended to more general network architectures. Learning rules are derived from the application of the expectation-maximization algorithm to the GSP-PPMS structure. Simulation results show the advantages of the proposed algorithm with respect to other schemes.

Algorithms↗

The evolutionary forces maintaining a wild polymorphism of Littorina saxatilis: model selection by computer simulations.

Two rocky shore ecotypes of Littorina saxatilis from north-west Spain live at different shore levels and habitats and have developed an incomplete reproductive isolation through size assortative mating. The system is regarded as an example of sympatric ecological speciation. Several experiments have indicated that different evolutionary forces (migration, assortative mating and habitat-dependent selection) play a role in maintaining the polymorphism. However, an assessment of the combined contributions of these forces supporting the observed pattern in the wild is absent. A model selection procedure using computer simulations was used to investigate the contribution of the different evolutionary forces towards the maintenance of the polymorphism. The agreement between alternative models and experimental estimates for a number of parameters was quantified by a least square method. The results of the analysis show that the fittest evolutionary model for the observed polymorphism is characterized by a high gene flow, intermediate-high reproductive isolation between ecotypes, and a moderate to strong selection against the nonresident ecotypes on each shore level. In addition, a substantial number of additive loci contributing to the selected trait and a narrow hybrid definition with respect to the phenotype are scenarios that better explain the polymorphism, whereas the ecotype fitnesses at the mid-shore, the level of phenotypic plasticity, and environmental effects are not key parameters.

Animals↗

Using nonlinear models in fMRI data analysis: model selection and activation detection.

There is an increasing interest in using physiologically plausible models in fMRI analysis. These models do raise new mathematical problems in terms of parameter estimation and interpretation of the measured data. In this paper, we show how to use physiological models to map and analyze brain activity from fMRI data. We describe a maximum likelihood parameter estimation algorithm and a statistical test that allow the following two actions: selecting the most statistically significant hemodynamic model for the measured data and deriving activation maps based on such model. Furthermore, as parameter estimation may leave much incertitude on the exact values of parameters, model identifiability characterization is a particular focus of our work. We applied these methods to different variations of the Balloon Model (Buxton, R.B., Wang, E.C., and Frank, L.R. 1998. Dynamics of blood flow and oxygenation changes during brain activation: the balloon model. Magn. Reson. Med. 39: 855-864; Buxton, R.B., Uludağ, K., Dubowitz, D.J., and Liu, T.T. 2004. Modelling the hemodynamic response to brain activation. NeuroImage 23: 220-233; Friston, K. J., Mechelli, A., Turner, R., and Price, C. J. 2000. Nonlinear responses in fMRI: the balloon model, volterra kernels, and other hemodynamics. NeuroImage 12: 466-477) in a visual perception checkerboard experiment. Our model selection proved that hemodynamic models better explain the BOLD response than linear convolution, in particular because they are able to capture some features like poststimulus undershoot or nonlinear effects. On the other hand, nonlinear and linear models are comparable when signals get noisier, which explains that activation maps obtained in both frameworks are comparable. The tools we have developed prove that statistical inference methods used in the framework of the General Linear Model might be generalized to nonlinear models.

Adult↗

Model selection in electromagnetic source analysis with an application to VEFs.

In electromagnetic source analysis, it is necessary to determine how many sources are required to describe the electroencephalogram or magnetoencephalogram adequately. Model selection procedures (MSPs) or goodness of fit procedures give an estimate of the required number of sources. Existing and new MSPs are evaluated in different source and noise settings: two sources which are close or distant and noise which is uncorrelated or correlated. The commonly used MSP residual variance is seen to be ineffective, that is it often selects too many sources. Alternatives like the adjusted Hotelling's test, Bayes information criterion and the Wald test on source amplitudes are seen to be effective. The adjusted Hotelling's test is recommended if a conservative approach is taken and MSPs such as Bayes information criterion or the Wald test on source amplitudes are recommended if a more liberal approach is desirable. The MSPs are applied to empirical data (visual evoked fields).

Algorithms↗

Evaluation of the benchmark dose method for dichotomous data: model dependence and model selection.

The benchmark dose (BMD) method was evaluated using the USEPA BMD software. Dose-response data on cleft palate and hydronephrosis for a number of related polyhalogenated aromatic compounds were obtained from the literature. According to chi(2) test statistics, each dichotomous USEPA model failed to adequately describe only 1 of 12 cleft palate data sets. For hydronephrosis, the models were discriminated to a higher extent according to global goodness-of-fit. NOAELs for cleft palate corresponded to BMDLs (the approximate lower confidence limit on the BMD) for extra risks in the range of 5% or below. Model dependence of the BMDL estimate was more pronounced at lower levels of benchmark response (BMR). A BMR of 5% (extra risk) is recommended for cleft palate since model differences at this level were limited for all data. In addition, at BMRs of 5-10% the BMDL for all models was little affected by the specified confidence limit size (in the 90-99% range). For BMDL determination a conservative model selection approach was applied. At the suggested level of BMR (5%) this procedure resulted in use of the same model (multistage model) for the cleft palate endpoint in general. Akaike's information criterion (AIC) was considered for comparison between models. Determination of appropriateness of use of such methods in dose-response applications requires further analysis.

Abnormalities, Drug-Induced↗

Unravelling the regulatory structure of biochemical networks using stimulus response experiments and large-scale model selection.

To unravel the complex in vivo regulatory interdependences of biochemical networks, experiments with the living organism are absolutely necessary. Stimulus response experiments (SREs) have become increasingly popular in recent years. The response of metabolite concentrations from all major parts of the central metabolism is monitored over time by modem analytical methods, producing several thousand data points. SREs are applied to determine enzyme kinetic parameters and to find unknown enzyme regulatory mechanisms. Owing to the complex regulatory structure of metabolic networks and the amount of measured data, the evaluation of an SRE has to be extensively supported by modelling. If the enzyme regulatory mechanisms are part of the investigation, a large number of models with different enzyme kinetics have to be tested for their ability to reproduce the observed behaviour. In this contribution, a systematic model-building process for data-driven exploratory modelling is introduced with the aim of discovering essential features of the biological system. The process is based on data pre-processing, correlation-based hypothesis generation, automatic model family generation, large-scale model selection and statistical analysis of the best-fitting models followed by an extraction of common features. It is illustrated by the example of the aromatic amino acid synthesis pathway in Escherichia coli.

Adaptation, Physiological↗

A joint analysis of quality of life and survival using a random effect selection model.

In studies of patients with advanced disease, longitudinal quality of life data may be truncated as a result of early death. Since survival and quality of life are likely to be related, modelling of the quality of life response needs to account for these different survival patterns. Here we discuss the application of a random effect selection model, in the form of a trivariate Normal model for the joint analysis of quality of life response (intercept and slope) and log survival time. Under certain assumptions this can give an unbiased description of the quality of life responses and valid inferences comparing treatment strategies in a clinical trial. It also indicates how quality of life and survival are related, by estimating the expected quality of life responses conditional on different survival times. Model parameters can be estimated using a restricted iterative generalized least-squares (RIGLS) procedure within standard software, extended to handle censoring of survival outcome using an EM algorithm. The model is applied to a physical quality of life score and survival data from a trial of treatment for patients with colorectal hepatic metastases. Survival differed between the treatment groups, and quality of life repsonse tended to be worse, both in initial level and change over time, for those patients who died earlier. The parameter estimates obtained agreed well with those from analysing the extended trial data set with complete survival information. Residual diagnostics used to check the necessary underlying assumptions of the model are exemplified. We conclude that such models can give an informative description of longitudinal responses when these are truncated by differential survival patterns.

Algorithms↗

Avoiding model selection bias in small-sample genomic datasets.

MOTIVATION: Genomic datasets generated by high-throughput technologies are typically characterized by a moderate number of samples and a large number of measurements per sample. As a consequence, classification models are commonly compared based on resampling techniques. This investigation discusses the conceptual difficulties involved in comparative classification studies. Conclusions derived from such studies are often optimistically biased, because the apparent differences in performance are usually not controlled in a statistically stringent framework taking into account the adopted sampling strategy. We investigate this problem by means of a comparison of various classifiers in the context of multiclass microarray data. RESULTS: Commonly used accuracy-based performance values, with or without confidence intervals, are inadequate for comparing classifiers for small-sample data. We present a statistical methodology that avoids bias in cross-validated model selection in the context of small-sample scenarios. This methodology is valid for both k-fold cross-validation and repeated random sampling.

Algorithms↗

Trading variance reduction with unbiasedness: the regularized subspace information criterion for robust model selection in kernel regression.

A well-known result by Stein (1956) shows that in particular situations, biased estimators can yield better parameter estimates than their generally preferred unbiased counterparts. This letter follows the same spirit, as we will stabilize the unbiased generalization error estimates by regularization and finally obtain more robust model selection criteria for learning. We trade a small bias against a larger variance reduction, which has the beneficial effect of being more precise on a single training set. We focus on the subspace information criterion (SIC), which is an unbiased estimator of the expected generalization error measured by the reproducing kernel Hilbert space norm. SIC can be applied to the kernel regression, and it was shown in earlier experiments that a small regularization of SIC has a stabilization effect. However, it remained open how to appropriately determine the degree of regularization in SIC. In this article, we derive an unbiased estimator of the expected squared error, between SIC and the expected generalization error and propose determining the degree of regularization of SIC such that the estimator of the expected squared error is minimized. Computer simulations with artificial and real data sets illustrate that the proposed method works effectively for improving the precision of SIC, especially in the high-noise-level cases. We furthermore compare the proposed method to the original SIC, the cross-validation, and an empirical Bayesian method in ridge parameter selection, with good results.

Bayes Theorem↗

Model selection for integrated recovery/recapture data.

Catchpole et al. (1998, Biometrics 54, 33-46) provide a novel scheme for integrating both recovery and recapture data analyses and derive sufficient statistics that facilitate likelihood computations. In this article, we demonstrate how their efficient likelihood expression can facilitate Bayesian analyses of these kinds of data and extend their methodology to provide a formal framework for model determination. We consider in detail the issue of model selection with respect to a set of recapture/recovery histories of shags (Phalacrocorax aristotelis) and determine, from the enormous range of biologically plausible models available, which best describe the data. By using reversible jump Markov chain Monte Carlo methodology, we demonstrate how this enormous model space can be efficiently and effectively explored without having to resort to performing an infeasibly large number of pairwise comparisons or some ad hoc stepwise procedure. We find that the model used by Catchpole et al. (1998) has essentially zero posterior probability and that, of the 477,144 possible models considered, over 60% of the posterior mass is placed on three neighboring models with biologically interesting interpretations.

Animals↗

Uncertainty modeling and model selection for geometric inference.

We first investigate the meaning of "statistical methods" for geometric inference based on image feature points. Tracing back the origin of feature uncertainty to image processing operations, we discuss the implications of asymptotic analysis in reference to "geometric fitting" and "geometric model selection" and point out that a correspondence exists between the standard statistical analysis and the geometric inference problem. Then, we derive the "geometric AIC" and the "geometric MDL" as counterparts of Akaike's AIC and Rissanen's MDL. We show by experiments that the two criteria have contrasting characteristics in detecting degeneracy.

Algorithms↗

Bayesian analysis and model selection for interval-censored survival data.

Interval-censored data occur in survival analysis when the survival time of each patient is only known to be within an interval and these censoring intervals differ from patient to patient. For such data, we present some Bayesian discretized semiparametric models, incorporating proportional and nonproportional hazards structures, along with associated statistical analyses and tools for model selection using sampling-based methods. The scope of these methodologies is illustrated through a reanalysis of a breast cancer data set (Finkelstein, 1986, Biometrics 42, 845-854) to test whether the effect of covariate on survival changes over time.

Algorithms↗

Model selection in magnetic resonance imaging measurements of vascular permeability: Gadomer in a 9L model of rat cerebral tumor.

Vasculature in and around the cerebral tumor exhibits a wide range of permeabilities, from normal capillaries with essentially no blood-brain barrier (BBB) leakage to a tumor vasculature that freely passes even such large molecules as albumin. In measuring BBB permeability by magnetic resonance imaging (MRI), various contrast agents, sampling intervals, and contrast distribution models can be selected, each with its effect on the measurement's outcome. Using Gadomer, a large paramagnetic contrast agent, and MRI measures of T(1) over a 25-min period, BBB permeability was estimated in 15 Fischer rats with day-16 9L cerebral gliomas. Three vascular models were developed: (1) impermeable (normal BBB); (2) moderate influx (leakage without efflux); and (3) fast leakage with bidirectional exchange. For data analysis, these form nested models. Model 1 estimates only vascular plasma volume, v(D), Model 2 (the Patlak graphical approach) v(D) and the influx transfer constant K(i). Model 3 estimates v(D), K(i), and the reverse transfer constant, k(b), through which the extravascular distribution space, v(e), is calculated. For this contrast agent and experimental duration, Model 3 proved the best model, yielding the following central tumor means (+/-s.d.; n = 15): v(D) = 0.07 +/- 0.03 for K(i) = 0.0105 +/- 0.005 min(-1) and v(e) = 0.10 +/- 0.04. Model 2 K(i) estimates were approximately 30% of Model 3, but highly correlated (r = 0.80, P < 0.0003). Sizable inhomogeneity in v(D), K(i), and k(b) appeared within each tumor. We conclude that employing nested models enables accurate assessment of transfer constants among areas where BBB permeability, contrast agent distribution volumes, and signal-to-noise vary.

Animals↗

Modeling selection for production traits under constant infection pressure.

This article presents a model describing the relationship between level of disease resistance and production under constant infection pressure. The model assumes that given a certain infection pressure, there is a threshold for resistance below which animals will stop producing, and that there is also a threshold for resistance above which animals produce at production potential. In between both thresholds animals will show a decrease in production, the size of decrease depending on the severity of infection and the level of resistance. The dynamic relationship between production and resistance when level of resistance changes, such as due to infection, is modeled both stochastically and deterministically. Selection started in a population with very poor level of resistance introduced in an environment with constant infection pressure. Mass selection on observed production was applied, which resulted in a nonlinear selection response for all three traits considered. When resistance is poor, selection for observed production results in increased level of resistance. With increasing level of resistance, selection response shifts to production potential and eventually selection for observed production is equivalent to selection for production potential. The rate at which resistance is improved depends on its heritability, the difference between both thresholds, and selection intensity. The model also revealed that when a zero correlation between resistance and production potential is assumed, the phenotypic correlation between resistance and observed production level increases for low levels of resistance and subsequently asymptotes to zero, whereas the phenotypic correlation between production potential and observed production asymptotes to 1.0. For most breeding schemes investigated, the deterministic model performed well in relation to the stochastic simulation results. Experimental results reported in literature support the model predictions.

Animal Husbandry↗

Bayesian model selection for genome-wide epistatic quantitative trait loci analysis.

The problem of identifying complex epistatic quantitative trait loci (QTL) across the entire genome continues to be a formidable challenge for geneticists. The complexity of genome-wide epistatic analysis results mainly from the number of QTL being unknown and the number of possible epistatic effects being huge. In this article, we use a composite model space approach to develop a Bayesian model selection framework for identifying epistatic QTL for complex traits in experimental crosses from two inbred lines. By placing a liberal constraint on the upper bound of the number of detectable QTL we restrict attention to models of fixed dimension, greatly simplifying calculations. Indicators specify which main and epistatic effects of putative QTL are included. We detail how to use prior knowledge to bound the number of detectable QTL and to specify prior distributions for indicators of genetic effects. We develop a computationally efficient Markov chain Monte Carlo (MCMC) algorithm using the Gibbs sampler and Metropolis-Hastings algorithm to explore the posterior distribution. We illustrate the proposed method by detecting new epistatic QTL for obesity in a backcross of CAST/Ei mice onto M16i.

Algorithms↗