Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

A model for macromolecular selection in complementary instructing systems.

In this paper we consider a model for the selection and evolution of biological macromolecules when their reproduction is based on complementary instruction. The model is an extension of one of Eigen's models for selection and takes in account explicitly the formation of both single stranded and double stranded molecular complexes. We construct exact solutions to the rate equations for the case of constant rate parameters and error distributions. Criteria for selection are discussed.

Biological Evolution↗

The prey localisation model: a stability analysis.

This paper analyzes the "Prey localisation Model" (House 1984), for animals that are unable to verge their eyes. The Prey localisation Model selects a single point or a portion of the scene in its visual space. In particular it imitates the behaviour of frogs and toads of selecting the closer target when two equally attractive prey are presented to it. The model achieves its goal by tightly coupling two prey selection processes, one for each eye, with lens accommodation. In this paper we offer a stability analysis of the model, and show how lens accommodation, i.e. adjustment in the focal length of the lens, biases the selection of the proximal target. We examine the properties of the model that are responsible for its behaviour and derive a set of conditions which guarantees the localisation of the correct target.

Animals↗

A model of weak selection in the infinite alleles framework.

Ewens (1972) proposed a model in the infinite allele framework for populations with neutrality of all alleles at a particular locus. This paper proposes a generalisation of Ewens' result for situations where there is a form of weak selection. The models considered here are continuous time, discrete state space Markov processes.

Alleles↗

Contrasting Bayesian analysis of survey data and clinical trials.

Although both surveys and clinical trials are amenable to Bayesian hierarchical modelling, the general aims, constraints and actual analysis of each can often vary considerably. First, examples are presented showing how Bayesian hierarchical modelling can be used to produce estimates for small areas from survey data and, also, how it can be used to combine data from clinical trials. Then, it will be shown how surveys and clinical trials may differ with respect to the presence of design effects/selection biases and with the ability to validate models. The impact of the design on modelling will be highlighted and a class of sample selection models will be shown to help alleviate the design's influence. Although surveys generally have enough data to validate many features of a model, clinical trials may not, leaving sensitivity analysis as a means to prior acceptance. Some design issues, contrasting Bayesian with frequentist methods, will also be discussed. Published in 2001 by John Wiley & Sons, Ltd.

Bayes Theorem↗

A comparison of several regression models for analysing cost of CABG surgery.

Investigators in clinical research are often interested in determining the association between patient characteristics and cost of medical or surgical treatment. However, there is no uniformly agreed upon regression model with which to analyse cost data. The objective of the current study was to compare the performance of linear regression, linear regression with log-transformed cost, generalized linear models with Poisson, negative binomial and gamma distributions, median regression, and proportional hazards models for analysing costs in a cohort of patients undergoing CABG surgery. The study was performed on data comprising 1959 patients who underwent CABG surgery in Calgary, Alberta, between June 1994 and March 1998. Ten of 21 patient characteristics were significantly associated with cost of surgery in all seven models. Eight variables were not significantly associated with cost of surgery in all seven models. Using mean squared prediction error as a loss function, proportional hazards regression and the three generalized linear models were best able to predict cost in independent validation data. Using mean absolute error, linear regression with log-transformed cost, proportional hazards regression, and median regression to predict median cost, were best able to predict cost in independent validation data. Since the models demonstrated good consistency in identifying factors associated with increased cost of CABG surgery, any of the seven models can be used for identifying factors associated with increased cost of surgery. However, the magnitude of, and the interpretation of, the coefficients vary across models. Researchers are encouraged to consider a variety of candidate models, including those better known in the econometrics literature, rather than begin data analysis with one regression model selected a priori. The final choice of regression model should be made after a careful assessment of how best to assess predictive ability and should be tailored to the particular data in question.

Aged↗

Complex limiting behaviour of multilocus genetic systems in cyclical environments.

Here we demonstrate that complex limiting behaviour (supercycles and chaotic-like phenomena) may arise in a rather broad and natural class of multilocus systems, both haploid and diploid, experiencing stabilizing selection with cyclically varying optima over a short period. These include loci with purely additive, dominant, or semidominant effects, with different types of their chromosome distribution. The observed complex dynamics appeared to manifest a certain stability with respect to disturbances of parameters specifying the structure of the selected system and environmental characteristics. This mode of multilocus dynamics by far exceeds the potential attainable under ordinary selection models resulting in simple behaviour. It may represent a novel evolutionary mechanism increasing genetic diversity over long time periods. This novel mechanism could contribute to the observation that biological diversity has increased over geological time regardless of the well-known massive extinctions.

Animals↗

Stationary gene frequency distribution in the environment fluctuating between two distinct states.

A general method is given to obtain a stationary distribution in a "stochastic" one-dimensional dynamical system in which an environmental parameter specifying the dynamical system is a stationary Markov process with only two states. By applying this method, the exact stationary gene frequency distribution is obtained for a genic selection model in the environment fluctuating between two distinct states. Several limiting stationary distributions are obtained therefrom, and one of them is shown to coincide with a stationary solution of the diffusion equation heuristically derived by us for more general cases. Discussion is given on the relationship between the diffusion equations obtained by various authors starting from discrete, non-overlapping generation models.

Animals↗

A simulation support system for solving large physiological models on microcomputers.

Although physiological modeling and computer simulation have become useful research tools to test new scientific theories and to design and analyze laboratory experiments, developing a new model can be a tedious process because the investigator must often write very complex and specific routines for data input and output. To facilitate the design of new models (as well as the use of existing models), we have developed MODSIM, a FORTRAN-based simulation support system for the IBM PC computer than can accommodate very large dynamic models having up to several thousand equations. It provides the investigator with utilities for continuous on-line graphical and/or tabular output, as well as facilities for dynamic interaction with the model. The user must only supply a model as a list of mathematical equations written in FORTRAN, along with the initial values of the model variables and parameters. The model is precompiled, compiled, and then linked to the MODSIM utilities. Without further programming, the user can then solve the model, select variables for graphical output, and stop the model at any time to analyze the data or to change a parameter before resuming the simulation. This simulation system makes it very easy to develop new models that actively interact with the experimental research of the investigator.

Computer Simulation↗

Examining the influence of drop-outs in a follow-up of maintained opiate users.

INTRODUCTION: In most longitudinal studies of problem opiate users, drop-outs are frequent, but not taken into account. However, missing data can induce important bias in parameters estimates. OBJECTIVE: The aim of this study was to examine the influence of drop-outs in the statistical analysis of a follow-up of opiate users in maintenance treatment. METHODS: Participants were 519 patients who had sought maintenance treatment between 1994 and 2001. Drug use was studied using the drug composite score of the Addiction Severity Index. A classical data analysis (linear mixed effects model for repeated measurements) was compared with a selection model, which consists, in this case, of a joint modelling of the score and of the drop-out probability in order to reduce bias induced by drop-outs. RESULTS: At 18 months, 38% of the patients were available for evaluation. Drop-outs were associated with low drug use and were informative. Each model showed that the score decreased over time and that it was associated with psychiatric problems. Unlike the classical method, the joint model showed no significant association between the score and age or treatment setting. CONCLUSIONS: These results show the importance of accounting for informative drop-outs in data analysis before drawing conclusions from such studies.

Adult↗

Molecular phylogeny of coleoid cephalopods (Mollusca: Cephalopoda) using a multigene approach; the effect of data partitioning on resolving phylogenies in a Bayesian framework.

The resolution of higher level phylogeny of the coleoid cephalopods (octopuses, squids, and cuttlefishes) has been hindered by homoplasy among morphological characters in conjunction with a very poor fossil record. Initial molecular studies, based primarily on small fragments of single mitochondrial genes, have produced little resolution of the deep relationships amongst coleoid cephalopod families. The present study investigated this issue using 3415 base pairs (bp) from three nuclear genes (octopine dehydrogenase, pax-6, and rhodopsin) and three mitochondrial genes (12S rDNA, 16S rDNA, and cytochrome oxidase I) from a total of 35 species (including representatives of each of the higher level taxa). Bayesian analyses were conducted on mitochondrial and nuclear genes separately and also all six genes together. Separate analyses were conducted with the data partitioned by gene, codon/rDNA, gene+codon/rDNA or not partitioned at all. In the majority of analyses partitioning the data by gene+codon was the appropriate model with partitioning by codon the second most selected model. In some instances the topology varied according to the model used. Relatively high posterior probabilities and high levels of congruence were present between the topologies resulting from the analysis of all Octopodiform (octopuses and vampire "squid") taxa for all six genes, and independently for the datasets of mitochondrial and nuclear genes. In contrast, the highest levels of resolution within the Decapodiformes (squids and cuttlefishes) resulted from analysis of nuclear genes alone. Different higher level Decapodiform topologies were obtained through the analysis of only the 1st+2nd codon positions of nuclear genes and of all three codon positions. It is notable that there is strong evidence of saturation among the 3rd codon positions within the Decapodiformes and this may contribute spurious signal. The results suggest that the Decapodiformes may have radiated earlier and/or had faster rates of evolution than the Octopodiformes. The following taxonomic conclusions are drawn from our analyses: (1) the order Octopoda and suborders Cirrata, Incirrata, and Oegopsida are monophyletic groups; (2) the family Spirulidae (Ram's horn squids) are the sister taxon to the family Sepiidae (cuttlefishes); (3) the family Octopodidae, as currently defined, is paraphyletic; (4) the superfamily Argonautoidea are basal within the suborder Incirrata; and (5) the benthic octopus genera Benthoctopus and Enteroctopus are sister taxa.

Animals↗

Regarding the sources of data analyzed with quantitative structure-skin permeability relationship methods (commentary on 'Investigation of the mechanism of flux across human skin in vitro by quantitative structure-permeability relationships').

We investigated the sources of data used in recently published predictive models of skin permeability. It was found that skin permeability coefficients for 63 compounds are poorly documented. We hypothesized that these coefficients were calculated using the simple two variable, three parameter 'Potts and Guy' regression equation and hence were not derived from experimental measurements. We therefore examined the distribution of residuals of these reported coefficients compared with the Potts and Guy predictions. The residuals cannot be described by a normal distribution. A substantial (51%) number of residuals equaled 0.00. Further analysis demonstrated that 89% (56 out of 63) of the skin permeability coefficients can be explained as being calculated by the Potts and Guy equation using different documented octanol-water partition coefficients, and/or transcription errors. The results strongly suggest that these 63 skin permeability coefficients are calculated and not experimentally determined-a conclusion subsequently confirmed by one of the developers of the data set. Continued use of these data would lead to biased model selection, underestimation of experimental variability, and overestimation of model predictive ability.

Linear Models↗

Libraries of multifunctional RNA conjugates for the selection of new RNA catalysts.

An in vitro selection system was developed for the selection of RNA molecules catalyzing bimolecular reactions between small reactants. The system is based on the direct selection protocol and involves libraries of multifunctional RNA conjugates rather than unmodified RNA transcripts. For the preparation of RNA conjugate libraries, a dinucleotide analog has been designed and synthesized containing a poly(ethylene glycol) linker with an embedded photocleavage site and a terminal attachment site for coupling potential reactants. Reactants are first coupled to the dinucleotide analog by activated ester chemistry and then ligated to the 3'-ends of enzymatically prepared RNA pool molecules, giving libraries of complex conjugates. Species that become attached to biotin on incubation with a biotinylated partner are isolated using streptavidin-derivatized matrices and then subjected to a photocleavage step. Selective cleavage of the linker releases only those RNA species in which reaction has taken place at the linker-coupled reactant, while products with the biotin attached to internal positions of the RNA part remain immobilized. Efficient photocleavage is achieved by laser irradiation at 355 nm, and the released RNAs are intact and amplifiable by reverse transcription. All steps are shown to be compatible with the overall selection procedure, as was shown by performing a model selection cycle. Besides allowing a broader scope of reaction types to be selected for, the strategy relieves the RNA from the requirement to possess substrate properties as well as catalytic activity, and the use of a cleavable linker will suppress the selection of catalysts for side reactions.

Photolysis↗

The latent structure of substance use among American Indian adolescents: an example using categorical variables.

Researchers frequently have only categorical data to analyze and cannot, for theoretical or methodological reasons, assume that the observed variables are discrete representations of an underlying continuous variable. We present latent class analysis as an alternative method of measuring latent variables in these circumstances. Latent class analysis does not require the assumptions of factor analyses about the nature of manifest and latent variables, but does allow the use of more precise model selection than techniques such as cluster analysis. We modeled the lifetime substance use of American Indian youth. The latent class model of American Indian teenagers' substance use had four classes: Abstaining, Predominantly Alcohol, Predominantly Alcohol and Marijuana, and Plural Substance. We then demonstrated the usefulness of this latent variable by using it to differentiate levels of several variables in a manner consistent with Social Cognitive Theory.

Adolescent↗

Evolution of advantageous alleles affecting population ecological characteristics in partially inbreeding populations.

The fate of advantageous alleles affecting intrinsic growth rate, carrying capacity or intra-specific competitive ability was examined in a partially inbreeding population. Generally, inbreeding had an effect on the evolution of advantageous alleles affecting population ecological characteristics. For example, in a specific underdominant case the number of stable internal equilibria decreased from two to one with only a slight degree of inbreeding. Equilibrium frequencies of stable internal equilibria and stability of fixation equilibria were also affected by the degree of inbreeding. For strictly advantageous alleles, inbreeding had the same qualitative effect on the fixation probability and mean fixation time as predicted in simpler selection models.

Alleles↗

Monthly model for genetic evaluation of laying hens. II. Random regression.

1. We investigated the use of monthly production records for genetic evaluation of laying hens, derived from a test day model with random regression in dairy cattle and compared it with other models. 2. Records of 6450 hens, daughters of 180 sires and 1335 dams, were analysed using a model with restricted maximum likelihood (REML): traits considered were monthly and cumulative egg production. Five models were studied: (1) random regression with covariates derived from the regression of Ali and Schaeffer (Canadian Journal of Animal Science, 67: 637-644, 1987) (RRMAS), (2) random regression with covariates derived from quartic polynomial (RRMP4), (3) fixed regression with covariates derived from Ali and Schaeffer (FRM), (4) multiple trait (MTM) and (5) cumulative (CM). 3. The models were compared on the basis of Spearman rank correlations of individual breeding values and sire breeding values estimated from subsets of full-sib split data. The hens (about 10% per generation) which ranked highest on their estimated breeding values from different models were compared phenotypically with their full records. 4. The estimates of heritability resulting from RRMP4 were biased upward from the estimates obtained from MTM, so this model was discarded. The heritabilities for monthly productions from RRMAS and MTM showed a similar pattern. They were high for the 1st month of production, decreased to their lowest value at about month 5 of production and increased again to the end of lay. 5. Spearman rank correlations between animal breeding values estimated by monthly models (RRMAS, FRM and MTM) were high, between 0.91 and 0.98, whereas those between estimates of monthly models and CM were lower, from 0.85 to 0.87. The correlations estimated either from intermittent months of measurements (odd vs even months) or full records were generally high, from 0.93 to 0.99. Information from odd months of production could be sufficient for cost-efficient recording schemes. The RRMAS generally had the highest correlation of sire breeding values between subsets of full-sib records, followed by MTM, RM and CM. Monthly models selected hens with higher productivity than the cumulative model. 6. In conclusion, genetic evaluation based on monthly production may be better than using cumulative production and RRMAS appeared to be the best among the models tested here.

Animals↗

Robust classification modeling on microarray data using misclassification penalized posterior.

MOTIVATION: Genome-wide microarray data are often used in challenging classification problems of clinically relevant subtypes of human diseases. However, the identification of a parsimonious robust prediction model that performs consistently well on future independent data has not been successful due to the biased model selection from an extremely large number of candidate models during the classification model search and construction. Furthermore, common criteria of prediction model performance, such as classification error rates, do not provide a sensitive measure for evaluating performance of such astronomic competing models. Also, even though several different classification approaches have been utilized to tackle such classification problems, no direct comparison on these methods have been made. RESULTS: We introduce a novel measure for assessing the performance of a prediction model, the misclassification-penalized posterior (MiPP), the sum of the posterior classification probabilities penalized by the number of incorrectly classified samples. Using MiPP, we implement a forward step-wise cross-validated procedure to find our optimal prediction models with different numbers of features on a training set. Our final robust classification model and its dimension are determined based on a completely independent test dataset. This MiPP-based classification modeling approach enables us to identify the most parsimonious robust prediction models only with two or three features on well-known microarray datasets. These models show superior performance to other models in the literature that often have more than 40-100 features in their model construction. AVAILABILITY: Our MiPP software program is available at the Bioconductor website (http://www.bioconductor.org).

Algorithms↗

Context-specific independence mixture modeling for positional weight matrices.

MOTIVATION: A positional weight matrix (PWM) is a statistical representation of the binding pattern of a transcription factor estimated from known binding site sequences. Previous studies showed that for factors which bind to divergent binding sites, mixtures of multiple PWMs increase performance. However, estimating a conventional mixture distribution for each position will in many cases cause overfitting. RESULTS: We propose a context-specific independence (CSI) mixture model and a learning algorithm based on a Bayesian approach. The CSI model adjusts complexity to fit the amount of variation observed on the sequence level in each position of a site. This not only yields a more parsimonious description of binding patterns, which improves parameter estimates, it also increases robustness as the model automatically adapts the number of components to fit the data. Evaluation of the CSI model on simulated data showed favorable results compared to conventional mixtures. We demonstrate its adaptive properties in a classical model selection setup. The increased parsimony of the CSI model was shown for the transcription factor Leu3 where two binding-energy subgroups were distinguished equally well as with a conventional mixture but requiring 30% less parameters. Analysis of the human-mouse conservation of predicted binding sites of 64 JASPAR TFs showed that CSI was as good or better than a conventional mixture for 89% of the TFs and for 70% for a single PWM model. AVAILABILITY: http://algorithmics.molgen.mpg.de/mixture.

Algorithms↗

Deleterious background selection with recombination.

An analytic expression for the expected nucleotide diversity is obtained for a neutral locus in a region with deleterious mutation and recombination. Our analytic results are used to predict levels of variation for the entire third chromosome of Drosophila melanogaster. The predictions are consistent with the low levels of variation that have been observed at loci near the centromeres of the third chromosome of D. melanogaster. However, the low levels of variation observed near the tips of this chromosome are not predicted using currently available estimates of the deleterious mutation rate and of selection coefficients. If considerably smaller selection coefficients are assumed, the low observed levels of variation at the tips of the third chromosome are consistent with the background selection model.

Animals↗