Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Estimation and model selection in constrained deconvolution.

We analyze in detail the estimation problem associated with the following problem. Given n noisy measurements (yi, i = 1, ..., n) of the response of a system to an input (A(t) where t indicates time), obtain an estimate of A(t) given a known K(t) (the unit impulse response function of the system) under the model: yi = integral of 0(ti) A(s)K(ti - s)ds + epsilon i where epsilon 1, ... ,epsilon n are independent identically distributed random variables with mean zero and common finite variance. In the solution to the problem, the unknown function is represented by a spline function, and the problem is recast in terms of (inequality constrained) linear regression. The main issues addressed are: (a) the comparison of different nonparametric regression methods in this context, and (b) how to do model selection, i.e., given a (finite) set of candidate spline functions, select the (possibly unique) best one using some (statistically based) selection criteria. Different spline candidate sets, and different asymptotic and resampling-based statistical selection criteria are compared by means of simulations. Due to the particular nature of the estimation problem, modifications to the criteria are suggested. Applications to simulated and real pharmacokinetics data are reported.

Animals↗

Modelling hybridoma cell growth and metabolism--a comparison of selected models and data.

Unstructured models for cell growth (cell specific growth and death rate) and metabolism (cell specific substrate uptake and metabolite production rates) of hybridoma cell lines were compared with special respect to significance, analytical error and range of validity. The diversity of the unstructured models cited reveals their mostly descriptive character compared to structured models. Bearing in mind this limited knowledge, empirical models can still serve as a valuable tool for process design. For understanding of the cell metabolism itself they might have been overemphasized in the past. For proper model design, care has to be taken to cover the whole range of process conditions. In particular if a process is to be run at very low substrate and high metabolite concentrations, chemostat cultures which have mostly been used for the model formulations, are not sufficient and have to be completed by, for example, fed-batch cultures.

Ammonia↗

Model inference or model selection: discussion of Klugkist, Laudy, and Hoijtink (2005).

I. Klugkist, O. Laudy, and H. Hoijtink (2005) presented a Bayesian approach to analysis of variance models with inequality constraints. Constraints may play 2 distinct roles in data analysis. They may represent prior information that allows more precise inferences regarding parameter values, or they may describe a theory to be judged against the data. In the latter case, the authors emphasized the use of Bayes factors and posterior model probabilities to select the best theory. One difficulty is that interpretation of the posterior model probabilities depends on which other theories are included in the comparison. The posterior distribution of the parameters under an unconstrained model allows one to quantify the support provided by the data for inequality constraints without requiring the model selection framework.

Analysis of Variance↗

Log-linear model selections in a rural dental health study.

This field study sought to measure the effects of dental delivery and school-based, dental health education on use of dental health care by children in grades K-6. We attempted to control for two potential confounding factors by an approximate randomization of children into treatment groups with stratification on grade and initial oral disease levels. A backward elimination log-linear model selection procedure for the 5-factor classification permitted tests for higher-order interaction, namely effect-modification, confounding and collapsibility. We found that the effect of dental health education on use of dental care depended on the mode of dental delivery.

Child↗

Model selection in binary trait locus mapping.

Quantitative trait locus (QTL) mapping methodology for continuous normally distributed traits is the subject of much attention in the literature. Binary trait locus (BTL) mapping in experimental populations has received much less attention. A binary trait by definition has only two possible values, and the penetrance parameter is restricted to values between zero and one. Due to this restriction, the infinitesimal model appears to come into play even when only a few loci are involved, making selection of an appropriate genetic model in BTL mapping challenging. We present a probability model for an arbitrary number of BTL and demonstrate that, given adequate sample sizes, the power for detecting loci is high under a wide range of genetic models, including most epistatic models. A novel model selection strategy based upon the underlying genetic map is employed for choosing the genetic model. We propose selecting the "best" marker from each linkage group, regardless of significance. This reduces the model space so that an efficient search for epistatic loci can be conducted without invoking stepwise model selection. This procedure can identify unlinked epistatic BTL, demonstrated by our simulations and the reanalysis of Oncorhynchus mykiss experimental data.

Animals↗

Inferring the location and effect of tumor suppressor genes by instability-selection modeling of allelic-loss data.

Cancerous tumor growth creates cells with abnormal DNA. Allelic-loss experiments identify genomic deletions in cancer cells, but sources of variation and intrinsic dependencies complicate inference about the location and effect of suppressor genes; such genes are the target of these experiments and are thought to be involved in tumor development. We investigate properties of an instability-selection model of allelic-loss data, including likelihood-based parameter estimation and hypothesis testing. By considering a special complete-data case, we derive an approximate calibration method for hypothesis tests of sporadic deletion. Parametric bootstrap and Bayesian computations are also developed. Data from three allelic-loss studies are reanalyzed to illustrate the methods.

Adenocarcinoma↗

A treatment selection model for weight reduction in adults with acquired brain injury: applications and preliminary findings.

This article presents a unique method for providing weight management assistance to persons who have experienced an acquired brain injury (ABI). Most of the available literature on this topic deals with weight loss methods for individuals who are not faced with the cognitive and behavioural challenges inherent in this population. A treatment selection protocol will be described that allows for appropriate selection of behavioural weight loss interventions. Interventions are based upon specific cognitive and behavioural difficulties that individuals with acquired brain injury may present. A detailed case study will also be presented depicting successful use of the treatment selection model with an adult male with an acquired brain injury.

Adult↗

Model selection methodology for inter-laboratory standardisation of antibody titres.

The aim of the European Sero-Epidemiology Network 2 (ESEN2) is to compare standardised serological results of vaccine preventable diseases across Europe. In order to adjust for laboratory and assay differences, the participant laboratories tested a standardisation panel for each disease and the results were plotted against those of a reference centre in order to obtain standardisation equations. We describe an algorithm for obtaining these equations which took into account outliers, censored data and model selection, and needed to be robust due to the variety of assays and diseases under investigation. If the standardisation equations explained sufficient variability, these were used to convert each country's main serum bank results to a common unitage, thereby allowing international comparisons to be made.

Algorithms↗

Quantitative trait nucleotide analysis using Bayesian model selection.

Although much attention has been given to statistical genetic methods for the initial localization and fine mapping of quantitative trait loci (QTLs), little methodological work has been done to date on the problem of statistically identifying the most likely functional polymorphisms using sequence data. In this paper we provide a general statistical genetic framework, called Bayesian quantitative trait nucleotide (BQTN) analysis, for assessing the likely functional status of genetic variants. The approach requires the initial enumeration of all genetic variants in a set of resequenced individuals. These polymorphisms are then typed in a large number of individuals (potentially in families), and marker variation is related to quantitative phenotypic variation using Bayesian model selection and averaging. For each sequence variant a posterior probability of effect is obtained and can be used to prioritize additional molecular functional experiments. An example of this quantitative nucleotide analysis is provided using the GAW12 simulated data. The results show that the BQTN method may be useful for choosing the most likely functional variants within a gene (or set of genes). We also include instructions on how to use our computer program, SOLAR, for association analysis and BQTN analysis.

Bayes Theorem↗

Mutation-selection models solved exactly with methods of statistical mechanics.

We reconsider deterministic models of mutation and selection acting on populations of sequences, or, equivalently, multilocus systems with complete linkage. Exact analytical results concerning such systems are few, and we present recent and new ones obtained with the help of methods from quantum statistical mechanics. We consider a continuous-time model for an infinite population of haploids (or diploids without dominance), with N sites each, two states per site, symmetric mutation and arbitrary fitness function. We show that this model is exactly equivalent to a so-called Ising quantum chain. In this picture, fitness corresponds to the interaction energy of spins, and mutation to a temperature-like parameter. The highly elaborate methods of statistical mechanics allow one to find exact solutions for non-trivial examples. These include quadratic fitness functions, as well as 'Onsager's landscape'. The latter is a fitness function which captures some essential features of molecular evolution, such as neutrality, compensatory mutations and flat ridges. We investigate the mean number of mutations, the mutation load, and the variance in fitness under mutation-selection balance. This also yields some insight into the 'error threshold' phenomenon, which occurs in some, but not all, examples.

Evolution, Molecular↗

Population genetics of polymorphism and divergence for diploid selection models with arbitrary dominance.

We develop a Poisson random-field model of polymorphism and divergence that allows arbitrary dominance relations in a diploid context. This model provides a maximum-likelihood framework for estimating both selection and dominance parameters of new mutations using information on the frequency spectrum of sequence polymorphisms. This is the first DNA sequence-based estimator of the dominance parameter. Our model also leads to a likelihood-ratio test for distinguishing nongenic from genic selection; simulations indicate that this test is quite powerful when a large number of segregating sites are available. We also use simulations to explore the bias in selection parameter estimates caused by unacknowledged dominance relations. When inference is based on the frequency spectrum of polymorphisms, genic selection estimates of the selection parameter can be very strongly biased even for minor deviations from the genic selection model. Surprisingly, however, when inference is based on polymorphism and divergence (McDonald-Kreitman) data, genic selection estimates of the selection parameter are nearly unbiased, even for completely dominant or recessive mutations. Further, we find that weak overdominant selection can increase, rather than decrease, the substitution rate relative to levels of polymorphism. This nonintuitive result has major implications for the interpretation of several popular tests of neutrality.

Computer Simulation↗

A frequency-dependent natural selection model for the evolution of social cooperation networks.

A model is presented for the evolution of several aspects of sociality based on reciprocal ties of social cooperation, modeling especially cooperative hunting behavior in carnivores. This model captures the possibility of a critical threshold in gene frequency, which, if reached, will lead to an explosion toward fixation of the "social" trait. This threshold phenomenon might be restated as follows: the precondition for evolution favorable to the specific form of social behavior considered is hard to satisfy, but-once this condition is satisfied-the tendency toward sociality is effectively irreversible. The simple model proposed appears to be highly robust, with most realistic changes additionally favoring the social gene.

Alleles↗

Model selection methodology in supervised learning with evolutionary computation.

The expressive power, powerful search capability, and the explicit nature of the resulting models make evolutionary methods very attractive for supervised learning applications in bioinformatics. However, their characteristics also make them highly susceptible to overtraining or to discovering chance relationships in the data. Identification of appropriate criteria for terminating evolution and for selecting an appropriately validated model is vital. Some approaches that are commonly applied to other modelling methods are not necessarily applicable in a straightforward manner to evolutionary methods. An approach to model selection is presented that is not unduly computationally intensive. To illustrate the issues and the technique two bioinformatic datasets are used, one relating to metabolite determination and the other to disease prediction from gene expression data.

Algorithms↗

The time of maximum effect for model selection in pharmacokinetic-pharmacodynamic analysis applied to frusemide.

AIMS: Both indirect-response models and effect-compartment models are used to describe the pharmacodynamics of drugs when there is a delay in the time course of the pharmacological effect in relation to the concentration of the drug. The aim of this study was to investigate whether the time of maximum response after single-dose administration at different dose levels could be used to distinguish between these models and to select the most appropriate pharmacokinetic-pharmacodynamic model for frusemide. METHODS: Three doses of frusemide, 10, 25 and 40 mg were given as rapid intravenous infusions to five healthy volunteers. Urine samples were collected for 5 h after dosing. Volume and sodium losses were isovolumetrically replaced with an intravenous rehydration fluid. Diuresis and natriuresis were modelled for all three doses simultaneously, applying both an indirect-response model and an effect-compartment model with the frusemide excretion rate as the pharmacokinetic input. RESULTS: The observed time of maximum diuretic and natriuretic response significantly increased with dose. This increase was well predicted by the indirect-response model, whereas the modelling with the effect-compartment model led to a poor prediction of the peaks. There was no difference between the observed and predicted time of maximum diuretic and natriuretic response using the indirect-response model, whereas the time of maximum response predicted by the effect-compartment model was significantly earlier than the time observed for the 25 mg (P < 0.05) and 40 mg (P < 0.05) doses. CONCLUSIONS: The time of maximum response to frusemide was better described using an indirect-response model than an effect-compartment model. Studying the time of maximum response after administration of different single doses of a drug may be used as a selective tool during pharmacokinetic-pharmacodynamic modelling.

Adult↗

Model selection in covariance structures analysis and the "problem" of sample size: a clarification.

Complex models for covariance matrices are structures that specify many parameters, whereas simple models require only a few. When a set of models of differing complexity is evaluated by means of some goodness of fit indices, structures with many parameters are more likely to be selected when the number of observations is large, regardless of other utility considerations. This is known as the sample size problem in model selection decisions. This article argues that this influence of sample size is not necessarily undesirable. The rationale behind this point of view is described in terms of the relationships among the population covariance matrix and 2 model-based estimates of it. The implications of these relationships for practical use are discussed.

Analysis of Variance↗

SNPs, haplotypes, and model selection in a candidate gene region: the SIMPle analysis for multilocus data.

Modern molecular techniques make discovery of numerous single nucleotide polymorphims (SNPs) in candidate gene regions feasible. Conventional analysis relies on either independent tests with each variant or the use of haplotypes in association analysis. The first technique ignores the dependencies between SNPs. The second, though it may increase power, often introduces uncertainty by estimating haplotypes from population data. Additionally, as the number of loci expands for a haplotype, ambiguity in interpretation increases for determining the underlying genetic components driving a detected association. Here, we present a genotype-level analysis to jointly model the SNPs via a SNP interaction model with phase information (SIMPle) to capture the underlying haplotype structure. This analysis estimates both the risk associated with each variant and the importance of phase between pairwise combinations of SNPs. Thus, rather than selecting between genotype- or haplotype-level approaches, the SIMPle method frames the analysis of multilocus data in a model selection paradigm, the aim to determine which SNPs, phase terms, and linear combinations best describe the relation between genetic variation and a trait of interest. To avoid unstable estimation due to sparse data and to incorporate both the dependencies among terms and the uncertainty in model selection, we propose a Bayes model averaging procedure. This highlights key SNPs and phase terms and yields a set of best representative models. Using simulations, we demonstrate the utility of the SIMPle model to identify crucial SNPs and underlying haplotype structures across a variety of causal models and genetic architectures.

Bayes Theorem↗