Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Investigation on the improvement of prediction by bootstrap model averaging.

OBJECTIVES: We illustrate a recently proposed two-step bootstrap model averaging (bootstrap MA) approach to cope with model selection uncertainty. The predictive performance is investigated in an example and in a simulation study. Results are compared to those derived from other model selection methods. METHODS: In the framework of the linear regression model we use the two-step bootstrap MA, which consists of a screening step to eliminate covariates thought to have no influence on the response, and a model-averaging step. We also apply the full model, variable selection using backward elimination based on Akaike's Information Criterion (AIC), the Bayes Information Criterion (BIC) and the bagging approach. The predictive performance is measured by the mean squared error (MSE) and the coverage of confidence intervals for the true response. RESULTS: We obtained similar results for all approaches in the example. In the simulation the MSE was reduced by all approaches in comparison to the full model. The smallest values are obtained for bootstrap MA. Only the bootstrap MA and the full model correctly estimated the nominal coverage. The backward elimination procedures led to substantial underestimation and bagging to an overestimation of the true coverage. The screening step of bootstrap MA eliminates most of the unimportant factors. CONCLUSION: The new bootstrap MA approach shows promising results for predictive performance. It increases practical usefulness by eliminating unimportant factors in the screening step.

Body Composition↗

Statistical test to compare the linkage model and the admixture model based on central limit results.

In the Admixture Model, the probability that an individual carries a certain allele at a specific marker depends on the allele frequencies in K ancestral populations and the proportion of the individual's genome originating from these populations. The markers are assumed to be independent. The Linkage Model is a Hidden Markov Model that extends the Admixture Model by incorporating linkage between neighboring loci. We prove consistency and asymptotic normality of maximum likelihood estimators for the ancestry of individuals in the Linkage Model, complementing earlier results by (Pfaff et al., 2004; Pfaffelhuber and Rohde, 2022; Heinzel, 2025) for the Admixture Model. These results are used to prove that a statistical test that allows for model selection between the Admixture Model and the Linkage Model is an asymptotic level-α-test. Finally, we demonstrate the practical relevance of our results by applying the test to real-world data from The 1000 Genomes Project Consortium (2015).

Genetic Linkage↗

The significance of over- and underdominance for the maintenance of genetic polymorphisms. I. Underdominance and stability.

It is pointed out that the standard selection models in population genetics all require some form of heterozygote advantage in fitness in order to guarantee the maintenance or stability of genetic polymorphisms. Even more recent results demonstrating the existence of stable two-locus polymorphisms with marginal underdominance at both loci are based on certain epistatically acting heterosis assumptions. This raises the question as to whether heterozygote advantage in fitness is indeed a generally valid principle of maintaining polymorphisms. To avoid ambiguity in definition of heterozygote advantage (overdominance) as it appears in multiallele or multilocus systems, a one-locus-two-allele model is considered. This model allows for sexually asymmetric selection and random mating. It is shown that the model produces globally stable polymorphisms exhibiting underdominance in fitness for a considerable and biologically reasonable range of selection values. Having thus properly refuted the general validity of the common overdominance principle, a modified version is suggested which covers the classical viability selection model and its extension to arbitrary, sexually asymmetric viability and fertility selection. This modified overdominance principle is based on the notion of fractional fitnesses and relates protectedness of biallelic polymorphisms to the extent to which each genotype reproduces its own type. The fact that the model treated displays frequency dependent fitnesses which may change in ranking while approaching equilibrium is discussed in relation to problems of the evolution of overdominance and underdominance.

Animals↗

ANN-QSAR model for selection of anticancer leads from structurally heterogeneous series of compounds.

Developing a model for predicting anticancer activity of any classes of organic compounds based on molecular structure is very important goal for medicinal chemist. Different molecular descriptors can be used to solve this problem. Stochastic molecular descriptors so-called the MARCH-INSIDE approach, shown to be very successful in drug design. Nevertheless, the structural diversity of compounds is so vast that we may need non-linear models such as artificial neural networks (ANN) instead of linear ones. SmartMLP-ANN analysis used to model the anticancer activity of organic compounds has shown high average accuracy of 93.79% (train performance) and predictability of 90.88% (validation performance) for the 8:3-MLP topology with different training and predicting series. This ANN model favourably compares with respect to a previous linear discriminant analysis (LDA) model [H. González-Díaz et al., J. Mol. Model 9 (2003) 395] that showed only 80.49% of accuracy and 79.34% of predictability. The present SmartMLP approach employed shorter training times of only 10h while previous models give accuracies of 70-89% only after 25-46 h of training. In order to illustrate the practical use of the model in bioorganic medicinal chemistry, we report the in silico prediction, and in vitro evaluation of six new synthetic tegafur analogues having IC(50) values in a broad range between 37.1 and 138 microgmL(-1) for leukemia (L1210/0) and human T-lymphocyte (Molt4/C8, CEM/0) cells. Theoretical predictions coincide very well with experimental results.

Animals↗

Evidence that positive selection drives Y-chromosome degeneration in Drosophila miranda.

Why does the Y chromosome harbor so few functional loci? Evolutionary theory predicts that Y chromosomes degenerate because they lack genetic recombination. Both positive and negative selection models have been invoked to explain this degeneration, as both can result in the recurrent fixation of linked deleterious mutations on a nonrecombining Y chromosome. To distinguish between these models, I investigated patterns of nucleotide variability along 37 kb of the recently formed neo-Y chromosome in Drosophila miranda. Levels of nucleotide variability on this chromosome are 30 times lower than in highly recombining portions of the genome. Both positive and negative selection models can result in reduced variability levels, but their effects on the frequency spectrum of mutations differ. Using coalescent simulations, I show that the patterns of nucleotide variability on the neo-Y chromosome are unlikely under deleterious mutation models (including background selection and Muller's ratchet) but are expected under recent positive selection. These results implicate positive selection as an important force driving the degeneration of Y chromosomes; adaptation at a few loci, possibly increasing male fitness, occurs at the cost of most other genes on this chromosome.

Animals↗

Regression-based sampling for persons with high health expenditures: evaluating accuracy and yield with the 1997 MEPS.

BACKGROUND: Given the high concentration of health care expenditures among a relatively small percentage of the population, the 1997 Medical Expenditure Panel Survey was designed to learn more about these high expenditure individuals by oversampling them. OBJECTIVE: Oversampling high expenditure individuals enables more precise estimation of what the nation's health care dollar buys and who pays it. It also enhances the ability to discern the causes of high health care expenses and the characteristics of the individuals who incur them. METHOD: Using the 1987 National Medical Expenditure Survey, a probabilistic model was developed to select households from the 1996 National Health Interview Survey likely to contain individuals incurring high levels of medical expenditures in the 1997 MEPS. The accuracy of the selection model, and the degree to which the high expenditure population was oversampled, are assessed with the 1997 MEPS data. RESULTS: Over half of the persons selected by the regression model were expected to have high health expenditures. Of the 456 persons selected by the model for oversampling, 257 individuals or 56.4% did, in fact, have high expenditures. Regression-based sampling increased the proportion of MEPS individuals with high expenditures from 14.3% without oversampling to 17.2% of the total cohort with oversampling (or from 938-1,126 persons). CONCLUSION: This paper demonstrates that a model-based approach to oversampling a high expenditure population, or any population with dynamic characteristics, can be highly successful in terms of sampling yield and accuracy.

Adult↗

Frequentist model-averaged estimators and tests for univariate twin models.

Parameter estimates from analyses of univariate twin data usually do not reflect the uncertainty due to the model selection phase of the data analysis. To address the effect of model selection uncertainty on parameter estimates, we introduce frequentist model-averaged estimators for univariate twin data analysis that use information-theoretic criteria to assign model weights. We conduct simulation studies to examine the performance of model-averaged estimators of additive genetic variance, and for tests for additive genetic variance based on model-averaged estimators. In simulation studies with small or moderate sample sizes, model-averaged estimators of additive genetic variance typically have lower mean-squared error than either (i) estimators from individual twin models, or (ii) estimators obtained from a decision procedure where the best-fitting model from likelihood-ratio testing is used to estimate additive genetic variance. For each sample size simulated, bootstrap tests based on model-averaged estimators have higher power to detect additive genetic variance than currently-used tests in most cases.

Analysis of Variance↗

The tyrp1-Tag/tyrp1-FGFR1-DN bigenic mouse: a model for selective inhibition of tumor development, angiogenesis, and invasion into the neural tissue by blockade of fibroblast growth factor receptor activity.

We describe herein a new transgenic mouse tumor model in which fibroblast growth factor (FGF) receptor activity is selectively inhibited. Tyrp1-Tag mice that develop early vascularized tumors of the retinal pigment epithelium were crossed with tyrp1-FGFR1-DN mice that express dominant-negative FGF receptors in the retinal pigment epithelium to generate bigenic mice. Initial angiogenesis-independent tumor growth progressed equally in tyrp1-Tag and bigenic mice with no significant differences in the number of dividing and apoptotic cells within the tumor. By contrast, at a later stage when tyrp1-Tag tumors rapidly expanded to fill the entire eye posterior chamber and migrate along the optic nerve toward the chiasma, bigenic tumors remained small and were poorly vascularized. Secondary tumors of small size developed in only 20% of bigenic mice by 1 month. Immunohistochemical analysis of secondary tumors from bigenic mice showed a reduction of angiogenesis and an increase in apoptosis in tumor cells. Tumor cells from bigenic mice expressed high levels of truncated FGF receptors and did not induce endothelial tube formation in vitro. All in all, this indicates that the tyrp1-Tag mouse may be a useful model to study selective tumor inhibition and the effect of antitumor therapy that targets a specific growth factor pathway. FGF receptors are required at the onset of tumor invasion and angiogenesis in ocular tumors and are good therapeutic targets in this model. The bigenic mouse may also constitute a useful model to answer more fundamental questions of cancer biology such as the mechanism of tumor escape.

Animals↗

Trophoblast-like human choriocarcinoma cells serve as a suitable in vitro model for selective cholesteryl ester uptake from high density lipoproteins.

As human choriocarcinoma cells display many of the biochemical and morphological characteristics reported for in utero invasive trophoblast cells we have studied cholesterol supply from high density lipoproteins (HDL) to these cells. Binding properties of 125I-labeled HDL subclass 3 (HDL3) at 4 degrees C were similar for BeWo, JAr, and Jeg3 choriocarcinoma cell lines while degradation rates at 37 degrees C were highest for BeWo. Calculating the selective cholesteryl ester (CE)-uptake as the difference between specific cell association of [3H]CE-labeled HDL3 and holoparticle association of 125I-labeled HDL3 revealed that in BeWo cells, the selective CE-uptake was slightly lower than holoparticle association. However, the pronounced capacity for specific cell association of [3H]CE-HDL3 and selective [3H]CE-uptake in excess of HDL3-holoparticle association, and cAMP-mediated enhanced cell association of [3H]CE-HDL3 in JAr and Jeg3 suggested the scavenger receptor class B, type I (SR-BI) to be responsible for this pathway. Abundant expression of SR-BI (but not SR-BII, a splice variant of SR-BI) could be observed in JAr and Jeg3 but not in BeWo cells using RT-PCR, Northern and Western blot analysis, and immunocytochemical technique. Adenovirus-mediated overexpression of SR-BI in all three choriocarcinoma cell lines resulted in an enhanced capacity for cell association of [3H]CE-HDL3 (20-fold in BeWo; fivefold in JAr and Jeg3). The fact that exogenous HDL3 remarkably increases proliferation in JAr and Jeg3 supports the notion that selective CE-uptake and subsequent intracellular generation of cholesterol is coupled to cellular growth. From our findings we propose that JAr and Jeg3 cells serve as a suitable in vitro model to study selective CE-supply to human placental cells.

Adenoviridae↗

Statistical limitations in functional neuroimaging. II. Signal detection and statistical inference.

The field of functional neuroimaging (FNI) methodology has developed into a mature but evolving area of knowledge and its applications have been extensive. A general problem in the analysis of FNI data is finding a signal embedded in noise. This is sometimes called signal detection. Signal detection theory focuses in general on issues relating to the optimization of conditions for separating the signal from noise. When methods from probability theory and mathematical statistics are directly applied in this procedure it is also called statistical inference. In this paper we briefly discuss some aspects of signal detection theory relevant to FNI and, in addition, some common approaches to statistical inference used in FNI. Low-pass filtering in relation to functional-anatomical variability and some effects of filtering on signal detection of interest to FNI are discussed. Also, some general aspects of hypothesis testing and statistical inference are discussed. This includes the need for characterizing the signal in data when the null hypothesis is rejected, the problem of multiple comparisons that is central to FNI data analysis, omnibus tests and some issues related to statistical power in the context of FNI. In turn, random field, scale space, non-parametric and Monte Carlo approaches are reviewed, representing the most common approaches to statistical inference used in FNI. Complementary to these issues an overview and discussion of non-inferential descriptive methods, common statistical models and the problem of model selection is given in a companion paper. In general, model selection is an important prelude to subsequent statistical inference. The emphasis in both papers is on the assumptions and inherent limitations of the methods presented. Most of the methods described here generally serve their purposes well when the inherent assumptions and limitations are taken into account. Significant differences in results between different methods are most apparent in extreme parameter ranges, for example at low effective degrees of freedom or at small spatial autocorrelation. In such situations or in situations when assumptions and approximations are seriously violated it is of central importance to choose the most suitable method in order to obtain valid results.

Biometry↗

Applications of a mixture survival model with covariates to the analysis of a depression prevention trial.

This paper presents a case study of model selection for survival analysis data. We use an approximate Bayesian method for model selection based on assessing the posterior probability of competing models given the data. We introduce the Schwarz criteria, an approximation to the logarithm of the Bayes factor, to provide an indication of evidence in favour of one model compared to another. Specifically, in the context of a depression prevention clinical trial we evaluate the efficacy of treatment in preventing or delaying the time to recurrence of depression, and evaluate how differences in the survival distributions between the two treatment groups depend on explanatory variables of interest. This investigation is based on a mixture survival model that explicitly incorporates the possibility of a surviving fraction.

Antidepressive Agents, Tricyclic↗

Demand for health care in Denmark: results of a national sample survey using contingent valuation.

In this paper we use willingness to pay (WTP) to elicit values for private insurance covering treatment for four different health problems. By way of obtaining these values, we test the viability of the contingent valuation method (CVM) and econometric techniques, respectively, as means of eliciting and analysing values from the general public. WTP responses from a Danish national sample survey, which was designed in accordance with existing guidelines, are analysed in terms of consistency and validity checks. Large numbers of zero responses are common in WTP studies, and are found here; therefore, the Heckman selectivity model and log-transformed OLS are employed. The selectivity model is rejected, but test results indicate that the lognormal model yields efficient and unbiased estimates. The results give confidence in the WTP estimates obtained and, more generally, in CVM as a means of valuing publicly provided goods and in econometrics as a tool for analysing WTP results containing many zero responses.

Adolescent↗

A model selection-based interval-mapping method for autopolyploids.

While extensive progress has been made in quantitative trait locus (QTL) mapping for diploid species, similar progress in QTL mapping for polyploids has been limited due to the complex genetic architecture of polyploids. To date, QTL mapping in polyploids has focused mainly on tetraploids with dominant and/or codominant markers. Here, we extend this view to include any even ploidy level under a dominant marker system. Our approach first selects the most likely chromosomal marker configurations using a Bayesian selection criterion and then fits an interval-mapping model to each candidate. Profiles of the likelihood-ratio test statistic and the maximum-likelihood estimates (MLEs) of parameters including QTL effects are obtained via the EM algorithm. Putative QTL are then detected using a resampling-based significance threshold, and the corresponding parental configuration is identified to be the underlying parental configuration from which the data are observed. Although presented via pseudo-doubled backcross experiments, this approach can be readily extended to other breeding systems. Our method is applied to single-dose restriction fragment autotetraploid alfalfa data, and the performance is investigated through simulation studies.

Algorithms↗

Range image segmentation using surface selection criterion.

In this paper, we address the problem of recovering the true underlying model of a surface while performing the segmentation. First, and in order to solve the model selection problem, we introduce a novel criterion, which is based on minimising strain energy of fitted surfaces. We then evaluate its performance and compare it with many other existing model selection techniques. Using this criterion, we then present a robust range data segmentation algorithm capable of segmenting complex objects with planar and curved surfaces. The presented algorithm simultaneously identifies the type (order and geometric shape) of each surface and separates all the points that are part of that surface. This paper includes the segmentation results of a large collection of range images obtained from objects with planar and curved surfaces. The resulting segmentation algorithm successfully segments various possible types of curved objects. More importantly, the new technique is capable of detecting the association between separated parts of a surface, which has the same Cartesian equation while segmenting a scene. This aspect is very useful in some industrial applications of range data analysis.

Algorithms↗

A population genetic model of selection that maintains specific trinucleotides at a specific location.

Periodic appearances of specific trinucleotides along the DNA sequence have been reported in the chicken core DNA, and the phenomenon has been suggested to be related to the supercoiling of DNA around nucleosomes. A population genetic model is constructed in which selection is operating to maintain specific trinucleotides at a specific location on the DNA sequence. Assuming low mutation rates, equilibrium probabilities of the appearances of respective trinucleotides were computed. Vague patterns appeared if the product of the effective size and the selection coefficient was 0.1-2.0. The genetic load and substitution rates in the equilibrium state were also computed. When the model was applied to the chicken DNA data, the product of the effective size and the selection coefficient was estimated to be 0.1-0.2. With this intensity of selection, the substitution rate was hardly different from that in the case without selection. However, the genetic load became fairly large. Considering the large number of times that DNA coils about nucleosomes, the number of trinucleotide sites must be very large, and thus the total load might be too large. Epistasis among these sites to reduce the total load is suggested to exist if selection is responsible for this periodic pattern observed in the chicken core DNA.

Alleles↗