Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

A comparison of nonlinear mixed-model analyses for a pediatric pharmacokinetic study.

Nonlinear mixed models are important tools for analyzing repeated measures data. In particular, these models are used for population pharmacokinetic analyses for estimating population pharmacokinetic parameters. As more clinical studies are performed for the advancement of treatment of pediatric patients, methodology is needed for comparing results from pharmacokinetic studies in pediatric patients and adult control groups. These pediatric studies introduce complexities to the design and analysis, including how analysis of sparse data affects the limitations of model selection. A case study is presented demonstrating that good communication with regulatory agencies and appropriate selection of analysis models are integral parts of completing population analyses for timely approval and labeling of drugs for treating pediatric patients.

Adult↗

Time squared: repeated measures on phylogenies.

Studies of gene expression profiles in response to external perturbation generate repeated measures data that generally follow nonlinear curves. To explore the evolution of such profiles across a gene family, we introduce phylogenetic repeated measures (PR) models. These models draw strength from 2 forms of correlation in the data. Through gene duplication, the family's evolutionary relatedness induces the first form. The second is the correlation across time points within taxonic units, individual genes in this example. We borrow a Brownian diffusion process along a given phylogenetic tree to account for the relatedness and co-opt a repeated measures framework to model the latter. Through simulation studies, we demonstrate that repeated measures models outperform the previously available approaches that consider the longitudinal observations or their differences as independent and identically distributed by using deviance information criteria as Bayesian model selection tools; PR models that borrow phylogenetic information also perform better than nonphylogenetic repeated measures models when appropriate. We then analyze the evolution of gene expression in the yeast kinase family using splines to estimate nonlinear behavior across 3 perturbation experiments. Again, the PR models outperform previous approaches and afford the prediction of ancestral expression profiles. To demonstrate PR model applicability more generally, we conclude with a short examination of variation in brain development across 4 primate species.

Algorithms↗

Generic biomass functions for Norway spruce in Central Europe--a meta-analysis approach toward prediction and uncertainty estimation.

To facilitate future carbon and nutrient inventories, we used mixed-effect linear models to develop new generic biomass functions for Norway spruce (Picea abies (L.) Karst.) in Central Europe. We present both the functions and their respective variance-covariance matrices and illustrate their application for biomass prediction and uncertainty estimation for Norway spruce trees ranging widely in size, age, competitive status and site. We collected biomass data for 688 trees sampled in 102 stands by 19 authors. The total number of trees in the "base" model data sets containing the predictor variables diameter at breast height (D), height (H), age (A), site index (SI) and site elevation (HSL) varied according to compartment (roots: n = 114, stem: n = 235, dry branches: n = 207, live branches: n = 429 and needles: n = 551). "Core" data sets with about 40% fewer trees could be extracted containing the additional predictor variables crown length and social class. A set of 43 candidate models representing combinations of lnD, lnH, lnA, SI and HSL, including second-order polynomials and interactions, was established. The categorical variable "author" subsuming mainly methodological differences was included as a random effect in a mixed linear model. The Akaike Information Criterion was used for model selection. The best models for stem, root and branch biomass contained only combinations of D, H and A as predictors. More complex models that included site-related variables resulted for needle biomass. Adding crown length as a predictor for needles, branches and roots reduced both the bias and the confidence interval of predictions substantially. Applying the best models to a test data set of 17 stands ranging in age from 16 to 172 years produced realistic allocation patterns at the tree and stand levels. The 95% confidence intervals (% of mean prediction) were highest for crown compartments (approximately +/- 12%) and lowest for stem biomass (approximately +/- 5%), and within each compartment, they were highest for the youngest and oldest stands, respectively.

Biomass↗

A comparative analysis of simulated and observed photosynthetic CO2 uptake in two coniferous forest canopies.

Gross canopy photosynthesis (P(g)) can be simulated with canopy models or retrieved from turbulent carbon dioxide (CO2) flux measurements above the forest canopy. We compare the two estimates and illustrate our findings with two case studies. We used the three-dimensional canopy model MAESTRA to simulate P(g) of two spruce forests differing in age and structure. Model parameter acquisition and model sensitivity to selected model parameters are described, and modeled results are compared with independent flux estimates. Despite higher photon fluxes at the site, an older German Norway spruce (Picea abies L. (Karst.)) canopy took up 25% less CO2 from the atmosphere than a young Scottish Sitka spruce (Picea sitchensis (Bong.) Carr.) plantation. The average magnitudes of P(g) and the differences between the two canopies were satisfactorily represented by the model. The main reasons for the different uptake rates were a slightly smaller quantum yield and lower absorptance of the Norway spruce stand because of a more clumped canopy structure. The model did not represent the scatter in the turbulent CO2 flux densities, which was of the same order of magnitude as the non-photosynthetically-active-radiation-induced biophysical variability in the simulated P(g). Analysis of residuals identified only small systematic differences between the modeled flux estimates and turbulent flux measurements at high vapor pressure saturation deficits. The merits and limitations of comparative analysis for quality evaluation of both methods are discussed. From this analysis, we recommend use of both parameter sets and model structure as a basis for future applications and model development.

Carbon Dioxide↗

Reparameterizing the pattern mixture model for sensitivity analyses under informative dropout.

Pattern mixture models are frequently used to analyze longitudinal data where missingness is induced by dropout. For measured responses, it is typical to model the complete data as a mixture of multivariate normal distributions, where mixing is done over the dropout distribution. Fully parameterized pattern mixture models are not identified by incomplete data; Little (1993, Journal of the American Statistical Association 88, 125-134) has characterized several identifying restrictions that can be used for model fitting. We propose a reparameterization of the pattern mixture model that allows investigation of sensitivity to assumptions about nonidentified parameters in both the mean and variance, allows consideration of a wide range of nonignorable missing-data mechanisms, and has intuitive appeal for eliciting plausible missing-data mechanisms. The parameterization makes clear an advantage of pattern mixture models over parametric selection models, namely that the missing-data mechanism can be varied without affecting the marginal distribution of the observed data. To illustrate the utility of the new parameterization, we analyze data from a recent clinical trial of growth hormone for maintaining muscle strength in the elderly. Dropout occurs at a high rate and is potentially informative. We undertake a detailed sensitivity analysis to understand the impact of missing-data assumptions on the inference about the effects of growth hormone on muscle strength.

Aged↗

An overview of a multimedia benchmarking analysis for three risk assessment models: RESRAD, MMSOILS, and MEPAS.

Multimedia modelers from the United States Environmental Protection Agency (EPA) and the United States Department of Energy (DOE) collaborated to conduct a detailed and quantitative benchmarking analysis of three multimedia models. The three models--RESRAD (DOE), MMSOILS (EPA), and MEPAS (DOE)--represent analytically-based tools that are used by the respective agencies for performing human exposure and health risk assessments. The study is performed by individuals who participate directly in the ongoing design, development, and application of the models. Model form and function are compared by applying the models to a series of hypothetical problems, first isolating individual modules (e.g., atmospheric, surface water, groundwater) and then simulating multimedia-based risk resulting from contaminant release from a single source to multiple environmental media. Study results show that the models differ with respect to environmental processes included (i.e., model features) and the mathematical formulation and assumptions related to the implementation of solutions. Depending on the application, numerical estimates resulting from the models may vary over several orders-of-magnitude. On the other hand, two or more differences may offset each other such that model predictions are virtually equal. The conclusion from these results is that multimedia models are complex due to the integration of the many components of a risk assessment and this complexity must be fully appreciated during each step of the modeling process (i.e., model selection, problem conceptualization, model application, and interpretation of results).

Air Pollutants↗

Hormesis: the dose-response revolution.

Hormesis, a dose-response relationship phenomenon characterized by low-dose stimulation and high-dose inhibition, has been frequently observed in properly designed studies and is broadly generalizable as being independent of chemical/physical agent, biological model, and endpoint measured. This under-recognized and -appreciated concept has the potential to profoundly change toxicology and its related disciplines with respect to study design, animal model selection, endpoint selection, risk assessment methods, and numerous other aspects, including chemotherapeutics. This article indicates that as a result of hormesis, fundamental changes in the concept and conduct of toxicology and risk assessment should be made, including (a) the definition of toxicology, (b) the process of hazard (e.g., including study design, selection of biological model, dose number and distribution, endpoint measured, and temporal sequence) and risk assessment [e.g., concept of NOAEL (no observed adverse effect level), low dose modeling, recognition of beneficial as well as harmful responses] for all agents, and (c) the harmonization of cancer and noncancer risk assessment.

Animals↗

Fitting physiological models to data.

Methods of fitting models to experimental data obtained from biological systems are reviewed. The Michaelis-Menten model is used as the working example of a familiar and well-studied biological model. Criteria for selecting models include goodness of fit, freedom from systematic errors, and simplicity. A given model is usually fitted and the optimal parameters determined by minimization of an objective function, usually the sum of the squared errors. Freedom from systematic errors is best judged graphically. The important mathematical methods of fitting models are derived from the calculus and include linear and quadratic programming. The latter leads to minimization of the sum of squares of errors. Optimal search procedures, which also perform this minimization, are surveyed. Other properties of models, such as the proper number of parameters and whether linearization is appropriate, are discussed. The specialized problems of biological models and data are considered.

Models, Biological↗

Kernel fisher discriminants for outlier detection.

The problem of detecting atypical objects or outliers is one of the classical topics in (robust) statistics. Recently, it has been proposed to address this problem by means of one-class SVM classifiers. The method presented in this letter bridges the gap between kernelized one-class classification and gaussian density estimation in the induced feature space. Having established the exact relation between the two concepts, it is now possible to identify atypical objects by quantifying their deviations from the gaussian model. This model-based formalization of outliers overcomes the main conceptual shortcoming of most one-class approaches, which, in a strict sense, are unable to detect outliers, since the expected fraction of outliers has to be specified in advance. In order to overcome the inherent model selection problem of unsupervised kernel methods, a cross-validated likelihood criterion for selecting all free model parameters is applied. Experiments for detecting atypical objects in image databases effectively demonstrate the applicability of the proposed method in real-world scenarios.

Journal Article↗

Interference of Competing Beneficial Mutations on Recombining Chromosomes.

Finding signatures of selective sweeps in genomes is a major goal of current population genomics, as it allows estimating the rate of beneficial mutations going to fixation and identifying the genes involved in selection. Models of recurrent selective sweeps traditionally assume that in chromosomal regions of normal recombination rates at most one beneficial allele is on the way to fixation. We review and extend here the theoretical studies on interference between closely linked beneficial mutations suggesting that this assumption may be violated. We show that interference between beneficial mutations may lead to substantially increased fixation times even in chromosomal regions of normal recombination rates. Furthermore, we discuss how interference can be detected in population genomic studies by analyzing genetic footprints of selective sweeps, and search for empirical evidence of interference in published datasets.

fixation times↗

Ethical dilemmas in nursing: the role of the nurse and perceptions of autonomy.

This study investigated decision making in ethical dilemmas and attitudes toward professional autonomy. It was based on Murphy's identification of three nurse-patient relationship models. The model identification was the result of Murphy's investigation of the levels of moral reasoning of nurse practitioners, from Kohlberg's theory of moral development. Autonomy is necessary for patient advocacy in Murphy's highest order model of nurse-patient relationship. 109 freshmen, 103 seniors, and 82 graduates (baccalaureate nursing) were examined for model selection, risk-taking, restrictions, and anxiety in the decision-making process in specific situations. Autonomy was measured independently. The most significant results indicated that freshmen were less likely to select the autonomous model of relationship, had lower attitudes toward professional nursing autonomy, and were less willing to take risks. Graduates were lower than either student group in their perceptions of restrictions and anxiety. The responses to each dilemma itself varied by situation in relation to the model preferred.

Adolescent↗

Normal ranges of neuropsychological tests for the diagnosis of Alzheimer's disease.

The diagnosis of early stage dementia is a highly complex process involving not only a somatic examination but also a neuropsychological assessment of the patient's cognitive capability. The American 'Consortium to Establish a Registry for Alzheimer's Disease' (CERAD) has proposed a set of tests in English which has been translated into German. This paper presents the statistical methodology applied to determine normal ranges adjusted for demographic variables for the German CERAD neuropsychological assessment battery (CERAD-NAB). The study population consists of participants of the Basel Study on the Elderly (Project BASEL) which aims at identifying preclinical markers of Alzheimer's disease. The normative sample has been defined by carefully excluding potentially relevant medical history and concomitant diseases and consists of 617 participants which are between 53 and 92 years old. Test results should be adjusted for gender, age, and years of education. For this purpose, a set of linear models including these predictors and subsets of their interactions and squares was evaluated for all 11 test scores derived from the CERAD-NAB battery. Model selection was based on the PRESS (predicted residual sum of squares) statistic. Although a strict application of this criterion selected 6 different models, a slight compromise allowed to fit all test scores by two models. In several tests of the CERAD-NAB many participants achieve maximal scores. Residuals of such test scores are heavily skewed. An arcsine transformation has been tuned to the data, so that residuals are close to a normal distribution, at least for residuals in the lower quartile which is relevant in diagnosing cognitive impairment. Test results are finally presented as z-scores which can be easily compared to a standard normal distribution. The evaluation of the CERAD-NAB is implemented on the Internet and in an Excel application.

Aged↗

[Methods of risk assessment from data of experimental carcinogenesis studies].

A comparison of the carcinogenic potencies of different substances and an assessment of the risk to exposed people can contribute significantly to the management of carcinogenic hazards including adequate regulatory decisions. A mathematical estimation of the risk caused by a certain height of exposure is usually done under the assumption of a linear dose-response relationship by means of an arithmetical factor which gives the ratio of risk to dose in the low-response range and which has to be found on the basis of epidemiological or experimental data. This work deals with calculations of risk/dose ratios from the data of three carcinogenicity tests by means of statistical methods. The influence of the selection of mathematical model, background assumption and the point of the function chosen for the calculation of the ratio was tested. The results show that an additivity assumption for the background leads to linearity of the curve in the low-response range such that the ratios calculated with different models (multistage, weibull, logit, probit) or for different points of the functions do not differ significantly within the data sets tested. Even under the assumption of independence of the background risk the influence of the model selected is still small if the ratios are calculated for the lowest experimental dose or for an exposure associated risk of 1%. Risk/dose ratios calculated by one of these methods utilize the information of several experimental groups. They seem to be well suited for a direct comparison of the potency of different carcinogens and as a basis for an assessment of the carcinogenic risk to humans, which can serve as a measure of the hazard of a single exposed person or may give an idea about the tumour incidences to be expected within an exposed population.

Administration, Oral↗

Hazard regression with interval-censored data.

In a recent paper, Kooperberg, Stone, and Truong (1995a) introduced hazard regression (HARE), in which linear splines and their tensor products are used to estimate the conditional log-hazard function based on possibly censored, positive response data and one or more covariates. Model selection is carried out in an adaptive fashion using maximum likelihood estimation of the unknown coefficients, Rao and Wald statistics to carry out stepwise addition and deletion of basis functions, and the Bayesian Information Criterion (BIC) to select the final model. In the present paper, the HARE methodology is extended to accommodate interval-censored data, time-dependent covariates, and cubic splines. The presence of interval-censored data means that the log-likelihood function may no longer be concave, presenting additional numerical challenges. The extended methodology is applied to a data set containing both interval-censoring and time-dependent covariates. The new software will be available in a future release of S-Plus.

Acquired Immunodeficiency Syndrome↗

Unexpected association between reproductive longevity and blood magnesium levels in a new model of selected mouse strains.

Two recently described mouse strains, with high (MGH) and low (MGL) blood magnesium (Mg) levels were obtained by selection over 19 generations. Both strains exhibit strong differences for characteristics generally known to be related to blood Mg levels, such as increased stress sensitivity and stress-induced aggressivity in MGL mice. In contrast, while experimental Mg deficiency due to low oral Mg intake has been shown to shorten life span and lower reproductive ability, reproductive longevity was longer in the MGL than in the MGH strain. Interestingly, the life spans of the two strains are very similar. Although this character could have been fixed in the strains by chance, with no relationship to the blood Mg level, the possibility of a causal link with the selection cannot be ruled out and is discussed. Regardless of the mechanisms at stake, the MGH and MGL strains appear to constitute a new model for the study of the relationships between reproductive longevity and blood Mg levels.

Aging↗

Stability of multivariable fractional polynomial models with selection of variables and transformations: a bootstrap investigation.

Sauerbrei and Royston have recently described an algorithm, based on fractional polynomials, for the simultaneous selection of variables and of suitable transformations for continuous predictors in a multivariable regression setting. They illustrated the approach by analyses of two breast cancer data sets. Here we extend their work by considering how to assess possible instability in such multivariable fractional polynomial models. We first apply the algorithm repeatedly in many bootstrap replicates. We then use log-linear models to investigate dependencies among the inclusion fractions for each predictor and among the simplified classes of fractional polynomial function chosen in the bootstrap samples. To further evaluate the results, we define measures of instability based on a decomposition of the variability of the bootstrap-selected functions in relation to a reference function from the original model. For each data set we are able to identify large, reasonably stable subsets of the bootstrap replications in which the functional forms of the predictors appear fairly stable. Despite the considerable flexibility of the family of fractional polynomials and the consequent risk of overfitting when several variables are considered, we conclude that the multivariable selection algorithm can find stable models.

Algorithms↗

An automated procedure for the extraction of metabolic network information from time series data.

Novel high-throughput measurement techniques in vivo are beginning to produce dense high-quality time series which can be used to investigate the structure and regulation of biochemical networks. We propose an automated information extraction procedure which takes advantage of the unique S-system structure and supports model building from time traces, curve fitting, model selection, and structure identification based on parameter estimation. The procedure comprises of three modules: model Generation, parameter estimation or model Fitting, and model Selection (GFS algorithm). The GFS algorithm has been implemented in MATLAB and returns a list of candidate S-systems which adequately explain the data and guides the search to the most plausible model for the time series under study. By combining two strategies (namely decoupling and limiting connectivity) with methods of data smoothing, the proposed algorithm is scalable up to realistic situations of moderate size. We illustrate the proposed methodology with a didactic example.

Algorithms↗

A model to select chemotherapy regimens for phase III trials for extensive-stage small-cell lung cancer.

BACKGROUND: Many more phase II studies have favorable outcomes than the subsequent phase III trials. We used historical data from phase II and phase III studies for patients with extensive-stage small-cell lung cancer (SCLC) to generate a statistical model to provide assistance in selecting chemotherapy regimens from phase II studies for subsequent use in phase III randomized studies. METHODS: Information from 21 phase III trials for patients with extensive-stage SCLC initiated during the period from 1972 through 1990 was reviewed to identify those that were preceded by phase II studies of the same regimen. We used data from all the trial pairs to develop a statistical model in which the number of patients, the median survival of patients, and the number of deaths observed in the phase II trial are used to estimate the statistical power of the subsequent phase III trial. All statistical tests were two-sided. RESULTS: Nine phase II studies were identified that preceded phase III trials of the same regimen. The regimens from two phase II studies with the greatest expected power in the phase III trial (0. 62 and 0.58) both demonstrated significantly prolonged survival when compared with standard treatment in subsequent phase III trials (P<. 001 and P =.002, respectively). The regimens from six of the other phase II studies, for which the median power expected in the phase III trial was 0.28 (range, 0.19-0.52), showed no difference when compared with standard treatment in a phase III trial. CONCLUSIONS: Phase II studies for particular regimens that have an expected power of greater than 0.55 provide a reasonable basis for proceeding with a phase III trial.

Antineoplastic Agents↗