Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Using the quantitative genetic threshold model for inferences between and within species.

Sewall Wright's threshold model has been used in modelling discrete traits that may have a continuous trait underlying them, but it has proven difficult to make efficient statistical inferences with it. The availability of Markov chain Monte Carlo (MCMC) methods makes possible likelihood and Bayesian inference using this model. This paper discusses prospects for the use of the threshold model in morphological systematics to model the evolution of discrete all-or-none traits. There the threshold model has the advantage over 0/1 Markov process models in that it not only accommodates polymorphism within species, but can also allow for correlated evolution of traits with far fewer parameters that need to be inferred. The MCMC importance sampling methods needed to evaluate likelihood ratios for the threshold model are introduced and described in some detail.

Bayes Theorem↗

On sample size and inference for two-stage adaptive designs.

Proschan and Hunsberger (1995, Biometrics 51, 1315-1324) proposed a two-stage adaptive design that maintains the Type I error rate. For practical applications, a two-stage adaptive design is also required to achieve a desired statistical power while limiting the maximum overall sample size. In our proposal, a two-stage adaptive design is comprised of a main stage and an extension stage, where the main stage has sufficient power to reject the null under the anticipated effect size and the extension stage allows increasing the sample size in case the true effect size is smaller than anticipated. For statistical inference, methods for obtaining the overall adjusted p-value, point estimate and confidence intervals are developed. An exact two-stage test procedure is also outlined for robust inference.

Biometry↗

[Statistical aspects in planning psychotropic drug trials (author's transl)].

It has become virtually unthinkable to objectively judge trials of comparisons without statistical inference. For studies involving psychopharmaceuticals the requirements have especially grown, since variations of efficacy are difficult to differentiate and objectivate. The basic principles of statistical planning must be applied to studies with psychopharmaceuticals. On the basis of these general principles the relevant points of the clinical and statistical models and their relationship to one another are discussed. Special aspects of the study objectives, the clinical model, the medical trial plan, the statistical model and design, the execution of the trials, statistical evaluation and its interpretation are presented. The setting up of hypotheses, balancing and maintenance of factors of disturbances, sample sizes, randomizing techniques and the confrontation of clinical relevance with statistical significance are some of the important points discussed.

Clinical Trials as Topic↗

Direct data manipulation for local decision analysis as applied to the problem of arsenic in drinking water from tube wells in Bangladesh.

A wide variety of tools are available, both parametric and nonparametric, for analyzing spatial data. However, it is not always clear how to translate statistical inferences into decision recommendations. This article explores the possibilities of estimating the effects of decision options using very direct manipulation of data, bypassing formal statistical analysis. We illustrate with the application that motivated this research, a study of arsenic in drinking water in nearly 5,000 wells in a small area in rural Bangladesh. We estimate the potential benefits of two possible remedial actions: (1) recommendations that people switch to nearby wells with lower arsenic levels; and (2) drilling new community wells. We use simple nonparametric clustering methods and estimate uncertainties using cross-validation.

Algorithms↗

BJ: an S-Plus program to fit linear regression models to censored data using the Buckley-James method.

Most researchers are familiar with ordinary multiple regression models, most commonly fitted using the method of least squares. The method of Buckley and James (J. Buckley, I. James, Linear regression with censored data, Biometrika 66 (1979) 429-436.) is an extension of least squares for fitting multiple regression models when the response variable is right-censored as in the analysis of survival time data. The Buckley-James method has been shown to have good statistical properties under usual regularity conditions (T.L. Lai, Z. Ying, Large sample theory of a modified Buckley-James estimator for regression analysis with censored data, Ann. Stat. 19 (1991) 1370-1402.). Nevertheless, even after 20 years of its existence, it is almost never used in practice. We believe that this is mainly due to lack of software and we describe here an S-Plus program that through its inclusion in a public domain function library fully exploits the power of the S-Plus programming environment. This environment provides multiple facilities for model specification, diagnostics, statistical inference, and graphical depiction of the model fit.

Data Interpretation, Statistical↗

A Bayesian change-point analysis of electromyographic data: detecting muscle activation patterns and associated applications.

Many facets of neuromuscular activation patterns and control can be assessed via electromyography and are important for understanding the control of locomotion. After spinal cord injury, muscle activation patterns can affect locomotor recovery. We present a novel application of reversible jump Markov chain Monte Carlo simulation to estimate activation patterns from electromyographic data. We assume the data to be a zero-mean, heteroscedastic process. The variance is explicitly modeled using a step function. The number and location of points of discontinuity, or change-points, in the step function, the inter-change-point variances, and the overall mean are jointly modeled along with the mean and variance from baseline data. The number of change-points is considered a nuisance parameter and is integrated out of the posterior distribution. Whereas current methods of detecting activation patterns are deterministic or provide only point estimates, ours provides distributional estimates of muscle activation. These estimates, in turn, are used to estimate physiologically relevant quantities such as muscle coactivity, total integrated energy, and average burst duration and to draw valid statistical inferences about these quantities.

Bayes Theorem↗

Statistical analysis of DNA sequences.

Developments in the statistical analysis of DNA sequence data since 1984 are reviewed. Mathematical methods employing dynamic programming or incorporating Markov chain theory have been developed to search sequences for regions of similarity and to align sequences. When the biological forces of mutation and genetic drift are included in models, distances between aligned sequences allow the construction of evolutionary trees. Theory based on models may lead to estimates of variation of parameter estimates and so give a means of assessing the statistical significance of observed patterns and relationships. The complexity of DNA sequences, however, suggests that most statistical inferences will rest on random permutations of sequences.

Base Sequence↗

A comparison of red blood cell thiopurine metabolites in children with acute lymphoblastic leukemia who received oral mercaptopurine twice daily or once daily: a Pediatric Oncology Group study (now The Children's Oncology Group).

INTRODUCTION: Mercaptopurine is an important antimetabolite for treatment of childhood acute lymphoblastic leukemia (ALL). It has been prescribed to be given daily without therapeutic monitoring of drug levels. After first-pass metabolism by hepatic xanthine oxidase (XO), mercaptopurine is converted into two major intracellular metabolites, thioguanine nucleotide (TGN) and methylated mercaptopurine metabolites (including methylated thioinosine nucleotides), which are cytotoxic in vitro. Its short plasma half-life and S-phase-dependent pharmacokinetics suggest that biologically active concentration and exposure duration may be critical to cell kill. METHODS: Pediatric Oncology Group (POG) 9605, a randomized, open label phase III study of standard-risk ALL, was designed to compare daily with twice-daily mercaptopurine during continuation therapy. Red blood cell (RBC) TGN and methylated mercaptopurine metabolite levels were measured as surrogates of leukemic cell levels in a randomly selected subset of patients. TGN and methylated mercaptopurine metabolites were analyzed quantitatively by high-performance liquid chromatography (HPLC) and reported in ng/8 x 10.8 RBC. Statistical inferences utilized multiple linear regression. RESULTS: One hundred eighteen patients received mercaptopurine 75 mg/m(2) daily and 108 received 37.5 mg/m(2)/dose twice daily. Descriptive statistics for the daily group showed the median TGN was 42 ng (mean and standard deviation [SD] = 48 +/- 35, quartiles 29-64). For the twice daily group, it was 40 ng (mean and SD = 40 +/- 27, quartiles 26-53). For methylated mercaptopurine metabolites, the daily group median was 2,020 ng (mean and SD = 2,278 +/- 1,559, quartiles 1,247-3,162); the twice daily group median was 1,275 ng (mean and SD = 1,580 +/- 1,240, quartiles 599-2,369). When adjusted for the covariables: actual dosage, days on study, age at diagnosis, white blood cell count, gender, Black race compared with not, and Hispanic compared with not, daily dosing resulted in significantly higher average methylated mercaptopurine metabolites by 668 (standard error [SE] = 179, P = 0.001) and a trend toward higher average TGNs by 6.2 (SE = 4.2, P = 0.14). CONCLUSIONS: Daily dosing of mercaptopurine resulted in higher mean red cell methylated mercaptopurine metabolites when compared to split (twice a day dosing). The data were inconclusive with respect to TGNs. The relationships of methylated mercaptopurine metabolites and TGNs to clinical outcomes will be elucidated as part of the maturing 9605 data.

Administration, Oral↗

Efficiencies of maximum likelihood methods of phylogenetic inferences when different substitution models are used.

Choice of a substitution model is a crucial step in the maximum likelihood (ML) method of phylogenetic inference, and investigators tend to prefer complex mathematical models to simple ones. However, when complex models with many parameters are used, the extent of noise in statistical inferences increases, and thus complex models may not produce the true topology with a higher probability than simple ones. This problem was studied using computer simulation. When the number of nucleotides used was relatively large (1000 bp), the HKY+Gamma model showed smaller d(T) topological distance between the inferred and the true trees) than the JC and Kimura models. In the cases of shorter sequences (300 bp) simpler model and search algorithm such as JC model and SA+NNI search were found to be as efficient as more complicated searches and models in terms of topological distances, although the topologies obtained under HKY+Gamma model had the highest likelihood values. The performance of relatively simple search algorithm SA+NNI was found to be essentially the same as that of more extensive SA+TBR search under all models studied. Similarly to the conclusions reached by Takahashi and Nei [Mol. Biol. Evol. 17 (2000) 1251], our results indicate that simple models can be as efficient as complex models, and that use of complex models does not necessarily give more reliable trees compared with simple models.

Algorithms↗

The study of long-term HIV dynamics using semi-parametric non-linear mixed-effects models.

Modelling HIV dynamics has played an important role in understanding the pathogenesis of HIV infection in the past several years. Non-linear parametric models, derived from the mechanisms of HIV infection and drug action, have been used to fit short-term clinical data from AIDS clinical trials. However, it is found that the parametric models may not be adequate to fit long-term HIV dynamic data. To preserve the meaningful interpretation of the short-term HIV dynamic models as well as to characterize the long-term dynamics, we introduce a class of semi-parametric non-linear mixed-effects (NLME) models. The models are non-linear in population characteristics (fixed effects) and individual variations (random effects), both of which are modelled semi-parametrically. A basis-based approach is proposed to fit the models, which transforms a general semi-parametric NLME model into a set of standard parametric NLME models, indexed by the bases used. The bases that we employ are natural cubic splines for easy implementation. The resulting standard NLME models are low-dimensional and easy to solve. Statistical inferences that include testing parametric against semi-parametric mixed-effects are investigated. Innovative bootstrap procedures are developed for simulating the empirical distributions of the test statistics. Small-scale simulation and bootstrap studies show that our bootstrap procedures work well. The proposed approach and procedures are applied to long-term HIV dynamic data from an AIDS clinical study.

Acquired Immunodeficiency Syndrome↗

A new parametric model for survival data with long-term survivors.

We develop a new parametric model using the three-parameter Burr XII distribution for the analysis of survival data with long-term survivors, which includes the previous Weibull mixture model as a special case. The new model is applied to the analysis of a set of leukaemia data for which previous attempts in the literature using traditional parametric models were unsatisfactory due to lack of fit. It is shown that the new model improves the fit to the leukaemia data significantly and is thus capable of providing more credible answers to a variety of statistical inference problems that are of interest to medical researchers and practitioners.

Humans↗

Statistical methods for the beta-binomial model in teratology.

The beta-binomial model is widely used for analyzing teratological data involving littermates. Recent developments in statistical analyses of teratological data are briefly reviewed with emphasis on the model. For statistical inference of the parameters in the beta-binomial distribution, separation of the likelihood introduces an likelihood inference. This leads to reducing biases of estimators and also to improving accuracy of empirical significance levels of tests. Separate inference of the parameters can be conducted in a unified way.

Animals↗

Population toxicokinetics of benzene.

In assessing the distribution and metabolism of toxic compounds in the body, measurements are not always feasible for ethical or technical reasons. Computer modeling offers a reasonable alternative, but the variability and complexity of biological systems pose unique challenges in model building and adjustment. Recent tools from population pharmacokinetics, Bayesian statistical inference, and physiological modeling can be brought together to solve these problems. As an example, we modeled the distribution and metabolism of benzene in humans. We derive statistical distributions for the parameters of a physiological model of benzene, on the basis of existing data. The model adequately fits both prior physiological information and experimental data. An estimate of the relationship between benzene exposure (up to 10 ppm) and fraction metabolized in the bone marrow is obtained and is shown to be linear for the subjects studied. Our median population estimate for the fraction of benzene metabolized, independent of exposure levels, is 52% (90% confidence interval, 47-67%). At levels approaching occupational inhalation exposure (continuous 1 ppm exposure), the estimated quantity metabolized in the bone marrow ranges from 2 to 40 mg/day.

Bayes Theorem↗

Likelihood ratios: a simple and flexible statistic for empirical psychologists.

Empirical studies in psychology typically employ null hypothesis significance testing to draw statistical inferences. We propose that likelihood ratios are a more straightforward alternative to this approach. Likelihood ratios provide a measure of the fit of two competing models; the statistic represents a direct comparison of the relative likelihood of the data, given the best fit of the two models. Likelihood ratios offer an intuitive, easily interpretable statistic that allows the researcher great flexibility in framing empirical arguments. In support of this position, we report the results of a survey of empirical articles in psychology, in which the common uses of statistics by empirical psychologists is examined. From the results of this survey, we show that likelihood ratios are able to serve all the important statistical needs of researchers in empirical psychology in a format that is more straightforward and easier to interpret than traditional inferential statistics.

Empirical Research↗

Issues in biomedical statistics: analysing 2 x 2 tables of frequencies.

How best to analyse statistically experimental results that are set out as a 2 x 2 table of frequencies has been debated by statisticians for more than 50 years. The main issue is what framework of statistical inference should be adopted. The design of most biomedical experiments that result in 2 x 2 tables of independent observations is compatible with the randomization model of inference and with the Fisher exact test. It is rare that the Neyman-Pearson population model is applicable and that a case can be made for using the Pearson chi 2 test, or others that refer a test statistic to the chi-squared distribution. Even then, the adjustments for the mismatch between the test statistic and the chi-squared distribution so as to control the risk of Type I error are complex that the Fisher test is probably a safer option (or Yates' correction to the Pearson test if there is no access to a computer). When the 2 x 2 table results from two sets of measurements having been made on the same group, the population model of inference is inapplicable and the exact form of the McNemar test should be used. Confidence intervals for differences in proportions, the likelihood ratio, or the odds ratio, refer to randomly sampled populations and are not compatible with the randomization model of inference.

Animals↗

Population toxicokinetics of tetrachloroethylene.

In assessing the distribution and metabolism of toxic compounds in the body, measurements are not always feasible for ethical or technical reasons. Computer modeling offers a reasonable alternative, but the variability and complexity of biological systems pose unique challenges in model building and adjustment. Recent tools from population pharmacokinetics, Bayesian statistical inference, and physiological modeling can be brought together to solve these problems. As an example, we modeled the distribution and metabolism of tetrachloroethylene (PERC) in humans. We derive statistical distributions for the parameters of a physiological model of PERC, on the basis of data from Monster et al. (1979). The model adequately fits both prior physiological information and experimental data. An estimate of the relationship between PERC exposure and fraction metabolized is obtained. Our median population estimate for the fraction of inhaled tetrachloroethylene that is metabolized, at exposure levels exceeding current occupational standards, is 1.5% [95% confidence interval (0.52%, 4.1%)]. At levels approaching ambient inhalation exposure (0.001 ppm), the median estimate of the fraction metabolized is much higher, at 36% [95% confidence interval (15%, 58%)]. This disproportionality should be taken into account when deriving safe exposure limits for tetrachloroethylene and deserves to be verified by further experiments.

Administration, Inhalation↗

Statistical uncertainty in the no-observed-adverse-effect level.

The no-observed-adverse-effect level (NOAEL) is a dose value that U.S. EPA reduces by uncertainty factors (UF) and modifying factors (MF) to obtain a reference dose (RfD) for input to regulatory decision making. Whether the true added risk at the NOAEL is below an acceptable level, however, is a source of statistical uncertainty itself. As several authors have previously noted, the probability that added risk at the NOAEL is not negligibly small increases as sample sizes decrease. This is because the definition of the NOAEL statistically controls for the chance of a false-positive error, but not for a false-negative error. The false-positive rate is the test level set by the user in testing for a statistically significant dose effect, typically 0.05. When it is held fixed, the increase in statistical uncertainty as sample size decreases produces an increase in the false-negative rate. Hence, the fewer data available for statistical inference, the higher the expected value of the NOAEL and the less toxic an agent is likely to appear. The solution lies in calculating the probability that a statistical procedure used will detect the maximum added risk acceptable for health regulation (the "power" at that added risk). If the observed response in a dose group is not significantly elevated relative to the control group, and the power for detecting a difference is low as well, then the statistical evidence is inconclusive. In such a case, additional data or other sources of information are needed for evaluating added risk. These concepts are illustrated for examples from the literature with dichotomous (quantal response) data and categorical (severity) data, using a new statistical procedure.

Research Design↗

Bayesian approach to searching for susceptibility genes: Gc2 and EsD1 alleles and multiple sclerosis.

Multiple sclerosis (MS) is one of the most common causes of neurological disability in early adulthood. The current literature is interested in identifying biological or DNA markers associated with genetic susceptibility to MS. The aim of this study is to investigate, by means of Bayesian statistical inference, whether the presence of Gc2 (Gc = group-specific component) and/or EsD1 (EsD = esterase D) alleles affects MS susceptibility. Gc and EsD are two classical genetic markers, being the first a serum protein polymorphism, the latter an isoenzyme polymorphism. The interest of the proposed statistical approach of searching for MS susceptibility genes relies on the analysis of two different functions, one function being inferred from our results on 56 unrelated patients from central Italy affected by MS, the other one from Italian and worldwide epidemiological data. The graphical analysis suggests that MS susceptibility is influenced by both Gc2 and EsD1 alleles; and EsD1 allele is more informative than Gc2. These results point out the advantages of the Bayesian approach in searching for susceptibility genes. Furthermore, the significant association between the considered alleles and the susceptibility to MS suggests possible hypotheses about the pathogenesis of the disease.

Bayes Theorem↗