Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Is multiple sclerosis an infectious disease? Inference about an input process based on the output.

The output process of an infinite-server queue with a Poisson process input is observed starting at time 0 with an empty queue. It is assumed that the service time distribution is known. This article discusses statistical inference about the input intensity. A controversial issue in the study of multiple sclerosis is addressed as a motivation for the model and methods developed.

Adolescent↗

On the scientific inference from clinical trials.

We have not been able to describe clearly how we generalize findings from a study to our own 'everyday patients'. This difficulty is not surprising, since generalization deals with how empirical observations are related to the growth of scientific knowledge, which is a major philosophical problem. An argument, sometimes used to discard evidence from a trial, is that the patient sample was too selected and therefore not 'representative' enough for the results to be meaningful for generalization. In this paper, we discuss issues of representativeness and generalizability. Other authors have shown that generalization cannot only depend on statistical inference. Then, how do randomized clinical trials contribute to the growth of knowledge? We discuss three aspects of the randomized clinical trial (Mant 1999), First, the trial is an empirical experiment set up to study the intervention on the question as specifically and as much in isolation from other -- biasing and confounding -- factors as possible (Rothman & Greenland 1998). Second, the trial is set up to challenge our prevailing hypotheses (or prejudices) and the trial is above all a help in error elimination (Popper 1992). Third, we need to learn to see new, unexpected and thought-provoking patterns in the data from a trial. Point one -- and partly point two -- refers to the paradigm of the controlled experiment in scientific method. How much a study contributes to our knowledge, with respect to points two and three, relates to its originality. In none of these respects is the representativeness of the patients, or the clinical situations, crucial for judging the study and its possible inferences. However, we also discuss that the biological domain of disease that was studied in a particular trial has to be taken into account. Thus, the inference drawn from a clinical study is not only a question of statistical generalization, but must include a jump from the world of experiences into the world of reason, assessment and theoretical judgement.

Biometry↗

Theory-based Bayesian models of inductive learning and reasoning.

Inductive inference allows humans to make powerful generalizations from sparse data when learning about word meanings, unobserved properties, causal relationships, and many other aspects of the world. Traditional accounts of induction emphasize either the power of statistical learning, or the importance of strong constraints from structured domain knowledge, intuitive theories or schemas. We argue that both components are necessary to explain the nature, use and acquisition of human knowledge, and we introduce a theory-based Bayesian framework for modeling inductive learning and reasoning as statistical inferences over structured knowledge representations.

Association Learning↗

Variance components analysis for pedigree-based censored survival data using generalized linear mixed models (GLMMs) and Gibbs sampling in BUGS.

Complex human diseases are an increasingly important focus of genetic research. Many of the determinants of these diseases are unknown and there is often a strong residual covariance between relatives even when all known genetic and environmental factors have been taken into account. This must be modeled correctly whether scientific interest is focused on fixed effects, as in an association analysis, or on the covariance structure itself. Analysis is straightforward for multivariate normally distributed traits, but difficulties arise with other types of trait. Generalized linear mixed models (GLMMs) offer a potentially unifying approach to analysis for many classes of phenotype including right censored survival times. This includes age-at-onset and age-at-death data and a variety of other censored traits. Markov chain Monte Carlo (MCMC) methods, including Gibbs sampling, provide a convenient framework within which such GLMMs may be fitted. In this paper, we use BUGS ("Bayesian inference using Gibbs sampling": a readily available, generic Gibbs sampler) to fit GLMMs for right-censored survival times in nuclear and extended families. We discuss parameter interpretation and statistical inference, and show how to circumvent a number of important theoretical and practical problems. Using simulated data, we show that model parameters are consistent. We further illustrate our methods using data from an ongoing cohort study. Finally, we propose that the random effects associated with a genetic component of variance (e.g., sigma(2)(A)) in a GLMM may be regarded as an adjusted "phenotype" and used as input to a conventional model-based or model-free linkage analysis. This provides a simple way to conduct a linkage analysis for a trait reflected in a right-censored survival time while comprehensively adjusting for observed confounders at the level of the individual and latent environmental effects shared across families.

Bayes Theorem↗

A method for determining zygosity of transgenic zebrafish by TaqMan real-time PCR.

When producing a genetically modified organism, intended genes are often integrated into a target genome by random insertions. Subsequently, it is often desirable to know the gene copy number of the transgenic organism and the zygosity of its offspring. Because of the random insertions, the estimation can be made only by quantitative measurement of the genes. Even though TaqMan real-time PCR has been used in gene expression analysis, it is routinely used to quantify differences larger than twofold or more than one PCR cycle. In this study, we employed TaqMan quantitative PCR to determine zygosity of transgenic fluorescent zebrafish in which a homozygote and a hemizygote differ by only twofold. We measured relative quantities of the transgene by taking the threshold cycle (Ct) of both the transgene and an internal control zebrafish genomic DNA. Using scatterplots and statistical inference, we demonstrated that homozygotes and hemizygotes could be differentiated unambiguously when multiple measurements were taken. We discuss the relationship between the repetitive measurements and TaqMan precision with a statistical model. The result illustrates that the method can be extended to some areas that require even higher precision such as determining the polyploidy of an organism.

Animals↗

Introducing SONS, a tool for operational taxonomic unit-based comparisons of microbial community memberships and structures.

The recent advent of tools enabling statistical inferences to be drawn from comparisons of microbial communities has enabled the focus of microbial ecology to move from characterizing biodiversity to describing the distribution of that biodiversity. Although statistical tools have been developed to compare community structures across a phylogenetic tree, we lack tools to compare the memberships and structures of two communities at a particular operational taxonomic unit (OTU) definition. Furthermore, current tests of community structure do not indicate the similarity of the communities but only report the probability of a statistical hypothesis. Here we present a computer program, SONS, which implements nonparametric estimators for the fraction and richness of OTUs shared between two communities.

Biodiversity↗

On some applications of Bayesian methods in cancer clinical trials.

The NCCTG randomized controlled clinical trial for the treatment of advanced colorectal carcinoma is a wonderful case study of the dynamic interplay between scientific learning and statistical inference. Ethical concerns for minimizing the number of patients assigned to an inferior treatment and interest in identifying subsets of patients for whom a treatment is most likely efficacious pose challenging problems for the practice of statistics. In the first part of this paper, I comment on the applications of Bayesian methods to these problems in the NCCTG trial as presented by Freedman and Spieglehalter and Dixon and Simon, respectively. In the second part of this paper, I discuss and illustrate a Bayesian approach to model sensitivity analysis with a particular focus on model specification and criticism. The Bayesian approach provides a formal methodology to assess the sensitivity of inferences to the inputs into an analysis so that it is possible to investigate the consequences of the specification of the model. I apply these methods to the specification and criticism of a class of survival models for the analysis of survival times in the NCCTG trial.

Antineoplastic Combined Chemotherapy Protocols↗

Statistical design and the analysis of gene expression microarray data.

Gene expression microarrays are an innovative technology with enormous promise to help geneticists explore and understand the genome. Although the potential of this technology has been clearly demonstrated, many important and interesting statistical questions persist. We relate certain features of microarrays to other kinds of experimental data and argue that classical statistical techniques are appropriate and useful. We advocate greater attention to experimental design issues and a more prominent role for the ideas of statistical inference in microarray studies.

Analysis of Variance↗

Deconvolution of event-related fMRI responses in fast-rate experimental designs: tracking amplitude variations.

Recent developments towards event-related functional magnetic resonance imaging has greatly extended the range of experimental designs. If the events occur in rapid succession, the corresponding time-locked responses overlap significantly and need to be deconvolved in order to separate the contributions of different events. Here we present a deconvolution approach, which is especially aimed at the analysis of fMRI data where sequence- or context-related responses are expected. For this purpose, we make the assumption of a hemodynamic response function (HDR) with constant yet not predefined shape but with possibly variable amplitudes. This approach reduces the number of variables to be estimated but still keeps the solutions flexible with respect to the shape. Consequently, statistical efficiency is improved. Temporal variations of the HDR strength are directly indicated by the amplitudes derived by the algorithm. Both the estimation efficiency and statistical inference are further supported by an improved estimation of the noise covariance. Using synthesized data sets, both differently shaped HDRs and varying amplitude factors were correctly identified. The gain in statistical sensitivity led to improved ratios of false- and true-positive detection rates for synthetic activations in these data. In an event-related fMRI experiment with a human subject, different HDR amplitudes could be derived corresponding to stimulation at different visual stimulus contrasts. Finally, in a visual spatial attention experiment we obtained different fMRI response amplitudes depending on the sequences of attention conditions.

Algorithms↗

Bayesian statistics in medical research: an intuitive alternative to conventional data analysis.

Statistical analysis of both experimental and observational data is central to medical research. Unfortunately, the process of conventional statistical analysis is poorly understood by many medical scientists. This is due, in part, to the counter-intuitive nature of the basic tools of traditional (frequency-based) statistical inference. For example, the proper definition of a conventional 95% confidence interval is quite confusing. It is based upon the imaginary results of a series of hypothetical repetitions of the data generation process and subsequent analysis. Not surprisingly, this formal definition is often ignored and a 95% confidence interval is widely taken to represent a range of values that is associated with a 95% probability of containing the true value of the parameter being estimated. Working within the traditional framework of frequency-based statistics, this interpretation is fundamentally incorrect. It is perfectly valid, however, if one works within the framework of Bayesian statistics and assumes a 'prior distribution' that is uniform on the scale of the main outcome variable. This reflects a limited equivalence between conventional and Bayesian statistics that can be used to facilitate a simple Bayesian interpretation based on the results of a standard analysis. Such inferences provide direct and understandable answers to many important types of question in medical research. For example, they can be used to assist decision making based upon studies with unavoidably low statistical power, where non-significant results are all too often, and wrongly, interpreted as implying 'no effect'. They can also be used to overcome the confusion that can result when statistically significant effects are too small to be clinically relevant. This paper describes the theoretical basis of the Bayesian-based approach and illustrates its application with a practical example that investigates the prevalence of major cardiac defects in a cohort of children born using the assisted reproduction technique known as ICSI (intracytoplasmic sperm injection).

Bayes Theorem↗

On the use of the Price equation.

This paper distinguishes two categories of questions that the Price equation can help us answer. The two different types of questions require two different disciplines that are related, but nonetheless move in opposite directions. These disciplines are probability theory on the one hand and statistical inference on the other. In the literature on the Price equation this distinction is not made. As a result of this, questions that require a probability model are regularly approached with statistical tools. In this paper, we examine the possibilities of the Price equation for answering questions of either type. By spending extra attention on mathematical formalities, we avoid the two disciplines to get mixed up. After that, we look at some examples, both from kin selection and from group selection, that show how the inappropriate use of statistical terminology can put us on the wrong track. Statements that are 'derived' with the help of the Price equation are, therefore, in many cases not the answers they seem to be. Going through the derivations in reverse can, however, be helpful as a guide how to build proper (probabilistic) models that do give answers.

Animals↗

[Clinical cases, trials, and the problem of the colored socks. Or the problem of induction in medicine].

The medical prevision is based on the induction, that is the logical procedure trying to get a general truth starting from pooling and analysis of single cases. A first type of induction is defined complete or mathematical, since it allows a switch from a finished to a wider and endless set of cases. The mathematical induction is a logical procedure, completely consequential and that's why the derived demonstration is always true. Instead the knowledge in medicine is founded upon empirical induction, inferred from the observation of a sufficiently wide number of real cases. But the conclusions which are induced empirically introduce an inevitable degree of uncertainty, because they cannot be deduced logically. Statistics is the tool employed for "weighing" somehow the degree of confidence in a conclusion induced from the observation of numerous single cases. But it presents the limit of quantifying global behaviors, so the generalization derived from statistical inference reflects only the final mean effects that occurred under some specific conditions. Here are presented the problems that arise when switching from conclusions induced by the observations of a wide sample to the management of the single patient. On the other side the limits, but also the complementary value of case reports, are presented.

Clinical Trials as Topic↗

Definition, interpretation and calculation of cost-effectiveness acceptability curves.

This paper discusses the definition, interpretation and computation of cost-effectiveness (CE) acceptability curves. A formal definition of the CE acceptability curve based on the net benefit approach is provided. The curve can be computed using parametric or non-parametric techniques and for both computational approaches we establish a formal relation between the CE acceptability curve and statistical inference based on confidence intervals and P values in CE analysis.

Cost-Benefit Analysis↗

Adjusting craniofacial correlations for technical error.

Reliability estimates provide a means of adjusting observed correlations for technical error. The results show that the true correlations among 11 craniofacial landmarks are consistently higher than observed values. Consequently, the regression slopes defining these relationships are also increased. The less reliable two measures are, and the closer their joint reliability approximates the observed correlation, the greater the expected change in true correlation. Adjusting craniofacial relationships for technical error may substantially increase the proportion of variation explained, and thereby alter statistical inferences drawn from results.

Cephalometry↗

Using SAS to conduct nonparametric residual bootstrap multilevel modeling with a small number of groups.

In multilevel modeling, researchers often encounter data with a relatively small number of units at the higher levels. As a result, of this and/or non-normality of the residuals, model parameter estimates, particularly the variance components and standard errors of parameter estimates at the group level, may be biased, thus the corresponding statistical inferences may not be trustworthy. This problem can be addressed by using bootstrap methods to estimate the standard errors of the parameter estimates for significance testing. This study illustrates how to use statistical analysis system (SAS) to conduct nonparametric residual bootstrap multilevel modeling. Specific SAS programs for such modeling are provided.

Models, Statistical↗

Statistics: can we prove an association for a rare complication?

BACKGROUND AND OBJECTIVES: Microcatheters for continuous spinal anesthesia were withdrawn from the market after an apparent increase in the incidence of cauda equina syndrome (CES) associated with spinal anesthesia after introduction of these catheters. The objective of this review is to evaluate the historical data on CES after spinal anesthesia and to compare the result to the recent data. METHODS: The literature on the use of statistics and on complications associated with spinal anesthesia was reviewed. Poisson probabilities were calculated to assess the probability of seeing the recent cases of CES, given the historical data. The sample size required for a prospective study was calculated. RESULTS: Statistics cannot "prove" an hypothesis but can only support or fail to support it. Use and interpretation of the p value are discussed, as are possible problems with interpretation of the p value. The difference between causal and statistical inference is discussed. Probabilities for the occurrence of CES after microcatheter use were calculated. The sample size for a prospective study of this problem was calculated; a large sample is required. CONCLUSIONS: Statistics alone cannot support an association of microcatheters with CES after spinal anesthesia. Additional considerations suggest a possible association, but further study is required.

Anesthesia, Spinal↗

Using complexity for the estimation of Bayesian networks.

Statistical inference of graphical models has become an important tool in the reconstruction of biological networks of the type which model, for example, gene regulatory interactions. In particular, the construction of a score-based Bayesian posterior density over the space of models provides an intuitive and computationally feasible method of assessing model uncertainty and of assigning statistical confidence to structural features. One problem which frequently occurs with this approach is the tendency to overestimate the degree of model complexity. Spurious graphical features obtained in this way may affect the inference in unpredictable ways, even when using scoring techniques, such as the Bayesian Information Criterion (BIC), that are specifically designed to compensate for overfitting. In this article we propose a simple adjustment to a BIC-based scoring procedure. The method proceeds in two steps. In the first step we derive an independent estimate of the parametric complexity of the model. In the second we modify the BIC score so that the mean parametric complexity of the posterior density is equal to the estimated value. The method is applied to a set of test networks, and to a collection of genes from the yeast genome known to possess regulatory relationships. A Bayesian network model with binary responses is employed. In the examples considered, we find that the number of spurious graph edges inferred is reduced, while the effect on the identification of true edges is minimal.

Algorithms↗

A simple imputation method for longitudinal studies with non-ignorable non-responses.

Missing data are a common problem in longitudinal studies in the health sciences. Motivated by data from the Muscatine Coronary Risk Factor (MCRF) study, a longitudinal study of obesity, we propose a simple imputation method for handling non-ignorable non-responses (i.e., when non-response is related to the specific values that should have been obtained) in longitudinal studies with either discrete or continuous outcomes. In the proposed approach, two regression models are specified; one for the marginal mean of the response, the other for the conditional mean of the response given non-response patterns. Statistical inference for the model parameters is based on the generalized estimating equations (GEE) approach. An appealing feature of the proposed method is that it can be readily implemented using existing, widely-available statistical software. The method is illustrated using longitudinal data on obesity from the MCRF study.

Algorithms↗