Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

From mere coincidences to meaningful discoveries.

People's reactions to coincidences are often cited as an illustration of the irrationality of human reasoning about chance. We argue that coincidences may be better understood in terms of rational statistical inference, based on their functional role in processes of causal discovery and theory revision. We present a formal definition of coincidences in the context of a Bayesian framework for causal induction: a coincidence is an event that provides support for an alternative to a currently favored causal theory, but not necessarily enough support to accept that alternative in light of its low prior probability. We test the qualitative and quantitative predictions of this account through a series of experiments that examine the transition from coincidence to evidence, the correspondence between the strength of coincidences and the statistical support for causal structure, and the relationship between causes and coincidences. Our results indicate that people can accurately assess the strength of coincidences, suggesting that irrational conclusions drawn from coincidences are the consequence of overestimation of the plausibility of novel causal forces. We discuss the implications of our account for understanding the role of coincidences in theory change.

Bayes Theorem↗

Problems in using age-stratum-specific reference rates for indirect standardization.

Disease risk in a study cohort can be compared with that in a reference population using the method of indirect standardization, adjusting for a difference in age distribution between the cohort and reference population. The common epidemiological practice is to use categorical age, typically in the form of 5-year age strata, in the standardization. This article discusses problems that arise owing to the categorization of age, including biased estimation and incorrect statistical inference on relative risk parameters. The same problems further extend to more general analyses using Poisson regression. We illustrate the problems using a hypothetical example and propose a simple remedy using linear splines. A slightly more elaborate method and its computer program are given in the appendix.

Age Distribution↗

The coverage of a random sample from a biological community.

A taxonomic group will frequently have a large number of species with small abundances. When a sample is drawn at random from this group, one is therefore faced with the problem that a large proportion of the species will not be discovered. A general definition of quantitative measures of "sample coverage" is proposed, and the problem of statistical inference is considered for two special cases, (1) the actual total relative abundance of those species that are represented in the sample, and (2) their relative contribution to the information index of diversity. The analysis is based on a extended version of the negative binomial species frequency model. The results are tabulated.

Animals↗

Nonparametric estimation of covariance structure in longitudinal data.

In longitudinal studies, the effect of various treatments over time is usually of prime interest. However, observations on the same subject are usually correlated and any analysis should account for the underlying covariance structure. A nonparametric estimate of the covariance structure is useful, either as a guide to the formulation of a parametric model or as the basis for formal inference without imposing parametric assumptions. The sample covariance matrix provides such an estimate when the data consist of a short sequence of measurements at a common set of time points on each of many subjects but is impractical when the data are severely unbalanced or when the sequences of measurements on individual subjects are long relative to the number of subjects. The variogram of residuals from a saturated model for the mean response has previously been suggested as a nonparametric estimator for covariance structure assuming stationarity. In this paper, we consider kernel weighted local linear regression smoothing of sample variogram ordinates and of squared residuals to provide a nonparametric estimator for the covariance structure without assuming stationarity. The value of the estimator as a diagnostic tool is demonstrated in two applications, one to a set of data concerning the blood pressure of newborn babies in an intensive care unit and the other to data on the time evolution of CD4 cell numbers in HIV seroconverters. The use of the estimator in more formal statistical inferences concerning the mean profiles requires further study.

Anti-HIV Agents↗

Maximum likelihood estimation of haplotype effects and haplotype-environment interactions in association studies.

The associations between haplotypes and disease phenotypes offer valuable clues about the genetic determinants of complex diseases. It is highly challenging to make statistical inferences about these associations because of the unknown gametic phase in genotype data. We describe a general likelihood-based approach to inferring haplotype-disease associations in studies of unrelated individuals. We consider all possible phenotypes (including disease indicator, quantitative trait, and potentially censored age at onset of disease) and all commonly used study designs (including cross-sectional, case-control, cohort, nested case-control, and case-cohort). The effects of haplotypes on phenotype are characterized by appropriate regression models, which allow various genetic mechanisms and gene-environment interactions. We present the likelihood functions for all study designs and disease phenotypes under Hardy-Weinberg disequilibrium. The corresponding maximum likelihood estimators are approximately unbiased, normally distributed, and statistically efficient. We provide simple and efficient numerical algorithms to calculate the maximum likelihood estimators and their variances, and implement these algorithms in a freely available computer program. Extensive simulation studies demonstrate that the proposed methods perform well in realistic situations. An application to the Carolina Breast Cancer Study reveals significant haplotype effects and haplotype-smoking interactions in the development of breast cancer.

Algorithms↗

An empirical analysis of eating disorders and anxiety disorders publications (1980-2000)--part II: Statistical hypothesis testing.

OBJECTIVE: The current study compared the eating disorder literature and the anxiety disorder literature in terms of statistical hypothesis testing features in 1980, 1990, and 2000. METHOD: Computer literature searches were conducted using PubMed and PsychInfo databases to identify relevant eating disorder and anxiety disorder articles published at each of the three time points. A total of 456 articles were randomly selected, including 228 articles each from the fields of eating disorders and anxiety disorders. Within each field, one third (76) of the articles were selected from each of the three time points. Two raters, from a team of eight trained raters, were randomly assigned to independently rate each article in terms of 75 separate methodologic features. In the current article, we will emphasize the findings about hypothesis testing and statistical analysis. Disagreements in ratings were resolved via consensus. Ratings were tabulated separately by field across the three time points. RESULTS: Few differences were observed between eating disorder and anxiety disorder publications in terms of statistical hypothesis testing features. Although increases were observed in both fields in a number of areas from 1980 to 2000, there remains a pervasive absence of many of the statistical hypothesis testing features recommended by the American Psychological Association Task Force on Statistical Inference. CONCLUSION: These results are discussed in terms of their implications for the fields of eating disorders and anxiety disorders, for researchers, for reviewers, and for professional journals and editorial boards.

Anxiety Disorders↗

A Bayesian approach to jointly estimate centre and treatment by centre heterogeneity in a proportional hazards model.

When multicentre clinical trial data are analysed, it has become more and more popular to look for possible heterogeneity in outcome between centres. However, beyond the investigation of such heterogeneity, it is also interesting to consider heterogeneity in treatment effect over centres. For time-to-event outcomes, this may be investigated by including a random centre effect and a random treatment by centre interaction in a Cox proportional hazards model. Assuming independence between the random effects, we propose a Bayesian approach to fit our proposed model. The parameters of interest are the variance components sigma(0) (2) and sigma(1) (2) of these random effects, which can be interpreted as a measure of centre and treatment effect over centres heterogeneity of the hazard. These variance components are estimated from their marginal posterior density after integrating out the fixed treatment effect and the random effects. As this integration cannot be performed analytically, the marginal posterior density is approximated using the Laplace integration technique. Statistical inference is then based on the characteristics of the posterior marginal density, such as the mode and the standard deviation. We demonstrate the proposed technique using data from a pooled database of seven EORTC bladder cancer clinical trials. Substantial centre and treatment effect over centres heterogeneity in disease-free interval was found.

Bayes Theorem↗

Clinical significance not statistical significance: a simple Bayesian alternative to p values.

OBJECTIVES: To take the common "Bayesian" interpretation of conventional confidence intervals to its logical conclusion, and hence to derive a simple, intuitive way to interpret the results of public health and clinical studies. DESIGN AND SETTING: The theoretical basis and practicalities of the approach advocated is at first explained and then its use is illustrated by referring to the interpretation of a real historical cohort study. The study considered compared survival on haemodialysis (HD) with that on continuous ambulatory peritoneal dialysis (CAPD) in 389 patients dialysed for end stage renal disease in Leicestershire between 1974 and 1985. Careful interpretation of the study was essential. This was because although it had relatively low statistical power, it represented all of the data that were available at the time and it had to inform a critical clinical policy decision: whether or not to continue putting the majority of new patients onto CAPD. MEASUREMENTS AND ANALYSIS: Conventional confidence intervals are often interpreted using subjective probability. For example, 95% confidence intervals are commonly understood to represent a range of values within which one may be 95% certain that the true value of whatever one is estimating really lies. Such an interpretation is fundamentally incorrect within the framework of conventional, frequency-based, statistics. However, it is valid as a statement of Bayesian posterior probability, provided that the prior distribution that represents pre-existing beliefs is uniform, which means flat, on the scale of the main outcome variable. This means that there is a limited equivalence between conventional and Bayesian statistics, which can be used to draw simple Bayesian style statistical inferences from a standard analysis. The advantage of such an approach is that it permits intuitive inferential statements to be made that cannot be made within a conventional framework and this can help to ensure that logical decisions are taken on the basis of study results. In the particular practical example described, this approach is applied in the context of an analysis based upon proportional hazards (Cox) regression. MAIN RESULTS AND CONCLUSIONS: The approach proposed expresses conclusions in a manner that is believed to be a helpful adjunct to more conventional inferential statements. It is of greatest value in those situations in which statistical significance may bear little relation to clinical significance and a conventional analysis using p values is liable to be misleading. Perhaps most importantly, this includes circumstances in which an important public health or clinical decision must be based upon a study that has unavoidably low statistical power. However, it is also useful in situations in which a decision must be based upon a large study that indicates that an effect that is highly statistically significant seems too small to be of practical relevance. In the illustrative example described, the approach helped in making a decision regarding the use of CAPD in Leicestershire during the latter half of the 1980s.

Bayes Theorem↗

Population HIV-1 dynamics in vivo: applicable models and inferential tools for virological data from AIDS clinical trials.

In this paper, we introduce a novel application of hierarchical nonlinear mixed-effect models to HIV dynamics. We show that a simple model with a sum of exponentials can give a good fit to the observed clinical data of HIV-1 dynamics (HIV-1 RNA copies) after initiation of potent antiviral treatments and can also be justified by a biological compartment model for the interaction between HIV and its host cells. This kind of model enjoys both biological interpretability and mathematical simplicity after reparameterization and simplification. A model simplification procedure is proposed and illustrated through examples. We interpret and justify various simplified models based on clinical data taken during different phases of viral dynamics during antiviral treatments. We suggest the hierarchical nonlinear mixed-effect model approach for parameter estimation and other statistical inferences. In the context of an AIDS clinical trial involving patients treated with a combination of potent antiviral agents, we show how the models may be used to draw biologically relevant interpretations from repeated HIV-1 RNA measurements and demonstrate the potential use of the models in clinical decision-making.

Acquired Immunodeficiency Syndrome↗

A design methodology for nonlinear systems containing parameter uncertainty.

A design methodology capable of dealing with nonlinear systems containing parameter uncertainty is presented. A generalized sensitivity analysis is incorporated which utilizes sampling of the parameter space and statistical inference. For a system with j adjustable and k nonadjustable parameters, this methodology (which includes an adaptive random search strategy) is used to determine the combination of j adjustable parameter values which maximizes the probability of the performance indices simultaneously satisfying design criteria given the uncertainty in the k nonadjustable parameters.

Algorithms↗

A Bayesian method for analysing spotted microarray data.

In the decade since their invention, spotted microarrays have been undergoing technical advances that have increased the utility, scope and precision of their ability to measure gene expression. At the same time, more researchers are taking advantage of the fundamentally quantitative nature of these tools with refined experimental designs and sophisticated statistical analyses. These new approaches utilise the power of microarrays to estimate differences in gene expression levels, rather than just categorising genes as up- or down-regulated, and allow the comparison of expression data across multiple samples. In this review, some of the technical aspects of spotted microarrays that can affect statistical inference are highlighted, and a discussion is provided of how several methods for estimating gene expression level across multiple samples deal with these challenges. The focus is on a Bayesian analysis method, BAGEL, which is easy to implement and produces easily interpreted results.

Algorithms↗

Multivariate group effect analysis in functional Magnetic Resonance Imaging.

In functional MRI (fMRI), analysis of multisubject data typically involves spatially normalizing (i.e. co-registering in a common standard space) all data sets and summarizing results in a single group activation map. This widely used approach does not explicitely account for between-subject anatomo-functional variability. Therefore, we propose a group effect analysis method which makes use of a multivariate model to select the main signal variations that are common to all subjects, while allowing final statistical inference on the individual scale. The normalization step is thus avoided and individual anatomo-functional features are preserved. The approach is evaluated by using simulated data and it is shown that sensitivity is drastically improved compared to more conventional individual analysis.

Algorithms↗

Reliability estimation based on system data with an unknown load share rule.

We consider a multicomponent load-sharing system in which the failure rate of a given component depends on the set of working components at any given time. Such systems can arise in software reliability models and in multivariate failure-time models in biostatistics, for example. A load-share rule dictates how stress or load is redistributed to the surviving components after a component fails within the system. In this paper, we assume the load share rule is unknown and derive methods for statistical inference on load-share parameters based on maximum likelihood. Components with (individual) constant failure rates are observed in two environments: (1) the system load is distributed evenly among the working components, and (2) we assume only the load for each working component increases when other components in the system fail. Tests for these special load-share models are investigated.

Biometry↗

A cladistic analysis of phenotype associations with haplotypes inferred from restriction endonuclease mapping. II. The analysis of natural populations.

Genes that code for products involved in the physiology of a phenotype are logical candidates for explaining interindividual variation in that phenotype. We present a methodology for discovering associations between genetic variation at such candidate loci (assayed through restriction endonuclease mapping) with phenotypic variation at the population level. We confine our analyses to DNA regions in which recombination is very rare. In this case, the genetic variation at the candidate locus can be organized into a cladogram that represents the evolutionary relationships between the observed haplotypes. Any mutation causing a significant phenotypic effect should be imbedded within the same historical structure defined by the cladogram. We showed, in the first paper of this series, how to use the cladogram to define a nested analysis of variance (NANOVA) that was very efficient at detecting and localizing phenotypically important mutations. However, the NANOVA of haplotype effects could only be applied to populations of homozygous genotypes. In this paper, we apply the quantitative genetic concept of average excess to evaluate the phenotypic effect of a haplotype or group of haplotypes stratified and contrasted according to the nested design defined by the cladogram. We also show how a permutational procedure can be used to make statistical inferences about the nested average excess values in populations containing heterozygous as well as homozygous genotypes. We provide two worked examples that investigate associations between genetic variation at or near the Alcohol dehydrogenase (Adh) locus and Adh activity in Drosophila melanogaster, and associations between genetic variation at or near some apolipoprotein loci and various lipid phenotypes in a human population.

Alcohol Dehydrogenase↗

Generalized linear latent variable models for repeated measures of spatially correlated multivariate data.

Observations of multiple-response variables across space and over time occur often in environmental and ecological studies. Compared to purely spatial models for a single response variable in the exponential family of distributions, fewer statistical tools are available for multiple-response variables that are not necessarily Gaussian. An exception is a common-factor model developed for multivariate spatial data by Wang and Wall (2003, Biostatistics 4, 569-582). The purpose of this article is to extend this multivariate space-only model and develop a flexible class of generalized linear latent variable models for multivariate spatial-temporal data. For statistical inference, maximum likelihood estimates and their standard deviations are obtained using a Monte Carlo EM algorithm. We also use a novel way to automatically adjust the Monte Carlo sample size, which facilitates the convergence of the Monte Carlo EM algorithm. The methodology is illustrated by an ecological study of red pine trees in response to bark beetle challenges in a forest stand of Wisconsin.

Algorithms↗

Analysis of clinical data with breached blindness.

In clinical trials, blinding is usually employed to prevent bias that may be introduced due to the knowledge of the identity of the treatment codes. This bias could alter the conclusion of statistical inference on the treatment effect. The purpose of this article is to propose a method for analysing clinical data with breached blindness. The example regarding the study of the effectiveness of an appetite suppressant in weight loss in obese woman as described in Brownell and Stunkard (Am. J. Psychiatry 1982; 139:1487-1489) is used to illustrate the application of the proposed methods.

Analysis of Variance↗

Jerome Cornfield's contributions to the conduct of clinical trials.

Jerome Cornfield's important contributions to the conduct of clinical trials are summarized here. They include consultative advice in the planning of many national trials, active collaboration in the conduct of many others, discussions of the role of classical and Bayesian methods of statistical inference in clinical trials, recommendations on data monitoring, contributions to the analysis of results of the University Group Diabetes Project, and efforts to assist the planning of coronary intervention trials with quantitative assessments of possible reductions in disease rates due to intervention on smoking and diet. An attempt is made to evaluate the impact of Cornfield's contributions to clinical trials.

Bayes Theorem↗

Testing anatomically specified hypotheses in functional imaging using cytoarchitectonic maps.

The statistical inference on functional imaging data is severely complicated by the embedded multiple testing problem. Defining a region of interest (ROI) where the activation is hypothesized a priori helps to circumvent this problem, since in this case the inference is restricted to fewer simultaneous tests, rendering it more sensitive. Cytoarchitectonic maps obtained from postmortem brains provide objective, a priori ROIs that can be used to test anatomically specified hypotheses about the localization of functional activations. We here analyzed three methods for the definition of ROIs based on probabilistic cytoarchitectonic maps. (1) ROIs defined by the volume assigned to a cytoarchitectonic area in the summary map of all areas (maximum probability map, MPM), (2) ROIs based on thresholding the individual probabilistic maps and (3) spherical ROIs build around the cytoarchitectonic center of gravity. The quality with which the thus defined ROIs represented the respective cytoarchitectonic areas as well as their sensitivity for detecting functional activations was subsequently statistically evaluated. Our data showed that the MPM method yields ROIs, which reflect most adequately the underlying anatomical hypotheses. These maps also show a high degree of sensitivity in the statistical analysis. We thus propose the use of MPMs for the definition of ROIs. In combination with thresholding based on the Gaussian random field theory, these ROIs can then be applied to test anatomically specified hypotheses in functional neuroimaging studies.

Artifacts↗