Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Risk, statistical inference, and the law of evidence: the use of epidemiological data in toxic tort cases.

Toxic torts are product liability cases dealing with alleged injuries due to chemical or biological hazards such as radiation, thalidomide, or Agent Orange. Toxic tort cases typically rely more heavily than other product liability cases on indirect or statistical proof of injury. There have been numerous theoretical analyses of statistical proof of injury in toxic tort cases. However, there have been only a handful of actual legal decisions regarding the use of such statistical evidence, and most of those decisions have been inconclusive. Recently, a major case from the Fifth Circuit, involving allegations that Benedectin (a morning sickness drug) caused birth defects, was decided entirely on the basis of statistical inference. This paper examines both the conceptual basis of that decision, and also the relationships among statistical inference, scientific evidence, and the rules of product liability in general.

Abnormalities, Drug-Induced↗

Statistical inference of chromosomal homology based on gene colinearity and applications to Arabidopsis and rice.

BACKGROUND: The identification of chromosomal homology will shed light on such mysteries of genome evolution as DNA duplication, rearrangement and loss. Several approaches have been developed to detect chromosomal homology based on gene synteny or colinearity. However, the previously reported implementations lack statistical inferences which are essential to reveal actual homologies. RESULTS: In this study, we present a statistical approach to detect homologous chromosomal segments based on gene colinearity. We implement this approach in a software package ColinearScan to detect putative colinear regions using a dynamic programming algorithm. Statistical models are proposed to estimate proper parameter values and evaluate the significance of putative homologous regions. Statistical inference, high computational efficiency and flexibility of input data type are three key features of our approach. CONCLUSION: We apply ColinearScan to the Arabidopsis and rice genomes to detect duplicated regions within each species and homologous fragments between these two species. We find many more homologous chromosomal segments in the rice genome than previously reported. We also find many small colinear segments between rice and Arabidopsis genomes.

Algorithms↗

Do statistical inferences allowing three alternative decisions give better feedback for environmentally precautionary decision-making?

Environmental policies and guidelines often specify standards for environmental indicators. The first part of this paper argues that, where compliance with these standards is assessed with the help of statistical inference, an inference employing a three-alternatives decision rule can provide more sensible feedback to environmental managers for precautionary decision-making. The second part of the paper shows how a three-alternatives statistical inference about compliance with a percentile standard might be applied to a small number of observations using a non-parametric binomial interval. This interval expression of uncertainty results in the sample size requirements for various percentile ranks becoming explicit.

Decision Making↗

Practical implications of modes of statistical inference for causal effects and the critical role of the assignment mechanism.

Causal inference in an important topic and one that is now attracting serious attention of statisticians. Although there exist recent discussions concerning the general definition of causal effects and a substantial literature on specific techniques for the analysis of data in randomized and nonrandomized studies, there has been relatively little discussion of modes of statistical inference for causal effects. This presentation briefly describes and contrasts four basic modes of statistical inference for causal effects, emphasizes the common underlying causal framework with a posited assignment mechanism, and describes practical implications in the context of an example involving the effects of switching from a name-band to a generic drug. A fundamental conclusion is that in such nonrandomized studies, sensitivity of inference to the assignment mechanism is the dominant issue, and it cannot be avoided by changing modes of inference, for instance, by changing from randomization-based to Bayesian methods.

Bayes Theorem↗

Statistical inference in generalized linear mixed models: a review.

We present a review of statistical inference in generalized linear mixed models (GLMMs). GLMMs are an extension of generalized linear models and are suitable for the analysis of non-normal data with a clustered structure. A GLMM contains parameters common to all clusters (fixed regression effects and variance components) and cluster-specific parameters. The latter parameters are assumed to be randomly drawn from a population distribution. The parameters of this population distribution (the variance components) have to be estimated together with the fixed effects. We focus on the case in which the cluster-specific parameters are normally distributed. The cluster-specific effects are integrated out of the likelihood so that the fixed effects and variance components can be estimated. Unfortunately, the integral over the cluster-specific effects is intractable for most GLMMs with a normal mixing distribution. Within a classical statistical framework, we distinguish between two broad classes of methods to handle this intractable integral: methods that rely on a numerical approximation to the integral and methods that use an analytical approximation to the integrand. Finally, we present an overview of available methods for testing hypotheses about the parameters of GLMMs.

Analysis of Variance↗

Statistical inference in evolutionary models of DNA sequences via the EM algorithm.

We describe statistical inference in continuous time Markov processes of DNA sequences related by a phylogenetic tree. The maximum likelihood estimator can be found by the expectation maximization (EM) algorithm and an expression for the information matrix is also derived. We provide explicit analytical solutions for the EM algorithm and information matrix.

Journal Article↗

Effects of randomization methods on statistical inference in disease cluster detection.

Monte Carlo methods are commonly used to assess the statistical significance of disease clusters. This usually involves permuting the observed outcome measure, such as the rate of disease, across the geographic units within the study area. When the variance of the disease rates is heterogeneous, however, randomizing the disease rate across the geographic units results in over-estimating the p-values in areas of low variance and under-estimating the p-values in areas of high variance. This bias results in under-ascertainment of clusters in urban areas and over-ascertainment of clusters in rural areas. As an alternative, randomizing the number of cases of disease or deaths proportional to the population at risk preserves the variance structure of the study area, therefore resulting in unbiased statistical inference. We compare results from randomizing rates with those from randomizing case counts, using county-level prostate cancer mortality data for the United States and ZIP-Code level prostate cancer incidence data for New York State, using the local Moran's I statistic.

Bias↗

Effects of misclassifications on statistical inferences in epidemiology.

Misclassification errors caused by imperfect sensitivity (U) and specificity (V) can affect statistical inferences in epidemiology. Such errors can lead to biases and increased standard errors in estimates of rates. Furthermore, low U and V can have a catastrophic effect on the power of a test to detect a change in rate, and, if U and V change even slightly as the rate changes, the effect on power may be dramatic.

Biometry↗

At what price significance? The effect of price estimates on statistical inference in economic evaluation.

Because data on resource utilization are now collected in many comparative trials of health interventions, statistical analysis of between-group differences in mean costs has become common. Statistical analyses of costs are generally performed conditional on a set of resource prices (or unit costs), thereby suppressing any uncertainty associated with those price estimates. Results presented here demonstrate that varying price estimates can have a non-negligible effect on statistical inference regarding between-group cost differences. Depending on the relative prices used in an analysis, between-group differences in total costs per patient may be either statistically significant or insignificant, regardless of whether differences in utilization of the underlying resources are statistically significant. These results highlight the importance of recognizing that evaluations based on patient-level economic data may be sensitive to assumptions regarding the values of unobserved variables, such as the relative prices of resources. Traditional methods of sensitivity analysis remain a valuable tool for analysing the implications of uncertainty around estimates of those unobserved variables.

Clinical Trials as Topic↗

Brain asymmetry: facts and fallacies in statistical inference.

In a recent article, C. M. Clark, R. Kessler, and R. Margolin (1985, Brain and Cognition, 4, 7-12) argue that standard statistical tests should not be applied to measures of regional cerebral metabolism. They offer no evidence that such measures uniquely escape the mathematical foundations of statistical inference but, instead, merely note the definitional relations among various descriptive and inferential statistics. They falsely claim that these relations invalidate conclusions drawn from statistical tests, and in making this claim, they commit several quite serious logical fallacies. These fallacies are examined.

Brain Mapping↗

Statistical inference by confidence intervals: issues of interpretation and utilization.

This article examines the role of the confidence interval (CI) in statistical inference and its advantages over conventional hypothesis testing, particularly when data are applied in the context of clinical practice. A CI provides a range of population values with which a sample statistic is consistent at a given level of confidence (usually 95%). Conventional hypothesis testing serves to either reject or retain a null hypothesis. A CI, while also functioning as a hypothesis test, provides additional information on the variability of an observed sample statistic (ie, its precision) and on its probable relationship to the value of this statistic in the population from which the sample was drawn (ie, its accuracy). Thus, the CI focuses attention on the magnitude and the probability of a treatment or other effect. It thereby assists in determining the clinical usefulness and importance of, as well as the statistical significance of, findings. The CI is appropriate for both parametric and nonparametric analyses and for both individual studies and aggregated data in meta-analyses. It is recommended that, when inferential statistical analysis is performed, CIs should accompany point estimates and conventional hypothesis tests wherever possible.

Bias↗

Statistical inference in the analysis of radiation compliance and its relation to treatment outcome.

A recent report in the literature was unable to associate the local--regional recurrence rates in breast cancer with deviations from the protocol's recommended dosages. The investigators proceeded to infer that there was "acceptable leeway" in the use of radiotherapy. The present manuscript examines the validity of certain statistical inferences drawn by these investigators.

Adenocarcinoma↗

Statistical inference for relative potency in bivariate dose-response assays with correlated responses.

The analysis of dose-response assays measuring two correlated responses is considered. Attention is given to statistical inference for the potency ratio. Results from a simulation study suggest that a post hoc adjustment for the correlation in parameter estimates obtained from univariate fits provides nearly as much power to detect differences in potency as a bivariate response model fit.

Algorithms↗

Statistical inferences about injury and persistence of environmentally stressed bacteria.

A standard technique for ascertaining the survival characteristics of bacteria after being environmentally stressed is to incubate the bacteria on both selective and non-selective media and count the colonies produced. Based on these colony counts, indexes of injury and persistence of the bacteria are calculated. To compare the stress of two different environments, a persistence ratio is calculated. In this paper, methods of statistical inference concerning these indexes and ratios are presented. These statistical methods use well-known procedures for analysis of binomial data and 2 times 2 table data, and are appropriate when the colony counts follow a Possion distribution.

Bacteria↗

Statistical inference using the g or K point pattern spatial statistics.

Spatial point pattern analysis provides a statistical method to compare an observed spatial pattern against a hypothesized spatial process model. The G statistic, which considers the distribution of nearest neighbor distances, and the K statistic, which evaluates the distribution of all neighbor distances, are commonly used in such analyses. One method of employing these statistics involves building a simulation envelope from the result of many simulated patterns of the hypothesized model. Specifically, a simulation envelope is created by calculating, at every distance, the minimum and maximum results computed across the simulated patterns. A statistical test is performed by evaluating where the results from an observed pattern fall with respect to the simulation envelope. However, this method, which differs from P. Diggle's suggested approach, is invalid for inference because it violates the assumptions of Monte Carlo methods and results in incorrect type I error rate performance. Similarly, using the simulation envelope to estimate the range of distances over which an observed pattern deviates from the hypothesized model is also suspect. The technical details of why the simulation envelope provides incorrect type I error rate performance are described. A valid test is then proposed, and details about how the number of simulated patterns impacts the statistical significance are explained. Finally, an example of using the proposed test within an exploratory data analysis framework is provided.

Data Interpretation, Statistical↗

Introduction to biostatistics: Part 4, statistical inference techniques in hypothesis testing.

Statistical methods used to test the null hypothesis are termed tests of significance. Selection of an appropriate test of significance is dependent on the type of data to be analyzed and the number of groups to be compared. Parametric tests of significance are based on the parameters, mean, standard deviation, and variance, and thus are used appropriately when interval or ratio data are analyzed. The t-test and analysis of variance (ANOVA) are examples of parametric tests of significance. Assumptions regarding the data to be analyzed when using the t-test or ANOVA include normality of the populations from which the sample data are drawn, homogeneity of the variances of the populations from which the sample data are drawn, and independence of the data points within a sample group. The t-test is the appropriate test of significance to use if there are only two groups to compare. If there are three or more groups to compare, ANOVA is the appropriate test. ANOVA holds the preset alpha level constant. While ANOVA will imply a significant difference between the groups compared, a multiple comparison test will define which of the three or more groups differ significantly.

Analysis of Variance↗

Statistical inference for simultaneous clustering of gene expression data.

Current methods for analysis of gene expression data are mostly based on clustering and classification of either genes or samples. We offer support for the idea that more complex patterns can be identified in the data if genes and samples are considered simultaneously. We formalize the approach and propose a statistical framework for two-way clustering. A simultaneous clustering parameter is defined as a function theta=Phi(P) of the true data generating distribution P, and an estimate is obtained by applying this function to the empirical distribution P(n). We illustrate that a wide range of clustering procedures, including generalized hierarchical methods, can be defined as parameters which are compositions of individual mappings for clustering patients and genes. This framework allows one to assess classical properties of clustering methods, such as consistency, and to formally study statistical inference regarding the clustering parameter. We present results of simulations designed to assess the asymptotic validity of different bootstrap methods for estimating the distribution of Phi(P(n)). The method is illustrated on a publicly available data set.

Cluster Analysis↗

Statistical inference for cancer trials with treatment switching.

In cancer clinical trials, it is not uncommon that some patients switched their treatments due to lack of efficacy and/or disease progression under ethical consideration. This treatment switch makes it difficult for the evaluation of the efficacy of the treatment under investigation. The current existing methods consider random treatment switch and do not take into consideration of prognosis and/or investigator's assessment that leads to patients' treatment switch. In this paper, we model patients' treatment switching effect in a latent event times model under parametric setting or a latent hazard rate model under the semi-parametric proportional hazard model. Statistical inference procedures under both models are provided. A simulation study is performed to investigate the performance of the proposed methods.

Antineoplastic Agents↗