Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Statistical analysis of a small set of time-ordered gene expression data using linear splines.

MOTIVATION: Recently, the temporal response of genes to changes in their environment has been investigated using cDNA microarray technology by measuring the gene expression levels at a small number of time points. Conventional techniques for time series analysis are not suitable for such a short series of time-ordered data. The analysis of gene expression data has therefore usually been limited to a fold-change analysis, instead of a systematic statistical approach. METHODS: We use the maximum likelihood method together with Akaike's Information Criterion to fit linear splines to a small set of time-ordered gene expression data in order to infer statistically meaningful information from the measurements. The significance of measured gene expression data is assessed using Student's t-test. RESULTS: Previous gene expression measurements of the cyanobacterium Synechocystis sp. PCC6803 were reanalyzed using linear splines. The temporal response was identified of many genes that had been missed by a fold-change analysis. Based on our statistical analysis, we found that about four gene expression measurements or more are needed at each time point.

Algorithms↗

Statistical considerations for a medical data base.

Some of the special statistical problems associated with the design and analysis of observational studies arising from medical data bases are reviewed. Particular attention is given to the adequacy of the collected data to provide information from which to draw valid statistical inferences.

Clinical Trials as Topic↗

Stochastic model of the overdispersion in the place cell discharge.

The spontaneous firing activity of the place cells reflects the position of an experimental animal in its arena. The firing rate is high inside a part of the arena, called the firing field, and low outside. It is generally accepted concept that this is the way in which the hippocampus stores a map of the environment. This well known fact was recently reinvestigated [Fenton, A.A., Muller, R.U., 1998. Proc. Natl. Acad. Sci. USA 95, 3182-3187] and it was found that while the activity was highly reliable in position, it did not retain the same reliability in time. The number of action potentials fired during different passes through the firing field were substantially different (overdispersion). We present a mathematical model based on a doubly stochastic Poisson process which is able to reproduce the experimental findings. Further, it enables us to propose specific statistical inference on the experiments in aim to verify data and model compatibility. The model permits to speculate about the neural mechanisms leading to the overdispersion in the activity of the hippocampal place cells. Namely, the statistical variation of the intensity of firing can be achieved, for example, by introducing a hierarchical structure into the local neural network.

Hippocampus↗

Prediction of splice sites with dependency graphs and their expanded bayesian networks.

MOTIVATION: Owing to the complete sequencing of human and many other genomes, huge amounts of DNA sequence data have been accumulated. In bioinformatics, an important issue is how to predict the complete structure of genes from the genomic DNA sequence, especially the human genome. A crucial part in the gene structure prediction is to determine the precise exon-intron boundaries, i.e. the splice sites, in the coding region. RESULTS: We have developed a dependency graph model to fully capture the intrinsic interdependency between base positions in a splice site. The establishment of dependency between two position is based on a chi2-test from known sample data. To facilitate statistical inference, we have expanded the dependency graph (which is usually a graph with cycles that make probabilistic reasoning very difficult, if not impossible) into a Bayesian network (which is a directed acyclic graph that facilitates statistical reasoning). When compared with the existing models such as weight matrix model, weight array model, maximal dependence decomposition, Cai et al.'s tree model as well as the less-studied second-order and third-order Markov chain models, the expanded Bayesian networks from our dependency graph models perform the best in nearly all the cases studied. AVAILABILITY: Software (a program called DGSplicer) and datasets used are available at http://csrl.ee.nthu.edu.tw/bioinf/ CONTACT: cclu@ee.nthu.edu.tw.

Bayes Theorem↗

Sampling size in the verification of manufactured-supplied air kerma strengths.

Quality control mandate that the air kerma strengths (S(K)) of permanent seeds be verified, this is usually done by statistics inferred from 10% of the seeds. The goal of this paper is to proposed a new sampling method in which the number of seeds to be measured will be set beforehand according to an a priori statistical level of uncertainty. The results are based on the assumption that the S(K) has a normal distribution. To demonstrate this, the S(K) of each of the seeds measured was corrected to ensure that the average S(K) of its sample remained the same. In this process 2030 results were collected and analyzed using a normal plot. In our opinion, the number of seeds sampled should be determined beforehand according to an a priori level of statistical uncertainty.

Air↗

[Statistically validated evaluation of clinical trials].

Data of clinical trials of medicinal products must be evaluated in statistically valid models. The statistical validity criteria are defined. Statistically invalid models will result in biased parameter and confidence interval estimations, erroneous statistical inferences and clinical interpretations. Finally, wrong decisions will call forth deleterious consequences in the judgement of the therapeutic effect and the frequency and severity of the adverse reactions of the tested new medicinal, and generic products. Statistically validated analyses will promote the international harmonization of the scientific evaluation of medicinal products according to the idea of the evidence-based-medicine. The study presents examples of clinical trials evaluated with a software checking statistical validity assumptions while performing evaluation of data.

Clinical Trials as Topic↗

Inference for the dependent competing risks model with masked causes of failure.

The competing risks model is useful in settings in which individuals/units may die/fail for different reasons. The cause specific hazard rates are taken to be piecewise constant functions. A complication arises when some of the failures are masked within a group of possible causes. Traditionally, statistical inference is performed under the assumption that the failure causes act independently on each item. In this paper we propose an EM-based approach which allows for dependent competing risks and produces estimators for the sub-distribution functions. We also discuss identifiability of parameters if none of the masked items have their cause of failure clarified in a second stage analysis (e.g. autopsy). The procedures proposed are illustrated with two datasets.

Algorithms↗

Combining voxel intensity and cluster extent with permutation test framework.

In a massively univariate analysis of brain image data, statistical inference is typically based on intensity or spatial extent of signals. Voxel intensity-based tests provide great sensitivity for high intensity signals, whereas cluster extent-based tests are sensitive to spatially extended signals. To benefit from the strength of both, the intensity and extent information needs to be combined. Various ways of combining voxel intensity and cluster extent are possible, and a few such combining methods have been proposed. Poline et al.'s [NeuroImage 16 (1997) 83] minimum P value approach is sensitive to signals whose either intensity or extent is significant. Bullmore et al.'s [IEEE Trans. Med. Imag. 18 (1999) 32] cluster mass method can detect signals whose intensity and extent are sufficiently large, even when they are not significant by intensity or extent alone. In this work, we study such combined inference methods using combining functions (Pesarin, F., 2001. Multivariate Permutation Tests. Wiley, New York) and permutation framework [Holmes et al., J. Cereb. Blood Flow Metab. 16 (1996) 7], which allow us to examine different ways of combining voxel intensity and cluster extent information without knowing their distribution. We also attempt to calibrate combined inference by using weighted combining functions, which adjust the test according to signals of interest. Furthermore, we propose meta-combining, a combining function of combining functions, which integrates strengths of multiple combining functions into a single statistic. We found that combined tests are able to detect signals that are not detected by voxel or cluster size test alone. We also found that the weighted combining functions can calibrate the combined test according to the signals of interest, emphasizing either intensity or extent as appropriate. Though not necessarily more sensitive than individual combining functions, the meta-combining function is sensitive to all types of signals and thus can be used as a single test summarizing all the combining functions.

Artifacts↗

Residual analysis in random regressions using SAS and S-PLUS.

A program package RRAP: Random Regression Residual Analysis Program using SAS [1] and S-PLUS [2] is available for performing random regression residual analysis. The PROCEDURE MIXED from SAS is used for statistical inference. Both elementary-level and individual-level residuals are used. The S-PLUS programs provide: (1) a transformation to orthogonalize the elementary-level correlated residuals for standard regression residual analyses; and (2) several statistics and plots for checking model assumptions, assessing model fitting and detecting outlying individuals. RRRAP starts with a SAS Macro RRRAPMAC on the data followed by a S-PLUS Program DoRRRAP on a UNIX system.

Models, Statistical↗

Analytic complexities associated with group therapy in substance abuse treatment research: problems, recommendations, and future directions.

In community-based alcoholism and drug abuse treatment programs, the vast majority of interventions are delivered in a group therapy context. In turn, treatment providers and funding agencies have called for more research on interventions delivered in groups in an effort to make the emerging empirical literature on the treatment of substance abuse more ecologically valid. Unfortunately, the complexity of data structures derived from therapy groups (because of member interdependence and changing membership over time) and the present lack of statistically valid and generally accepted approaches to analyzing these data have had a significant stifling effect on group therapy research. This article (a) describes the analytic challenges inherent in data generated from therapy groups, (b) outlines common (but flawed) analytic and design approaches investigators often use to address these issues (e.g., ignoring group-level nesting, treating data from therapy groups with changing membership as fully hierarchical), and (c) provides recommendations for handling data from therapy groups using presently available methods. In addition, promising data-analytic frameworks that may eventually serve as foundations for the development of more appropriate analytic methods for data from group therapy research (i.e., nonhierarchical data modeling, pattern-mixture approaches) are also briefly described. Although there are other substantial obstacles that impede rigorous research on therapy groups (e.g., evaluation and measurement of group process, limited control over treatment delivery ingredients), addressing data-analytic problems is critical for improving the accuracy of statistical inferences made from research on ecologically valid group-based substance abuse interventions.

Humans↗

Bayesian inference on biopolymer models.

MOTIVATION: Most existing bioinformatics methods are limited to making point estimates of one variable, e.g. the optimal alignment, with fixed input values for all other variables, e.g. gap penalties and scoring matrices. While the requirement to specify parameters remains one of the more vexing issues in bioinformatics, it is a reflection of a larger issue: the need to broaden the view on statistical inference in bioinformatics. RESULTS: The assignment of probabilities for all possible values of all unknown variables in a problem in the form of a posterior distribution is the goal of Bayesian inference. Here we show how this goal can be achieved for most bioinformatics methods that use dynamic programming. Specifically, a tutorial style description of a Bayesian inference procedure for segmentation of a sequence based on the heterogeneity in its composition is given. In addition, full Bayesian inference algorithms for sequence alignment are described. AVAILABILITY: Software and a set of transparencies for a tutorial describing these ideas are available at http://www.wadsworth.org/res&res/bioinfo/

Bayes Theorem↗

A Monte Carlo evaluation of three statistical methods used in path analysis.

Results of a Monte Carlo study to investigate the properties of three statistical methods used extensively in path analysis of family data are presented. All three methods are based on the maximum likelihood principle and involve the assumptions of multivariate normality and large sample (asymptotic) statistical properties. The methods differ, however, in the specification of the likelihood function. Given a set of correlation estimates, method 1 maximizes the likelihood function under the stipulation that the estimates are independent. Method 2 differs from the former by allowing for covariances among the correlation estimators. Method 3 involves (direct) maximization of the likelihood function for the individual family observations assuming multivariate normality for the vector of family observations. The Monte Carlo study investigated validity of the test statistics and confidence intervals and evaluated the relative efficiency and bias of the parameter estimates based on 1,000 replications of each of several simulation conditions. The effects of violating the two basic assumptions, multivariate normality and asymptotic theory, were investigated by comparing results for non-normally vs normally distributed family data and for small vs large sample sizes. It is shown that method 3 provides valid statistical inferences under multivariate normality and that it is generally robust against minor departures from normality. Method 2 is also robust against minor deviations from normality, but it is sensitive to small sample sizes. Method 1 yields highly conservative test statistics under all conditions studied.

Genetics, Medical↗

Cost-effectiveness analysis when the WTA is greater than the WTP.

The incremental cost effectiveness ratio has long been the standard parameter of interest in the assessment of the cost-effectiveness of a new treatment. However, due to concerns with interpretability and statistical inference, authors have suggested using the willingness-to-pay for a unit of health benefit to define the incremental net benefit as an alternative. The incremental net benefit has a more consistent interpretation and is amenable to routine statistical procedures. These procedures rely on the fact that the willingness-to-accept compensation for a loss of a unit of health benefit (at some cost saving) is the same as the willingness-to-pay for it. Theoretical and empirical evidence suggest, however, that in health care the willingness-to-accept is about twice as much as the willingness-to-pay. We use Bayesian methods to provide a statistical procedure for the cost-effectiveness comparison of two arms of a randomized clinical trial that allows the willingness-to-pay and the willingness-to-accept to have different values. An example is provided.

Bayes Theorem↗

Calculating percentage prediction error: a user's note.

The equations of calculation of percentage prediction error (percentage prediction error = [equation: see text] x 100 or percentage prediction error = [equation: see text] x 100) and similar equations have been widely used. However, not much is known about the property of this type of equation and the caution which should be taken into account when using this type of equation. Moreover, little is known about the power of percentage prediction error as statistical inference. In the present study we address these points in the use of this type of equation.

Bias↗

Gene expression analysis with the parametric bootstrap.

Recent developments in microarray technology make it possible to capture the gene expression profiles for thousands of genes at once. With this data researchers are tackling problems ranging from the identification of 'cancer genes' to the formidable task of adding functional annotations to our rapidly growing gene databases. Specific research questions suggest patterns of gene expression that are interesting and informative: for instance, genes with large variance or groups of genes that are highly correlated. Cluster analysis and related techniques are proving to be very useful. However, such exploratory methods alone do not provide the opportunity to engage in statistical inference. Given the high dimensionality (thousands of genes) and small sample sizes (often <30) encountered in these datasets, an honest assessment of sampling variability is crucial and can prevent the over-interpretation of spurious results. We describe a statistical framework that encompasses many of the analytical goals in gene expression analysis; our framework is completely compatible with many of the current approaches and, in fact, can increase their utility. We propose the use of a deterministic rule, applied to the parameters of the gene expression distribution, to select a target subset of genes that are of biological interest. In addition to subset membership, the target subset can include information about relationships between genes, such as clustering. This target subset presents an interesting parameter that we can estimate by applying the rule to the sample statistics of microarray data. The parametric bootstrap, based on a multivariate normal model, is used to estimate the distribution of these estimated subsets and relevant summary measures of this sampling distribution are proposed. We focus on rules that operate on the mean and covariance. Using Bernstein's Inequality, we obtain consistency of the subset estimates, under the assumption that the sample size converges faster to infinity than the logarithm of the number of genes. We also provide a conservative sample size formula guaranteeing that the sample mean and sample covariance matrix are uniformly within a distance epsilon > 0 of the population mean and covariance. The practical performance of the method using a cluster-based subset rule is illustrated with a simulation study. The method is illustrated with an analysis of a publicly available leukemia data set.

Journal Article↗

Analysis and interpretation of cost data in randomised controlled trials: review of published studies.

OBJECTIVE: To review critically the statistical methods used for health economic evaluations in randomised controlled trials where an estimate of cost is available for each patient in the study. DESIGN: Survey of published randomised trials including an economic evaluation with cost values suitable for statistical analysis; 45 such trials published in 1995 were identified from Medline. MAIN OUTCOME MEASURES: The use of statistical methods for cost data was assessed in terms of the descriptive statistics reported, use of statistical inference, and whether the reported conclusions were justified. RESULTS: Although all 45 trials reviewed apparently had cost data for each patient, only 9 (20%) reported adequate measures of variability for these data and only 25 (56%) gave results of statistical tests or a measure of precision for the comparison of costs between the randomised groups. Only 16 (36%) of the articles gave conclusions which were justified on the basis of results presented in the paper. No paper reported sample size calculations for costs. CONCLUSIONS: The analysis and interpretation of cost data from published trials reveal a lack of statistical awareness. Strong and potentially misleading conclusions about the relative costs of alternative therapies have often been reported in the absence of supporting statistical evidence. Improvements in the analysis and reporting of health economic assessments are urgently required. Health economic guidelines need to be revised to incorporate more detailed statistical advice.

Cost-Benefit Analysis↗

Regression analysis of multivariate grouped survival data.

Multivariate failure time data arise when each study subject may experience several types of event or when there are clusterings of observational units such that failure times within the same cluster are correlated. The failure times are often subject to interval grouping or have truly discrete measurements. In this paper, the marginal distribution for each discrete failure time variable is formulated by a grouped-data version of the proportional hazards model while the dependence structure is unspecified. Generalized estimating equations in the spirit of Liang and Zeger (1986, Biometrika 73, 13-22) are proposed to estimate the regression parameters and survival probabilities. The resulting estimators are consistent and asymptotically normal. Robust estimators for the limiting covariance matrices are constructed. Simulation studies demonstrate that the asymptotic approximations are adequate for practical use and that ignoring the intracluster dependence in the variance-covariance estimation would lead to invalid statistical inference. A psychological experiment is provided for illustration.

Age Factors↗

The power of analysis: statistical perspectives. Part 2.

A discussion of the importance of statistical power in research is presented accompanied by nomograms for determining sample size and statistical power for the Student's paired and unpaired t tests with a Type I error of 5%. A brief review of statistical inference is presented. Some findings from Part I are reviewed.

Clinical Trials as Topic↗