Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

The use of nonparametric statistics in quantitative electron microscopy.

Parametric statistical methods assume samples that have a normal distribution and representative sample sizes (i.e. n >20). Quantitative electron microscopy is inherently restricted to small sample sizes and a priori there is no way to know if the expression of the ligand being studied has a normal distribution. Thus to make statistical inferences based on data generated by quantitative electron microscopy using parametric methods may not be justified. Nonparametric statistical methods offer a tool for the evaluation of data that do not meet the criteria for analysis by parametric methods. In this report I show the utility of using nonparametric statistical methods for the analysis of data generated by quantitative electron microscopy.

Animals↗

Simultaneous inference for longitudinal data with detection limits and covariates measured with errors, with application to AIDS studies.

In AIDS studies such as HIV viral dynamics, statistical inference is often complicated because the viral load measurements may be subject to left censoring due to a detection limit and time-varying covariates such as CD4 counts may be measured with substantial errors. Mixed-effects models are often used to model the response and the covariate processes in these studies. We propose a unified approach which addresses the censoring and measurement errors simultaneously. We estimate the model parameters by a Monte-Carlo EM algorithm via the Gibbs sampler. A simulation study is conducted to compare the proposed method with the usual two-step method and a naive method. We find that the proposed method produces approximately unbiased estimates with more reliable standard errors. A real data set from an AIDS study is analysed using the proposed method.

Acquired Immunodeficiency Syndrome↗

Incremental net benefit in randomized clinical trials.

There are three approaches to health economic evaluation for comparing two therapies. These are (i) cost minimization, in which one assumes or observes no difference in effectiveness, (ii) incremental cost-effectiveness, and (iii) incremental net benefit. The latter can be expressed either in units of effectiveness or costs. When analysing data from a clinical trial, expressing incremental net benefit in units of cost allows the investigator to examine all three approaches in a single graph, complete with the corresponding statistical inferences. Furthermore, if costs and effectiveness are not censored, this can be achieved using common two-sample statistical procedures. The above will be illustrated using two examples, one with censoring and one without.

Amiodarone↗

Structural analysis of sterol distributions in the plasma membrane of living cells.

Although plasma membrane (PM) cholesterol-rich and -poor domains have been isolated by subcellular fractionation, the real-time arrangement of cholesterol in such domains in living cells is still unclear. Therefore, dehydroergosterol (DHE), a naturally occurring fluorescent sterol, was incorporated into cultured L-cell fibroblasts. Two PM markers, the enhanced cyan fluorescent protein (ECFP-Mem) and 3'-dioctadecyloxacarbocyanine perchlorate [DiOC(18)(3)], were used to distinguish DHE localized at the PM of living cells. Spatial enrichment of DHE in the PM of living cells was visualized in real time by multiphoton laser scanning microscopy (MPLSM). Quantitative models and image-processing techniques were developed for statistical analysis of the distribution of DHE within the PM. The PM was resolved from the cytoplasm in a two-step process, and a smooth trajectory reference of the PM was refined by statistical regression and moments-based techniques. Thus, DHE intensities over the PM were measured following the major DHE intensity distributions. Spatial distributions of DHE within the PM were examined by a statistical inference technique, complete spatial randomness (CSR). For PM regions densely populated with DHE, the distributions of DHE exhibited statistical arrangements that were not spatial random (i.e., homogeneous Poisson process) or regular but, instead, exhibited strong cluster patterns. In effect, real-time MPLSM imaging data for the first time demonstrated that sterol enrichment occurred in clustered regions in the PM, consistent with the existence of cholesterol-rich domains in the plasma membrane of living cells.

Animals↗

[Comparison of two measurement methods with a gold standard as applied to 20-MHz-sonography and clinical palpation for ascertaining the thickness of pigmented skin tumours].

AIM: This paper focuses on different statistical methods for comparing two measurement methods with an additionally available gold standard. A given data example is used as the basis of the calculations. METHOD: We provide a complementary statistical analysis of a study presented by Hoffmann et al. on sonometric and palpatory measurements of the size of pigmented skin tumours in 681 patients. RESULTS: For comparing two measurement methods with respect to a gold standard, several statistical parameters assessing one measurement method can be used. In addition, there are further descriptive and some inference-statistical methods available. CONCLUSION: If there is a suitable categorization of the measurements, the comparison of the methods should be performed using the positive predictive values and kappa coefficients as descriptive measures. Moreover, the McNemar test can be used for comparing the differential accuracy of allocation. When investigating continuous measurements, a comparison using mere correlation analyses can lead to false conclusions. Therefore, we recommend the direct analysis of the individual measurement errors by means of numerical and graphical representations. The absolute values of the measurement errors can be compared using the sign test for paired samples.

Biometry↗

Robust accurate identification of peptides (RAId): deciphering MS2 data using a structured library search with de novo based statistics.

MOTIVATION: The key to MS -based proteomics is peptide sequencing. The major challenge in peptide sequencing, whether library search or de novo, is to better infer statistical significance and better attain noise reduction. Since the noise in a spectrum depends on experimental conditions, the instrument used and many other factors, it cannot be predicted even if the peptide sequence is known. The characteristics of the noise can only be uncovered once a spectrum is given. We wish to overcome such issues. RESULTS: We designed RAId to identify peptides from their associated tandem mass spectrometry data. RAId performs a novel de novo sequencing followed by a search in a peptide library that we created. Through de novo sequencing, we establish the spectrum-specific background score statistics for the library search. When the database search fails to return significant hits, the top-ranking de novo sequences become potential candidates for new peptides that are not yet in the database. The use of spectrum-specific background statistics seems to enable RAId to perform well even when the spectral quality is marginal. Other important features of RAId include its potential in de novo sequencing alone and the ease of incorporating post-translational modifications.

Algorithms↗

Methods for improving regression analysis for skewed continuous or counted responses.

Standard inference procedures for regression analysis make assumptions that are rarely satisfied in practice. Adjustments must be made to insure the validity of statistical inference. These adjustments, known for many years, are used routinely by some health researchers but not by others. We review some of these methods and give an example of their use in a health services study for a continuous and a count outcome. For the continuous outcome, we describe re-transformation using the smear factor, accounting for missing cases via multiple imputation and attrition weights and improving results with bootstrap methods. For the count outcome, we describe zero inflated Poisson and negative binomial models and the two-part model to account for overabundance of zero values. Recent advances in computing and software development have produced user-friendly computer programs that enable the data analyst to improve prediction and inference based on regression analysis.

Data Interpretation, Statistical↗

Soft tissue differentiation using multiband signatures of high resolution ultrasonic transmission tomography.

In this paper, we are interested in soft tissue differentiation by multiband images obtained from the High-Resolution Ultrasonic Transmission Tomography (HUTT) system using a spectral target detection method based on constrained energy minimization (CEM). We have developed a new tissue differentiation method (called "CEM filter bank") consisting of multiple CEM filters specially designed for detecting multiple types of tissues. Statistical inference on the output of the CEM filter bank is used to make a decision based on the maximum statistical significance rather than the magnitude of each CEM filter output. We test and validate this method through three-dimensional interphantom/intraphantom soft tissue classification where target profiles obtained from an arbitrary single slice are used for differentiation over multiple other tomographic slices. The performance of the proposed classifier is assessed using receiver operating characteristic analysis. We also apply our method to classify tiny structures inside a bovine kidney and sheep kidneys. Using the proposed method we can detect physical objects and biological tissues such as styrofoam balls, chicken tissue, calyces, and vessel-duct successfully.

Algorithms↗

Statistical analysis of nonlinear structural equation models with continuous and polytomous data.

A general nonlinear structural equation model with mixed continuous and polytomous variables is analysed. A Bayesian approach is proposed to estimate simultaneously the thresholds, the structural parameters and the latent variables. To solve the computational difficulties involved in the posterior analysis, a hybrid Markov chain Monte Carlo method that combines the Gibbs sampler and the Metropolis-Hasting algorithm is implemented to produce the Bayesian solution. Statistical inferences, which involve estimation of parameters and their standard errors, residuals and outliers analyses, and goodness-of-fit statistics for testing the posited model, are discussed. The proposed procedure is illustrated by a simulation study and a real example.

Humans↗

Weighing the results of differing 'low dose' studies of the mouse prostate by Nagel, Cagen, and Ashby: quantification of experimental power and statistical results.

Differing experimental findings with respect to "low dose" responses in the mouse prostate after in utero exposure have generated considerable controversy. An analysis of such controversies requires a broad strength and weight of the evidence approach. For example, a National Toxicology Program review panel acquired the raw data from nearly 50 studies and then statistically reanalyzed these data in a common and comparable approach. However, the statistical power of the various studies was not calculated and the quantitative p values were not reported in this reanalysis. Such calculations and values address vital strength- and weight-of-the-evidence questions: (1) how sensitive were the various studies to detect changes in prostate weight, particularly the negative replicate studies and (2) what were the p values; were negative studies robust or only marginal in their inability to find an effect? We first examined the statistical power of the studies to detect a positive effect on prostate weight. Preliminary calculations indicated that the two subsequent replicating studies were indeed more sensitive to changes in prostate weight in comparison to the original study, having reasonable power to detect an effect at only 50% of the response reported in the original study. Additional calculations were performed using the raw data available from one negative replicating study and the methods recommended by the statistics subpanel of the original review. This analysis used Dunnett's multiple comparison procedure for groups with p<0.05 to infer statistical significance, employed an analysis-of-covariance model with body weight as a covariate, and addressed litter as a nested random effect. The quantitated p values for this replicated study, comparing the two Bisphenol A treatment groups (2 and 20 microg/kg/day) to the control, were 0.821 and 0.972, respectively. This indicates this study was indeed robust in finding no treatment-related effect. Thus, the weight and strength of the evidence, based on sensitivity and quantitative p value, was that it is highly unlikely for this negative replicating study to have missed a true effect. In the future, we recommend a similar use of statistical power analysis for those designing experimental studies and for those conducting weight-of-the-evidence reviews, and we also recommend the clear quantitation and reporting of p values to support the review's interpretation and conclusions.

Algorithms↗

Inference for clinical trials with some protocol amendments.

The use of adaptive methods in clinical development has become very popular in recent years due to its flexibility in modifying trial procedures and/or statistical procedures of on-going clinical trials. Modifications to trial procedures are usually documented by protocol amendments. However, the actual patient population after protocol amendments could deviate from the originally targeted patient population. In addition, protocol amendments made based on accrued data of the on-going trial may distort the sampling distribution of the statistic designed for the case of no protocol change. In this article, we model the population deviations due to protocol amendments using some covariates and study how to develop a valid statistical inference procedure. An example concerning an asthma trial is presented for illustration.

Algorithms↗

When effect sizes disagree: the case of r and d.

The increased use of effect sizes in single studies and meta-analyses raises new questions about statistical inference. Choice of an effect-size index can have a substantial impact on the interpretation of findings. The authors demonstrate the issue by focusing on two popular effect-size measures, the correlation coefficient and the standardized mean difference (e.g., Cohen's d or Hedges's g), both of which can be used when one variable is dichotomous and the other is quantitative. Although the indices are often practically interchangeable, differences in sensitivity to the base rate or variance of the dichotomous variable can alter conclusions about the magnitude of an effect depending on which statistic is used. Because neither statistic is universally superior, researchers should explicitly consider the importance of base rates to formulate correct inferences and justify the selection of a primary effect-size statistic.

Data Interpretation, Statistical↗

Survival analysis in clinical trials: past developments and future directions.

The field of survival analysis emerged in the 20th century and experienced tremendous growth during the latter half of the century. The developments in this field that have had the most profound impact on clinical trials are the Kaplan-Meier (1958, Journal of the American Statistical Association 53, 457-481) method for estimating the survival function, the log-rank statistic (Mantel, 1966, Cancer Chemotherapy Report 50, 163-170) for comparing two survival distributions, and the Cox (1972, Journal of the Royal Statistical Society, Series B 34, 187-220) proportional hazards model for quantifying the effects of covariates on the survival time. The counting-process martingale theory pioneered by Aalen (1975, Statistical inference for a family of counting processes, Ph.D. dissertation, University of California, Berkeley) provides a unified framework for studying the small- and large-sample properties of survival analysis statistics. Significant progress has been achieved and further developments are expected in many other areas, including the accelerated failure time model, multivariate failure time data, interval-censored data, dependent censoring, dynamic treatment regimes and causal inference, joint modeling of failure time and longitudinal data, and Baysian methods.

Biometry↗

Simultaneous inference in epidemiological studies.

Some difficulties encountered in using and interpreting significance tests in both exploratory and hypothesis testing epidemiological studies are discussed. Special consideration is given to the problems of simultaneous statistical inference--how are inferences to be modified when many significance tests are performed on the same set of data? Although some partial solutions are available, greater emphasis on estimation methods and less use of and reliance on significance testing in epidemiological studies is more appropriate.

Epidemiology↗

On the use of permutation in and the performance of a class of nonparametric methods to detect differential gene expression.

MOTIVATION: Recently a class of nonparametric statistical methods, including the empirical Bayes (EB) method, the significance analysis of microarray (SAM) method and the mixture model method (MMM), have been proposed to detect differential gene expression for replicated microarray experiments conducted under two conditions. All the methods depend on constructing a test statistic Z and a so-called null statistic z. The null statistic z is used to provide some reference distribution for Z such that statistical inference can be accomplished. A common way of constructing z is to apply Z to randomly permuted data. Here we point our that the distribution of z may not approximate the null distribution of Z well, leading to possibly too conservative inference. This observation may apply to other permutation-based nonparametric methods. We propose a new method of constructing a null statistic that aims to estimate the null distribution of a test statistic directly. RESULTS: Using simulated data and real data, we assess and compare the performance of the existing method and our new method when applied in EB, SAM and MMM. Some interesting findings on operating characteristics of EB, SAM and MMM are also reported. Finally, by combining the idea of SAM and MMM, we outline a simple nonparametric method based on the direct use of a test statistic and a null statistic.

Algorithms↗

MLE and Bayesian inference of age-dependent sensitivity and transition probability in periodic screening.

This article extends previous probability models for periodic breast cancer screening examinations. The specific aim is to provide statistical inference for age dependence of sensitivity and the transition probability from the disease free to the preclinical state. The setting is a periodic screening program in which a cohort of initially asymptomatic women undergo a sequence of breast cancer screening exams. We use age as a covariate in the estimation of screening sensitivity and the transition probability simultaneously, both from a frequentist point of view and within a Bayesian framework. We apply our method to the Health Insurance Plan of Greater New York study of female breast cancer and give age-dependent sensitivity and transition probability density estimates. The inferential methodology we develop is also applicable when analyzing studies of modalities for early detection of other types of progressive chronic diseases.

Adult↗

Statistics and environmental policy: case studies from long-term environmental monitoring data.

Environmental objectives are statements of policy which are intended to be assessed using information from a monitoring program. An environmental monitoring program has to be adequate in its quality and quantity of data so that the environmental objectives can be assessed. Also, the resulting data should be able to contribute information towards decisions to modify policy at a later time if desirable. However, monitoring programs can fail to return satisfactory information for policymakers because future statistical needs have not been anticipated, potential confounding factors were not considered, or sampling protocols did not specify suitable randomization. A key intermediate role exists for the use of statistical inference in providing a logical framework for using monitoring data to test hypotheses about fulfillment of environmental objectives. The undertaking of a statistical inferential approach to monitoring design can result in data which are more general in their interpretation and, thus, more useful as input to policy development and review. Some case studies from the Victorian EPA monitoring programs will be presented to illustrate these points.

Environmental Monitoring↗

A bayesian statistical algorithm for RNA secondary structure prediction.

A Bayesian approach for predicting RNA secondary structure that addresses the following three open issues is described: (1) the need for a representation of the full ensemble of probable structures; (2) the need to specify a fixed set of energy parameters; (3) the desire to make statistical inferences on all variables in the problem. It has recently been shown that Bayesian inference can be employed to relax or eliminate the need to specify the parameters of bioinformatics recursive algorithms and to give a statistical representation of the full ensemble of probable solutions with the incorporation of uncertainty in parameter values. In this paper, we make an initial exploration of these potential advantages of the Bayesian approach. We present a Bayesian algorithm that is based on stacking energy rules but relaxes the need to specify the parameters. The algorithm returns the exact posterior distribution of the number of destabilizing loops, stacking energy matrices, and secondary structures. The algorithm generates statistically representative structures from the full ensemble of probable secondary structures in exact proportion to the posterior probabilities. Once the forward recursions for the algorithm are completed, the backward recursive sampling executes in O(n) time, providing a very efficient approach for generating representative structures. We demonstrate the utility of the Bayesian approach with several tRNA sequences. The potential of the approach for predicting RNA secondary structures and presenting alternative structures is illustrated with applications to the Escherichia coli tRNA(Ala) sequence and the Xenopus laevis oocyte 5S rRNA sequence.

Algorithms↗