Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Combining biomarkers to detect disease with application to prostate cancer.

In early detection of disease, combinations of biomarkers promise improved discrimination over diagnostic tests based on single markers. An example of this is in prostate cancer screening, where additional markers have been sought to improve the specificity of the conventional Prostate-Specific Antigen (PSA) test. A marker of particular interest is the percent free PSA. Studies evaluating the benefits of percent free PSA reflect the need for a methodological approach that is statistically valid and useful in the clinical setting. This article presents methods that address this need. We focus on and-or combinations of biomarker results that we call logic rules and present novel definitions for the ROC curve and the area under the curve (AUC) that are applicable to this class of combination tests. Our estimates of the ROC and AUC are amenable to statistical inference including comparisons of tests and regression analysis. The methods are applied to data on free and total PSA levels among prostate cancer cases and matched controls enrolled in the Physicians' Health Study.

Aged↗

Parsimonious modelling of water and suspended sediment flux from nested catchments affected by selective tropical forestry.

The ability to model the suspended sediment flux (SSflux) and associated water flow from terrain affected by selective logging is important to the establishment of credible measures to improve the ecological sustainability of forestry practices. Recent appreciation of the impact of parameter uncertainty on the statistical credibility of complex models with little internal state validation supports the use of more parsimonious approaches such as data-based mechanistic (DBM) modelling. The DBM approach combines physically based understanding with model structure identification based on transfer functions and objective statistical inference. Within this study, these approaches have been newly applied to rainfall-SSflux response. The dynamics of the sediment system, together with the rainfall-river flow system, were monitored at five nested contributory areas within a 44 ha headwater region in Malaysian Borneo. The data series analysed covered a whole year at a 5 min resolution, and were collected during a period some five to six years after selective timber harvesting had ceased. Physically based and statistical interpretation of these data was possible given the wealth of contemporary and past hydrogeomorphic data collected within the same region. The results indicated that parsimonious, three-parameter models of rainfall-river flow and rainfall-SSflux for the whole catchment describe 80 and 90% of the variance, respectively, and that parameter changes between scales could be explained in physically meaningful terms. Indeed, the modelling indicated some new conceptual descriptions of the river flow and sediment-generation systems. An extreme rainstorm having a 10-20 year return period was present within the data series and was shown to generate new mass movements along the forestry roads that had a differential impact on the monitored contributory areas. Critically, this spatially discrete behaviour was captured by the modelling and may indicate the potential use of DBM approaches for (i) predicting the differential effect of alternative forestry practices, (ii) estimating uncertainty in the behaviour of ungauged areas and (iii) forecasting river flow and SSflux in terrain with temporal changes in rainfall regime and forestry impacts.

Ecosystem↗

Optimal predictions in everyday cognition.

Human perception and memory are often explained as optimal statistical inferences that are informed by accurate prior probabilities. In contrast, cognitive judgments are usually viewed as following error-prone heuristics that are insensitive to priors. We examined the optimality of human cognition in a more realistic context than typical laboratory studies, asking people to make predictions about the duration or extent of everyday phenomena such as human life spans and the box-office take of movies. Our results suggest that everyday cognitive judgments follow the same optimal statistical principles as perception and memory, and reveal a close correspondence between people's implicit probabilistic models and the statistics of the world.

Cognition↗

The likelihood ratio for comparing means when a portion of the subjects fail to respond.

Results of clinical studies are often obscured by the fact that some of the subjects improve with treatment while others do not. As a consequence, for example, we may obtain a contaminated distribution in a treatment group, having one component similar to the entire distribution for a control group and the other shifted by the treatment effect. Maximum-likelihood estimation and the likelihood ratio test for investigating the proportion of responders together with the treatment effect are based on asymptotic theory and use iterative maximization techniques. We investigate the sampling distribution of the estimates and the test and demonstrate their use in statistical inference. Our results suggest that the chi-square distribution with 1.4 degrees of freedom provides a useful test criterion even when the sample sizes are as small as 10. Only in obvious testing situations is the t-test adequate.

Clinical Trials as Topic↗

Using the Past to Predict the Present: Confidence Intervals for Regression Equations in Phylogenetic Comparative Methods.

Two phylogenetic comparative methods, independent contrasts and generalized least squares models, can be used to determine the statistical relationship between two or more traits. We show that the two approaches are functionally identical and that either can be used to make statistical inferences about values at internal nodes of a phylogenetic tree (hypothetical ancestors), to estimate relationships between characters, and to predict values for unmeasured species. Regression equations derived from independent contrasts can be placed back onto the original data space, including computation of both confidence intervals and prediction intervals for new observations. Predictions for unmeasured species (including extinct forms) can be made increasingly accurate and precise as the specificity of their placement on a phylogenetic tree increases, which can greatly increase statistical power to detect, for example, deviation of a single species from an allometric prediction. We reexamine published data for basal metabolic rates (BMR) of birds and show that conventional and phylogenetic allometric equations differ significantly. In new results, we show that, as compared with nonpasserines, passerines exhibit a lower rate of evolution in both body mass and mass-corrected BMR; passerines also have significantly smaller body masses than their sister clade. These differences may justify separate, clade-specific allometric equations for prediction of avian basal metabolic rates.

allometry↗

[Bayesian thinking on its way into medical statistics?].

BACKGROUND: Bayesian statistical analysis is a paradigm quite different from traditional statistical inference. We wanted to show the usefulness of this approach for some medical problems. MATERIALS AND METHODS: We started with Bayes equation as it is used for estimating the probability of illness based on a specific laboratory test. We also looked into a recent Cochrane report on mammography that accepted two studies as valid and five others as biased. In comparison we used examples of clinical trials from other areas that have been misinterpreted by the use of a traditional statistical approach only. RESULTS: We found that by taking into account our prior beliefs about the likely effects on breast cancer mortality of routine radiological screening programmes, the new data fit well into an estimate of a 5% mortality reduction with a 77% chance that there is a positive effect of screening. INTERPRETATION: Bayesian statistics is helpful in making decisions on the basis of experimental evidence by taking into account our prior knowledge, whereas p-values in traditional statistics only give information on how often we will end up with a false positive conclusion in the long run.

Bayes Theorem↗

A shift from significance test to hypothesis test through power analysis in medical research.

Medical research literature until recently, exhibited substantial dominance of the Fisher's significance test approach of statistical inference concentrating more on probability of type I error over Neyman-Pearson's hypothesis test considering both probability of type I and II error. Fisher's approach dichotomises results into significant or not significant results with a P value. The Neyman-Pearson's approach talks of acceptance or rejection of null hypothesis. Based on the same theory these two approaches deal with same objective and conclude in their own way. The advancement in computing techniques and availability of statistical software have resulted in increasing application of power calculations in medical research and thereby reporting the result of significance tests in the light of power of the test also. Significance test approach, when it incorporates power analysis contains the essence of hypothesis test approach. It may be safely argued that rising application of power analysis in medical research may have initiated a shift from Fisher's significance test to Neyman-Pearson's hypothesis test procedure.

Biomedical Research↗

Missing... presumed at random: cost-analysis of incomplete data.

When collecting patient-level resource use data for statistical analysis, for some patients and in some categories of resource use, the required count will not be observed. Although this problem must arise in most reported economic evaluations containing patient-level data, it is rare for authors to detail how the problem was overcome. Statistical packages may default to handling missing data through a so-called 'complete case analysis', while some recent cost-analyses have appeared to favour an 'available case' approach. Both of these methods are problematic: complete case analysis is inefficient and is likely to be biased; available case analysis, by employing different numbers of observations for each resource use item, generates severe problems for standard statistical inference. Instead we explore imputation methods for generating 'replacement' values for missing data that will permit complete case analysis using the whole data set and we illustrate these methods using two data sets that had incomplete resource use information.

Algorithms↗

A generalized estimating equations approach to quantitative trait locus detection of non-normal traits.

To date, most statistical developments in QTL detection methodology have been directed at continuous traits with an underlying normal distribution. This paper presents a method for QTL analysis of non-normal traits using a generalized linear mixed model approach. Development of this method has been motivated by a backcross experiment involving two inbred lines of mice that was conducted in order to locate a QTL for litter size. A Poisson regression form is used to model litter size, with allowances made for under- as well as over-dispersion, as suggested by the experimental data. In addition to fixed parity effects, random animal effects have also been included in the model. However, the method is not fully parametric as the model is specified only in terms of means, variances and covariances, and not as a full probability model. Consequently, a generalized estimating equations (GEE) approach is used to fit the model. For statistical inferences, permutation tests and bootstrap procedures are used. This method is illustrated with simulated as well as experimental mouse data. Overall, the method is found to be quite reliable, and with modification, can be used for QTL detection for a range of other non-normally distributed traits.

Animals↗

Against Popperized epidemiology.

The recommendation of Popper's philosophy of science should be adopted by epidemiologists is disputed. Reference is made to other authors who have shown that the most constructive elements in Popper's ideas have been advocated by earlier philosophers and have been used in epidemiology without abandoning inductive reasoning. It is argued that Popper's denigration of inductive methods is particularly harmful to epidemiology. Inductive reasoning and statistical inference play a key role in the science; it is suggested that unfamiliarity with these ideas contributes to widespread misunderstanding of the function of epidemiology. Attention is drawn to a common fallacy involving correlations between three random variables. The prevalence of the fallacy may be related to confusion between deductive and inductive logic.

Epidemiology↗

Statistical methods and software for the analysis of highthroughput reverse genetic assays using flow cytometry readouts.

Highthroughput cell-based assays with flow cytometric readout provide a powerful technique for identifying components of biologic pathways and their interactors. Interpretation of these large datasets requires effective computational methods. We present a new approach that includes data pre-processing, visualization, quality assessment, and statistical inference. The software is freely available in the Bioconductor package prada. The method permits analysis of large screens to detect the effects of molecular interventions in cellular systems.

Databases, Factual↗

The relative risk in a cohort study with Poisson cases.

This paper deals with making statistical inference about the relative risk (or risk ratio) in a cohort (or prospective) study with dichotomous exposure when the number of cases is a Poisson distributed variable. The exact procedure for testing the null hypothesis for the relative risk and the exact computation of its confidence interval for a single 2 X 2 table is presented. Maximum likelihood methods and the homogeneity test are presented for the common risk ratio when data is stratified in several 2 X 2 tables. These methods are based upon a sufficient statistic and therefore are considered proper statistical alternatives to the more descriptive epidemiological measures such as (in)directly standardized mortality (morbidity) ratios. All computations can be done on a programmable pocket calculator. With the HP-41 CV more than 70 strata can be distinguished.

Adolescent↗

Statistical genetics concepts and approaches in schizophrenia and related neuropsychiatric research.

Statistical genetics is a research field that focuses on mathematical models and statistical inference methodologies that relate genetic variations (ie, naturally occurring human DNA sequence variations or "polymorphisms") to particular traits or diseases (phenotypes) usually from data collected on large samples of families or individuals. The ultimate goal of such analysis is the identification of genes and genetic variations that influence disease susceptibility. Although of extreme interest and importance, the fact that many genes and environmental factors contribute to neuropsychiatric diseases of public health importance (eg, schizophrenia, bipolar disorder, and depression) complicates relevant studies and suggests that very sophisticated mathematical and statistical modeling may be required. In addition, large-scale contemporary human DNA sequencing and related projects, such as the Human Genome Project and the International HapMap Project, as well as the development of high-throughput DNA sequencing and genotyping technologies have provided statistical geneticists with a great deal of very relevant and appropriate information and resources. Unfortunately, the use of these resources and their interpretation are not straightforward when applied to complex, multifactorial diseases such as schizophrenia. In this brief and largely nonmathematical review of the field of statistical genetics, we describe many of the main concepts, definitions, and issues that motivate contemporary research. We also provide a discussion of the most pressing contemporary problems that demand further research if progress is to be made in the identification of genes and genetic variations that predispose to complex neuropsychiatric diseases.

Chromosome Mapping↗

Intent-to-treat analysis for clinical trials: use of data collected after termination of treatment protocol.

Following patients for a period of time after termination of treatment protocols is a common practice in clinical trials of drug treatments. After termination of the protocol treatment, patients are usually provided with medical care as deemed appropriate, e.g. different doses of the same drug, augmentation with other drugs, or different drugs. Data collected on patients when they are not in the treatment protocol are termed "off-treatment" data to be differentiated from "on-treatment" data that are collected while patients are in the treatment protocol. The purpose of the present paper is to describe some recent statistical methodological advances in the use of "off-treatment" data for intent-to-treat (IT) analysis using mixture models [Biometrics 52 (1996) 1002]. Two-piece spline models conditional on the protocol treatment dropout time are developed first. The two pieces of the spline model represent the on-treatment and off-treatment segments of data and are joined at the dropout time. The weighted average of the two-piece conditional models across realizations of the dropout times provides the mixture model for the pragmatic IT analysis. The weights are the estimates of the probabilities of the dropout times. The mixture model is amenable to an explanatory analysis which assumes that the patients remain on their assigned treatments. The model allows the parameters of the splines to depend on the dropout time. Statistical inference is based on bootstrap sampling procedures. We have illustrated this methodology using data from a drug trial comparing nortriptyline and paroxetine in the treatment of major depression in older patients.

Clinical Trials as Topic↗

Statistical validity for testing associations between genetic markers and quantitative traits in family data.

In genetic analysis it is often of interest to analyze associations between traits of unknown genetic etiology and genetic markers from pedigree data. Statistical methods that assume independence of pedigree members cannot be used because they disregard the statistical dependencies of members in a pedigree. For quantitative traits, a regression model proposed by George and Elston [Genet Epidemiol 4:193-201, 1987] uses an asymptotic likelihood ratio test and incorporates a correlation structure that allows for statistical dependence among the pedigree members. The statistical validity of this test is assessed for finite samples by measuring the discrepancy between the empirical and theoretical chi-square distributions. The variance of the mean of the dependent variable is determined to be related to this discrepancy and can be used to determine whether a pedigree structure is large enough for making valid statistical inferences on the basis of the asymptotic test. A multi-generational pedigree of 200 or so individuals should in many cases be sufficient for valid results when using the asymptotic likelihood ratio test for the association between markers and continuous traits.

Family↗

Exchangeability in multivariate Markov chain models.

Time-homogeneous Markov chain models with state space [0, 1]k are useful in analysis of binary follow-up data on k individuals that interact. The number of parameters increases exponentially with k so more restrictive models are imperative for statistical inference. The hypothesis that the matrix of transition probabilities is invariant under permutation of individuals is discussed. It is shown that if individuals are exchangeable, then the process counting the number of individuals occupying a given state is a Markov chain. This reduction of data is sufficient if either at most a single individual may change state between two consecutive time points or if a state is absorbing. Similar results are obtained for exchangeability within two subgroups. Inference in the multivariate process reduces to a univariate problem if individuals are independent given the group's previous response. It is shown how conditional independence could be tested assuming exchangeability. The different hypotheses re examined in an analysis of the occurrence of bacteria in milk samples of Danish dairy cattle.

Animals↗

Estimating haplotype-disease associations with pooled genotype data.

The genetic dissection of complex human diseases requires large-scale association studies which explore the population associations between genetic variants and disease phenotypes. DNA pooling can substantially reduce the cost of genotyping assays in these studies, and thus enables one to examine a large number of genetic variants on a large number of subjects. The availability of pooled genotype data instead of individual data poses considerable challenges in the statistical inference, especially in the haplotype-based analysis because of increased phase uncertainty. Here we present a general likelihood-based approach to making inferences about haplotype-disease associations based on possibly pooled DNA data. We consider cohort and case-control studies of unrelated subjects, and allow arbitrary and unequal pool sizes. The phenotype can be discrete or continuous, univariate or multivariate. The effects of haplotypes on disease phenotypes are formulated through flexible regression models, which allow a variety of genetic hypotheses and gene-environment interactions. We construct appropriate likelihood functions for various designs and phenotypes, accommodating Hardy-Weinberg disequilibrium. The corresponding maximum likelihood estimators are approximately unbiased, normally distributed, and statistically efficient. We develop simple and efficient numerical algorithms for calculating the maximum likelihood estimators and their variances, and implement these algorithms in a freely available computer program. We assess the performance of the proposed methods through simulation studies, and provide an application to the Finland-United States Investigation of NIDDM Genetics Study. The results show that DNA pooling is highly efficient in studying haplotype-disease associations. As a by-product, this work provides valid and efficient methods for estimating haplotype-disease associations with unpooled DNA samples.

Algorithms↗

Detecting differential gene expression with a semiparametric hierarchical mixture method.

Mixture modeling provides an effective approach to the differential expression problem in microarray data analysis. Methods based on fully parametric mixture models are available, but lack of fit in some examples indicates that more flexible models may be beneficial. Existing, more flexible, mixture models work at the level of one-dimensional gene-specific summary statistics, and so when there are relatively few measurements per gene these methods may not provide sensitive detectors of differential expression. We propose a hierarchical mixture model to provide methodology that is both sensitive in detecting differential expression and sufficiently flexible to account for the complex variability of normalized microarray data. EM-based algorithms are used to fit both parametric and semiparametric versions of the model. We restrict attention to the two-sample comparison problem; an experiment involving Affymetrix microarrays and yeast translation provides the motivating case study. Gene-specific posterior probabilities of differential expression form the basis of statistical inference; they define short gene lists and false discovery rates. Compared to several competing methodologies, the proposed methodology exhibits good operating characteristics in a simulation study, on the analysis of spike-in data, and in a cross-validation calculation.

Algorithms↗