Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

A theory of cortical responses.

This article concerns the nature of evoked brain responses and the principles underlying their generation. We start with the premise that the sensory brain has evolved to represent or infer the causes of changes in its sensory inputs. The problem of inference is well formulated in statistical terms. The statistical fundaments of inference may therefore afford important constraints on neuronal implementation. By formulating the original ideas of Helmholtz on perception, in terms of modern-day statistical theories, one arrives at a model of perceptual inference and learning that can explain a remarkable range of neurobiological facts.It turns out that the problems of inferring the causes of sensory input (perceptual inference) and learning the relationship between input and cause (perceptual learning) can be resolved using exactly the same principle. Specifically, both inference and learning rest on minimizing the brain's free energy, as defined in statistical physics. Furthermore, inference and learning can proceed in a biologically plausible fashion. Cortical responses can be seen as the brain's attempt to minimize the free energy induced by a stimulus and thereby encode the most likely cause of that stimulus. Similarly, learning emerges from changes in synaptic efficacy that minimize the free energy, averaged over all stimuli encountered. The underlying scheme rests on empirical Bayes and hierarchical models of how sensory input is caused. The use of hierarchical models enables the brain to construct prior expectations in a dynamic and context-sensitive fashion. This scheme provides a principled way to understand many aspects of cortical organization and responses. The aim of this article is to encompass many apparently unrelated anatomical, physiological and psychophysical attributes of the brain within a single theoretical perspective. In terms of cortical architectures, the theoretical treatment predicts that sensory cortex should be arranged hierarchically, that connections should be reciprocal and that forward and backward connections should show a functional asymmetry (forward connections are driving, whereas backward connections are both driving and modulatory). In terms of synaptic physiology, it predicts associative plasticity and, for dynamic models, spike-timing-dependent plasticity. In terms of electrophysiology, it accounts for classical and extra classical receptive field effects and long-latency or endogenous components of evoked cortical responses. It predicts the attenuation of responses encoding prediction error with perceptual learning and explains many phenomena such as repetition suppression, mismatch negativity (MMN) and the P300 in electroencephalography. In psychophysical terms, it accounts for the behavioural correlates of these physiological phenomena, for example, priming and global precedence. The final focus of this article is on perceptual learning as measured with the MMN and the implications for empirical studies of coupling among cortical areas using evoked sensory responses.

Biophysical Phenomena↗

Tree-structured gatekeeping tests in clinical trials with hierarchically ordered multiple objectives.

This paper discusses a new class of multiple testing procedures, tree-structured gatekeeping procedures, with clinical trial applications. These procedures arise in clinical trials with hierarchically ordered multiple objectives, for example, in the context of multiple dose-control tests with logical restrictions or analysis of multiple endpoints. The proposed approach is based on the principle of closed testing and generalizes the serial and parallel gatekeeping approaches developed by Westfall and Krishen (J. Statist. Planning Infer. 2001; 99:25-41) and Dmitrienko et al. (Statist. Med. 2003; 22:2387-2400). The proposed testing methodology is illustrated using a clinical trial with multiple endpoints (primary, secondary and tertiary) and multiple objectives (superiority and non-inferiority testing) as well as a dose-finding trial with multiple endpoints.

Antihypertensive Agents↗

Statistical tests in experimental psychiatric research.

It is pointed out that the subjects used in psychiatric research experiments are usually drawn in such a way as to invalidate many of the commonly applied statistical tests. The necessarily non-statistical component of any inference is noted, and the area where one might hope for an exact statistical inference is identified. A class of tests permitting such inferences is described. Their theoretical and practical advantages are outlined.

Humans↗

Protein molecular function prediction by Bayesian phylogenomics.

We present a statistical graphical model to infer specific molecular function for unannotated protein sequences using homology. Based on phylogenomic principles, SIFTER (Statistical Inference of Function Through Evolutionary Relationships) accurately predicts molecular function for members of a protein family given a reconciled phylogeny and available function annotations, even when the data are sparse or noisy. Our method produced specific and consistent molecular function predictions across 100 Pfam families in comparison to the Gene Ontology annotation database, BLAST, GOtcha, and Orthostrapper. We performed a more detailed exploration of functional predictions on the adenosine-5'-monophosphate/adenosine deaminase family and the lactate/malate dehydrogenase family, in the former case comparing the predictions against a gold standard set of published functional characterizations. Given function annotations for 3% of the proteins in the deaminase family, SIFTER achieves 96% accuracy in predicting molecular function for experimentally characterized proteins as reported in the literature. The accuracy of SIFTER on this dataset is a significant improvement over other currently available methods such as BLAST (75%), GeneQuiz (64%), GOtcha (89%), and Orthostrapper (11%). We also experimentally characterized the adenosine deaminase from Plasmodium falciparum, confirming SIFTER's prediction. The results illustrate the predictive power of exploiting a statistical model of function evolution in phylogenomic problems. A software implementation of SIFTER is available from the authors.

Adenosine Deaminase↗

The quota: "an equally serious problem" for us all.

The author explores the question--was a quota applied to the admission of women to the medical school of the University of Toronto? She uses a variety of sources, including class photographs, archives, statistics, personal recollections, and oral history. The existence of a quota from 1944 to 1968 is inferred from statistical patterns and confirmed by a surprising source. In closing, the historiographic implications of this project for research on other politically charged topics are considered.

Canada↗

Interval estimation for a difference between intraclass kappa statistics.

Model-based inference procedures for the kappa statistic have developed rapidly over the last decade. However, no method has yet been developed for constructing a confidence interval about a difference between independent kappa statistics that is valid in samples of small to moderate size. In this article, we propose and evaluate two such methods based on an idea proposed by Newcombe (1998, Statistics in Medicine, 17, 873-890) for constructing a confidence interval for a difference between independent proportions. The methods are shown to provide very satisfactory results in sample sizes as small as 25 subjects per group. Sample size requirements that achieve a prespecified expected width for a confidence interval about a difference of kappa statistic are also presented.

Computer Simulation↗

Relative risk estimation and inference using a generalized logrank statistic.

When comparing two survival distributions with proportional hazard functions, the logrank test is optimal for testing the null hypothesis that the constant hazard ratio (relative risk) is one. In this paper, we focus on (i) testing for departures from a relative risk other than one, and (ii) estimation of the relative risk. The standard tool to address both (i) and (ii) is the Cox proportional hazards model. However, the performance of the Cox model can be less than optimal with small samples. We show why this is the case, and propose a simple alternative method of estimation and inference based on a generalized logrank (GLR) statistic. While the GLR and Cox model approaches are asymptotically similar, empirical results reveal that the GLR approach is notably more efficient than the Cox model when the number of subjects is small (< 100 subjects per treatment group). An example based on survival times of cervical cancer patients is used to illustrate the proposed methodology.

Animals↗

Influences on inferences. Effect of errors in data on statistical evaluation.

BACKGROUND: Inadvertent random and systemic errors introduced into data sets and manipulation of data are well-defined sources of discrepancies in statistical evaluation of clinical trials. In this study, the authors show the influence of errors on the widely used statistical result, P values. METHODS: Using data from a retrospective study of patients with Hodgkin disease treated at the University of Minnesota between 1970 and 1984 and observed to 1988, we introduced various errors into the data to study the impact on results. RESULTS: Inadvertent random and systemic errors affect statistical results. Data entry and transcription errors, vague definitions of endpoints and prognostic factors, and the omission and selection of patients are examples of frequent errors that affect statistical evaluation. CONCLUSION: The results and inferences of many studies are sensitive to systemic errors and data manipulation. Great care must be given to the clear definitions of terms, exclusion and inclusion criteria, group assignments, treatment protocols, and the subgroups on which statistical analysis is performed. Clinicians and statisticians must work together to improve the performance and interpretation of clinical trials.

Clinical Trials as Topic↗

[Models of causal inference: critical analysis of the use of statistics in epidemiology].

The foundations on which the concept of risk has been constructed are discussed. A description of Rubin's model of causal inference, which was first developed in the domain of applied statistics, and later incorporated into a branch of epidemiology, is taken as the starting point. Analysis of the premisses of causal inference brings to light the logical stages in the construction of the concept of risk, allowing it to be understood "from the inside". The abovementioned branch of statistics and epidemiology seeks to demonstrate that statistics can infer causality instead of simply revealing statistical associations; the model gives the basis for estimating that which way be defined as the effect of a cause. Using this procedural distinction between causal inference and association, the model also seeks to differentiate between the epidemiologial dimension of concepts and the merely statistical dimension. This leads to greater complexity when handing the concepts of interation and coofounding. The redective aspects inherent in this methodological construction of risk are here high lighted. Thus, whether applied to individual or populational inferences, this methodological construction imposes limits that need to be taken into account in its theoretical and practical application to epidemiology.

Causality↗

A hierarchical approach to inferences concerning interobserver agreement for multinomial data.

We consider inference methods for interobserver agreement studies characterized by two raters and several outcome categories that one can naturally combine to address a series of questions of a priori interest. We propose a new method based on a series of nested, statistically independent inferences, each corresponding to a binary outcome variable obtained by combining a substantively relevant subset of the original categories. We conduct the inferences using a goodness-of-fit procedure that extends the approach of Donner and Eliasziw. The methodology presented is an alternative to methodology that places each of the outcome categories on an equal footing in estimating interobserver agreement for multinomial data. We provide two examples.

Carcinoma↗

Fundamentals of cDNA microarray data analysis.

Microarray technology is a powerful approach for genomics research. The multi-step, data-intensive nature of this technology has created an unprecedented informatics and analytical challenge. It is important to understand the crucial steps that can affect the outcome of the analysis. In this review, we provide an overview of the contemporary trend on various main analysis steps in the microarray data analysis process, which includes experimental design, data standardization, image acquisition and analysis, normalization, statistical significance inference, exploratory data analysis, class prediction and pathway analysis, as well as various considerations relevant to their implementation.

Animals↗

A statistical model for interpreting computerized dynamic posturography data.

Computerized dynamic posturography (CDP) is widely used for assessment of altered balance control. CDP trials are quantified using the equilibrium score (ES), which ranges from zero to 100, as a decreasing function of peak sway angle. The problem of how best to model and analyze ESs from a controlled study is considered. The ES often exhibits a skewed distribution in repeated trials, which can lead to incorrect inference when applying standard regression or analysis of variance models. Furthermore, CDP trials are terminated when a patient loses balance. In these situations, the ES is not observable, but is assigned the lowest possible score--zero. As a result, the response variable has a mixed discrete-continuous distribution, further compromising inference obtained by standard statistical methods. Here, we develop alternative methodology for analyzing ESs under a stochastic model extending the ES to a continuous latent random variable that always exists, but is unobserved in the event of a fall. Loss of balance occurs conditionally, with probability depending on the realized latent ES. After fitting the model by a form of quasi-maximum-likelihood, one may perform statistical inference to assess the effects of explanatory variables. An example is provided, using data from the NIH/NIA Baltimore Longitudinal Study on Aging.

Adult↗

Likelihood analysis for the ratio of means of two independent log-normal distributions.

Existing methods for comparing the means of two independent skewed log-normal distributions do not perform well in a range of small-sample settings such as a small-sample bioavailability study. In this article, we propose two likelihood-based approaches-the signed log-likelihood ratio statistic and modified signed log-likelihood ratio statistic-for inference about the ratio of means of two independent log-normal distributions. More specifically, we focus on obtaining p-values for testing the equality of means and also constructing confidence intervals for the ratio of means. The performance of the proposed methods is assessed through simulation studies that show that the modified signed log-likelihood ratio statistic is nearly an exact approach even for very small samples. The methods are also applied to two real-life examples.

Biological Availability↗

Analysis of repeated pregnancy outcomes.

Women tend to repeat reproductive outcomes, with past history of an adverse outcome being associated with an approximate two-fold increase in subsequent risk. These observations support the need for statistical designs and analyses that address this clustering. Failure to do so may mask effects, result in inaccurate variance estimators, produce biased or inefficient estimates of exposure effects. We review and evaluate basic analytic approaches for analysing reproductive outcomes, including ignoring reproductive history, treating it as a covariate or avoiding the clustering problem by analysing only one pregnancy per woman, and contrast these to more modern approaches such as generalized estimating equations with robust standard errors and mixed models with various correlation structures. We illustrate the issues by analysing a sample from the Collaborative Perinatal Project dataset, demonstrating how the statistical model impacts summary statistics and inferences when assessing etiologic determinants of birth weight.

Adolescent↗

Uncertainty modeling and model selection for geometric inference.

We first investigate the meaning of "statistical methods" for geometric inference based on image feature points. Tracing back the origin of feature uncertainty to image processing operations, we discuss the implications of asymptotic analysis in reference to "geometric fitting" and "geometric model selection" and point out that a correspondence exists between the standard statistical analysis and the geometric inference problem. Then, we derive the "geometric AIC" and the "geometric MDL" as counterparts of Akaike's AIC and Rissanen's MDL. We show by experiments that the two criteria have contrasting characteristics in detecting degeneracy.

Algorithms↗

Novel multilocus measure of linkage disequilibrium to estimate past effective population size.

Linkage disequilibrium (LD) between densely spaced, polymorphic genetic markers in humans and other species contains information about historical population size. Inferring past population size is of interest both from an evolutionary perspective (e.g., testing the "out of Africa" hypothesis of human evolution) and to improve models for mapping of disease and quantitative trait genes. We propose a novel multilocus measure of LD, the chromosome segment homozygosity (CSH). CSH is defined for a specific chromosome segment, up to the full length of the chromosome. In computer simulations CSH was generally less variable than the r(2) measure of LD, and variability of CSH decreased as the number of markers in the chromosome segment was increased. The essence and utility of our novel measure is that CSH over long distances reflects recent effective population size (N), whereas CSH over small distances reflects the effective size in the more distant past. We illustrate the utility of CSH by calculating CSH from human and dairy cattle SNP and microsatellite marker data, and predicting N at various times in the past for each species. Results indicated an exponentially increasing N in humans and a declining N in dairy cattle. CSH is a valuable statistic for inferring population histories from haplotype data, and has implications for mapping of disease loci.

Animals↗

Use of multiple markers in population-based molecular epidemiologic studies of tuberculosis.

SETTING: Many epidemiologic studies of tuberculosis are being conducted worldwide. Fingerprinting with a secondary marker in strains with fewer than six IS6110-hybridizing bands enhances the tracking of strains, but its impact on population-level inferences has not been well studied. OBJECTIVE: To investigate the effects of secondary genotyping for low-copy Mycobacterium tuberculosis isolates with polymorphic guanine-cytosine-rich repetitive sequence (PGRS) on epidemiologic inferences in population-based research settings. DESIGN: For San Francisco tuberculosis cases (1991-1996), clusters were defined by IS6110 alone and by PGRS/IS6110 to 1) estimate recent transmission, 2) evaluate the theoretical influence of bacterial population parameters on these estimates, and 3) assess risk factors for recent transmission. RESULTS: Secondary typing on low-copy strains (20.3% of all isolates) decreased the estimate of recent transmission from 29.1% to 25.3% (P = 0.03). The most influential parameters in determining whether supplemental genotyping results in different estimates were the proportion of low-copy strains and the amount of clustering. Risk factors for recent transmission were identical for both definitions of clustering. CONCLUSION: The statistical and inferred effects of secondary genotyping of M. tuberculosis seem to depend on the proportion of low-copy strains in the population. When this proportion is low or when few secondary patterns match, supplemental genotyping may yield minimal insight into population-level investigations.

Adult↗