Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Galton's legacy to research on intelligence.

In the 1999 Galton Lecture for the annual conference of The Galton Institute, the author summarizes the main elements of Galton's ideas about human mental ability and the research paradigm they generated, including the concept of 'general' mental ability, its hereditary component, its physical basis, racial differences, and methods for measuring individual differences in general ability. Although the conclusions Galton drew from his empirical studies were seldom compelling for lack of the needed technology and methods of statistical inference in his day, contemporary research has generally borne out most of Galton's original and largely intuitive ideas, which still inspire mainstream scientific research on intelligence.

Eugenics↗

On the choice of computational unit in statistical analysis.

It is stated that in trials where experimental units on different levels (sites, patients, etc.) are employed, the highest level unit should be used as computational unit when computing standard errors and in statistical inference. Using a lower level unit will underestimate the standard error and the level of significance (P-value). A numerical illustration is presented.

Dental Scaling↗

Randomized single-subject experimental designs.

Books on single-subject methodology tend to focus on traditional operant research techniques and thus provide little or no discussion of random introduction of treatments and statistical tests based on such randomization, i.e. randomization tests. Those books are the principal references to which researchers must turn for a comprehensive coverage of single-subject methodology, and so many researchers are likely to be unaware of the relevance of randomization (random assignment of treatment times to treatments) and randomization tests to single-subject experimentation. That is unfortunate because randomization is necessary in order to draw valid statistical inferences about treatment effects. The role of randomization in providing control over major threats to internal validity is explained in this article, and a number of randomized single-subject designs and their applications are provided. Appropriate rank tests are specified, and sources of free software for other, more complex, statistical tests are given.

Child, Preschool↗

Bayesian inference with probabilistic population codes.

Recent psychophysical experiments indicate that humans perform near-optimal Bayesian inference in a wide variety of tasks, ranging from cue integration to decision making to motor control. This implies that neurons both represent probability distributions and combine those distributions according to a close approximation to Bayes' rule. At first sight, it would seem that the high variability in the responses of cortical neurons would make it difficult to implement such optimal statistical inference in cortical circuits. We argue that, in fact, this variability implies that populations of neurons automatically represent probability distributions over the stimulus, a type of code we call probabilistic population codes. Moreover, we demonstrate that the Poisson-like variability observed in cortex reduces a broad class of Bayesian inference to simple linear combinations of populations of neural activity. These results hold for arbitrary probability distributions over the stimulus, for tuning curves of arbitrary shape and for realistic neuronal variability.

Algorithms↗

Inferences from alarming events.

An extreme event, such as a nuclear accident, an earthquake, a cluster of adverse reactions to a particular drug, or excessive breakdowns of some class of equipment, frequently focuses attention for the first time on an important issue. By then, however, data on the incidence and magnitudes of relevant past events may be unavailable or too costly to reconstruct. Using a simple probability model, we derive methods for drawing statistical inferences based only on the magnitude of the first event noticed and the amount of exposure before this event occurred. We assume that an event is noticed only when its magnitude exceeds some threshold, and we develop methods of inference that are valid even when this threshold is unknown. One tempting but incorrect approach is to treat the magnitude of the observed event as if it were the threshold, forgetting that smaller magnitudes might have been noticed as well. The biases that arise when this mistake is made turn out to be substantial; risks can easily be overstated by a factor of 3.

Disasters↗

Generalized linear mixture models for handling nonignorable dropouts in longitudinal studies.

This paper presents a method for analysing longitudinal data when there are dropouts. In particular, we develop a simple method based on generalized linear mixture models for handling nonignorable dropouts for a variety of discrete and continuous outcomes. Statistical inference for the model parameters is based on a generalized estimating equations (GEE) approach (Liang and Zeger, 1986). The proposed method yields estimates of the model parameters that are valid when nonresponse is nonignorable under a variety of assumptions concerning the dropout process. Furthermore, the proposed method can be implemented using widely available statistical software. Finally, an example using data from a clinical trial of contracepting women is used to illustrate the methodology.

Journal Article↗

Genetic variance components analysis for binary phenotypes using generalized linear mixed models (GLMMs) and Gibbs sampling.

The common complex diseases such as asthma are an important focus of genetic research, and studies based on large numbers of simple pedigrees ascertained from population-based sampling frames are becoming commonplace. Many of the genetic and environmental factors causing these diseases are unknown and there is often a strong residual covariance between relatives even after all known determinants are taken into account. This must be modelled correctly whether scientific interest is focused on fixed effects, as in an association analysis, or on the covariances themselves. Analysis is straightforward for multivariate Normal phenotypes, but difficulties arise with other types of trait. Generalized linear mixed models (GLMMs) offer a potentially unifying approach to analysis for many classes of phenotype including multivariate Normal traits, binary traits, and censored survival times. Markov Chain Monte Carlo methods, including Gibbs sampling, provide a convenient framework within which such models may be fitted. In this paper, Bayesian inference Using Gibbs Sampling (a generic Gibbs sampler; BUGS) is used to fit GLMMs for multivariate Normal and binary phenotypes in nuclear families. BUGS is easy to use and readily available. We motivate a suitable model structure for Normal phenotypes and show how the model extends to binary traits. We discuss parameter interpretation and statistical inference and show how to circumvent a number of important theoretical and practical problems that we encountered. Using simulated data we show that model parameters seem consistent and appear unbiased in smaller data sets. We illustrate our methods using data from an ongoing cohort study.

Binomial Distribution↗

Introduction to biostatistics: Part 1, Basic concepts.

Statistical methods commonly used to analyze data presented in journal articles should be understood by both medical scientists and practicing clinicians. Inappropriate data analysis methods have been reported in 42% to 78% of original publications in critical reviews of selected medical journals. The only way to halt researchers' misuse of statistics and improve the clinician's knowledge of statistics is through education. This is the first of a six-part series of articles intended to provide the reader with a basic, yet fundamental knowledge of common biomedical statistical methods. The series will cover basic concepts of statistical analysis, descriptive statistics, statistical inference theory, comparison of means, chi 2, and correlational and regression techniques. A conceptual explanation will accompany discussion of the appropriate use of these techniques.

Biometry↗

Volumetric three-dimensional recognition of biological microorganisms using multivariate statistical method and digital holography.

We present a new statistical approach to real-time sensing and recognition of microorganisms using digital holographic microscopy. We numerically produce many section images at different depths along a longitudinal direction from the single digital hologram of three-dimensional (3D) microorganisms in the Fresnel domain. For volumetric 3D recognition, the test pixel points are randomly selected from the section image; this procedure can be repeated with different specimens of the same microorganism. The multivariate joint density functions are calculated from the pixel values of each section image at the same random pixel points. The parameters of the statistical distributions are compared using maximum likelihood estimation and statistical inference algorithms. The performance of the proposed system is illustrated with preliminary experimental results.

Artificial Intelligence↗

A statistical method for the analysis of positron emission tomography neuroreceptor ligand data.

A method for voxel by voxel statistical inference of PET radioligand receptor studies is presented. This method is aimed at detecting differences in radioligand binding between baseline and activation scans. It uses nonlinear least squares theory to estimate the ligand-receptor model parameters and utilizes the residuals to calculate their associated variance. The approach both increases the degrees of freedom for statistical testing and produces more accurate estimates of the standard deviation of the parameters. This technique is applicable to any ligand with a validated compartmental model, whether reversibly or irreversibly bound. The method was investigated and compared with a simple voxel-wise t test. Both simulated and real PET data for the dopamine D(1) receptor ligand [(11)C]SCH 23390 were used to assess the method. The assumptions implicit in the residuals methods were validated. The residuals method was found to be more sensitive than a simple t test, while not producing false-positive results. In addition, we showed that this method reliably differentiates changes in radioligand binding from the effects of changes in cerebral blood flow.

Adult↗

Regression analysis with missing covariate data using estimating equations.

In regression analysis, missing covariate data has been among the most common problems. Frequently, practitioners adopt the so-called complete-case analysis, i.e., performing the analysis on only a complete dataset after excluding records with missing covariates. Performing a complete-case analysis is convenient with existing statistical packages, but it may be inefficient since the observed outcomes and covariates on those records with missing covariates are not used. It can even give misleading statistical inference if missing is not completely at random. This paper introduces a joint estimating equation (JEE) for regression analysis in the presence of missing observations on one covariate, which may be thought of as a method in a general framework for the missing covariate data problem proposed by Robins, Rotnitzky, and Zhao (1994, Journal of the American Statistical Association 89, 846-866). A generalization of JEE to more than one such covariate is discussed. The JEE is generally applicable to estimating regression coefficients from a regression model, including linear and logistic regression. Provided that the missing covariate data is either missing completely at random or missing at random (in addition to mild regularity conditions), estimates of regression coefficients from the JEE are consistent and have an asymptotic normal distribution. Simulation results show that the asymptotic distribution of estimated coefficients performs well in finite samples. Also shown through the simulation study is that the validity of JEE estimates depends on the correct specification of the probability function that characterizes the missing mechanism, suggesting a need for further research on how to robustify the estimation from making this nuisance assumption. Finally, the JEE is illustrated with an application from a case-control study of diet and thyroid cancer.

Biometry↗

Detection of convergent and parallel evolution at the amino acid sequence level.

Adaptive evolution at the molecular level can be studied by detecting convergent and parallel evolution at the amino acid sequence level. For a set of homologous protein sequences, the ancestral amino acids at all interior nodes of the phylogenetic tree of the proteins can be statistically inferred. The amino acid sites that have experienced convergent or parallel changes on independent evolutionary lineages can then be identified by comparing the amino acids at the beginning and end of each lineage. At present, the efficiency of the methods of ancestral sequence inference in identifying convergent and parallel changes is unknown. More seriously, when we identify convergent or parallel changes, it is unclear whether these changes are attributable to random chance. For these reasons, claims of convergent and parallel evolution at the amino acid sequence level have been disputed. We have conducted computer simulations to assess the efficiencies, of the parsimony and Bayesian methods of ancestral sequence inference in identifying convergent and parallel-change sites. Our results showed that the Bayesian method performs better than the parsimony method in identifying parallel changes, and both methods are inefficient in identifying convergent changes. However, the Bayesian method is recommended for estimating the number of convergent-change sites because it gives a conservative estimate. We have developed statistical tests for examining whether the observed numbers of convergent and parallel changes are due to random chance. As an example, we reanalyzed the stomach lysozyme sequences of foregut fermenters and found that parallel evolution is statistically significant, whereas convergent evolution is not well supported.

Amino Acid Sequence↗

Approximating the coalescent with recombination.

The coalescent with recombination describes the distribution of genealogical histories and resulting patterns of genetic variation in samples of DNA sequences from natural populations. However, using the model as the basis for inference is currently severely restricted by the computational challenge of estimating the likelihood. We discuss why the coalescent with recombination is so challenging to work with and explore whether simpler models, under which inference is more tractable, may prove useful for genealogy-based inference. We introduce a simplification of the coalescent process in which coalescence between lineages with no overlapping ancestral material is banned. The resulting process has a simple Markovian structure when generating genealogies sequentially along a sequence, yet has very similar properties to the full model, both in terms of describing patterns of genetic variation and as the basis for statistical inference.

Chromosomes↗

Inference of population history using a likelihood approach.

We introduce an approach to revealing the likelihood of different population histories that utilizes an explicit model of sequence evolution for the DNA segment under study. Based on a phylogenetic tree reconstruction method we show that a Tamura-Nei model with heterogeneous mutation rates is a fair description of the evolutionary process of the hypervariable region I of the mitochondrial DNA from humans. Assuming this complex model still allows the estimation of population history parameters, we suggest a likelihood approach to conducting statistical inference within a class of expansion models. More precisely, the likelihood of the data is based on the mean pairwise differences between DNA sequences and the number of variable sites in a sample. The use of likelihood ratios enables comparison of different hypotheses about population history, such as constant population size during the past or an increase or decrease of population size starting at some point back in time. This method was applied to show that the population of the Basques has expanded, whereas that of the Biaka pygmies is most likely decreasing. The Nuu-Chah-Nulth data are consistent with a model of constant population.

Base Composition↗

Extracting a maximum of useful information from statistical research data.

The practice of statistical inference in psychological research is critically reviewed. Particular emphasis is put on the fast pace of change from the sole reliance on null hypothesis significance testing (NHST) to the inclusion of effect size estimates, confidence intervals, and an interest in the Bayesian approach. We conclude that these developments are helpful for psychologists seeking to extract a maximum of useful information from statistical research data, and that seven decades of criticism against NHST is finally having an effect.

Bayes Theorem↗

Association analysis of polymorphisms in serotonin 1B receptor (HTR1B) gene with heroin addiction: a comparison of molecular and statistically estimated haplotypes.

OBJECTIVES: 5-Hydroxytryptamine (serotonin)-1B receptors (HTR1B) may play an important role in psychiatric disorders and drug and alcohol dependence. In this study we report on genotype, molecular haplotype and statistically estimated haplotype analyses of previously identified polymorphisms in positions -261T>G, -161A>T, 129C>T, 861G>C and 1180A>G of the HTR1B gene in ethnically diverse populations (African-Americans, Caucasians, Hispanics and Asians) including 235 former heroin addicts and 161 control subjects from New York City. The objectives were to test for an association of molecular and statistically estimated haplotypes and genotypes in HTR1B gene with heroin addiction and to compare results provided by molecular and statistically estimated haplotyping methods. METHODS: Genotype analysis was performed using a standard TaqMan protocol. Molecular haplotype analysis of the subset of polymorphisms consisting of -261T>G, -161A>T and 129C>T was performed using a protocol specially designed by our group, using fluorescent PCR. This is based on use of allele-specific primers complementary to flanking polymorphisms and a fluorescently labeled sequence-specific TaqMan probe set complementary to an internal polymorphism of the haplotype region. Every individual's statistically inferred haplotype pair agreed with the individual's haplotype pair determined by molecular haplotyping. RESULTS AND CONCLUSION: A point-wise significant association of haplotype pairs containing allele G at position 1180 with protective effect from heroin addiction in Caucasians was found. A point-wise nominally significant association of allele 1180G with a protective effect from heroin addiction was found in Caucasians. Statistically significant differences across four ethnic groups in control subjects for allelic frequencies of -261T>G and -161A>T were found.

Black or African American↗

Martingale methods for the analysis of epidemic data.

After explaining why martingale methods play an important role in statistical inference for parameters of epidemic models, we give a tutorial introduction to these methods in the more familiar context of data on independent and identically distributed survival times. In this simpler setting we introduce requisite results from martingale theory and demonstrate that martingale methods simply lead to well known estimates and their standard errors. We then turn to the context of epidemics, and illustrate how martingale methods can be used to derive a method of inference for the infection potential in a simple model with removal of infectives. The resulting method involves only simple computations and we demonstrate that the method applies under much more general assumptions. There follows a critical review of several applications of martingale methods for the analysis of infectious disease data. It emerges that the approach provides simple methods of inference in some situations where standard methods of inference are not available, or are too cumbersome. The range of applications seems limited, but new applications continue to be found. Little has been done to confirm high efficiency of martingale methods in epidemic applications.

Communicable Diseases↗

Inference for smooth curves in longitudinal data with application to an AIDS clinical trial.

We discuss a longitudinal study where data for many subjects are collected at irregular intervals. The study is a randomized trial of HIV infected subjects and the response variable of interest is serum neopterin. The mean of the outcome variable, taken over patients in each treatment group, is assumed to follow a smooth curve. Piecewise cubic polynomials with a moderate number of knots are used to model the curves. A general parametric form is assumed for the covariance structure. Maximum penalized likelihood estimation is used to smooth the over-parameterized curves. Statistical inference for the mean curves, including confidence bands and hypothesis tests, is discussed. Two approaches, one using a Bayesian interpretation of the penalized likelihood and the other based on the asymptotic distribution of the maximum penalized likelihood estimates, are discussed and contrasted. The properties of the confidence bands obtained from these two approaches are evaluated by examining their coverage rates in a simulation study.

Acquired Immunodeficiency Syndrome↗