Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

The age of the common ancestor of eukaryotes and prokaryotes: statistical inferences.

In this paper, a simple distance measure was used to estimate the age (T) of the common ancestor of eukaryotes and prokaryotes which takes the rate variation among sites and the pattern of amino acid substitutions into account. Our new estimate of T based on Doolittle et al.'s data is about 2.5 billion years ago (Ga), with 95% confidence interval from 2.1 to 2.9 Ga. This result indicates (1) that Doolittle et al.'s estimate (approximately 2.0 Ga) seems too recent, and (2) that the traditional view about the divergence time between eukaryotes and prokaryotes (T0 = 3.5 Ga) can be rejected at the 0.1% significance level.

Amino Acid Sequence↗

Statistical inference and model selection for the 1861 Hagelloch measles epidemic.

A stochastic epidemic model is proposed which incorporates heterogeneity in the spread of a disease through a population. In particular, three factors are considered: the spatial location of an individual's home and the household and school class to which the individual belongs. The model is applied to an extremely informative measles data set and the model is compared with nested models, which incorporate some, but not all, of the aforementioned factors. A reversible jump Markov chain Monte Carlo algorithm is then introduced which assists in selecting the most appropriate model to fit the data.

Adolescent↗

Statistical inference for the area under the receiver operating characteristic curve in the presence of random measurement error.

The area under the receiver operating characteristic curve is the most commonly used measure of the ability of a biomarker to distinguish between two populations. Some markers are subject to substantial measurement error. Under normality assumptions, the authors develop a confidence interval procedure for the area under the receiver operating characteristic curve that adjusts for measurement error. This procedure assumes the availability of data from a reliability study of the biomarker. A simulation study was used to check the validity of the proposed confidence interval. Furthermore, it was shown that not adjusting for measurement error could result in a serious understatement of the effectiveness of the biomarker.

Analysis of Variance↗

A comparison of two methods for making statistical inferences on Nei's measure of genetic distance.

The delta and jackknife methods can be used to estimate Nei's measure of genetic distance and calculate confidence intervals for this estimate. Computer stimulations were used to study the bias and variance of each estimator and the accuracy of the corresponding approximate 95% confidence intervals. The simulations were conducted using 3 sets of data and several sample sizes. The results showed: (1) the jackknife reduced bias; (2) in 8 out of 9 cases the variance and mean square error of the jackknife estimator were less; (3) a second order jackknife reduced the bias the most but suffered a corresponding increase in variance; (4) both the first order jackknife and delta methods yielded intervals whose confidence levels were approximately equal but less than 95%.

Alleles↗

On differential variability of expression ratios: improving statistical inference about gene expression changes from microarray data.

We consider the problem of inferring fold changes in gene expression from cDNA microarray data. Standard procedures focus on the ratio of measured fluorescent intensities at each spot on the microarray, but to do so is to ignore the fact that the variation of such ratios is not constant. Estimates of gene expression changes are derived within a simple hierarchical model that accounts for measurement error and fluctuations in absolute gene expression levels. Significant gene expression changes are identified by deriving the posterior odds of change within a similar model. The methods are tested via simulation and are applied to a panel of Escherichia coli microarrays.

Algorithms↗

Beyond statistical inference: a decision theory for science.

Traditional null hypothesis significance testing does not yield the probability of the null or its alternative and, therefore, cannot logically ground scientific decisions. The decision theory proposed here calculates the expected utility of an effect on the basis of (1) the probability of replicating it and (2) a utility function on its size. It takes significance tests--which place all value on the replicability of an effect and none on its magnitude--as a special case, one in which the cost of a false positive is revealed to be an order of magnitude greater than the value of a true positive. More realistic utility functions credit both replicability and effect size, integrating them for a single index of merit. The analysis incorporates opportunity cost and is consistent with alternate measures of effect size, such as r2 and information transmission, and with Bayesian model selection criteria. An alternate formulation is functionally equivalent to the formal theory, transparent, and easy to compute.

Data Interpretation, Statistical↗

New explicit expressions for relative frequencies of single-nucleotide polymorphisms with application to statistical inference on population growth.

We present new methodology for calculating sampling distributions of single-nucleotide polymorphism (SNP) frequencies in populations with time-varying size. Our approach is based on deriving analytical expressions for frequencies of SNPs. Analytical expressions allow for computations that are faster and more accurate than Monte Carlo simulations. In contrast to other articles showing analytical formulas for frequencies of SNPs, we derive expressions that contain coefficients that do not explode when the genealogy size increases. We also provide analytical formulas to describe the way in which the ascertainment procedure modifies SNP distributions. Using our methods, we study the power to test the hypothesis of exponential population expansion vs. the hypothesis of evolution with constant population size. We also analyze some of the available SNP data and we compare our results of demographic parameters estimation to those obtained in previous studies in population genetics. The analyzed data seem consistent with the hypothesis of past population growth of modern humans. The analysis of the data also shows a very strong sensitivity of estimated demographic parameters to changes of the model of the ascertainment procedure.

Data Interpretation, Statistical↗

Statistical inference for correlated data in ophthalmologic studies.

In ophthalmologic studies, each subject usually contributes important information for each of two eyes and the values from the two eyes are generally highly correlated. Previous studies showed that test procedures for binary paired data that ignore the presence of intraclass correlation could lead to inflated significance levels. Furthermore, it is possible that asymptotic versions of these procedures that take the intraclass correlation into account could also produce unacceptably high type I error rates when the sample size is small or the data structure is sparse. We propose two alternatives for these situations, namely the exact unconditional and approximate unconditional procedures. According to our simulation results, the exact procedures usually produce extremely conservative empirical type I error rates. That is, the corresponding type I error rates could greatly underestimate the pre-assigned nominal level (e.g. (empirical type I error rate/nominal type I error rate) 0.8). On the other hand, the approximate unconditional procedures usually yield empirical type I error rates close to the pre-chosen nominal level. We illustrate our methodologies with a data set from a retinal detachment study.

Biometry↗

Comparison of three methods for generating group statistical inferences from independent component analysis of functional magnetic resonance imaging data.

PURPOSE: To evaluate the relative effectiveness of three previously proposed methods of performing group independent component analysis (ICA) of functional magnetic resonance imaging (fMRI) data. MATERIALS AND METHODS: Data were generated via computer simulation. Components were added to a varying number of subjects between 1 and 20, and intersubject variability was simulated for both the added sources and their associated time courses. Three methods of group ICA analyses were performed: across-subject averaging, subject-wise concatenation, and row-wise concatenation (e.g., across time courses). RESULTS: Concatenating across subjects provided the best overall performance in terms of accurate estimation of the sources and associated time courses. Averaging across subjects provided accurate estimation (R > 0.9) of the time courses when the sources were present in a sufficient fraction (about 15%) of 100 subjects. Concatenating across time courses was shown not to be a feasible method when unique sources were added to the data from each subject, simulating the effects of motion and susceptibility artifacts. CONCLUSION: Subject-wise concatenation should be used when computationally feasible. For studies involving a large number of subjects, across-subject averaging provides an acceptable alternative and reduces the computational load.

Computer Simulation↗

Computational anatomy and neuropsychiatric disease: probabilistic assessment of variation and statistical inference of group difference, hemispheric asymmetry, and time-dependent change.

Three components of computational anatomy (CA) are reviewed in this paper: (i) the computation of large-deformation maps, that is, for any given coordinate system representations of two anatomies, computing the diffeomorphic transformation from one to the other; (ii) the computation of empirical probability laws of anatomical variation between anatomies; and (iii) the construction of inferences regarding neuropsychiatric disease states. CA utilizes spatial-temporal vector field information obtained from large-deformation maps to assess anatomical variabilities and facilitate the detection and quantification of abnormalities of brain structure in subjects with neuropsychiatric disorders. Neuroanatomical structures are divided into two types: subcortical structures-gray matter (GM) volumes enclosed by a single surface-and cortical mantle structures-anatomically distinct portions of the cerebral cortical mantle layered between the white matter (WM) and cerebrospinal fluid (CSF). Because of fundamental differences in the geometry of these two types of structures, image-based large-deformation high-dimensional brain mapping (HDBM-LD) and large-deformation diffeomorphic metric matching (LDDMM) were developed for the study of subcortical structures and labeled cortical mantle distance mapping (LCMDM) was developed for the study of cortical mantle structures. Studies of neuropsychiatric disorders using CA usually require the testing of hypothesized group differences with relatively small numbers of subjects per group. Approaches that increase the power for testing such hypotheses include methods to quantify the shapes of individual structures, relationships between the shapes of related structures (e.g., asymmetry), and changes of shapes over time. Promising preliminary studies employing these approaches to studies of subjects with schizophrenia and very mild to mild Alzheimer's disease (AD) are presented.

Algorithms↗

Statistical inference for a linear function of medians: confidence intervals, hypothesis testing, and sample size requirements.

When the distribution of the response variable is skewed, the population median may be a more meaningful measure of centrality than the population mean, and when the population distribution of the response variable has heavy tails, the sample median may be a more efficient estimator of centrality than the sample mean. The authors propose a confidence interval for a general linear function of population medians. Linear functions have many important special cases including pairwise comparisons, main effects, interaction effects, simple main effects, curvature, and slope. The confidence interval can be used to test 2-sided directional hypotheses and finite interval hypotheses. Sample size formulas are given for both interval estimation and hypothesis testing problems.

Humans↗

Statistical inference in a two-compartment model for hematopoiesis.

We present a method for parameter estimation in a two-compartment hidden Markov model of the first two stages of hematopoiesis. Hematopoiesis is the specialization of stem cells into mature blood cells. As stem cells are not distinguishable in bone marrow, little is known about their behavior, although it is known that they have the ability to self-renew or to differentiate to more specialized (progenitor) cells. We observe progenitor cells in samples of bone marrow taken from hybrid cats whose cells contain a natural binary marker. With data consisting of the changing proportions of this binary marker over time from several cats, estimates for stem cell self-renewal and differentiation parameters are obtained using an estimating equations approach.

Animals↗

Statistical inference for familial disease clusters.

In many epidemiologic studies, the first indication of an environmental or genetic contribution to the disease is the way in which the diseased cases cluster within the same family units. The concept of clustering is contrasted with incidence. We assume that all individuals are exchangeable except for their disease status. This assumption is used to provide an exact test of the initial hypothesis of no familial link with the disease, conditional on the number of diseased cases and the distribution of the sizes of the various family units. New parametric generalizations of binomial sampling models are described to provide measures of the effect size of the disease clustering. We consider models and an example that takes covariates into account. Ascertainment bias is described and the appropriate sampling distribution is demonstrated. Four numerical examples with real data illustrate these methods.

Adult↗