Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Case-control studies of association in structured or admixed populations.

Case-control tests for association are an important tool for mapping complex-trait genes. But population structure can invalidate this approach, leading to apparent associations at markers that are unlinked to disease loci. Family-based tests of association can avoid this problem, but such studies are often more expensive and in some cases--particularly for late-onset diseases--are impractical. In this review article we describe a series of approaches published over the past 2 years which use multilocus genotype data to enable valid case-control tests of association, even in the presence of population structure. These tests can be classified into two categories. "Genomic control" methods use the independent marker loci to adjust the distribution of a standard test statistic, while "structured association" methods infer the details of population structure en route to testing for association. We discuss the statistical issues involved in the different approaches and present results from simulations comparing the relative performance of the methods under a range of models.

Alleles↗

A free energy principle for the brain.

By formulating Helmholtz's ideas about perception, in terms of modern-day theories, one arrives at a model of perceptual inference and learning that can explain a remarkable range of neurobiological facts: using constructs from statistical physics, the problems of inferring the causes of sensory input and learning the causal structure of their generation can be resolved using exactly the same principles. Furthermore, inference and learning can proceed in a biologically plausible fashion. The ensuing scheme rests on Empirical Bayes and hierarchical models of how sensory input is caused. The use of hierarchical models enables the brain to construct prior expectations in a dynamic and context-sensitive fashion. This scheme provides a principled way to understand many aspects of cortical organisation and responses. In this paper, we show these perceptual processes are just one aspect of emergent behaviours of systems that conform to a free energy principle. The free energy considered here measures the difference between the probability distribution of environmental quantities that act on the system and an arbitrary distribution encoded by its configuration. The system can minimise free energy by changing its configuration to affect the way it samples the environment or change the distribution it encodes. These changes correspond to action and perception respectively and lead to an adaptive exchange with the environment that is characteristic of biological systems. This treatment assumes that the system's state and structure encode an implicit and probabilistic model of the environment. We will look at the models entailed by the brain and how minimisation of its free energy can explain its dynamics and structure.

Afferent Pathways↗

A new statistical method for haplotype reconstruction from population data.

Current routine genotyping methods typically do not provide haplotype information, which is essential for many analyses of fine-scale molecular-genetics data. Haplotypes can be obtained, at considerable cost, experimentally or (partially) through genotyping of additional family members. Alternatively, a statistical method can be used to infer phase and to reconstruct haplotypes. We present a new statistical method, applicable to genotype data at linked loci from a population sample, that improves substantially on current algorithms; often, error rates are reduced by > 50%, relative to its nearest competitor. Furthermore, our algorithm performs well in absolute terms, suggesting that reconstructing haplotypes experimentally or by genotyping additional family members may be an inefficient use of resources.

Algorithms↗

Analysis of cost data in randomized trials: an application of the non-parametric bootstrap.

Health economic evaluations are now more commonly being included in pragmatic randomized trials. However a variety of methods are being used for the presentation and analysis of the resulting cost data, and in many cases the approaches taken are inappropriate. In order to inform health care policy decisions, analysis needs to focus on arithmetic mean costs, since these will reflect the total cost of treating all patients with the disease. Thus, despite the often highly skewed distribution of cost data, standard non-parametric methods or use of normalizing transformations are not appropriate. Although standard parametric methods of comparing arithmetic means may be robust to non-normality for some data sets, this is not guaranteed. While the randomization test can be used to overcome assumptions of normality, its use for comparing means is still restricted by the need for similarly shaped distributions in the two groups. In this paper we show how the non-parametric bootstrap provides a more flexible alternative for comparing arithmetic mean costs between randomized groups, avoiding the assumptions which limit other methods. Details of several bootstrap methods for hypothesis tests and confidence intervals are described and applied to cost data from two randomized trials. The preferred bootstrap approaches are the bootstrap-t or variance stabilized bootstrap-t and the bias corrected and accelerated percentile methods. We conclude that such bootstrap techniques can be recommended either as a check on the robustness of standard parametric methods, or to provide the primary statistical analysis when making inferences about arithmetic means for moderately sized samples of highly skewed data such as costs.

Cognitive Behavioral Therapy↗

Technical note: computing tests of fixed effects in a restricted class of mixed models.

Inferences about fixed effects in mixed linear models are important in a variety of animal science studies. The statistical theory for making such inferences is well known, and if the variance components are known up to a proportionality constant, then optimal exact tests can be performed. Computing the test statistics, however, can still be problematic when the random effects have many levels. In practice, approximate tests that are easily computed but less efficient are usually employed. This article describes reduction in error sum of squares procedures for performing the exact test and for computing associated confidence intervals. By taking advantage of iterative algorithms for solving Henderson's mixed-model equations, the tests can be performed without inverting the covariance matrix or computing a generalized inverse of the mixed-model coefficient matrix. The procedures are illustrated on an animal model that has three random effects, two with 1,372 levels and one with 450 levels.

Algorithms↗

Tests for genetic association using family data.

We use likelihood-based score statistics to test for association between a disease and a diallelic polymorphism, based on data from arbitrary types of nuclear families. The Nonfounder statistic extends the transmission disequilibrium test (TDT) to accommodate affected and unaffected offspring, missing parental genotypes, phenotypes more general than qualitative traits, such as censored survival data and quantitative traits, and residual correlation of phenotypes within families. The Founder statistic compares observed or inferred parental genotypes to those expected in the general population. Here the genotypes of affected parents and those with many affected offspring are weighted more heavily than unaffected parents and those with few affected offspring. We illustrate the tests by applying them to data on a polymorphism of the SRD5A2 gene in nuclear families with multiple cases of prostate cancer. We also use simulations to compare the power of these family-based statistics to that of the score statistic based on Cox's partial likelihood for censored survival data, and find that the family-based statistics have considerably more power when there are many untyped parents. The software program FGAP for computing test statistics is available at http://www.stanford.edu/dept/HRP/epidemiology/FGAP.

Alleles↗

Trial-to-trial variability of cortical evoked responses: implications for the analysis of functional connectivity.

OBJECTIVES: The time series of single trial cortical evoked potentials typically have a random appearance, and their trial-to-trial variability is commonly explained by a model in which random ongoing background noise activity is linearly combined with a stereotyped evoked response. In this paper, we demonstrate that more realistic models, incorporating amplitude and latency variability of the evoked response itself, can explain statistical properties of cortical potentials that have often been attributed to stimulus-related changes in functional connectivity or other intrinsic neural parameters. METHODS: Implications of trial-to-trial evoked potential variability for variance, power spectrum, and interdependence measures like cross-correlation and spectral coherence, are first derived analytically. These implications are then illustrated using model simulations and verified experimentally by the analysis of intracortical local field potentials recorded from monkeys performing a visual pattern discrimination task. To further investigate the effects of trial-to-trial variability on the aforementioned statistical measures, a Bayesian inference technique is used to separate single-trial evoked responses from the ongoing background activity. RESULTS: We show that, when the average event-related potential (AERP) is subtracted from single-trial local field potential time series, a stimulus phase-locked component remains in the residual time series, in stark contrast to the assumption of the common model that no such phase-locked component should exist. Two main consequences of this observation are demonstrated for statistical measures that are computed on the residual time series. First, even though the AERP has been subtracted, the power spectral density, computed as a function of time with a short sliding window, can nonetheless show signs of modulation by the AERP waveform. Second, if the residual time series of two channels co-vary, then their cross-correlation and spectral coherence time functions can also be modulated according to the shape of the AERP waveform. Bayesian estimation of single-trial evoked responses provides further proof that these time-dependent statistical changes are due to remnants of the evoked phase-locked component in the residual time series. CONCLUSIONS: Because trial-to-trial variability of the evoked response is commonly ignored as a contributing factor in evoked potential studies, stimulus-related modulations of power spectral density, cross-correlation, and spectral coherence measures is often attributed to dynamic changes of the connectivity within and among neural populations. This work demonstrates that trial-to-trial variability of the evoked response must be considered as a possible explanation of such modulation.

Animals↗

Statistical methods for the meta-analysis of cluster randomization trials.

Cluster randomization trials have become a very attractive research strategy, particularly for the evaluation of health service interventions. The need to conduct meta-analyses of such trials is also becoming more common. However, as with cluster randomization trials in general, such analyses raise special methodologic challenges. In this paper, we discuss and illustrate several statistical approaches that might be applied to a meta-analysis of cluster randomization trials, each of which has a binary endpoint. Statistical methods for constructing inferences for a summary intervention odds ratio include those based on Mantel-Haenszel procedures, the ratio estimator approach, Woolf procedures and generalized estimating equations using robust variance estimation. The advantages and disadvantages of each method are discussed in the context of an example.

Cluster Analysis↗

CASTOR: clustering algorithm for sequence taxonomical organization and relationships.

Given a set of related proteins, two important problems in biology are the inference of protein subsets such that members of one subset share a common function and the identification of protein regions that possess functional significance. The former is typically approached by hierarchical bottom-up clustering based on pairwise sequence similarity and various linkage rules. The latter is typically approached in a supervised manner, based on global multiple sequence alignment. However, the two problems are inextricably linked, since functional subsets are usually characterized by distinctive functional regions. This paper introduces CASTOR, an automatic and unsupervised system that addresses both problems simultaneously and efficiently. It identifies protein regions that are likely to have functional significance by discovering and refining statistically significant motifs. It infers likely functional protein subsets and their relationships based on the presence of the discovered motifs in a top-down and recursive manner, allowing the identification of both hierarchical and nonhierarchical subset relationships. This is, to our knowledge, the first system that approaches both problems simultaneously in a top-down, systematic manner. CASTOR's performance is evaluated against the G-protein coupled receptor superfamily. The identified protein regions lead to a taxonomical organization of this superfamily that is in remarkable agreement with a biologically motivated one and which outperforms those produced by bottom-up clustering methods. We also find that conventional hierarchical representations may fail to accurately describe the complexity of evolutionary development responsible for the final organization of a complex protein family. In particular, many functional relationships governing distant subfamilies of such a protein family may not be represented hierarchically.

Algorithms↗

The accuracy of statistical methods for estimation of haplotype frequencies: an example from the CD4 locus.

Haplotype analysis has become increasingly important for the study of human disease as well as for reconstruction of human population histories. Computer programs have been developed to estimate haplotype frequencies statistically from marker phenotypes in unrelated individuals. However, there currently are few empirical reports on the accuracy of statistical estimates that must infer linkage phase. We have analyzed haplotypes at the CD4 locus on chromosome 12 that consist of a short tandem-repeat polymorphism and an Alu insertion/deletion polymorphism located 9.8 kb apart, in 398 individuals from 10 geographically diverse sub-Saharan African populations. Haplotype frequency estimates obtained using gene counting based on molecularly haplotyped (phase-known) data were compared with haplotype frequency estimates obtained using the expectation-maximization algorithm. We show that the estimated frequencies of common haplotypes do not differ significantly with the use of phase-known versus phase-unknown data. However, rare haplotypes are occasionally miscalled when their presence/absence must be inferred. Thus, for those research questions for which the common haplotypes are most important, frequency estimates based on the phase-unknown marker-typing results from unrelated individuals will be sufficient. However, in cases where knowledge of rare haplotypes is critical, molecular haplotyping will be necessary to determine linkage phase unambiguously.

Africa South of the Sahara↗

Colored noise and computational inference in neurophysiological (fMRI) time series analysis: resampling methods in time and wavelet domains.

Even in the absence of an experimental effect, functional magnetic resonance imaging (fMRI) time series generally demonstrate serial dependence. This colored noise or endogenous autocorrelation typically has disproportionate spectral power at low frequencies, i.e., its spectrum is (1/f)-like. Various pre-whitening and pre-coloring strategies have been proposed to make valid inference on standardised test statistics estimated by time series regression in this context of residually autocorrelated errors. Here we introduce a new method based on random permutation after orthogonal transformation of the observed time series to the wavelet domain. This scheme exploits the general whitening or decorrelating property of the discrete wavelet transform and is implemented using a Daubechies wavelet with four vanishing moments to ensure exchangeability of wavelet coefficients within each scale of decomposition. For (1/f)-like or fractal noises, e.g., realisations of fractional Brownian motion (fBm) parameterised by Hurst exponent 0 < H < 1, this resampling algorithm exactly preserves wavelet-based estimates of the second order stochastic properties of the (possibly nonstationary) time series. Performance of the method is assessed empirically using (1/f)-like noise simulated by multiple physical relaxation processes, and experimental fMRI data. Nominal type 1 error control in brain activation mapping is demonstrated by analysis of 13 images acquired under null or resting conditions. Compared to autoregressive pre-whitening methods for computational inference, a key advantage of wavelet resampling seems to be its robustness in activation mapping of experimental fMRI data acquired at 3 Tesla field strength. We conclude that wavelet resampling may be a generally useful method for inference on naturally complex time series.

Artifacts↗

Measuring statistical literacy.

This study considers the measurement of Statistical Literacy understanding that goes beyond the basic chance and data skills and knowledge in the mathematics curriculum. This understanding requires application of mathematical skills in a range of contextual situations and draws on aspects of statistics, such as variation and inference, which may not be explicit in the school curriculum. The study reports the outcomes from tests of Statistical Literacy given to 673 students from Grades 5 to 10. It confirms the nature and structure of a previously identified construct of Statistical Literacy and proposes three subgroups of items that address aspects of Statistical Literacy that might usefully be measured by classroom teachers.

Adolescent↗

Statistical criteria in FMRI studies of multisensory integration.

Inferences drawn from functional magnetic resonance imaging (fMRI) studies are dependent on the statistical criteria used to define different brain regions as "active" or "inactive" under the experimental manipulation. In fMRI studies of multisensory integration, additional criteria are used to classify a subset of the active brain regions as "multisensory." Because there is no general agreement in the literature on the optimal criteria for performing this classification, we investigated the effects of seven different multisensory statistical criteria on a single test dataset collected as human subjects performed auditory, visual, and auditory- visual object recognition. Activation maps created using the different criteria differed dramatically. The classification of the superior temporal sulcus (STS) was used as a performance measure, because a large body of converging evidence demonstrates that the STS is important for auditory-visual integration. A commonly proposed criterion, "supra-additivity" or "super-additivity", which requires the multisensory response to be larger than the summed unisensory responses, did not classify STS as multisensory. Alternative criteria, such as requiring the multisensory response to be larger than the maximum or the mean of the unisensory responses, successfully classified STS as multisensory. This practical demonstration strengthens theoretical arguments that the super-additivity is not an appropriate criterion for all studies of multisensory integration. Moreover, the importance of examining evoked fMRI responses, whole brain activation maps, maps from multiple individual subjects, and mixed-effect group maps are discussed in the context of selecting statistical criteria.

Acoustic Stimulation↗

Analyses of cost data in economic evaluations conducted alongside randomized controlled trials.

OBJECTIVE: The adoption and diffusion of new medical treatments depend increasingly on evidence of costs and cost-effectiveness. This evidence is increasingly being generated from economic data collected in randomized clinical trials. The objective of this article is to evaluate the statistical methods used for analysis of cost data in economic evaluations conducted alongside randomized controlled trials. METHODS: Systematic review of economic evaluations based on patient-level cost or resource-use data collected in randomized trials was published in 2003. One hundred fifteen articles were identified from the MEDLINE database. The use of statistical methods for 1) joint comparison of costs and effects and assessment of stochastic uncertainty, 2) incremental cost estimation, and 3) handling of incomplete or censored cost data was evaluated. RESULTS: Only 42 (37%) of the 115 economic evaluations presented a cost-effectiveness ratio or estimated net benefits and 24 (57%) of these reported the uncertainty of this statistic. A comparison of costs alone was more common with 92 (80%) of the 115 studies statistically comparing costs between treatment groups. Of these, about two-thirds (62; 68%) used at least one statistical test appropriate for drawing inferences for arithmetic means. Incomplete cost data were reported in 67 (58%) studies with only two using a published statistical approach for handling censored cost data. CONCLUSION: The quality of statistical methods used in economic evaluations conducted alongside randomized controlled trials was poor in the majority of studies published in 2003. Adoption of appropriate statistical methods is required before the results from such studies can consistently provide valid information to decision-makers.

Cost-Benefit Analysis↗

A knowledge-based information system for monitoring drug levels.

The expert system shell SMR has been enhanced to include information system routines for designing data screens and providing facilities for data entry, storage, retrieval, queries and descriptive statistics. The data for inference making is abstracted from the data base record and inserted into a data array to which the knowledge base is applied to derive the appropriate advice and comments. The enhanced system has been used to develop an intelligent information system for monitoring serum drug levels which includes evaluation of temporal changes and production of specialized printed reports. The module for digoxin has been fully developed and validated. To demonstrate the extension to other drugs a module for phenytoin was constructed with only a rudimentary knowledge base. Data from the request forms together with the S-digoxin results are entered into the data base by the department secretary. The day's results are then reviewed by the clinical pharmacologist. For each case, previous results may be displayed and are taken into account by the system in the decision process. The knowledge base is applied to the data to formulate an evaluative comment on the report returned to the requestor. The report includes a semi-graphic presentation of the current and previous results and either the system's interpretation or one entered by the pharmacologist if he does not agree with it. The pharmacologist's comment is also recorded in the data base for future retrieval, analysis and possible updating of the knowledge base. The system is now undergoing testing and evaluation under routine operations in the clinical pharmacology service. It is a prototype for other applications in both laboratory and clinical medicine currently under development at Uppsala University Hospital. This system may thus provide a vehicle for a more intensive penetration of knowledge-based systems in practical medical applications.

Data Interpretation, Statistical↗

Effects of pdgf-bb on rat dermal fibroblast behavior in mechanically stressed and unstressed collagen and fibrin gels.

The dose-response effects of platelet-derived growth factor BB (PDGF-BB) on rat dermal fibroblast (RDF) behavior in mechanically stressed and unstressed type I collagen and fibrin were investigated using quantitative assays developed in our laboratory. In chemotaxis experiments, RDFs responded optimally (P < 0.05) to a gradient of 10 ng/ml PDGF-BB in both collagen and fibrin. In separate experiments, the migration of RDFs and the traction exerted by RDFs in the presence of PDGF-BB (0, 0.1, 1, 10, or 100 ng/ml) were assessed simultaneously in the presence or absence of stress. RDF migration increased significantly (P < 0.05) at doses of 10 and 100 ng/ml PDGF-BB in collagen and fibrin in the presence and absence of stress. In contrast, the effects of PDGF-BB on RDF traction depended on the gel type and stress state. PDGF-BB decreased fibroblast traction in stressed collagen, but increased traction in unstressed collagen (P < 0.05). No statistical conclusion could be inferred for stressed fibrin, but increasing PDGF-BB decreased traction in unstressed fibrin (P < 0.05). These results demonstrate the complex response of fibroblasts to environmental cues and suggest that mechanical resistance to compaction may be a crucial element in dictating fibroblast behavior.

Animals↗

Origin of extra chromosome in Patau syndrome.

Five live-born infants with Patau syndrome were studied for the nondisjunctional origin of the extra chromosome. Transmission modes of chromosomes 13 from parents to a child were determined using both QFQ- and RFA-heteromorphisms as markers, and the origin was ascertained in all of the patients. The extra chromosome had originated in nondisjunction at the maternal first meiotic division in two patients, at the maternal second meiosis in other two, and at the paternal first meiosis in the remaining one. Summarizing the results of the present study, together with those of the previous studies on a liveborn and abortuses with trisomy 13, nondisjunction at the maternal and the paternal meiosis occurred in this trisomy in the ratio of 14:3. This ratio is not statistically different from that inferred from the previous studies for Down syndrome. These findings suggest that there may be a fundamental mechanism common to the occurrence of nondisjunction in the acrocentric trisomies.

Chromosome Banding↗

Knowledge management as a decision support method: a diagnostic workup strategy application.

We have explored the potential of a computer-based approach called "knowledge management" to aid in clinical problem solving and education. The major features of the approach are its ability to support flexible and immediate access by a user to relevant knowledge and annotation and organization of the knowledge for personal use and subsequent retrieval. We illustrate this approach with its application to diagnostic workup strategy problems. In this application, knowledge may be in the form of static narrative text, diagrams, pictures, graphs, tables, flow charts, or bibliographic citations. Other more dynamic forms of knowledge may be the result of simulations, "what if" analyses or modeling, quantitative mathematical or statistical calculation, or heuristic inference. User assessment has demonstrated the system's ease of use and user perception of its desirability, but underscores the need for a "critical mass" of knowledge before such an approach will be widely utilized.

Decision Support Techniques↗