Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

An empirical comparison of case-control and trio based study designs in high throughput association mapping.

Motivated by high throughput genotyping technology, our aim in this study was to experimentally compare the power and accuracy of case-control and family trio based approaches for haplotype based, large scale, association gene mapping. We compared trio based and case-control study designs in different disease models, and partitioned the performance differences into separate components: those from the sample ascertainment, the effective sample size, and the haplotyping approaches. For systematic and controlled tests, we simulated a rapidly expanding and relatively young isolated population. The experiments were also replicated with real asthma data. We used computationally efficient methods that scale up to large amounts of both markers and individuals. Mapping is based on a haplotype association test for haplotypes of 1-10 markers. For population based haplotype reconstruction, we use HaploRec, and compare it to both a simple trio based inference and true haplotypes. Firstly and surprisingly, statistically inferred population based haplotypes can be equally powerful as true haplotypes. Secondly, as expected, the effective sample size has a clear effect on both gene detection power and mapping accuracy. Thirdly, the sample ascertainment method does not have much effect on mapping accuracy. Finally, an interesting side result is that the simple haplotype association test clearly outperformed exhaustive allelic transmission disequilibrium tests. The results suggest that the case-control design is a powerful alternative to the more laborious family based ascertainment approach, especially for large datasets, and wherever population stratification can be controlled.

Algorithms↗

The effects of normalization on the correlation structure of microarray data.

BACKGROUND: Stochastic dependence between gene expression levels in microarray data is of critical importance for the methods of statistical inference that resort to pooling test-statistics across genes. It is frequently assumed that dependence between genes (or tests) is sufficiently weak to justify the proposed methods of testing for differentially expressed genes. A potential impact of between-gene correlations on the performance of such methods has yet to be explored. RESULTS: The paper presents a systematic study of correlation between the t-statistics associated with different genes. We report the effects of four different normalization methods using a large set of microarray data on childhood leukemia in addition to several sets of simulated data. Our findings help decipher the correlation structure of microarray data before and after the application of normalization procedures. CONCLUSION: A long-range correlation in microarray data manifests itself in thousands of genes that are heavily correlated with a given gene in terms of the associated t-statistics. By using normalization methods it is possible to significantly reduce correlation between the t-statistics computed for different genes. Normalization procedures affect both the true correlation, stemming from gene interactions, and the spurious correlation induced by random noise. When analyzing real world biological data sets, normalization procedures are unable to completely remove correlation between the test statistics. The long-range correlation structure also persists in normalized data.

Algorithms↗

Characterizing stimulus-response functions using nonlinear regressors in parametric fMRI experiments.

Parametric study designs proved very useful in characterizing the relationship between experimental parameters (e.g., word presentation rate) and regional cerebral blood flow in positron emission tomography studies. In a previous paper we presented a method that fits nonlinear functions of stimulus or task parameters to hemodynamic responses, using second-order polynomial expansions. Here we expand this approach to model nonlinear relationships between BOLD responses and experimental parameters, using fMRI. We present a framework that allows this technique to be implemented in the context of the general linear model employed by statistical parametric mapping (SPM). Statistical inferences, in this instance, are based on F statistics and in this respect we emphasize the use of corrected P values for F fields (i.e., SPM¿F¿). The approach is illustrated with a fMRI study that looked at the effect of increasing auditory word-presentation rate. Our parametric design allowed us to characterize different forms of rate-dependent responses in three critical regions: (i) bilateral frontal regions showed a categorical response to the presence of words irrespective of rate, suggesting a role for this region in establishing cognitive (e.g., attentional) set; (ii) in bilateral occipitotemporal regions activations increased linearly with increasing word rate; and (iii) posterior auditory association cortex exhibited a nonlinear (inverted U) relationship to word rate.

Arousal↗

Statistical analysis: the need, the concept, and the usage.

In general, better understanding of the need and usage of statistics would benefit the medical community in India. This paper explains why statistical analysis is needed, and what is the conceptual basis for it. Ophthalmic data are used as examples. The concept of sampling variation is explained to further corroborate the need for statistical analysis in medical research. Statistical estimation and testing of hypothesis which form the major components of statistical inference are construed. Commonly reported univariate and multivariate statistical tests are explained in order to equip the ophthalmologist with basic knowledge of statistics for better understanding of research data. It is felt that this understanding would facilitate well designed investigations ultimately leading to higher quality practice of ophthalmology in our country.

Data Interpretation, Statistical↗

Summarization, smoothing, and inference in epidemiologic analysis. 1991 Ipsen Lecture, Hindsgavl, Denmark.

In a recent article (Epidemiology 1990; 1: 421-429) I resurrected some historical criticisms of conventional statistics in non-randomized, non-randomly sampled studies, and suggested some improvements to current practice in response to these criticisms. Here, I propose that some resolution can be achieved by separating data analysis into summarization, smoothing, and inferential phases. Methods of statistical inference are in fact smoothing methods, as are many methods of descriptive statistics, and as such can be viewed as pattern-recognition devices. Scientific inference is not a statistical process, but instead concerns derivation of explanations for patterns detected by statistical methods. Improvements could be made to all three phases simply by keeping the phases distinct.

Analysis of Variance↗

Investigating potential risk factors for seasonal variation: an example using graphical and spectral analysis methods based on the production of milk components in dairy cattle.

The objective of this study was to illustrate methods for investigating factors associated with seasonality, using milk-component production as an example. Milk-protein and fat percentages showed a seasonal pattern; percentages were lowest during June and July and highest in October and November. Graphical methods were used to compare herd calving patterns to seasonal production patterns and spectral analysis were used to compare seasonal production patterns between farm groups with different management practices. For the comparison of seasonality of production and herd calving patterns, data was obtained from archival records for all cows enrolled in Dairy Herd Improvement (DHI) milk recording in Ontario, Canada from 1990 to 1994. For comparisons of seasonality and management practices, monthly protein and fat percentages were obtained from the Dairy Farmers of Ontario from March 1985 to July 1994. Management information was obtained from responses to questionnaires completed by 364 dairy producers in Ontario. Graphical analyses provided a visualization of the relationship between herd calving patterns and seasonality of production-however, graphical methods alone did not allow statistical inferences to be made. Spectral analyses provided a formal statistical test of the null hypothesis of no association between an independent variable (farm management type) and seasonal production pattern in the data over time, provided that the outcome followed the same seasonal pattern regardless of covariate levels under the null hypothesis.

Animal Husbandry↗

Efficacy and duration of salmeterol powder inhalation in protecting against exercise-induced bronchoconstriction.

BACKGROUND: The protective effect of a new long acting beta 2-agonist, salmeterol, against exercise-induced bronchoconstriction has been documented when given as inhaled aerosol. OBJECTIVE: The aim of the present study was to examine the duration of the protective effect of a single dose of salmeterol, 50 micrograms, inhaled as dry powder, against exercise-induced bronchoconstriction. METHODS: Sixteen patients with reproducible exercise-induced bronchoconstriction were challenged on a treadmill on two prestudy visits and six study days. The patients were challenged 4, 8, and 12 hours postdosing. The study was designed as a double blind placebo-controlled randomized crossover trial. RESULTS: Statistically significant differences in % maximum fall in PEFR were found at four and eight hours postdosing, in favor of salmeterol. At 12 hours postdosing, no clear statistical inference was possible, owing to the presence of a statistical carry-over effect; however, significant differences in area under curve in favor of salmeterol were found at 4, 8, and 12 hours postdosing. CONCLUSION: Salmeterol, 50 micrograms, dry powder inhalation had a protective effect against exercise-induced bronchoconstriction up to 12 hours postdosing, as compared with placebo. No adverse effects were identified.

Administration, Inhalation↗

Incorporating individual error rate into association test of unmatched case-control design.

OBJECTIVES: Genotyping error commonly occurs and could reduce the power and bias statistical inference in genetics studies. In addition to genotypes, some automated biotechnologies also provide quality measurement of each individual genotype. We studied the relationship between the quality measurement and genotyping error rate. Furthermore, we propose two association tests incorporating the genotyping quality information with the goal to improve statistical power and inference. METHODS: 50 pairs of DNA sample duplicates were typed for 232 SNPs by BeadArray technology. We used scatter plot, smoothing function and generalized additive models to investigate the relationship between genotype quality score (q) and inconsistency rate (ĩ) among duplicates. We constructed two association tests: (1) weighted contingency table test (WCT) and (2) likelihood ratio test (LRT) to incorporate individual genotype error rate (epsilon(i)), in unmatched case-control setting. RESULTS: In the 50 duplicates, we found q and ĩ were in strong negative association, suggesting the genotypes with low quality score were more likely to be mistyped. The WCT improved the statistical power and partially corrects the bias in point estimation. The LRT offered moderate power gain, but was able to correct the bias in odds ratio estimation. The two new methods also performed favorably in some scenarios when epsilon(i) was mis-specified. CONCLUSIONS: With increasing number of genetic studies and application of automated genotyping technology, there is a growing need to adequately account for individual genotype error rate in statistical analysis. Our study represents an initial step to address this need and points out a promising direction for further research.

Case-Control Studies↗

Inferring extinction from a sighting record.

The extinctions of plant and animal species are almost never observed directly, but must be inferred from sighting records. This paper reviews some methods for statistical inference about the extinction of a single species based on a record of its sightings.

Algorithms↗

Statistical inversion for medical x-ray tomography with few radiographs: II. Application to dental radiology.

Diagnostic and operational tasks in dental radiology often require three-dimensional information that is difficult or impossible to see in a projection image. A CT-scan provides the dentist with comprehensive three-dimensional data. However, often CT-scan is impractical and, instead, only a few projection radiographs with sparsely distributed projection directions are available. Statistical (Bayesian) inversion is well-suited approach for reconstruction from such incomplete data. In statistical inversion, a priori information is used to compensate for the incomplete information of the data. The inverse problem is recast in the form of statistical inference from the posterior probability distribution that is based on statistical models of the projection data and the a priori information of the tissue. In this paper, a statistical model for three-dimensional imaging of dentomaxillofacial structures is proposed. Optimization and MCMC algorithms are implemented for the computation of posterior statistics. Results are given with in vitro projection data that were taken with a commercial intraoral x-ray sensor. Examples include limited-angle tomography and full-angle tomography with sparse projection data. Reconstructions with traditional tomographic reconstruction methods are given as reference for the assessment of the estimates that are based on the statistical model.

Algorithms↗

Ecological statistics of Gestalt laws for the perceptual organization of contours.

Although numerous studies have measured the strength of visual grouping cues for controlled psychophysical stimuli, little is known about the statistical utility of these various cues for natural images. In this study, we conducted experiments in which human participants trace perceived contours in natural images. These contours are automatically mapped to sequences of discrete tangent elements detected in the image. By examining relational properties between pairs of successive tangents on these traced curves, and between randomly selected pairs of tangents, we are able to estimate the likelihood distributions required to construct an optimal Bayesian model for contour grouping. We employed this novel methodology to investigate the inferential power of three classical Gestalt cues for contour grouping: proximity, good continuation, and luminance similarity. The study yielded a number of important results: (1) these cues, when appropriately defined, are approximately uncorrelated, suggesting a simple factorial model for statistical inference; (2) moderate image-to-image variation of the statistics indicates the utility of general probabilistic models for perceptual organization; (3) these cues differ greatly in their inferential power, proximity being by far the most powerful; and (4) statistical modeling of the proximity cue indicates a scale-invariant power law in close agreement with prior psychophysics.

Form Perception↗

Issues in applied statistics for public health bioterrorism surveillance using multiple data streams: research needs.

The objective of this report is to provide a basis to inform decisions about priorities for developing statistical research initiatives in the field of public health surveillance for emerging threats. Rapid information system advances have created a vast opportunity of secondary data sources for information to enhance the situational and health status awareness of populations. While the field of medical informatics and initiatives to standardize healthcare-seeking encounter records continue accelerating, it is necessary to adapt analytic and statistical methodologies to mature in sync with sibling information science technologies. One major right-of-passage for statistical inference is to advance the optimal application of analytic methodologies for using multiple data streams in detecting and characterizing public health population events of importance. This report first describes the problem in general and the data context, then delineates more specifically the practical nature of the problem and the related issues. Approaches currently applied to data with time-series, statistical process control and traditional inference concepts are described with examples in the section on Statistics and the Role of the Analytic Surveillance Data Monitor. These are the techniques that are providing substance to surveillance professionals and enabling use of multiple data streams. The next section describes use of a more complex approach that takes temporal as well as spatial dimensions into consideration for detection and situational awareness regarding event distributions. The space-time statistic has successfully been used to detect and track public health events of interest. Important research questions which are summarized at the end of this report are described in more detail with respect to the methodological application in the respective sections. This was thought to help elucidate the research requirements as summarized later in the report. Following the description of the space-time scan statistical application; this report extends to a less traditional area of promise given what has been observed in recent application of analytic methods. Bayesian networks (BNs) represent a conceptual step with advantages of flexibility for the public health surveillance community. Progression from traditional to the more extending statistical concepts in the context of the dynamic status quo of responsibility and challenge, leads to a conclusion consisting of categorical research needs. The report is structured by design to inform judgment about how to build on practical systems to achieve better analytic outcomes for public health surveillance. There are references to research issues throughout the sections with a summarization at the end, which also includes items previously unmentioned in the report.

Algorithms↗

Testing for anatomically specified regional effects.

We present a simple method that allows statistical inferences to be made about the significance of regional effects in statistical parametric maps (SPMs) when the approximate location of the effect is specified in advance. The test can be thought of as analogous to assessing activations with uncorrected P values based on the height of SPMs but, in this instance, using the spatial extent or volume of the nearest activated region. The advantage of the current test is that it eschews a correction for multiple comparisons even though the exact location of the expected activation may not be known.

Brain Mapping↗

Design and analysis of intra-subject variability in cross-over experiments.

Recently, interest has grown in the development of inferential techniques to compare treatment variabilities in the setting of a cross-over experiment. In particular, comparison of treatments with respect to intra-subject variability has greater interest than has inter-subject variability. We begin with a presentation of a general approach for statistical inference within a cross-over design. We discuss three different statistical models where model choice depends on the design and assumptions about carry-over effects. Each model incorporates t-variate random subject effects, where t is the number of treatments. We develop maximum likelihood (ML) and restricted maximum likelihood (REML) approaches to derive parameter estimators and we consider a special case in which closed-form expressions for the variance component estimators are available. Finally, we illustrate the methodologies with the analysis of data from three examples.

Area Under Curve↗

The area between curves (ABC)--measure in nutritional anthropometry.

This paper considers a statistic--recently suggested by Mora--for the deviation of a sample distribution from a reference distribution which typically arises in anthropometry when using the nutritional indicators height/age, weight/age or weight/height. The statistic measures the area between curves (ABC) and stands for the mass of the sample distribution which is not covered by the reference distribution. The paper provides a statistical framework for the ABC and includes some minor corrections of Mora's original paper. For the normal distribution situation with common or different variances, formulae are derived which include a partition of ABC into parts corresponding to malnourished and well-nourished groups. However, the main result is a non-parametric generalization of the ABC, motivated by the fact that the nutritional indicators often have skewed distributions with heavier left tails. Non-parametric statistical inference is provided by linking the ABC to the Kolmogorov-Smirnov statistic.

Analysis of Variance↗

Identification of significant periodic genes in microarray gene expression data.

BACKGROUND: One frequent application of microarray experiments is in the study of monitoring gene activities in a cell during cell cycle or cell division. A new challenge for analyzing the microarray experiments is to identify genes that are statistically significantly periodically expressed during the cell cycle. Such a challenge occurs due to the large number of genes that are simultaneously measured, a moderate to small number of measurements per gene taken at different time points, and high levels of non-normal random noises inherited in the data. RESULTS: Based on two statistical hypothesis testing methods for identifying periodic time series, a novel statistical inference approach, the C&G procedure, is proposed to effectively screen out statistically significantly periodically expressed genes. The approach is then applied to yeast and bacterial cell cycle gene expression data sets, as well as to human fibroblasts and human cancer cell line data sets, and significantly periodically expressed genes are successfully identified. CONCLUSION: The C&G procedure proposed is an effective method for identifying statistically significant periodic genes in microarray time series gene expression data.

Algorithms↗

Health aspects of extra-aural noise research.

The WHO definition of "health" is critically discussed in its broad context. Decision making in noise policy has to be made in the evaluation range between social and physical well-being. The term "adverse" is a crucial one in the process of risk characterization. In toxicological terms it refers to the single event itself; in psychosocial terms it refers to the relative number of people affected. The evidence of the association between community noise and cardiovascular outcomes is evaluated. The results of epidemiological studies in this field can be used for decision making when assessing maximum acceptable noise levels in the community. Since dose response relationships were mostly studied with respect to road traffic noise, inferences have to be made with respect to aircraft noise. Issues of statistical inferring are discussed.

Cardiovascular Diseases↗

Evaluating markers for selecting a patient's treatment.

Selecting the best treatment for a patient's disease may be facilitated by evaluating clinical characteristics or biomarker measurements at diagnosis. We consider how to evaluate the potential impact of such measurements on treatment selection algorithms. For example, magnetic resonance neurographic imaging is potentially useful for deciding whether a patient should be treated surgically for Carpal Tunnel Syndrome or should receive less-invasive conservative therapy. We propose a graphical display, the selection impact (SI) curve that shows the population response rate as a function of treatment selection criteria based on the marker. The curve can be useful for choosing a treatment policy that incorporates information on the patient's marker value exceeding a threshold. The SI curve can be estimated using data from a comparative randomized trial conducted in the population as long as treatment assignment in the trial is independent of the predictive marker. Estimating the SI curve is therefore part of a post hoc analysis to determine whether the marker identifies patients that are more likely to benefit from one treatment over another. Nonparametric and parametric estimates of the SI curve are proposed in this article. Asymptotic distribution theory is used to evaluate the relative efficiencies of the estimators. Simulation studies show that inference is straightforward with realistic sample sizes. We illustrate the SI curve and statistical inference for it with data motivated by an ongoing trial of surgery versus conservative therapy for Carpal Tunnel Syndrome.

Algorithms↗