Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Regression analysis of biomedical research data based on a repeated measure or cluster sample.

Research in biology and medicine often entails estimating the effect of an exposure variable X on a response variable Y from a cluster sample, that is, where X and Y may be measured repeatedly from the same subject (cluster), or X and Y may be measured from two or more subjects based on some related grouping such as a litter or household (cluster). Typically, regression analysis of the data is performed ignoring the subject or grouping (cluster) identity. This analytical approach has two drawbacks. First, the statistical inference of the regression coefficient of Y on X (beta) ignored cluster identity will likely be biased. More serious though, is that beta ignored cluster identity will likely be quite discrepant from the average within-cluster beta. It is the latter that is of relevance to the research question. Indeed it is not clear what beta ignored cluster identity really conveys. We describe a multiple regression model for the analysis of data from a cluster sample. The model treats the cluster as a nominal confounding variable to be adjusted. The idea is to represent the cluster by a set of dummy variables to be included as explanatory variables in the regression model. This model gives the average within-cluster beta with valid statistical inference. Numeric examples were used to illustrate the application of the dummy variable multiple regression model. This statistical method can be implemented by any software package that includes multiple regression, and virtually all commercial packages include this procedure.

Cluster Analysis↗

Generalized least-squares method applied to fMRI time series with empirically determined correlation matrix.

Functional magnetic resonance imaging (fMRI) time series analysis and statistical inferences about the effect of a cognitive task on the regional cerebral blood flow (rCBF) are largely based on the linear model. However, this method requires that the error vector is a gaussian variable with an identity correlation matrix. When this assumption cannot be accepted, statistical inferences can be made using generalized least squares. In this case, knowledge of the covariance matrix of the error vector is needed. In the present report, we propose a method that needs stationarity of the autocorrelation function but is more flexible than autoregressive model of order p (AR(p)) models because it is not necessary to predefine a relation between coefficients of the correlation matrix. We tested this method on sets of simulated data (with presence of an effect of interest or not) representing a time series with a monotonically decreasing autocorrelation function. This time series mimicked an experiment using a random event-related design that does not create correlation between scans. The autocorrelation function is empirically determined and used to reconstitute the correlation matrix as the toeplitz matrix built from the autocorrelation function. When applied to simulated time series with no effect of interest, this method allows the determination of F values corresponding to the accurate false positive level. Moreover, when applied to time series with an effect of interest, this method gives a density function of F values which allows the rejection of the null hypothesis. This method provides a flexible but interpretable time domain noise model.

Attention↗

A blocking Gibbs sampling method to detect major genes with phenotypic data from a diallel mating.

Diallel mating is a frequently used design for estimating the additive and dominance genetic (polygenic) effects involved in quantitative traits observed in the half- and full-sib progenies generated in plant breeding programmes. Gibbs sampling has been used for making statistical inferences for a mixed-inheritance model (MIM) that includes both major genes and polygenes. However, using this approach it has not been possible to incorporate the genetic properties of major genes with the additive and dominance polygenic effects in a diallel mating population. A parent block Gibbs sampling method was developed in this study to make statistical inferences about the major gene and polygenic effects on quantitative traits for progenies derived from a half-diallel mating design. Using simulated data sets with different major and polygenic effects, the proposed method accurately estimated the major and polygenic effects of quantitative traits, and possible genotypes of parents and progenies. The impact of specifying different prior distributions was examined and was found to have little effect on inference on the posterior distribution. This approach was applied to an experimental data set of Loblolly pine (Pinus taeda L.) derived from a 6-parent half-diallel mating. The result indicated that there might be a recessive major gene affecting height growth in this diallel population.

Chromosome Mapping↗

Exact finite-sample significance and confidence regions for goodness-of-fit statistics in one-way multinomials.

As multinomial processing tree models become more popular in psychology, appropriate methods for statistical inference also become more necessary. Conventional methods are all based on the asymptotic chi-square approximation to the exact multinomial test, but the accuracy of this approximation has been shown to be poor in the usual small-sample situation. This paper describes an efficient algorithm that allows the exact multinomial test to be applied without incurring prohibitive computational costs. The algorithm is well suited to addressing three different aspects of statistical inference in one-way multinomials: (i) evaluating the exact significance of a multinomial test; (ii) determining significance at a given preset level; and (iii) enumerating exact confidence regions. Examples are given that illustrate how this algorithm accomplishes each of these tasks, and an analysis of its computational cost in some small-sample situations is also provided.

Algorithms↗

Sampling variability and estimates of density dependence: a composite-likelihood approach.

It is well known that sampling variability, if not properly taken into account, affects various ecologically important analyses. Statistical inference for stochastic population dynamics models is difficult when, in addition to the process error, there is also sampling error. The standard maximum-likelihood approach suffers from large computational burden. In this paper, I discuss an application of the composite-likelihood method for estimation of the parameters of the Gompertz model in the presence of sampling variability. The main advantage of the method of composite likelihood is that it reduces the computational burden substantially with little loss of statistical efficiency. Missing observations are a common problem with many ecological time series. The method of composite likelihood can accommodate missing observations in a straightforward fashion. Environmental conditions also affect the parameters of stochastic population dynamics models. This method is shown to handle such nonstationary population dynamics processes as well. Many ecological time series are short, and statistical inferences based on such short time series tend to be less precise. However, spatial replications of short time series provide an opportunity to increase the effective sample size. Application of likelihood-based methods for spatial time-series data for population dynamics models is computationally prohibitive. The method of composite likelihood is shown to have significantly less computational burden, making it possible to analyze large spatial time-series data. After discussing the methodology in general terms, I illustrate its use by analyzing a time series of counts of American Redstart (Setophaga ruticilla) from the Breeding Bird Survey data, San Joaquin kit fox (Vulpes macrotis mutica) population abundance data, and spatial time series of Bull trout (Salvelinus confluentus) redds count data.

Algorithms↗

The effect of filter size on VBM analyses of DT-MRI data.

Voxel-based morphometry (VBM) has been used to analyze diffusion tensor MRI (DT-MRI) data in a number of studies. In VBM, following spatial normalization, data are smoothed to improve the validity of statistical inferences and to reduce inter-individual variation. However, the size of the smoothing filter used for VBM of DT-MRI data is highly variable across studies. For example, a literature review revealed that Gaussian smoothing kernels ranging in size (full width at half maximum) from zero to 16 mm have been used in DT-MRI VBM type studies. To investigate the effect of varying filter size in such analyses, whole brain DT-MRI data from 14 schizophrenic patients were compared with those of 14 matched control subjects using VBM, when the filter size was varied from zero to 16 mm. Within this range of smoothing, four different conclusions regarding apparent patient control differences could be made: (i) no significant patient-control differences; (ii) reduced FA in right superior temporal gyrus (STG) in patients; (iii) reduced FA in both right STG and left cerebellum in patients; and (iv) reduced FA only in left cerebellum in patients. These findings stress the importance of recognizing the effect of the matched filter theorem on VBM analyses of DT-MRI data. Finally, we investigated whether one of the underlying assumptions of parametric VBM, i.e., the normality of the residuals, is met. Our results suggest that, even with moderate smoothing, a large number of voxels within central white matter regions may have non-normally distributed residuals thus making valid statistical inferences with a parametric approach problematic in these areas.

Adult↗

Effects of sampling regime on the mean and variance of home range size estimates.

1. Although the home range is a fundamental ecological concept, there is considerable debate over how it is best measured. There is a substantial literature concerning the precision and accuracy of all commonly used home range estimation methods; however, there has been considerably less work concerning how estimates vary with sampling regime, and how this affects statistical inferences. 2. We propose a new procedure, based on a variance components analysis using generalized mixed effects models to examine how estimates vary with sampling regime. 3. To demonstrate the method we analyse data from one study of 32 individually marked roe deer and another study of 21 individually marked kestrels. We subsampled these data to simulate increasingly less intense sampling regimes, and compared the performance of two kernel density estimation (KDE) methods, of the minimum convex polygon (MCP) and of the bivariate ellipse methods. 4. Variation between individuals and study areas contributed most to the total variance in home range size. Contrary to recent concerns over reliability, both KDE methods were remarkably efficient, robust and unbiased: 10 fixes per month, if collected over a standardized number of days, were sufficient for accurate estimates of home range size. However, the commonly used 95% isopleth should be avoided; we recommend using isopleths between 90 and 50%. 5. Using the same number of fixes does not guarantee unbiased home range estimates: statistical inferences differ with the number of days sampled, even if using KDE methods. 6. The MCP method was highly inefficient and results were subject to considerable and unpredictable biases. The bivariate ellipse was not the most reliable method at low sample sizes. 7. We conclude that effort should be directed at marking more individuals monitored over long periods at the expense of the sampling rate per individual. Statistical results are reliable only if the whole sampling regime is standardized. We derive practical guidelines for field studies and data analysis.

Animals↗

DNA microarray profiling to identify angiotensin-responsive genes in vascular smooth muscle cells: potential mediators of vascular disease.

Angiotensin II (Ang II) induces changes in vessel structure by its capacity to activate genes that are coupled to signaling pathways such as extracellular signal-regulated kinase (ERK), p38, and phosphatidylinositol 3-kinase (PI3K). Using a DNA microarray containing 5088 genes and expressed sequence tags, we initially established a database of replicated experiments (n=4) to define the variances in mRNA expression in response to Ang II versus vehicle treatment. We observed a wide range of values for the coefficients of variation in a gene-specific manner. Guided by power calculations, we used statistical inference on a sufficient number of experimental replicates to minimize the number of false-negatives and define a subset of Ang II-responsive genes (P<0.05). To further characterize the molecular circuitry that couples Ang II stimulation with mRNA expression, we assessed expression profiles in the presence and absence of inhibitors of ERK, p38, and PI3K. Using two different methods of computational cluster analysis, we identified a subset of six matricellular proteins (eg, osteopontin and plasminogen activator inhibitor-1) that are coordinately upregulated by Ang II via an ERK/p38-dependent pathway. In addition, these cluster analyses identified calpactins I and II as novel Ang II-responsive genes. Given that Ang II promotes vascular lesion formation, we examined whether this matricellular gene cluster was also coordinately regulated in vivo. Indeed, we demonstrate that both calpactin I and osteopontin are upregulated in response to vascular injury. Taken together, the combined use of DNA microarrays, statistical inference, and cluster analysis identified novel, coordinately regulated Ang II-responsive genes that may mediate vascular lesion formation.

Angiotensins↗

Significance testing, interval estimation or Bayesian inference: comments to "Extracting a maximum of useful information from statistical research data" by S. Sohlberg and G. Andersson.

Statistical inference plays an important part in the formation of scientific knowledge in psychology. Starting from a paper by Sohlberg and Andersson (2005; Scandinavian Journal of Psychology, 46, 69-77) these issues are discussed. It is argued that interval estimates are easy to understand and that they are more suitable than significance testing for most problems. Bayesian inference is a coherent description of the information building process. With some examples it is shown that null hypothesis significance testing is full of contradictions. Finally, some other important issues like convenience sampling and model selection are shortly mentioned.

Bayes Theorem↗

Efficiently measuring recognition performance with sparse data.

We examine methods for measuring performance in signal-detection-like tasks when each participant provides only a few observations. Monte Carlo simulations demonstrate that standard statistical techniques applied to a d' analysis can lead to large numbers of Type I errors (incorrectly rejecting a hypothesis of no difference). Various statistical methods were compared in terms of their Type I and Type II error (incorrectly accepting a hypothesis of no difference) rates. Our conclusions are the same whether these two types of errors are weighted equally or Type I errors are weighted more heavily. The most promising method is to combine an aggregate d' measure with a percentile bootstrap confidence interval, a computer-intensive nonparametric method of statistical inference. Researchers who prefer statistical techniques more commonly used in psychology, such as a repeated measures t test, should use gamma (Goodman & Kruskal, 1954), since it performs slightly better than or nearly as well as d'. In general, when repeated measures t tests are used, gamma is more conservative than d': It makes more Type II errors, but its Type I error rate tends to be much closer to that of the traditional .05 alpha level. It is somewhat surprising that gamma performs as well as it does, given that the simulations that generated the hypothetical data conformed completely to the d' model. Analyses in which H--FA was used had the highest Type I error rates. Detailed simulation results can be downloaded from www.psychonomic.org/archive/Schooler-BRM-2004.zip.

Computer Simulation↗

A cumulative specificity model for proteases from human immunodeficiency virus types 1 and 2, inferred from statistical analysis of an extended substrate data base.

Statistical analysis of an expanded data base of regions in viral polyproteins and in non-viral proteins that are sensitive to hydrolysis by the protease from human immunodeficiency virus (HIV) type 1 has generated a model which characterizes the substrate specificity of this retroviral enzyme. The model leads to an algorithm for predicting protease-susceptible sites from primary structure. Amino acids in each of the sites from P4 to P4' are tabulated for 40 protein substrates, and the frequency of occurrence for each residue is compared to the natural abundance of that amino acid in a selected data set of globular proteins. The results suggest that the highest stringency for particular amino acid residues is at the P2, P1, and P2' positions of the substrate. The broad specificity of the HIV-1 protease appears to be a consequence of its being able to bind productively substrates in which interactions with only a few Pi or Pi' side-chains need be optimized. The analysis, extended to 22 protein segments cleaved by the HIV-2 protease, delineates marked differences in specificity from that of the HIV-1 enzyme.

Actins↗

Characterization of dose-response relationships inferred by statistically significant trend tests.

A method is proposed for classifying various experimental outcomes associated with statistically significant trend tests according to a set of sequential testing within a family of closed (under intersections) one-sided tests. The intent of the procedure is to characterize the general shape of implied dose-response relationships, taking care neither to inflate the false-positive (Type I) error rate by overtesting, nor to sacrifice power by overadjusting for multiple comparisons.

Animals↗

Randomization, statistics, and causal inference.

This paper reviews the role of statistics in causal inference. Special attention is given to the need for randomization to justify causal inferences from conventional statistics, and the need for random sampling to justify descriptive inferences. In most epidemiologic studies, randomization and random sampling play little or no role in the assembly of study cohorts. I therefore conclude that probabilistic interpretations of conventional statistics are rarely justified, and that such interpretations may encourage misinterpretation of nonrandomized studies. Possible remedies for this problem include deemphasizing inferential statistics in favor of data descriptors, and adopting statistical techniques based on more realistic probability models than those in common use.

Bayes Theorem↗

The power of the Z statistic: implications for trauma research and quality assurance review.

The Z statistic can be used to test whether the observed number of survivors in a specific trauma population is significantly different from what would be expected based on the Major Trauma Outcome Study (MTOS) norms. However, as with any statistic, inferences based on the Z statistic should be made with care. This is particularly true when a non-significant Z statistic is observed. The purpose of this paper, using data from a large, urban trauma registry, is to illustrate how the power of the Z statistic, or its ability to detect a difference between observed and expected survival, is influenced by the magnitude of the difference, the direction of the difference, the survival probability distribution of the study population, and the sample size. The implications for trauma research and quality assurance review are discussed.

Craniocerebral Trauma↗

A gene-environment interaction between inferred kallikrein genotype and potassium.

Urinary kallikrein excretion has been shown statistically to be partially determined by a major gene in large Utah pedigrees with the use of segregation analysis. A previous twin analysis of environmental factors influencing urinary kallikrein level showed that urinary potassium twin differences were strongly related to differences in urinary kallikrein. The present study uses 769 individuals in 58 Utah pedigrees to analyze the association of urinary potassium with urinary kallikrein within statistically inferred kallikrein genotypes. Fitting genotype-specific curves relating urinary kallikrein level to 12-hour urinary potassium amount within a major gene, polygene, and common environment model, we showed a significant statistical urinary potassium interaction with the inferred major gene for kallikrein (P = .0002). The heterozygotes (with a frequency of 50%) had a significant association between urinary kallikrein and potassium (slope, 0.51 +/- 0.04 SD), whereas there was no association with potassium in the low homozygotes, suggesting a genetic defect involving the kallikrein response to potassium. The model predicted that an increase in urinary potassium excretion of 0.8 SD above the mean in these pedigrees would be associated with high kallikrein levels in the heterozygotes similar to the high homozygotes. A decrease of 1.3 SD in urinary potassium excretion in heterozygous individuals was associated with kallikrein levels similar to the homozygous individuals with low kallikrein. Because in the steady state urinary potassium represents dietary potassium intake, this study suggests that an increase in dietary potassium intake in 50% of these pedigree members, estimated to be heterozygous at the kallikrein locus, would be associated with an increase in an underlying genetically determined low kallikrein level.(ABSTRACT TRUNCATED AT 250 WORDS)

Adolescent↗

Directed Convergence in Stable Percept Acquisition.

We view a perceptual capacity as a nondeductive inference, represented as a function from a set of premises to a set of conclusions. The application of the function to a single premise to produce a single conclusion is called a "percept" or "instantaneous percept." We define a stable percept as a convergent sequence of instantaneous percepts. Assuming that the sets of premises and conclusions are metric spaces, we introduce a strategy for acquiring stable percepts, called directed convergence. We consider probabilistic inferences, where the premise and conclusion sets are spaces of probability measures, and in this context we study Bayesian probabilistic/recursive inference. In this type of Bayesian inference the premises are probability measures, and the prior as well as the posterior is updated nontrivially at each iteration. This type of Bayesian inference is distinguished from classical Bayesian statistical inference where the prior remains fixed, and the posterior evolves by conditioning on successively more punctual premises. We indicate how the directed convergence procedure may be implemented in the context of Bayesian probabilistic/recursive inference. We discuss how the L(infinity) metric can be used to give numerical control of this type of Bayesian directed convergence. Copyright 2001 Academic Press.

Journal Article↗