Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Characterizing dynamic brain responses with fMRI: a multivariate approach.

In this paper we present a multivariate analysis of evoked hemodynamic responses and their spatiotemporal dynamics as measured with fast fMRI. This analysis uses standard multivariate statistics (MANCOVA) and the general linear model to make inferences about effects of interest and canonical variates analysis (CVA) to describe the important features of these effects. We have used these techniques to characterize the form of hemodynamic transients that are evoked during a cognitive or sensorimotor task. In particular we do not assume that the neural or hemodynamic response reaches some "steady state" but acknowledge that these physiological changes could show profound task-dependent adaptation and time-dependent changes during the task. To address this issue we have modeled hemodynamic responses using appropriate temporal basis functions and estimated their exact form within the general linear model using MANCOVA. We do not propose that this analysis is a particularly powerful way to make inferences about functional specialization (or more generally functional anatomy) because it only provides statistical inferences about the distributed (whole brain) responses evoked by different conditions. However, its application to characterizing the temporal aspects of evoked hemodynamic responses reveals some compelling and somewhat unexpected perspectives on transient but stereotyped responses to changes in cognitive or sensorimotor processing. The most remarkable observation is that these responses can be biphasic and show profound differences in their form depending on the extant task or condition. Furthermore these differences can be seen in the absence of changes in mean signal.

Arousal↗

Statistical phylogeography.

While studies of phylogeography and speciation in the past have largely focused on the documentation or detection of significant patterns of population genetic structure, the emerging field of statistical phylogeography aims to infer the history and processes underlying that structure, and to provide objective, rather than ad hoc explanations. Methods for parameter estimation are now commonly used to make inferences about demographic past. Although these approaches are well developed statistically, they typically pay little attention to geographical history. In contrast, methods that seek to reconstruct phylogeographic history are able to consider many alternative geographical scenarios, but are primarily nonstatistical, making inferences about particular biological processes without explicit reference to stochastically derived expectations. We advocate the merging of these two traditions so that statistical phylogeographic methods can provide an accurate representation of the past, consider a diverse array of processes, and yet yield a statistical estimate of that history. We discuss various conceptual issues associated with statistical phylogeographic inferences, considering especially the stochasticity of population genetic processes and assessing the confidence of phylogeographic conclusions. To this end, we present some empirical examples that utilize a statistical phylogeographic approach, and then by contrasting results from a coalescent-based approach to those from Templeton's nested cladistic analysis (NCA), we illustrate the importance of assessing error. Because NCA does not assess error in its inferences about historical processes or contemporary gene flow, we performed a small-scale study using simulated data to examine how our conclusions might be affected by such unconsidered errors. NCA did not identify the processes used to simulate the data, confusing among deterministic processes and the stochastic sorting of gene lineages. There is as yet insufficient justification of NCA's ability to accurately infer or distinguish among alternative processes. We close with a discussion of some unresolved problems of current statistical phylogeographic methods to propose areas in need of future development.

Animals↗

Proceedings of the SMBE Tri-National Young Investigators' Workshop 2005. Accurate inference and estimation in population genomics.

Both intra- and interspecific genomic comparisons have revealed local similarities in the level and frequency of mutational variation, as well as in patterns of gene expression. This autocorrelation between measurements leads to violations of assumptions of independence in many statistical methods, resulting in misleading and incorrect inferences. Here I show that autocorrelation can be due to many factors and is present across the genome. Using a one-dimensional spatial stochastic model, I further show how previous results can be employed to correct for autocorrelation along chromosomes in population and comparative genomics research. When multiple hypothesis tests are autocorrelated, I demonstrate that a simple correction can lead to increased power in statistical inference. I present a preliminary analysis of population genomic data from Drosophila simulans to show the ubiquity of autocorrelation and applicability of the methods proposed here.

Animals↗

Inferred HLA haplotype information for donors from hematopoietic stem cells donor registries.

Human leukocyte antigen (HLA) matching remains a key issue in the outcome of transplantation. In hematopoietic stem cell transplantation with unrelated donors, the matching for compatible donors is based on the HLA phenotype information. In familial transplantation, the matching is achieved at the haplotype level because donor and recipient share the block-transmitted major histocompatibility complex region. We present a statistical method based on the HLA haplotype inference to refine the HLA information available in an unrelated situation. We implement a systematic statistical inference of the haplotype combinations at the individual level. It computes the most likely haplotype pair given the phenotype and its probability. The method is validated on 301 phase-known phenotypes from CEPH families (Centre d'Etude du Polymorphisme Humain). The method is further applied to 85,933 HLA-A B DR typed unrelated donors from the French Registry of hematopoietic stem cells donors (France Greffe de Moelle). The average value of prediction probability is 0.761 (SD 0.199) ranging from 0.26 to 1. Correlations between phenotype characteristics and predictions are also given. Homozygosity (OR = 2.08; [2.02-2.14] p <10(-3)) and linkage disequilibrium (p <10(-3)) are the major factors influencing the quality of prediction. Limits and relevance of the method are related to limits of haplotype estimation. Relevance of the method is discussed in the context of HLA matching refinement.

Algorithms↗

A multivariate study of some lung function tests at different age groups in healthy Indian males.

Considerable attempts have been made to study the changes in lung function in man in different age groups using univariate statistical techniques in which the lung function tests were assumed to be independent of each other. Actually the lung function tests are well correlated with each other and, thus, the inferences drawn on the basis of univariate statistical analysis may be misleading due to the violation of the assumption of independence. On the other hand, simultaneous changes in lung function in man in different age groups cannot be tested using univariate statistical techniques. Keeping in view such shortcomings of the univariate statistical techniques, an attempt has been made in the present investigation to study simultaneous changes in some lung function tests [viz. vital capacity (VC), forced vital capacity (FVC), forced expiratory volume for one second (FEV1), expiratory reserve volume (ERV), inspiratory capacity (IC) and maximum voluntary ventilation (MVV)] at different age groups (viz. 21-25, 26-30, 31-35, 36-40, 41-45, 46-50, 51-55, 56-60 and 61-70 years) in healthy Indian males using multivariate statistical techniques (viz. Wilks' statistic (A) and Mahalanobis' D2 statistic) for drawing valid statistical inferences. It is concluded that remarkable significant changes take place in lung function after the age of forty years.

Adult↗

Approximate Bayesian inference for random effects meta-analysis.

Whilst meta-analysis is becoming a more commonplace statistical technique, Bayesian inference in meta-analysis requires complex computational techniques to be routinely applied. We consider simple approximations for the first and second moments of the parameters of a Bayesian random effects model for meta-analysis. These computationally inexpensive methods are based on simple analytical formulae that provide an efficient tool for a qualitative analysis and a quick numerical estimation of posterior quantities. They are shown to lead to sensible approximations in two examples of meta-analyses and to be in broad agreement with the more computationally intensive Gibbs sampling.

Antibiotic Prophylaxis↗

Using crime statistics in nursing.

1. Inferences about crime only may be made when comparing a particular place (state, city) to other similar places. 2. Crime and deviance statistics are unique in that one must have an adequate population base to make inferences that are meaningful. If the population base is not adequate, an inference cannot be made unless this condition is remedied. 3. Use only rates of occurrences per 100,000 people per year as a standard for comparison. If some other rate is used, then this fact must be clearly stated and the reason for deviating adequately explained.

Crime↗

On making multiple comparisons in clinical and experimental pharmacology and physiology.

1. It is a central thesis of this review that in clinical and experimental pharmacology and physiology the goal of statistical analysis should be to minimize the risk of making any false-positive inferences from the results of an experiment (experimentwise Type I error). 2. It is common in clinical and experimental pharmacology and physiology for the effects of several treatments to be tested within a single experiment. Specific intercomparisons of these several effects, made in a pairwise or more complex fashion, inflates the risk of making false-positive inferences unless special statistical procedures are used. 3. A number of multiple comparison procedures is described and their ability to control experimentwise Type I error is evaluated critically. 4. When only a few (less than 5) of all possible pairwise or more complex comparisons are made between treatment groups, the Dunn-Sidák procedure provides maximum protection against excessive experimentwise Type I error and is very convenient to use. 5. When a control group is compared with all other treatment groups in a pairwise fashion, especially when the number of groups is large, the Dunnett procedure is more powerful than the Dunn-Sidák. 6. If investigators insist on making all possible pairwise comparisons among treatment groups, the Tukey-Kramer procedure provides maximum protection against false-positive inferences but inflates the Type II error rate. If it is especially important to avoid Type II error then the more complicated, stepwise procedures of the Ryan-Peritz-Welsch variety should be considered.

Analysis of Variance↗

Support with clarity: a proper trend in medical statistics.

BACKGROUND: The aim of this study was to establish effective methods to review and evaluate, and to emphasize in support with clarity as a proper trend in medical statistics. METHODS: The clinical research material used in this study is stemmed from JAMC. STUDY: I, N is 220 subjects, from Pattern of Coronary Arterial Distribution and Its Relation to Coronary Artery Diameter, Z A Kaimkhani, MM Ali, AMA Faruqi JAMC Jan-Mar 2005; 17(1): 40-3. STUDY: II, N is 105 patients, from Sclerotherapy Plus Octreotide Versus Sclerotherapy Alone In The Management Of Gastro-Oesophageal Variceal Hemorrhage. HA Shah, K Mumtaz, W Jafri, S Abid, S Hamid, A Ahmad, Z Abbas. JAMC Jan-Mar 2005; 17(1): 10-4. 2 Systemic review and evaluation with statistical principles is to be used. ASSESSMENT AND DISCUSSION: The reports of 2 clinical researches from JAMC are assessed, and both used for demonstrating the characteristics and pitfalls statistically. CONCLUSION: Before reaching any significant difference in statistics, hopefully, all clinicians will be able to deal with the data to be measured by selecting proper statistical models as the best as we can in order to gain appropriate inference in medical statistics.

Humans↗

Compositional statistics: an improvement of evolutionary parsimony and its application to deep branches in the tree of life.

We present compositional statistics, a new method of phylogenetic inference, which is an extension of evolutionary parsimony. Compositional statistics takes account of the base composition of the compared sequences by using nucleotide positions that evolutionary parsimony ignores. It shares with evolutionary parsimony the features of rate invariance and the fundamental distinction between transitions and transversions. Of the presently available methods of phylogenetic inference, compositional statistics is based on the fewest and mildest assumptions about the mode of DNA sequence evolution. It is therefore applicable to phylogenetic studies of the most distantly related organisms or molecules. This was illustrated by analyzing conservative positions in the DNA sequences of the large subunit of RNA polymerase from three archaebacterial groups, a eubacterium, a chloroplast, and the three eukaryotic polymerases. Internally consistent results, which are in accord with our knowledge of organelle origin and archaebacterial physiology, were achieved.

Amino Acid Sequence↗

Fatal alcohol poisoning: medico-legal practices and mortality statistics.

Compilation of mortality statistics from death certificate data is based on international and national conventions which in certain situations result in the underlying cause-of-death other than that established and reported by the physician. The present study compares all fatal alcohol poisonings in 1997 as registered on forensic toxicological grounds at the accredited central laboratory and as presented in the national cause-of-death statistics, according to the underlying cause-of-death, by applying international statistical rules and principles in ICD-10. Four groups were formed, and case frequencies in each group were obtained from forensic toxicological data, group "T51" for acute poisonings due to alcohol alone, and group "Comb" for acute alcohol poisonings combined with some drug, medicament or other biological substance, and from cause-of-death statistics data, group "X45", for deaths from alcohol poisoning, and group "F102" for those medico-legal fatal alcohol poisoning deaths which at the statistics office were inferred to be due to alcoholism. The study shows that in Finland the officially compiled statistics on fatal alcohol poisonings, when compared with medico-legal statements based on forensic toxicological examinations, were underrepresented by 31.4% in 1997. About two-thirds of this underrepresentation is explained by preferring, as the underlying cause-of-death, alcoholism to acute alcohol poisoning, and about one-third by preferring, in cases of acute combined poisonings, the drug component to the alcohol. From 1998 onwards, more emphasis has been put on the alcohol component when coding medico-legally proven accidental deaths from simultaneous poisoning with alcohol and a medicinal agent. This change in coding practices presumably explains the subsequent decline in the annual underrepresentation rate of alcohol poisoning in mortality statistics to the level of 15-16%. It is concluded that the present ICD rules inevitably lead to underrepresentation of alcohol poisonings in the mortality statistics, and conceptual and practical proposals for future procedures are made.

2-Propanol↗

Haplotype inference in random population samples.

Contemporary genotyping and sequencing methods do not provide information on linkage phase in diploid organisms. The application of statistical methods to infer and reconstruct linkage phase in samples of diploid sequences is a potentially time- and labor-saving method. The Stephens-Smith-Donnelly (SSD) algorithm is one such method, which incorporates concepts from population genetics theory in a Markov chain-Monte Carlo technique. We applied a modified SSD method, as well as the expectation-maximization and partition-ligation algorithms, to sequence data from eight loci spanning >1 Mb on the human X chromosome. We demonstrate that the accuracy of the modified SSD method is better than that of the other algorithms and is superior in terms of the number of sites that may be processed. Also, we find phase reconstructions by the modified SSD method to be highly accurate over regions with high linkage disequilibrium (LD). If only polymorphisms with a minor allele frequency >0.2 are analyzed and scored according to the fraction of neighbor relations correctly called, reconstructions are 95.2% accurate over entire 100-kb stretches and are 98.6% accurate within blocks of high LD.

Algorithms↗

Statistical significance--a misconstrued notion in medical research.

The P-value is the significance probability of obtaining a value of the test statistic that is as extreme, in relation to the null hypothesis, as that observed. Medical researchers may, in some situations, disagree on its appropriate use or on its interpretation as a summary measure of consistency with the null hypothesis in a particular data set. More informative statistical measures such as the likelihood ratio and the Bayesian posterior probability have been suggested for drawing inferences from clinical trials and epidemiologic studies. Causal inference is not statistical in nature; rather it strives to provide scientific explanations or criticisms of proposed explanations that would describe the observed data pattern. In this context, it is important to remember that a finding may not be medically important, or a causal hypothesis may even not be true even if a study shows a significant P-value.

Bayes Theorem↗

Grass evolution inferred from chromosomal rearrangements and geometrical and statistical features in RNA structure.

The grasses (Poaceae) represent a monophyletic lineage that arose about 70 million years ago. The lineage contains about 10,000 species that differ widely in morphology and physiology. Species show striking differences in genome size, a feature important in the context of conservation of gene content and order (synteny and colinearity) and in the extension of genomic information directly from one grass species to another using comparative approaches. Grass diversification has been a contentious issue, as the exact branching order of the various subfamilies has been difficult to establish with standard methods. This motivated an evolutionary study of deep phylogenetic relationships based on the structure of coding and non-coding RNA molecules and on chromosomal rearrangements. Phylogenetic relationships in the grass family were inferred directly from the structure of RNA using cladistic principles and considerations in statistical mechanics. Coded attributes describing topological and thermodynamic information embedded in RNA molecules were treated as linearly ordered multi-state characters and were polarized by fixing the direction of character transformation toward molecular order. Intrinsically rooted phylogenies derived from the structure of signal recognition particle (SRP) RNA, the mRNA encoded by the early nodulation gene enod40, the small subunit of ribosomal RNA (rRNA), and the internal transcribed spacer ITS1 of rRNA established an order for the diversification of major grass lineages, suggesting a sister relationship of the Pooideae and the PACCAD clade. This same conclusion was reached when large-scale chromosomal rearrangements derived from the comparative genetic mapping of cereal genomes were studied. Chromosomal complements aligned in the most parsimonious manner allowed identification and coding of characters depicting chromosomal translocations, insertions, and linkage block arrangements and the reconstruction of phylogenetic trees based on large-scale chromosomal structure. Congruent reconstruction of deep branching relationships using geometrical and statistical features of RNA structure and orthology and large scale chromosomal recombination events support assumptions of polarization in character argumentation, and fail to falsify the claim that extant grass chromosomes can be considered combinations of linkage blocks of an ancestor of the rice genome. Congruence also suggests that the universal tendency toward order in RNA and the search for the most parsimonious organization of be genome architecture appear to be mutually supported drivers of molecular evolution. The study clarifies the relationship of major clades in the grasses, shows that phylogenetic history can be reconstructed effectively from the combinatorial exchange of chromosomal linkage blocks, and reveals considerable phylogenetic signal embedded in the structure of signal polypeptide-coding mRNA molecules, describing an instance where mRNA structure is the subject of strong evolutionary constraint.

Base Pairing↗

A fast and flexible statistical model for large-scale population genotype data: applications to inferring missing genotypes and haplotypic phase.

We present a statistical model for patterns of genetic variation in samples of unrelated individuals from natural populations. This model is based on the idea that, over short regions, haplotypes in a population tend to cluster into groups of similar haplotypes. To capture the fact that, because of recombination, this clustering tends to be local in nature, our model allows cluster memberships to change continuously along the chromosome according to a hidden Markov model. This approach is flexible, allowing for both "block-like" patterns of linkage disequilibrium (LD) and gradual decline in LD with distance. The resulting model is also fast and, as a result, is practicable for large data sets (e.g., thousands of individuals typed at hundreds of thousands of markers). We illustrate the utility of the model by applying it to dense single-nucleotide-polymorphism genotype data for the tasks of imputing missing genotypes and estimating haplotypic phase. For imputing missing genotypes, methods based on this model are as accurate or more accurate than existing methods. For haplotype estimation, the point estimates are slightly less accurate than those from the best existing methods (e.g., for unrelated Centre d'Etude du Polymorphisme Humain individuals from the HapMap project, switch error was 0.055 for our method vs. 0.051 for PHASE) but require a small fraction of the computational cost. In addition, we demonstrate that the model accurately reflects uncertainty in its estimates, in that probabilities computed using the model are approximately well calibrated. The methods described in this article are implemented in a software package, fastPHASE, which is available from the Stephens Lab Web site.

Calibration↗

An embedding-based framework enables statistical testing of gene-set function hypotheses inferred by large language models.

Emerging large language models (LLMs) can infer gene functions directly from gene lists, enabling hypothesis generation without predefined gene sets. However, these LLM-derived predictions are qualitative, and principled statistical validation is lacking. Here, we develop an embedding-based statistical framework that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs. We benchmark seven state-of-the-art embedding models using curated and retrieval-augmented literature-derived gene descriptions across diverse biological contexts. OpenAI's text-embedding-3-large and Google's gemini-embedding-001 perform best, capturing gene-gene functional relationships in 88.7-92.5% of Gene Ontology biological processes and approximately 98.6% of canonical pathways. In gene-function association analyses, these models achieve high sensitivity (95.2-98.4%) and specificity (72.7-84.3%). Through contamination analysis and evaluation using experimentally informed protein assembly gene sets, our framework distinguishes biologically meaningful LLM-inferred hypotheses from noise, outperforming confidence-based inference and conventional enrichment analysis. We further develop the open-source R package DEGEmbedR and demonstrate its utility for interpreting a drug perturbation-derived differentially expressed gene (DEG) signature lacking significant conventional enrichment results. Together, these results establish LLM-derived embeddings as a quantitative foundation for functional genomics and the statistical validation of LLM-based gene function inference.

Large Language Models↗

Connective molecular pathways of experimental bladder inflammation.

Inflammation is an inherent response of the organism that permits its survival despite constant environmental challenges. The process normally leads to recovery from injury and to healing. However, if targeted destruction and assisted repair are not properly phased, chronic inflammation can result in persistent tissue damage. To better understand the inflammatory process, we recently introduced a profiling methodology to identify common genes involved in bladder inflammation. The method represents a complementation to the classic quantification of inflammation and provides information regarding the early, intermediate, and late events in gene regulation. However, gene profiling fails to describe the molecular pathways and their interconnections involved in the particular inflammatory response. The present work introduces a new statistical technique for inferring functional interconnections between inflammatory pathways underlying classic models of bladder inflammation and permits the modeling of the inflammatory network. This new statistical method is based on variants of cluster analysis, Boolean networking, differential equations, Bayesian networking, and partial correlation. By applying partial correlation analysis, we developed mosaics of gene expression that permitted a global visualization of common and unique pathways elicited by different stimuli. The significance of these processes was tested from both biological and statistical viewpoints. We propose that connective mosaic may represent the necessary simplification step to visualize cDNA array results.

Animals↗

Simultaneous inference for generalized linear models with unmeasured confounders.

Tens of thousands of simultaneous hypothesis tests are routinely performed in genomic studies to identify differentially expressed genes. However, due to unmeasured confounders, many standard statistical approaches may be substantially biased. This paper investigates the large-scale hypothesis testing problem for multivariate generalized linear models in the presence of confounding effects. Under arbitrary confounding mechanisms, we propose a unified statistical estimation and inference framework that harnesses orthogonal structures and integrates linear projections into three key stages. It begins by disentangling marginal and uncorrelated confounding effects to recover the latent coefficients. Subsequently, latent factors and primary effects are jointly estimated through lasso-type optimization. Finally, we incorporate projected and weighted bias-correction steps for hypothesis testing. Theoretically, we establish the identification conditions of various effects and non-asymptotic error bounds. We show effective Type-I error control of asymptotic-tests as sample and response sizes approach infinity. Numerical experiments demonstrate that the proposed method controls the false discovery rate by the Benjamini-Hochberg procedure and is more powerful than alternative methods. By comparing single-cell RNA-seq counts from two groups of samples, we demonstrate the suitability of adjusting confounding effects when significant covariates are absent from the model.

Hidden variables↗