Search PubMedSearch

SEARCH · Search PubMed

Results for “Statistical Bootstrap”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Statistical methods for analyzing developmental toxicity data.

A description and review of methods for performing per-litter analyses involving extrabinomial proportion response is provided. It is stressed that the litter should be regarded as the appropriate experimental unit for quantitative analysis in studies for teratogenic or heritable mutagenic effects. Attention is directed at statistical identification of possible treatment effects, such as a positive dose response to a chemical stimulus. The methods range from distribution-free, nonparametric analyses to models involving parametric distributions such as the beta-binomial density. It is seen that most current methods require computer implementation. When concern is raised over misspecification of assumptions critical to the statistical analysis, it is argued that relatively parameter-free methods are appropriate for use. These include statistical bootstrapping and rank-based analyses.

Models, Statistical

The evolutionary relationships among known life forms.

Sequences of small subunit (SSU) and large subunit (LSU) ribosomal RNA genes from archaebacteria, eubacteria, and the nucleus, chloroplasts, and mitochondria of eukaryotes have been compared in order to identify the most conservative positions. Aligned sets of these positions for both SSU and LSU rRNA have been used to generate tree diagrams relating the source organisms/organelles. Branching patterns were evaluated using the statistical bootstrapping technique. The resulting SSU and LSU trees are remarkably congruent and show a high degree of similarity with those based on alternative data sets and/or generated by different techniques. In addition to providing insights into the evolution of prokaryotic and eukaryotic (nuclear) lineages, the analysis reported here provides, for the first time, an extensive phylogeny of the mitochondrial lineage.

Base Sequence

Intercorrelations of regional cerebral glucose metabolic rates in Alzheimer's disease.

Patterns of cerebral metabolic correlations were compared between 21 Alzheimer's disease patients and 21 healthy age-matched controls in the resting state. Cerebral metabolic rates for glucose were determined by positron emission tomography using [18F]2-fluoro-2-deoxy-D-glucose. Partial correlation coefficients, controlling for whole brain glucose metabolism, were evaluated between pairs of regional glucose metabolic rates in 59 brain regions. Reliable correlation coefficients were obtained with the 'jackknife' and 'bootstrap' statistical procedures. Compared with healthy controls, the Alzheimer patients had significantly fewer reliable partial correlation coefficients between frontal and parietal lobe regions, and more reliable correlations between the cerebellum and temporal lobe. The number of reliable correlations between many bilaterally symmetric brain regions was reduced in the Alzheimer patients, as compared with controls. These results suggest that in the early stages of Alzheimer's disease there is a breakdown of the organized functional activity between the two cerebral hemispheres, and between parietal and frontal lobe structures.

Aged

Statistical properties of bootstrap estimation of phylogenetic variability from nucleotide sequences: II. Four taxa without a molecular clock.

The statistical properties of sample estimation and bootstrap estimation of phylogenetic variability from a sample of nucleotide sequences were studied by considering model trees of three taxa with an outgroup. The cases of constant and varying rates of nucleotide substitution were compared. From sequences obtained by simulation, phylogenetic trees were constructed by using the maximum parsimony (MP) and neighbor-joining (NJ) methods. The effectiveness and consistency of the MP method were studied in terms of proportions of informative sites. The results of simulation showed that bootstrap estimation of the confidence level for an inferred phylogeny can be used even under unequal rates of evolution if the rate differences are not large so that the MP method is not misleading. The condition under which the MP method becomes misleading (inconsistent) is more stringent for slowly evolving sequences than for rapidly evolving ones, and it also depends on the length of the internal branch. If the rate differences are large so that the MP method becomes consistently misleading, then bootstrap estimation will reinforce an erroneous conclusion on topology. Similar conclusions apply to the NJ method with uncorrected distances. The NJ method with corrected distances performs poorly when the sequence length is short but can avoid the inconsistency problem if the sequence length is long and if the distances can be estimated accurately.

Base Sequence

Statistical properties of bootstrap estimation of phylogenetic variability from nucleotide sequences. I. Four taxa with a molecular clock.

The statistical properties of sample estimation and bootstrap estimation of phylogenetic variability from a sample of nucleotide sequences are studied by using model trees of three taxa with an outgroup and by assuming a constant rate of nucleotide substitution. The maximum-parsimony method of tree reconstruction is used. An analytic formula is derived for estimating the sequence length that is required if P, the probability of obtaining the true tree from the sampled sequences, is to be equal to or higher than a given value. Bootstrap estimation is formulated as a two-step sampling procedure: (1) sampling of sequences from the evolutionary process and (2) resampling of the original sequence sample. The probability that a bootstrap resampling of an original sequence sample will support the true tree is found to depend on the model tree, the sequence length, and the probability that a randomly chosen nucleotide site is an informative site. When a trifurcating tree is used as the model tree, the probability that one of the three bifurcating trees will appear in > or = 95% of the bootstrap replicates is < 5%, even if the number of bootstrap replicates is only 50; therefore, the probability of accepting an erroneous tree as the true tree is < 5% if that tree appears in > or = 95% of the bootstrap replicates and if more than 50 bootstrap replications are conducted. However, if a particular bifurcating tree is observed in, say, < 75% of the bootstrap replicates, then it cannot be claimed to be better than the trifurcating tree even if > or = 1,000 bootstrap replications are conducted. When a bifurcating tree is used as the model tree, the bootstrap approach tends to overestimate P when the sequences are very short, but it tends to underestimate that probability when the sequences are long. Moreover, simulation results show that, if a tree is accepted as the true tree only if it has appeared in > or = 95% of the bootstrap replicates, then the probability of failing to accept any bifurcating tree can be as large as 58% even when P = 95%, i.e., even when 95% of the samples from the evolutionary process will support the true tree. Thus, if the rate-constancy assumption holds, bootstrapping is a conservative approach for estimating the reliability of an inferred phylogeny for four taxa.

Phylogeny

Identifying Single-Cell Expression Quantitative Trait Loci Using a Bootstrap Penalized Hurdle Model.

BACKGROUND: Expression quantitative trait loci (eQTL) analysis links genetic variants to gene expression levels, helping to uncover how genetic variation contributes to gene regulation. While traditional eQTL analyses rely on bulk RNA-seq data, recent advances in single-cell RNA sequencing (scRNA-seq) have made it possible to detect cell-type-specific eQTLs. However, the inherent sparsity and heterogeneity of scRNA-seq data present major challenges for standard modeling approaches. METHODS: In this paper, we propose a novel statistical framework, Bootstrap Penalized Hurdle regression model (BPHurdle), designed specifically for scRNA-seq data. BPHurdle employs a hurdle modeling framework, where a logistic component accounts for the excess zeros in single-cell expression data, and a Poisson component jointly evaluates the effects of multiple SNPs on positive gene expression levels. RESULTS: Through simulation studies, we show that BPHurdle achieves high accuracy and robustness in identifying regulatory variants. We further demonstrate its utility on a real dataset through a case study focusing on a subset of differentially expressed genes, where it successfully identifies reliable cell-type-specific eQTLs. CONCLUSIONS: Overall, BPHurdle offers an advanced and flexible approach for single-cell eQTL mapping, providing deeper insight into the genetic regulation of gene expression at cellular resolution.

Quantitative Trait Loci

POISE: Spectral Inference of Parent-of-Origin Effects in Unlabeled Genomic Data.

MOTIVATION: Parent of Origin Effects (POEs), where the effect of an an allele on a phenotype differs based on maternal or paternal inheritance implicated in growth, metabolism, and neurodevelopment. Traditional tests for POEs require family data to determine parental origins of transmitted alleles. Given that such studies are expensive and time consuming compared to genome-wide association studies (GWAS), tests that function absent inheritance information are highly desirable. We develop a method, based on community detection from machine learning, that infers POEs via a spectral decomposition, obtains confidence intervals via a non-parametric bootstrap, and safeguards against confounding by non POE sources of variation. We refer to our method as Parent of Origin Inference via Spectral Estimation (POISE). RESULTS: We demonstrate that POISE is well-calibrated under both Gaussian and heavy-tailed noise in simulation studies, with improved robustness to true POEs compared to existing covariance-based tests. POISE provides per-trait effect estimates with bias-corrected bootstrap confidence intervals and incorporates an information-theoretic minimum detectable effect size that filters unreliable estimates, conferring robustness to covariance-deflating variance QTL. We then apply POISE to GWAS data from the UK Biobank using BMI, LDL cholesterol, and HDL cholesterol. POISE recovers established POE loci and identifies 134 additional variants at genes implicated in lipid metabolism, immune regulation, and growth. AVAILABILITY AND IMPLEMENTATION: The code for this method in Python is available at https://github.com/bystrogenomics/POISE.

Community Detection

Family-Wise Error Rate Control in Clinical Trials With Overlapping Populations.

We consider clinical trials with multiple, overlapping patient populations that test multiple treatment policies specifically tailored to these populations. Such designs may lead to multiplicity issues, as false statements will affect several populations. For type I error control, often the family-wise error rate (FWER) is controlled, which is the probability to reject at least one true null hypothesis. If the joint distribution of the test statistics is known, the FWER level can be exhausted by determining critical values or adjusted-levels. The adjustment is typically done under the common ANOVA assumptions. However, the performed tests are then only valid under the rather strong assumption of homogeneous null effects, that is, when the null hypothesis applies to all subpopulations and their intersections. We show that under cancelling null effects, when heterogeneous effects cancel out in some or all subpopulations, this procedure does not provide FWER control. We also suggest different alternatives and compare them in terms of FWER control and their power.

Humans

Bootstrapping: applications to psychophysiology.

This paper presents the statistical technique known as the bootstrap to the general audience of psychophysiologists. The bootstrap, introduced by Efron (1979), allows data analysts to study the distribution of sample statistics that might otherwise be too complicated to consider. The technique, which requires simple calculations, involves drawing repeated samples (with replacement) from the empirical--or the actual--data distribution and then building a distribution for a statistic by calculating a value of the statistic for each sample. The bootstrap can be used to obtain confidence intervals, standard errors, and even higher moments for the statistic. It is similar to the well-known jackknife of Quenouille and Tukey. After discussing the history and theory of both the bootstrap and the jackknife, we illustrate the use of the bootstrap in the statistical analysis of correlation coefficients and the general linear model.

Algorithms

A bootstrap analysis of four in vitro short-term test performances.

The present analysis is aimed at estimating the confidence intervals of a number of association measures that describe the relationships of 4 in vitro short-term tests with rodent carcinogenicity, as well as with each other. The measures considered were: sensitivity, specificity and accuracy of the short-term tests with respect to chemical carcinogens, and performance dissimilarity indices (Hamming distances). The analysis refers to Salmonella, mouse lymphoma L5178Y cell mutation, chromosomal aberrations and sister-chromatid exchanges in Chinese hamster ovary cells, and is based on the data generated in the frame of the U.S. National Toxicology Program (NTP). It exploits the properties of a statistical technique, called bootstrap, to derive from only one sample of chemicals the variability intervals of the associations that the biological systems (mutagenicity assays and rodent carcinogenicity) would show in the 'universe' of the chemical compounds. The combination of the bootstrap technique with multivariate statistical methods pointed to a remarkable robustness and reliability of the information derived from the NTP data base, and provided descriptive insights into the data.

Animals

Statistical analysis of the extended Hansen method using the bootstrap technique.

In this study, simple bootstrap techniques are combined with the extended Hansen solubility approach to calculate biases, standard errors, and confidence limits of the partial solubility parameters and to obtain bias-corrected values for these solubility parameters. The bootstrap method is rather new in its application to problems in the pharmaceutical sciences and, therefore, is described here in some detail. This method provides measures of the statistical variation of ratios of regression coefficients without making unwarranted assumptions about data variability. The bootstrap can be used in many statistical packages such as MINITAB, SPSS, SAS, BMDP, or GLIM, all of which are widely available, and could be useful in other areas of the pharmaceutical sciences where regression analysis is employed.

Computer Simulation

Testing separate families of segregation hypotheses: bootstrap methods.

Aspects of the statistical modeling and assessment of hypotheses concerning quantitative traits in genetics research are discussed. It is suggested that a traditional approach to such modeling and hypothesis testing, whereby competing models are "nested" in an effort to simplify their probabilistic assessment, can be complimented by an alternative statistical paradigm - the separate-families-of-hypotheses approach to segregation analysis. Two bootstrap-based methods are described that allow testing of any two, possibly non-nested, parametric genetic hypotheses. These procedures utilize a strategy in which the unknown distribution of a likelihood ratio-based test statistic is simulated, thereby allowing the estimation of critical values for the test statistic. Though the focus of this paper concerns quantitative traits, the strategies described can be applied to qualitative traits as well. The conceptual advantages and computational ease of these strategies are discussed, and their significance levels and power are examined through Monte Carlo experimentation. It is concluded that the separate-families-of-hypotheses approach, when carried out with the methods described in this paper, not only possesses some favorable statistical properties but also is well suited for genetic segregation analysis.

Alleles

Assessing diagnostic tests once an optimal cutoff point has been selected.

The specificity and sensitivity of a quantitative diagnostic test depends on the chosen cutoff point. The common practice of selecting a cutoff point that maximizes the specificity plus the sensitivity, as judged from the observed test results, is studied here by simulation. Test performance is on average assessed too optimistically by this procedure--a phenomenon of importance when sample sizes are small. For example, the average positive bias is up to 15% of the test performance for sample sizes of 25. Furthermore, binomial calculated standard errors of specificity and sensitivity estimates are incorrect. A Monte Carlo statistical method--the "bootstrap procedure"--is applied to correct for bias and to estimate standard errors, including the standard error of the optimal cutoff point. Independent and paired comparisons of two diagnostic tests are also considered when optimal cutoff points have been selected. For this purpose, binomial statistical tests behave satisfactorily. Examples of power functions are presented.

Bile Acids and Salts

The surface area of monomeric proteins: significance of power law behavior.

The coefficients in a power law fit of accessible area versus molecular weight for high-resolution monomeric protein structures are assessed with respect to statistical accuracy using bootstrap analyses, and with respect to physical significance using model systems and the concept of roughness or fractal structure of the protein surface.

Models, Statistical

Quantifying lung structure. Experimental design and biologic variation in various models of lung injury.

The lung is a complex organ composed of a large number of different cell types of varying size and shape. Quantification of lung structure requires an understanding of how the distribution of specific cells and their characteristics affect the accuracy of measurement made on them and how to optimize experimental design for a morphometric study. We have studied lung structural modifications in a variety of lung injuries over the last decade. Extensive quantitative data from EM morphometric studies of pulmonary tissue have been collected. These data provide a unique opportunity to study the accuracy and efficiency of methods used to quantitate lung structure. We present and discuss novel computation-intensive methods for the estimation of biologic variability, sampling error, and measurement error. A new concept, unnested analysis of variance for stratified sampling and the use of computer-based methods for statistical analysis (the bootstrap method) and optimizing experimental design (nonlinear minimization procedure) are described in this report. Examples of experimental designs with their corresponding levels of accuracy and cost are also provided. The number of samples needed for a given level of precision is affected by the volume density of the structure being measured. The most important determinant for the overall accuracy of a morphometric study is the number of animals studied. Biologic variations between samples within an animal and among animals can vary significantly as a function of the model of injury studied.

Animals

Comprehensive evaluation of ACMG/AMP-based variant classification tools.

MOTIVATION: The American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) guidelines represent the gold standard for clinical variant interpretation. Despite the widespread adoption of ACMG/AMP guidelines, a comprehensive comparison of the software tools designed to implement them has been lacking. This represents a significant gap, as clinicians require evidence-based guidance on which tools to use in their practice. RESULTS: We benchmarked four ACMG/AMP-based tools (Franklin, InterVar, TAPES, Genebe) selected from 22 tools, and compared their performance with LIRICAL, a top-performing phenotype-driven tool, using 151 expert-curated datasets from Mendelian disorders. Selection criteria included free availability, VCF compatibility, operational reliability, and not being disease-specific. Our evaluation framework assessed top-N accuracy (N&#x2009;=&#x2009;1, 5, 10, 20, 50), retention rates, precision, recall, F1 scores, and area under the curve (AUC). Statistical validation employed bootstrap confidence intervals (n&#x2009;=&#x2009;1000) and Friedman tests. LIRICAL (68.21%) and Franklin (61.59%) demonstrated superior top-10 variant prioritization accuracy in Mendelian disorders, significantly outperforming other tools (P&#x2009;=&#x2009;.0000). Results demonstrate that tools with advanced phenotypic integration significantly outperform those relying primarily on genomic features. AVAILABILITY AND IMPLEMENTATION: All data and source code required to reproduce the findings of this study are openly available in the Code Ocean repository at https://doi.org/10.24433/CO.6562438.v1.

Software