Search PubMedSearch

SEARCH · Search PubMed

Results for “Statistical Bootstrap”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Comprehensive evaluation of ACMG/AMP-based variant classification tools.

MOTIVATION: The American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) guidelines represent the gold standard for clinical variant interpretation. Despite the widespread adoption of ACMG/AMP guidelines, a comprehensive comparison of the software tools designed to implement them has been lacking. This represents a significant gap, as clinicians require evidence-based guidance on which tools to use in their practice. RESULTS: We benchmarked four ACMG/AMP-based tools (Franklin, InterVar, TAPES, Genebe) selected from 22 tools, and compared their performance with LIRICAL, a top-performing phenotype-driven tool, using 151 expert-curated datasets from Mendelian disorders. Selection criteria included free availability, VCF compatibility, operational reliability, and not being disease-specific. Our evaluation framework assessed top-N accuracy (N = 1, 5, 10, 20, 50), retention rates, precision, recall, F1 scores, and area under the curve (AUC). Statistical validation employed bootstrap confidence intervals (n = 1000) and Friedman tests. LIRICAL (68.21%) and Franklin (61.59%) demonstrated superior top-10 variant prioritization accuracy in Mendelian disorders, significantly outperforming other tools (P = .0000). Results demonstrate that tools with advanced phenotypic integration significantly outperform those relying primarily on genomic features. AVAILABILITY AND IMPLEMENTATION: All data and source code required to reproduce the findings of this study are openly available in the Code Ocean repository at https://doi.org/10.24433/CO.6562438.v1.

Software

Using logistic regression to estimate the adjusted attributable risk of low birthweight in an unmatched case-control study.

Other authors have shown how to estimate attributable risk based on stratification. In this paper, we show how to estimate adjusted attributable risks, standard errors, and confidence intervals from an unmatched case-control study that has population-based controls and uses the logistic regression model to estimate relative risk. We apply the method to data from a case-control study of low birthweight. The method is conceptually simple, has no assumptions beyond those of the logistic model, makes use of computer-intensive statistical techniques (the bootstrap), and extends to interactions. A Fortran computer program to carry out the computations is available from the authors upon request.

Case-Control Studies

A quantitative technique for characterizing tasks in psychophysiology studies.

A model was designed to specify the components of blood pressure (BP) in a reactivity study. The model considered six components for BP in response to tasks: the average resting BP across all subjects, a given subject's deviation from that average, the average task effect for a specific task, a given subject's deviation from the average task effect, the average effect of repeated assessments of BP within a given task, and an error term. The model and data from 71 adult men were used to estimate the components represented by averages. The variances and covariances of the components were represented by deviations from this average. Utilizing the likelihood ratio statistic with the bootstrap null distribution, the model gives a reasonable representation of the data. Alternative models were also tested; however, they fell short of representing the data well. The fitting and testing of components in models like ours may offer some guidance in the design of future reactivity studies.

Adult

Molecular phylogeny of plethodonine salamanders and hylid frogs: statistical analysis of protein comparisons.

The bootstrapping method of determining confidence in the topology of phylogenetic trees has been applied to electrophoretic protein data for two groups of amphibians: salamanders of two North American genera (Aneides and Plethodon) of the tribe Plethodontini and Holarctic hylid frogs. Some current methods of phylogenetic reconstruction for electrophoretic protein data have been evaluated by comparing the trees obtained from molecular data sets with available morphological data. Molecular data on the phylogenetic relationships of Aneides and Plethodon, data obtained from electrophoretic and immunological studies, indicate that Aneides probably was derived from western Plethodon subsequent to the separation of eastern and western Plethodon. Thus Plethodon very likely is a paraphyletic genus. The extremely low rate of morphological evolution in Plethodon compared with that in Aneides causes difficulty in indicating their evolutionary relationships taxonomically because there are no synapomorphic morphological characters that define either eastern or western Plethodon, whereas there are several for the genus Aneides. Thus molecular data alone probably indicate the evolutionary relationships of the species in these genera. Highton and Larson's (1979) arrangement of species of Plethodon into eight species groups is supported. The topologies of the unweighted pair-group method using arithmetic means (UPGMA) and distance Wagner trees were compared with independent morphological and molecular data on the relationships of the 28 plethodonine species. It was found that UPGMA trees indicate relationships that are more in agreement with other information than are those provided by distance Wagner trees. The use of the bootstrap technique indicates that the topologies of UPGMA trees are better supported statistically than are the topologies of distance Wagner trees. Moreover, different addition criteria produce a variety of distance Wagner trees with different topologies, each with several groupings that are not supported statistically. It is concluded that considerable caution should be used in interpreting the topology of distance Wagner trees. Very similar results were obtained with a second data set on 30 taxa of Holarctic hylid frogs. Trees obtained by the neighbor-joining method are more in agreement with UPGMA phenograms and other data, so this method of phylogenetic reconstruction may be useful to systematists not willing to assume constant rates of evolution.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence

The defibrillation success rate versus energy relationship: Part II--Estimation with the "bootstrap".

Seventy or so defibrillation trials were typically attempted to determine the relationship between defibrillation success rate and energy (DSRE). Clinically, it may be desirable to estimate the DSRE relationship with fewer trials. We used the statistical resampling technique called the "bootstrap" to determine the number of defibrillation trials necessary for an accurate estimation of the DSRE relationship. The bootstrap technique assumes that the observed database is the maximum likelihood sample of the estimated population. The observed database is repeatedly resampled to produce a large bootstrap data-base and the bootstrap best estimate of a statistic is determined. DSRE data were obtained from ten dogs (20.5 +/- 1.5 kg). We bootstrapped our experimental DSRE data by two methods: (1) randomly choosing with replacement a specified number of defibrillation trials per energy; and (2) randomly choosing with replacement a specified number of defibrillation trials per bootstrap replication. For both bootstrap techniques, 100 replications were made. We performed a linear regression analysis on the bootstrap success rates and the observed success rates determined from 71.0 +/- 6.8 defibrillation attempts from each of the ten dogs. We concluded that 28 defibrillation trials are necessary to estimate the observed DSRE relationship with a correlation coefficient of 0.95.

Animals

Statistical analysis of trends in urban ozone air quality.

The purpose of this paper is to demonstrate the use of some statistical methods for examining trends in ambient ozone air quality downwind of major urban areas. To this end, daily maximum 1-hr ozone concentrations measured over New Jersey, metropolitan New York City and Connecticut for the period 1980 to 1989 were assembled and analyzed. This paper discusses the application of the bootstrap method, extreme value statistics and a nonparametric test for evaluating trends in urban ozone air quality. The results indicate that although there is an improvement in ozone air quality downwind of New York City, there has been little change in ozone levels upwind of New York City during this ten-year period.

Air Pollutants

Effects of fronto-occipital artificial cranial vault modification on the cranial base and face.

Artificial reshaping of the cranial vault has been practiced by many human groups and provides a natural experiment in which the relationships of neurocranial, cranial base, and facial growth can be investigated. We test the hypothesis that fronto-occipital artificial reshaping of the neurocranial vault results in specific changes in the cranial base and face. Fronto-occipital reshaping results from the application of pads or a cradle board which constrains cranial vault growth, limiting growth between the frontal and occipital and allowing compensatory growth of the parietals in a mediolateral direction. Two skeletal series including both normal and artificially modified crania are analyzed, a prehistoric Peruvian Ancon sample (47 normal, 64 modified crania) and a Songish Indian sample from British Columbia (6 normal, 4 modified). Three-dimensional coordinates of 53 landmarks were measured with a diagraph and used to form 9 finite elements as a prelude to finite element scaling analysis. Finite element scaling was used to compare average normal and modified crania and the results were evaluated for statistical significance using a bootstrap test. Fronto-occipitally reshaped Ancon crania are significantly different from normal in the vault, cranial base, and face. The vault is compressed along an anterior-superior to posterior-inferior axis and expanded along a mediolateral axis in modified individuals. The cranial base is wider and shallower in the modified crania and the face is foreshortened and wider with the anterior orbital rim moving inferior and posterior towards the cranial base. The Songish crania display a different modification of the vault and face, indicating that important differences may exist in the morphological effects of fronto-occipital reshaping from one group to another.

British Columbia

A method for screening the quality of hospital care using administrative data: preliminary validation results.

Applying a computerized algorithm to administrative data to help assess the quality of hospital care is intriguing. As Iezzoni and colleagues point out, there are major differences of opinion as to the worth of such efforts. This article significantly advances the state of the art in using administrative data to screen for potential quality-of-care problems. In addition, this work on identifying complications of care goes well beyond the emphasis of many government organizations on hospital mortality rates. One question, however, not raised in the paper is: What is a practical upper limit to the sensitivity and specificity in comparing computerized screen results with the consensus judgments of a group of independent physicians? Advanced statistical techniques (such as bootstrapping) might be used to estimate the stability of consensus judgments by physician groups. When the judgments of two groups of physicians are compared with each other, the resulting sensitivity and specificity will not be .99! In addition, more training of members of the physician panels would probably have increased interrater reliability. While acknowledging this problem, the researchers' detailed analysis of the panel results is intriguing and represents a model for such studies. It is hoped that the authors will follow up on the avenues opened here. Furthermore, what degree of accuracy is necessary to identify facilities with higher-than-expected rates of complications? The authors discuss problems involved in using administrative data to target hospitals and departments for more costly in-depth reviews of quality. It is hoped that the promising findings that are reported here will be validated in other studies. Certainly their algorithms should find a ready audience in insurers and hospitals willing to try them out. Finally, should we expect additional research to lead to improvement in the authors' algorithms? I believe the algorithms will prove difficult to improve upon; but perhaps we should not worry about this. At some point, however, the cost of trying to identify and correct quality problems in "minimally outlier" hospitals will exceed the benefits, particularly given alternative uses for the funds. Might we now be close the the "flat of the curve" in the development of such systems for identification of quality problems? This issue should be discussed much further in future studies.

Abstracting and Indexing

[Statistical evaluation of residue data for the assessment of predicted values (half-life, withdrawal time) as an example of toltrazuril and enrofloxacin in trout].

The half-lives and withdrawal times of the veterinary drugs Toltrazuril and Enrofloxacin in trout have been assessed by statistical analysis. Confidence intervals were computed using a normal distribution of residual data and an empirical distribution by the Bootstrap method. Both methods produced similar statistics for the two drugs. Simulation of the residue data according to the regression lines of the decay curves has shown that the Bootstrap method is better for use when the residue patterns are not distributed normally. Using confidence intervals, a statistical mean of withdrawal times can be assessed. Taking into account the decays of the individual antibiotics in all treated trout, tolerance intervals for the regression lines are obtained: the calculated 10-20% longer withdrawal time includes values for which the antibiotic concentration in 95% of the treated trout is decreased below the tolerance level of Toltrazuril or Enrofloxacin.

4-Quinolones

Assessing agreement.

Formal evaluation of the ability of clinicians and researchers to agree, for example, on the clinical assessment of patients, increasingly is becoming important. Two measures of agreement, kappa and the intraclass correlation coefficient, are described and illustrated. The calculation of confidence intervals that correspond to these statistics by means of the "bootstrap" method also is discussed.

Analysis of Variance

Detection of outlying data in bioavailability/bioequivalence studies.

This paper considers the problem of detecting outlying data in bioavailability/bioequivalence studies. We define outlying subjects as those whose responses in bioavailability to all formulations differ from the rest of the subjects. We also define an outlying observation as the response in bioavailability of a subject to a particular formulation which is grossly different from the average bioavailability of that formulation calculated from all subjects. We propose two test procedures. The first, based on two-sample Hotelling T2, is to detect possible outlying subjects. The second, based on residuals from formulation means, is to identify possible outlying observations within subjects. Both procedures take into account the covariance structure of the responses to formulations, dependence of test statistics, and multiplicity of test procedures. We apply the Monte Carlo or bootstrap simulation to evaluate the sampling distributions of test statistics. An example from a 3-way crossover bioequivalence study illustrates the two procedures.

Biological Availability

The bootstrap and identification of prognostic factors via Cox's proportional hazards regression model.

This paper describes the use of the bootstrap, a new computer-based statistical methodology, to help validate a regression model resulting from the fitting of Cox's proportional hazards model to a set of censored survival data. As an example, we define a prognostic model for outcome in childhood acute lymphocytic leukemia with the Cox model and use of a training set of 224 patients. To validate the accuracy of the model, we use a bootstrap resampling technique to mimic the population under study in two stages. First, we select the important prognostic factors via a stepwise regression procedure with 100 bootstrap samples. Secondly we estimate the corresponding regression parameters for these important factors with 400 bootstrap samples. The bootstrap result suggests that the model constructed from the training set is reasonable.

Child

A rule for the early detection of chronic changes in cystic fibrosis patient status.

A statistical decision-making system has been developed which will predict the clinical status of a patient with cystic fibrosis based on daily self measurements obtained at home. The data for the study were collected from CF patients within 7-12 years of age. Thirty-two participants recorded four daily measurements (weight, vital capacity, breathing rate, and resting pulse) and one weekly measurement (height). In addition to the 4 daily measured values, the clinical status of each patient at his/her most recent previous clinic visit was used as a predictor variable. The measured values were used as the basis for the development of a discriminant rule. The goal of the rule was to determine whether each patient's clinical status was deteriorating, stable, or improving at the time of the most recent set of weekly measurements. Three types of analysis were performed: linear discriminant analysis, quadratic discriminant analysis, and nearest neighbor. Quadratic discriminant analysis provided the best discrimination due to the differences in the covariance matrices among the populations. The rule was able to correctly classify 77% of the 103 cases in the learning set. To further evaluate the rule, both a weighted classification percentage and weighted kappa statistic were calculated for the rule. Bootstrapping was used to predict the performance of the rule on the population with results of 77% correctly classified overall.

Body Weight

Psychological Capital and Perceived Stress in Nurses: The Mediating Role of Need for Recovery and Recovery Experiences.

AIM: To examine whether recovery experiences and need for recovery mediate the association between psychological capital and perceived stress in nurses. DESIGN: Cross-sectional online survey. METHODS: A total of 184 nurses currently practicing in France completed an anonymous online questionnaire administered in January 2023. Participants completed the French versions of the Psychological Capital Questionnaire, the Recovery Experience Scale, the Need for Recovery Scale, and the Perceived Stress Scale. Two mediation models were estimated using ordinary least squares regression with percentile bootstrap inference for the indirect effects. RESULTS: Psychological capital was positively associated with recovery experiences and negatively associated with need for recovery. Need for recovery was strongly and positively associated with perceived stress, while psychological capital showed no direct association with perceived stress. The indirect effect of psychological capital on perceived stress through need for recovery was statistically detectable, whereas a complementary indirect effect through recovery experiences was in the expected direction but did not meet the bootstrap criterion for statistical significance. CONCLUSION: In this cross-sectional sample, psychological capital was associated with lower perceived stress primarily through its association with reduced need for recovery rather than through a direct association with stress. Interventions intended to protect nurse well-being should therefore be conceived as combining individual resource-building with organizational architectures that enable effective recovery. IMPLICATION FOR NURSING PRACTICE: Psychological capital should be seen as one element in integrated interventions rather than a stand-alone solution to nursing stress. Brief recovery-focused interventions for nurses need to be combined with scheduling practices and organizational policies that protect rest and make recovery feasible, a configuration that appears more promising than psychological capital training alone. REPORTING METHOD: The study followed the STROBE Statement for the reporting of cross-sectional studies. NO PATIENT OR PUBLIC CONTRIBUTION: This study focused on nurses as study participants. No patient or public stakeholder was involved in the design, conduct, or interpretation of the research, since the research question concerns occupational psychological resources rather than clinical practice or care delivery.

Humans

Demonstration of the reproducibility of treatment efficacy from a single multicenter trial.

According to the Food and Drug Administration's Guidelines for the Format and Content of the Clinical and Statistical Sections of New Drug Applications, approval of a new drug "should be supported by more than one well-controlled trial and carried out by independent investigators. This interpretation is consistent with the general scientific demand for replicability." Nevius has described a four-point proposal for assessing statistical evidence in a single multicenter trial. Briefly, these four points are: (1) combined analysis shows significant results, (2) consistency over centers in terms of direction, (3) consistency over centers in terms of producing nominally significant results in centers with sufficient power, and (4) evidence of efficacy after adjustment for multiple comparisons. What is not clear from Nevius' proposal is how to quantify whether the amount of evidence in a single multicenter trial is equivalent to that from two separate trials. It is proposed that the post hoc subdivision of a multicenter trial may address this issue if the inherent multiple testing problem is accommodated. A minimax statistic is developed to test the hypothesis that the effect of the drug has been reproduced in a single multicenter trial. Monte Carlo simulation is used to generate the distribution of the minimax statistic under the null and several alternative hypotheses. Data from a multicenter trial are used to demonstrate the technique. Bootstrapping is used to determine the null distribution of the minimax statistic.

Humans

Factors underlying variation in spontaneous and clastogen-induced sister chromatid exchanges and chromosome breakage frequencies.

The latent "factors" influencing spontaneous and clastogen-induced genetic damage, measured by rates of sister chromatid exchange (SCE) and chromosome breakage (CB), were investigated in a small sample of 20 unrelated, healthy individuals. The covariation of spontaneous and clastogen-induced (bleomycin [BLM], streptonigrin [SN], mitomycin-C [MMC], 4-nitroquinoline-1-oxide [4NQO]) SCEs and CBs was analyzed by maximum-likelihood factor analysis. A single-factor model resulted in large standardized regression coefficients of measured variables on the factor for spontaneous and BLM- and SN-induced SCE frequencies, and a modest regression coefficient for MMC-induced SCEs. A two-factor model, after varimax rotation, yielded one factor strongly associated with spontaneous and BLM- and SN-induced SCE frequencies, and a second factor associated with spontaneous and BLM- and SN-induced CBs. A bootstrap analysis of this data set indicated the statistical significance of one regression coefficient (i.e., P less than or equal to 0.05) and borderline significance (0.07 less than or equal to P less than or equal to 0.11) of three other regression coefficients on the first factor, to be interpreted as an effector of SCE frequencies. However, for the second factor, none of the bootstrapped regression coefficients was significant (P greater than 0.22). Due to the modest sample utilized in this study, the validity of this model should be further explored using additional, larger data sets.

4-Nitroquinoline-1-oxide

Rate and mode differences between nuclear and mitochondrial small-subunit rRNA genes in mushrooms.

Sequences from homologous regions of the nuclear and mitochondrial small-subunit rRNA genes from 10 members of the mushroom order Boletales were used to construct evolutionary trees and to compare the rates and modes of evolution. Trees constructed independently for each gene by parsimony and tested by bootstrap analysis have identical topologies in all statistically significant branches. Examination of base substitutions revealed that the nuclear gene is biased toward C-T transitions and that the distribution of transversions in the mitochondrial gene is strongly effected by an A-T bias. When only homologous regions of the two genes were compared, base substitutions per nucleotide were roughly 16-fold greater in the mitochondrial gene. The difference in the frequency of length mutations was at least as great but was impossible to estimate accurately because of their absence in the nuclear gene. Maximum likelihood was used to show that base-substitution rates vary dramatically among the branches. A significant part of the rate inconstancy was caused by an accelerated nuclear rate in one branch and a retarded mitochondrial rate in a different branch. A second part of the rate variability involved a consistent inconstancy: short branches exhibit ratios of mitochondrial to nuclear divergences of less than 1, while longer branches had ratios of approximately 4:1-8:1. This pattern suggests a systematic error in the branch length calculation. The error may be related to the simplicity of the divergence estimates, which assumes that all base positions have an equal probability of change.

Base Composition

Detection of two-component mixtures of lognormal distributions in grouped, doubly truncated data: analysis of red blood cell volume distributions.

We have examined the statistical requirements for the detection of mixtures of two lognormal distributions in doubly truncated data when the sample size is large. The expectation-maximization algorithm was used for parameter estimation. A bootstrap approach was used to test for a mixture of distributions using the likelihood ratio statistic. Analysis of computer simulated mixtures showed that as the ratio of the difference between the means to the minimum standard deviation increases, the power for detection also increases and the accuracy of parameter estimates improves. These procedures were used to examine the distribution of red blood cell volume in blood samples. Each distribution was doubly truncated to eliminate artifactual frequency counts and tested for best fit to a single lognormal distribution or a mixture of two lognormal distributions. A single population was found in samples obtained from 60 healthy individuals. Two subpopulations of cells were detected in 25 of 27 mixtures of blood prepared in vitro. Analyses of mixtures of blood from 40 patients treated for iron-deficiency anemia showed that subpopulations could be detected in all by 6 weeks after onset of treatment. To determine if two-component mixtures could be detected, distributions were examined from untransfused patients with refractory anemia. In two patients with inherited sideroblastic anemia a mixture of microcytic and normocytic cells was found, while in the third patient a single population of microcytic cells was identified. In two family members previously identified as carriers of inherited sideroblastic anemia, mixtures of microcytic and normocytic subpopulations were found. Twenty-five patients with acquired myelodysplastic anemia were examined. A good fit to a mixture of subpopulations containing abnormal microcytic or macrocytic cells was found in two. We have demonstrated that with large sample sizes, mixtures of distributions can be detected even when distributions appear to be unimodal. These statistical techniques provide a means to characterize and quantify alterations in erythrocyte subpopulations in anemia but could also be applied to any set of grouped, doubly truncated data to test for the presence of a mixture of two lognormal distributions.

Adolescent