Search PubMedSearch

SEARCH · Search PubMed

Results for “Bootstrap resampling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

32 records · Page 2Linked to original sources

Coagulation activation is associated with genomic-instability-related features in TP53-mutated AML and MDS: routine laboratory patterns beyond classical disseminated intravascular coagulation.

BACKGROUND: Disseminated intravascular coagulation (DIC) is a serious complication of acute myeloid leukemia (AML) associated with poor prognosis. In TP53-mutated AML and myelodysplastic syndrome (MDS), however, the classical ISTH criteria rarely identify overt DIC, although bleeding and thrombotic complications are well documented in acute leukaemia. We hypothesized that these patients exhibit a lower-grade, subclinical coagulation activation that is associated with the underlying genomic-instability-related features of TP53-mutant disease. METHODS: We retrospectively analyzed 107 consecutive patients with TP53-mutated AML (n = 52) or high-risk MDS (MDS, n = 55), median age 65 years, diagnosed and initially evaluated at our centre between 2018 and 2025. Seven routine coagulation markers and 46 co-mutated genes were evaluated for associations with overall survival (OS) using univariate and multivariable Cox regression, continuous dose-response modeling, and unsupervised k-means clustering. Internal validity was assessed by 1000 bootstrap resamples. RESULTS: Overt DIC according to ISTH criteria was rare (15%). Subclinical activation was common: 50% of patients had a D-dimer &#x2265;1&#xa0;&#x3bc;g/mL, 41% a fibrinogen &#x2265;4&#xa0;g/L, and 29% an INR &#x2265;1.2. In univariate analysis, D-dimer, fibrinogen, INR, prothrombin time, and activated partial thromboplastin time were each associated with OS (HR 1.33-1.38 per SD; all p < 0.05). Complex karyotype correlated with higher D-dimer (median 1.39 vs. 0.60&#xa0;&#x3bc;g/mL, p = 0.022) and fibrinogen (3.91 vs. 2.53&#xa0;g/L, p = 0.007), while TP53 variant allele frequency (VAF) showed modest positive correlations with D-dimer (&#x3c1; = 0.21), INR (&#x3c1; = 0.27), and PT (&#x3c1; = 0.27; all p < 0.05). Clustering identified three coagulation phenotypes: Silent (51%), Thrombo-inflammatory (31%), and Consumption-like (18%), showing a graded but statistically non-significant gradient in molecular features and a stepwise decline in median OS (14, 10 and 8 months; log-rank p = 0.041). After adjustment for complex karyotype, TP53 VAF, and favorable co-mutation count, the Consumption-like phenotype was associated with a non-significant increased risk (HR 1.83, 95% CI 0.92-3.65, p = 0.084), whereas favorable co-mutation pathways remained independently protective (HR 0.56, 95% CI 0.35-0.90, p = 0.016). CONCLUSION: In TP53-mutated AML/MDS, coagulation activation intensity is associated with the degree of genomic instability. The three phenotypes may add biological resolution beyond classical DIC and cytogenetic risk groups, but represent laboratory patterns rather than validated bleeding or thrombosis prediction tools. However, after accounting for genomic features, phenotypes were not independent predictors of outcome, with complex karyotype, TP53 VAF, and favorable co-mutation count driving prognosis. Because treatment intensity and other clinical confounders were not available, these survival associations are hypothesis-generating. Coagulation profiling remains inexpensive, widely accessible, and offers a practical window into disease biology that warrants prospective validation.

TP53

Genetic counseling in rare syndromes: a resampling method for determining an approximate confidence interval for gene location with linkage data from a single pedigree.

Multipoint linkage analysis is a powerful method for mapping a rare disease gene on the human gene map despite limited genotype and pedigree data. However, there is no standard procedure for determining a confidence interval for gene location by using multipoint linkage analysis. A genetic counselor needs to know the confidence interval for gene location in order to determine the uncertainty of risk estimates provided to a consultant on the basis of DNA studies. We describe a resampling, or "bootstrap," method for deriving an approximate confidence interval for gene location on the basis of data from a single pedigree. This method was used to define an approximate confidence interval for the location of a gene causing nonsyndromal X-linked mental retardation in a single pedigree. The approach seemed robust in that similar confidence intervals were derived by using different resampling protocols. Quantitative bounds for the confidence interval were dependent on the genetic map chosen. Once an approximate confidence interval for gene location was determined for this pedigree, it was possible to use multipoint risk analysis to estimate risk intervals for women of unknown carrier status. Despite the limited genotype data, the combination of the resampling method and multipoint risk analysis had a dramatic impact on the genetic advice available to consultants.

Chromosome Mapping

Machine learning-guided risk stratification in elderly AML based on genomic, immunophenotypic and therapeutic profiles.

BACKGROUND: Elderly patients with acute myeloid leukemia (AML) exhibit considerable biological and clinical heterogeneity, hindering precise prognosis. Existing prognostic systems inadequately capture the complexity of elderly AML due to their reliance on data from younger cohorts and omission of key factors like immunophenotypic markers and therapeutic profiles. This study aimed to develop and internally validate a machine learning-based prognostic model specifically tailored to elderly AML patients. METHODS: A total of 156 patients were analyzed using a two-stage modeling strategy. Clinical and genomic variables were modeled first, followed by independent analysis of immunophenotypic features. Feature selection was performed using multilayer perceptron (MLP) and random forest (RF), while multivariate Cox regression was used for final model construction. Internal validation was conducted using 1000 bootstrap iterations to assess model stability and performance. RESULTS: The model demonstrated strong predictive performance, with a concordance index (C-index) of 0.702. Time-dependent area under the curve (AUC) and calibration plots confirmed accurate prediction of 1-, 3-, and 5-year overall survival. Decision curve analysis indicated favorable net benefit across a range of threshold probabilities. Key independent prognostic factors identified included TP53 mutations, high CD13 expression, and IDH2 mutations. CONCLUSION: This model provides a robust and interpretable tool for individualized risk stratification in elderly AML. By integrating genomic, immunophenotypic, and therapeutic variables, it may help optimize treatment decisions and improve outcomes for this vulnerable population. Future efforts should focus on external validation and integration of dynamic biomarkers.

Humans

Statistical properties of bootstrap estimation of phylogenetic variability from nucleotide sequences. I. Four taxa with a molecular clock.

The statistical properties of sample estimation and bootstrap estimation of phylogenetic variability from a sample of nucleotide sequences are studied by using model trees of three taxa with an outgroup and by assuming a constant rate of nucleotide substitution. The maximum-parsimony method of tree reconstruction is used. An analytic formula is derived for estimating the sequence length that is required if P, the probability of obtaining the true tree from the sampled sequences, is to be equal to or higher than a given value. Bootstrap estimation is formulated as a two-step sampling procedure: (1) sampling of sequences from the evolutionary process and (2) resampling of the original sequence sample. The probability that a bootstrap resampling of an original sequence sample will support the true tree is found to depend on the model tree, the sequence length, and the probability that a randomly chosen nucleotide site is an informative site. When a trifurcating tree is used as the model tree, the probability that one of the three bifurcating trees will appear in > or = 95% of the bootstrap replicates is < 5%, even if the number of bootstrap replicates is only 50; therefore, the probability of accepting an erroneous tree as the true tree is < 5% if that tree appears in > or = 95% of the bootstrap replicates and if more than 50 bootstrap replications are conducted. However, if a particular bifurcating tree is observed in, say, < 75% of the bootstrap replicates, then it cannot be claimed to be better than the trifurcating tree even if > or = 1,000 bootstrap replications are conducted. When a bifurcating tree is used as the model tree, the bootstrap approach tends to overestimate P when the sequences are very short, but it tends to underestimate that probability when the sequences are long. Moreover, simulation results show that, if a tree is accepted as the true tree only if it has appeared in > or = 95% of the bootstrap replicates, then the probability of failing to accept any bifurcating tree can be as large as 58% even when P = 95%, i.e., even when 95% of the samples from the evolutionary process will support the true tree. Thus, if the rate-constancy assumption holds, bootstrapping is a conservative approach for estimating the reliability of an inferred phylogeny for four taxa.

Phylogeny

Human mesor-hypertensive chronorisk.

Twelve endocrine variables in blood from a small number of clinically healthy adult women were sampled systematically around the clock and the seasons. Pattern discrimination methods singled out certain hormone values in certain seasons as classifiers for a high vs low risk of developing diseases associated with a high blood pressure. Further evidence in support of such classifiers is obtained on data from adolescent, menstrually cycling young adults and post-menopausal women, here analyzed as pool of series, with the scope of the data from any one age group greatly extended by a resampling procedure, namely, by bootstrapping. This mathematical approach was carried out on data series around the clock and seasons on several hormones as well as systolic and diastolic blood pressure. Classifier roles were strongly supported for plasma aldosterone and thyroid stimulating hormone, originally by an analysis of variance and, in the case of aldosterone, by circannual cosinor analysis and by numerical resampling. Circannual bootstrapping, a procedure recommended for broad routine use as a safeguard for hypothesis testing, was also done for plasma cortisol, dehydro-epi-androsterone sulfate and prolactin, variables for which (parametric) analyses of variance and cosinors did not reveal any difference between groups at high and low cardiovascular risk. In these instances, bootstrapping results are tentative and await further analyses. Results show the ability of circannual bootstrapping to detect outliers. Identification of classifiers provides cost-effective endocrine checks complementing the targeted automatic monitoring of blood pressure. Circannual indices for risk evaluation are, however, costly in several ways since it takes at least a year and quite a few samples to estimate them reliably. Accordingly, we also extended the scope of previous results by the application of an added procedure for circadian bootstrapping. With circadian as well as circannual bootstrapping, we here illustrate a major potential component of a system of chrono-engineering for health maintenance. This system should start with focus on the newborn. The results on adults here analyzed are likely to be more prominent in the neonate, to the extent that they are genetic in origin, yet amenable to modification by the extra-uterine environment.

Adolescent

Simultaneous small-sample multivariate Bernoulli confidence intervals.

A technique based on the bootstrap is presented for assessing the simultaneous confidence level of k small-sample confidence intervals for multivariate Bernoulli marginal frequencies. The small-sample intervals used are those of Clopper and Pearson (1934, Biometrika 26, 404-413) and require iterative computation. To estimate the simultaneous confidence level, the multivariate Bernoulli vectors are resampled via the bootstrap and the Clopper-Pearson intervals recomputed on each pseudosample. The bootstrap estimate is then the proportion of times (computed via Monte Carlo) that all the k intervals computed by resampling contain the original sample frequencies. The technique is applied to single-sample HLA data.

Analysis of Variance

The defibrillation success rate versus energy relationship: Part II--Estimation with the "bootstrap".

Seventy or so defibrillation trials were typically attempted to determine the relationship between defibrillation success rate and energy (DSRE). Clinically, it may be desirable to estimate the DSRE relationship with fewer trials. We used the statistical resampling technique called the "bootstrap" to determine the number of defibrillation trials necessary for an accurate estimation of the DSRE relationship. The bootstrap technique assumes that the observed database is the maximum likelihood sample of the estimated population. The observed database is repeatedly resampled to produce a large bootstrap data-base and the bootstrap best estimate of a statistic is determined. DSRE data were obtained from ten dogs (20.5 +/- 1.5 kg). We bootstrapped our experimental DSRE data by two methods: (1) randomly choosing with replacement a specified number of defibrillation trials per energy; and (2) randomly choosing with replacement a specified number of defibrillation trials per bootstrap replication. For both bootstrap techniques, 100 replications were made. We performed a linear regression analysis on the bootstrap success rates and the observed success rates determined from 71.0 +/- 6.8 defibrillation attempts from each of the ten dogs. We concluded that 28 defibrillation trials are necessary to estimate the observed DSRE relationship with a correlation coefficient of 0.95.

Animals

An individualized nomogram for predicting progression-free survival in systemic anaplastic large cell lymphoma: a multicenter, retrospective, and internally validated study.

OBJECTIVES: To develop an individualized nomogram for predicting disease progression risk in systemic anaplastic large cell lymphoma (sALCL). METHODS: Independent predictors of progression-free survival (PFS) were identified using Cox regression in a multicenter retrospective cohort of 109 sALCL patients (2010-2022). These were incorporated into a three-factor nomogram, evaluated via bootstrapped internal validation (1000 resamples), ROC analysis, C-index, decision curve analysis (DCA), and clinical impact curve (CIC). RESULTS: A total of 29 PFS events occurred during a median follow-up of 31 months. Multivariable modelling selected serum &#x3b2;2-microglobulin elevation, extranodal disease, and front-line chemotherapy choice (CHOP versus CHOPE or BV+CHP) as autonomous progression drivers. Upon internal bootstrap validation, the nomogram yielded strong prognostic accuracy, achieving AUCs of 0.81, 0.85 and 0.87 for 1-, 3- and 5-year progression-free survival, alongside a corrected C-index of 0.779 (95% CI: 0.699 - 0.861). Calibration plots showed close agreement between predicted and observed outcomes, while DCA confirmed superior net clinical benefit versus conventional IPI or Ann Arbor stratification across multiple decision thresholds. CONCLUSION: This first sALCL-specific nomogram integrates clinical and treatment variables to provide personalized PFS risk estimation. While internally validated, this exploratory, observation-based tool requires external validation and recalibration in prospective cohorts before clinical implementation.

Humans

Bootstrapped confidence intervals for the Cox model using a linear relative risk form.

A linear relative risk form for the Cox model is sometimes more appropriate than the usual exponential form. The usual asymptotic confidence interval may not have the appropriate coverage, however, due to flatness of the likelihood in the neighbourhood of beta. For a single continuous covariate, we derive bootstrapped confidence intervals with use of two resampling methods. The first resamples the original data and yields both one-step and fully iterated estimates of beta. The second resamples the score and information quantities at each failure time to yield a one-step estimate. We computed the bootstrapped confidence intervals by three different methods and compared these intervals to one based on the asymptotic standard error and to a likelihood-based interval. The bootstrapped intervals did not perform well and underestimated the true coverage in most cases.

Computer Simulation

Bootstrap variance estimators for the parameters of small-sample sensory-performance functions.

The bootstrap method, due to Bradley Efron, is a powerful, general method for estimating a variance or standard deviation by repeatedly resampling the given set of experimental data. The method is applied here to the problem of estimating the standard deviation of the estimated midpoint and spread of a sensory-performance function based on data sets comprising 15-25 trials. The performance of the bootstrap estimator was assessed in Monte Carlo studies against another general estimator obtained by the classical "combination-of-observations" or incremental method. The bootstrap method proved clearly superior to the incremental method, yielding much smaller percentage biases and much greater efficiencies. Its use in the analysis of sensory-performance data may be particularly appropriate when traditional asymptotic procedures, including the probit-transformation approach, become unreliable.

Animals

Mitochondrial DNA evolution in primates: transition rate has been extremely low in the lemur.

Based on mitochondrial DNA (mt-DNA) sequence data from a wide range of primate species, branching order in the evolution of primates was inferred by the maximum likelihood method of Felsenstein without assuming rate constancy among lineages. Bootstrap probabilities for the maximum likelihood tree topology among alternatives were estimated without performing a maximum likelihood estimation for each resampled data set. Variation in the evolutionary rate among lineages was examined for the maximum likelihood tree by a method developed by Kishino and Hasegawa. From these analyses it appears that the transition rate of mtDNA evolution in the lemur has been extremely low, only about 1/10 that in other primate lines, whereas the transversion rate does not differ significantly from that of other primates. Furthermore, the transition rate in catarrhines, except the gibbon, is higher than those in the tarsier and in platyrrhines, and the transition rate in the gibbon is lower than those in other catarrhines. Branching dates in primate evolution were estimated by a molecular clock analysis of mtDNA, taking into account the rate of variation among different lines, and the results were compared with those estimated from nuclear DNA. Under the most likely model, where the evolutionary rate of mtDNA has been uniform within a great apes/human clade, human/chimpanzee clustering is preferred to the alternative branching orders among human, chimpanzee, and gorilla.

Animals

iMTSS: an integrated framework for biology- and patient-driven prognosis in myelofibrosis undergoing transplantation.

BACKGROUND: Allogeneic hematopoietic cell transplantation is the only curative treatment for myelofibrosis, but failure occurs by two mechanistically distinct routes: relapse of the neoplasm, which reflects its underlying genetics, and non-relapse mortality, which reflects whether the patient and graft tolerate the procedure. Established prognostic systems either lack molecular granularity or were derived in the non-transplant setting, and all collapse these two routes into a single survival estimate. None can indicate why an individual patient is at risk, or which class of intervention might reduce that risk. OBJECTIVE: To determine why an individual patient is at risk and to develop and validate an integrated framework that quantifies biology- and patient-driven prognosis. STUDY DESIGN: We analyzed 1,550 adults undergoing first allogeneic transplantation for primary or secondary myelofibrosis across international centers, the largest genomically annotated transplant cohort in this disease. The cohort was split into development (n=930) and validation (n=620) sets. Overall survival was modeled by Cox regression; relapse and non-relapse mortality were modeled as competing events by Fine-Gray subdistribution-hazard regression at 2 years. Discrimination was assessed by the concordance index with bootstrap confidence intervals. The molecular contribution was quantified by variance decomposition of, and robustness to the analytic choices was examined by resampling. RESULTS: A genetically defined disease-intrinsic axis, including TP53 allelic state, RAS pathway mutations, ASXL1 and driver genotype, blasts and blood counts, predicted 2 year relapse incidence (validation concordance 0.69, 95% CI 0.63 to 0.74), whereas a non-overlapping host and structural axis, including portal vein thrombosis, donor type, patients' performance status, and age predicted 2-year non-relapse mortality (0.63, 95% CI 0.59 to 0.68). The two scores shared only 3.4% of their variance, indicating that a patient's disease genetics carried almost no information about non-relapse mortality. Variance decomposition showed that TP53 allelic state alone accounted for 30% of the relapse score. Recombined, the framework discriminated overall survival (concordance 0.640, 95% CI 0.616 to 0.662) better than every established prognostic system. For proof of concept, 3 risk groups separated in the validation cohort, with 5 year survival of 72%, 58%, and 39% (P<0.001), and the models were well calibrated. CONCLUSIONS: Relapse and non-relapse mortality after transplantation for myelofibrosis are governed by distinct dimensions. Estimating both outcomes independently with genetic and clinical information, in addition to overall survival, establishes an individualized basis for transplant decision-making. The calculator is openly available (https://imtss-calculator.com).

mortality

Inference of horizontal genetic transfer from molecular data: an approach using the bootstrap.

Inconsistencies in taxonomic relationships implicit in different sets of nucleic acid sequences potentially result from horizontal transfer of genetic material between genomes. A nonparametric method is proposed to determine whether such inconsistencies are statistically significant. A similarity coefficient is calculated from ranked pairwise identities and evaluated against a distribution of similarity coefficients generated from resampled data. Subsequent analyses of partial data sets, obtained by the elimination of individual taxa, identify particular taxa to which the significance may be attributed, and can sometimes help in distinguishing horizontal genetic transfer from inconsistencies due to convergent evolution or variation in evolutionary rate. The method was successfully applied to data sets that were not found to be significantly different with existing methods that use comparisons of phylogenetic trees. The new statistical framework is also applicable to the inference of horizontal transfer from restriction fragment length polymorphism distributions and protein sequences.

Animals

Estimation and inference in pharmacokinetic models: the effectiveness of model reformulation and resampling methods for functions of parameters.

It is well known that high parameter estimate correlations and asymptotic variance estimates can cause estimation and inference problems in the analysis of pharmacokinetic models. In this paper we show that analysis of three important functions of pharmacokinetic parameters, the half-life, mean residence time, and the area under the curve, can sometimes be greatly improved by reformulating the model to address collinearity and by using the bootstrap to form confidence intervals. The resultant estimators can be more accurate than the original ones, and resultant confidence intervals can be narrower. Of the three measures, the half-life estimator is much better behaved than the estimators of mean residence time and area under the curve under collinearity, suggesting that it (or measures like it) should be used more often.

Analysis of Variance