Search PubMedSearch

SEARCH · Search PubMed

Results for “Method arbitration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2 recordsLinked to original sources

Disagreement-informed arbitration for gene regulatory network inference: A score-level meta-classifier and a diagnostic typology of inter-method conflict.

Gene regulatory network inference methods routinely disagree about individual edges, and practitioners resolve those conflicts by choosing one method or averaging them all. We ask whether the conflict can instead be arbitrated per edge. A gradient-boosted classifier is trained on the raw scores that ten inference methods-correlation-based, information-theoretic, sparse-regression and tree-ensemble, including GENIE3, GRNBoost2, CLR and ARACNe-assign to each candidate regulator-target pair, so that the weight given to each method varies from edge to edge. Across six single-cell perturbation screens spanning four cell types, arbitration improves on mean ensembling by +0.056 AUROC on Adamson and +0.083 on Shifrut under target-grouped cross-validation. The evaluation protocol turns out to matter more than the model. Edge-level cross-validation, standard in this literature, inflates apparent gains by 0.060 AUROC through target-gene leakage-comparable to the entire honest improvement. The effect is far larger for methods that represent genes implicitly: a supervised graph-attention link predictor trained on identical folds scores AUROC 0.930 under edge-level cross-validation, better than anything else we evaluate, and 0.533 once target genes are held out. Any method that parameterises genes is exposed, which covers most graph- and embedding-based approaches. A five-category typology of inter-method conflict localises where arbitration pays off, with the largest gains on edges where the methods disagree and the smallest where they already agree, while adding nothing as model input; we therefore report it as a diagnostic instrument rather than a modelling contribution. We also characterise what the ground truth measures: most perturbed genes in widely used screens are not transcription factors, and a mediation screen bounds how much of the perturbation response can be direct.

Ensemble methods

Artificial intelligence-supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs.

BACKGROUND: Most European population mammography screening programs rely on double reading with arbitration, a model that delivers mortality benefit but is increasingly challenged by radiologist workload, variable specificity, and interval cancers. Artificial intelligence (AI) is being evaluated to support or optimize these established European screening pathways. PURPOSE: To synthesize prospective or program-embedded evaluations of AI conducted within European-style population screening programs and to estimate exploratory program-level absolute risk differences (RDs) per 1000 examinations for cancer detection rate (CDR) and recall. MATERIALS AND METHODS: We performed a prespecified, focused evidence synthesis of three large studies embedded within routine population screening programs operating under European-relevant workflows: MASAI (randomized AI-supported risk triage within a national program), ScreenTrustCAD (prospective paired-reader evaluation with AI as an independent reader in a double-reading framework), and PRAIM (nationwide decision-referral implementation). Outcomes were harmonized as AI-control RDs per 1000 examinations. Random-effects pooling used Hartung-Knapp-Sidik-Jonkman models. For the paired-reader design, sensitivity analyses applied a Kish effective sample-size approach across plausible within-examination correlations (ρ = 0.3-0.8). Positive predictive value (PPV) and workflow/time outcomes were summarized descriptively. RESULTS: Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I2 ≈ 12%), consistent with a modest directional increase with borderline statistical uncertainty. The pooled recall RD was -0.6 per 1000 (95% CI -3.1 to +2.1; I2 ≈ 41-43%), indicating no consistent recall increase across screening programs. Where reported, PPV was higher with AI-supported screening. Efficiency signals included 44.3% fewer total readings in MASAI and shorter reading times for AI-normal examinations in PRAIM; in PRAIM, a program-level safety-net mechanism recovered 204 cancers that would otherwise have been missed. CONCLUSION: In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals. These findings suggestAI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution.

Humans