Search PubMedSearch

SEARCH · Search PubMed

Results for “Models, Dental”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

721 records · Page 3Linked to original sources

Proteomic characterization of the acquired enamel pellicle under acidic challenges at early and mature formation stages.

OBJECTIVES: This study aimed to characterize acquired enamel pellicle (AEP) proteomic changes after exposure to citric acid (CA) and hydrochloric acid (HCl) under different pellicle formation times (3 and 120&#x202f;min) in the same volunteers. DESIGN: Nine healthy volunteers participated in this randomized crossover in vivo study. The AEP was allowed to form for 3 or 120&#x202f;min and subsequently exposed for 10&#x202f;s to deionized water (control), 1% CA (pH 2.5), or 0.01&#x202f;M HCl (pH 2.0). Pellicle samples were collected, followed by protein extraction, tryptic digestion, and analysis by nanoliquid chromatography (nanoLC) coupled to mass spectrometry (MS) with MSE (data-independent acquisition; nanoLC-MS&#x1d31;). Label-free quantitative proteomics were performed for relative quantification using t-test (p&#x202f;<&#x202f;0.05). RESULTS: At 120&#x202f;min, CA exposure markedly reduced several typical AEP proteins, especially acidic proline-rich proteins (PRPs). Conversely, basic PRPs were upregulated, suggesting acid-resistance protein signature. At 3&#x202f;min, basal-layer proteins (PRPs, cystatins, histatins and mucins) were more abundant. Hemoglobins increased 6-8-fold (up to 150-fold in 3&#x202f;min control), suggesting association with early pellicle formation and an acid-resistant protein signature. CA exposures for 120&#x202f;min also upregulated typical AEP proteins (PRPs, mucins, cystatins, immunoglobulins), while HCl exposure depleted albumins and lactotransferrin. CONCLUSION: Intrinsic and extrinsic acids induce distinct proteomic signatures in the AEP. Hemoglobin and PRPs appear consistently enriched in the early pellicle layer, reflecting an initial acid-resistant protein signature. These findings provide new insights into the molecular remodeling of the AEP following intrinsic and extrinsic acid exposure, highlighting proteins potentially involved in early-stage pellicle formation.

Humans

Flap Versus Tunneling for Horizontal Ridge Augmentation With FDBA and i-PRF: A Randomized Controlled Clinical Trial.

AIM: This study evaluated the efficacy of conventional flap and tunneling techniques for horizontal alveolar ridge augmentation using freeze-dried bone allograft (FDBA) particles combined with injectable platelet-rich fibrin (i-PRF). MATERIALS AND METHODS: Forty-five patients were randomly allocated to one of three groups (n&#x2009;=&#x2009;15 each): conventional flap (CF), tunneling with membrane (TM), or tunneling without membrane (TnM). Preoperative ridge width was measured via cone beam computed tomography (CBCT). All augmentation procedures incorporated FDBA and i-PRF; an absorbable collagen membrane was applied in the CF and TM groups. Follow-up assessments, including CBCT imaging and histomorphometric analysis, were conducted 6&#x2009;months postoperatively. For normally distributed data, ANOVA with Tukey's post hoc test and paired samples t-test were applied. Non-normally distributed data were analyzed using Kruskal-Wallis, Mann-Whitney U, and Wilcoxon signed-rank tests. RESULTS: Statistical analysis was performed on 43 patients. All groups demonstrated an increase in ridge width after 6&#x2009;months. At the 2&#x2009;mm level, the mean width gain was 1.28&#x2009;mm (95% CI: 0.17 to 2.40) in the TM group, 2.85&#x2009;mm (95% CI: 1.80 to 3.89) in the TnM group, and 1.95&#x2009;mm (95% CI: 1.07 to 2.83) in the CF group. However, statistical analysis revealed no significant intergroup variation (p&#x2009;>&#x2009;0.05). Histomorphometric assessments similarly demonstrated comparable outcomes across all groups, with no statistically significant differences observed (p&#x2009;>&#x2009;0.05). CONCLUSION: Within the limitations of this study, the tunneling technique, regardless of membrane use, appears to be a clinically viable alternative to the conventional flap method for horizontal alveolar ridge augmentation. However, further studies with longer follow-up periods are required to substantiate these findings. TRIAL REGISTRATION: irct.behdasht.gov.ir identifier: IRCT 20101204005305N21.

Humans

Outcomes and Associated Prognostic Factors for Orthograde Canal Obturation Using Ortho MTA III: A Randomised Prospective Clinical Trial.

AIM: To prospectively compare treatment outcomes for orthograde canal obturation using Ortho MTA III (OMTA) with the continuous wave of compaction (CWC) using gutta-percha (GP) and AH Plus sealer, and to identify associated predictive factors. METHODOLOGY: Informed consent was obtained (110 patients), and single- or two-rooted permanent teeth (n&#x2009;=&#x2009;120) diagnosed with pulp necrosis (or previously treated) and asymptomatic apical periodontitis or chronic apical abscess (periapical index, PAI&#x2009;&#x2265;&#x2009;3) were randomly assigned to two groups (n&#x2009;=&#x2009;60/group). The canals were prepared to a minimal apical size #40 (ISO) based on their initial file size, disinfected and obturated by either CWC or OMTA using an enhanced disinfection protocol (GP disinfected, new gloves after each intraoperative radiograph and before starting obturation). Clinical and periapical radiographic examinations were conducted by two calibrated, independent endodontists during follow-up periods of at least 12&#x2009;months. Success rates and associated predictive factors (tooth-, operator- and patient-related) were analysed statistically using binary and multiple logistic regression (p&#x2009;<&#x2009;0.05). RESULTS: The median recall period was 30&#x2009;months (14-48&#x2009;months), and 104 teeth were finally analysed (recall rate: 86.67%). No significant differences in success rate were observed between the groups (p&#x2009;>&#x2009;0.05) under both loose (OMTA: 88.24%, CWC: 83.02%) and strict criteria (OMTA: 64.71%, CWC: 58.49%). Multivariate analysis revealed that age (OR&#x2009;=&#x2009;5.735, 95% CI: 1.286-25.577, p&#x2009;=&#x2009;0.022), periapical lesion size (OR&#x2009;=&#x2009;6.596, 95% CI: 1.397-31.138, p&#x2009;=&#x2009;0.017) and PAI score (OR&#x2009;=&#x2009;2.081, 95% CI: 1.047-4.136, p&#x2009;=&#x2009;0.036) were significant predictors of treatment failure. CONCLUSIONS: Orthograde obturation of infected canals with Ortho MTA III demonstrated comparable success and treatment outcomes to those filled with GP and sealer by CWC, supporting its potential as a clinically viable alternative for the obturation of infected root canals. TRIAL REGISTRATION: cris.nih.go.kr registration number: KCT00099939.

Humans

Statistical test to compare the linkage model and the admixture model based on central limit results.

In the Admixture Model, the probability that an individual carries a certain allele at a specific marker depends on the allele frequencies in K ancestral populations and the proportion of the individual's genome originating from these populations. The markers are assumed to be independent. The Linkage Model is a Hidden Markov Model that extends the Admixture Model by incorporating linkage between neighboring loci. We prove consistency and asymptotic normality of maximum likelihood estimators for the ancestry of individuals in the Linkage Model, complementing earlier results by (Pfaff et al., 2004; Pfaffelhuber and Rohde, 2022; Heinzel, 2025) for the Admixture Model. These results are used to prove that a statistical test that allows for model selection between the Admixture Model and the Linkage Model is an asymptotic level-&#x3b1;-test. Finally, we demonstrate the practical relevance of our results by applying the test to real-world data from The 1000 Genomes Project Consortium (2015).

Genetic Linkage

Modelling the effects of biological intervention in a dynamical gene network.

Cellular response to environmental and internal signals can be modeled by dynamical gene regulatory networks (GRN). In the literature, three main classes of gene network models can be distinguished: (1) non-quantitative (or data-based) models which do not describe the probability distribution of gene expressions; (2) quantitative models which fully describe the probability distribution of all genes co-expression; and (3) mechanistic models which allow for a causal interpretation of gene interactions. We propose two rigorous frameworks to model gene alteration in a dynamical GRN, depending on whether the network model is quantitative or mechanistic. We explain how these models can be used for design of experiment, or, if additional alteration data are available, for validation purposes or to improve the parameter estimation of the original model. We apply these methods to the Gaussian graphical model, which is quantitative but non-mechanistic, and to mechanistic models of Bayesian networks and penalized linear regression.

Gene Regulatory Networks

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

An integrated multiscale air quality modelling framework for industrial park pollution: Linking local emissions to regional transport.

Capturing the spatiotemporal distribution of pollutants in industrial parks remains challenging for regional air quality models because of their coarse resolution (3 km), resulting in uncertainties in local emission quantification. To address this, we developed the Integrated Multiscale Air Quality Modelling System for Industry (IAQMS-Industry), coupling the regional Nested Air Quality Prediction Modelling System (NAQPMS) with a city-scale chemical transport model. This framework integrates point-source locations and Gaussian plume dispersion to simulate particulate matter with a diameter smaller than 2.5 micrometres (PM2.5) at 100 m resolution. Applied to the Beijing Yi Zhuang and Tangshan industrial parks and evaluated against observations. The coupled model achieved a normalized mean bias (NMB) ranging from 3.1 % to 6.2 %, improving upon NAQPMS (-16.9 % to -7.7 %). Spatial analysis revealed that coarse regional grids underestimated the PM2.5&#x200b; concentrations at industrial sites by smoothing gradients, whereas IAQMS-Industry successfully resolved spatial patterns. Industrial point emissions accounted for 22.9 %-26.4 % of PM2.5 in the coupled model, which was significantly greater than the regional model estimates of 1.6 %-13.7 %. These findings indicate that regional models overestimate pollutant dispersion processes in industrial parks while underestimating local industrial impacts. By explicitly resolving point-source dynamics and linking them to regional transport, IAQMS-Industry provides a robust tool for designing targeted emission controls in industrial cities and balancing local air quality improvements with minimized regional pollution outflow. This study underscores the necessity of multiscale modelling for accurate source apportionment and informed environmental governance in industrial zones.

Air Pollution

Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.

Mixed-type outcome data occur when the outcome variable's distribution is a mixture of both continuous and discrete ordinal variables. Such mixed-type outcomes are common in biomedical, psychological, and the health sciences, particularly for variables having either a detection or quantitation limit. When interest lies in identifying a combination of genomic features associated with a mixed-type outcome, any method used would require a variable selection strategy for high-dimensional data. Unfortunately, few variable selection methods exist for modeling a mixed-type outcome when the covariate space is high dimensional. This study develops a high-dimensional penalized cumulative probability model (CPM), to allow for the identification of genomic features associated with mixed-type outcome of interest. We demonstrated how such model may be estimated using the iterative penalization procedure-the generalized monotone incremental forward stagewise (GMIFS) algorithm. The Model-X knockoffs procedure was combined with the estimation algorithm to control the false discovery rates (FDR) when performing variable selection. Through extensive simulation studies, our penalized CPM was shown to outperform alternative methods in terms of controlled variable selection performance by achieving high statistical power with the FDR being controlled at the target level. We demonstrate the utility of our method by applying it to predict estimated glomeruli filtration rate (eGFR) in kidney transplant recipients at 24&#x2009;months post-transplant using baseline gene expression data as predictors. Our CPM model identified five genes associated with this mixed-type outcome which have important links to renal disease, which may provide prognostic guidance for kidney transplantation recipients.

Models, Statistical

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Revealing the Shared Genetic Architecture of Metabolic Dysfunction-Associated Steatotic Liver Disease-Related Traits Through Genomic Structural Equation Modeling.

Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.

Humans

Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases.

PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.

Humans

Rational design of high-productivity perfusion processes for CHO Cells: From growth inhibitory strategies to model-driven optimization.

While perfusion culture for Chinese hamster ovary (CHO) cells offers advantages such as continuous operation and flexibility, it suffers from product loss through cell bleeding and difficulties in reaching high productivity due to sustained rapid cell growth. Growth inhibitory strategies are widely used to enhance productivity in fed&#x2011;batch processes; however, their practical implementation and comparative effectiveness in perfusion processes remain insufficiently explored. Meanwhile, process development often relies on costly trial&#x2011;and&#x2011;error approaches. Here, we systematically compared three growth inhibitory strategies in perfusion culture-low cell&#x2011;specific perfusion rate (CSPR), sodium butyrate, and mild hypothermia-with respect to cell growth, metabolism, productivity, and product quality. Genome&#x2011;scale metabolic flux sampling analysis revealed that low&#x2011;CSPR and sodium butyrate induce a convergent up&#x2011;regulation of energy metabolism, correlating with greater gains in specific productivity (qp). Building on this insight, we developed a growth&#x2011;kinetic model for the combined low&#x2011;CSPR + butyrate strategy, incorporating parameter uncertainty. This model&#x2011;guided framework enabled the rational design of two distinct high&#x2011;productivity perfusion processes: a sustained mode that achieved robust long&#x2011;term stability alongside substantial productivity gains, and a high&#x2011;intensity mode that pushed qp and daily volumetric titer to their maxima, with increases of up to 108.94% and 190.36%, respectively, in a model CHO cell line with a moderate baseline productivity. Our study provides a proof&#x2011;of&#x2011;concept framework for perfusion intensification, from strategy selection to rational process design.

Animals

Risk prediction models for blood transfusion in patients undergoing total hip and knee arthroplasty: a systematic review and meta-analysis.

OBJECTIVE: To systematically review and evaluate published risk prediction models for perioperative blood transfusion in patients undergoing total hip or knee arthroplasty (THA/TKA). METHODS: We systematically searched PubMed, Web of Science, the Cochrane Library, and Embase from inception to May 31, 2025. Two researchers independently screened the literature, extracted data, and assessed the risk of bias and applicability using the Prediction model Risk Of Bias Assessment Tool (PROBAST). The area under the receiver operating characteristic curve (AUC) values were pooled via a meta-analysis using Stata 18.0. RESULTS: d Fourteen studies containing 36 prediction models were included. The incidence of blood transfusion among THA/TKA patients ranged from 3.2% to 30.8%. Preoperative hemoglobin (Hb) level, tranexamic acid (TXA) use, operative duration, intraoperative blood loss, and age were the most frequently incorporated predictors. Model sensitivity ranged from 58% to 94.5%, and specificity ranged from 71.3% to 94%. Meta-analysis showed that the pooled AUC value of the 13 validated models was 0.87 (95% CI: 0.85-0.90), suggesting good discriminatory performance. All models were rated as having a high risk of bias. The applicability of four studies was rated as unclear. CONCLUSION: Although the included studies demonstrated promising discriminative ability of prediction models for blood transfusion in THA/TKA, all were assessed as having a high risk of bias using the PROBAST tool. Therefore, future research should prioritize the development of models with larger sample sizes, rigorous study designs, and multicenter external validation.

Humans

Development and Validation of a Predictive Model for Identification of Cognitive Impairment Risk in Older Adults with Subjective Cognitive Decline&#xff1a;A Longitudinal Study.

BACKGROUND: Subjective cognitive decline (SCD) is a transitional state between objective cognitive impairment and cognitively intact mental status, providing a critical window for implementing preventive interventions to delay objective cognitive decline. AIMS: We aimed to develop a predictive model for SCD progression in older adults with mild cognitive impairment (MCI). This model will facilitate the identification of risk factors and establishment of targeted interventions for community-based SCD management. METHODS: Data from the China Health and Retirement Longitudinal Study (CHARLS) was utilized in this study, extracting 18 indicators. Potential predictors selected through univariate Cox regression and LASSO regression analyses were sequentially incorporated into a multivariable Cox regression model. A nomogram was constructed to establish a predictive model. Model validation encompassed Area Under Curve (AUC) metrics for discriminative capacity, complemented by quantitative assessments using calibration curve analysis for precision verification and decision curve analysis (DCA) for clinical utility evaluation. RESULTS: A total of 1099 older adults with SCD were included in the final analysis, of whom 114 (10.3%) developed MCI. Multivariable Cox regression identified residence, marital status, educational level, social participation, gait speed, and baseline cognitive function. The model demonstrated time-dependent AUC values of 0.885, 0.830, 0.839, and 0.836 in the training set when evaluating discriminative capacity at 2-, 4-, 7-, and 9-year, respectively. The predictive model showed excellent predictive ability according to AUC, calibration curve, and DCA. CONCLUSIONS: A predictive model was created to estimate the risk of developing MCI in older individuals with SCD, offering clinician-actionable intervention benchmarks for preventive care.

Humans

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29&#x2009;709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85)&#x2009;and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans