Search PubMedSearch

SEARCH · Search PubMed

Results for “probability calibration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

73 recordsLinked to original sources

RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware probability analysis.

Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.

Atrial Fibrillation

Risk Factors and Predictive Model for Postoperative High Myopia in Children Undergoing Congenital Cataract Surgery With Intraocular Lens Implantation.

PURPOSE: To identify risk factors associated with the development of high myopia following congenital cataract surgery and to establish a robust predictive model. DESIGN: Retrospective clinical cohort study. SUBJECTS: This retrospective study included 106 pediatric patients who underwent congenital cataract surgery with primary IOL implantation (mean follow-up 8.19 years). The model was externally validated in an independent cohort of 72 patients with a mean follow-up of 7.83 years. METHODS: Preoperative and postoperative ocular biometric parameters were collected. Risk factors for postoperative high myopia were analyzed using Cox proportional hazards regression, which served as the basis for model construction. The predictive performance of the model was rigorously evaluated for discrimination and calibration. Discriminative ability was quantified using Harrell's C-index and the area under the receiver operating characteristic curve (AUC). Model calibration was assessed via calibration plots by comparing predicted probabilities with actual observed outcomes. Internal validation was performed using a bootstrapping method (500 iterations) to ensure model stability and adjust for potential overfitting. RESULTS: An initial postoperative refraction of <+0.75D, and a higher IOL Power to Axial length Ratio (IOL/AL ratio) were identified as significant risk factors for the development of postoperative high myopia. Shorter preoperative axial length was associated with a greater magnitude of postoperative myopic shift. The predictive model demonstrated robust performance, achieving a C-index of 0.711 (internal validation C-index: 0.713). The area under the receiver operating characteristic curve (AUC) values for predicting high myopia at 5 and 10 years were 0.858 and 0.745, respectively. Furthermore, calibration curves demonstrated excellent agreement between the predicted and observed outcomes throughout the follow-up period. In external validation, the model achieved a C-index of 0.825, 5-year AUC of 0.833, and 10-year AUC of 0.713. CONCLUSIONS: Our analysis established that initial postoperative refraction <+0.75D, and an elevated IOL/AL ratio are key determinants of high myopia risk following surgery. Shorter preoperative axial length was associated with a greater magnitude of postoperative myopic shift. This predictive framework provides clinicians with a practical tool to optimize preoperative IOL selection and identify high-risk infants who require vigilant myopia prevention and balanced amblyopia management.

Humans

Beyond Photometric Consistency: Addressing Loss Insensitivity to Depth Noise in Endoscopic Estimation via Error Calibration.

Self-supervised monocular depth estimation in endoscopy is fundamentally constrained by the ill-posed nature of photometric supervision. In this work, we identify a critical yet overlooked cause of this ambiguity: the inherent insensitivity of photometric loss to depth noise. To overcome this intrinsic limitation, we propose Depth Error Calibration Learning (DECL), a two-stage framework that suppresses prediction variance and mitigates residual errors in self-supervised depth estimation. In Stage I (Variance Reduction), a cyclic depth generation strategy produces multiple depth hypotheses for the input image. The per-pixel empirical variance is quantified and integrated into a dedicated variance loss term, which penalizes inconsistent predictions and encourages the network to generate more stable and reliable depth estimates. In Stage II (Bias Calibration), an image-conditioned diffusion model refines the Stage-I depth prior and mitigates structured residuals through iterative denoising, thereby improving geometric accuracy and global consistency. Extensive experiments on three public endoscopic datasets demonstrate that DECL achieves consistent improvements over representative self-supervised monocular depth estimation methods under the evaluated protocols. Moreover, ablation studies on two representative backbones indicate that DECL is not restricted to a single network implementation, while broader validation on additional backbone families remains necessary. The source code is publicly available at https://github.com/DavidLuBit/EndoDenoising.

Journal Article

Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.

Mixed-type outcome data occur when the outcome variable's distribution is a mixture of both continuous and discrete ordinal variables. Such mixed-type outcomes are common in biomedical, psychological, and the health sciences, particularly for variables having either a detection or quantitation limit. When interest lies in identifying a combination of genomic features associated with a mixed-type outcome, any method used would require a variable selection strategy for high-dimensional data. Unfortunately, few variable selection methods exist for modeling a mixed-type outcome when the covariate space is high dimensional. This study develops a high-dimensional penalized cumulative probability model (CPM), to allow for the identification of genomic features associated with mixed-type outcome of interest. We demonstrated how such model may be estimated using the iterative penalization procedure-the generalized monotone incremental forward stagewise (GMIFS) algorithm. The Model-X knockoffs procedure was combined with the estimation algorithm to control the false discovery rates (FDR) when performing variable selection. Through extensive simulation studies, our penalized CPM was shown to outperform alternative methods in terms of controlled variable selection performance by achieving high statistical power with the FDR being controlled at the target level. We demonstrate the utility of our method by applying it to predict estimated glomeruli filtration rate (eGFR) in kidney transplant recipients at 24&#x2009;months post-transplant using baseline gene expression data as predictors. Our CPM model identified five genes associated with this mixed-type outcome which have important links to renal disease, which may provide prognostic guidance for kidney transplantation recipients.

Models, Statistical

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Efficacy of smartphone- and bibliotherapy-delivered multicomponent lifestyle medicine interventions for probable depression: A three-arm randomized controlled trial.

BACKGROUND: This study examined the efficacy of smartphone- (AG) and bibliotherapy-delivered (BG) lifestyle medicine (LM) interventions compared with a waitlist control group (WLG) in reducing depressive symptoms. METHODS: A total of 122 adults with probable depression were randomized to AG (n&#xa0;=&#xa0;41), BG (n&#xa0;=&#xa0;40), or WLG (n&#xa0;=&#xa0;41). AG and BG received the same core 8-week multicomponent LM intervention via a smartphone application or booklets, respectively. The core content included lifestyle psychoeducation, physical activity, diet and nutrition, stress and sleep management, goal-setting, and motivational techniques. Outcomes were assessed at baseline and immediate post-intervention (Week 9) in all groups, with 1-month (Week 13) and 3-month (Week 21) follow-ups conducted in the intervention groups only. RESULTS: At Week 9, AG (d&#xa0;=&#xa0;0.89) and BG (d&#xa0;=&#xa0;0.64) showed significantly greater reductions in depressive symptoms than WLG, with within-group improvements maintained at 1- and 3-month follow-ups (ps&#xa0;<.001, d&#xa0;=&#xa0;0.77-0.96). Clinically significant improvement was achieved by 78% of AG and 55% of BG participants, both significantly higher than WLG (19.5%; ps&#xa0;<.001). Compared with WLG, both interventions yielded greater improvements in overall lifestyle and physical activity (d&#xa0;=&#xa0;0.59-0.91) at Week 9. The AG showed additional benefits for perceived stress, health responsibility, nutrition, spiritual growth, and stress management (d&#xa0;=&#xa0;0.58-0.73), whereas BG uniquely improved insomnia symptoms (d&#xa0;=&#xa0;0.83). CONCLUSION: Smartphone- and bibliotherapy-delivered LM interventions are efficacious for managing probable depression. Further RCTs comparing them with established treatments are warranted.

Humans

Effect of Bariatric Surgery on Improvement of Retinal Microvasculature in Patients With Obesity: A Systematic Review and Meta-Analysis.

Obesity is associated with adverse retinal microvascular changes, including narrower central retinal arteriolar equivalent (CRAE), wider central retinal venular equivalent (CRVE) and lower arteriovenous ratio (AVR). Although bariatric surgery improves cardiometabolic risk, its effect on retinal microvascular calibre remains uncertain. We systematically searched the Cochrane Library, PubMed, ScienceDirect and Scopus up to June 2026. Prospective cohort studies reporting CRAE, CRVE or AVR before and after bariatric surgery in patients with obesity were included. Risk of bias was assessed using the Newcastle-Ottawa Scale. Pooled mean differences (MD) with 95% confidence intervals (CI) were calculated using a random-effects model. Eight prospective cohort studies including 283 patients were included. Seven studies contributed to the CRAE and CRVE analyses, and seven contributed to the AVR analysis. Bariatric surgery was associated with a significant increase in CRAE (MD&#x2009;=&#x2009;+4.02&#x2009;&#x3bc;m, CI [0.26-7.78], p&#x2009;=&#x2009;0.04). Sensitivity analysis excluding one study showed a stronger and more consistent effect (MD&#x2009;=&#x2009;+5.41&#x2009;&#x3bc;m, CI [4.59-6.22], p&#x2009;<&#x2009;0.01). CRVE significantly decreased after surgery (MD&#x2009;=&#x2009;-6.75&#x2009;&#x3bc;m, CI [-7.42 to -6.07], p&#x2009;<&#x2009;0.01), while AVR significantly increased (MD&#x2009;=&#x2009;+0.03, CI [0.02-0.05], p&#x2009;<&#x2009;0.01). Sensitivity analyses suggested that the direction of effect was generally robust, although between-study heterogeneity was present. Bariatric surgery may be associated with favourable retinal microvascular changes in patients with obesity, reflected by increased CRAE and AVR and decreased CRVE.

Humans

Microbial allies in a cotton pest: A descriptive account of associated microbiota dynamics in Dysdercus cingulatus across development.

BACKGROUND: Hemipteran insects harbour several symbiotic partners, mainly bacteria, which play pivotal roles for hosts like dietary provision, support overall physiology, xenobiotic degradation and manipulate/regulate behaviour. Most of these symbionts usually reside and operate from the digestive tracts of the animals. Cotton is one of the major cash crops in India and Dysdercus cingulatus (D. cingulatus) though a secondary pest, is causing significant destruction of cotton bolls, poor lint quality and reduce oil content of seeds. Premature opening of cotton bolls often leads to bacterial and fungal infections, thus resulting in extensive economic loss worldwide. D. cingulatus is a hemimetabolous insect that comprises of developmental stages like egg, nymph (5 instar stages), and adult. The present work explored the ontogeny specific diversity in the associated microbiota and predicted their probable functional inputs in D. cingulatus. RESULTS: The data obtained using 16S rRNA gene sequencing (NovaSeq 6000) revealed presence of members of Proteobacteria (65.83%), Firmicutes (24%), Actinobacteria (10%) phyla throughout the ontogeny of D. cingulatus. Highest alpha diversity of these symbiotic bacteria was recorded in the third instar nymphs in contrast to rest of the developmental stages. Among all the observed genera, Stenotrophomonas, Hungatella and Glutamicibacter were predominant from egg to adult stages. MicFunPred, a tool used for predicting the probable functional inputs of these symbionts, hinted at their probable stage specific contribution in crucial biochemical pathways such as polyketide biosynthesis, ascorbate/aldarate metabolism, pentose phosphate and glyoxylate cycles, steroid hormone and peptidoglycan biosynthesis, and glycolysis/pyruvate metabolism. CONCLUSIONS: The primary investigations on the ontogenetic composition and diversity of associated microbiota, suggest dynamic shifts in D. cingulatus, concurrent with their probable functions/roles in the host development and metabolism. To the best of our knowledge, this is the first report on symbiotic microbiota variation across the developmental stages of D. cingulatus that provides preliminary descriptive observations that may guide future functional and experimental investigations into microbiota-based pest management.

Animals

Could the preoperative urethral curve be used to predict immediate urinary continence following Retzius-sparing robot-assisted radical prostatectomy? A retrospective multi-center study.

PURPOSE: Immediate urinary continence (UC) recovery following Retzius-sparing robot-assisted radical prostatectomy (RS-RARP) remains highly variable, highlighting the need for reliable preoperative prediction. We aimed to develop and validate models to identify patients likely to achieve immediate UC recovery following RS-RARP. MATERIALS AND METHODS: A total of 580 prostate cancer patients who underwent RS-RARP from four medical centers were assigned to a training set (n=348), an internal validation set (n=103) and an external validation set (n=129). Independent predictors were identified through univariate analysis and LASSO regression. A nomogram was constructed using multivariate logistic regression. Its performance was evaluated with receiver operating characteristic (ROC) curve, calibration curves, and decision curve analysis. RESULTS: Immediate UC recovery was observed in 84.5% (294/348) of patients in the training cohort, 80.6% (83/103) in the internal validation cohort, and 81.4% (105/129) in the external validation cohort, respectively. Multivariate analysis identified membranous urethral length (MUL) (OR=1.23, P=0.029) and urethral curvature (OR=2.84, P<0.001) as independent predictors, while prostate volume (PV) (OR=0.84, P <0.001) as a protective factor. The nomogram integrating MUL, PV, and urethral curvature demonstrated superior predictive accuracy, with an AUC of 0.87 (95% CI, 0.83-0.91) in the training cohort. The bootstrap-corrected calibration slope was 0.96, and the Brier score was 0.08.&#xa0;Calibration curves and decision curve analysis confirmed the predictive accuracy and clinical utility of the nomogram. CONCLUSIONS: Our study introduces a novel quantitative method for assessing urethral curvature. The mpMRI-based model, integrating urethral curvature and prostate spatial configuration, offers enhanced predictive accuracy for postoperative immediate UC recovery.

Humans

AI-driven snapshot hyperspectral imaging for on-line sorting systems in food industry: From real-time sensing to intelligent decision-making.

High-throughput food sorting requires rapid, non-destructive detection of external defects, foreign materials, and internal quality attributes in heterogeneous food matrices. Conventional scanning hyperspectral imaging may suffer from motion-induced spatial-spectral mismatches, whereas snapshot hyperspectral imaging (S-HSI) captures spectral images within a single integration time. However, its advantage is limited by trade-offs in resolution, signal-to-noise ratio (SNR), reconstruction uncertainty, and calibration stability, which are further amplified by variable tissue structure, surface reflection, moisture, and fat distribution in foods. This review critically examines artificial intelligence (AI)-driven S-HSI for on-line food sorting within a sensing-representation-decision-execution framework. Compact architectures are compared according to their physical constraints, food-sorting suitability, and ability to support mapping between spectral responses and physicochemical quality attributes. AI strategies are reviewed for spectral reconstruction, image restoration, spatial-spectral representation, band selection, uncertainty-aware decision-making, and edge implementation. AI can partially compensate for snapshot-specific limitations, but current evidence remains largely limited to laboratory or prototype studies. Future work should link system performance to food safety and quality outcomes by reporting throughput, decision latency, calibration drift, missed-detection risk, false-rejection cost, and closed-loop sorting success.

Hyperspectral Imaging

A mechanism-guided framework for prioritizing membrane-interaction anti-Vibrio peptides from peptidomics data.

A mechanism-guided framework for prioritizing membrane-interaction antimicrobial peptide candidates from proteomics-derived peptide mixtures is presented. The framework integrates conservative machine-learning-based antimicrobial peptide (AMP) screening with a literature-derived membrane-interaction plausibility (MAP) assessment and a data-driven membrane-interaction ranking function (AIPx), followed by structural visualization for interpretability. MAP encodes physicochemical characteristics commonly associated with peptide-membrane interaction and provides a graded plausibility assessment. Building upon this physicochemically interpretable framework, AIPx ranks peptides using feature weights calibrated from experimentally characterized anti-Vibrio peptides, where minimum inhibitory concentration (MIC) values are used as a coarse-grained ranking reference rather than a direct prediction target. In a peptidomics-based peptide fractionation study targeting Vibrio spp., AIPx exhibited a consistent relationship with experimentally observed antibacterial activity. Distributional analysis revealed that peptide fractions exhibiting high anti-Vibrio activity are characterized by enrichment of high-ranking peptides rather than by AMP abundance alone. By structuring AMP identification and prioritization as sequential stages, the MAP&#xa0;+&#xa0;AIPx framework enables interpretable and experimentally actionable candidate selection by reducing biologically implausible candidates. The framework facilitates species-oriented prioritization of AMP candidates, addressing a key challenge in antimicrobial peptide discovery where activity may depend on target-specific membrane characteristics. Moreover, the approach is extensible through species-specific calibration and supports interpretable, mechanism-informed prioritization in antimicrobial peptide discovery.

Proteomics

Association between Kidney Tubular Secretory Clearance with Cognitive Function among Adults with CKD in the Systolic Blood Pressure Intervention Trial.

BACKGROUND: Persons with CKD are disproportionally affected with cognitive impairment, yet the pathophysiology linking the two conditions is unclear. Because kidney tubule secretion is essential for clearance of medications, uremic toxins, and metabolites, we hypothesized that worse tubular secretion would be associated with reduced cognitive function in CKD. METHODS: The Systolic Blood Pressure Intervention Trial tested a systolic blood pressure target <120 mmHg vs. <140 mmHg in hypertensive individuals at high cardiovascular risk. In paired blood and urine specimens from 1,937 participants with eGFR <60 ml/min/1.73m2, we measured 10 endogenous tubule-secreted metabolites and calculated a urine/plasma ratio for each, then averaged these to generate a summary secretion score. We used unadjusted and multivariable-adjusted linear regression and mixed models to evaluate cross-sectional and longitudinal associations of the secretion score with the Montreal Cognitive Assessment, Digit Symbol Coding, and Logical Memory immediate and delayed tests-measured at baseline and months 24 and 48 of follow-up. Multivariable Cox regression evaluated associations with incident probable dementia and mild cognitive impairment, adjudicated by prespecified criteria. RESULTS: Mean age was 73 &#xb1; 9 years, 41% were women, mean eGFR was 48.2 &#xb1; 11.4 ml/min/1.73m2, median albuminuria was 14.8 [7.1-48.6] mg/g. Lower secretion score was associated with a 0.06 higher adjusted logical memory delayed score (95% CI: 0.02, 0.11) but not with other cognitive tests at baseline or longitudinal cognitive decline. After a median 4.1 years of follow-up, 118 developed probable dementia and 187 developed mild cognitive impairment. Each 1-SD lower secretion score was associated with lower risk of probable dementia (HR 0.78, 95% CI: 0.63, 0.98) but not mild cognitive impairment. CONCLUSIONS: Among Systolic Blood Pressure Intervention Trial participants with CKD, lower estimated tubular secretion was not associated with worse cognition at baseline or during longitudinal follow-up.

Journal Article

Repeated scoring with the adult appendicitis score improves the sensitivity and the specificity of appendicitis diagnosis in patients with early equivocal signs of appendicitis: a secondary analysis.

PURPOSE: The utilization of computed tomography in the early stage of acute appendicitis may result in overdiagnosis and unnecessarily expose patients to ionising radiation. The Adult Appendicitis Score (AAS) can be used to select patients for imaging. Observation and re-scoring in the DIAMOND trial reduced the need for imaging. Now, we wanted to determine if the change in AAS (&#x2206;AAS) can serve as a diagnostic tool to select patients for imaging even more precisely. METHODS: Eighty-eight patients with early equivocal appendicitis participated in the observation arm of the DIAMOND trial. The data for these patients were reanalysed, and &#x2206;AAS during the observation was calculated. The baseline AAS, final AAS, and the change in C-reactive protein (&#x2206;CRP) were selected as reference standards. RESULTS: Eighty-three patients with complete data were included in the analysis. The AUROC (Area Under the Receiver Operating Characteristic) values are as follows: &#x2206;AAS, 0.932 (95% CI 0.868-0.996); baseline AAS, 0.629 (95% CI 0.498-0.760); final AAS, 0.936 (95% CI 0.886-0.987); and &#x2206;CRP, 0.796 (95% CI 0.696-0.897). Using receiver operating characteristic curves, we established the thresholds for low (AAS&#x2009;&#x2264;&#x2009;-2), intermediate (AAS -1 to 0), and high (AAS&#x2009;&#x2265;&#x2009;1) probability of appendicitis. The negative predictive value for the low-probability group and the positive predictive value for the high-probability group concerning acute appendicitis were 97% and 94%, respectively. CONCLUSION: Patients with equivocal signs of appendicitis may benefit from short observation and the calculation of &#x2206;AAS to reduce overdiagnosis and exposure to excessive imaging. REGISTRATION: The DIAMOND trial was officially registered on ClinicalTrials.gov (NCT02742402) on April 13, 2016.

Adult

Modelling the effects of biological intervention in a dynamical gene network.

Cellular response to environmental and internal signals can be modeled by dynamical gene regulatory networks (GRN). In the literature, three main classes of gene network models can be distinguished: (1) non-quantitative (or data-based) models which do not describe the probability distribution of gene expressions; (2) quantitative models which fully describe the probability distribution of all genes co-expression; and (3) mechanistic models which allow for a causal interpretation of gene interactions. We propose two rigorous frameworks to model gene alteration in a dynamical GRN, depending on whether the network model is quantitative or mechanistic. We explain how these models can be used for design of experiment, or, if additional alteration data are available, for validation purposes or to improve the parameter estimation of the original model. We apply these methods to the Gaussian graphical model, which is quantitative but non-mechanistic, and to mechanistic models of Bayesian networks and penalized linear regression.

Gene Regulatory Networks

Electro-clinical efficacy and safety of midazolam in neonatal seizures: a systematic review with individual level exploratory analysis of gestational age-related treatment response.

UNLABELLED: Neonatal seizures are the most common neurological emergency during the neonatal period and are associated with increased mortality and adverse neurodevelopmental outcomes. Despite current recommendations supporting phenobarbital as first-line therapy, seizure control remains suboptimal in a large proportion of neonates, prompting the use of second-line antiseizure medications. Midazolam is increasingly administered in refractory neonatal seizures but evidence regarding its electro-clinical efficacy and safety remains limited and heterogeneous. To systematically review the available evidence on the electro-clinical efficacy and safety of midazolam in neonatal seizures and to perform an exploratory individual-level analysis investigating the association between gestational age and treatment response. A systematic review was conducted according to PRISMA 2020 guidelines. Studies including neonates with EEG- or aEEG-confirmed seizures treated with midazolam were included. Binary logistic regression was performed to assess the individual-level association between gestational age and treatment response. Eleven studies involving 146 neonates treated with midazolam were included. Electro-clinical response was observed in 101/146 neonates (69.2%), while seizure cessation was achieved in 61/146 neonates (41.8%). In an exploratory complete-case logistic regression analysis, higher gestational age appeared to be associated with a greater probability of electro-clinical response. The predicted probability curve crossed the 50% response probability at approximately 36.5&#xa0;weeks of gestation. Hypotension was the most frequently reported adverse event, while respiratory depression, sedation-related effects, and transient EEG/aEEG suppression were reported less frequently. CONCLUSIONS: Midazolam may have a role as an add-on antiseizure medication in neonatal seizures, particularly in refractory cases. However, the evidence remains limited by heterogeneity in study design, EEG monitoring strategies, outcome definitions, and incomplete individual-level data. The observed association between gestational age and response is hypothesis-generating and requires prospective validation. WHAT IS KNOWN: &#x2022; Phenobarbital often provides incomplete seizure control in neonates, making second-line antiseizure therapies necessary in refractory cases. &#x2022; Evidence supporting midazolam for neonatal seizures remains limited and heterogeneous. WHAT IS NEW: &#x2022; This systematic review summarizes the electro-clinical efficacy and safety of midazolam and includes an exploratory patient-level analysis suggesting that higher gestational age may be associated with improved treatment response. &#x2022; These findings support further prospective studies on developmental determinants of response to GABAergic therapy.

Humans

Development and Validation of a Predictive Model for Identification of Cognitive Impairment Risk in Older Adults with Subjective Cognitive Decline&#xff1a;A Longitudinal Study.

BACKGROUND: Subjective cognitive decline (SCD) is a transitional state between objective cognitive impairment and cognitively intact mental status, providing a critical window for implementing preventive interventions to delay objective cognitive decline. AIMS: We aimed to develop a predictive model for SCD progression in older adults with mild cognitive impairment (MCI). This model will facilitate the identification of risk factors and establishment of targeted interventions for community-based SCD management. METHODS: Data from the China Health and Retirement Longitudinal Study (CHARLS) was utilized in this study, extracting 18 indicators. Potential predictors selected through univariate Cox regression and LASSO regression analyses were sequentially incorporated into a multivariable Cox regression model. A nomogram was constructed to establish a predictive model. Model validation encompassed Area Under Curve (AUC) metrics for discriminative capacity, complemented by quantitative assessments using calibration curve analysis for precision verification and decision curve analysis (DCA) for clinical utility evaluation. RESULTS: A total of 1099 older adults with SCD were included in the final analysis, of whom 114 (10.3%) developed MCI. Multivariable Cox regression identified residence, marital status, educational level, social participation, gait speed, and baseline cognitive function. The model demonstrated time-dependent AUC values of 0.885, 0.830, 0.839, and 0.836 in the training set when evaluating discriminative capacity at 2-, 4-, 7-, and 9-year, respectively. The predictive model showed excellent predictive ability according to AUC, calibration curve, and DCA. CONCLUSIONS: A predictive model was created to estimate the risk of developing MCI in older individuals with SCD, offering clinician-actionable intervention benchmarks for preventive care.

Humans

Development and validation of a liquid chromatography-tandem mass spectrometry method for the quantification of twenty-five steroids in equine serum.

Steroids are potential biomarkers for monitoring equine pregnancy. However, immunoassays currently used for their quantification suffer from cross-reactivity and limited specificity, thus requiring more accurate methods. This study reports the development and validation of a robust liquid chromatography-tandem mass spectrometry (LC-MS/MS) method for simultaneous quantification of 25 steroids covering the main biosynthetic pathways of progestogens, corticosteroids, androgens, and estrogens. Steroids were extracted by protein precipitation followed by evaporation, derivatization, and reconstitution before LC-MS/MS analysis. A surrogate matrix was used for calibration and validation to avoid endogenous interference. Validation was performed according to and partly adapted from Clinical and Laboratory Standards Institute guidelines (CLSI), including linearity, trueness, precision, limits of detection and quantification, measurement uncertainty, recovery, matrix effects, carryover, selectivity, and stability. Calibration curves were fitted using the best-performing weighted linear or quadratic regression model, yielding excellent linearity (R2&#xa0;>&#xa0;0.990), trueness between -9.0% and 2.3%, and intra- and inter-day precision <6.3%. Lower limits of quantification ranged from 2.07 to 2250&#xa0;pg/mL depending on physiological analytes concentration. Extraction recovery averaged 24.3-114.9%, matrix effects were acceptable, and accuracy ranged from 94.4% to 98.9%. No carryover or interferences were detected. Measurement uncertainty remained <15%. This study presents the first LC-MS/MS method partially validated per CLSI criteria for the quantification of 24 steroids in equine serum. The method offers a sensitive and specific alternative to immunoassays and provides a robust tool for equine steroid profiling with potential applications in pregnancy monitoring, placentitis diagnosis, and fetal sex determination.

Animals

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning