Search PubMedSearch

PubMed · 42663140

Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.

Abstract

Mixed-type outcome data occur when the outcome variable's distribution is a mixture of both continuous and discrete ordinal variables. Such mixed-type outcomes are common in biomedical, psychological, and the health sciences, particularly for variables having either a detection or quantitation limit. When interest lies in identifying a combination of genomic features associated with a mixed-type outcome, any method used would require a variable selection strategy for high-dimensional data. Unfortunately, few variable selection methods exist for modeling a mixed-type outcome when the covariate space is high dimensional. This study develops a high-dimensional penalized cumulative probability model (CPM), to allow for the identification of genomic features associated with mixed-type outcome of interest. We demonstrated how such model may be estimated using the iterative penalization procedure-the generalized monotone incremental forward stagewise (GMIFS) algorithm. The Model-X knockoffs procedure was combined with the estimation algorithm to control the false discovery rates (FDR) when performing variable selection. Through extensive simulation studies, our penalized CPM was shown to outperform alternative methods in terms of controlled variable selection performance by achieving high statistical power with the FDR being controlled at the target level. We demonstrate the utility of our method by applying it to predict estimated glomeruli filtration rate (eGFR) in kidney transplant recipients at 24 months post-transplant using baseline gene expression data as predictors. Our CPM model identified five genes associated with this mixed-type outcome which have important links to renal disease, which may provide prognostic guidance for kidney transplantation recipients.

Explore related subjects

Keep this discovery

BibTeXRIS

Shuai Sun, Valeria R Mas, Kellie J Archer. 2026. Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.. https://doi.org/10.1002/sim.70723

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

Future promise, current clinical ambiguity: a systematic review of machine learning algorithm outputs predicting risk of cardiovascular disease.

OBJECTIVE: To examine whether the outputs of machine learning algorithms designed to predict risk of cardiovascular disease (CVD) address known deficiencies of the Framingham Risk Score (FRS) and improve risk estimates. METHODS: For this critical review, Medline, Embase and IEEE were searched from inception to 1 January 2025. Included were studies describing machine learning algorithms designed to specifically compare output of cardiovascular risk assessment with the FRS. Commentaries, letters, unpublished work or non-peer-reviewed papers were excluded.Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, two reviewers screened titles and abstracts independently, then populated a purpose-built data extraction form. A subsequent qualitative thematic analysis focused on algorithms' strengths, added value, potential harms, unintended consequences and equity implications.The main outcome assessed was whether, among healthy adults, the algorithm improved CVD risk prediction relative to the FRS. RESULTS: Of 707 studies retrieved, 29 met inclusion criteria. 23 reported improved predictive ability relative to the FRS. Most datasets and/or medical records used included sociodemographic predictors of CVD not included among FRS inputs. Some added costly diagnostic tests like CT angiography to FRS screening indicators. When they were defined, inputs and outcomes such as hypertension or myocardial infarction did not always adhere to FRS values. Statistical significance was generally taken as a proxy for clinical significance. Some algorithms overestimated the number at risk compared with the FRS without discussing whether that larger proportion might be at risk of overdiagnosis rather than CVD, while a few decreased the proportion found to be at risk. CONCLUSIONS: Use of artificial intelligence to improve accuracy of risk assessment for CVD demonstrates the technological capacity to merge known sociodemographic predictors with biologic variables and examine non-linear interactions among these. Still needed to achieve patient benefit is clinical insight, adherence to screening principles and cost-benefit assessment of inputs selected.

Humans

Machine learning vs. traditional methods for predicting postoperative cardiac complications after non-cardiac surgery: a systematic review and Bayesian network meta-analysis.

INTRODUCTION: Accurate prediction of peri-operative cardiac complications is critical to optimise pre-operative decision-making. Traditional risk prediction scores, such as the Revised Cardiac Risk Index, show only modest discrimination. Machine learning can model complex, non-linear relationships but their predictive performance compared with traditional scores remains unclear. METHODS: We performed a systematic review and Bayesian network meta-analysis. The primary outcome was postoperative adverse cardiac events following non-cardiac surgery. Prediction models were assessed relative to the Revised Cardiac Risk Index. As many studies evaluated multiple versions of each model type, the highest performing ('best version') and lowest performing ('worst version') results were analysed. Models were ranked using the surface under the cumulative ranking curve (SUCRA). RESULTS: Thirteen studies evaluating 54 models and 927,113 patients were included. Machine learning approaches generally outperformed traditional risk scores. Automated machine learning ranked highest (SUCRA 96.6) showed the greatest improvement in the best version analysis (mean difference (MD) 0.28 (95%CrI 0.16-0.40)) and remained superior in the sensitivity analysis (MD 0.30 (95%CrI 0.14-0.45)). Gradient boosting models showed superior performance over the Revised Cardiac Risk Index across analysis (best version: MD 0.20 (95%CrI 0.14-0.26), worst version: MD 0.18 (95%CrI 0.12-0.25), SUCRA 82.4). The Gupta Perioperative Risk for Myocardial Infarction or Cardiac Arrest score outperformed the Revised Cardiac Risk Index in the best version analysis (MD 0.16 (95%CrI 0.01-0.32)). Between-study heterogeneity was low. None of the included studies externally validated their machine learning models and only six were judged to be at low risk of bias. DISCUSSION: Most machine learning models showed better discrimination than traditional risk scores, with automated machine learning and gradient boosting models ranking highest. However, study quality, calibration reporting and absence of external validation limit immediate clinical adoption. Prospective, multicentre evaluation is required before integration of these models into peri-operative practice.

Humans

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (Ψ) represents one of the most abundant and conserved RNA modifications. Ψ provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of Ψ sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel Ψ site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA Ψ-site prediction. The Ψ modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA Ψ-site prediction. Meta-PseU offers a new framework for robust Ψ-site identification by using long sequences.

Pseudouridine