Search PubMedSearch

SEARCH · Search PubMed

Results for “clinical proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Clinical proteomics in inborn errors of metabolism: from biomarker discovery to implementation.

INTRODUCTION: Inborn errors of metabolism (IEMs) are rare, heterogeneous disorders traditionally diagnosed through genetic testing, enzyme assays, and metabolite measurements. However, these tools often do not fully explain phenotypic variability, organ involvement, disease progression, or treatment response. Clinical proteomics provides a complementary functional layer by capturing changes in protein abundance, proteoforms, post-translational modifications (PTM), and biological pathways, offering insights beyond genotype- and metabolite-based approaches. AREAS COVERED: This review examines the role of high-resolution mass spectrometry and computational proteomics in biomarker discovery and clinical decision-making for IEMs. It focuses on their contribution to diagnosis, variant interpretation, patient stratification, and treatment monitoring. Disease-specific applications are discussed, with the strongest evidence in lysosomal storage disorders, mitochondrial diseases, congenital disorders of glycosylation, and selected neurodegenerative or renal metabolic conditions. The literature search was performed in PubMed, Scopus, Web of Science, and Google Scholar, covering peer-reviewed articles available up to 2026, with emphasis on methodological advances and translational applications in clinical proteomics for IEMs. EXPERT OPINION: Proteomics will not replace established diagnostic tools, but it can help address clinically actionable questions in selected contexts. Translation into clinical practice will require standardized workflows, multicenter validation, clinically anchored endpoints, and integration with other omics approaches.

Humans

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

Deciphering mitochondrial metabolic vulnerabilities in ovarian clear cell carcinoma with mass spectrometry-based clinical proteomics.

INTRODUCTION: Ovarian clear cell carcinoma (OCCC) is a rare gynecologic malignancy with a high mortality rate and a lack of response to standard chemotherapy. Despite the functional association between the loss of ARID1A and mitochondrial dependency, the clinical translation of mitochondria-targeted therapies in OCCC has been hindered by a substantial disconnect between biological insight and therapeutic application. There is an urgent, unmet need to identify novel, more specific and effective therapies targeting the mitochondria-related molecular vulnerabilities of ARID1A-mutant OCCC. AREAS COVERED: This critical perspective is informed by results from PubMed literature searches and recent webinars and presentations providing insight into opportunities for mass spectrometry (MS)-based proteomic approaches to enhance and accelerate the clinical translation of mitochondria-targeted therapies in OCCC. EXPERT OPINION: The MS-based proteomic analysis of clinically-relevant experimental models of OCCC will provide a unique opportunity to progress beyond simplified preclinical models and incorporate the full spectrum of patient-specific systemic and microenvironmental factors that may influence therapeutic response, including the adipocyte-related metabolic dependencies of OCCC. Targeted MS is a precise and robust approach that can be applied to verify these novel, mechanistic insights into how mitochondria-targeted therapies intersect with tumor metabolism in OCCC.

Humans

Longitudinal clinical proteomics reveals pneumonia type-specific protein biomarkers and autoantibodies.

Community-acquired pneumonia is a major cause of morbidity and mortality globally. Specific molecular endotypes are currently not well defined, and different viral or bacterial pathogens may trigger specific host responses and pathogenic mechanisms. We performed longitudinal proteomic profiling of bronchoalveolar lavage fluid and plasma from bacterial, influenza, and SARS-CoV-2-driven pneumonia. Our analysis revealed highly pneumonia type-specific proteomic signatures, including COVID-19-specific antibodies locally produced in the lung. These antibodies showed biased immunoglobulin V-domain usage, linked to a CD69/CD83 plasma cell state associated with disease severity and degree of autoimmunity. Using mass spectrometry-driven autoantibody profiling in 2 independent COVID-19 cohorts, we identified 177 putative autoantibodies targeting extracellular matrix, nuclear, and immune-related proteins. Of note, temporal changes in autoantibody profiles correlated with clinical markers of inflammation, organ dysfunction, and duration of hospitalization. These findings highlight the autoimmune aspects of COVID-19 and provide potential biomarkers and therapeutic targets to help improve patient outcomes.

Humans

Proteomics at scale: Bottlenecks and opportunities for early-career researchers in a fast developing field.

The field of proteomics has rapidly evolved over the last five years enabled by rapid advances in instrumentation and computation. At the same time, the proteomics community is also growing. This is reflected by the increasing participation in international conferences such as those organized by the European Proteomics Association and the Human Proteome Organization. These events provide early-career researchers with unique opportunities to exchange ideas, develop collaborations, and build networks that support professional development. One such network is the Young Proteomics Investigators Club, a European initiative supported by European Proteomics Association and led by early-career researchers. In this Community-Driven project, we investigate recent trends in proteomics by screening conference abstracts and evaluating the session attendance at Human Proteome Organization Congresses and European Proteomics Association conferences. Based on these analyses, we identified five areas that, from our perspective, are shaping the current trends in proteomics: clinical proteomics, proteomics of post-translational modifications, single-cell proteomics, systems biology and multi-omics, and computational proteomics. For each area, we highlight both unique challenges and identify a common theme: a shift from exploratory studies with manageable sample numbers towards large screenings and cohorts and the generation of big data, which often comes with the lack of computational support, organizational networks, and infrastructure. In this light, we describe the unique challenges and opportunities faced by early-career researchers. We point to actionable directions for enabling reproducible and transparent proteomics as well as community-driven projects and initiatives, which are often providing training and support. SIGNIFICANCE: In this perspective, the Young Proteomics Investigators Club (YPIC) discusses advances in analytical developments and computational approaches in proteomics research. Based on empirical analysis of recent European Proteomics Association conference and Human Proteome Organization congresses contributions, we identify clinical, single-cell, post-translational and systems-level proteomics as the research areas that have gained most momentum in the last three to five years. What makes this work distinctive is that it is written by and for early-career researchers, thereby uniquely identifying where momentum, challenges, and unmet needs converge for the newest generation of proteomics researchers. Rather than cataloguing advances, we examine the widening gap between what modern proteomics can generate and what individual researchers can realistically process, validate, and interpret. We describe specific structural barriers including access to high performance computing, limited formal training in scalable data analysis, the need for unified benchmarking standards and navigating clinical collaboration frameworks. We then highlight opportunities for the field, such as community-curated benchmarks, interdisciplinary mentorship models, and shared computational infrastructure. By making these challenges explicit from an early-career researchers standpoint, we aim to inform how training, funding, and community initiatives can be shaped to support the next generation of proteomics researchers.

Proteomics

Integrating Imaging-Derived Clinical Endotypes with Plasma Proteomics and External Polygenic Risk Scores Enhances Coronary Microvascular Disease Risk Prediction.

Coronary microvascular disease (CMVD) is an underdiagnosed but significant contributor to the burden of ischemic heart disease, characterized by angina and myocardial infarction. The development of risk prediction models such as polygenic risk scores (PRS) for CMVD has been limited by a lack of large-scale genome-wide association studies (GWAS). However, there is significant overlap between CMVD and enrollment criteria for coronary artery disease (CAD) GWAS. In this study, we developed CMVD PRS models by selecting variants identified in a CMVD GWAS and applying weights from an external CAD GWAS, using CMVD-associated loci as proxies for the genetic risk. We integrated plasma proteomics, clinical measures from perfusion PET imaging, and PRS to evaluate their contributions to CMVD risk prediction in comprehensive machine and deep learning models. We then developed a novel unsupervised endotyping framework for CMVD from perfusion PET-derived myocardial blood flow data, revealing distinct patient subgroups beyond traditional case-control definitions. This imaging-based stratification substantially improved classification performance alongside plasma proteomics and PRS, achieving AUROCs between 0.65 and 0.73 per class, significantly outperforming binary classifiers and existing clinical models, highlighting the potential of this stratification approach to enable more precise and personalized diagnosis by capturing the underlying heterogeneity of CMVD. This work represents the first application of imaging-based endotyping and the integration of genetic and proteomic data for CMVD risk prediction, establishing a framework for multimodal modeling in complex diseases.

Cardiovascular Disease

Unlocking the Circulating Proteome: Toward Clinical Translation.

Blood-based proteomics is approaching a translational inflection point. Driven by advances in measurement technologies, rapid expansion of analytical capabilities, and growing adoption across research and medical communities, there is increasing demand for clinically actionable biomarkers. As the field transitions away from purely large-scale discovery-oriented studies toward more informed, targeted, application-driven analyses, the generation of proteomic data is no longer the bottleneck. Instead, the central challenge is to translate these measurements into robust, reproducible, and clinically meaningful insights. In this Review, we assess recent technological and methodological developments, evaluate persistent preanalytical and interpretative limitations, and outline the key steps required for clinical translation. We focus on three deeply interconnected dimensions: the capabilities and constraints of current measurement platforms, the role of computational and machine learning approaches in extracting biological and clinical signals, and the emergence of large-scale population studies that create new opportunities for validation and generalization. Finally, we discuss a forward-looking vision in which proteomics plays a central role in dynamic, multilayered omics frameworks, where integration with genomics, temporal profiling, and imaging can deepen our understanding of health, disease, and therapeutic response.

Humans

Comprehensive Assessment of the Intrinsic Pancreatic Microbiome.

OBJECTIVE: To sought comprehensively profile tissue and cyst fluid in patients with benign, precancerous, and cancerous conditions of the pancreas to characterize the intrinsic pancreatic microbiome. BACKGROUND: Small studies in pancreatic ductal adenocarcinoma (PDAC) and intraductal papillary mucinous neoplasm (IPMN) have suggested that intrapancreatic microbial dysbiosis may drive malignant transformation. METHODS: Pancreatic samples were collected at the time of resection from 109 patients. Samples included tumor tissue (control, n = 20; IPMN, n = 20; PDAC, n = 19) and pancreatic cyst fluid (IPMN, n = 30; serous cystadenomas, n = 10; mucinous cystic neoplasm, n = 10). Assessment of bacterial DNA by quantitative polymerase chain reaction and 16S ribosomal RNA gene sequencing was performed. Downstream analyses determined the relative abundances of individual taxa between groups and compared intergroup diversity. Whole-genome sequencing data from 140 patients with PDAC in the National Cancer Institute's Clinical Proteomic Tumor Analysis Consortium were analyzed to validate findings. RESULTS: Sequencing of pancreatic tissue yielded few microbial reads regardless of diagnosis, and analysis of pancreatic tissue showed no difference in the abundance and composition of bacterial taxa between normal pancreas, IPMN, or PDAC groups. Low-grade and high-grade dysplasia IPMN were characterized by low bacterial abundances with no difference in tissue composition and a slight increase in Pseudomonas and Sediminibacterium in high-grade dysplasia cyst fluid. Decontamination analysis using the Clinical Proteomic Tumor Analysis Consortium database confirmed a low-biomass, low-diversity intrinsic pancreatic microbiome that did not differ by pathology. CONCLUSIONS: Our analysis of the pancreatic microbiome demonstrated very low intrinsic biomass that is relatively conserved across diverse neoplastic conditions and thus unlikely to drive malignant transformation.

Humans

Pan-Cancer Quantification of Driver Alteration Transmission Across Molecular Layers Reveals Limited Propagation to Protein Abundance.

Precision oncology relies primarily on DNA-level alterations for therapeutic decisions, but the extent to which driver mutations propagate to protein abundance has not been systematically evaluated. Here, I developed a regression-based transmission score (TS_R 2) to quantify driver alteration signal propagation across DNA, mRNA, and protein layers. Applying this framework to matched genomic, transcriptomic, proteomic, and phosphoproteomic data from 754 Clinical Proteomic Tumor Analysis Consortium (CPTAC) tumors across seven cancer types, I analyzed 86 driver gene-cancer type pairs, of which 83 were evaluable for the full two-layer transmission score. I employed covariate-adjusted regression for each molecular transition, assessing significance via permutation testing (n = 1000). Mixed-effects modeling then partitioned gene-intrinsic from cancer-type-dependent effects. Only 5 of 83 evaluable pairs (6%) demonstrated high transmission (TS_R 2 > 0.05), with receptor tyrosine kinases (EGFR, FGFR2) exemplifying this class. The primary bottleneck occurred at the mutation-mRNA transition, not mRNA-protein translation. Gene identity accounted for 49% of transmission efficiency variance, nearly double the contribution of cancer type (29%). Copy number alterations transmitted signals 13.8-fold more efficiently than point mutations, and truncating mutations showed higher transmission than missense variants (Wilcoxon p = 0.005). Microsatellite instability attenuated mRNA-protein transmission in UCEC and COAD. These findings demonstrate that many driver alterations show limited propagation to protein abundance. This challenges DNA-only interpretations in precision oncology and provides a framework for integrated functional driver prioritization.

Humans

A pan-cancer analysis of MEX3D in human tumors.

BACKGROUND: MEX3D, a member of the MEX3 RNA-binding protein family, has emerged as a potential regulatory molecule in cancer. However, its role across different tumor types remains largely unexplored. METHODS: We conducted a pan-cancer analysis of MEX3D using transcriptomic and proteomic data from the Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), and Clinical Proteomic Tumor Analysis Consortium (CPTAC). Expression patterns, clinical correlations, survival outcomes, genetic alterations, RNA modification associations, immune infiltration, and functional enrichment were systematically evaluated. RESULTS: MEX3D was significantly dysregulated in numerous cancers at both mRNA and protein levels. Its expression correlated with tumor stage in ACC, LIHC, OV, SKCM, and THCA. Elevated MEX3D expression was associated with poor overall survival (OS) and disease-specific survival (DSS) in multiple malignancies, including ACC, LGG, LUAD, and MESO. Genetic alteration analysis revealed frequent amplifications and mutations, particularly in SARC and OV. MEX3D was positively correlated with RNA modification-related genes (m1A, m5C, m6A) and immune regulatory genes such as CD276, TGFB1, VEGFA, and ICOSLG. Additionally, MEX3D expression showed significant associations with tumor mutational burden (TMB), microsatellite instability (MSI), and cancer-associated fibroblast infiltration. Functional enrichment analyses indicated that MEX3D-related genes are involved in reproductive cellular processes, RNA binding, the Hippo signaling pathway, and microRNA-related oncogenic pathways. CONCLUSION: This pan-cancer analysis highlights the heterogeneous expression and cancer-specific prognostic significance of MEX3D. MEX3D is associated with immune infiltration, immune regulatory genes, RNA modification-related genes, TMB/MSI, and pathways involved in gene regulation and tumor progression. These findings suggest that MEX3D may participate in cancer-specific post-transcriptional and microenvironmental regulatory networks.

Biomarker

PDZ-binding kinase promotes ovarian cancer cell proliferation and invasion via CCNB1 regulation.

BACKGROUND: Ovarian cancer is one of the most lethal gynecological malignancies, characterized by late diagnosis, frequent recurrence, and high mortality. PDZ-binding kinase (PBK), a serine/threonine kinase of the mitogen-activated protein kinase kinase (MAPKK) family, has been implicated in the tumorigenesis of multiple cancers, yet its role in ovarian cancer remains incompletely characterized. This study aimed to investigate the effect of PBK on the proliferation and invasion of ovarian cancer cells. METHODS: The expression of PBK and cyclin B1 (CCNB1) in normal ovarian tissues and ovarian cancer tissues was analyzed using online databases including Gene Expression Profiling Interactive Analysis 2 (GEPIA2), Clinical Proteomic Tumor Analysis Consortium (CPTAC), and Kaplan-Meier Plotter. Clinical tissue specimens were collected to detect the expression of PBK and CCNB1 by immunohistochemistry. Quantitative real-time polymerase chain reaction (PCR) was performed to detect PBK messenger RNA (mRNA) expression levels in clinical specimens and cell lines. Western blot was used to detect PBK protein expression in ovarian cancer cell lines. ES2 and A2780 cells with higher PBK expression were selected to construct PBK knockdown cell lines using lentiviral interference vectors. Cell Counting Kit-8 (CCK-8) assay, colony formation assay, and 5-ethynyl-2'-deoxyuridine (EdU) assay were performed to explore the effect of PBK knockdown on cell proliferation. Transwell assay was used to investigate the effect on cell invasion. The Cancer Genome Atlas (TCGA) and Kyoto Encyclopedia of Genes and Genomes (KEGG) databases were utilized to analyze PBK-related pathways and predict CCNB1 as the gene most closely related to PBK. RESULTS: PBK was significantly overexpressed in ovarian cancer tissues and cell lines compared with normal controls, and high PBK expression was associated with poor overall survival (OS) and progression-free survival (PFS). Knockdown of PBK expression inhibited the proliferation, colony formation, and invasion of ovarian cancer cells. Bioinformatics analysis revealed that CCNB1 was significantly overexpressed in ovarian cancer and high CCNB1 expression was associated with poor OS. CCNB1 was also significantly highly expressed in ovarian cancer tissues as validated by immunohistochemistry and was associated with lymph node metastasis. PBK and CCNB1 expression showed a significant positive correlation in TCGA ovarian cancer datasets. Knockdown of PBK inhibited CCNB1 expression in ovarian cancer cells. CONCLUSIONS: PBK promotes ovarian cancer cell proliferation and invasion. PBK knockdown leads to CCNB1 downregulation. These findings suggest that CCNB1 contributes to PBK-mediated oncogenic effects and identify the PBK-CCNB1 axis as a potential therapeutic target for ovarian cancer treatment.

PDZ-binding kinase (PBK)

The molecular similarity landscape of preclinical cancer models to patient tumors.

Selecting appropriate preclinical models is fundamental for translational oncology, yet a large-scale, multi-omic quantitative comparison of their similarity to primary human tumors is lacking. To address this, we integrated transcriptomic, proteomic, and genomic profiles from over 10,000 primary tumors from The Cancer Genome Atlas (TCGA) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC), alongside 4,000 preclinical models. Using a robust computational framework, we revealed a clear hierarchy of transcriptomic and proteomic similarity to patient tumors: with patient-dervied xenografts (PDXs) having greater transcriptomic and proteomic similarity to patient tumors (>) compared with patient-derived organoids (PDOs), which are equal in hierarchy to that of PDX-dervied organoids (PDXOs) > cell lines. We also quantified high molecular conservation (Pearson correlation coefficient = 0.96) across paired in vitro to in vivo platform (organoids to PDX) transitions. Furthermore, genomic analysis demonstrated that whole-exome sequencing (WES) outperforms RNA-seq in detecting DNA variants, and it identified a clonal complexity hierarchy (cell lines > PDXOs > PDXs > PDOs) reflecting the effect of passaging history on intratumor heterogeneity. Ultimately, this study delivers a comprehensive quantitative benchmark, establishing a population-level hierarchy of molecular similarity between preclinical models and primary tumors and providing a data-driven reference for model selection. These findings offer a data-driven framework for selecting models that balance biological representativeness with experimental practicality.

Humans

Bridging the Gap From Proteomics Technology to Clinical Application: Highlights From the 68th Benzon Foundation Symposium.

The 68th Benzon Foundation Symposium brought together leading experts to explore the integration of mass spectrometry-based proteomics and artificial intelligence to revolutionize personalized medicine. This report highlights key discussions on recent technological advances in mass spectrometry-based proteomics, including improvements in sensitivity, throughput, and data analysis. Particular emphasis was placed on plasma proteomics and its potential for biomarker discovery across various diseases. The symposium addressed critical challenges in translating proteomic discoveries to clinical practice, including standardization, regulatory considerations, and the need for robust "business cases" to motivate adoption. Promising applications were presented in areas such as cancer diagnostics, neurodegenerative diseases, and cardiovascular health. The integration of proteomics with other omics technologies and imaging methods was explored, showcasing the power of multimodal approaches in understanding complex biological systems. Artificial intelligence emerged as a crucial tool for the acquisition of large-scale proteomic datasets, extracting meaningful insights, and enhancing clinical decision-making. By fostering dialog between academic researchers, industry leaders in proteomics technology, and clinicians, the symposium illuminated potential pathways for proteomics to transform personalized medicine, advancing the cause of more precise diagnostics and targeted therapies.

Proteomics

Development of a Fit-For-Purpose Multi-Marker Panel for Early Diagnosis of Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) suffers from a lack of an effective diagnostic method, which hampers improvement in patient survival. Carbohydrate antigen 19-9 (CA19-9) is the only FDA-approved blood biomarker for PDAC, yet its clinical utility is limited due to suboptimal performance. Liquid chromatography-mass spectrometry (LC-MS) has emerged as a burgeoning technology in clinical proteomics for the discovery, verification, and validation of novel biomarkers. A plethora of protein biomarker candidates for PDAC have been identified using LC-MS, yet few has successfully transitioned into clinical practice. This translational standstill is owed partly to insufficient considerations of practical needs and perspectives of clinical implementation during biomarker development pipelines, such as demonstrating the analytical robustness of proposed biomarkers which is critical for transitioning from research-grade to clinical-grade assays. Moreover, the throughput and cost-effectiveness of proposed assays ought to be considered concomitantly from the early phases of the biomarker pipelines for enhancing widespread adoption in clinical settings. Here, we developed a fit-for-purpose multi-marker panel for PDAC diagnosis by consolidating analytically robust biomarkers as well as employing a relatively simple LC-MS protocol. In the discovery phase, we comprehensively surveyed putative PDAC biomarkers from both in-house data and prior studies. In the verification phase, we developed a multiple-reaction monitoring (MRM)-MS-based proteomic assay using surrogate peptides that passed stringent analytical validation tests. We adopted a high-throughput protocol including a short gradient (<10&#xa0;min) and simple sample preparation (no depletion or enrichment steps). Additionally, we developed our assay using serum samples, which are usually the preferred biospecimen in clinical settings. We developed predictive models based on our final panel of 12 protein biomarkers combined with CA19-9, which showed improved diagnostic performance compared to using CA19-9 alone in discriminating PDAC from non-PDAC controls including healthy individuals and patients with benign pancreatic diseases. A large-scale clinical validation is underway to demonstrate the clinical validity of our novel panel.

Humans

GICPIdb: an archival repository of multimodal data focusing on pathological images for gastrointestinal cancers.

INTRODUCTION: Deep learning (DL) shows great potential for predicting biomarkers from routine histopathological slides of gastrointestinal (GI) cancers. Yet most existing models are validated on limited patient cohorts, while pathological image annotation and molecular marker standardization demand substantial professional expertise. To address these gaps, we constructed the Gastrointestinal Cancer Pathological Image Archive (GICPIdb, gicpidb.shubuzuo.top), a dedicated database and web platform covering seven major GI cancer types. METHODS: High-quality hematoxylin and eosin (H&E)-stained whole-slide images were collected from multiple sources and uniformly processed. Image annotations were performed by board-certified pathologists following standardized protocols. GICPIdb offers five interactive web modules for data uploading, quality control, feature extraction, online annotation and AI-based prediction. Its intuitive interface supports data browsing, retrieval, visualization and downloading. RESULTS: The database houses 2,863 pathologist-annotated, uniformly processed, high-quality H&E stained images collected from 2,655 patients. Of these, 1,699 patients were sourced from The Cancer Genome Atlas (TCGA), 182 from the Clinical Proteomic Tumor Analysis Consortium (CPTAC), and 424 from China-Japan Friendship Hospital and 350 from Chifeng Municipal Hospital in Inner Mongolia, China. It also integrates data on over 50 key molecular markers (e.g., MSI, TMB) and prognostic labels related to survival, recurrence and metastasis. DISCUSSION: GICPIdb aims to promote the development of DL-driven AI tools for cancer research and clinical translation. The multi-institutional data collection and standardized annotation pipeline are expected to enhance the generalizability and reproducibility of AI-based prediction models across diverse patient populations.

deep learning

UALCAN Mobile, an app for cancer proteogenomic data analysis.

Cancer is a complex disease affecting various organs and is a major cause of death worldwide. During cancer initiation, disease progression, and tumor metastasis, various genomic and proteomic alterations are observed. Recent technological advances have led to the generation of large amounts of molecular data, including genomics and transcriptomics. These large-scale datasets can be utilized to analyze and identify sub-class-specific cancer biomarkers and targets. However, there is a need for the development of user-friendly tools for large-scale data analysis, disseminating the analyzed data in a visualizable format to cancer researchers with no programming skills. We developed UALCAN, a comprehensive platform that allows users to integrate disparate data to better understand the genes, proteins, and pathways perturbed in cancer and make discoveries of potential biomarkers and targets. In the current study, we describe the development of the UALCAN Mobile application (app) that will provide cancer transcriptomic data obtained from The Cancer Genome Atlas (TCGA) project to evaluate protein-coding gene expression based on various stratifications, including stage, grade, race, gender, and molecular-subtypes across over 30 types of cancers. In addition, the UALCAN mobile provides data analysis options for epigenetic changes due to DNA promoter methylation and Clinical Proteomic Tumor Analysis Consortium (CPTAC) cancer proteomic data. The app provides access to large cancer molecular datasets on the go. To find changes in the expression of causative genes and proteins and to identify biomarkers and therapeutic targets, UALCAN mobile app will be extremely valuable. The "UALCAN Mobile" app is free to use and can be downloaded from both the iOS/Apple and the Android Play Store and has been downloaded over 100 times in each of iOS and android app stores.

app

A weakly supervised deep learning-based recurrence prediction and risk stratification of lung adenocarcinoma from pathology whole-slide images.

BACKGROUND: Accurate prediction of postoperative recurrence in lung adenocarcinoma (LUAD) is essential for guiding clinical decision-making and improving patient outcomes. Although various predictive models have been developed, most rely on complex genomic analyses and high-dimensional clinical data. The complexity of these approaches substantially limits their feasibility for routine clinical use. To address this clinical challenge, this study aims to predict postoperative recurrence using routinely available hematoxylin and eosin (H&E)-stained images and characterize the associated biological features. METHODS: A total of 329 patients who underwent curative resection at the First Affiliated Hospital of Wenzhou Medical University (FHWMU) were retrospectively enrolled and randomly assigned to training and internal validation cohorts in a 7:3 ratio. An independent external validation cohort comprising 70 patients from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) was included. Three patch-level feature extractors (Inception_V3, ResNet18, and DenseNet121) were evaluated within a weakly supervised multiple-instance learning (MIL) framework incorporating automated region-of-interest (ROI) detection on segmented whole-slide images (WSIs). Model performance was assessed using the area under the receiver operating characteristic curve (AUC), Kaplan-Meier (KM) survival analysis, and multivariable Cox proportional hazards regression. Transcriptomic profiling and gene set enrichment analysis (GSEA) were conducted to investigate biological differences between risk groups. RESULTS: The model achieved AUCs of 0.923 in the training cohort, 0.891 in the internal validation cohort, and 0.847 in the external validation cohort. The model effectively stratified patients into high- and low-risk groups with significantly different recurrence-free survival (RFS) across all cohorts (all P&#x2009;<&#x2009;0.001) and retained prognostic value within AJCC stages I-III. Transcriptomic analyses revealed consistent enrichment of cell cycle-related pathways and neutrophil extracellular trap (NET) formation in high-risk patients across both institutional and CPTAC cohorts, aligning with distinct biological profiles of the model-derived risk stratification. CONCLUSIONS: This weakly supervised deep learning framework enables accurate and externally validated prediction of postoperative recurrence in LUAD using routinely available histopathological images, and integration of histopathological features with molecular analyses enhances biological interpretability. This work provides a clinically accessible and cost-effective tool for postoperative risk assessment in LUAD patients.

Humans

Deep learning-based multimodal pathogenomics integration for precision cancer prognosis.

BACKGROUND: Recent studies have revealed valuable prognostic insights in haematoxylin and eosin (H&E)-stained histological sections and transcriptomic profiles, suggesting potential applications in machine learning. However, existing methods lack sufficient intra- and inter-modal interactions, and face challenges in clinical validation due to incomplete multimodal data. METHODS: We proposed PathoGems (PathoGenomics-based integrative survival prediction), a weakly-supervised, interpretable multimodal learning framework that integrates histology and genomic profiles for precise cancer prognosis prediction. To evaluate the robustness of PathoGems, we initially curated a dataset of 1965 cases across four cohorts from The Cancer Genome Atlas (TCGA), including breast, colorectal, glioblastoma, and esophageal cancers. For external validation, PathoGems was further evaluated on four independent cohorts, consisting of 76 breast cancer and 41 esophageal squamous cell carcinoma cases from Zhejiang Cancer Hospital, as well as 102 colorectal cancer and 58 glioblastoma cases from the Clinical Proteomic Tumor Analysis Consortium (CPTAC). RESULTS: PathoGems effectively stratified patients into favorable and unfavorable risk groups, revealing significant differences in histological patterns, genomic features, and overall survival (log-rank test, p&#x2009;<&#x2009;0.05). Moreover, the model&#x2019;s predictions are further supported by visualization and transcriptomic analysis, enhancing interpretability and reliability. CONCLUSIONS: By fusing histological and clinicogenomic multimodal models, PathoGems will provide a solid foundation for developing an innovative tool that aids clinicians in making informed decisions and selection personalized treatment strategies for cancer patients.

Humans