Search PubMedSearch

SEARCH · Search PubMed

Results for “Biomedical data processing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

18 recordsLinked to original sources

NeuroOmics-Net: An interpretable multimodal deep learning framework for Alzheimer's disease diagnosis and progression prediction using neuroimaging, EEG, and genomic data.

Accurate diagnosis and progression prediction of Alzheimer's disease (AD) remain challenging due to the heterogeneous nature of the disease, which involves structural brain degeneration, electrophysiological dysfunction, and molecular dysregulation. Most existing deep learning approaches rely on a single modality or limited multimodal combinations, thereby failing to capture the complex cross-domain interactions underlying AD progression. Furthermore, the scarcity of large-scale datasets containing synchronized neuroimaging, electrophysiological, and genomic measurements restricts the development of comprehensive multimodal diagnostic systems. To address these challenges, this study proposes NeuroOmics-Net, a multimodal deep learning framework for Alzheimer's disease analysis that integrates structural magnetic resonance imaging (sMRI), electroencephalography (EEG), and gene expression data. The proposed framework combines a Hierarchical Multi-View Encoder (HME) for modality-specific feature extraction, a Cross-Omics Attention Fusion (CAF) module for adaptive integration of complementary biomarkers, and a Disease Progression Graph Learning (DPGL) module for modeling progression-related relationships across biological domains. To facilitate cross-modal integration from independent cohorts, Regularized Canonical Correlation Analysis (RCCA) is employed to align heterogeneous feature representations within a shared latent space. Experiments were conducted using publicly available datasets from ADNI, PhysioNet, and GEO repositories comprising 1120 diagnosis-aligned samples. The proposed framework achieved 94.3% classification accuracy and an AUC of 0.975 for distinguishing normal controls (NC), mild cognitive impairment (MCI), and Alzheimer's disease subjects, while attaining 93.7% accuracy for predicting conversion from stable mild cognitive impairment (sMCI) to progressive mild cognitive impairment (pMCI). However, a fairness sensitivity analysis using stratified demographic reweighting revealed accuracy ranging from 90.8% (low-education, high-comorbidity proxy subgroup) to 96.1% (low-risk, high-reserve proxy subgroup), a demographic parity gap of 5.3 percentage points, indicating that overall accuracy reflects a performance ceiling in a relatively homogeneous research cohort rather than a realistic estimate for demographically diverse clinical populations. Comparative evaluations demonstrated consistent improvements over state-of-the-art unimodal and multimodal deep learning models. Interpretability analysis further identified clinically relevant biomarkers, including hippocampal and entorhinal atrophy, theta-alpha EEG alterations, and APOE-associated molecular pathways. Because sMRI, EEG, and gene expression data were sourced from separate, unpaired cohorts with no subjects possessing all three synchronized measurements, all reported cross-modal associations reflect population-level statistical correspondence across diagnosis-matched groups rather than within-subject physiological coupling; no claim of intra-individual causal cross-modal interaction is made. These findings demonstrate that NeuroOmics-Net provides an effective computer-aided framework for multimodal biomedical data processing and Alzheimer's disease analysis. By integrating neuroimaging, electrophysiological, and genomic information, the proposed approach enables accurate diagnosis, progression prediction, and biologically interpretable decision support for clinical and translational applications.

Humans

Target and biomarker exploration portal for drug discovery.

MOTIVATION: The discovery of novel drug targets and precision biomarkers remains a major challenge in drug development, with traditional differential expression analysis often overlooking key regulatory proteins. Here, we present a novel, web-based bioinformatics tool, the Target and Biomarker Exploration Portal (TBEP), designed to accelerate the drug discovery process by integrating large-scale biomedical data with network analysis techniques. RESULTS: TBEP harnesses machine-learning approaches to mine and combine multimodal datasets, including human genetics, functional genomics, and protein-protein interaction networks, to decode causal disease mechanisms and uncover novel therapeutic targets and precision biomarkers for specific phenotypes. A unique feature of the tool is its ability to process large-scale data in real-time, facilitated by an efficient cloud-based architecture. Additionally, the tool incorporates an integrated large language model (LLM), which assists researchers in exploring and interpreting complex biological relationships within the generated networks and multi-omics data using natural language (English). By offering an intuitive, interactive interface, the LLM enhances the exploration of biological insights, making it easier for scientists to derive actionable conclusions. This powerful integration of network analysis, multi-omics data, and LLM provides a robust framework for accelerating the identification of novel drug targets. AVAILABILITY AND IMPLEMENTATION: The tool is publicly available at https://tbep.missouri.edu. The source code, documentation and installation instructions are available at GitHub repository: https://github.com/mizzoudbl/tbep.

Drug Discovery

BAV-LLPS: a database of bacterial, archaea, and virus liquid-liquid phase separation proteins.

MOTIVATION: Liquid-liquid phase separation (LLPS) is a key process underlying the formation of biomolecular condensates, such as membrane-less organelles, that compartmentalize biochemical processes inside the cells. While LLPS has been extensively studied in eukaryotes, its role in bacteria, archaea, and viruses remains far less characterized. Recent studies in bacteria have revealed that LLPS-driven condensates play critical roles in RNA processing, stress response, and pathogenicity. Similarly, many viruses exploit LLPS to facilitate crucial steps in their infection cycles, including viral entry, genome replication, assembly, and host immune evasion. RESULTS: In this work, we introduce a hand-curated database of LLPS proteins from bacteria, archaea, and viruses (BAV-LLPS Database). This resource, extended through sequence similarity searches, comprises over 5000 proteins and integrates diverse data including biological annotations, sequence features, predicted disordered regions, LLPS per site probability, and AlphaFold2-based structural models. Additionally, our web server enables users to explore both the curated and homologous derived datasets, providing a platform to uncover evolutionary relationships and intrinsic and differential properties of LLPS proteins across various taxonomic groups. This work seeks to deepen our understanding of LLPS mechanisms beyond eukaryotic organisms, emphasizing their significance across diverse life forms. It also aims to foster the development of specialized predictive tools that will facilitate the exploration and characterization of LLPS processes in a wide array of living organisms, thereby contributing to advancements in both fundamental biological research and applied biomedical sciences. AVAILABILITY AND IMPLEMENTATION: BAV-LLPS DB is freely accessible at https://bav-llps-db.bioinformatica.org/. The data can be retrieved from the website. The source code of the database can be downloaded from https://bav-llps-db.bioinformatica.org/download.

Databases, Protein

Pre-Meta: priors-augmented retrieval for LLM-based metadata generation.

MOTIVATION: While high-throughput sequencing technologies have dramatically accelerated genomic data generation, the manual processes required for dataset annotation and metadata creation impede the efficient discovery and publication of these resources across disparate public repositories. Large language models (LLMs) have the potential to streamline dataset profiling and discovery. However, their current limitations in generalizing across specialized knowledge domains, particularly in fields such as biomedical genomics, prevent them from fully realizing this potential. This article presents Pre-Meta, an LLM-agnostic and domain-independent data annotation pipeline with an enriched retrieval procedure that leverages related priors-such as pre-generated metadata tags and ontologies-as auxiliary information to improve the accuracy of automated metadata generation. RESULTS: Validated using five selected metadata fields sampled across 1500 papers, the Pre-Meta assisted annotation experiment-without finetuning and prompt optimization-demonstrates a systemic improvement in the annotation task: shown through a 23%, 72%, and 75% accuracy gain from conventional RAG adoptions of GPT-4o mini, Llama 8B, and Mistral 7B respectively. AVAILABILITY AND IMPLEMENTATION: The code, data access, and scripts are available at: https://github.com/SINTEF-SE/LLMDap.

Metadata

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models

Managing workflow executions with WESkit.

SUMMARY: In biomedical research, managing computational workflows across numerous projects-with varying parameters, tools, and environments-creates major challenges in scalability, reproducibility, and collaboration. Here, we present WESkit, an implementation of the Global Alliance for Genomics and Health (GA4GH) Workflow Execution Service (WES) interface, designed to streamline the execution, monitoring, and documentation of data processing workflows. It addresses the complexities involved in managing numerous executions with varying parameters across diverse research projects. Supporting both Snakemake and Nextflow, the system enables consistent automation and centralized monitoring, which benefits research groups aiming for long-term reproducibility and scalable collaboration. Its suitability for larger teams and service units is further enhanced by seamless integration into cloud environments, contributing to the GA4GH cloud framework. AVAILABILITY AND IMPLEMENTATION: The software WESkit is available under MIT license at the GitLab repository (https://gitlab.com/one-touch-pipeline/weskit). The WESkit main repository is archived at Software Heritage (https://archive.softwareheritage.org/browse/origin/directory/?origin_url=https://gitlab.com/one-touch-pipeline/weskit/api.git) and can be found using "one-touch-pipeline/weskit" term in the search section.

Workflow

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design

EMTscore infers divergent EMT pathways from omics data and enables rapid screening for EMT-associated gene sets.

MOTIVATION: Quantitative analyses of epithelial-mesenchymal transition (EMT) have been widely used in several areas of biomedical sciences due to its importance in development and cancer progression, but its multi-contextual nature requires standardization and implementation of gene set scoring methods beyond capacities of conventional tools. RESULTS: We developed EMTscore, a package that provides an efficient implementation of unbiased scoring methods for multiple EMT pathways using individual single-cell or bulk omics data, and the package allows rapid screening for cellular processes correlated with EMT. AVAILABILITY AND IMPLEMENTATION: EMTscore is available from GitHub https://github.com/wenmm/EMTscore under the GNU General Public License, and is uploaded on Zenodo with a DOI 10.5281/zenodo.19487376.

Epithelial-Mesenchymal Transition

PEARL: integrative multi-omics classification and omics feature discovery via deep graph learning.

MOTIVATION: Integrating multi-omics data provides valuable insights into biological processes by capturing information across multiple molecular layers, enabling a comprehensive understanding of complex diseases and driving advancements in precision medicine. However, existing computational methods for multi-omics integration face significant challenges, such as low reliability and poor generalizability, due to the high dimensionality and low sample size nature of omics data. RESULTS: To address these challenges, we present PEARL (Pearson-Enhanced spectrAl gRaph convoLutional networks), a novel deep graph learning method for biomedical classification and functional important omics features identification. PEARL leverages a simple yet effective learning architecture to achieve superior and robust performance in high-dimensional, low-sample-size multi-omics settings. Our results demonstrate that PEARL significantly outperforms existing state-of-the-art methods on both synthetic and real biomedical datasets. Furthermore, applied to Alzheimer's disease (AD) brain multi-omics data, features prioritized by PEARL lead to functionally important genes that demonstrate significant enrichment in AD-related pathways. These findings highlight PEARL's practical utility in biomedical research and its potential to enhance biological interpretability in multi-omics studies. AVAILABILITY AND IMPLEMENTATION: The source code of our computational framework is available at https://github.com/zqq121017/PEARL.

Multiomics

Zebrafish as a versatile model in biomedical research, from disease modeling to regenerative medicine: a review.

Zebrafish are an effective animal model widely utilized in biomedical research. They are known for their rapid reproduction and substantial genetic similarity to humans. Their transparent embryos directly enable the visualization of developmental processes and disease progression. This makes zebrafish invaluable for studying a broad range of human diseases, including cancer, cardiovascular disorders, and neurodegenerative conditions. Compared with other vertebrate models, zebrafish offer several advantages, including ease of genome editing, cost-effective maintenance, and suitability for high-throughput drug screening. Recent advancements have expanded the use of zebrafish in disease modeling and regenerative medicine, providing deeper insights into the genetic and cellular mechanisms underlying human pathologies. Zebrafish provide a robust platform for evaluating the safety, efficacy, and regenerative potential of both natural and synthetic biomaterials, including hydroxyapatite, bioactive glass nanoparticles, and bioceramics. This capability facilitates the creation of artificial tissues that closely resemble native structures. Additionally, integrating artificial intelligence technologies has improved automated data analysis and phenotyping in zebrafish studies, enhancing both accuracy and throughput. This review highlights current applications of zebrafish in disease modeling, drug discovery, regenerative medicine, and biomaterial assessment, emphasizing their evolving role as a versatile preclinical platform supported by advanced genetic and computational tools.

Animals

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings.

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Humans

Drug target ontology to classify and integrate drug discovery data.

BACKGROUND: One of the most successful approaches to develop new small molecule therapeutics has been to start from a validated druggable protein target. However, only a small subset of potentially druggable targets has attracted significant research and development resources. The Illuminating the Druggable Genome (IDG) project develops resources to catalyze the development of likely targetable, yet currently understudied prospective drug targets. A central component of the IDG program is a comprehensive knowledge resource of the druggable genome. RESULTS: As part of that effort, we have developed a framework to integrate, navigate, and analyze drug discovery data based on formalized and standardized classifications and annotations of druggable protein targets, the Drug Target Ontology (DTO). DTO was constructed by extensive curation and consolidation of various resources. DTO classifies the four major drug target protein families, GPCRs, kinases, ion channels and nuclear receptors, based on phylogenecity, function, target development level, disease association, tissue expression, chemical ligand and substrate characteristics, and target-family specific characteristics. The formal ontology was built using a new software tool to auto-generate most axioms from a database while supporting manual knowledge acquisition. A modular, hierarchical implementation facilitate ontology development and maintenance and makes use of various external ontologies, thus integrating the DTO into the ecosystem of biomedical ontologies. As a formal OWL-DL ontology, DTO contains asserted and inferred axioms. Modeling data from the Library of Integrated Network-based Cellular Signatures (LINCS) program illustrates the potential of DTO for contextual data integration and nuanced definition of important drug target characteristics. DTO has been implemented in the IDG user interface Portal, Pharos and the TIN-X explorer of protein target disease relationships. CONCLUSIONS: DTO was built based on the need for a formal semantic model for druggable targets including various related information such as protein, gene, protein domain, protein structure, binding site, small molecule drug, mechanism of action, protein tissue localization, disease association, and many other types of information. DTO will further facilitate the otherwise challenging integration and formal linking to biological assays, phenotypes, disease models, drug poly-pharmacology, binding kinetics and many other processes, functions and qualities that are at the core of drug discovery. The first version of DTO is publically available via the website http://drugtargetontology.org/ , Github ( http://github.com/DrugTargetOntology/DTO ), and the NCBO Bioportal ( http://bioportal.bioontology.org/ontologies/DTO ). The long-term goal of DTO is to provide such an integrative framework and to populate the ontology with this information as a community resource.

Biological Ontologies

Evaluating 12 automated, whole-genome sequencing analysis pipelines for Mycobacterium tuberculosis complex: a comparative study.

BACKGROUND: Reliance on complex, custom-built bioinformatics pipelines is a barrier to the implementation of whole-genome sequencing (WGS) of Mycobacterium tuberculosis in high-burden settings in some low-income and middle-income countries (LMICs). Automated analysis pipelines could address this inequity in access to WGS-based diagnostics and surveillance. This study aimed to systematically evaluate the performance and usability of publicly available WGS pipelines for M tuberculosis. METHODS: We identified automated M tuberculosis WGS analysis pipelines through searches of PubMed and GitHub from database inception up to Aug 31, 2024. Accuracy, cost, accessibility, and scalability were assessed for each pipeline. We evaluated the accuracy of genotypic drug susceptibility testing (gDST) using publicly available sequences with phenotypic susceptibility data for 12 antituberculosis drugs. We estimated pooled sensitivity and specificity for each pipeline, across all drugs, by conducting a bivariate meta-analysis, with random effects representing between-drug variability. Lineage classifications were compared, and a previously epidemiologically well-characterised dataset was used to compare measures of genomic relatedness. FINDINGS: Among 28 candidate pipelines, 16 were excluded as they were unmaintained and inexecutable. 12 pipelines (11 compatible with Illumina and four compatible with Nanopore), all free to use, were included for evaluation. Six pipelines processed and stored data remotely, but for five of these six, scalability was limited by the need to upload sequences through web portals. For local processing pipelines, scalability was dependent on substantial local computational resources, data storage capacity, and command-line interfaces that limited user-friendliness. Only one of six remote-processing pipelines removed human DNA sequences before server upload. gDST was similarly accurate across ten of 11 Illumina-compatible pipelines and three of four Nanopore-compatible pipelines. All pipelines classified the main lineages consistently, although there were differences at sublineage resolution. Outputs from three of four pipelines reporting genomic relatedness were compatible with commonly cited single nucleotide polymorphism difference thresholds. INTERPRETATION: Numerous automated analysis pipelines capable of enhancing equity in M tuberculosis WGS are available. Given the overall similarities between the pipelines evaluated in this study in terms of gDST performance, lineage classification, and genomic relatedness inference, non-functional attributes such as availability, accessibility, scalability, and privacy could represent the point of difference for prospective users in LMICs with a high burden of tuberculosis. FUNDING: The Rhodes Trust, Wellcome, Ellison Institute of Technology, and the UK National Institute for Health and Care Research Oxford Biomedical Research Centre.

Mycobacterium tuberculosis

ClarID: A Human-Readable and Compact Identifier Specification for Biomedical Metadata Integration.

BACKGROUND: In biomedical research, subjects and biospecimens are commonly tracked using simple IDs or UUIDs, which guarantee uniqueness but convey no embedded semantic information. Contextual metadata (such as tissue type, diagnosis, or assay) is often stored separately, making integration, cohort selection, and downstream analysis cumbersome. While structured barcoding systems exist in large consortia (e.g., TCGA, GTEx) or domain-specific contexts (e.g., SPREC, GOLD), no unified, extensible framework currently spans both subjects and biosamples in a human- and machine-readable way. METHODS: We developed ClarID, a domain-agnostic specification that supports two identifier formats: (i) a human-readable form (e.g., 'CNAG_Test-HomSap-00001-LIV-TUM-RNA-C22.0-TRT-P1W' that encodes key metadata such as project, species, subject_id, tissue, assay, disease, timepoint and duration (from that event); and (ii) a compact version named 'stub' (e.g., 'CT01001LTR0N401T1W') optimized for filenames, pipelines, and labeling.ClarID is implemented through an open-source command-line tool, ClarID-Tools, which processes tabular metadata files (CSV/TSV) and uses a YAML-based codebook to generate, decode, and validate identifiers, as well as to create and read QR codes. The tool supports bulk and single-sample processing and allows easy integration with institutional workflows. RESULTS: To demonstrate ClarID's utility, we applied it to datasets from the Genomic Data Commons (GDC), generating interpretable identifiers for more than 113,000 clinical records (subjects) and 4,255 biospecimen records. All materials, including pre-processing scripts, input and encoded data, are publicly available and fully reproducible via the accompanying GitHub repository and Google Colab. CONCLUSIONS: ClarID fills a critical gap between opaque accession numbers and rich metadata schemas by embedding key context directly into structured identifiers. It enhances traceability, facilitates downstream analysis, and remains adaptable to project-specific needs through a configurable codebook. The accompanying ClarID-Tools software is freely available, together with full documentation and reproducible pipelines, at https://github.com/CNAG-Biomedical-Informatics/clarid-tools.

Biosample identifiers

Artificial Intelligence for Colorectal Surgeons-Part II: Research Applications, Challenges in Adoption, and Practical Resources.

BACKGROUND: This is part II of a 2-part series examining artificial intelligence in colorectal surgery. Part I established foundational concepts and clinical applications. Implementation, however, requires understanding research methodologies, available resources, and the specific challenges currently limiting widespread adoption. These topics are the focus of part II. OBJECTIVE: To examine artificial intelligence's transformation of surgical research, provide practical implementation resources, address adoption challenges, and explore future directions in colorectal surgery. METHODS: Comprehensive literature review focusing on artificial intelligence research methodology, implementation barriers, educational resources, and emerging technologies relevant to colorectal surgeons. RESULTS: Artificial intelligence streamlines clinical trial design through predictive modeling and natural language processing, reducing enrollment challenges that contribute to failed or inadequate trial accrual. Machine learning enables heterogeneity analysis within clinical trials, identifying treatment-responsive subgroups. Foundation models unlock analysis of unstructured electronic health record data at scale. Professional societies and universities offer specialized artificial intelligence education programs, with open-access data sets facilitating research participation. However, implementation faces multifaceted challenges: technical infrastructure demands, with real-time processing requiring dedicated graphics processing unit clusters; regulatory frameworks struggling with continuously evolving algorithms; undefined liability distribution for artificial intelligence-assisted decisions; algorithmic bias risking health care disparities; and the "black box" problem limiting clinical trust. Economic barriers include substantial initial costs without clear reimbursement pathways. Future directions include multimodal artificial intelligence integrating imaging, genomics, and histopathology; cognitive robotic systems with real-time decision support; digital twin technology for patient-specific surgical simulation; and global surgical artificial intelligence networks enabling distributed learning across institutions. CONCLUSIONS: Although artificial intelligence offers transformative potential for colorectal surgery research and practice, successful implementation requires addressing technical, regulatory, ethical, and economic challenges. The surgeon's evolving role demands both traditional expertise and computational fluency. Future advances in multimodal integration, autonomous systems, and global collaboration will fundamentally reshape surgical practice but will require thoughtful implementation prioritizing patient benefit and clinical value.

Humans

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans

Adeno-Associated Virus Gene Therapy Translation: Lessons from Early Regulatory Meetings.

The Platform Vector-Gene Therapy (PaVe-GT) program is a National Institutes of Health (NIH) initiative that aims to develop adeno-associated virus (AAV) gene therapies for four monogenic rare diseases, two organic acidemias and two congenital myasthenic syndromes. PaVe-GT's platform-based approach identifies and diminishes redundancies and applies efficiencies in preclinical, clinical, and regulatory activities. The program's hypothesis is that implementing these efficiencies can accelerate clinical trial initiation. Based on its platform-centric experience and public-serving mission, the PaVe-GT program actively shares its scientific and regulatory learnings with the public to benefit the development of similar gene therapy products for rare diseases. PaVe-GT's first investigational AAV gene therapy candidate is AAV serotype 9 human propionyl-CoA carboxylase alpha subunit (AAV9-hPCCA) for propionic acidemia caused by PCCA deficiency, which received initial feedback from the Food and Drug Administration (FDA) in an INitial Targeted Engagement for Regulatory Advice on CBER/Center for Drug Evaluation and Research (CDER) ProducTs (INTERACT) meeting. Upon further product development that took into consideration the FDA's initial advice, the program obtained the Agency's feedback in pre-investigational new drug (IND) (Type B) and Type C meetings. Here, we share our experience from these meetings, including strategy, preparation, pre- and post-meeting feedback from the FDA, and lessons learned during the AAV9-hPCCA regulatory process, which the program plans to apply across the PaVe-GT platform. Topics discussed in the regulatory meetings included animal model and efficacy studies, toxicology study plans, manufacturing of the investigational AAV product, and clinical trial design. The main lessons learned from the pre-IND and Type C meetings for AAV9-hPCCA are: (1) Pharmacology/Toxicology studies in a single rodent species are sufficient for filing an initial IND; (2) FDA feedback guides product quality improvements and early development of a quantitative potency assay; (3) use of biomarkers as potential surrogate endpoints in a future efficacy trial benefits from collection of data in the natural history study and the first-in-human Phase 1/2 study; and (4) evidence from the Phase 1/2 clinical trial could be leveraged to support a license application. Lightly redacted regulatory documents and comprehensive templates developed by the PaVe-GT team are available on the PaVe-GT website.

Dependovirus