Search PubMedSearch

SEARCH · Search PubMed

Results for “AI discovery”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Rare variant analysis of whole genome sequenced juvenile idiopathic arthritis multiplex pedigrees identifies rare variants in NOD2 and ACVR1.

Juvenile idiopathic arthritis is a complex rheumatic disease that is influenced by environmental and genetic factors. Linkage and genome-wide association studies have identified genes that contribute to the risk of developing juvenile idiopathic arthritis but are limited in their ability to identify disease-risk variants of large effect. Penetrant, heritable risk variants can be detected in high-risk families, but such cases are uncommon due to the low prevalence of juvenile idiopathic arthritis. This study utilizes whole-genome sequencing of 23 multiplex families, the largest such cohort to date, to discover variants and genes relevant to JIA pathogenesis. Pathogenic variants in NOD2 associated with Blau syndrome, an ultra-rare Mendelian inflammatory disorder, are the most recurrent variants in the cohort, consistent with previous reports that milder presentations of Blau syndrome are oftentimes misdiagnosed as juvenile idiopathic arthritis. For the first time, however, rare variants in ACVR1 and SMAD6, integral components of the Bone Morphogenic Protein pathway, are found to be associated with juvenile idiopathic arthritis. Identified ACVR1 variants map to critical protein domains. AlphaFold modeling predicts that the ACVR1 interaction with its inhibitor OGT is disrupted by these variants, indicating that the patient-mutated protein has a gain-of-function phenotype. Drosophila melanogaster expressing either a wild-type or patient-mutated version of ACVR1 exhibit embryonic lethality, with the mutant exhibiting 1.4-fold greater lethality than wild-type. The combination of family-based cohorts for gene discovery, AI-based computational tools, and animal model studies for tests of variant function underscores shared disease pathogenesis between JIA and monogenic disorders of immunity and connective tissue.

Arthritis, Juvenile

Decoding cancer with artificial intelligence: Transforming research, diagnosis, and therapy with future insights.

Cancer remains one of the leading global health burdens, with increasing complexity in genomic, imaging, and clinical datasets presenting significant challenges for effective management. Artificial intelligence (AI) has emerged as a powerful tool to address these challenges by enabling pattern recognition, knowledge integration, and data-driven decision-making. This review highlights recent advances in the application of AI across cancer research, diagnosis, and therapy. In research, AI accelerates drug discovery and repurposing, enhances genomic data interpretation, and facilitates biomarker identification through multi-omics integration. In diagnosis, AI has demonstrated high technical performance in radiology for lesion detection and image segmentation, in pathology for tumour grading and molecular prediction, and in liquid biopsy for non-invasive biomarker analysis. In therapy, AI supports precision medicine by predicting treatment responses, monitoring disease progression, and optimizing clinical trial design. Despite these advances, barriers such as data heterogeneity, algorithmic bias, interpretability, and regulatory challenges remain. Future directions, including explainable AI, federated learning, multimodal modelling, and digital twins, hold promise for translating AI-driven innovations into routine oncology practice. Significance Statement This review provides a timely synthesis of recent (2020-2025) advances in artificial intelligence across cancer research, diagnosis, and therapy, highlighting applications in drug discovery, genomics, multi-omics biomarker identification, and clinical decision-making. By integrating technological progress with translational and clinical relevance, this work serves as a valuable resource for bridging AI innovation with precision oncology practice. As a narrative review, the literature was identified through targeted PubMed, Scopus, and Google Scholar searches, combining terms for artificial intelligence, machine learning, and deep learning with cancer-related keywords, with priority given to peer-reviewed studies published between 2020 and 2025, seminal earlier works, and official regulatory or guideline documents. Within each domain, representative studies were selected to illustrate methodological diversity, clinical context, and current translational readiness rather than to provide exhaustive coverage of an extremely rapidly evolving field.

Artificial intelligence

Community-driven advances in computational mass spectrometry: The perspective of EuBIC-MS members.

Advances in data acquisition, artificial intelligence, and integrative bioinformatics are driving the rapid evolution of computational mass spectrometry, and in turn, transforming modern proteomics, metabolomics, and lipidomics. These developments have greatly increased the scale and complexity of mass spectrometry data, underscoring the importance of evolving accurate, transparent, efficient and reproducible data processing workflows. Addressing these challenges requires collaborative innovation that brings together expertise in software engineering, statistics, and biology. The European Bioinformatics Community for Mass Spectrometry (EuBIC-MS), an initiative of the European Proteomics Association (EuPA), fosters a culture of open, community-driven development through its biennial Developers Meetings and Winter Schools. This commentary summarizes the scientific background and outcomes of the EuBIC-MS Developers Meeting 2025, which took place in Novacella, Italy. Three keynote presentations highlighted major frontiers in the field: deep proteome and phosphoproteome profiling, text mining for protein-protein interaction extraction, and scalable proteomics for AI-driven drug discovery. Seven community-selected hackathons addressed emerging challenges such as single-cell proteomics data analysis, FAIR metadata extraction, deep learning frameworks, R-Python interoperability, and DIA validation. Together, these efforts demonstrate the potential for scientific and technical innovation to arise from open collaboration, and highlight how community-driven initiatives can accelerate progress in computational mass spectrometry. SIGNIFICANCE: Modern proteomics increasingly depends on computational advances to translate complex, high-dimensional data into biological knowledge. The EuBIC-MS Developers Meeting 2025 exemplifies how community-driven collaboration can directly accelerate this process by bringing together experts from bioinformatics, statistics, and experimental proteomics to co-develop open, interoperable, and reproducible analytical tools. By fostering shared software frameworks, transparent benchmarking, and collaborative problem solving, the EuBIC-MS community helps ensure that technological innovation translates into reliable biological insights. This collaborative model strengthens the foundation for quantitative, system-level understanding of proteomes and establishes a sustainable path for integrating artificial intelligence and next-generation data acquisition into routine biological discovery. This commentary shows some current highlights in the field of computational mass spectrometry and community-based approaches undertaken during the most recent Developers Meeting to solve these challenges. The approaches discussed and initiated during the meeting - ranging from deep proteome profiling and phosphosite mapping to text mining, single-cell data analysis, and FAIR metadata extraction - address key bottlenecks that currently limit the biological interpretability and comparability of proteomics data.

Mass Spectrometry

Artificial Intelligence for Natural Products Discovery and Development.

Natural products (NPs) remain a cornerstone of modern drug discovery, offering stereochemical complexity and diverse bioactivities that precisely modulate therapeutic targets, refined through billions of years of evolution. However, their research has long been hindered by inefficient, empirical workflows, high resource consumption, structural complexity, and the "multicomponent, multi-target" nature of their mechanisms. The exponential growth of genomic, metabolomic, and spectral data has overwhelmed conventional analytical methods, exposing critical bottlenecks in handling high-dimensional, heterogeneous datasets that exceed human interpretive capacity. Artificial intelligence (AI) is emerging as a transformative paradigm to address these challenges, integrating multi-omics and chemical data to shift NP research from fragmented empiricism toward mechanism-driven, precision-oriented development. By leveraging deep learning architectures- including graph neural networks, Transformers, and diffusion-based generative models-AI enables systematic decoding of NP biosynthesis, automated structure elucidation, rational target identification, knowledge extraction from vast unstructured scientific literature, and de novo molecular design. This review comprehensively surveys recent advances in AI applications across the full NP discovery and development pipeline, encompassing genome mining, structure-based and ligand-based virtual screening, multimodal structural characterization, lead optimization, and biosynthetic pathway engineering. We further examine the emerging roles of protein-centric, molecule- centric, and multimodal foundation models, as well as large language models, in bridging genotype-to-chemotype gaps and unlocking unstructured scientific knowledge. Finally, we discuss critical challenges including data scarcity, representational limitations for complex stereochemistry, physical plausibility in generative models, and the urgent need for experimental validation, while outlining future directions toward autonomous experimentation, closed-loop optimization, and human-AI collaborative discovery.

Artificial intelligence

Identification of compounds that repress DUX4 expression in facioscapulohumeral muscular dystrophy.

Facioscapulohumeral muscular dystrophy (FSHD) is caused by epigenetic dysregulation of the disease locus, leading to pathogenic misexpression of DUX4 in skeletal muscle. Thus, most FSHD therapeutic approaches target DUX4. Our previous study identified the chromatin remodeling factor BAZ1A (bromodomain adjacent to zinc finger domain protein 1A) as a promising target for therapeutic development. Here we used an artificial intelligence-based screening pipeline to identify molecules predicted to bind the BAZ1A bromodomain, and validated hit compounds using FSHD-specific assays in FSHD myocytes. One compound, termed C06, emerged as a potent repressor of DUX4 and DUX4 target gene expression. Interestingly, while C06 exhibited binding to BAZ1A in vitro, it can also inhibit multiple kinases, including p38α, an upstream activator of DUX4. Despite this, at low doses C06 was an equally effective and more specific repressor of DUX4 than losmapimod, which is a robust and specific p38 inhibitor. At low concentrations, C06 returns the DUX4 gene expression signature to a healthier profile without major effects on the muscle transcriptome. Thus, C06 is a useful tool for potent and specific DUX4 suppression, and a viable candidate for further development. Our results highlight both the utility and limitations of AI for targeted drug discovery, and the importance of using an FSHD-specific functional screening strategy for selecting relevant candidates.

Muscular Dystrophy, Facioscapulohumeral

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis

Uncovering heterogeneous effects via localized feature selection.

Identifying features that interact to trigger disease, while accounting for heterogeneity across diverse populations, is essential for the development of precision and targeted medicine. Despite the availability of vast and complex health-related datasets, most existing works focus on identifying disease-associated features at the population level or within a few subpopulations, often overlooking individual-level heterogeneity within these groups. To address this limitation, we propose a framework that utilizes localized test statistics to identify disease-associated features tailored to individual profiles. Our method leverages the recently developed knockoffs methodology to control the noise level of the selection set so that the results are replicable. Moreover, it allows for the discovery of hidden heterogeneous effects within the data, as demonstrated in an application to single-cell RNA sequencing data for Alzheimer's disease. By aggregating localized feature selection results, our framework also enables powerful population-level feature selection. Our framework provides a powerful tool for exploratory studies of precision medicine, offering the potential to generate novel hypotheses for confirmatory biological experiments.

Alzheimer Disease

Listening forward: emerging roles of bioacoustics in ecology, evolution, and conservation.

Bioacoustics is increasingly shifting from a mostly descriptive pursuit to one that can anticipate ecological change. Recent innovations-from autonomous recording units and edge-computing sensors to speech-inspired feature extraction and machine-learning techniques like transfer learning, unsupervised discovery, and explainable AI-are transforming the study of animal communication. These advances let us work at scales previously difficult to imagine. Automated species recognition, individual identification, and even tracking cultural evolution over decades are now within reach. Entire ecosystem soundscapes can be mapped with unprecedented resolution. Looking ahead, global listening networks, adaptive acoustic indices, and live biodiversity dashboards seem increasingly realistic. We may soon build digital models that simulate communication networks under future scenarios. Closer integration with genomics, physiology, and robotics could link vocal traits to their genetic, physiological, and ecological drivers. Challenges remain, including data governance, acoustic privacy, and equitable access to the planet's sonic heritage. Bioacoustics may be on the way to becoming a predictive, integrative science - one particularly well suited to monitoring, interpreting, and helping safeguard life's communication systems in a rapidly changing world.

Animals

ELISA (Embedding-Linked Interactive Single-cell Agent): an interpretable hybrid generative Artificial Intelligence agent for expression-grounded discovery in single-cell genomics.

Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain opaque to natural language. Here, we introduce ELISA (Embedding-Linked Interactive Single-cell Agent), an interpretable framework that unifies single-cell generative pretrained transformer expression embeddings with biomedical bidirectional encoder representations from transformers-based semantic retrieval and large-language model (LLM)-mediated interpretation for interactive single-cell discovery. An automatic query classifier routes inputs to gene marker scoring, semantic matching, or reciprocal rank fusion pipelines depending on whether the query is a gene signature, natural language concept, or mixture of both. Integrated analytical modules perform pathway activity scoring across 60+ gene sets, ligand-receptor interaction prediction using 280+ curated pairs, condition-aware comparative analysis, and cell-type proportion estimation, all operating directly on embedded data without access to the original count matrix. Benchmarked across six diverse scRNA-seq datasets spanning inflammatory lung disease, pediatric and adult cancers, organoid models, healthy tissue, and neurodevelopment, ELISA significantly outperforms CellWhisperer, a classical lexical retriever (BM25), and a random baseline in cell type retrieval (combined permutation test, $p < 2\times 10^{-5}$ for each), with particularly large gains on gene-signature queries (Cohen's $d = 5.98$ for mean reciprocal rank). ELISA replicates published biological findings (mean composite score 0.88), and generates candidate hypotheses through grounded LLM reasoning, bridging the gap between transcriptomic data exploration and biological discovery.

Generative Artificial Intelligence

A comprehensive review of AI innovations for tackling antimicrobial resistance.

Antimicrobial resistance (AMR) represents a major global public health concern, rendering available antimicrobials ineffective and leading to infections that are difficult to treat. Artificial intelligence (AI) has been increasingly applied across the AMR continuum, including resistance prediction, rapid diagnostics, new antimicrobial discovery, drug repurposing, antimicrobial surveillance, and clinical decision support. In this review, we aim to highlight recent developments in the use of artificial intelligence (AI) to address antimicrobial resistance (AMR). In addition, we review computational methods that help interpret genomic, phenomic, clinical, and epidemiological data to support the development of treatment strategies and novel antimicrobial agents. The key issues addressed include data quality, model interpretability, external validation, regulatory requirements, privacy, and fairness. While AI is not a complete solution to AMR, it can certainly strengthen the global AMR response by complementing key areas of AMR such as antimicrobial stewardship, infection prevention, laboratory diagnostics, and global surveillance.

Antimicrobial resistance (AMR)

Artificial intelligence agents and agentic artificial intelligence applied to precision medicine.

Precision medicine seeks to individualise care by integrating multimodal biomedical data, yet most deployed clinical artificial intelligence (AI) remains assistive, providing predictions without managing workflows or adapting autonomously. Agentic AI, built on large language models (LLMs), has emerged as a paradigm characterised by autonomy, goal-directed reasoning, memory, planning and tool use. This review synthesises evidence on agentic AI and LLMs applied to precision medicine, encompassing drug discovery, genomics, oncology, rare disease diagnostics and clinical pharmacology. This review also examines architectural components, recent validation milestones and emerging challenges, including hallucination, sociodemographic bias and evolving regulatory frameworks across the FDA, the EU AI Act and the WHO.

agentic AI

Artificial intelligence in healthcare and medicine: clinical applications, therapeutic advances, and future perspectives.

Healthcare systems worldwide face growing challenges, including rising costs, workforce shortages, and disparities in access and quality, particularly in low- and middle-income countries. Artificial intelligence (AI) has emerged as a transformative tool capable of addressing these issues by enhancing diagnostics, treatment planning, patient monitoring, and healthcare efficiency. AI's role in modern medicine spans disease detection, personalized care, drug discovery, predictive analytics, telemedicine, and wearable health technologies. Leveraging machine learning and deep learning, AI can analyze complex data sets, including electronic health records, medical imaging, and genomic profiles, to identify patterns, predict disease progression, and recommend optimized treatment strategies. AI also has the potential to promote equity by enabling cost-effective, resource-efficient solutions in low-resource and remote settings, such as mobile diagnostics, wearable biosensors, and lightweight algorithms. Successful deployment requires addressing critical challenges, including data privacy, algorithmic bias, model interpretability, regulatory oversight, and maintaining human clinical oversight. Emphasizing scalable, ethical, and evidence-driven implementation, key strategies include clinician training in AI literacy, adoption of resource efficient tools, global collaboration, and robust regulatory frameworks to ensure transparency, safety, and accountability. By complementing rather than replacing healthcare professionals, AI can reduce errors, optimize resources, improve patient outcomes, and expand access to quality care. This review emphasizes the responsible integration of AI as a powerful catalyst for innovation, sustainability, and equity in healthcare delivery worldwide.

Humans

Out-of-the-box bioinformatics capabilities of large language models (LLMs).

Large Language Models (LLMs), AI agents and co-scientists promise to accelerate scientific discovery across fields ranging from chemistry to biology. Bioinformatics- the analysis of DNA, RNA and protein sequences plays a crucial role in biological research and is especially amenable to AI-driven automation given its computational nature. Here, we assess the bioinformatics capabilities of three popular general-purpose LLMs on a set of tasks covering basic analytical questions that include code writing and multi-step reasoning in the domain. Utilizing questions from Rosalind, a bioinformatics educational platform, we compare the performance of the LLMs vs. humans on 104 questions undertaken by 110 to 68,760 individuals globally. GPT-3.5 provided correct answers for 59/104 (58%) questions, while Llama-3-70B and GPT-4o answered 49/104 (47%) correctly. GPT-3.5 was the best performing in most categories, followed by Llama-3-70B and then GPT-4o. 71% of the questions were correctly answered by at least one LLM. The best performing categories included DNA analysis, while the worst performing were sequence alignment/comparative genomics and genome assembly. Overall, LLMs performance mirrored that of humans with lower performance in tasks in which humans had low performance and vice versa. However, LLMs also failed in some instances where most humans were correct and, in a few cases, LLMs excelled where most humans failed. To the best of our knowledge, this presents the first assessment of general purpose LLMs on basic bioinformatics tasks in distinct areas relative to the performance of hundreds to thousands of humans. LLMs provide correct answers to several questions that require use of biological knowledge, reasoning, statistical analysis and computer code.

Journal Article

BioMedGraphica: An All-in-One Platform for Joint Textual Biomedical Prior Knowledge and Numeric Graph Generation.

Multi-omic data analysis is essential for scientific discovery in precision medicine. However, translating statistical results of omic data analysis into novel scientific hypothesis remains a significant challenge. Human experts must manually review analysis results and generate new hypothesis based on extensive and inter-connected biomedical prior knowledge, which is subjective and not scalable. While large language models (LLMs) can accelerate the discovery, their reasoning improves when grounded in structured, auditable and comprehensive biomedical prior knowledge. Biomedical knowledge, however, is scattered across heterogeneous databases that use diverse and inconsistent nomenclature systems, making it difficult to integrate resources into a unified format for scalable analysis. This fragmentation limits the ability of AI systems to fully leverage biomedical data for scientific discovery. To address these challenges, we developed BioMedGraphica , an all-in-one platform that harmonizes fragmented biomedical resources by integrating 11 entity types and 30 relation types from 43 databases into a unified knowledge graph containing 2,306,921 entities and 27,232,091 relations. In addition, to the best of our knowledge, this is the first work to propose a novel Textual-Numeric Graph (TNG) data-structure for multi-omics data analysis. In TNG, textual information captures prior biological knowledge (e.g., transcription start sites, functions, mechanisms), while numeric values represent quantitative biomedical features, and the integrated relations can help uncover mechanisms. By bridging prior knowledge with user-specific data, TNG is a novel and ideal data-structure for the development of graph foundation models, with the potential to improve prediction performance and interpretability, while also augmenting LLMs by supplying graph-structured mechanistic context to strengthen reasoning. The details for BioMedGraphica code can be accessed by github link: https://github.com/FuhaiLiAiLab/BioMedGraphica and BioMedGraphica knowledge graph data can be downloaded from huggingface dataset: https://huggingface.co/datasets/FuhaiLiAiLab/BioMedGraphica.

biomedical knowledge graph

Plant cis-regulatory grammar: Decoding the multidimensional code of transcriptional regulation for programmable crop engineering.

Cis-regulatory elements (CREs) orchestrate the spatiotemporal precision of gene expression that underlies plant development, adaptation, and domestication. Decoding the cis-regulatory grammar of plant genomes remains a central challenge in modern biology, with profound implications for programmable crop engineering. Here, recent conceptual and technological advances are synthesized to reshape our understanding of plant CREs. This review first argues that CRE function is not only an intrinsic property of DNA sequence alone but also emerges from a multidimensional context, including chromatin accessibility, histone modifications, three-dimensional genome topology, and cell type-specific regulatory landscapes. Furthermore, the convergence of single-cell epigenomics, high-throughput functional assays, and CRISPR-based dissection has begun to unravel this contextual grammar, revealing the computational principles governing transcriptional regulation. Critically, we propose that artificial intelligence (AI) platforms are catalyzing an ongoing transition from descriptive discovery to predictive engineering, wherein these platforms outperform natural evolution in designing synthetic CREs. Finally, a roadmap is outlined toward a plant regulatory grammar foundation model, which will enable truly predictive engineering of gene expression when fine-tuned for specific tasks. Collectively, the integration of single-cell resolution maps, precise genome editing, AI-driven design, and regulatory-compliant delivery systems promises to transform our ability to reprogram plant gene regulation for next-generation agriculture, bridging the gap between foundational regulatory biology and tangible crop improvement.

artificial intelligence

The AI Revolution: Shaping the Present and Future of Pharmaceutical Research and Development.

The transformative role of artificial intelligence (AI) in the pharmaceutical industry is examined, with a focus on its significant contributions to drug discovery, development, and clinical trial processes. It highlights the inefficiencies and high costs associated with traditional drug development and explores how AI and machine learning (ML) can enhance these processes by analyzing extensive biological datasets. The historical context of AI in pharmaceutical development is examined, noting how advances in computational power and data accessibility have facilitated innovative methodologies, such as predictive analytics and natural language processing. Contemporary trends reveal the integration of AI technologies in drug design, repurposing, and patient response forecasting. This study also addresses the challenges of participant recruitment for clinical trials and proposes AI-driven solutions to optimize patient selection and data management. Furthermore, it discusses AI's role in tailored medicine, emphasizing its potential for advancing precision therapy through targeted drug development and personalized treatment strategies. The importance of digital tools, genomic data analysis, and AI-driven imaging technologies for customizing therapeutic approaches is underscored, along with the regulatory and ethical challenges posed by AI deployment in healthcare. This study illustrates the complexities of AI applications in the pharmaceutical sector, offering insights into both successful and unsuccessful initiatives. The findings suggest that the digitalization of the pharmaceutical industry and enhanced AI integration hold promise for developing safer and more effective therapeutic strategies, while also identifying obstacles to their widespread adoption and optimal functionality.

Artificial intelligence