Search PubMedSearch

SEARCH · Search PubMed

Results for “foundation model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Biological Foundation Models for Complex Disease Research and Clinical Translation.

Complex diseases, including cancer, rare genetic disorders, neurodevelopmental and psychiatric conditions, and neurodegenerative diseases, arise from interactions among genetic variation, gene regulation, and cellular states that are difficult to capture using a single data type or biological scale. Biological foundation models address this challenge by treating nucleotides and genes as tokens and learning representations that can be transferred to downstream biomedical and clinical tasks. In this review, we examine two major model classes, genomic sequence foundation models and cell foundation models, and compare their tokenization strategies, model architectures, pretraining objectives, and adaptation methods. We summarize their emerging applications in regulatory variant interpretation, disease-associated cell-state analysis, drug-response prediction, and therapeutic target discovery across complex diseases. We distinguish applications supported by experimental or retrospective validation from those that remain primarily computational or conceptual. We further discuss key challenges to clinical translation, including multimodal data integration, model interpretability, benchmarking, patient-specific prediction, and privacy protection. We highlight future opportunities to integrate biological foundation models with emerging frameworks of medical digital twins, agentic AI, and federated learning. By linking model design to translational goals, this review provides a practical framework for evaluating biological foundation models and their readiness for complex disease research and clinical use.

biological foundation model

Causal circuit tracing reveals distinct computational architectures in single-cell foundation models: inhibitory dominance, biological coherence, and cross-model convergence.

MOTIVATION: Sparse autoencoders (SAEs) decompose foundation-model activations into interpretable features, but the model-internal causal interactions between those features (i.e. what ablating one feature does to the others, as distinct from the biological causal structure of the underlying cells)-and how those model-internal relationships relate to biological structure-are uncharacterized in single-cell foundation models. RESULTS: We introduce model-internal causal circuit tracing-zeroing one SAE feature at a source layer and measuring the resulting change in all downstream SAE features, for each of 120 source features-and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96&#xa0;892 ablation-derived edges, 80&#xa0;191 forward passes). On annotation-selected source features, edges share GO/KEGG/Reactome/STRING/TRRUST ontology terms at 50.9%-68.5%, a 2.9-6.2&#xd7; enrichment over a configuration-preserving permutation null (P<.002); on 20 randomly sampled source features this attenuates to 21.5%-26.3%-still 2.5-3.1&#xd7; above null-quantifying the annotation-selection contribution. Inhibitory dominance (fraction of ablation edges with d<0, i.e. source activation supports downstream target) is 65.5%-89.4%. scGPT produces larger raw per-edge effects (mean |d|=1.40 versus 1.05); after feature-share normalization, Geneformer is stronger (paired gene-pair ratio 0.64 on 33&#xa0;301 shared pairs). Cross-model consensus yields 1142 architecture-invariant domain pairs (ordered pairs of GO biological-process categories "A&#x2192;B" each connected by at least one ablation edge in both models; 10.6&#xd7; enrichment over permutation null; P<.001). Circuit edge magnitude explains <1% of the variance in marginal driver-gene coexpression on the same cells (R2=0.010, n=31&#xa0;176): the graph encodes structure beyond bivariate correlation. Against a matched-cell-type ENCODE ChIP-seq prior, circuit-predicted transcription factor (TF)&#x2192;target pairs are enriched 2.06&#xd7; (Fisher OR 5.84), markedly higher than 1.12&#xd7; against TRRUST; direct ChIP-seq-supported target pairs show 10-30&#xd7; larger CRISPRi sign-bias-corrected excess than indirect pairs. Gene-level CRISPRi validation on Replogle K562 and the noncancer RPE1 arm (and a true primary-T-cell control from Shifrut E, Carnevale J, Tobin V et&#xa0;al. Genome-wide CRISPR screens in primary human T cells reveal key regulators of immune function. Cell 2018; 175: 1958-71.e15) after sign-bias correction shows excess over baseline of +0.03 and +0.35 percentage points on K562 and RPE1, respectively (baseline already 52%-56% from sign marginals); effect-magnitude Spearman correlations &#x3c1;&#x2248;0. Bootstrap and per-cell-type stability (N&#x2208;{50,100,200}; B cell, CD4&#xa0;+ T, macrophage) give Pearson r&#x2265;0.97 on shared edges with 100% sign agreement; edge Jaccard grows monotonically with sample size. The circuit graph is therefore highly reproducible as an effect-size map, cell type specific in edge identity, consistent with coexpression encoding, and weakly but detectably enriched for ChIP-seq-supported direct regulatory edges. AVAILABILITY AND IMPLEMENTATION: https://github.com/Biodyn-AI/bio-sae-circuits (Python). Archival DOI: 10.5281/zenodo.19,633,166 (Zenodo).

Humans

Orthrus: Towards Evolutionary and Functional RNA Foundation Models.

In the face of rapidly accumulating genomic data, our ability to accurately predict key mature RNA properties that underlie transcript function and regulation remains limited. Pre-trained genomic foundation models offer an avenue to adapt learned RNA representations to biological prediction tasks. However, existing genomic foundation models are trained using strategies borrowed from textual domains that do not leverage biological domain knowledge. Here, we introduce Orthrus, a Mamba-based mature RNA foundation model pre-trained using a novel self-supervised contrastive learning objective with biological augmentations. Orthrus is trained by maximizing embedding similarity between curated pairs of RNA transcripts, where pairs are formed from splice isoforms of 10 model organisms and transcripts from orthologous genes in 400+ mammalian species from the Zoonomia Project. This training objective results in a latent representation that clusters RNA sequences with functional and evolutionary similarities. We find that the generalized mature RNA isoform representations learned by Orthrus significantly outperform genomic foundation models on mRNA property prediction tasks, and requires only a fraction of fine-tuning data to do so. Finally, we show that Orthrus is capable of capturing divergent biological function of individual transcript isoforms.

Journal Article

MutBERT: probabilistic genome representation improves genomics foundation models.

MOTIVATION: Understanding the genomic foundation of human diversity and disease requires models that effectively capture sequence variation, such as single nucleotide polymorphisms (SNPs). While recent genomic foundation models have scaled to larger datasets and multi-species inputs, they often fail to account for the sparsity and redundancy inherent in human population data, such as those in the 1000 Genomes Project. SNPs are rare in humans, and current masked language models (MLMs) trained directly on whole-genome sequences may struggle to efficiently learn these variations. Additionally, training on the entire dataset without prioritizing regions of genetic variation results in inefficiencies and negligible gains in performance. RESULTS: We present MutBERT, a probabilistic genome-based masked language model that efficiently utilizes SNP information from population-scale genomic data. By representing the entire genome as a probabilistic distribution over observed allele frequencies, MutBERT focuses on informative genomic variations while maintaining computational efficiency. We evaluated MutBERT against DNABERT-2, various versions of Nucleotide Transformer, and modified versions of MutBERT across multiple downstream prediction tasks. MutBERT consistently ranked as one of the top-performing models, demonstrating that this novel representation strategy enables better utilization of biobank-scale genomic data in building pretrained genomic foundation models. AVAILABILITY AND IMPLEMENTATION: https://github.com/ai4nucleome/mutBERT.

Humans

Accelerating inference in genomic and proteomic foundation models via speculative decoding.

MOTIVATION: Genomic and protein foundation models (GFMs and PFMs) have demonstrated strong performance in learning the language of DNA and proteins, but their use in large-scale sequence generation is limited by the latency of autoregressive decoding. Because every token triggers a forward pass of a large Transformer, whose inference is relatively slow, long-sequence generation quickly becomes costly. RESULTS: In this work we adapt speculative decoding to a representative GFM: the DNA model DNAGPT and two representative PFMs: ProGen2 and ProtGPT2. We implement a probabilistic variant of speculative decoding, in which a lightweight draft model proposes short token spans and a larger target model verifies or corrects them in parallel, while preserving the target model's sampling distribution. Across all three models we systematically study the effect of speculation window length, temperature, draft architecture and prompt length, and we benchmark tokens per second over multiple runs per configuration. Speculative decoding yields consistent speedups over standard key-value cached decoding, with maximum observed speedup reaching 100% increase, while average gains across models ranging between 20% and 40% (e.g. 1.2&#xd7;-1.4&#xd7;), without changing the underlying target model predictions. Our results show that speculative decoding is a practical and model-agnostic strategy for accelerating genomic and proteomic sequence generation without sacrificing prediction quality. AVAILABILITY AND IMPLEMENTATION: All code and results are freely available at https://github.com/Georgakopoulos-Soares-lab/BioSpecDec.

Genomics

Bridging Ancestry Gaps in Genomic Risk Prediction with Tabular Foundation Models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Ancestry Continuum

Multimodal Deep Learning and Foundation Models for Early Detection and Forecasting of Plant Diseases.

Plant diseases destroy 20-40% of global food production annually, posing a critical threat to food security for a projected population of 9.7 billion by 2050. Conventional diagnostic approaches relying on expert visual assessment are slow, costly, and unsuitable for modern agricultural scales. While deep convolutional neural networks demonstrated early promise, single-modality, image-centric systems consistently fail under real-world field conditions characterized by variable lighting, co-occurring infections, and cultivar diversity. This review synthesizes a decade of progress across four interconnected frontiers: the evolution of deep learning architectures for plant disease detection; the adaptation of foundation models including CLIP, SAM, and DINOv2 to agricultural contexts; the development of multimodal fusion frameworks integrating imagery, environmental, genomic, and hyperspectral data; and the transition from static disease diagnosis to descriptive comparison of reported metrics, which suggested that multimodal approaches frequently reported improved diagnostic performance relative to corresponding single-modality baselines, although direct cross-study comparison was limited by methodological heterogeneity. A systematic review following PRISMA guidelines identifies eligible comparative studies. Descriptive comparison of reported performance metrics across these studies indicated that multimodal approaches generally achieved higher accuracy and sensitivity than single-modality models, particularly for pre-symptomatic disease detection. Eight critical research gaps are identified, including the absence of a unified agricultural foundation model and limited climate-aware forecasting under non-stationary climate projections. A structured research agenda is proposed to accelerate translation from laboratory performance to globally equitable, field-deployable crop protection systems.

convolutional neural networks

Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms.

Contrastive learning is a widely adopted self-supervised pretraining strategy, yet its dependence on cohort composition remains underexplored. We present Contrasting by Augmented Patient Electrocardiograms (CAPE) foundation model and pretrain on four cohorts (n&#x2009;=&#x2009;5,203,269), from diverse populations across three continents (North America, South America, Asia). We systematically assess how cohort demographics, health status, and population diversity influence the downstream performance for prediction tasks also including two additional cohorts from another continent (Europe). We find that downstream performance depends on the distributional properties of the pretraining cohort, including demographics and health status. Moreover, while pretraining with a multi-centre, demographically diverse cohort improves in-distribution accuracy, it reduces out-of-distribution (OOD) generalisation of our contrastive approach by encoding cohort-specific artifacts. To address this, we propose the In-Distribution Batch (IDB) strategy, which preserves intra-cohort consistency during pretraining, discourages learning of spurious cohort-specific features, and instead promotes clinically meaningful variability within cohorts. This leads to improved out-of-distribution robustness, with gains of 9-40% in downstream label prediction performance. This work provides insights into pretraining strategies for more clinically deployable and generalisable foundation models.

Journal Article

Real-world deployment of a fine-tuned pathology foundation model for lung cancer biomarker detection.

Artificial intelligence models using digital histopathology slides stained with hematoxylin and eosin offer promising, tissue-preserving diagnostic tools for patients with cancer. Despite their advantages, their clinical utility in real-world settings remains unproven. Assessing EGFR mutations in lung adenocarcinoma demands rapid, accurate and cost-effective tests that preserve tissue for genomic sequencing. PCR-based assays provide rapid results but with reduced accuracy compared with next-generation sequencing and require additional tissue. Computational biomarkers leveraging modern foundation models can address these limitations. Here we assembled a large international clinical dataset of digital lung adenocarcinoma slides (N&#x2009;=&#x2009;8,461) to develop a computational EGFR biomarker. Our model fine-tunes an open-source foundation model, improving task-specific performance with out-of-center generalization and clinical-grade accuracy on primary and metastatic specimens (mean area under the curve: internal 0.847, external 0.870). To evaluate real-world clinical translation, we conducted a prospective silent trial of the biomarker on primary samples, achieving an area under the curve of 0.890. The artificial-intelligence-assisted workflow reduced the number of rapid molecular tests needed by up to 43% while maintaining the current clinical standard performance. Our retrospective and prospective analyses demonstrate the real-world clinical utility of a computational pathology biomarker.

Humans

NextVir: Enabling classification of tumor-causing viruses with genomic foundation models.

MOTIVATION: Oncoviruses, pathogens known to cause or increase the risk of cancer, include both common viruses such as human papillomaviruses and rarer pathogens such as human T-lymphotropic viruses. Computational methods for detecting viral DNA from data acquired by modern DNA sequencing technologies have enabled studies of the association between oncoviruses and cancers. Those studies are rendered particularly challenging when multiple species of oncovirus are present in a tumor sample. In such scenarios, merely detecting the presence of a sequencing read of viral origin is insufficiently informative-instead, a more precise characterization of the viral content in the sample is required. RESULTS: We address this need with NextVir, to our knowledge the first multi-class viral classification framework that adapts genomic foundation models to detecting and classifying sequencing reads of oncoviral origin. Specifically, NextVir explores several foundation models-DNABERT-S, Nucelotide Transformer, and HyenaDNA-and efficiently fine-tunes them to enable accurate identification of the sequencing reads' origin. The results demonstrate superior performance of the proposed framework over existing deep learning methods and suggest downstream potential for foundational models in genomics.

Humans

Bridging ancestry gaps in genomic risk prediction with tabular foundation models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Humans

Foundation model enables interpretable open and error-tolerant searching for mass spectrometry-based proteomics.

MOTIVATION: Mass spectrometry-based proteomics allows studying all proteins of a sample on a molecular level. However, mass spectra are noisy and contain complex patterns, making them inherently challenging to analyze with algorithmic approaches. In terms of the protein sequence landscape, most recent bottom-up MS-based proteomics studies consider either a diverse pool of post-translational modifications, employ large databases-as in metaproteomics or proteogenomics, study multiple isoforms of proteins, include unspecific cleavage sites or even combinations thereof. All this makes peptide and protein identifications challenging. RESULTS: Here, we present a foundation model, called yHydra, that jointly embeds spectra and peptides. This allows us to implement various downstream tasks and search modes in Euclidean space. We implement an open search which allows querying multiple ten-thousands of spectra against millions of peptides. Furthermore, we implement an error-tolerant search for identifying additional proteoforms that are not included in off-the-shelf reference proteomes. Our foundation model provides meaningful embeddings, as we interpret learned peptide embeddings in comparison to the peptide's physico-chemical properties. Hydra's open search, assigns delta masses to each identification which allows to unrestrictedly characterize post-translational modifications. The error-tolerant mode of yHydra can be used as post-processing to existing search engines or as a standalone. yHydra is evaluated on several real life data sets for the identification of modified peptide sequences and shows up to 25% increase in peptide identification at constant false discovery rate compared to the current state-of-the-art. AVAILABILITY AND IMPLEMENTATION: Code is available on Gitlab: https://gitlab.com/dacs-hpi/yHydra, and https://gitlab.com/dacs-hpi/yHydra_train.

Proteomics

A Foundation Model Based CT Biomarker for Non-Invasive Prediction of Response to Neoadjuvant Immunochemotherapy in Non-Small Cell Lung Cancer.

Predicting pathological complete response (pCR) to neoadjuvant immunochemotherapy in non-small cell lung cancer (NSCLC) is clinically important yet remains challenging. Here, we introduce a foundation model-derived computed tomography (CT) imaging biomarker established from a multi-center cohort of 702 patients. Specifically, we developed and validated a non-invasive baseline CT-based model for risk stratification of pathological response. To address scanner and protocol heterogeneity, we first built a 3D Vision Mamba-based CT super-resolution model trained on 2494 cases for image standardization. We then fine-tuned a lung cancer-specific CT foundation model from a pretrained 3D model (VoCo) using 6643 chest CT scans. Finally, we constructed a multi-task Swin Transformer that jointly performs risk stratification and segments tumors to generate the imaging biomarker. Across five centers, the model achieved consistently strong generalization (AUC: 0.75-0.87) for pCR prediction. Genomic analysis revealed that the biomarker was independent of tumor mutational burden but significantly associated with TP53 mutations, suggesting an association with a radiogenomic phenotype related to this alteration. Together, these results demonstrate a generalizable and biologically meaningful foundation model-based biomarker for non-invasive risk stratification of pathological response in NSCLC.

Female

scPlantLLM: A Foundation Model for Exploring Single-cell Expression Atlases in Plants.

Single-cell RNA sequencing (scRNA-seq) provides unprecedented insights into plant cellular diversity by enabling high-resolution analyses of gene expression at the single-cell level. However, the complexity of scRNA-seq data, including challenges in batch integration, cell type annotation, and gene regulatory network (GRN) inference, demands advanced computational approaches. To address these challenges, we developed scPlantLLM, a Transformer model trained on millions of plant single-cell data points. Using a sequential pretraining strategy incorporating masked language modeling and cell type annotation tasks, scPlantLLM generates robust and interpretable single-cell data embeddings. When applied to Arabidopsis thaliana datasets, scPlantLLM excels in clustering, cell type annotation, and batch integration, achieving an accuracy of up to 0.91 in zero-shot learning scenarios. Furthermore, the model demonstrates an ability to identify biologically meaningful GRNs and subtle cellular subtypes, showcasing its potential to advance plant biology research. Compared to traditional methods, scPlantLLM outperforms in key metrics such as adjusted rand index (ARI), normalized mutual information (NMI), and silhouette score (SIL), highlighting its superior clustering accuracy and biological relevance. scPlantLLM represents a foundation model for exploring plant single-cell expression atlases, offering unprecedented capabilities to resolve cellular heterogeneity and regulatory dynamics across diverse plant systems. The code used in this study is available at https://github.com/compbioNJU/scPlantLLM.

Single-Cell Analysis

Foundation model reveals the shared organization of transcription and topologically associating domains.

The three-dimensional organization of chromatin into topologically associating domains (TADs) may impact gene regulation by bringing distant genes into contact. However, studies of TADs' function and their influence on transcription have been constrained by ambiguities in TAD boundary definitions and challenges in directly measuring their regulatory effects. We overcome these limitations by developing species-level consensus TAD maps for human and mouse by using a bag-of-genes approach that exposes an emergent regulatory structure. To quantify TAD-mediated relationships, we use a foundation model trained on 33 million transcriptomes to define a contextual similarity metric that captures higher-order relationships missed by co-expression. We find that TADs are regions of elevated co-regulation, with our framework yielding testable hypotheses about chromatin organization across cellular states. This TAD-linked enhancement is strongest during early development and declines with aging, while cancer cells show distinct TAD usage that shifts with chemotherapy. Together, these findings suggest that chromatin organization acts through probabilistic rather than deterministic mechanisms.

Humans

Foundation model based multimodal transformer framework for survival analysis in HER2 stratified breast cancer.

Objective. To improve survival prediction for HER2-positive breast cancer by integrating histopathological, molecular, and clinical data using a multimodal transformer framework.Approach. We propose a multimodal transformer framework for breast cancer survival prediction using HER2 stratified (SurvMBC), a foundation model-enhanced architecture that fuses three data modalities: whole-slide images, clinical narratives, and molecular features. Tumor microenvironment features are extracted using a pathology language and image pre-training (PLIP), clinical narratives are processed with BioBERT, and miRNA expression plus DNA methylation data are embedded using Gen2Vec. These representations are integrated through a cross-modal transformer with attention mechanisms for survival prediction.Main results. The model was evaluated on 1,095 HER2-positive breast cancer patients from The Cancer Genome Atlas. SurvMBC achieved a concordance index (C-index) of 0.857 (95% CI: 0.834, 0.880), a low integrated Brier score, and a strong inverse negative binomial log-likelihood. Risk stratification based on model outputs significantly separated high- and low-risk groups (log-rankp< 0.01) and showed strong associations with tumor stage, grade, and hormone receptor status (allp< 0.05).Significance. SurvMBC demonstrates the effectiveness of multimodal fusion in addressing tumor heterogeneity and improving prognostic accuracy. The attention-based integration enables context-aware learning of survival-relevant features across modalities, supporting individualized risk stratification and risk-adaptive treatment planning for HER2 stratified breast cancer patients.

Breast Neoplasms

GBFN: A gated bimodal fusion network leveraging foundation model embeddings for cancer drug sensitivity prediction.

Despite recent progress in deep learning for cancer drug sensitivity prediction, many existing models still rely on task-specific representation learning or relatively simple multimodal fusion, which may limit their ability to capture complex drug-cell interactions. To address this issue, we developed GBFN, a gated bimodal fusion network for continuous IC50 prediction that integrates pretrained drug and cell-line representations. Specifically, drug embeddings were obtained from SMI-TED, whereas cell-line embeddings were derived from transcriptomic profiles using BulkFormer. These two modalities were then combined through a dimension-wise gated fusion module and used to predict IC50 values in matched drug-cell line pairs. On the CCLE-based benchmark, GBFN outperformed representative neural baselines, including GraphDRP, TGSA, and TransEDRP, and achieved the best overall performance, with an R&#xb2; of 0.8714 and an RMSE of 0.8938. Moreover, ablation analysis showed that the model using drug features and cell-line expression data with gated fusion performed better than the corresponding model using direct concatenation, indicating that the improvement was associated with the fusion strategy rather than with the input modalities alone. In addition, cell-line expression data were more informative than mutation data in the present setting, and adding mutation data to the model using drug features and expression data did not further improve performance. Across major cancer types, GBFN maintained generally high cell-line-level predictive performance, and perturbation-based attribution identified biologically relevant transcriptomic programs in selected drug-cell line settings. Together, these findings support GBFN as a compact and effective framework for continuous drug response prediction.

Humans

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics