Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “foundation model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Real-world deployment of a fine-tuned pathology foundation model for lung cancer biomarker detection.

Artificial intelligence models using digital histopathology slides stained with hematoxylin and eosin offer promising, tissue-preserving diagnostic tools for patients with cancer. Despite their advantages, their clinical utility in real-world settings remains unproven. Assessing EGFR mutations in lung adenocarcinoma demands rapid, accurate and cost-effective tests that preserve tissue for genomic sequencing. PCR-based assays provide rapid results but with reduced accuracy compared with next-generation sequencing and require additional tissue. Computational biomarkers leveraging modern foundation models can address these limitations. Here we assembled a large international clinical dataset of digital lung adenocarcinoma slides (N = 8,461) to develop a computational EGFR biomarker. Our model fine-tunes an open-source foundation model, improving task-specific performance with out-of-center generalization and clinical-grade accuracy on primary and metastatic specimens (mean area under the curve: internal 0.847, external 0.870). To evaluate real-world clinical translation, we conducted a prospective silent trial of the biomarker on primary samples, achieving an area under the curve of 0.890. The artificial-intelligence-assisted workflow reduced the number of rapid molecular tests needed by up to 43% while maintaining the current clinical standard performance. Our retrospective and prospective analyses demonstrate the real-world clinical utility of a computational pathology biomarker.

Humans↗

NextVir: Enabling classification of tumor-causing viruses with genomic foundation models.

MOTIVATION: Oncoviruses, pathogens known to cause or increase the risk of cancer, include both common viruses such as human papillomaviruses and rarer pathogens such as human T-lymphotropic viruses. Computational methods for detecting viral DNA from data acquired by modern DNA sequencing technologies have enabled studies of the association between oncoviruses and cancers. Those studies are rendered particularly challenging when multiple species of oncovirus are present in a tumor sample. In such scenarios, merely detecting the presence of a sequencing read of viral origin is insufficiently informative-instead, a more precise characterization of the viral content in the sample is required. RESULTS: We address this need with NextVir, to our knowledge the first multi-class viral classification framework that adapts genomic foundation models to detecting and classifying sequencing reads of oncoviral origin. Specifically, NextVir explores several foundation models-DNABERT-S, Nucelotide Transformer, and HyenaDNA-and efficiently fine-tunes them to enable accurate identification of the sequencing reads' origin. The results demonstrate superior performance of the proposed framework over existing deep learning methods and suggest downstream potential for foundational models in genomics.

Humans↗

Bridging ancestry gaps in genomic risk prediction with tabular foundation models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Humans↗

A foundational model of time for heterogeneous clinical databases.

Differences among the database representations of clinical data are a major barrier to the integration of databases and to the sharing of decision-support applications across databases. Prior research on resolving data heterogeneity has not addressed specifically the types of mismatches found in various timestamping approaches for clinical data. Such temporal mismatches, which include time-unit differences among timestamps, must be overcome before many applications can use these data to reason about diagnosis, therapy, or prognosis. In this paper, we present an analysis of the types of temporal mismatches that exist in databases. To formalize these various approaches to timestamping, we provide a foundational model of time. This model gives us the semantics necessary to encode the temporal dimensions of clinical data in legacy databases and to transform such heterogeneous data into a uniform temporal representation suitable for decision support. We have implemented this foundational model as an extension to our Chronus system, which provides clinical decision-support applications the ability to match temporal patterns in clinical databases. We discuss the uniqueness of our approach in comparison with other research on representing and querying clinical data with varying timestamp representations.

Databases as Topic↗

OQAFMA Querying agent for the Foundational Model of Anatomy: a prototype for providing flexible and efficient access to large semantic networks.

The development of large semantic networks, such as the UMLS, which are intended to support a variety of applications, requires a flexible and efficient query interface for the extraction of information. Using one of the source vocabularies of UMLS as a test bed, we have developed such a prototype query interface. We first identify common classes of queries needed by applications that access these semantic networks. Next, we survey StruQL, an existing query language that we adopted, which supports all of these classes of queries. We then describe the OQAFMA Querying Agent for the Foundational Model of Anatomy (OQAFMA), which provides an efficient implementation of a subset of StruQL by pre-computing a variety of indices. We describe how OQAFMA leverages database optimization by converting StruQL queries to SQL. We evaluate the flexibility and efficiency of our implementation using English queries written by anatomists. This evaluation verifies that OQAFMA provides flexible, efficient access to one such large semantic network, the Foundational Model of Anatomy, and suggests that OQAFMA could be an efficient query interface to other large biomedical knowledge bases, such as the Unified Medical Language System.

Abstracting and Indexing↗

Foundation model enables interpretable open and error-tolerant searching for mass spectrometry-based proteomics.

MOTIVATION: Mass spectrometry-based proteomics allows studying all proteins of a sample on a molecular level. However, mass spectra are noisy and contain complex patterns, making them inherently challenging to analyze with algorithmic approaches. In terms of the protein sequence landscape, most recent bottom-up MS-based proteomics studies consider either a diverse pool of post-translational modifications, employ large databases-as in metaproteomics or proteogenomics, study multiple isoforms of proteins, include unspecific cleavage sites or even combinations thereof. All this makes peptide and protein identifications challenging. RESULTS: Here, we present a foundation model, called yHydra, that jointly embeds spectra and peptides. This allows us to implement various downstream tasks and search modes in Euclidean space. We implement an open search which allows querying multiple ten-thousands of spectra against millions of peptides. Furthermore, we implement an error-tolerant search for identifying additional proteoforms that are not included in off-the-shelf reference proteomes. Our foundation model provides meaningful embeddings, as we interpret learned peptide embeddings in comparison to the peptide's physico-chemical properties. Hydra's open search, assigns delta masses to each identification which allows to unrestrictedly characterize post-translational modifications. The error-tolerant mode of yHydra can be used as post-processing to existing search engines or as a standalone. yHydra is evaluated on several real life data sets for the identification of modified peptide sequences and shows up to 25% increase in peptide identification at constant false discovery rate compared to the current state-of-the-art. AVAILABILITY AND IMPLEMENTATION: Code is available on Gitlab: https://gitlab.com/dacs-hpi/yHydra, and https://gitlab.com/dacs-hpi/yHydra_train.

Proteomics↗

Challenges in converting frame-based ontology into OWL: the Foundational Model of Anatomy case-study.

A description logics representation of the Foundational Model of Anatomy (FMA) in the Web Ontology Language (OWL-DL) would allow developers to combine it with other OWL ontologies, and would provide the benefit of being able to access generic reasoning tools. However, the FMA is currently represented in a frame language. The differences between description logics and frames are not only syntactic, but also semantic. We analyze some theoretical and computational limitations of converting the FMA into OWL-DL. Namely, some of the constructs used in the FMA do not have a direct equivalent in description logics, and a complete conversion of the FMA in description logics is too large to support reasoning. Therefore, an OWL-DL representation of the FMA would have to be optimized for each application. We propose a solution based on OWL-Full, a superlanguage of OWL-DL, that meets the expressiveness requirements and remains application-independent. Specific simplified OWL-DL representations can then be generated from the OWL-Full model by applications. We argue that this solution is easier to implement and closer to the application needs than an integral translation, and that the latter approach would only make the FMA maintenance more difficult.

Anatomy↗

A prototype natural language interface to a large complex knowledge base, the Foundational Model of Anatomy.

We describe a constrained natural language interface to a large knowledge base, the Foundational Model of Anatomy (FMA). The interface, called GAPP, handles simple or nested questions that can be parsed to the form, subject-relation-object, where subject or object is unknown. With the aid of domain-specific dictionaries the parsed sentence is converted to queries in the StruQL graph-searching query language, then sent to a server we developed, called OQAFMA, that queries the FMA and returns output as XML. Preliminary evaluation shows that GAPP has the potential to be used in the evaluation of the FMA by domain experts in anatomy.

Anatomy↗

A Foundation Model Based CT Biomarker for Non-Invasive Prediction of Response to Neoadjuvant Immunochemotherapy in Non-Small Cell Lung Cancer.

Predicting pathological complete response (pCR) to neoadjuvant immunochemotherapy in non-small cell lung cancer (NSCLC) is clinically important yet remains challenging. Here, we introduce a foundation model-derived computed tomography (CT) imaging biomarker established from a multi-center cohort of 702 patients. Specifically, we developed and validated a non-invasive baseline CT-based model for risk stratification of pathological response. To address scanner and protocol heterogeneity, we first built a 3D Vision Mamba-based CT super-resolution model trained on 2494 cases for image standardization. We then fine-tuned a lung cancer-specific CT foundation model from a pretrained 3D model (VoCo) using 6643 chest CT scans. Finally, we constructed a multi-task Swin Transformer that jointly performs risk stratification and segments tumors to generate the imaging biomarker. Across five centers, the model achieved consistently strong generalization (AUC: 0.75-0.87) for pCR prediction. Genomic analysis revealed that the biomarker was independent of tumor mutational burden but significantly associated with TP53 mutations, suggesting an association with a radiogenomic phenotype related to this alteration. Together, these results demonstrate a generalizable and biologically meaningful foundation model-based biomarker for non-invasive risk stratification of pathological response in NSCLC.

Female↗

Foundational model of neuroanatomy: implications for the Human Brain Project.

In order to meet the need for a controlled terminology in neuroinformatics, we have integrated the extensive terminology of NeuroNames into the Foundational Model of anatomy. We illustrate the application of foundational principles for the establishment of an inheritance hierarchy, which accommodates anatomical attributes of neuroanatomical concepts and provides the foundation to which other information may be linked.

Brain↗

scPlantLLM: A Foundation Model for Exploring Single-cell Expression Atlases in Plants.

Single-cell RNA sequencing (scRNA-seq) provides unprecedented insights into plant cellular diversity by enabling high-resolution analyses of gene expression at the single-cell level. However, the complexity of scRNA-seq data, including challenges in batch integration, cell type annotation, and gene regulatory network (GRN) inference, demands advanced computational approaches. To address these challenges, we developed scPlantLLM, a Transformer model trained on millions of plant single-cell data points. Using a sequential pretraining strategy incorporating masked language modeling and cell type annotation tasks, scPlantLLM generates robust and interpretable single-cell data embeddings. When applied to Arabidopsis thaliana datasets, scPlantLLM excels in clustering, cell type annotation, and batch integration, achieving an accuracy of up to 0.91 in zero-shot learning scenarios. Furthermore, the model demonstrates an ability to identify biologically meaningful GRNs and subtle cellular subtypes, showcasing its potential to advance plant biology research. Compared to traditional methods, scPlantLLM outperforms in key metrics such as adjusted rand index (ARI), normalized mutual information (NMI), and silhouette score (SIL), highlighting its superior clustering accuracy and biological relevance. scPlantLLM represents a foundation model for exploring plant single-cell expression atlases, offering unprecedented capabilities to resolve cellular heterogeneity and regulatory dynamics across diverse plant systems. The code used in this study is available at https://github.com/compbioNJU/scPlantLLM.

Single-Cell Analysis↗

Foundation model reveals the shared organization of transcription and topologically associating domains.

The three-dimensional organization of chromatin into topologically associating domains (TADs) may impact gene regulation by bringing distant genes into contact. However, studies of TADs' function and their influence on transcription have been constrained by ambiguities in TAD boundary definitions and challenges in directly measuring their regulatory effects. We overcome these limitations by developing species-level consensus TAD maps for human and mouse by using a bag-of-genes approach that exposes an emergent regulatory structure. To quantify TAD-mediated relationships, we use a foundation model trained on 33 million transcriptomes to define a contextual similarity metric that captures higher-order relationships missed by co-expression. We find that TADs are regions of elevated co-regulation, with our framework yielding testable hypotheses about chromatin organization across cellular states. This TAD-linked enhancement is strongest during early development and declines with aging, while cancer cells show distinct TAD usage that shifts with chemotherapy. Together, these findings suggest that chromatin organization acts through probabilistic rather than deterministic mechanisms.

Humans↗

Experimental evaluation of an elastic foundation model to predict contact pressures in knee replacements.

Computational wear prediction is an attractive concept for evaluating new total knee replacement designs prior to physical testing and implementation. An important hurdle to such technology is the lack of in vivo contact pressure predictions. To address this issue, this study evaluates a computationally efficient simulation approach that combines the advantages of rigid and deformable body modeling. The hybrid method uses rigid body dynamics to predict body positions and orientations and elastic foundation theory to predict contact pressures between general three-dimensional surfaces. To evaluate the method, we performed static pressure experiments with a commercial knee implant in neutral alignment using flexion angles of 0, 30, 60, and 90 degrees and loads of 750, 1500, 2250, and 3000N. Using manufacturer CAD geometry for the same implant, an elastic foundation model with linear or nonlinear polyethylene material properties was implemented within a commercial multibody dynamics software program. The model's ability to predict experimental peak and average contact pressures simultaneously was evaluated by performing dynamic simulations to find the static configuration. Both the linear and nonlinear material models predicted the average contact pressure data well, while only the linear material model could simultaneously predict the trends in the peak contact pressure data. This novel modeling approach is sufficiently fast and accurate to be used in design sensitivity and optimization studies of knee implant mechanics and ultimately wear.

Compressive Strength↗

The evolving neuroanatomical component of the Foundational Model of Anatomy.

In order to meet the need for an expressive ontology in neuroinformatics, we have integrated the extensive terminologies of NeuroNames and Terminologia Anatomica into the Foundational Model of Anatomy (FMA). We have enhanced the FMA to accommodate information unique to neuronal structures, such as axonal input/output relationships.

Anatomy↗

An approach to the anatomical correlation of species through the Foundational Model of Anatomy.

The increasing need for extrapolating information from one species to another has been highlighted by contemporary research in bioinformatics, genomics, proteomics, and animal models of human disease, as well as other fields. We propose an approach to correlating the anatomy of Homo sapiens with selected species, using the Foundational Model of Anatomy (FMA) as a framework, and graph matching as a method, for determining similarities and differences in the nodes and relationships (edges) defined by the attributed graph of the FMA. We illustrate our approach by comparing anatomical structures of mouse and human that present prototypical mapping problems.

Anatomy↗

Experience in reasoning with the foundational model of anatomy in OWL DL.

The objective of this study is to compare description logics (DLs) and frames for representing large-scale biomedical ontologies and reasoning with them. The ontology under investigation is the Foundational Model of Anatomy (FMA). We converted it from its frame-based representation in Protégé into OWL DL. The OWL reasoner Racer helped identify unsatisfiable classes in the FMA. Support for consistency checking is clearly an advantage of using DLs rather than frames. The interest of reclassification was limited, due to the difficulty of defining necessary and sufficient conditions for anatomical entities. The sheer size and complexity of the FMA was also an issue.

Computational Biology↗

Proposed classification of cells in the Foundational Model of Anatomy.

A logical and principled representation of cell types and their component parts could serve as a framework for correlating the various ontologies that are emerging in bioinformatics with a focus on cells and subcellular biological entities. In order to address this need we have extended the Foundational Model of Anatomy (FMA)1,2 from macroscopic to cellular and subcellular anatomical entities. The poster will provide a live demonstration of this implementation.

Anatomy↗

The potential of the digital anatomist foundational model for assuring consistency in UMLS sources.

Inconsistent anatomical concept representation can be identified in anatomy textbooks and hard copy term lists, as well as in UMLS source vocabularies and other controlled medical terminologies. In this report we select some examples of inconsistent representations of anatomical concepts, and illustrate how these inconsistencies can be explained and reconciled by the Digital Anatomist Foundational Model. We use this process for gaining a measure of the validity of the logic-based Model.

Anatomy↗