Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Foundation model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

A formal theory for spatial representation and reasoning in biomedical ontologies.

OBJECTIVE: The objective of this paper is to demonstrate how a formal spatial theory can be used as an important tool for disambiguating the spatial information embodied in biomedical ontologies and for enhancing their automatic reasoning capabilities. METHOD AND MATERIALS: This paper presents a formal theory of parthood and location relations among individuals, called Basic Inclusion Theory (BIT). Since biomedical ontologies are comprised of assertions about classes of individuals (rather than assertions about individuals), we define parthood and location relations among classes in the extended theory Basic Inclusion Theory for Classes (BIT+Cl). We then demonstrate the usefulness of this formal theory for making the logical structure of spatial information more precise in two ontologies concerned with human anatomy: the Foundational Model of Anatomy (FMA) and GALEN. RESULTS: We find that in both the FMA and GALEN, class-level spatial relations with different logical properties are not always explicitly distinguished. As a result, the spatial information included in these biomedical ontologies is often ambiguous and the possibilities for implementing consistent automatic reasoning within or across ontologies are limited. CONCLUSION: Precise formal characterizations of all spatial relations assumed by a biomedical ontology are necessary to ensure that the information embodied in the ontology can be fully and coherently utilized in a computational environment. This paper can be seen as an important beginning step toward achieving this goal, but much more work along these lines is required.

Anatomy↗

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis↗

Amplification of Terminologia anatomica by French language terms using Latin terms matching algorithm: a prototype for other language.

OBJECTIVE: Terminologia anatomica is the new standard in anatomical terminology. This terminology is available only in Latin and English and its worldwide adoption is subject to the addition of terms from others languages. On the other hand, Nomina anatomica, the previous standard, has been widely translated. Aim of this work was to append foreign terms to Terminologia by using similarity-matching algorithm between its Latin terms and those from Nomina. METHODS: A semi-automatic matching of Latin terms from Terminologia with those of Nomina was performed using a string-to-string distance algorithm and manual assessment. We used a French-Latin version of Nomina together with Terminologia and we suggested French terms for Terminologia. Coverage was evaluated by the number of exact and approximate matches. A target of 78% was set due to the higher number of terms in Terminologia compared to Nomina. Relevance was estimated by manually comparing the meanings of the English and French terms related to the same Latin term. The question was whether they refer to the same anatomical structure. RESULTS: Exact or approximate matches were found for 5982 terms (76.5%) of Terminologia. Our results indicated that more than 75% of the terms from Terminologia came from Nomina, most of them were left unchanged and all were used with the same meaning. CONCLUSION: This method produces relevant results, reaching our 78% target. The method is based only on Latin terms and can be used for other languages. We consider this work as a starting point for adding terms to other knowledge sources, such as the foundational model of anatomy or the Unified Medical Language System (UMLS).

Algorithms↗

Desiderata for domain reference ontologies in biomedicine.

Domain reference ontologies represent knowledge about a particular part of the world in a way that is independent from specific objectives, through a theory of the domain. An example of reference ontology in biomedical informatics is the Foundational Model of Anatomy (FMA), an ontology of anatomy that covers the entire range of macroscopic, microscopic, and subcellular anatomy. The purpose of this paper is to explore how two domain reference ontologies--the FMA and the Chemical Entities of Biological Interest (ChEBI) ontology, can be used (i) to align existing terminologies, (ii) to infer new knowledge in ontologies of more complex entities, and (iii) to manage and help reasoning about individual data. We analyze those kinds of usages of these two domain reference ontologies and suggest desiderata for reference ontologies in biomedicine. While a number of groups and communities have investigated general requirements for ontology design and desiderata for controlled medical vocabularies, we are focusing on application purposes. We suggest five desirable characteristics for reference ontologies: good lexical coverage, good coverage in terms of relations, compatibility with standards, modularity, and ability to represent variation in reality.

Anatomy↗

A safety-centric perspective on innovation and risk in the use of artificial intelligence in genomics.

Adopting a safety-centric approach, this article explores how generative artificial intelligence (AI), and more specifically, foundation models for biological sequences, can exacerbate data quality issues, technical biases, and dual-use potential, particularly in critical applications such as clinical genetics, precision medicine, and pathogen engineering. This work centres on how misuse risks emerge throughout the innovation pipeline and how these intersect with the growing accessibility of generative genomic models. Particular attention is given to dual-use governance and infrastructure hardening in sequence analysis workflows. The work aims to provide scientists, regulators, and policymakers with a toolkit to discuss beneficial innovation in genomic AI while maintaining robust safeguards against harm and misuse.

Genomics↗

Jingjing Zhai and Edward S. Buckler.

Dr. Laura Zahn asked the authors, Dr. Jingjing Zhai and Dr. Edward (Ed) S. Buckler, to tell us about their research relating to their Cell Genomics paper, "PlantCAD2: A DNA foundation model for interpreting genomes across flowering plants."

Genomics↗

ELISA (Embedding-Linked Interactive Single-cell Agent): an interpretable hybrid generative Artificial Intelligence agent for expression-grounded discovery in single-cell genomics.

Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain opaque to natural language. Here, we introduce ELISA (Embedding-Linked Interactive Single-cell Agent), an interpretable framework that unifies single-cell generative pretrained transformer expression embeddings with biomedical bidirectional encoder representations from transformers-based semantic retrieval and large-language model (LLM)-mediated interpretation for interactive single-cell discovery. An automatic query classifier routes inputs to gene marker scoring, semantic matching, or reciprocal rank fusion pipelines depending on whether the query is a gene signature, natural language concept, or mixture of both. Integrated analytical modules perform pathway activity scoring across 60+ gene sets, ligand-receptor interaction prediction using 280+ curated pairs, condition-aware comparative analysis, and cell-type proportion estimation, all operating directly on embedded data without access to the original count matrix. Benchmarked across six diverse scRNA-seq datasets spanning inflammatory lung disease, pediatric and adult cancers, organoid models, healthy tissue, and neurodevelopment, ELISA significantly outperforms CellWhisperer, a classical lexical retriever (BM25), and a random baseline in cell type retrieval (combined permutation test, $p < 2\times 10^{-5}$ for each), with particularly large gains on gene-signature queries (Cohen's $d = 5.98$ for mean reciprocal rank). ELISA replicates published biological findings (mean composite score 0.88), and generates candidate hypotheses through grounded LLM reasoning, bridging the gap between transcriptomic data exploration and biological discovery.

Generative Artificial Intelligence↗

PharaCon: a new framework for identifying bacteriophages via conditional representation learning.

MOTIVATION: Identifying bacteriophages (phages) within metagenomic sequences is essential for understanding microbial community dynamics. Transformer-based foundation models have been successfully employed to address various biological challenges. However, these models are typically pre-trained with self-supervised tasks that do not consider label variance in the pre-training data. This presents a challenge for phage identification as pre-training on mixed bacterial and phage data may lead to information bias due to the imbalance between bacterial and phage samples. RESULTS: To overcome this limitation, we proposed a novel conditional BERT framework that incorporates label classes as special tokens during pre-training. Specifically, our conditional BERT model attaches labels directly during tokenization, introducing label constraints into the model's input. Additionally, we introduced a new fine-tuning scheme that enables the conditional BERT to be effectively utilized for classification tasks. This framework allows the BERT model to acquire label-specific contextual representations from mixed sequence data during pre-training and applies the conditional BERT as a classifier during fine-tuning, and we named the fine-tuned model as PharaCon. We evaluated PharaCon against several existing methods on both simulated sequence datasets and real metagenomic contig datasets. The results demonstrate PharaCon's effectiveness and efficiency in phage identification, highlighting the advantages of incorporating label information during both pre-training and fine-tuning. AVAILABILITY AND IMPLEMENTATION: The source code and associated data can be accessed at https://github.com/Celestial-Bai/PharaCon.

Bacteriophages↗

Investigating implicit knowledge in ontologies with application to the anatomical domain.

Knowledge in biomedical ontologies can be explicitly represented (often by means of semantic relations), but may also be implicit, i.e., embedded in the concept names and inferable from various combinations of semantic relations. This paper investigates implicit knowledge in two ontologies of anatomy: the Foundational Model of Anatomy and GALEN. The methods consist of extracting the knowledge explicitly represented, acquiring the implicit knowledge through augmentation and inference techniques, and identifying the origin of each semantic relation. The number of relations (12 million in FMA and 4.6 million in GALEN), broken down by source, is presented. Major findings include: each technique provides specific relations; and many relations can be generated by more than one technique. The application of these findings to ontology auditing, validation, and maintenance is discussed, as well as the application to ontology integration.

Artificial Intelligence↗

The United Kingdom National Healthy School Standard: a framework for strengthening the school nurse role.

The purpose of this review is to analyze the school nursing role within the National Healthy School Standard (NHSS) in the United Kingdom with a view toward clarifying and strengthening the role of school nurses globally. Within the National Healthy School Standard framework, school nurses serve an integral role in linking health and education partnerships to promote effective school health programs. School nurse contributions to the National Healthy School Standard, as well as barriers and supports, are discussed. Additionally, the methods school nurses implement to partner, to manage service delivery, and to work with schools are outlined. The central role of school nurses within the National Healthy School Standard framework provides a guide for school nurses in the United States to demonstrate their importance as key players in healthy schools that promote health and education. The framework deserves recognition as a foundational model to help strengthen both the school nurse role and school health programs around the world.

Cooperative Behavior↗

The Continuity Trap in Data Science Health Research.

Secondary use is now the ordinary condition of data science health research rather than an exception to it. Electronic health records collected for clinical care become prediction tools and inputs for generative AI; imaging archives become foundation-model corpora; genomic datasets become resources for polygenic risk scores; and legacy biospecimens become renewable, indefinitely distributable cell lines. Governance has responded by emphasizing verifiable instruments such as provenance logs, repository approvals, broad-consent forms, data-use agreements, model cards, records of processing, and locality-preserving architectures. These instruments are necessary, and they answer real questions about lineage, privacy, institutional responsibility, and accountability, but they are not sufficient to establish that a present use remains ethically justified. We define ethical continuity as the persistence of normatively relevant relationships between the original conditions of data generation or material collection and subsequent downstream uses, such that current uses remain justifiable in light of the expectations, permissions, meanings, and relational obligations present at entrustment. We then define the Continuity Trap as a review-stage governance error in which a salient signal of continuity in one domain is treated as sufficient evidence of ethical continuity overall, causing inquiry into the remaining domains to close prematurely. The trap is not ordinary noncompliance, ethics creep, or a demand for universal rereview; it is a cross-domain inference error that can arise even in careful, good-faith review. We distinguish it from proxy closure, of which it is a continuity-specific subtype, and from Goodhart's and Campbell's laws, which describe how measures degrade once they become targets. We operationalize ethical continuity across 4 domains: provenance, semantics, authorization, and relational standing, developed in our Representational Veracity framework, and we show that these domains can diverge as data are linked, transformed, modeled, and redeployed. We identify the institutional mechanisms-provenance privilege, descriptor sedimentation, authorization fossilization, and community effacement-that cause auditable signals to be overread, and we examine how the US Health Insurance Portability and Accountability Act (HIPAA) of 1996, the General Data Protection Regulation, the European Health Data Space, US Food and Drug Administration guidance, the US National Institute of Standards and Technology (NIST) AI Risk Management Framework, and federated-learning governance can reduce risk while still inducing continuity traps. We apply the framework to consent and nonconsent settings, including public health, immunization, syndromic, and wastewater surveillance, polygenic risk scores, induced pluripotent stem cells, federated learning, and health-related large language models. The policy implication is trigger-based continuity review: rather than rereviewing every reuse, investigators and reviewers should identify the weakest continuity domain at the present data stage and impose a domain-matched safeguard, recorded in a short continuity statement. This reframing is intended for the committees, repositories, funders, and governance bodies that decide whether reuse may proceed, and it matters most in cross-border and low-resource settings. Provenance should begin ethical review; it should not end it.

Data Science↗

Insights into Tardigrade Damage-Suppression Protein, Dsup.

Tardigrades are microscopic invertebrates capable of surviving extreme environmental conditions through unique molecular adaptations. Among the proteins implicated in their remarkable resilience is a novel protein known as damage suppressor (Dsup), a key factor in protecting cellular DNA from elevated levels of radiation. Since its discovery, numerous studies have explored the biochemical, structural, and functional properties of Dsup. In this review, we summarize the current knowledge surrounding these properties and describe several proposed mechanisms by which Dsup may confer protection. For each proposed mechanism, we outline the foundational model, present supporting evidence, and highlight critical gaps in our understanding. Taken together, we believe that Dsup likely employs multiple complementary mechanisms to protect DNA. Finally, we discuss emerging applications of Dsup and Dsup-inspired technologies for human health. Overall, this review synthesizes our current understanding and provides a framework to guide future investigations into this remarkable protein.

Animals↗

Representation of temporal indeterminacy in clinical databases.

Temporal indeterminancy is common in clinical medicine because the time of many clinical events is frequently not precisely known. Decision support systems that reason with clinical data may need to deal with this indeterminancy. This indeterminacy support must have a sound foundational model so that other system components may take advantage of it. In particular, it should operate in concert with temporal abstraction, a feature that is crucial in several clinical decision support systems that our group has developed. We have implemented a temporal query system called Tzolkin that provides extensive support for the temporal indeterminancies found in clinical medicine, and have integrated this support with our temporal abstraction mechanism. The resulting system provides a simple, yet powerful approach for dealing with temporal indeterminancy and temporal abstraction.

Databases as Topic↗

A formal method to resolve temporal mismatches in clinical databases.

Overcoming data heterogeneity is essential to the transfer of decision-support programs to legacy databases and to the integration of data in clinical repositories. Prior methods have focused primarily on problems of differences in terminology and patient identifiers, and have not addressed formally the problem of temporal data heterogeneity, even though time is a necessary element in storing, manipulating, and reasoning about clinical data. In this paper, we present a method to resolve temporal mismatches present in clinical databases. This method is based on a foundational model of time that can formalize various temporal representations. We use this temporal model to define a novel set of twelve operators that can map heterogeneous time-stamped data into a uniform temporal scheme. We present an algorithm that uses these mapping operators, and we discuss our implementation and evaluation of the method as a software program called Synchronus.

Algorithms↗

The role of definitions in biomedical concept representation.

The Foundational Model (FM) of anatomy, developed as an anatomical enhancement of UMLS, classifies anatomical entities in a structural context. Explicit definitions have played a critical role in the establishment of FM classes. Essential structural properties that distinguish a group of anatomical entities serve as the differentiate for defining classes. These, as well as other structural attributes, are introduced as template slots in Protégé, a frame-based knowledge acquisition system, and are inherited by descendants of the class. A set of desiderata has evolved during the instantiation of the FM for formulating definitions. We contend that 1. these desiderata generalize to non-anatomical domains and 2. satisfying them in constituent vocabularies of UMLS would enhance the quality of information retrievable through UMLS.

Anatomy↗

Scale and context: issues in ontologies to link health- and bio-informatics.

Bridging levels of scale and context are key problems for integrating Bio- and Health Informatics. Formal, logic-based ontologies using expressive formalisms are naturally "fractal" and provide new methods to support these aims. The basic notion of composition can be used to bridge scales; axioms can be used to carry implicit information; specific context markers can be included in definitions; and a hierarchy of semantic links can be used to represent subtle differences in point of view. Experience with OpenGALEN, the UK Drug Ontology and new experiments with the Gene Ontology and Foundational Model of Anatomy suggest that these are powerful tools provide practical solutions.

Medical Informatics↗

Dynamic control of B lymphocyte development in the bursa of fabricius.

The chicken is a foundational model for immunology research and continues to be a valuable animal for insights into immune function. In particular, the bursa of Fabricius can provide a useful experimental model of the development of B lymphocytes. Furthermore, an understanding of avian immunity has direct practical application since chickens are a vital food source. Recent work has revealed some of the molecular interactions necessary to allow proper repertoire diversification in the bursa while enforcing quality control of the lymphocytes produced, ensuring that functional cells without self-reactive immunoglobulin receptors populate the peripheral immune organs. Our laboratory has focused on the function of chB6, a novel molecule capable of inducing rapid apoptosis in bursal B cells. Our recent work on chB6 will be presented and placed in the context of other recent studies of B cell development in the bursa.

Animals↗

Aligning representations of anatomy using lexical and structural methods.

OBJECTIVE: The objective of this experiment is to develop methods for aligning two representations of anatomy (the Foundational Model of Anatomy and GALEN) at the lexical and structural level. METHODS: The alignment consists of the following four steps: 1)acquiring terms, 2) identifying anchors (i.e., shared concepts) lexically, 3) acquiring explicit and implicit semantic relations, and 4) identifying anchors structurally. RESULTS: 2,353 anchors were identified by lexical methods, of which 91% were supported by structural evidence. No evidence was found for 7.5%of the anchors and 1.5% received negative evidence. DISCUSSION: The importance of taking advantage of implicit domain knowledge acquired through complementation,augmentation, and inference is discussed.

Anatomy↗