Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Foundation model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

An operations model of psychosocial structure and function and of psychotherapy.

Currently, personality theory and clinical psychology have a fairly substantial tradition of promoting a strongly scientific basis for clinical work and theorizing. However, an appropriate foundation model has been difficult to identify and establish. A theory of human operations, here proposed, may provide such an elementary model. The theory is rooted in the organizational and industrial field known as operations, which is a highly systematic, precise, flexible, scientific approach to the understanding and management of human goal-seeking action in the broadest sense. The proposed model includes the classical humanistic, clinical, and decision theoretic notions of values, cognition, emotions, ego, behavior, objectives, outcomes, feedback, and defenses. These notions are placed within an overall operations frame of reference and developed in such a manner that they can be used to assess human clinical problems and to design therapeutic interventions. The strengths and limitations of the model are discussed.

Cognition↗

Flashzoi: an enhanced Borzoi for accelerated genomic analysis.

MOTIVATION: Accurately predicting how DNA sequence drives gene regulation and how genetic variants alter gene expression is a central challenge in genomics. Borzoi, which models over ten thousand genomic assays including RNA-seq coverage from over half a megabase of sequence context alone promises to become an important foundation model in regulatory genomics, both for massively annotating variants and for further model development. However, the currently used relative positional encodings limit Borzoi's computational efficiency. RESULTS: We present Flashzoi, an enhanced Borzoi model that leverages rotary positional encodings and FlashAttention-2. This achieves over 3-fold faster training and inference and up to 2.4-fold reduced memory usage, while maintaining or improving accuracy in modeling various genomic assays including RNA-seq coverage, predicting variant effects, and enhancer-promoter linking. Flashzoi's improved efficiency facilitates large-scale genomic analyses and opens avenues for exploring more complex regulatory mechanisms and modeling. AVAILABILITY AND IMPLEMENTATION: The Flashzoi model architecture is part of the MIT-licensed borzoi-pytorch package, can be found at https://github.com/johahi/borzoi-pytorch and installed via pip. Model weights for all four Flashzoi and Borzoi replicates are available at https://huggingface.co/johahi under the MIT license. The code has been archived at https://zenodo.org/records/15669913.

Genomics↗

Agentomics: an agentic system that autonomously develops novel state-of-the-art solutions for biomedical machine learning tasks.

MOTIVATION: Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack flexibility, while large language models (LLMs) struggle to consistently deliver reproducible machine learning codebases, and existing LLM Agent-powered solutions lag behind human-engineered ML models. RESULTS: Here, we introduce Agentomics, an autonomous LLM-powered agentic system for end-to-end ML experimentation. Given a biomedical dataset, Agentomics implements various ML modeling strategies, and produces a ready-to-use ML model. Agentomics introduces strict validation checkpoints for standard ML development steps, allowing gradual development on top of working code with defined interfaces and validated artifacts. Further, it offers native support for biomedical foundation models that can be leveraged during experimentation. The generic nature of Agentomics allows the user to create ML solutions for a large variety of datasets and use various LLMs. We evaluate Agentomics across 20 datasets from the domains of Protein Engineering, Drug Discovery, and Regulatory Genomics. When benchmarked against other agentic systems, Agentomics outperformed them in all tested domains. When benchmarked against human expert solutions, Agentomics generated novel state-of-the-art models for 11/20 established benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Agentomics is implemented in Python. Source code and documentation are freely available at: https://github.com/BioGeMT/Agentomics-ML.

Machine Learning↗

scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution.

Understanding how regulatory DNA elements shape gene expression across individual cells is a fundamental challenge in genomics. Joint RNA-seq and epigenomic profiling provides opportunities to build unifying models of gene regulation capturing sequence determinants across steps of gene expression. However, current models, developed primarily for bulk omics data, fail to capture the cellular heterogeneity and dynamic processes revealed by single-cell multi-modal technologies. Here, we introduce scooby, the first framework to model scRNA-seq coverage and scATAC-seq insertion profiles along the genome from sequence at single-cell resolution. For this, we leverage the pre-trained multi-omics profile predictor Borzoi as a foundation model, equip it with a cell-specific decoder, and fine-tune its sequence embeddings. Specifically, we condition the decoder on the cell position in a precomputed single-cell embedding resulting in strong generalization capability. Applied to a hematopoiesis dataset, scooby recapitulates cell-specific expression levels of held-out genes, and identifies regulators and their putative target genes through in silico motif deletion. Moreover, accurate variant effect prediction with scooby allows for breaking down bulk eQTL effects into single-cell effects and delineating their impact on chromatin accessibility and gene expression. We anticipate scooby to aid unraveling the complexities of gene regulation at the resolution of individual cells.

Journal Article↗

Biosensor immunosurface engineering inspired by B-cell membrane-bound antibodies: modeling and analysis of multivalent antigen capture by immobilized antibodies.

Immobilized antibodies are used by many biosensors and diagnostic tests as specific receptors for the presence of targeted substances in clinical, biological, or environmental samples. The antibodies used in these devices are the soluble form of the antibodies presented on the B-cell membrane: they have the same specificity, but they may differ from those presented on the B cell in orientation, flexibility, mobility, and support-membrane properties. These properties influence the formation of noncovalent bonds between the pathogen antigenic determinants (epitopes) and the amino acids of the antibodies. This paper extends the theoretical modeling foundation addressing multivalent antigen binding to cell surface receptors to account for local and far-field antibody surface density effects, immobilized antibodies, and the flexibility and range of motion of immobilized antibodies. An analysis of the derived model provides insight into the design of biosensor immunosurfaces to enhance pathogen capture capability.

Antigen-Antibody Complex↗

Curriculum focus: traditional dental education confronts the new biology and social responsibility.

The traditional dental education base does not suggest that graduates are prepared only for private general practice. It promises much, much more. It is the beginning of what any graduate may wish or want it to be. In fact it is necessary for new graduates contemplating private practice to seek more in the way of practice management if they plan to be successful in that model. Such is also the case for additional study if one wants to pursue a specialty; further training in the scientific method to undertake research; or further study in education for a teaching career. The predoctoral program promises to be just a beginning, but a sound and sensible beginning. The current foundation model of the predoctoral curriculum continues to serve society and the profession very well. It provides the necessary grounding in the fundamentals of basic medical sciences, clinical biological sciences, and behavioral sciences upon which any graduate can build and pursue a career beyond what is thought to be the traditional role model of the private practice of general dentistry. Faculty members and dental programs can choose, if they desire, to tailor mission statements to reflect a changing emphasis. This could help dental applicants become more thoughtful and self-selective in considering one of many different career options available. Given the emerging demographic profile and trends in needs, the curriculum emphasis could be altered to encourage a change in the balance or proportion of graduates who undertake career options in a postdoctoral experience, in public health, research, teaching, administration, the military, hospital, or specialty practice.(ABSTRACT TRUNCATED AT 250 WORDS)

Curriculum↗

Plant cis-regulatory grammar: Decoding the multidimensional code of transcriptional regulation for programmable crop engineering.

Cis-regulatory elements (CREs) orchestrate the spatiotemporal precision of gene expression that underlies plant development, adaptation, and domestication. Decoding the cis-regulatory grammar of plant genomes remains a central challenge in modern biology, with profound implications for programmable crop engineering. Here, recent conceptual and technological advances are synthesized to reshape our understanding of plant CREs. This review first argues that CRE function is not only an intrinsic property of DNA sequence alone but also emerges from a multidimensional context, including chromatin accessibility, histone modifications, three-dimensional genome topology, and cell type-specific regulatory landscapes. Furthermore, the convergence of single-cell epigenomics, high-throughput functional assays, and CRISPR-based dissection has begun to unravel this contextual grammar, revealing the computational principles governing transcriptional regulation. Critically, we propose that artificial intelligence (AI) platforms are catalyzing an ongoing transition from descriptive discovery to predictive engineering, wherein these platforms outperform natural evolution in designing synthetic CREs. Finally, a roadmap is outlined toward a plant regulatory grammar foundation model, which will enable truly predictive engineering of gene expression when fine-tuned for specific tasks. Collectively, the integration of single-cell resolution maps, precise genome editing, AI-driven design, and regulatory-compliant delivery systems promises to transform our ability to reprogram plant gene regulation for next-generation agriculture, bridging the gap between foundational regulatory biology and tangible crop improvement.

artificial intelligence↗

Stroke treatment economic model (STEM): predicting long-term costs from functional status.

BACKGROUND AND PURPOSE: Stroke is a debilitating disease with long-term social and economic consequences. As new therapies for acute ischemic stroke are forthcoming, there is an increasing need to understand their long-term economic implications. To address this need, a stroke economic model was created. METHODS: The model consists of 3 modules. A short-term module incorporates short-term clinical trial data. A long-term module composed of several Markov submodels predicts patient transitions among various locations over time. The modules are connected via a bridge component that groups the survivors at the end of the short-term module according to their functional status and location. Examples of analyses that can be conducted with this model are provided with the use of data from 2 international trials. For illustration, UK unit costs were estimated. RESULTS: With the trial data in the short-term module, the short-term management cost is estimated to be pound8326 (US $13,649 [USD]). Hospital stay was the major cost driver. By the end of the trials, there was a pronounced difference in the distribution of patient locations between functional groups. It is predicted in the long-term module that the subsequent cost amounts to pound75 985 (124,564 USD) for a major and pound27,995 (45,893 USD) for a minor stroke. CONCLUSIONS: Linking functional recovery at the end of short-term treatment with patients' treatment and residential locations allows this model to estimate the long-term economic impact of stroke interventions. Using patient location instead of the more common natural history as the model foundation allows quantification of the long-term impact to become data driven and hence increases confidence in the results.

Cost of Illness↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings.

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Sequence Analysis, DNA↗

DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e., DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 23 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. Model, codes, and data is publicly available at https://github.com/MAGlCS-LAB/DNABERT_S.

Journal Article↗

Toward real-time quantification of driving risks: a systematic review and research agenda of risk field theory.

In complex traffic systems, driving risk often evolves in a continuous and progressive manner prior to crash occurrence. How to effectively represent and analyze such latent risk states remains a central challenge in traffic safety research. In recent years, risk field-based approaches have introduced spatial and spatiotemporal continuous modeling paradigms, providing new perspectives for characterizing the distribution of traffic risk and its dynamic evolution. Motivated by the rapid growth of this research area and the lack of a systematic synthesis, this paper presents a comprehensive review of studies applying risk field theory to driving safety and traffic risk analysis. Following the PRISMA guidelines, relevant literature was collected through multi-database searches and analyzed using a combination of bibliometric analysis and qualitative review. The review systematically summarizes the theoretical foundations, modeling elements, data sources, analytical methods, and application domains of risk field-related research. Particular attention is given to studies that conceptualize traffic risk as a continuous field, complemented by a broader review of traffic risk factor literature to identify key elements and analytical dimensions involved in risk field modeling. On this basis, the paper synthesizes research progress in major application areas, including traffic safety state representation, driving behavior analysis, traffic conflict assessment, and autonomous driving and human-machine cooperative systems. Differences and commonalities among existing studies are compared in terms of modeling strategies, data support, and application scenarios. Through this systematic review, the paper clarifies the main research themes and methodological trends of risk field-based studies, providing a structured framework for understanding the evolution and application of this approach and offering methodological insights for risk perception modeling and safety-oriented decision support in intelligent transportation systems (ITS).

Humans↗

Understanding the sources of performance in deep drug response models reveals insights and improvements.

MOTIVATION: Anti-cancer drug response prediction (DRP) using cancer cell lines (CLs) is crucial in stratified medicine and drug discovery. Recently, new deep learning models for DRP have improved performance over their predecessors. However, different models use different input data types and architectures making it hard to find the source of these improvements. Here we consider published DRP models that report state-of-the-art performance predicting continuous response values. These models take chemical structures of drugs and omics profiles of CLs as input. RESULTS: By experimenting with these models and comparing with our simple baselines, we show that no performance comes from drug features, instead, performance is due to the transcriptomics CL profiles. Furthermore, we show that, depending on the testing type, much of the current reported performance is a property of the training target values. We address these limitations by creating BinaryET and BinaryCB that predict binary drug response values, guided by the hypothesis that this reduces the noise in the drug efficacy data. Thus, better aligning them with biochemistry that can be learnt from the input data. BinaryCB leverages a chemical foundation model, while BinaryET is trained from scratch using a transformer-type architecture. We show that these models learn useful chemical drug features, which is the first time this has been demonstrated for multiple testing types to our knowledge. We further show binarizing the drug response values causes the models to learn useful chemical drug features. We also show that BinaryET improves performance over BinaryCB, and the published models that report state-of-the-art performance. AVAILABILITY AND IMPLEMENTATION: Code is available from https://github.com/Nik-BB/Understanding_DRP_models.

Humans↗

Levels of computer education for professional nursing. Development of a prototype graduate course.

Leveling computer education for professional nurses at the undergraduate and graduate levels is presented via a four-tiered model. Foundational course content at level 2 for all graduate nursing students preparing for advanced nursing roles is described. A quasi-experimental study (N = 80) demonstrated the effectiveness of the course in terms of the attitude changes toward computerization and perceived knowledge about computer applications in nursing.

Attitude of Health Personnel↗

The digital anatomist structural abstraction: a scheme for the spatial description of anatomical entities.

In this paper, we propose a generalized scheme for the symbolic description of the spatial attributes of anatomical entities. The power of the scheme lies in the ability to model the spatial objects at the highest level of granularity: information can be obtained at the desired level of detail needed for a given application. This scheme uses the topological classes of point, line, surface, and volume to represent zero-D, one-D, two-D and three-D objects. A spatial object participates as a node in three complementary networks; the topology network, the part-of network, and the spatial associations network. The topology network describes a spatial object in terms of its boundaries, the part-of network describes a spatial object in terms of its parts, and the spatial associations network describes the spatial object in terms of its relationships to other spatial objects. All three of the networks can be used in combination or alone to answer queries to the spatial information system. The Digital Anatomist Structural Abstraction together with the other components of the Digital Anatomist Foundational Model will provide the information for describing and reasoning about anatomical entities.

Anatomy↗

BioMedGraphica: An All-in-One Platform for Joint Textual Biomedical Prior Knowledge and Numeric Graph Generation.

Multi-omic data analysis is essential for scientific discovery in precision medicine. However, translating statistical results of omic data analysis into novel scientific hypothesis remains a significant challenge. Human experts must manually review analysis results and generate new hypothesis based on extensive and inter-connected biomedical prior knowledge, which is subjective and not scalable. While large language models (LLMs) can accelerate the discovery, their reasoning improves when grounded in structured, auditable and comprehensive biomedical prior knowledge. Biomedical knowledge, however, is scattered across heterogeneous databases that use diverse and inconsistent nomenclature systems, making it difficult to integrate resources into a unified format for scalable analysis. This fragmentation limits the ability of AI systems to fully leverage biomedical data for scientific discovery. To address these challenges, we developed BioMedGraphica , an all-in-one platform that harmonizes fragmented biomedical resources by integrating 11 entity types and 30 relation types from 43 databases into a unified knowledge graph containing 2,306,921 entities and 27,232,091 relations. In addition, to the best of our knowledge, this is the first work to propose a novel Textual-Numeric Graph (TNG) data-structure for multi-omics data analysis. In TNG, textual information captures prior biological knowledge (e.g., transcription start sites, functions, mechanisms), while numeric values represent quantitative biomedical features, and the integrated relations can help uncover mechanisms. By bridging prior knowledge with user-specific data, TNG is a novel and ideal data-structure for the development of graph foundation models, with the potential to improve prediction performance and interpretability, while also augmenting LLMs by supplying graph-structured mechanistic context to strengthen reasoning. The details for BioMedGraphica code can be accessed by github link: https://github.com/FuhaiLiAiLab/BioMedGraphica and BioMedGraphica knowledge graph data can be downloaded from huggingface dataset: https://huggingface.co/datasets/FuhaiLiAiLab/BioMedGraphica.

biomedical knowledge graph↗

Noah's Ark-Red Cross Foundation: a Swedish model.

During the Spring of 1991, the author spent many weeks at the Noah's Ark-Red Cross Foundation, a support service for HIV infected persons, and their families and friends, located in Stockholm, Sweden. The purpose was to study, through interviews, observation and participation, the foundation's interactive model in order to discover what makes it work and share that knowledge with other professionals. The Noah's Ark Model consists of three spheres of activity: service, including reception services and the volunteer programme; information and education, including the Hot Line and the Newsletter; and counselling and support, including the guest house. Staff from each area interact freely with and participate in the activities of other areas. The foundation also utilizes the services of carefully trained volunteers. This use of volunteers makes it unique in Sweden. It is the dedication and flexibility of the staff and volunteers that make this model work. The report of the study follows.

Acquired Immunodeficiency Syndrome↗

Anatomical reasoning in the informatics age: Principles, ontologies, and agendas.

Reasoning about anatomy shares historical scientific roots with formal logic and artificial intelligence. With advances in computer-based intelligent programming, high-level biological structural knowledge may be exploited directly for biomedical research, clinical tasks, and educational applications. We consider the special nature of anatomical domain knowledge, emphasizing the complex concepts and semantics that must be represented in the development of ontologies, formally structured databases of biological information. We review the evolution of the fundamental scientific principles of logic and artificial intelligence needed for building machines that can make use of anatomical knowledge. We look at methods for compiling ontologies and compare the structural designs of the Foundational Model of Anatomy and Open GALEN ontologies. We further consider issues related to mapping developing anatomy resources with other biological ontologies in genomics, proteomics, and physiology. Although early results are promising, considerable resources and continuing effort must be committed to completing and extending anatomical ontologies for the ultimate success of computer-based anatomical reasoning. Anat Rec (Part B: New Anat) 289B:72-84, 2006. (c) 2006 Wiley-Liss, Inc.

Artificial Intelligence↗

Supporting communication in young children with developmental disabilities.

The behavior of parents, adult caregivers, and peers comprises the critical features of community support for the development of communication in young children with developmental disabilities. In a bio-ecological model of development, communication development is the result of the interactions of individuals with specific characteristics, in particular contexts over time. From the perspective of this model, foundational findings of intervention research to current views of communication development in children with developmental disabilities are summarized. The contributions of individual child characteristics to child-caregiver interactions that support language development are illustrated based on research with children who have autism, Williams syndrome, Down syndrome, and children who use augmentative communication systems. Parent-child interaction and the quality and quantity of parent talk are discussed as factors in children's language development. The effects of young children's delayed language on their interactions with peers, the contributions of peers to children's language learning and use, and the critical features of classroom settings that support child language development are reviewed. MRDD Research Reviews 7:143-150, 2001.

Child↗