Search PubMedSearch

SEARCH · Search PubMed

Results for “Generative models”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis

scSurv: a deep generative model for single-cell survival analysis.

MOTIVATION: Single-cell omics analysis has unveiled the heterogeneity of various cell types within tumors. However, no methodology currently reveals how this heterogeneity influences cancer patient survival at single-cell resolution. Here, we introduce scSurv, combining a Cox proportional hazards model with a deep generative model of single-cell transcriptome, to estimate individual cellular contributions to clinical outcomes. RESULTS: The accuracy of scSurv was validated using both simulated and real datasets. This method identifies cells associated with favorable or adverse prognoses and extracts genes correlated with their contribution levels. In melanoma, scSurv reproduces known prognostic macrophage classifications and facilitates hazard mapping through spatial transcriptomics in renal cell carcinoma. We also identified genes consistently associated with prognosis across multiple cancers and demonstrated the applicability of this method to infectious diseases. scSurv is a novel framework for quantifying the heterogeneity of individual cellular effects on clinical outcomes. AVAILABILITY: The implementation of scSurv is available on GitHub (https://github.com/3254c/scSurv) and Zenodo (https://doi.org/10.5281/zenodo.17793054).

Humans

A probabilistic generative model for quantification of DNA modifications enables analysis of demethylation pathways.

We present a generative model, Lux, to quantify DNA methylation modifications from any combination of bisulfite sequencing approaches, including reduced, oxidative, TET-assisted, chemical-modification assisted, and methylase-assisted bisulfite sequencing data. Lux models all cytosine modifications (C, 5mC, 5hmC, 5fC, and 5caC) simultaneously together with experimental parameters, including bisulfite conversion and oxidation efficiencies, as well as various chemical labeling and protection steps. We show that Lux improves the quantification and comparison of cytosine modification levels and that Lux can process any oxidized methylcytosine sequencing data sets to quantify all cytosine modifications. Analysis of targeted data from Tet2-knockdown embryonic stem cells and T cells during development demonstrates DNA modification quantification at unprecedented detail, quantifies active demethylation pathways and reveals 5hmC localization in putative regulatory regions.

5-Methylcytosine

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

Generative model for the first cell fate bifurcation in mammalian development.

The first cell fate bifurcation in mammalian development directs cells toward either the trophectoderm (TE) or inner cell mass (ICM) compartments in pre-implantation embryos. This decision is regulated by the subcellular localization of a transcriptional co-activator YAP and takes place over several progressively asynchronous cleavage divisions. As a result of this asynchrony and variable arrangement of blastomeres, reconstructing the dynamics of the TE/ICM cell specification from fixed embryos is extremely challenging. To address this, we developed a live-imaging approach and applied it to measure pairwise dynamics of nuclear YAP and its direct target genes, CDX2 and SOX2, which are key transcription factors of the TE and ICM, respectively. Using these datasets, we constructed a generative model of the first cell fate bifurcation, which reveals the time-dependent statistics of the TE and ICM cell allocation. In addition to making testable predictions for the joint dynamics of the full YAP/CDX2/SOX2 motif, the model revealed the stochastic nature of the induction timing of the key cell fate determinants and identified the features of YAP dynamics that are necessary or sufficient for this induction. Notably, temporal heterogeneity was particularly prominent for SOX2 expression among ICM cells. As heterogeneities within the ICM have been linked to the initiation of the second cell fate decision in the embryo, understanding the origins of this variability is of key significance. The presented approach reveals the dynamics of the first cell fate choice and lays the groundwork for dissecting the next cell fate decisions in mouse development.

Animals

CoxFormer enables spatial omics inference with multimodal generative modeling.

Gene co-expression maps transcriptome-wide gene-gene relationships, yet high-quality estimates cover less than half the genome. Meanwhile, spatial omics either profiles restricted in situ panels or lacks cellular resolution. Extending co-expression transcriptome-wide could overcome these limitations by inferring unassayed gene expression at subcellular resolution. Here we show that CoxFormer integrates literature-derived gene knowledge with co-expression networks from bulk tissues and large-scale single-cell atlases to learn 512-dimensional representations for 32,016 human genes. These embeddings capture functional gene relationships and serve as a generative prior for spatial inference across platforms and modalities. Without requiring a matched single-cell RNA-sequencing reference, CoxFormer supports four applications beyond measured genes: histology-based expression imputation, gene activity prediction from chromatin accessibility, subcellular super-resolution inference, and pathological region detection. Together, CoxFormer extends gene embedding from gene- and cell-level tasks to whole-transcriptome spatial inference, providing a unified framework for biological analysis beyond the limited gene coverage of current spatial omics technologies.

Humans

Generative AI Models in Time-Varying Biomedical Data: Scoping Review.

BACKGROUND: Trajectory modeling is a long-standing challenge in the application of computational methods to health care. In the age of big data, traditional statistical and machine learning methods do not achieve satisfactory results as they often fail to capture the complex underlying distributions of multimodal health data and long-term dependencies throughout medical histories. Recent advances in generative artificial intelligence (AI) have provided powerful tools to represent complex distributions and patterns with minimal underlying assumptions, with major impact in fields such as finance and environmental sciences, prompting researchers to apply these methods for disease modeling in health care. OBJECTIVE: While AI methods have proven powerful, their application in clinical practice remains limited due to their highly complex nature. The proliferation of AI algorithms also poses a significant challenge for nondevelopers to track and incorporate these advances into clinical research and application. In this paper, we introduce basic concepts in generative AI and discuss current algorithms and how they can be applied to health care for practitioners with little background in computer science. METHODS: We surveyed peer-reviewed papers on generative AI models with specific applications to time-series health data. Our search included single- and multimodal generative AI models that operated over structured and unstructured data, physiological waveforms, medical imaging, and multi-omics data. We introduce current generative AI methods, review their applications, and discuss their limitations and future directions in each data modality. RESULTS: We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and reviewed 155 articles on generative AI applications to time-series health care data across modalities. Furthermore, we offer a systematic framework for clinicians to easily identify suitable AI methods for their data and task at hand. CONCLUSIONS: We reviewed and critiqued existing applications of generative AI to time-series health data with the aim of bridging the gap between computational methods and clinical application. We also identified the shortcomings of existing approaches and highlighted recent advances in generative AI that represent promising directions for health care modeling.

Artificial Intelligence

CRISPR/Cpf1-mediated knockout of FLG in human induced pluripotent stem cells generates a model for studying epidermal barrier dysfunction.

Loss of filaggrin (FLG) function impairs skin barrier formation and contributes to common inflammatory skin diseases. In this study, we established a FLG knockout human induced pluripotent stem cell (iPSC) line based on KOLF2.1 J using CRISPR/Cas12a (Cpf1)-mediated genome editing. A guide RNA targeting exon 2 introduced a homozygous mutation, which was confirmed by sequencing. The edited cells maintained typical pluripotent stem cell morphology, expressed key undifferentiated markers, and retained the ability to differentiate into all three germ layers. Karyotype and copy number variation (CNV) analyses confirmed genomic stability and parental origin; the cells were free of mycoplasma. This cell line enables studies of FLG-associated skin biology and pathology.

Humans

Model-directed generation of artificial CRISPR-Cas13a guide RNA sequences improves nucleic acid detection.

CRISPR guide RNA sequences deriving exactly from natural sequences may not perform optimally in every application. Here we implement and evaluate algorithms for designing maximally fit, artificial CRISPR-Cas13a guides with multiple mismatches to natural sequences that are tailored for diagnostic applications. These guides offer more sensitive detection of diverse pathogens and discrimination of pathogen variants compared with guides derived directly from natural sequences and illuminate design principles that broaden Cas13a targeting.

CRISPR-Cas Systems

GENKI: A generative framework for scalable and robust metabolic kinetic modeling.

GENKI (Generative ENsemble KPI-Informed) is a variational autoencoder-based framework for large-scale kinetic modeling of metabolism. Developed for metabolic engineering applications, GENKI is designed to improve the recovery of kinetically feasible models that reproduce experimentally observed phenotypes under genetic and environmental perturbations. The framework is trained on feasible kinetic model ensembles and uses phenotype-based key performance indicators (KPIs), derived from multi-omics and bioprocess data, to label and enrich models according to their agreement with mutant and condition-specific observations. This enables targeted generation of biologically relevant parameter sets with improved predictive performance. Crucially, GENKI recovers kinetic parameter sets that jointly reproduce wild-type and multiple perturbed physiologies within a single model. We apply GENKI to large-scale kinetic models of Escherichia coli and Saccharomyces cerevisiae under enzyme perturbations and oxygen shifts. In both systems, GENKI enriches kinetic ensembles with models that more accurately reproduce experimentally observed physiologies across multiple perturbations and conditions. GENKI therefore provides a practical framework for perturbation-aware kinetic model refinement within iterative Design-Build-Test-Learn workflows.

DBTL

Toward AI Virtual Cells for Hepatology: Representation, Generation, Dynamics, and Intervention in Single-Cell Models.

``Single-cell and spatial atlases describe the healthy and diseased liver at high resolution, including lobular hepatocyte zonation, fibrotic macrophage-stellate niches, cholangiocyte reactions, immune remodeling, and hepatocellular carcinoma ecosystems. These maps show where cell states occur but do not, by themselves, predict whether liver injury will progress or how the liver will respond to an untested drug, toxicant, or genetic perturbation. In this review, we organize current approaches toward an AI Virtual Cell (AIVC) for the liver into three complementary modeling routes. Generative models represent cell states, dynamics and transport models infer state transitions, and pretrained or foundation models test whether learned representations transfer across donors, etiologies, disease stages, and platforms. Perturbation-response prediction serves as a cross-cutting assessment of whether these layers can predict responses to untested genetic, chemical, inflammatory, or metabolic interventions. Available evidence can be categorized as direct liver validation, liver-included benchmarks, general single-cell evidence, and conceptual applications. Published models demonstrate individual components, including atlas integration, inferred trajectories, transferable representations, and retrospective response programs. However, these models do not constitute a prospectively validated liver simulator. At minimum, evaluation should include donor-, etiology-, stage-, platform-, and perturbation-level hold-outs. Model performance should be reported using response direction, recovery of differentially expressed genes and rare states, and calibrated uncertainty. Claims about tissue- or function-level prediction additionally require independent spatial, histologic, metabolic, and functional readouts. Near-term use should prioritize experiment selection and hypothesis generation, whereas clinical decision support remains a longer-term objective.

AI Virtual Cell

PRISM-G: an interpretable privacy scoring framework for assessing risk in synthetic human genome data.

MOTIVATION: Synthetic genomic data promises broader data access, but unresolved privacy risks remain a major concern. Existing evaluations often rely on similarity-based metrics that measure proximity between real and synthetic genomes, overlooking additional mechanisms through which genomic information may leak. RESULTS: We introduce PRISM-G, a model-agnostic framework that quantifies privacy exposure in synthetic genomic data across three complementary components: proximity to real genomes in genetic-coordinate space, replay of familial or population-structure patterns, and trait-linked exposure through rare variants and membership-inference signals. These components are normalized and combined through a risk-averse aggregation into a single 0-100 PRISM-G score. By pairing PRISM-G with downstream utility metrics, the framework also enables analysis of privacy-utility trade-offs across generative models. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT solver (Genomator). Our results show that privacy vulnerabilities arise along different axes across models and marker densities, demonstrating that a single similarity-based metric is insufficient to characterize genomic privacy risk. AVAILABILITY AND IMPLEMENTATION: The source code of PRISM-G is available at https://github.com/alejocrojo09/prismg.

Humans

Looked but didn't see: inattentional blindness and yes-bias confabulation in vision-language models.

Previous work showed that many participants fail to notice a gorilla in a video of people playing basketball. Another study found that 83% of trained radiologists failed to report a gorilla figure inserted into a chest CT nodule-search task, even though eye-tracking revealed that most observers had foveated the figure. We ask whether a similar phenomenon exists in contemporary vision-language models (VLMs). We find that (i) VLMs are capable of spotting the gorilla in both still-frame images and videos of lung CT scans; (ii) models display inattentional blindness, which varies according to model generation and type of stimulus presented; (iii) Gemini-3.1-Pro outperforms most other flagship and open-weight VLMs at identifying the presence or absence of the gorilla. We additionally ran a segmentation experiment utilizing two different model classes: a generalist (SAM 3), which found the gorilla but produced little to no results for anatomy-based prompts; a medical specialist (BiomedParse), which produced more promising anatomy-based results but flagged "gorilla" on gorilla-free control videos on 82% of frames. The behavioral signature of inattentional blindness reproduces in VLMs, but a unique confabulation failure mode means that any "did the model see X" claim requires signal-detection analysis with a matched-control false-alarm baseline.

Journal Article

BioEMMA: Automated Generation of Model-Specific Escher-Compatible Maps from KEGG Pathways.

Genome-scale metabolic models are widely used to investigate cellular metabolism, but their interpretation and comparison are limited by the lack of reproducible pathway-level visualizations with a common spatial organization. This study presents BioEMMA, a Python-based tool for the automated generation of model-specific metabolic pathway maps in the Escher JSON format using coordinate information from curated KEGG pathway maps. BioEMMA parses KGML files, map reaction and metabolite identifiers to model database namespaces, filters pathway elements according to an input SBML model, adds non-primary metabolites, reconstructs Escher-compatible layouts, and supports flux visualization. The tool was integrated into a reproducible BioUML workflow for metabolic model reconstruction. BioEMMA was evaluated using the e_coli_core model and the KEGG glycolysis/gluconeogenesis pathway while generating a model-specific map with overlaid FBA fluxes. It was then applied to compare E. coli reconstructions generated by gapseq, ModelSEEDpy, and Reconstructor across three central carbon metabolism pathways. To broaden the evaluation, BioEMMA was applied using 87 prokaryotic BiGG models and three eukaryotic models. The analysis revealed pathway-specific differences in reaction coverage, shared and model-specific reactions, and predicted flux activity. BioEMMA therefore provides a reproducible framework for pathway-level visualization and comparison of genome-scale metabolic reconstructions within a common spatial coordinate system.

Escher maps

Creating bottom-up RNA transfer vehicles from synthetic protein assemblies.

Evolution guides biological systems to populate ecological niches, with viruses among the most successful examples of this principle. Viruses evolved over billions of years to efficiently transfer genetic information. Although viruses are highly diverse, most have converged towards remarkable similarity in the size and shape of their capsids1,2. By contrast, generative models for protein design enable the creation of protein architectures that are absent from nature3-5. Here we investigate whether protein assemblies designed by artificial intelligence can be functionalized to construct nucleic acid transport vehicles that are independent of evolutionary trajectories. By combining natural protein domains with synthetic protein assemblies, we create more than 100 bottom-up RNA transfer vehicles with unique sizes and shapes. These vehicles surpass the RNA transfer efficiency of widely used delivery vehicles by several orders of magnitude. In addition, we demonstrate that their tropism can be programmed by incorporation of computationally designed peptide binders and use them to deliver therapeutically relevant cargo RNAs into a wide range of cellular models. We show the in vivo biodistribution of one of these vehicles in a mouse at near-single-cell resolution, confirm its safety, and use it to perform a gene-editing treatment strategy for Duchenne muscular dystrophy in patient-derived cells and a pig. Our work demonstrates how proteins created by generative artificial intelligence can be harnessed for the rational engineering of RNA transport systems with the desired properties by overcoming the limitations of natural protein diversity.

Journal Article

Molecular Pathways, Target Landscape, and Translational Models in Heart Failure with Preserved Ejection Fraction.

Heart failure with preserved ejection fraction (HFpEF) is a substantial global health burden and the greatest unmet medical need for cardiovascular diseases. It is marked by pronounced clinical heterogeneity and complex multi-system pathophysiology with limited therapeutic options. Progress in developing effective therapeutics is constrained by the inadequacy of experimental models to fully recapitulate the multifactorial nature of the disease. Recent evidence underscores the significant involvement of inflammatory, oxidative, and mitochondrial pathways in the pathogenesis of HFpEF, with non-coding RNAs and epigenetic regulation serving as crucial modulators and prospective therapeutic targets. This review maps the HFpEF target landscape, while critically assessing the mechanistic contributions, translational fidelity, and limitations of existing in vivo and in vitro models. Further, advances are noted among the in vitro technologies, including human cardiac organoids and engineered heart tissues integrated with high-throughput multi-omics and computational modeling, enabling in-depth examination of HFpEF mechanisms. Finally, we underscore the necessity of integrative, systems-level approaches and multi-marker strategies to enhance translational relevance, improve risk stratification, and accelerate development of mechanism-based therapies. Collectively, this review supports phenotypic-guided and mechanism-informed therapeutic development for HFpEF, and provides a roadmap for next generation model development and therapeutic innovation.

Humans

Embed-Search-Align: DNA sequence alignment using Transformer models.

MOTIVATION: DNA sequence alignment, an important genomic task, involves assigning short DNA reads to the most probable locations on an extensive reference genome. Conventional methods tackle this challenge in two steps: genome indexing followed by efficient search to locate likely positions for given reads. Building on the success of Large Language Models in encoding text into embeddings, where the distance metric captures semantic similarity, recent efforts have encoded DNA sequences into vectors using Transformers and have shown promising results in tasks involving classification of short DNA sequences. Performance at sequence classification tasks does not, however, guarantee sequence alignment, where it is necessary to conduct a genome-wide search to align every read successfully, a significantly longer-range task by comparison. RESULTS: We bridge this gap by developing a "Embed-Search-Align" (ESA) framework, where a novel Reference-Free DNA Embedding (RDE) Transformer model generates vector embeddings of reads and fragments of the reference in a shared vector space; read-fragment distance metric is then used as a surrogate for sequence similarity. ESA introduces: (i) Contrastive loss for self-supervised training of DNA sequence representations, facilitating rich reference-free, sequence-level embeddings, and (ii) a DNA vector store to enable search across fragments on a global scale. RDE is 99% accurate when aligning 250-length reads onto a human reference genome of 3 gigabases (single-haploid), rivaling conventional algorithmic sequence alignment methods such as Bowtie and BWA-Mem. RDE far exceeds the performance of six recent DNA-Transformer model baselines such as Nucleotide Transformer, Hyena-DNA, and shows task transfer across chromosomes and species. AVAILABILITY AND IMPLEMENTATION: Please see https://anonymous.4open.science/r/dna2vec-7E4E/readme.md.

Sequence Analysis, DNA