Search PubMedSearch

SEARCH · Search PubMed

Results for “Generative models”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Experience with a Fourier method for determining the extracellular potential fields of excitable cells with cylindrical geometry.

In this chapter, well-known solutions that utilize a Fourier transform method for determining the extracellular, volume-conductor potential distribution surrounding elongated excitable cells of cylindrical geometry are reformulated as a discrete Fourier transform (DFT) problem, which subsequently permits the volume-conductor problem to be viewed as an equivalent linear-filtering problem. This DFT formulation is fast and computationally efficient. In addition, it lends itself to the application of some rather well-known techniques in linear systems theory (e.g., the DFT for convolution and least mean-square (Wiener) filtering for optimal prediction of a signal in random noise). Two specific examples are employed to demonstrate the utility of this discrete Fourier method: (1) the single, isolated, active nerve fiber in an essentially infinite volume conductor and (2) the isolated, active nerve trunk in a similar type of extracellular medium. In each of these, our DFT method is employed to obtain both the classical "forward" and "inverse" potential solutions for each volume conductor problem. In the case where the single, active nerve fiber is the bioelectric source in the volume conductor, simulated action-potential data from an invertebrate giant axon is utilized, and potentials at various points in the extracellular medium are calculated. The calculated potential distributions in axial distance z, at various radial distances r, are consistent with well-known experimental fact. When the active nerve trunk acts as the bioelectric source, the DFT method provides calculated potential distributions that are fairly consistent with experimental data under a variety of experimental conditions. For example, in these experiments, a special, isolated frog spinal cord preparation is used that permits separate or combined stimulation of the motor and sensory nerve fiber components of the attached sciatic nerve trunk. By manipulating the stimulus intensity applied to the motor (ventral) or appropriate sensory (dorsal) roots of the spinal cord, a variety of multiphasic extracellular volume-conductor potentials can be recorded from the sciatic nerve. The excellent agreement of model-generated and experimental data, regardless of the complexity of surface potential waveform, tends to validate the modeling assumptions and offer encouragement that this computationally efficient DFT method may be usefully employed in volume-conductor problems where both the bioelectric source, and the surrounding volume conductor, are of a much more complicated nature.

Action Potentials

PRISM-G: an interpretable privacy scoring framework for assessing risk in synthetic human genome data.

MOTIVATION: Synthetic genomic data promises broader data access, but unresolved privacy risks remain a major concern. Existing evaluations often rely on similarity-based metrics that measure proximity between real and synthetic genomes, overlooking additional mechanisms through which genomic information may leak. RESULTS: We introduce PRISM-G, a model-agnostic framework that quantifies privacy exposure in synthetic genomic data across three complementary components: proximity to real genomes in genetic-coordinate space, replay of familial or population-structure patterns, and trait-linked exposure through rare variants and membership-inference signals. These components are normalized and combined through a risk-averse aggregation into a single 0-100 PRISM-G score. By pairing PRISM-G with downstream utility metrics, the framework also enables analysis of privacy-utility trade-offs across generative models. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT solver (Genomator). Our results show that privacy vulnerabilities arise along different axes across models and marker densities, demonstrating that a single similarity-based metric is insufficient to characterize genomic privacy risk. AVAILABILITY AND IMPLEMENTATION: The source code of PRISM-G is available at https://github.com/alejocrojo09/prismg.

Humans

Looked but didn't see: inattentional blindness and yes-bias confabulation in vision-language models.

Previous work showed that many participants fail to notice a gorilla in a video of people playing basketball. Another study found that 83% of trained radiologists failed to report a gorilla figure inserted into a chest CT nodule-search task, even though eye-tracking revealed that most observers had foveated the figure. We ask whether a similar phenomenon exists in contemporary vision-language models (VLMs). We find that (i) VLMs are capable of spotting the gorilla in both still-frame images and videos of lung CT scans; (ii) models display inattentional blindness, which varies according to model generation and type of stimulus presented; (iii) Gemini-3.1-Pro outperforms most other flagship and open-weight VLMs at identifying the presence or absence of the gorilla. We additionally ran a segmentation experiment utilizing two different model classes: a generalist (SAM 3), which found the gorilla but produced little to no results for anatomy-based prompts; a medical specialist (BiomedParse), which produced more promising anatomy-based results but flagged "gorilla" on gorilla-free control videos on 82% of frames. The behavioral signature of inattentional blindness reproduces in VLMs, but a unique confabulation failure mode means that any "did the model see X" claim requires signal-detection analysis with a matched-control false-alarm baseline.

Journal Article

BioEMMA: Automated Generation of Model-Specific Escher-Compatible Maps from KEGG Pathways.

Genome-scale metabolic models are widely used to investigate cellular metabolism, but their interpretation and comparison are limited by the lack of reproducible pathway-level visualizations with a common spatial organization. This study presents BioEMMA, a Python-based tool for the automated generation of model-specific metabolic pathway maps in the Escher JSON format using coordinate information from curated KEGG pathway maps. BioEMMA parses KGML files, map reaction and metabolite identifiers to model database namespaces, filters pathway elements according to an input SBML model, adds non-primary metabolites, reconstructs Escher-compatible layouts, and supports flux visualization. The tool was integrated into a reproducible BioUML workflow for metabolic model reconstruction. BioEMMA was evaluated using the e_coli_core model and the KEGG glycolysis/gluconeogenesis pathway while generating a model-specific map with overlaid FBA fluxes. It was then applied to compare E. coli reconstructions generated by gapseq, ModelSEEDpy, and Reconstructor across three central carbon metabolism pathways. To broaden the evaluation, BioEMMA was applied using 87 prokaryotic BiGG models and three eukaryotic models. The analysis revealed pathway-specific differences in reaction coverage, shared and model-specific reactions, and predicted flux activity. BioEMMA therefore provides a reproducible framework for pathway-level visualization and comparison of genome-scale metabolic reconstructions within a common spatial coordinate system.

Escher maps

Creating bottom-up RNA transfer vehicles from synthetic protein assemblies.

Evolution guides biological systems to populate ecological niches, with viruses among the most successful examples of this principle. Viruses evolved over billions of years to efficiently transfer genetic information. Although viruses are highly diverse, most have converged towards remarkable similarity in the size and shape of their capsids1,2. By contrast, generative models for protein design enable the creation of protein architectures that are absent from nature3-5. Here we investigate whether protein assemblies designed by artificial intelligence can be functionalized to construct nucleic acid transport vehicles that are independent of evolutionary trajectories. By combining natural protein domains with synthetic protein assemblies, we create more than 100 bottom-up RNA transfer vehicles with unique sizes and shapes. These vehicles surpass the RNA transfer efficiency of widely used delivery vehicles by several orders of magnitude. In addition, we demonstrate that their tropism can be programmed by incorporation of computationally designed peptide binders and use them to deliver therapeutically relevant cargo RNAs into a wide range of cellular models. We show the in vivo biodistribution of one of these vehicles in a mouse at near-single-cell resolution, confirm its safety, and use it to perform a gene-editing treatment strategy for Duchenne muscular dystrophy in patient-derived cells and a pig. Our work demonstrates how proteins created by generative artificial intelligence can be harnessed for the rational engineering of RNA transport systems with the desired properties by overcoming the limitations of natural protein diversity.

Journal Article

Molecular Pathways, Target Landscape, and Translational Models in Heart Failure with Preserved Ejection Fraction.

Heart failure with preserved ejection fraction (HFpEF) is a substantial global health burden and the greatest unmet medical need for cardiovascular diseases. It is marked by pronounced clinical heterogeneity and complex multi-system pathophysiology with limited therapeutic options. Progress in developing effective therapeutics is constrained by the inadequacy of experimental models to fully recapitulate the multifactorial nature of the disease. Recent evidence underscores the significant involvement of inflammatory, oxidative, and mitochondrial pathways in the pathogenesis of HFpEF, with non-coding RNAs and epigenetic regulation serving as crucial modulators and prospective therapeutic targets. This review maps the HFpEF target landscape, while critically assessing the mechanistic contributions, translational fidelity, and limitations of existing in vivo and in vitro models. Further, advances are noted among the in vitro technologies, including human cardiac organoids and engineered heart tissues integrated with high-throughput multi-omics and computational modeling, enabling in-depth examination of HFpEF mechanisms. Finally, we underscore the necessity of integrative, systems-level approaches and multi-marker strategies to enhance translational relevance, improve risk stratification, and accelerate development of mechanism-based therapies. Collectively, this review supports phenotypic-guided and mechanism-informed therapeutic development for HFpEF, and provides a roadmap for next generation model development and therapeutic innovation.

Humans

Embed-Search-Align: DNA sequence alignment using Transformer models.

MOTIVATION: DNA sequence alignment, an important genomic task, involves assigning short DNA reads to the most probable locations on an extensive reference genome. Conventional methods tackle this challenge in two steps: genome indexing followed by efficient search to locate likely positions for given reads. Building on the success of Large Language Models in encoding text into embeddings, where the distance metric captures semantic similarity, recent efforts have encoded DNA sequences into vectors using Transformers and have shown promising results in tasks involving classification of short DNA sequences. Performance at sequence classification tasks does not, however, guarantee sequence alignment, where it is necessary to conduct a genome-wide search to align every read successfully, a significantly longer-range task by comparison. RESULTS: We bridge this gap by developing a "Embed-Search-Align" (ESA) framework, where a novel Reference-Free DNA Embedding (RDE) Transformer model generates vector embeddings of reads and fragments of the reference in a shared vector space; read-fragment distance metric is then used as a surrogate for sequence similarity. ESA introduces: (i) Contrastive loss for self-supervised training of DNA sequence representations, facilitating rich reference-free, sequence-level embeddings, and (ii) a DNA vector store to enable search across fragments on a global scale. RDE is 99% accurate when aligning 250-length reads onto a human reference genome of 3 gigabases (single-haploid), rivaling conventional algorithmic sequence alignment methods such as Bowtie and BWA-Mem. RDE far exceeds the performance of six recent DNA-Transformer model baselines such as Nucleotide Transformer, Hyena-DNA, and shows task transfer across chromosomes and species. AVAILABILITY AND IMPLEMENTATION: Please see https://anonymous.4open.science/r/dna2vec-7E4E/readme.md.

Sequence Analysis, DNA

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries.

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism's own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

5' Untranslated Regions

Malnutrition and adverse outcomes after spine surgery: a systematic review and meta-analysis.

BACKGROUND CONTEXT: Malnutrition is linked to adverse surgical outcomes, but its impact in spine surgery remains unclear due to inconsistent findings and heterogeneous definitions, including use of serum albumin, prealbumin, lymphocyte count, the Geriatric Nutritional Risk Index, and the Prognostic Nutritional Index. We conducted a systematic review and meta-analysis to evaluate the relationship between malnutrition and postoperative outcomes in spine surgery. PURPOSE: To systematically evaluate the association between preoperative malnutrition and postoperative outcomes in patients undergoing spine surgery. STUDY DESIGN: Systematic review and meta-analysis. PATIENT SAMPLE: Patients undergoing elective or urgent spine surgery across included observational studies comparing malnourished vs well-nourished cohorts. OUTCOME MEASURES: Primary outcomes included postoperative mortality and overall surgical complications. Secondary outcomes included infectious complications (sepsis, urinary tract infection, wound complications), delirium, reoperation, 30-day and 90-day readmission, and prolonged length of hospital stay. METHODS: A systematic search of PubMed, Embase, Cochrane Library, and Web of Science was performed on April 7, 2025, following PRISMA guidelines. Studies directly comparing postoperative outcomes in malnourished vs well-nourished spine surgery patients were included. A random-effects model generated pooled odds ratios for complications. Outcomes assessed included mortality, surgical complications, infectious outcomes, readmission, reoperation, delirium, prolonged length of stay, and wound complications. RESULTS: Of 2,851 screened articles, 37 met the inclusion criteria, encompassing 16,987 malnourished patients. Malnutrition was associated with significantly increased odds of mortality (OR: 4.05, 95% CI [2.97-5.54]), delirium (OR: 3.95, 95% CI [2.49-6.27]), sepsis (OR: 2.77, 95% CI [2.31-3.33]), surgical complications (OR: 1.79, 95% CI [1.57-2.04]), urinary tract infection (OR: 1.81, 95% CI [1.59-2.06), wound complications (OR: 2.10, 95% CI [1.80-2.45]), reoperation (OR: 1.70, 95% CI [1.46-1.97]), prolonged length of hospital stay (OR: 3.46, 95% CI [2.57-4.65]), 30-day readmission (OR: 1.59, 95% CI [1.36-1.86]), and 90-day readmission (OR: 2.13, 95% CI [1.67-2.71]). CONCLUSIONS: Malnutrition was consistently associated with adverse outcomes after spine surgery. Routine nutritional assessment and targeted preoperative optimization should be considered a standard component of perioperative spine care to help reduce postoperative complications and improve recovery.

Humans

Insights into the Catalytic Activity of a Metagenome-Derived Urethanase.

The discovery of urethanases shows an opportunity to access the biotechnological recycling of polyurethane-based plastics (PURs), widely used in the manufacture of everyday materials. However, the mechanistic understanding of these enzymes remains under debate. In this work, we report a QM/MM-based mechanistic study of the metagenome-derived urethanase UMG-SP2 catalyzing the degradation of a urethane-like model compound, 4-nitrophenyl benzylcarbamate (pNC). A high-quality structural model generated with AlphaFold2, prior to the availability of the crystal structure, accurately captured the Ser-Ser-Lys catalytic triad characteristic of amidase signature enzymes. Highly accurate constant-pH nonequilibrium molecular dynamics and Monte Carlo (neMD/MC) simulations provided the full titration curve of active site Lys, explaining the need for alkaline media for the enzyme to be active. The generation of the free energy landscape, obtained by means of free energy perturbation methods with the M06-2X DFT functional describing the QM region of the full system, reveals an esterase-like three-step mechanism of UMG-SP2, i.e., acylation, hydrolysis, and decarboxylation, with all steps being kinetically feasible. Our computational results show very good agreement with experimental kinetic data, with a calculated free energy barrier of 21.2 kcal·mol-1 for the rate-determining step compared to 22.9 kcal·mol-1 derived from the experimentally measured turnover frequency (TOF). The present results also open the door for the final decarboxylation occurring in the solution after the release of the product of the hydrolysis step or within the active site. These findings provide an atomistic insight into the urethanase function and establish a robust framework for the future design of biocatalysts targeting polyurethane degradation.

Metagenome

Vinigrol Tricyclic Scaffold Biosynthesis Employs an Atypical Terpene Cyclase and a Multipotent Cyclization Cascade.

Vinigrol (1) is a fungal diterpenoid consisting of a decahydro-1,5-butanonaphthalene ring system with no analogs in nature. Despite immense efforts in synthetic studies, the vinigrol biosynthesis pathway remains largely unknown. Herein, we identified a biosynthetic gene cluster for 1 and fully elucidated the biosynthetic pathway. By employing an AlphaFold-generated model structure, we identified the possible catalytic residues of the noncanonical terpene cyclase and analyzed their function by site-directed mutagenesis. We found that the G340A mutation opened a cryptic pathway for an unprecedented tetracyclic diterpene, defined here as virgarene. Retro-biosynthetic theoretical analysis provided a solid foundation for the complex cyclization pathway for the vinigrol scaffold, its chemical transformation to a structurally distinct bonnadiene, and redirection of the enzymatic cyclization cascade to virgarene. Close inspection of the terpene cyclization pathway via integrated experimental and theoretical approaches would allow efficient exploration of novel terpenoid chemistries.

Cyclization

Enhancing pan-cancer spatial transcriptomics at single-cell resolution with stPainter.

Subcellular spatial transcriptomics can resolve tissue architecture at cellular scale, but sparse gene panels and limited detection sensitivity constrain downstream analysis. Existing enhancement methods often require tissue-matched single-cell RNA sequencing (scRNA-seq) references and dataset-specific retraining. Here we show that stPainter, a conditional generative model pretrained on a pan-cancer scRNA-seq atlas, can enhance spatial transcriptomics data without matched references or retraining. Using a latent diffusion architecture guided by Stochastic Differential Equations (SDE), stPainter reconstructs expanded expression profiles from sparse measurements and produces latent representations for clustering and cell-state analysis. When we apply stPainter upon 6 spatial transcriptomics datasets of different cancer types, we demonstrate that our model empowers downstream biological analyses, including fine-grained subpopulation clustering and pathway enrichment. Comparison with spatially resolved proteomics (CODEX) provided independent support for regional agreement between imputed cellular compositions and protein-level tissue organization. These results establish stPainter as a scalable approach for analyzing tumor microenvironments without auxiliary sequencing data.

Spatial Transcriptomics

The discrete-time kinetic model analysis of DNA content distributions in experimental tumour cells.

A method was developed to analyse and characterize FMF measurements of DNA content distribution, utilizing the discrete time kinetic (DTK) model for cell kinetics analysis. The DTK model determines the time sequence of the cell age distribution during the proliferation of a tumor cell population and simulates the distribution pattern of the DNA content of cells in each age compartment of the cell cycle. The cells in one age compartment are distributed and spread into several compartments of the DNA content distribution to allow for different rates of DNA synthesis and instrument dispersion effects. It is assumed that the DNA content of cells in each age compartment has a Gaussian distribution. Thus, for a given cell age distribution the DNA content distribution depends on two parameters of the cells in each age compartment: the average DNA content and its coefficient of variation. As the DTK model generates the best fit DNA content distribution to the FMF measurement data, it enables one to estimate specific values of these two parameters in each stage of the cell cycle and to determine the fraction of cells in each cycle phase. The method was utilized to fit FMf measurements of DNA content distributions and to analyse their relationship tothe cell kinetic parameters, namely cell loss rate, cell cycle times and grwoth graction of exponentially growing Chinese hamster ovary cells in vitro and, also, with a wide range of coeffficients of variation, of the L1210 ascites tumour during the growth period.

Cell Cycle

Mechanism-Driven Diagnostic Development: A Specimen-Aware Framework Illustrated by Colorectal Cancer and Solid Tumours.

Translational oncology has moved rapidly from histopathology and single-analyte biomarkers toward multi-dimensional molecular profiling. Yet many clinically deployed tests still use reductionist biomarker strategies that under-represent cancer complexity. This review examines whether a mechanistic, multi-layered, and specimen-aware approach can improve cancer detection, classification, prognosis, minimal residual disease (MRD) assessment, and therapeutic selection. Evidence across solid tumours shows that genomic alterations alone incompletely explain tumour state, metastatic behaviour, immune evasion, or therapeutic vulnerability. Integrated genome and transcriptome analyses, proteogenomics, single-cell atlases, fragmentomic, methylation based cell-free DNA assays, metabolomics and microbiome assessments reveal clinically relevant biology that single modality tests cannot determine. Minimally invasive collected specimens can extend access to screening, diagnosis and longitudinal monitoring, but the choice of specimen should be matched to disease biology and analytes that represent mechanisms of oncogenesis. However, translation remains constrained by pre-analytical variability, contamination, differences in tumour shedding behaviour, clonal haematopoiesis, translation of generated models, incomplete external validation and uncertain downstream clinical utility for emerging platforms. This review provides a commentary on the future of cancer diagnostics, the considerations and barriers to clinical translation, the relationship between utility and dimensionality of biomarkers assessed and the emerging rationale towards mechanistically grounded integrated models.

biomarkers

Construction and evaluation of an independently generated transgenic mouse model carrying mutated human HRAS genes for short-term carcinogenicity assessment.

The study aimed to construct and evaluate an independently generated transgenic mouse model applied to the short-term carcinogenicity assessment. Mutated human HRAS fragment containing an intron point-mutation was inserted into C57BL/6JGpt mice via bacterial artificial chromosome transgenic technology, eventually generating BALB/c;B6J-Tg(hHRAS)16/Gpt mice, abbreviated as HRAS mice. The inserted human HRAS fragment in HRAS mice was characterized, revealing five tandem copies at chromosome 19. Baseline profiles, including biochemical, hematological, immunophenotypic, survival, and carcinogenic data of HRAS mice, were collected. To evaluate the tumor susceptibility in HRAS mice, we applied N-Nitroso-N-methylurea (MNU) to HRAS mice in a short-term carcinogenicity assessment conducted according to Good Laboratory Practice. The genetic characteristics of HRAS mice include five tandem arrays of mutated human HRAS fragments located in genomic coordinate 7,755,606 on chromosome 19 and the duplication of a 9-kilobase genome sequence (genomic coordinate 7,755,606-7,746,509) located on chromosome 19. HRAS mice showed a relatively lower incidence and range of spontaneous tumor formation during long-term observation compared to CByB6F1-Tg(HRAS)2Jic (Tg.rasH2) transgenic mice. The short-term carcinogenicity assessment showed a strong tumor response to MNU, with high incidences of lymphoma (≥ 90%) and stomach squamous cell papilloma (≥ 90%) in both male and female HRAS mice. The HRAS mice showed susceptibility to MNU and exhibited baseline characteristics distinct from those of Tg.rasH2 mice. The co-expression of HRAS and MKI67 at the cellular localization level was found in neoplasms of HRAS mice. These findings preliminarily evaluated the feasibility of HRAS mice applied to the short-term carcinogenicity assessment.

Animals

Further characterization of Sendai virus DI-RNAs: a model for their generation.

Sendai virus DI-RNAs which contain complementary ends have been characterized as follows. First, the complementary ends of three DI-RNAs, although somewhat different in size (110-150 base pairs), contain sequences that are both identical to each other and to the 5' end of the nondefective (ND) genome. Second, almost all the sequences contained sequences that are both identical to each other and to the 5' end of the nondefective (ND) genome. Second, almost all the sequences contained in the DI-RNAs derive from sequences that are contiguous to the 5' end of the ND genome. The ND genome, on the other hand, does not contain any sequences that are complementary to its 5' end. A genetic map and a model for the generation of the Sendai DI-RNAs are presented.

Base Sequence

[Possibilities of the clinical comparison of diseases with a hereditary burden in the descendant generation using the model of "parents-children" ill with schizophrenia].

A special statistic method of confrontation of diseases with aggravation hereditary in the group "parents - children" is proposed. This method can be used under clinical formalization of the diseases. It is shown on a model group parents - children (118 pairs) suffering with schizophrenia, that statistical confrontation makes possible to work out group and individual prognoses in the descending generation. It is found that invariability in the descending generation is provided by a small number of stable indices, other symptoms being variable. Statistic description of the disease symptoms in the descending generation is of interest for planning genetic interpretation.

Adult