Search PubMedSearch

SEARCH · Search PubMed

Results for “Perturbation data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A transcription factor regulatory atlas for activity inference and perturbation prediction.

Inferring transcription factor (TF) activity from transcriptomes and predicting transcriptome-wide responses to TF perturbations remain challenging, in part because available TF-mRNA resources often face a trade-off between precision and coverage and typically lack signed regulatory information. Here, we present TFActProfiler, a TF-mRNA resource and computational framework that learns signed, quantitative TF-mRNA regulatory coefficients by integrating heterogeneous prior evidence (ChIP-based, motif-based, and curated TF-mRNA annotations) with large-scale bulk and single-cell RNA-seq atlases. TFActProfiler contains 2 606 176 signed TF-mRNA interactions and improves TF activity inference in TF knockdown benchmarks relative to widely used regulon resources while retaining broad TF and target coverage. In addition, because the same learned regulatory coefficients can be used to model downstream transcriptional effects, TFActProfiler enables prediction of transcriptome-wide gene expression responses to TF knockdown without training on task-matched perturbation data. When perturbation datasets are available, TFActProfiler can be further refined to achieve performance comparable to state-of-the-art machine-learning baselines. By providing a direction-aware representation of TF-mRNA regulation for both activity inference and perturbation-response modeling, TFActProfiler supports systematic dissection of gene regulatory programs across diverse cellular contexts.

Transcription Factors

Decoding heterogeneous single-cell perturbation responses.

Understanding how cells respond differently to perturbation is crucial in cell biology, but existing methods often fail to accurately quantify and interpret heterogeneous single-cell responses. Here we introduce the perturbation-response score (PS), a method to quantify diverse perturbation responses at a single-cell level. Applied to single-cell perturbation datasets such as Perturb-seq, PS outperforms existing methods in quantifying partial gene perturbations. PS further enables single-cell dosage analysis without needing to titrate perturbations, and identifies 'buffered' and 'sensitive' response patterns of essential genes, depending on whether their moderate perturbations lead to strong downstream effects. PS reveals differential cellular responses on perturbing key genes in contexts such as T cell stimulation, latent HIV-1 expression and pancreatic differentiation. Notably, we identified a previously unknown role for the coiled-coil domain containing 6 (CCDC6) in regulating liver and pancreatic cell fate decisions. PS provides a powerful method for dose-to-function analysis, offering deeper insights from single-cell perturbation data.

Single-Cell Analysis

Gene regulatory network structure informs the distribution of perturbation effects.

Gene regulatory networks (GRNs) govern many core developmental and biological processes underlying human complex traits. Even with broad-scale efforts to characterize the effects of molecular perturbations and interpret gene coexpression, it remains challenging to infer the architecture of gene regulation in a precise and efficient manner. Key properties of GRNs, like hierarchical structure, modular organization, and sparsity, provide both challenges and opportunities for this objective. Here, we seek to better understand properties of GRNs using a new approach to simulate their structure and model their function. We produce realistic network structures with a novel generating algorithm based on insights from small-world network theory, and we model gene expression regulation using stochastic differential equations formulated to accommodate modeling molecular perturbations. With these tools, we systematically describe the effects of gene knockouts within and across GRNs, finding a subset of networks that recapitulate features of a recent genome-scale perturbation study. With deeper analysis of these exemplar networks, we consider future avenues to map the architecture of gene expression regulation using data from cells in perturbed and unperturbed states, finding that while perturbation data are critical to discover specific regulatory interactions, data from unperturbed cells may be sufficient to reveal regulatory programs.

Gene Regulatory Networks

[Beta-lactoglobulin AB fluorescence under different physico-chemical conditions. Denaturation by urea and organic solvents].

Dependences of different fluorescence parameters of bovine beta-lactoglobulin AB on the concentrations of urea (pH 2.8-8.8), ethanol (pH 2.1-10.2), and dioxane (pH 5.3) have been investigated. The denaturation properties (the free energy and the stoichiometry of denaturative interaction) are highly dependent on pH values. The data obtained indicate that the hydrophobic interactions are the determining forces in the stabilization process of the beta-lactoglobulin molecule. The relative contribution of these interactions lowers with pH rise. The denaturation of beta-lactoglobulin AB proceeds through two stages under conditions when the protein octamer exists. Up to 30 vol.% of ethanol and dioxane, the penetration of the organic molecules into the external parts of the protein globule takes place. At the concentration of the solvent exceeding 50 vol.% structural transitions are observed. The comparison of fluorescence and perturbation spectral data enables one to localise tryptophan residues in the protein more precisely. The results of this and former reports lead to hypothesis that beta-lactoglobulin may serve as a transporter of some substances which are unstable to acidic media.

Animals

Relaxation spectra of yeast hexokinases. Isomerization of the enzyme.

Yeast hexokinase isozymes P1 and P11 exhibit a pH dependent, rapid relaxation process at 15 degrees C at enzyme concentrations of 100-474 muM and over a pH range of 6-8. The process was detected by equilibrium temperature jump spectroscopy using the indicator probe phenol red. The value of 1/tau varies from about 6 ms-1 at pH 8 for both isozymes to 50 ms-1 for P1 and 85 ms-1 for P11 at pH 6. The data are consistent with a mechanism involving an enzyme isomerization coupled to an ionization. The forward rate constant for the isomerization of the proposed mechanism varies between 3 and 7 ms-1; the ratio of the reverse rate constant to the ionization Ka is between 0.5 and 2 X 10(11) M-1 S-1; the estimated pKa varies between 5.5 and 6.1. The ranges of values in rate constants and pKa represent variations observed between preparations of the same isozyme and between isozymes. The isomerization rate is at least 50 times faster than catalysis under all conditions and the pKa is lower than that controlling activity. The rate of isomerization is unchanged by addition of sugar and nucleotide ligands, but the amplitude of the process is perturbed. These data imply that isomerizing and ionizing forms are sensitive to events at the active site. These equilibria between forms of hexokinase are fast enough, and have the right properties, to be important to the mechanism and regulation of the enzyme.

Adenosine Triphosphate

Disagreement-informed arbitration for gene regulatory network inference: A score-level meta-classifier and a diagnostic typology of inter-method conflict.

Gene regulatory network inference methods routinely disagree about individual edges, and practitioners resolve those conflicts by choosing one method or averaging them all. We ask whether the conflict can instead be arbitrated per edge. A gradient-boosted classifier is trained on the raw scores that ten inference methods-correlation-based, information-theoretic, sparse-regression and tree-ensemble, including GENIE3, GRNBoost2, CLR and ARACNe-assign to each candidate regulator-target pair, so that the weight given to each method varies from edge to edge. Across six single-cell perturbation screens spanning four cell types, arbitration improves on mean ensembling by +0.056 AUROC on Adamson and +0.083 on Shifrut under target-grouped cross-validation. The evaluation protocol turns out to matter more than the model. Edge-level cross-validation, standard in this literature, inflates apparent gains by 0.060 AUROC through target-gene leakage-comparable to the entire honest improvement. The effect is far larger for methods that represent genes implicitly: a supervised graph-attention link predictor trained on identical folds scores AUROC 0.930 under edge-level cross-validation, better than anything else we evaluate, and 0.533 once target genes are held out. Any method that parameterises genes is exposed, which covers most graph- and embedding-based approaches. A five-category typology of inter-method conflict localises where arbitration pays off, with the largest gains on edges where the methods disagree and the smallest where they already agree, while adding nothing as model input; we therefore report it as a diagnostic instrument rather than a modelling contribution. We also characterise what the ground truth measures: most perturbed genes in widely used screens are not transcription factors, and a mediation screen bounds how much of the perturbation response can be direct.

Ensemble methods

Predicting cellular responses to perturbation across diverse contexts with State.

While machine learning models offer potential for predicting transcriptomic effects of perturbation, they currently struggle to generalize across cellular contexts. Here, we introduce State, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. State is trained using single-cell gene expression data to predict perturbation effects across sets of cells. State improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling, and chemical perturbations with significantly improved accuracy compared with baselines. Its cell embeddings trained on observational data from 167 million cells enable the identification of strong perturbations in cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that can be used to evaluate future models. Overall, the performance and flexibility of State set the stage for scaling the development of AI models of cell state.

Machine Learning

BMDx2: A Tool for Integrating Toxicogenomics-Based Dose-Dependency Analysis and AOP-Based Mechanistic Insights.

Despite the advent of mechanistic toxicology using omics data to link molecular perturbations with systemic outcomes, regulatory toxicology still lacks the application of mechanism-anchored metrics from such data. This is partially because traditional gene-centric analysis often falls short of linking molecular changes to adverse outcomes. To address this gap, BMDx2, an open-source tool that transforms multi-dose toxicogenomics datasets into quantitative, mechanistic evidence for human chemical safety assessment is developed. BMDx2 couples benchmark-dose modeling with Adverse Outcome Pathway (AOP) enrichment to derive transcriptomic-based points of departure, enabling potency ranking, chemical prioritization, and mechanistically anchored explanations of the effect of chemical exposures. BMDx2 can process a broad range of data, including DNA microarray and RNA sequencing studies. Here, case studies are used to illustrate the versatility of BMDx2 in characterizing the mechanism of action of chemicals. An initial case study on carbon nanotubes exposure applies integrative analysis of transcriptomics and genome-wide DNA methylation data, uncovering cellular reprogramming processes underlying fibrosis. A second case study on bleomycin exposure demonstrate how transcriptomic data alone can be mapped to fibrosis-related AOPs in a standardized, regulatory appropriate manner. Together, these examples show how BMDx2 supports the regulatory application of toxicogenomics and accelerates mechanism-based chemical safety evaluation.

Toxicogenetics

Mapping transcriptional responses to cellular perturbation dictionaries with RNA fingerprinting.

Single-cell perturbation dictionaries provide systematic measurements of how cells respond to genetic and chemical perturbations, and create the opportunity to assign causal interpretations to observational data. Here, we introduce RNA fingerprinting, a statistical framework that maps transcriptional responses from new experiments onto reference perturbation dictionaries. RNA fingerprinting learns denoised perturbation "fingerprints" from single-cell data, then probabilistically assigns query cells to one or more candidate perturbations while accounting for uncertainty. We benchmark our method across ground-truth datasets, demonstrating accurate assignments at single-cell resolution, scalability to genome-wide screens, and the ability to resolve combinatorial perturbations. We demonstrate its broad utility across diverse biological settings: identifying context-specific regulators of p53 under ribosomal stress, characterizing drug mechanisms of action and dose-dependent off-target effects, and uncovering cytokine-driven B cell heterogeneity during secondary influenza infection in vivo. Together, these results establish RNA fingerprinting as a versatile framework for interpreting single-cell datasets by linking cellular states to the underlying perturbations which generated them.

Journal Article

AI proteomics: from protein identification to virtual cells.

Artificial intelligence (AI) is transforming scientific research, including proteomics. In this Perspective, we highlight key mass spectrometry (MS)-based proteomics areas where AI is driving innovation, ranging from protein identification to building AI virtual cells. These include improving peptide and protein identification and quantification; characterizing protein-protein interactions and protein complexes; advancing spatial and perturbation proteomics; integrating multi-omics data; and, ultimately, enabling AI virtual cells. Finally, we call for global collaboration among data producers, data consumers and other stakeholders to establish an AI-friendly ecosystem for MS-based proteomics, laying the foundation for transformative advancements in proteomics driven by AI.

Proteomics

Habitat Specialisation Impacts Clownfish Demographic Resilience to Pleistocene Sea-Level Fluctuations.

Habitat fragmentation and loss are key threats to biodiversity, yet their impacts on marine species remain poorly understood. Clownfishes, which rely on sea anemones for shelter and reproduction, provide an interesting model to explore how ecological specialisation mediates species responses to habitat perturbations. We used whole-genome data from 382 individuals across 10 species with varying host specialisations to reconstruct demographic histories and infer spatial genetic structure to assess the impact of Pleistocene sea-level fluctuations. Generalist species, associated with multiple hosts, maintained stable effective population sizes () and population connectivity during habitat fragmentation, reflecting resilience to environmental instability. In contrast, specialists experienced severedeclines and genetic structuring, driven by their dependence on specific hosts, without signs of population recovery following habitat reconnection. Spatial genomic analyses identified the Indonesian Through-Flow as a key dispersal corridor and the Coral Triangle as a critical hub of genetic diversity, while continental shelves and extensive open ocean regions appeared as barriers to gene flow. Our findings reveal how host specialisation shapes clownfish population dynamics, emphasising the importance of incorporating ecological dependencies into conservation assessments and deepening our understanding of species responses to ecological constraints and environmental changes over evolutionary timescales.

Animals

Iterative, multimodal, and scalable single-cell profiling for discovery and characterization of signaling regulators.

Cell signaling plays a critical role in regulating cellular state, yet uncovering regulators of signaling pathways and understanding their molecular consequences remains challenging. Here, we present an iterative experimental and computational framework to identify and characterize regulators of signaling proteins, using the mTOR marker phosphorylated RPS6 (pRPS6) as a case study. We present a customized workflow that uses the 10x Flex assay to jointly profile intracellular protein levels, transcriptomes, and CRISPR perturbations in single cells. We use this to generate a "glossary" dataset of paired protein-RNA measurements across targeted perturbations, which we leverage to train a predictive model of pRPS6 levels based solely on transcriptomic data. Applying this model to a genome-wide Perturb-seq dataset enables in silico screening for pRPS6 and nominates novel regulators of mTOR signaling. Experimental validation confirms these predictions and reveals mechanistic diversity among hits, including changes in signaling output driven by anabolic activity, cellular proliferation and multiple stress pathways. Our work demonstrates how integrated experimental and computational approaches provide a scalable framework for multimodal phenotyping and discovery.

Journal Article

Active learning of enhancer and silencer regulatory grammar in photoreceptors.

Cis-regulatory elements (CREs) direct gene expression in health and disease, and models that can accurately predict their activities from DNA sequences are crucial for biomedicine. Deep learning represents one emerging strategy to model the regulatory grammar that relates CRE sequence to function. However, these models require training data on a scale that exceeds the number of CREs in the genome. We address this problem using active machine learning to iteratively train models on multiple rounds of synthetic DNA sequences assayed in live mammalian retinas. During each round of training the model actively selects sequence perturbations to assay, thereby efficiently generating informative training data. We iteratively trained a model that predicts the activities of sequences containing binding motifs for the photoreceptor transcription factor Cone-rod homeobox (CRX) using an order of magnitude less training data than current approaches. The model's internal confidence estimates of its predictions are reliable guides for designing sequences with high activity. The model correctly identified critical sequence differences between active and inactive sequences with nearly identical transcription factor binding sites, and revealed order and spacing preferences for combinations of motifs. Our results establish active learning as an effective method to train accurate deep learning models of cis-regulatory function after exhausting naturally occurring training examples in the genome.

Journal Article

UALCAN Mobile, an app for cancer proteogenomic data analysis.

Cancer is a complex disease affecting various organs and is a major cause of death worldwide. During cancer initiation, disease progression, and tumor metastasis, various genomic and proteomic alterations are observed. Recent technological advances have led to the generation of large amounts of molecular data, including genomics and transcriptomics. These large-scale datasets can be utilized to analyze and identify sub-class-specific cancer biomarkers and targets. However, there is a need for the development of user-friendly tools for large-scale data analysis, disseminating the analyzed data in a visualizable format to cancer researchers with no programming skills. We developed UALCAN, a comprehensive platform that allows users to integrate disparate data to better understand the genes, proteins, and pathways perturbed in cancer and make discoveries of potential biomarkers and targets. In the current study, we describe the development of the UALCAN Mobile application (app) that will provide cancer transcriptomic data obtained from The Cancer Genome Atlas (TCGA) project to evaluate protein-coding gene expression based on various stratifications, including stage, grade, race, gender, and molecular-subtypes across over 30 types of cancers. In addition, the UALCAN mobile provides data analysis options for epigenetic changes due to DNA promoter methylation and Clinical Proteomic Tumor Analysis Consortium (CPTAC) cancer proteomic data. The app provides access to large cancer molecular datasets on the go. To find changes in the expression of causative genes and proteins and to identify biomarkers and therapeutic targets, UALCAN mobile app will be extremely valuable. The "UALCAN Mobile" app is free to use and can be downloaded from both the iOS/Apple and the Android Play Store and has been downloaded over 100 times in each of iOS and android app stores.

app

A double-negative prostate cancer subtype is vulnerable to SWI/SNF-targeting degrader molecules.

Proteolysis targeting chimera (PROTAC) therapies degrading SWI/SNF ATPases offer a novel approach to interfere with androgen receptor (AR) signaling in AR-dependent castration-resistant prostate cancer (CRPC-AR). To explore the utility of SWI/SNF therapy beyond AR-sensitive CRPC, we investigated SWI/SNF-targeting agents in AR-negative CRPC. SWI/SNF targeting PROTAC treatment of cell lines and organoid models reduced the viability of not only CRPC-AR but also WNT-signaling dependent AR-negative CRPC (CRPC-WNT). The CRPC-WNT subgroup represents 11% of around 400,000 cases of CRPC worldwide who die yearly of CRPC. We discovered that SWI/SNF ATPase SMARCA4 depletion interfered with the master transcriptional regulator TCF7L2 (TCF4) in CRPC-WNT. Functionally, TCF7L2 maintains proliferation via the MAPK signaling axis in this subtype of CRPC. These data suggest a mechanistic rationale for interventions that perturb the DNA binding of the pro-proliferative TCF7L2 transcription factor (TF) and/or direct MAPK signaling inhibition in the CRPC-WNT subclass of advanced prostate cancer.

Journal Article

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans

MAdLandExpression: integrating sexual reproduction into the Physcomitrium patens expression atlas.

Physcomitrium patens is a bryophyte model system particularly valuable for evolutionary developmental and comparative genomics studies. Sexual reproduction in bryophytes offers unique insights into the evolution of land plant reproduction. Unlike seed plants, bryophytes have a dominant gametophyte phase and provide significant advantages for studying sexual reproduction, such as the possibility to maintain embryo-lethal mutants through vegetative propagation or the presence of motile male gametes. More than 25 years after the first publications of transcriptomic data for P. patens, expression data of most developmental stages of P. patens as well as its responses to various biotic and abiotic perturbations have been represented by microarrays or RNA-seq datasets. To facilitate the use of such data, we introduce the MAdLandExpression atlas as a successor of PEATmoss (Physcomitrium Expression Atlas Tool), integrating its 109 P. patens expression experiments and expanding it with 20 recently published RNA-seq samples of sexual reproduction stages, thus completing the coverage of the P. patens life cycle. The MAdLandExpression atlas also introduces new features for data visualization and analysis, such as the comparison of samples from multiple datasets and gene set normalization. Using this tool, the sexual reproduction dataset was analyzed, identifying genes potentially important for egg and sperm cell development, and confirming the behavior of known key genes in sexual development observed in previous studies.

Bryopsida

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis