Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Cell annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A cross-species analysis of the rodent uterotrophic program: elucidation of conserved responses and targets of estrogen signaling.

Physiological, morphological, and transcriptional alterations elicited by ethynyl estradiol in the uteri of Sprague-Dawley rats and C57BL/6 mice were assessed using comparable study designs, microarray platforms, and analysis methods to identify conserved estrogen signaling networks. Comparative analysis identified 153 orthologous gene pairs that were positively correlated, suggesting conserved transcriptional targets important in uterine proliferation. Functional annotation for these responses were associated with angiogenesis, water and solute transport, cell cycle control, redox control, DNA replication, protein synthesis and transport, xenobiotic metabolism, cell-cell communication, energetics, and cholesterol and fatty acid regulation. The identification of conserved temporal expression patterns of these orthologs provides experimental support for the transfer of functional annotation from mouse orthologs to 44 previously unannotated rat expressed sequence tags based on their homology and co-expression patterns. The identification of comparable temporal phenotypic responses linked to related gene expression profiles demonstrates the ability of systematic comparative genomic assessments to elucidate important conserved mechanisms in rodent estrogen signaling during uterine proliferation.

Animals↗

Overexpression of MDR1 using a retroviral vector differentially regulates genes involved in detoxification and apoptosis and confers radioprotection.

Overexpression of P-glycoprotein (P-gp), the product of the MDR1 (multidrug resistance 1) gene, might complement chemotherapy and radiotherapy in the treatment of tumors. However, for safety and mechanistic reasons, it is important to know whether MDR1 overexpression influences the expression of other genes. Therefore, we analyzed differential gene expression in cells of the human lymphoblast cell line TK6 retrovirally transduced with MDR1 using the GeneChip Human Genome U133 Plus2.0 (Affymetrix). Sixty-one annotated genes showed a significant change in expression (P < 10(-4)) in MDR1-overexpressing cells compared to untransduced cells and cells transduced with a control virus expressing the neomycin phosphotransferase gene. Several genes coding for proteins involved in detoxification and exocytosis showed approximately 1.4- 4-fold increases in transcript levels (e.g. ALDH1A, UNC13). Additionally, pro-apoptosis genes were down-regulated (e.g. twofold for CASP1, 2.5-fold for NALP7) with concomitant increased expression of the potential anti-apoptosis gene AKT3. In functional assays the influence of MDR1 overexpression on apoptosis signaling was further corroborated by showing reduced rates of apoptosis in response to irradiation in TK6 cells transduced with MDR1. In conclusion, the resistant phenotype of MDR1-mediated P-gp-overexpressing cells is associated with differential expression of genes coding for metabolic and apoptosis-related proteins. These results have important implications for understanding the mechanisms by which MDR1 gene therapy can protect normal tissues from radiation- or chemotherapy-induced damage during tumor treatment.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

A four-dimensional digital image archiving system for cell lineage tracing and retrospective embryology.

The paper describes a digital image archiving system for time-lapse microscopy. The system uses an MS-DOS compatible computer to store video images while simultaneously controlling a stepping motor. In a typical experiment, images might be taken at 30 s intervals in each of 25 consecutive focal planes. A system with 2.5 Gbyte disk capacity can store approximately 18,000 full frame images: 6 h recording at maximum resolution. Once recorded, images series stored on disk can be 'played back' in any order. Generally, images from a single focal plane are displayed consecutively in either forward or reverse time. The focal plane can be shifted during playback, allowing individual cells to be followed as they move between focal planes. To facilitate the annotation and interpretation of the real-time images, a mouse-driven interface allows users to define and follow individual objects (e.g. cells). The recorded image series can be achieved inexpensively using standard digital tape backup hardware. In this laboratory, the system has been particularly useful for tracing embryonic cell lineages and cell migrations. Detailed system specifications, including source code, compiled programs, hardware requirements and users manual are available directly from the author or by anonymous FTP (ciw1.ciwemb.edu).

Algorithms↗

Annotations and functional analyses of the rice WRKY gene superfamily reveal positive and negative regulators of abscisic acid signaling in aleurone cells.

The WRKY proteins are a superfamily of regulators that control diverse developmental and physiological processes. This family was believed to be plant specific until the recent identification of WRKY genes in nonphotosynthetic eukaryotes. We have undertaken a comprehensive computational analysis of the rice (Oryza sativa) genomic sequences and predicted the structures of 81 OsWRKY genes, 48 of which are supported by full-length cDNA sequences. Eleven OsWRKY proteins contain two conserved WRKY domains, while the rest have only one. Phylogenetic analyses of the WRKY domain sequences provide support for the hypothesis that gene duplication of single- and two-domain WRKY genes, and loss of the WRKY domain, occurred in the evolutionary history of this gene family in rice. The phylogeny deduced from the WRKY domain peptide sequences is further supported by the position and phase of the intron in the regions encoding the WRKY domains. Analyses for chromosomal distributions reveal that 26% of the predicted OsWRKY genes are located on chromosome 1. Among the dozen genes tested, OsWRKY24, -51, -71, and -72 are induced by abscisic acid (ABA) in aleurone cells. Using a transient expression system, we have demonstrated that OsWRKY24 and -45 repress ABA induction of the HVA22 promoter-beta-glucuronidase construct, while OsWRKY72 and -77 synergistically interact with ABA to activate this reporter construct. This study provides a solid base for functional genomics studies of this important superfamily of regulatory genes in monocotyledonous plants and reveals a novel function for WRKY genes, i.e. mediating plant responses to ABA.

Abscisic Acid↗

Large-scale identification of proteins expressed in mouse embryonic stem cells.

A protein subset expressed in the mouse embryonic stem (ES) cell line, E14-1, was characterized by mass spectrometry-based protein identification technology and data analysis. In total, 1790 proteins including 365 potential nuclear and 260 membrane proteins were identified from tryptic digests of total cell lysates. The subset contained a variety of proteins in terms of physicochemical characteristics, subcellular localization, and biological function as defined by Gene Ontology annotation groups. In addition to many housekeeping proteins found in common with other cell types, the subset contained a group of regulatory proteins that may determine unique ES cell functions. We identified 39 transcription factors including Oct-3/4, Sox-2, and undifferentiated embryonic cell transcription factor I, which are characteristic of ES cells, 88 plasma membrane proteins including cell surface markers such as CD9 and CD81, 44 potential proteinaceous ligands for cell surface receptors including growth factors, cytokines, and hormones, and 100 cell signaling molecules. The subset also contained the products of 60 ES-specific and 41 stemness genes defined previously by the DNA microarray analysis of Ramalho-Santos et al. (Ramalho-Santos et al., Science 2002, 298, 597-600), as well as a number of components characteristic of differentiated cell types such as hematopoietic and neural cells. We also identified potential post-translational modifications in a number of ES cell proteins including five Lys acetylation sites and a single phosphorylation site. To our knowledge, this study provides the largest proteomic dataset characterized to date for a single mammalian cell species, and serves as a basic catalogue of a major proteomic subset that is expressed in mouse ES cells.

Animals↗

Mapping the immune-genetic architecture of Epstein-Barr virus-related phenotypes and multiple sclerosis through a single-cell genetic framework for target prioritization and pharmacologic hypothesis generation.

BACKGROUND: Multiple sclerosis (MS) is a severe neuroinflammatory disease causing substantial long-term disability. Strong epidemiologic evidence links Epstein-Barr virus (EBV) exposure with MS risk, but genetic evidence for immune target prioritization in EBV-related phenotypes remains limited. METHODS: We integrated single-cell cis-eQTL data from 14 immune cell types with GWASs of an EBV-related clinical phenotype and MS using a single-cell Mendelian randomization framework with colocalization analyses. Candidate eGenes were evaluated in independent cohorts. For multi-SNP instruments, we performed heterogeneity, pleiotropy, MR-Egger, weighted median, mode-based, and MR-PRESSO sensitivity analyses. We also conducted phenome-wide association analyses and queried DrugBank to annotate candidate compounds targeting prioritized genes. RESULTS: We prioritized 43 immune-cell-specific candidate eGenes with convergent genetic support, including 6 for the EBV-related phenotype and 37 for MS. SERPINB1 in NK cells was associated with increased risk of the EBV-related phenotype, whereas HLA-G was associated with decreased risk. For MS, APOM and MSH5 showed protective associations, while AHI1 showed cell-type-dependent, bidirectional associations across immune lineages. Colocalization and independent cohort evaluation supported these findings. Among FDR-significant multi-SNP associations, MR-Egger intercept tests did not indicate directional pleiotropy, although a small subset showed heterogeneity or MR-PRESSO signals. Phenome-wide analyses identified no significant adverse phenotypic associations among evaluable genes at the prespecified threshold. DrugBank annotation nominated sodium nitroprusside, fasudil, artenimol, and choline as hypothesis-generating compounds for experimental follow-up. CONCLUSIONS: This study provides a single-cell genetic framework for prioritizing immune-cell-specific candidate targets for EBV-related phenotypes and MS, and nominates genetically supported targets and pharmacologic hypotheses for experimental investigation.

Humans↗

Bio-medical entity extraction using support vector machines.

OBJECTIVE: Support vector machines (SVMs) have achieved state-of-the-art performance in several classification tasks. In this article we apply them to the identification and semantic annotation of scientific and technical terminology in the domain of molecular biology. This illustrates the extensibility of the traditional named entity task to special domains with large-scale terminologies such as those in medicine and related disciplines. METHODS AND MATERIALS: The foundation for the model is a sample of text annotated by a domain expert according to an ontology of concepts, properties and relations. The model then learns to annotate unseen terms in new texts and contexts. The results can be used for a variety of intelligent language processing applications. We illustrate SVMs capabilities using a sample of 100 journal abstracts texts taken from the {human, blood cell, transcription factor} domain of MEDLINE. RESULTS: Approximately 3400 terms are annotated and the model performs at about 74% F-score on cross-validation tests. A detailed analysis based on empirical evidence shows the contribution of various feature sets to performance. CONCLUSION: Our experiments indicate a relationship between feature window size and the amount of training data and that a combination of surface words, orthographic features and head noun features achieve the best performance among the feature sets tested.

Algorithms↗

Functional screening for proapoptotic genes by reverse transfection cell array technology.

Application of mathematical algorithms to sequenced whole genomes revealed a large number of predicted genes, requiring functional assays for their characterization in a high-throughput manner. Here, we report on the development of a screening assay, which is based on reverse transfection of cellular arrays and subsequent analysis of cell morphology to identify novel proapoptotic genes. Expression plasmids containing full-length cDNAs were cotransfected with the reporter plasmid pEYFP to screen for apoptotic body formation, based on EYFP fluorescence. The assay was validated and applied to 382 human sequence-verified full-length open reading frames, most of them of unknown function. In this initial screening, proapoptotic effects could be demonstrated for 10 of these genes. For 6 of them apoptosis induction could be confirmed both by TUNEL assay and by FACS analysis of cells stained according to Nicoletti: 1 gene was not yet annotated for an apoptotic function (ST6GAL2), while 5 genes were without annotated function (FLJ20551, CXorf12, FAM105A, TMEM66, C19orf4). Our study demonstrates the potential of this method to characterize functionally genes of unknown function in a highly parallel format.

Apoptosis↗

Tribus: semi-automated discovery of cell identities and phenotypes from multiplexed imaging and proteomic data.

MOTIVATION: Multiplexed imaging and single-cell analysis are increasingly applied to investigate the tissue spatial ecosystems in cancer and other complex diseases. Accurate single-cell phenotyping based on marker combinations is a critical but challenging task due to (i) low reproducibility across experiments with manual thresholding, and, (ii) labor-intensive ground-truth expert annotation required for learning-based methods. RESULTS: We developed Tribus, an interactive knowledge-based classifier for multiplexed images and proteomic datasets that avoids hard-set thresholds and manual labeling. We demonstrated that Tribus recovers fine-grained cell types, matching the gold standard annotations by human experts. Additionally, Tribus can target ambiguous populations and discover phenotypically distinct cell subtypes. Through benchmarking against three similar methods in four public datasets with ground truth labels, we show that Tribus outperforms other methods in accuracy and computational efficiency, reducing runtime by an order of magnitude. Finally, we demonstrate the performance of Tribus in rapid and precise cell phenotyping with two large in-house whole-slide imaging datasets. AVAILABILITY AND IMPLEMENTATION: Tribus is available at https://github.com/farkkilab/tribus as an open-source Python package.

Proteomics↗

DNA replication-timing analysis of human chromosome 22 at high resolution and different developmental states.

Duplication of the genome during the S phase of the cell cycle does not occur simultaneously; rather, different sequences are replicated at different times. The replication timing of specific sequences can change during development; however, the determinants of this dynamic process are poorly understood. To gain insights into the contribution of developmental state, genomic sequence, and transcriptional activity to replication timing, we investigated the timing of DNA replication at high resolution along an entire human chromosome (chromosome 22) in two different cell types. The pattern of replication timing was correlated with respect to annotated genes, gene expression, novel transcribed regions of unknown function, sequence composition, and cytological features. We observed that chromosome 22 contains regions of early- and late-replicating domains of 100 kb to 2 Mb, many (but not all) of which are associated with previously described chromosomal bands. In both cell types, expressed sequences are replicated earlier than nontranscribed regions. However, several highly transcribed regions replicate late. Overall, the DNA replication-timing profiles of the two different cell types are remarkably similar, with only nine regions of difference observed. In one case, this difference reflects the differential expression of an annotated gene that resides in this region. Novel transcribed regions with low coding potential exhibit a strong propensity for early DNA replication. Although the cellular function of such transcripts is poorly understood, our results suggest that their activity is linked to the replication-timing program.

Cell Differentiation↗

Expression analysis of pancreatic cancer cell lines reveals association of enhanced gene transcription and genomic amplifications at the 8q22.1 and 8q24.22 loci.

Despite tremendous effort and progress in the diagnostics of pancreatic cancer with respect to imaging techniques and molecular genetics, only very few patients can be cured by surgery leading to a 5-year survival rate of only 3%. Especially the lack of chemotherapeutical options in this entity requires a better understanding of the molecular mechanisms leading to pancreatic carcinoma growth and progression in order to develop novel treatment regimens. To identify signaling pathways that are critical for this tumor entity, we compared six well-established pancreatic cancer cell lines (Capan-1, Capan-2, HUP-T3, HUP-T4, KCL-MOH, PaTu-8903) with colon cancer cell lines and tumor cell lines of non-epithelial origin by expression profiling. For this purpose we employed Human Genome Focus Arrays representing about 8500 well annotated human genes. We identified 353 genes with significantly high expression in the group of pancreatic carcinomas. Based on Gene Ontology annotations these genes are especially involved in Rho protein signal transduction, proteasome activator activity, cell motility, apoptotic program, and cell-cell adhesion processes indicating these pathways to be interesting candidates for the design of targeted therapies. Most pancreatic carcinomas are characterized by mutations in the TP53 and the KRAS genes and the absence of microsatellite instability, which could also be confirmed for our panel of pancreatic carcinoma cell lines. Looking for individual differences within this group that may be responsible for more or less aggressive behavior, we identified genomic amplifications at the 8q22.1 and the 8q24.22 loci to be associated with enhanced gene transcription. Because we have previously shown that gains of genomic material from the long arm of chromosome 8 have an adverse effect on the outcome of pancreatic carcinoma patients, we conclude that functional analysis of amplified genes at 8q22 and/or 8q24 may lead to an improved understanding of pancreatic carcinoma progression.

Apoptosis↗

An enzyme that regulates ether lipid signaling pathways in cancer annotated by multidimensional profiling.

Hundreds, if not thousands, of uncharacterized enzymes currently populate the human proteome. Assembly of these proteins into the metabolic and signaling pathways that govern cell physiology and pathology constitutes a grand experimental challenge. Here, we address this problem by using a multidimensional profiling strategy that combines activity-based proteomics and metabolomics. This approach determined that KIAA1363, an uncharacterized enzyme highly elevated in aggressive cancer cells, serves as a central node in an ether lipid signaling network that bridges platelet-activating factor and lysophosphatidic acid. Biochemical studies confirmed that KIAA1363 regulates this pathway by hydrolyzing the metabolic intermediate 2-acetyl monoalkylglycerol. Inactivation of KIAA1363 disrupted ether lipid metabolism in cancer cells and impaired cell migration and tumor growth in vivo. The integrated molecular profiling method described herein should facilitate the functional annotation of metabolic enzymes in any living system.

Carbamates↗

The application of new software tools to quantitative protein profiling via isotope-coded affinity tag (ICAT) and tandem mass spectrometry: I. Statistically annotated datasets for peptide sequences and proteins identified via the application of ICAT and tandem mass spectrometry to proteins copurifying with T cell lipid rafts.

Lipid rafts were prepared according to standard protocols from Jurkat T cells stimulated via T cell receptor/CD28 cross-linking and from control (unstimulated) cells. Co-isolating proteins from the control and stimulated cell preparations were labeled with isotopically normal (d0) and heavy (d8) versions of the same isotope-coded affinity tag (ICAT) reagent, respectively. Samples were combined, proteolyzed, and resultant peptides fractionated via cation exchange chromatography. Cysteine-containing (ICAT-labeled) peptides were recovered via the biotin tag component of the ICAT reagents by avidin-affinity chromatography. On-line micro-capillary liquid chromatography tandem mass spectrometry was performed on both avidin-affinity (ICAT-labeled) and flow-through (unlabeled) fractions. Initial peptide sequence identification was by searching recorded tandem mass spectrometry spectra against a human sequence data base using SEQUEST software. New statistical data modeling algorithms were then applied to the SEQUEST search results. These allowed for discrimination between likely "correct" and "incorrect" peptide assignments, and from these the inferred proteins that they collectively represented, by calculating estimated probabilities that each peptide assignment and subsequent protein identification was a member of the "correct" population. For convenience, the resultant lists of peptide sequences assigned and the proteins to which they corresponded were filtered at an arbitrarily set cut-off of 0.5 (i.e. 50% likely to be "correct") and above and compiled into two separate datasets. In total, these data sets contained 7667 individual peptide identifications, which represented 2669 unique peptide sequences, corresponding to 685 proteins and related protein groups.

Amino Acid Sequence↗

Investigation on glycosylation patterns of proteins from human liver cancer cell lines based on the multiplexed proteomics technology.

Glycosylation, a very important post-translational modification of proteins, is increasingly coming into notice. However, large-scale, throughput investigations on glycosylated proteins are few. We applied a sensitive and fast fluorescence-based multiplexed proteomics (MP) technology which included two-dimensional gel electrophoresis (2-DE) followed by the fluorescence staining of glycoprotein and mass spectrometry identification for the purpose of constructing glycoprotein databases of the typical human hepatocellular carcinoma cell lines including Hep3B cell line without metastasis and MHCC97H with highly metastatic potential as well as the control non-tumor Chang liver cell. 74+/-2 (n=3), 78+/-3 (n=3) and 72+/-5 (n=3) glycoprotein spots were detected on 2-DE gels from Chang liver, Hep3B and MHCC97H cell sample using this MP technique, respectively. In all, 80 glycoproteins from three cell lines were successfully identified via peptide mass profiling using MALDI-TOF-MS/MS and the identified glycoproteins were annotated to our databases. In addition, we also found the glycosylation pattern differences among these three cell lines. The protein glycosylation alteration would be have great significance for the diagnosis of HCC and prediction of its metastasis. This study described the construction of glycosylation patterns of proteins and glycoproteome databases of human liver cells by the novel technological platform. The glycoproteome databases also provide essential basis for following study.

Animals↗

Focusing of gene expression as the basis of stem cell differentiation.

In a prior report (Stem Cells Dev 14(4):354-366, 2005), we employed two-dimensional gel electrophoresis followed by advanced proteomics and the Database for Annotation, Visualization and Integrated Discovery (DAVID) to compare the protein expression profiles of mesenchymal stem cells to that of fully differentiated osteoblasts. These data were reported to advance technical approaches to define the basis of differentiation, but also led us to suggest that osteogenic differentiation of stem cells may result from the focusing of gene expression in functional clusters (e.g., calcium-regulated signaling proteins or adherence proteins) rather than simply from the induced expression of new genes, as many have assumed. Here, we have employed these analytical techniques to compare protein expression by mesenchymal stem cells directly with that of cells derived from them after induced osteogenic differentiation. Our results support the concept of gene focusing as the basis of differentiation. Specifically, induced differentiation results in a decrease in the number of mesenchymal cell markers and calcium-mediated signaling molecules expressed by their differentiated progeny. This effect was seen in parallel to increased expression of specific extracellular matrix (ECM) molecules and their receptors. These results strongly imply that changes in the ECM have a direct impact on stem cell differentiation, and that osteogenic differentiation of stem cells directed by matrix clues results from focusing of the expression of genes involved in Ca2+-dependent signaling pathways.

Calcium Signaling↗

Differential expression of genes related to HFE and iron status in mouse duodenal epithelium.

Iron absorption, distribution, use, and storage are thought to be tightly regulated since altered iron stores may lead to cellular damage and disease. HFE, the hereditary hemochromatosis gene product, is expressed in the crypts of the duodenum, but the molecular mechanism by which it contributes to the inhibition of iron absorption is still unknown. In this study we aimed to identify transcriptional profiles in the duodenal epithelium of Hfe(-/-) mice. We used dedicated microarrays to compare gene expression among the duodenum of Hfe(-/-) mice, induced iron overload mice, and control mice. We found 151 differentially expressed genes and unknown sequences between Hfe(-/-) mice and normal littermates. Gene profiling revealed a gene subset more specific for Hfe inactivation. The functional annotation of upregulated genes highlighted that mucus production and cell maintenance may account for the influence of Hfe on epithelium integrity and luminal iron uptake.

Animals↗

Genome-Wide Characterization of &#x3b2;-Glucosidase (TaBGLU) Genes in Bread Wheat and Their Expression Under Drought, Cold, and Combined Stress.

Glycoside hydrolase 1 (GH1) &#x3b2;-glucosidases were known to activate hormone conjugates and defense metabolites, yet their genomic organization and stress-response dynamics in wheat remained incompletely defined. We therefore performed an integrated characterization of TaBGLUs spanning phylogeny, gene structure and conserved motifs, subcellular localization, promoter cis-elements, Gene Ontology enrichment, protein-protein interaction networks, and targeted expression profiling. Wheat TaBGLUs partitioned into well-supported clades that shared canonical GH1 catalytic residues and a largely conserved motif scaffold. Subcellular localization predictions indicated predominant nuclear and chloroplast targeting, with a smaller cohort directed to secretory or endomembrane compartments. Promoters were enriched for light-responsive, hormone-related (ABA, JA/SA, auxin, GA) and stress-associated (MYB/WRKY, heat, low temperature) cis-elements, and functional annotations were consistent with roles in carbohydrate and cell-wall metabolism, hormone homeostasis, and defense. Network analysis revealed a densely connected TaBGLU submodule embedded within broader carbohydrate and defense interaction networks, suggesting coordinated or cooperative functions. Expression profiling under cold, drought, and combined drought and cold demonstrated broad stress inducibility, with early activation detected by 6 h, cold-responsive maxima typically at 12 h, drought-responsive peaks predominating at 24 h, and combined stress eliciting both earlier and more sustained expression maxima between 12-24 h. Representative strongly responsive genes included TaBGLU20, TaBGLU44, TaBGLU6, and TaBGLU23, which showed pronounced late induction under combined stress, TaBGLU30, which exhibited an earlier combined-stress peak, and TaBGLU12, which displayed a marked late drought-specific response. Taken together, this integrated genomic, regulatory, and expression atlas refined the wheat BGLU repertoire relative to previous gene model inventories, highlighted candidate TaBGLUs with central network positions and strong stress inducibility, and provided concrete entry points for functional validation and breeding for improved stress resilience.

Triticum↗

A transcription factor regulatory atlas for activity inference and perturbation prediction.

Inferring transcription factor (TF) activity from transcriptomes and predicting transcriptome-wide responses to TF perturbations remain challenging, in part because available TF-mRNA resources often face a trade-off between precision and coverage and typically lack signed regulatory information. Here, we present TFActProfiler, a TF-mRNA resource and computational framework that learns signed, quantitative TF-mRNA regulatory coefficients by integrating heterogeneous prior evidence (ChIP-based, motif-based, and curated TF-mRNA annotations) with large-scale bulk and single-cell RNA-seq atlases. TFActProfiler contains 2&#x2009;606&#x2009;176 signed TF-mRNA interactions and improves TF activity inference in TF knockdown benchmarks relative to widely used regulon resources while retaining broad TF and target coverage. In addition, because the same learned regulatory coefficients can be used to model downstream transcriptional effects, TFActProfiler enables prediction of transcriptome-wide gene expression responses to TF knockdown without training on task-matched perturbation data. When perturbation datasets are available, TFActProfiler can be further refined to achieve performance comparable to state-of-the-art machine-learning baselines. By providing a direction-aware representation of TF-mRNA regulation for both activity inference and perturbation-response modeling, TFActProfiler supports systematic dissection of gene regulatory programs across diverse cellular contexts.

Transcription Factors↗