Search PubMedSearch

SEARCH · Search PubMed

Results for “machine learning regulatory prediction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Saturating the eQTL map in Drosophila: Genome-wide patterns of cis and trans regulation of transcriptional variation in outbred populations.

Most genetic polymorphisms associated with complex traits are found in non-coding regions of the genome. Characterizing their effect presents a formidable challenge, and expression quantitative trait locus (eQTLs) mapping has been a key approach to do so. As comprehensive eQTL maps are available only for a few species, here we developed the Drosophila outbred synthetic population (Dros-OSP) and used it to characterize the landscape of transcriptional regulation in Drosophila melanogaster. We collected head and body transcriptomes and genomes from 1,286 outbred flies and mapped local and distant eQTLs for 98% of the genes. We characterized the network organization of the transcriptome across tissues and described the properties of local and distal eQTLs in terms of genetic diversity, heritability, connectivity, and pleiotropy. These results provide new insights into the genetic basis of transcriptional regulation in the fruit fly and offer a new mapping resource that will expand the possibilities currently available for the Drosophila community.

Animals

Massively parallel approaches for characterizing noncoding functional variation in human evolution.

The genetic differences underlying unique phenotypes in humans compared to our closest primate relatives have long remained a mystery. Similarly, the genetic basis of adaptations between human groups during our expansion across the globe is poorly characterized. Uncovering the downstream phenotypic consequences of these genetic variants has been difficult, as a substantial portion lies in noncoding regions, such as cis-regulatory elements (CREs). Here, we review recent high-throughput approaches to measure the functions of CREs and the impact of variation within them. CRISPR screens can directly perturb CREs in the genome to understand downstream impacts on gene expression and phenotypes, while massively parallel reporter assays can decipher the regulatory impact of sequence variants. Machine learning has begun to be able to predict regulatory function from sequence alone, further scaling our ability to characterize genome function. Applying these tools across diverse phenotypes, model systems, and ancestries is beginning to revolutionize our understanding of noncoding variation underlying human evolution.

Humans

Engineering Bacillus Subtilis for Efficient Biosynthesis of Riboflavin: Current Knowledge and Future Perspectives.

Riboflavin is an essential water-soluble vitamin that serves as a precursor for the biosynthesis of the flavin cofactors FMN and FAD, which play pivotal roles in numerous redox and energy metabolism reactions. With the growing global demand for sustainable vitamin production, microbial fermentation has become an attractive alternative to chemical synthesis due to its environmental and economic advantages. Among microbial hosts, Bacillus subtilis has emerged as a leading cell factory for riboflavin production owing to its GRAS status, well-characterized genetics, and efficient protein secretion system. This review provides a comprehensive overview of recent advances in metabolic engineering strategies to enhance riboflavin biosynthesis in B. subtilis. Key topics include strengthening biosynthetic and precursor pathways, relieving feedback inhibition, balancing metabolic flux and cell growth, employing adaptive laboratory evolution, and utilizing omics-guided optimization and 13C metabolic flux analysis. Moreover, the integration of synthetic biology tools such as riboswitch engineering, regulatory element design, and high-throughput screening has significantly accelerated strain improvement. Despite remarkable progress, challenges remain in achieving precise regulatory control, optimizing multi-gene expression, and enhancing genome integration efficiency. Future research combining multi-omics data, synthetic regulatory design, and machine learning-driven predictive modeling is expected to further advance the development of intelligent B. subtilis cell factories. However, the practical implementation of these systems remains constrained by the metabolic burden of overproduction and the lack of universal regulatory models that can predict strain performance across varying industrial scales.

Bacillus subtilis

Conserved HSFA1-dependent chromatin dynamics drive heat stress responses in plants.

Eukaryotic organisms remodel chromatin landscapes to regulate gene expression in response to environmental stress. In plants, heat stress (HS) induces widespread chromatin changes, yet the role of heat shock transcription factors (HSFs) in chromatin remodeling and their evolutionary conservation remains unclear. Using Marchantia polymorpha Mphsf mutants and Arabidopsis thaliana Athsfa1s mutants, we identify HSFA1 as a key regulator of HS-induced cis-regulatory element (CRE) accessibility, a mechanism conserved across land plants, mice, and humans. Gene regulatory network modeling reveals parallel transcription factor subnetworks, with MpWRKY10 and MpABI5B acting as indirect and negative HS regulators. We further showed that ABA modulates gene expression in an HSFA1-dependent manner without inducing chromatin remodeling. Finally, we develop a machine learning framework integrating chromatin accessibility and CRE information to predict gene expression across species, revealing stress-responsive regulatory logic at the transcriptional level. These findings provide insights into how TFs coordinate chromatin architecture to drive stress adaptation.

Heat-Shock Response

Machine learning-based analysis of the impact of 5' untranslated region on protein expression.

The 5' untranslated region (5'UTR) plays a crucial regulatory role in messenger RNA (mRNA), with modified 5'UTRs extensively utilized in vaccine production, gene therapy, etc. Nevertheless, manually optimizing 5'UTRs may encounter difficulties in balancing the effects of various cis-elements. Consequently, multiple 5'UTR libraries have been created, and machine learning models have been employed to analyze and predict translation efficiency (TE) and protein expression, providing insights into critical regulatory features. On the one hand, these screening libraries, based on TE and mean ribosome load, struggle to accurately quantify protein expression; on the other hand, a precise method for quantifying 5'UTRs necessitates a significantly costlier library. To resolve this dilemma, we constructed a library utilizing firefly luciferase as the reporter to measure accurate protein expression. In addition, we optimized the library construction method by clustering mRNA sequences to reduce redundant data and minimize the size of the dataset. This dual strategy by increasing accuracy and reducing dataset size was found to be effective in predicting the 5'UTRs from the PC3 cell line.

5' Untranslated Regions

HINN: Hierarchical Input Neural Network identifies multi-omics biomarker for cognitive decline.

Understanding complex diseases requires models that can integrate diverse layers of biological data while yielding insights that are biologically interpretable. Although multi-omics integration with machine learning (ML) has advanced disease prediction and biomarker discovery, most existing approaches overlook the hierarchical and regulatory relationships that connect these molecular layers. Here, we present the Hierarchical Input Neural Network (HINN), a deep learning framework that incorporates known cross-omics relationships directly into its architecture, capturing the flow of information from genomics to epigenomics, transcriptomics, and downstream biological processes. By embedding these relationships, HINN improves both predictive performance and biological interpretability. We applied HINN to blood-derived multi-omics data from individuals with Alzheimer's disease or mild cognitive impairment to predict cognitive scores from standardized assessments. HINN outperformed both baseline and state-of-the-art models and pinpointed multi-omics biomarkers-including SNPs and promoter-region CpG sites in ATP6V1C1 and RCHY1 -that were significantly correlated with plasma p-Tau181 levels. These features map to biologically relevant processes with potential implications for cognitive decline. Our findings demonstrate how combining deep learning with biological knowledge can uncover interpretable, blood-based biomarkers for cognitive decline due to complex diseases such as Alzheimer's. All code and data are openly available at https://github.com/bozdaglab/HINN.

Alzheimer’s disease

Comprehensive analysis of diagnostic biomarkers related to histone acetylation in acute myocardial infarction.

BACKGROUND: Acute myocardial infarction (AMI) has become a serious disease that endangers human health, with high morbidity and mortality. Numerous studies have reported histone acetylation can result in the occurrence of cardiovascular diseases. This article aims to explore the potential biomarkers of histone acetylation regulatory genes (ARGs) in AMI patients. METHODS: Five AMI datasets were downloaded from the Gene Expression Omnibus (GEO) database. Next, ARG-related genes were gathered by gene set variation analysis (GSVA) and Spearman's correlation analysis. Subsequently, weighted gene co-expression network analysis (WGCNA) was performed to identify the module genes related to histone acetylation regulation. In the GSE60993 and GSE48060 datasets, the common differentially expressed genes (DEGs) between AMI and control samples were screened. Importantly, the intersecting genes were obtained by overlapping ARGs-related genes, common DEGs, and module genes. Then, the biomarkers in AMI were determined by machine learning, receiver operating characteristic (ROC) curves, and quantitative PCR (qPCR). In addition, immune analysis, drug prediction, molecular docking, and the lncRNA-miRNA-mRNA regulatory network targeting the biomarkers were analyzed, respectively. RESULTS: Here, a total of 18 intersecting genes were identified by overlapping 7,349 ARGs-related genes, 5,565 module genes, and 25 common DEGs. Further, five biomarkers (AQP9, HLA-DQA1, MCEMP1, NKG7, and S100A12) were obtained, and a nomogram was constructed and verified based on these biomarkers. Notably, the biomarkers were significantly associated with CD8 T cells and neutrophils. In addition, the drugs related to biomarkers were predicted, and ATOGEPANT with the molecular target (S100A12) had a high binding affinity (docking score = -10 kcal/mol). CONCLUSION: AQP9, HLA-DQA1, MCEMP1, NKG7, and S100A12 were identified as biomarkers related to ARGs in AMI, which provides a new perspective to study the relationship between ARGs and AMI.

Humans

ExoShorkie: predicting RNA-seq coverage of exogenous genomes in yeast by transfer learning.

MOTIVATION: Predicting the RNA-seq coverage of native and exogenous sequences is central to many molecular- and synthetic-biology applications. Substantial progress has been made in developing methods to predict the RNA-seq coverage of native genomic sequences, with the recently developed Shorkie achieving state-of-the-art performance in yeast. However, prediction performance of these methods over exogenous DNA is still unknown. Recent studies measured RNA-seq coverage of large exogenous genomes in yeast, providing a unique opportunity to train machine-learning models on a large exogenous sequence space and to improve both prediction performance and our understanding of regulatory mechanisms. RESULTS: We introduce ExoShorkie, a method we developed by extending Shorkie through transfer learning across multiple exogenous RNA-seq datasets. We demonstrate that ExoShorkie significantly improves prediction performance on held-out exogenous genomes and outperforms both a native-genome-trained Shorkie baseline and Yorzoi, the only competing method for predicting exogenous RNA-seq coverage in yeast, in cross-validation and in leave-one-genome-out evaluations. Furthermore, through interpretability analyses we reveal biologically meaningful regulatory motifs and distinct regulatory rules in exogenous genomes in yeast, providing new insights into transcriptional regulation. AVAILABILITY AND IMPLEMENTATION: ExoShorkie is available at https://github.com/OrensteinLab/ExoShorkie.

Genome, Fungal

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article

An encyclopedia of human enhancer-gene regulatory interactions.

Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1-6. Here we create and evaluate a resource of more than 92 million enhancer-gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element-gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study (GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer-promoter contacts, additional features that guide enhancer-promoter communication include promoter class and enhancer-enhancer synergy. These genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.

Humans

A transcription factor regulatory atlas for activity inference and perturbation prediction.

Inferring transcription factor (TF) activity from transcriptomes and predicting transcriptome-wide responses to TF perturbations remain challenging, in part because available TF-mRNA resources often face a trade-off between precision and coverage and typically lack signed regulatory information. Here, we present TFActProfiler, a TF-mRNA resource and computational framework that learns signed, quantitative TF-mRNA regulatory coefficients by integrating heterogeneous prior evidence (ChIP-based, motif-based, and curated TF-mRNA annotations) with large-scale bulk and single-cell RNA-seq atlases. TFActProfiler contains 2 606 176 signed TF-mRNA interactions and improves TF activity inference in TF knockdown benchmarks relative to widely used regulon resources while retaining broad TF and target coverage. In addition, because the same learned regulatory coefficients can be used to model downstream transcriptional effects, TFActProfiler enables prediction of transcriptome-wide gene expression responses to TF knockdown without training on task-matched perturbation data. When perturbation datasets are available, TFActProfiler can be further refined to achieve performance comparable to state-of-the-art machine-learning baselines. By providing a direction-aware representation of TF-mRNA regulation for both activity inference and perturbation-response modeling, TFActProfiler supports systematic dissection of gene regulatory programs across diverse cellular contexts.

Transcription Factors

3D epigenome of glial cell types in developing human cortex.

The human cortex is complex and heterogeneous, undergoing extensive expansion during development1,2. Our prior study of neurogenesis, including radial glia (RG), intermediate progenitor cells, excitatory neurons and interneurons demonstrated that chromatin looping underlies transcriptional regulation for lineage-specific genes, shedding light on how non-coding genetic variants contribute to neuropsychiatric disorders by means of cell-type-specific gene regulation3. RG have a crucial role in generating cellular diversity through both neurogenesis and gliogenesis and can be further classified into ventricular RG (vRG) and outer RG (oRG)4,5. Given their significance in cortical development, we conducted a comprehensive three-dimensional (3D) epigenomic analysis of four main glial populations, including vRG, oRG, oligodendrocyte precursor cells and microglia, from the mid-gestational human neocortex. By integrating gene expression, chromatin accessibility, DNA methylation and 3D chromatin interactions, we identified cell-type-specific candidate cis-regulatory elements (cCREs) and validated their regulatory function using transgenic mouse embryos. Using machine learning, we prioritized 112 schizophrenia risk variants within glia cCREs and further confirmed the predicted vRG enhancer disruption by the rs4449074 risk allele in vivo. Finally, oRG cCREs are enriched for human accelerated regions compared with other cCREs and a subset of human accelerated regions show activity differences from their chimpanzee orthologues that interact with genes involved in neuronal development. Our findings advance the understanding of human-specific gene regulation during corticogenesis.

Journal Article

Decoding gene regulation in plant genomes with artificial intelligence.

One of the central goals of plant functional genomics is to uncover regulatory mechanisms that shape agriculturally important traits to inform crop improvement. Recent advances in machine learning (ML) and artificial intelligence (AI), especially Large Language Models (LLMs), have greatly transformed our ability to derive regulatory information from complex genomics data. This review starts with a brief introduction of recent advances in AI and ML. We then present a plant-focused synthesis of emerging applications of AI- and LLM tools to: (i) predict epigenomic features, regulatory DNA elements, and gene expressions; (ii) infer gene regulatory network; and (iii) estimate post-transcriptional regulation.

Artificial intelligence

Active learning of enhancer and silencer regulatory grammar in photoreceptors.

Cis-regulatory elements (CREs) direct gene expression in health and disease, and models that can accurately predict their activities from DNA sequences are crucial for biomedicine. Deep learning represents one emerging strategy to model the regulatory grammar that relates CRE sequence to function. However, these models require training data on a scale that exceeds the number of CREs in the genome. We address this problem using active machine learning to iteratively train models on multiple rounds of synthetic DNA sequences assayed in live mammalian retinas. During each round of training the model actively selects sequence perturbations to assay, thereby efficiently generating informative training data. We iteratively trained a model that predicts the activities of sequences containing binding motifs for the photoreceptor transcription factor Cone-rod homeobox (CRX) using an order of magnitude less training data than current approaches. The model's internal confidence estimates of its predictions are reliable guides for designing sequences with high activity. The model correctly identified critical sequence differences between active and inactive sequences with nearly identical transcription factor binding sites, and revealed order and spacing preferences for combinations of motifs. Our results establish active learning as an effective method to train accurate deep learning models of cis-regulatory function after exhausting naturally occurring training examples in the genome.

Journal Article

Popcorn: prediction of short coding and noncoding genomic sequences in prokaryotes.

SUMMARY: The most challenging prokaryotic genes to identify often correspond to short ORFs (sORFs) encoding small proteins or to noncoding RNAs. RNA-seq experiments commonly evince small transcripts that do not correspond to annotated genes and are candidates for novel coding sORFs or small regulatory RNAs, but it can be difficult to accurately assess whether the numerous small transcripts are coding or not. We present Popcorn (PrOkaryotic Prediction of Coding OR Noncoding), a novel machine learning method for determining whether prokaryotic sequences are coding or noncoding. We find that Popcorn is effective in distinguishing coding from noncoding sequences, including coding sORFs and noncoding RNAs. AVAILABILITY AND IMPLEMENTATION: Freely available for use on the web at https://cs.wellesley.edu/∼btjaden/Popcorn. Source code available at https://github.com/btjaden/Popcorn and https://doi.org/10.5281/zenodo.15120075.

Open Reading Frames

Sequence optimization targeting mRNA stability enhances monoclonal antibody titers in CHO cells.

This study presents a DNA sequence optimization approach that integrates mRNA stability as a tunable design parameter to enhance monoclonal antibody expression in Chinese hamster ovary (CHO) cells. A comprehensive combinatorial library of synonymous coding-sequence variants of an IgG1 light chain was integrated as single copies at a defined genomic locus in CHO cells with identical regulatory elements. Steady-state mRNA abundance, quantified by deep sequencing of gDNA and mRNA, served as a proxy for mRNA stability. These data were used to train a machine learning model that predicts mRNA abundance from coding sequence using embeddings from a pre-trained nucleotide transformer. This abundance predictor, together with established translational metrics, was incorporated into a genetic algorithm for multi-objective codon optimization. As proof-of-concept, we optimized sequences encoding Trastuzumab to either maximize or minimize the abundance criterion and obtained benchmark sequences from two commercial providers. Using targeted integration, we generated CHO cell lines and measured protein titer and cell-specific productivity. Sequences optimized for high abundance significantly increased intracellular mRNA levels (+41%), protein titer (+59%), and cell-specific productivity (+85%) relative to low-abundance designs, while viable cell densities remained comparable. Compared to commercial benchmarks, high-abundance sequences achieved significantly higher titer (+70%) and cell-specific productivity (+98%). These findings establish mRNA stability as a practical and complementary design parameter for codon optimization in monoclonal antibody production, with potential applicability to other proteins and expression systems.

CHO

Transcriptomics-based exploration of ubiquitination-related biomarkers and potential molecular mechanisms in laryngeal squamous cell carcinoma.

BACKGROUND: One of the most common and prevalent cancers is laryngeal squamous cell carcinoma (LSCC), which poses a great threat to the life and health of the patient. Nonetheless, it has been demonstrated that ubiquitination is crucial for the development and course of LSCC. Therefore, it is particularly important to identify biomarkers for ubiquitination-related genes (UbRGs) in LSCC. METHODS: Differentially expressed genes (DEGs) in the LSCC versus controls were obtained by differential expression analysis. Also, key modular genes associated with LSCC were obtained using weighted gene co-expression network analysis (WGCNA). Next, DEGs, key module genes, and UbRGs were taken to intersect to obtain candidate genes. And then machine algorithms were to screen potential biomarkers, further their diagnostic value were analyzed and validated. Then, therapeutic agents for biomarkers were predict. In addition, the regulatory networks of the biomarkers were mapped. The expression levels of biomarkers were detected in clinical samples using reverse transcription-quantitative PCR (RT-qPCR). RESULTS: A total of eight candidate genes were acquired by the overlap 1,911 DEGs, the key modular genes of WGCNA, and 1,393 UbRGs. A sum of four biomarkers (WDR54, KAT2B, NBEAL2 and LNX1) were identified by two machine learning, then these four biomarkers were validated in GSE127165 and the expression trend was consistent with TCGA-LSCC, they were recorded as biomarkers. Moreover, the accuracy of the biomarkers in predicting clinical aspects of LSCC was confirmed by the receiver operating characteristic (ROC) curves. Subsequently, cancers such as malignant neoplasms, colorectal cancers, tumors, and primary malignant neoplasms were significantly associated with the biomarkers, which further suggests that these four biomarkers were strongly associated with cancer. Meanwhile, the drugs garcinol, cocaine, and triazolam, among others, used for LSCC treatment were predicted. Finally, transcription factors (TFs) (BRD4, MYC, AR, and CTCF) were predicted to regulate the biomarkers. RT-qPCR assays illustrated that the expression trends of KAT2B, LNX1 and NBEAL2 remained consistent with the dataset. CONCLUSION: The identification of four biomarkers (WDR54, KAT2B, NBEAL2 and LNX1) associated with UbRGs could ultimately serve as a predictive clinical diagnosis of LSCC and provide insight into the molecular mechanisms of LSCC.

Humans

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping