Search PubMedSearch

SEARCH · Search PubMed

Results for “complex network analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Acupoint Selection Patterns and Potential Mechanisms of Acupuncture in Knee Osteoarthritis: A Combined Data Mining and Network Pharmacology Study.

OBJECTIVE: To identify the core acupoint prescription and Kellgren-Lawrence (K-L) grade-dependent compatibility patterns of acupuncture for KOA through complex network analysis, and to predict the potential molecular mechanisms underlying the core prescription via network pharmacology. METHODS: Literature was retrieved from PubMed, EMbase, Cochrane Library, Web of Science, CNKI, Wanfang, VIP, and SinoMed (inception to September 3, 2025). Frequency, association rule, complex network, and K-L grade subgroup analyses were applied. Potential targets of the core prescription were identified via network pharmacology and intersected with disease targets from OMIM, Therapeutic Target, GeneCards, and DrugBank. A protein-protein interaction (PPI) network was constructed, and Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed to explore the potential molecular mechanisms. RESULTS: We included 522 studies, yielding 582 prescriptions involving 123 acupoints. The core prescription comprised 24 acupoints, including Dubi (ST35), Neixiyan (EX-LE4), Liangqiu (ST34), Xuehai (SP10), Zusanli (ST36), Yanglingquan (GB34), Yinlingquan (SP9), among others. K-L subgroup analysis revealed ST35, GB34, SP9, and SP10 as universal core acupoints. The mild-to-moderate subgroup mainly used local acupoints, while the moderate-to-severe subgroup centered on ST35, with increased distal acupoint usage and higher degree values. Network pharmacology analysis identified 77 overlapping targets. Core targets included tumor necrosis factor (TNF), interleukin 6 (IL6), interleukin 1 beta (IL1B), tumor protein p53 (TP53), matrix metallopeptidase 9 (MMP9), signal transducer and activator of transcription 3 (STAT3), transforming growth factor beta 1 (TGFB1), caspase 3 (CASP3), and B-cell lymphoma 2 (BCL2), which were enriched in inflammation and immunity, cartilage metabolism, and tissue repair pathways. CONCLUSION: The core acupoint prescription for KOA features local acupoints combined with distal ones, exhibiting distinct patterns across K-L grades. Our computational findings suggest that core acupoints may potentially delay knee joint degeneration by synergistically regulating inflammation, cartilage metabolism, apoptosis, and tissue repair, although these predictions require experimental validation. These findings provide preliminary evidence and a theoretical basis for standardized clinical point selection and further mechanistic research.

KOA

Amplicon and metagenomic sequencing reveal thifluzamide drive rhizosphere microbial structural shifts and functional adaption.

Thifluzamide (TF) is a widely used phenyl urea fungicide in rice production; however, its impacts on the structural composition and functional dynamics of the rhizosphere microbiome remain poorly understood. Here, we systematically investigated the effects of TF on the structure, interactions, and functional potential of the rice (Oryza sativa L.) rhizosphere microbiome using integrated amplicon sequencing and metagenomic approaches. TF application significantly altered both bacterial and fungal community composition, bacterial diversity was markedly reduced, whereas fungal diversity increased. With bacterial diversity markedly reduced while fungal diversity increased. Beta-diversity analyses revealed strong treatment-driven community separation, indicating pronounced TF-induced microbial restructuring. Co-occurrence network analysis demonstrated reduced complexity and connectivity in bacterial networks but increased negative co-occurrence patterns within fungal communities, suggesting contrasting stability responses between microbial kingdoms. Metagenomic profiling further revealed substantial functional shifts, including the differential enrichment of KEGG and COG pathways associated with xenobiotic metabolism. Notably, while total ARG abundance remained stable, TF exposure altered the resistome profile by selectively enriching specific classes of antibiotic resistance genes (ARGs), biocide resistance genes (BRGs), and mobile genetic elements (MGEs). Strong positive correlations between MGEs and ARGs highlighted an elevated potential for horizontal gene transfer. Metagenome-assembled genome (MAG) analysis identified specific TF-enriched bacterial taxa, including Methylophilus, Sulfurospirillum, and Azospirillum, which harbored genes involved in pesticide degradation and xenobiotic transformation. Collectively, these findings demonstrate that TF profoundly reshapes the rice rhizosphere microbiome by altering microbial diversity, interaction networks, resistance gene profiles, and functional capacities. This study provides genomic insights into fungicide-microbiome interactions, underscoring the potential ecological implications associated with TF application, while identifying candidate microbial taxa that may contribute to pesticide degradation and rhizosphere microecology resilience.

Rhizosphere

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

Gene behaviors-based network enrichment analysis and its application to reveal immune disease pathways enriched with COVID-19 severity-specific gene networks.

MOTIVATION: Gene network analysis is essential for understanding the complex mechanisms underlying diseases, which often involve disruptions in molecular networks rather than individual genes. Despite the availability of large-scale omics datasets and computational tools for gene network analysis, interpretation of the biological relevance of these extensive networks remains challenging. RESULTS: We propose a novel computational strategy, gene behaviors-based network enrichment analysis, which systematically identifies functional pathways enriched in phenotype-specific gene networks. Our novel method incorporates comprehensive network characteristics, i.e. gene expression levels, edge strengths, and structural patterns of edges, to rank genes based on activity and assess pathway enrichment, effectively identifying functional pathways enriched within these networks. Through simulation studies, our strategy demonstrated superior performance compared with that of existing methods in identifying enriched pathways. We applied this strategy to whole-blood RNA-seq data from 1102 COVID-19 samples provided by the Japan COVID-19 Task Force. The analysis revealed immune disease pathways enriched with COVID-19 severity-specific gene networks, including "Systemic lupus erythematosus" in asymptomatic and severe samples and "Inflammatory bowel disease," "Primary immunodeficiency," and "Rheumatoid arthritis" in mild samples. Key biomarkers of COVID-19, such as CXCL8, S100A9, and HLA class I genes, have been identified as critical hub genes and the main players within these networks. AVAILABILITY AND IMPLEMENTATION: Code is available in Figshare (https://doi.org/10.6084/m9.figshare.29093648.v3).

COVID-19

Genome-wide characterization of the sugar transporter protein family identifies candidate genes for bacterial wilt resistance breeding in tobacco.

Sugar transporter proteins (STPs) play pivotal roles in hexose allocation and plant stress responses. However, systematic characterization of the STP family in tobacco (Nicotiana tabacum) and its involvement in Ralstonia solanacearum resistance remains unclear. In this study, 37 NtSTP genes were identified and classified into six groups, with Group VI being the most conserved and Group V exhibiting dicot-specific expansion. Gene structure and conserved motif analyses revealed that most NtSTP members possess the typical MFS_STP domain, although variations in exon-intron organization and motif composition suggested functional divergence. Tandem duplication (TD) served as the primary driver of NtSTP family expansion, and Ka/Ks values of all paralogous pairs were less than 1, indicative of purifying selection. Promoter cis-element analysis revealed a complex regulatory network involving hormone signaling (ABA, JA, SA, GA, ET), stress responses, and light signaling. RT-qPCR expression profiling revealed that ten NtSTP genes (NtSTP1, 5, 7, 21, 22, 24, 26, 27, 28, and 29) exhibited significant transcriptional upregulation upon R. solanacearum infection. Specifically, NtSTP5, NtSTP7, NtSTP21, NtSTP22, NtSTP24, NtSTP26, and NtSTP27 peaked at 12 h post-inoculation (hpi), whereas NtSTP1, NtSTP28, and NtSTP29 reached their highest expression levels at 24 hpi. By contrast, NtSTP6, NtSTP13, and NtSTP30 displayed reduced expression upon R. solanacearum infection. These expression patterns indicate functional diversification within the NtSTP family and imply that these members may be transcriptionally modulated during plant responses to R. solanacearum. The present work provides preliminary and valuable candidate gene resources that may facilitate future disease resistance breeding programs in tobacco.

NtSTP gene family

Research on multi-trait genome association study method based on Shannon information entropy.

BACKGROUND: Genetic analysis of complex traits is crucial for elucidating disease mechanisms and biological inheritance processes. However, traditional Genome-wide Association Study (GWAS) for single trait often fail to capture the synergistic effects of genetic loci on multiple traits. METHODS: This study proposes a method for analyzing the association between multiple traits and gene regions based on Shannon information entropy. Innovatively, Shannon information entropy is introduced to integrate gene region information as genetic entropy, thereby constructing an Inverse Shannon Entropy-Multi-Trait Association Analysis of Gene Region genetic model (InvSE-MTAGR). Furthermore, a partial regression test is applied to the model to establish the Inverse Partial Shannon Entropy-Multi-Trait Association Analysis of Gene Region method (InvPSE-MTAGR). When performing multi-trait analysis with InvSE-MTAGR, the method achieved statistical significance by accumulating minor effects, thereby enhancing the ability to identify pleiotropic gene regions. RESULTS: The simulation results showed that the proposed multi-trait gene region association analysis method performed well in terms of both Type I error rate control and statistical power. Leveraging tomato and sorghum datasets for validation, the proposed multi-trait gene region association analysis method based on Shannon information entropy accurately pinpointed most of the gene regions harboring candidate genes. CONCLUSION: The study reveals the advantage of multi-trait method in integrating weak-effect pleiotropic signals and capturing the correlation among traits, which provides an efficient theoretical tool for dynamic analysis of complex multi-trait genetic networks and multi-target collaborative breeding of crops.

Genome-Wide Association Study

Comparative transcriptomic analysis of the gills and hepatopancreas of freshwater-cultured Litopenaeus vannamei under chronic nitrite stress.

To investigate the differences in molecular responses between the gills and hepatopancreas of freshwater-cultured Litopenaeus vannamei under chronic nitrite stress, a 30-day chronic stress experiment was conducted with a control group and a stress group. Transcriptomic analysis of the gills and hepatopancreas was performed using Illumina sequencing; differentially expressed genes (DEGs) were identified, and GO, KEGG, GSEA, PPI, and RT-qPCR validation were carried out. The results showed that 196 DEGs (161 up-regulated and 35 down-regulated) were identified in the gills, and 287 DEGs (199 up-regulated and 88 down-regulated) in the hepatopancreas, with only 18 DEGs shared between the two tissues. DEGs in the gills were enriched in oxidoreductase activity, glycerophospholipid metabolism, and tyrosine metabolism; DEGs in the hepatopancreas were enriched in lipid transporter activity, phagosome, ECM-receptor interaction, and riboflavin metabolism. GSEA revealed significant suppression of the mTOR pathway in the gills and the Polycomb complex pathway in the hepatopancreas. PPI network analysis identified hub genes P5CS and eEF2 in the gills, and PER, TUBB1, SHMT, and TUBB4B in the hepatopancreas. RT-qPCR validation was consistent with the RNA-seq results (R2 = 0.764). This study indicates that, under chronic nitrite stress, the gill response is centered on redox regulation and inhibition of growth metabolism, whereas the hepatopancreas response primarily involves lipid transport, cytoskeletal remodeling, and phagosome activation. The two tissues synergistically adapt through fundamental biosynthetic and motor protein pathways. This research provides molecular evidence for deciphering the nitrite tolerance mechanisms in freshwater-cultured shrimp.

Animals

Sparse spectral graph analysis and its application to gastric cancer drug resistance-specific molecular interplays identification.

Uncovering acquired drug resistance mechanisms has garnered considerable attention as drug resistance leads to treatment failure and death in patients with cancer. Although several bioinformatics studies developed various computational methodologies to uncover the drug resistance mechanisms in cancer chemotherapy, most studies were based on individual or differential gene expression analysis. However the single gene-based analysis is not enough, because perturbations in complex molecular networks are involved in anti-cancer drug resistance mechanisms. The main goal of this study is to reveal crucial molecular interplay that plays key roles in mechanism underlying acquired gastric cancer drug resistance. To uncover the mechanism and molecular characteristics of drug resistance, we propose a novel computational strategy that identified the differentially regulated gene networks. Our method measures dissimilarity of networks based on the eigenvalues of the Laplacian matrix. Especially, our strategy determined the networks' eigenstructure based on sparse eigen loadings, thus, the only crucial features to describe the graph structure are involved in the eigenanalysis without noise disturbance. We incorporated the network biology knowledge into eigenanalysis based on the network-constrained regularization. Therefore, we can achieve a biologically reliable interpretation of the differentially regulated gene network identification. Monte Carlo simulations show the outstanding performances of the proposed methodology for differentially regulated gene network identification. We applied our strategy to gastric cancer drug-resistant-specific molecular interplays and related markers. The identified drug resistance markers are verified through the literature. Our results suggest that the suppression and/or induction of COL4A1, PXDN and TGFBI and their molecular interplays enriched in the Extracellular-related pathways may provide crucial clues to enhance the chemosensitivity of gastric cancer. The developed strategy will be a useful tool to identify phenotype-specific molecular characteristics that can provide essential clues to uncover the complex cancer mechanism.

Stomach Neoplasms

Calcium sensors and their interacting protein kinases: genomics of the Arabidopsis and rice CBL-CIPK signaling networks.

Calcium signals mediate a multitude of plant responses to external stimuli and regulate a wide range of physiological processes. Calcium-binding proteins, like calcineurin B-like (CBL) proteins, represent important relays in plant calcium signaling. These proteins form a complex network with their target kinases being the CBL-interacting protein kinases (CIPKs). Here, we present a comparative genomics analysis of the full complement of CBLs and CIPKs in Arabidopsis and rice (Oryza sativa). We confirm the expression and transcript composition of the 10 CBLs and 25 CIPKs encoded in the Arabidopsis genome. Our identification of 10 CBLs and 30 CIPKs from rice indicates a similar complexity of this signaling network in both species. An analysis of the genomic evolution suggests that the extant number of gene family members largely results from segmental duplications. A phylogenetic comparison of protein sequences and intron positions indicates an early diversification of separate branches within both gene families. These branches may represent proteins with different functions. Protein interaction analyses and expression studies of closely related family members suggest that even recently duplicated representatives may fulfill different functions. This work provides a basis for a defined further functional dissection of this important plant-specific signaling system.

Amino Acid Sequence

A Comprehensive Analysis of Differential Protein Expression in the Plasma of Rheumatoid Arthritis Patients Utilizing Data-Independent Acquisition (DIA) Proteomics Technology.

BACKGROUND: Rheumatoid Arthritis (RA) is a Prevalent Autoimmune Disorder Affecting Millions of People Worldwide. A Thorough Understanding of Its Clinical and Pathological Features Is Essential to Improve Patient Outcomes. METHODS: This Study Combined Data-Independent Acquisition Proteomics and Enzyme-Linked Immunosorbent Assay (ELISA) to Identify and Validate Potential Plasma Protein Biomarkers for the Early Diagnosis of RA. RESULTS: Differential Proteomic Analysis Identified Differentially Expressed Proteins Between Patients With RA and Healthy Controls and Characterized Their Functions. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes Enrichment Analyses Were Performed to Explore Protein Functions and Associated Biological Pathways. The STRING Database and the Metascape Platform Were Used to Conduct an in-Depth Analysis of the Protein-Protein Interaction Network, Highlighting the Functional Attributes and Interconnections of Upregulated Proteins and Identifying Key Protein Complexes Involved in RA. ELISA Analysis of Plasma Samples Revealed Significantly Elevated SERPINA3 Levels in Patients With RA, Which Were Positively Correlated With Disease Activity Indicators-Including Erythrocyte Sedimentation Rate, C-Reactive Protein, and Disease Activity Score 28-But Were Not Correlated With Rheumatoid Factor or Its Subtypes. CONCLUSIONS: This Study Provides New Insights and Identifies Potential Biomarkers for the Early Diagnosis of RA.

Humans

transFusion: a novel comprehensive platform for integration analysis of single-cell and spatial transcriptomics.

MOTIVATION: Understanding spatial organization, intercellular interactions, and regulatory networks within the spatial context of tissues is crucial for uncovering complex biological processes and disease mechanisms. Spatial transcriptomics technologies have revolutionized this field by enabling the spatially resolved profiling of gene expression. 10× Visium has emerged as the predominant spatial technology, but its low resolution and the complexity of integrating multimodal datasets present significant analytical challenges, particularly for researchers with limited computational and statistical expertise. Current spatial transcriptomics analysis platforms generally fall short of effectively integrating multimodal data and maximizing the utility of spatial information-such as uncovering complex cellular spatial dependencies, multimodal gradient patterns, and spatial coexpression of ligand-receptor pairs and regulatory networks related to disease or biological states-thereby limiting their ability to provide comprehensive end-to-end analytical workflows when analyzing 10× Visium data. RESULTS: To address these limitations, we developed transFusion, a novel, advanced web-based platform specializing in the most comprehensive and effective integration analysis of scRNA-seq and 10× Visium spatial transcriptomics data. transFusion offers 12 key functions, from basic visualization to advanced analyses, including intercellular dependency analysis, ligand-receptor coexpression identification and visualization, and spatial multimodal gradient variation patterns. Two case studies were used to demonstrate transFusion's capabilities in exploring tissue architecture, intercellular communication, dependency networks, and multimodal gradient variation patterns with minimal computational skills and statistical expertise. transFusion provides a flexible and powerful framework for multimodal data integration analysis. AVAILABILITY AND IMPLEMENTATION: transFusion is freely available at https://github.com/WQLin8/transFusion.

Spatial Transcriptomics

Nested co-expression network analysis identifies compact gene clusters in a black box.

MOTIVATION: Digital analysis of biological systems requires methods capable of identifying both broad and nested gene modules reflecting complex biological processes. Existing transcriptomic methods often miss compact gene sets corresponding to subprocesses in specialized cell types, limiting insights into functional heterogeneity. RESULTS: We present Nested-WGCNA, a two-stage unsupervised network analysis algorithm designed to identify coarse-grained and fine-grained gene modules. Applied to bulk RNA-Seq data, Nested-WGCNA reveals stable modules reproducible across datasets. When validated against scRNA-Seq data, these modules correspond to both major and minor immune cell subtypes. Application to immunotherapy response datasets uncovers predictive and prognostic biomarkers, highlighting its utility in treatment stratification and biomarker discovery. AVAILABILITY: The NestedWGCNA source code and analysis pipeline are available on GitHub (https://github.com/ilyada/NestedWGCNA) and archived on Zenodo (https://doi.org/10.5281/zenodo.18959244).

Algorithms

Machine learning vs. traditional methods for predicting postoperative cardiac complications after non-cardiac surgery: a systematic review and Bayesian network meta-analysis.

INTRODUCTION: Accurate prediction of peri-operative cardiac complications is critical to optimise pre-operative decision-making. Traditional risk prediction scores, such as the Revised Cardiac Risk Index, show only modest discrimination. Machine learning can model complex, non-linear relationships but their predictive performance compared with traditional scores remains unclear. METHODS: We performed a systematic review and Bayesian network meta-analysis. The primary outcome was postoperative adverse cardiac events following non-cardiac surgery. Prediction models were assessed relative to the Revised Cardiac Risk Index. As many studies evaluated multiple versions of each model type, the highest performing ('best version') and lowest performing ('worst version') results were analysed. Models were ranked using the surface under the cumulative ranking curve (SUCRA). RESULTS: Thirteen studies evaluating 54 models and 927,113 patients were included. Machine learning approaches generally outperformed traditional risk scores. Automated machine learning ranked highest (SUCRA 96.6) showed the greatest improvement in the best version analysis (mean difference (MD) 0.28 (95%CrI 0.16-0.40)) and remained superior in the sensitivity analysis (MD 0.30 (95%CrI 0.14-0.45)). Gradient boosting models showed superior performance over the Revised Cardiac Risk Index across analysis (best version: MD 0.20 (95%CrI 0.14-0.26), worst version: MD 0.18 (95%CrI 0.12-0.25), SUCRA 82.4). The Gupta Perioperative Risk for Myocardial Infarction or Cardiac Arrest score outperformed the Revised Cardiac Risk Index in the best version analysis (MD 0.16 (95%CrI 0.01-0.32)). Between-study heterogeneity was low. None of the included studies externally validated their machine learning models and only six were judged to be at low risk of bias. DISCUSSION: Most machine learning models showed better discrimination than traditional risk scores, with automated machine learning and gradient boosting models ranking highest. However, study quality, calibration reporting and absence of external validation limit immediate clinical adoption. Prospective, multicentre evaluation is required before integration of these models into peri-operative practice.

Humans

Unravelling the transcriptomic characteristics of bronchoalveolar lavage in post-covid pulmonary fibrosis.

BACKGROUND: Post-Covid Pulmonary Fibrosis (PCPF) has emerged as a significant global issue associated with a poor quality of life and significant morbidity. Currently, our understanding of the molecular pathways of PCPF is limited. Hence, in this study, we performed whole transcriptome sequencing of the RNA isolated from the bronchoalveolar lavage (BAL) samples of PCPF and compared it with idiopathic pulmonary fibrosis (IPF) and non-ILD (Interstitial Lung Disease) control to understand the gene expression profile and associated pathways. METHODS: BAL samples from PCPF (n = 3), IPF (n = 3), and non-ILD Control (n = 3) (individuals with apparent healthy lung without interstitial lung disease) groups were obtained and RNA were isolated for whole transcriptomic sequencing. Differentially Expressed Genes (DEGs) were determined followed by functional enrichment analysis and qPCR validation. RESULTS: A panel of differentially expressed genes were identified in bronchoalveolar lavage fluid cells (BALF) of PCPF as compare to control and IPF. Our analysis revealed dysregulated pathways associated with cell cycle regulation, immune responses, and neuroinflammatory processes. Real-time validation further supported these findings. The PPI network and module analysis shed light on potential biomarkers and underscore the complex interplay of molecular mechanisms in PCPF. The comparison of PCPF and IPF identified a significant downregulation of pathways that were more prominent in IPF. CONCLUSION: This investigation provides crucial insights into the molecular mechanism of PCPF and also outlines avenues for prospective research and the development of therapeutic approaches.

Humans

Predicting host tropism in influenza a viruses: insights from multi-segment nucleotide signatures.

BACKGROUND: Influenza A virus (IAV) poses a significant public health threat due to its cross-species transmission and complex host adaptation mechanisms. This study integrated whole-genome data from avian, human, swine, and bovine IAV strains, using machine learning to predict viral host tropism based on nucleotide site features and to identify key sites driving host adaptation along with their synergistic effects. METHODS: A total of 64,000 IAV sequences from avian, human, swine, and bovine hosts were analyzed to build host-prediction models. A four-class classification framework (avian, human, swine, bovine) was constructed using nucleotide site features from all eight genomic segments (PB2, PB1, PA, HA, NP, NA, MP, NS). Eight machine learning algorithms (logistic regression, decision tree, random forest, SVM, KNN, gradient boosting, XGBoost, LightGBM) were benchmarked via 10-fold stratified cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, AUPRC, and AUC. SHAP (SHapley Additive exPlanations) analysis prioritized critical nucleotide sites, while bivariate association tests identified synergistic/antagonistic interactions between sites. Nucleotide composition profiles were compared across host groups using hierarchical clustering and heatmap visualization. RESULTS: The XGBoost algorithm demonstrated the best and most stable performance, achieving an AUC value of over 0.95 in distinguishing human-derived sequences from non-human ones. SHAP analysis identified the top 20 critical nucleotide sites for each gene segment, such as sites 46 and 698 in the NS segment. Nucleotide composition analysis revealed high similarity between human and swine sequences in the HA and PB2 segments, and between avian and bovine sequences. The HA segment was particularly challenging in differentiating human from swine strains. Bivariate site association analysis uncovered significant synergistic or antagonistic effects between key sites within gene segments, forming complex networks. For instance, in the NS segment, a positive prediction contribution was observed when sites 371, 698, and 419 were all G. CONCLUSIONS: This study advances our mechanistic understanding of IAV host adaptation, identifies molecular determinants for zoonotic risk stratification, and establishes a scalable machine learning framework for predicting viral host tropism through nucleotide signature analysis, thereby enhancing surveillance strategies and informing preventive measures against emerging viral threats.

Influenza A virus

GSK3B inhibition partially reverses brain ethanol-induced transcriptomic changes in C57BL/6J mice: Expression network co-analysis with human genome-wide association studies.

Alcohol use disorder (AUD) is a chronic behavioral disease with greater than 50% of its risk due to complex genetic contributions. Existing pharmacological and behavioral treatments for AUD are minimally effective and underutilized. Animal model behavioral genetics and human genome-wide association studies have begun to identify individual genes contributing to the progressive compulsive consumption of ethanol that occurs with AUD, promising possible new therapeutic targets. Our laboratory has previously identified Gsk3b as a central member in a network of ethanol-responsive genes in mouse prefrontal cortex, which altered ethanol consumption with genetic manipulation and was also significantly associated with risk for alcohol dependence in human genome-wide association studies. Here we perform detailed brain RNA sequencing transcriptomic studies to characterize a highly specific and clinically available GSK3B pharmacological inhibitor, tideglusib, as a possible therapeutic for clinical trials on treatment of AUD. A model of chronic intermittent ethanol consumption was used to study gene expression changes in prefrontal cortex and nucleus accumbens in the presence or absence of tideglusib treatment. Multivariate analysis of differentially expressed genes showed that tideglusib largely reversed ethanol- induced expression changes for two prominent clusters of genes in both prefrontal cortex and nucleus accumbens. Bioinformatic analysis showed these genes to have prominent roles in neuronal functioning and synaptic activity. Additionally, mouse brain differential gene expression data was analyzed together with human protein-protein interaction and genome-wide association studies on AUD to derive networks responding to tideglusib and relevant to human genetic risk for alcohol dependence. These studies identified discrete networks significantly enriched with genes provisionally associated with AUD, and provide key information on central hubs of such networks. Together these studies document tideglusib as a major modulator of chronic ethanol consumption-evoked brain gene expression signatures, and identify possible new targets for therapeutic modulation of AUD.

Journal Article

In silico screening of anti-atherosclerotic compounds from Morus alba leaves by machine learning and network pharmacology.

OBJECTIVE: This study integrates machine learning with network pharmacology, molecular docking, and molecular dynamics simulations to screen bioactive compounds from Mulberry leaves and elucidate their potential mechanisms against atherosclerosis (AS). METHODS: A training dataset of anti-AS active compounds was compiled and encoded as Morgan fingerprints. Three machine learning classifiers, specifically Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XG-Boost), were constructed and evaluated using multiple performance metrics. Potential active components from Mulberry leaves and AS-related targets were retrieved, followed by protein-protein interaction network construction and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Molecular docking was then performed to evaluate binding affinities between core targets and candidate compounds, and the most stable complex was subjected to molecular dynamics simulations using GROMACS (2025). RESULTS: The RF model achieved superior performance (accuracy= 0.8354, F1 = 0.8408, AUC = 0.9119) with 100% external validation accuracy. Thirteen anti-AS candidates were prioritized from mulberry leaves, four of which have been previously documented. Network pharmacology revealed AKT1 and IL6 as core targets, enriched in pathways such as endocrine resistance. Molecular docking and dynamics simulations confirmed strong binding between oxysanguinarine and AKT1, with the complex exhibiting high stability. CONCLUSION: The RF model provides a reliable computational tool for prioritizing anti-AS compounds from Mulberry leaves. The integrated analysis reveals that Mulberry leaves exert anti-atherosclerotic effects through multi-target (e.g., AKT1, IL6) and multi-pathway (e.g., PI3K-Akt) mechanisms, offering a framework for further experimental validation.

Morus

MiNEApy: enhancing enrichment network analysis in metabolic networks.

MOTIVATION: Modeling genome-scale metabolic networks (GEMs) helps understand metabolic fluxes in cells at a specific state under defined environmental conditions or perturbations. Elementary flux modes (EFMs) are powerful tools for simplifying complex metabolic networks into smaller, more manageable pathways. However, the enumeration of all EFMs, especially within GEMs, poses significant challenges due to computational complexity. Additionally, traditional EFM approaches often fail to capture essential aspects of metabolism, such as co-factor balancing and by-product generation. The previously developed Minimum Network Enrichment Analysis (MiNEA) method addresses these limitations by enumerating alternative minimal networks for given biomass building blocks and metabolic tasks. MiNEA facilitates a deeper understanding of metabolic task flexibility and context-specific metabolic routes by integrating condition-specific transcriptomics, proteomics, and metabolomics data. This approach offers significant improvements in the analysis of metabolic pathways, providing more comprehensive insights into cellular metabolism. RESULTS: Here, I present MiNEApy, a Python package reimplementation of MiNEA, which computes minimal networks and performs enrichment analysis. I demonstrate the application of MiNEApy on both a small-scale and a genome-scale model of the bacterium Escherichia coli, showcasing its ability to conduct minimal network enrichment analysis using minimal networks and context-specific data. AVAILABILITY AND IMPLEMENTATION: MiNEApy can be accessed at: https://github.com/vpandey-om/mineapy.

Metabolic Networks and Pathways