Search PubMedSearch

SEARCH · Search PubMed

Results for “Mutual information”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

14 recordsLinked to original sources

Mutual Information-based Prognostic Biomarker Discovery in Cancer Genomics: Conceptual Framework and Representative Applications of MI-POG.

Mutual information (MI)-based approaches have increasingly been applied to cancer genomics; however, their use for genome-wide prognostic biomarker discovery remains relatively underexplored. The present article summarizes the conceptual workflow of Mutual Information-based Prognostic Omics Gene (MI-POG) based on previously published applications in breast cancer, lower-grade glioma, and other cancer datasets. The framework consists of clinical endpoint discretization, genome-wide MI-based screening, candidate ranking, and downstream validation using conventional survival-analysis approaches. Previous MI-POG applications identified solute carrier family 20 member 1 (SLC20A1) as a prognostic biomarker in hormone receptor-positive breast cancer. Elevated SLC20A1 expression was associated with unfavorable survival outcomes and was independently validated in the Molecular Taxonomy of Breast Cancer International Consortium (METABRIC) cohort. Methodological analyses demonstrated how survival endpoints can be integrated into an information-theoretic framework through fixed-time outcome discretization, enabling model-independent assessment of molecular-clinical dependencies. Applications across multiple cancer datasets suggested the potential applicability of the framework across biologically distinct tumor types, although further validation will be required to establish its robustness and generalizability. In conclusion, MI-POG can be formalized as an information-theoretic framework for genome-wide identification of prognostic biomarkers by quantifying molecular-clinical dependencies using mutual information. Representative applications from previously published studies suggest that MI-POG may complement conventional survival-analysis approaches and provide a useful strategy for biomarker discovery, although additional benchmarking and prospective validation will be required.

Humans

Stress-induced altered expression of hippocampal nuclear and mitochondrial encoded genes in rats and cross-species genetic associations reveal molecular links to depression.

BACKGROUND: Mitochondria play a pivotal role in energy production, and their dysfunction not only hampers cells' ability to meet energy requirements but also contributes to the impairment of neural plasticity, a critical feature of depressive disorders. In this study, mitochondrial cross-omics analysis was carried out in the hippocampus of restraint rats to understand the role of mitochondria in depression pathophysiology. METHODS: The expression profiles of hippocampal mitochondrial and nuclear-encoded genes in mitochondrial fractions from restraint and handled control rats were obtained using high-throughput RNA sequencing. Weighted gene co-expression network analysis (WGCNA) was used to identify the gene co-expression and pathways associated with the restraint phenotype. Mutual Information Network algorithm tools Arance, CLR, and MRNET were additionally used to screen the functional modules and hub genes and their similarity with the WGCNA-based network analysis. Finally, cross-species homology followed by gene association analysis was conducted to obtain SNPs and haplotypes related to depression phenotype. RESULTS: A significant proportion of mitochondrial and nuclear-encoded genes showed differential regulation in the hippocampus of restraint rats. WGCNA and Mutual Information Network analysis yielded distinct functional modules significantly related to restraint phenotype. Further network analysis revealed distinct co-expression patterns associated with differentially expressed genes associated with these modules. Cross-species analysis showed 39 significantly associated SNPs with the depression phenotype, where the most significant SNP, rs10899570, was located within the TENM4 gene. Further, rs1573529 and rs10899570 were distributed into the linkage disequilibrium block where SNPs were highly correlated. Subsequent haplotype analysis showed that rs1573529 and rs10899570 were significantly associated with depressive behavior. CONCLUSIONS: The study demonstrates a significant impact of restraint stress on mitochondrial functions and genetic association, suggesting their critical role in depression pathophysiology.

Animals

A Graph Contrastive Learning Method for Enhancing Genome Recovery in Complex Microbial Communities.

Accurate genome binning is essential for resolving microbial community structure and functional potential from metagenomic data. However, existing approaches-primarily reliant on tetranucleotide frequency (TNF) and abundance profiles-often perform sub-optimally in the face of complex community compositions, low-abundance taxa, and long-read sequencing datasets. To address these limitations, we present MBGCCA, a novel metagenomic binning framework that synergistically integrates graph neural networks (GNNs), contrastive learning, and information-theoretic regularization to enhance binning accuracy, robustness, and biological coherence. MBGCCA operates in two stages: (1) multimodal information integration, where TNF and abundance profiles are fused via a deep neural network trained using a multi-view contrastive loss, and (2) self-supervised graph representation learning, which leverages assembly graph topology to refine contig embeddings. The contrastive learning objective follows the InfoMax principle by maximizing mutual information across augmented views and modalities, encouraging the model to extract globally consistent and high-information representations. By aligning perturbed graph views while preserving topological structure, MBGCCA effectively captures both global genomic characteristics and local contig relationships. Comprehensive evaluations using both synthetic and real-world datasets-including wastewater and soil microbiomes-demonstrate that MBGCCA consistently outperforms state-of-the-art binning methods, particularly in challenging scenarios marked by sparse data and high community complexity. These results highlight the value of entropy-aware, topology-preserving learning for advancing metagenomic genome reconstruction.

canonical correlation analysis

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping

scPlantLLM: A Foundation Model for Exploring Single-cell Expression Atlases in Plants.

Single-cell RNA sequencing (scRNA-seq) provides unprecedented insights into plant cellular diversity by enabling high-resolution analyses of gene expression at the single-cell level. However, the complexity of scRNA-seq data, including challenges in batch integration, cell type annotation, and gene regulatory network (GRN) inference, demands advanced computational approaches. To address these challenges, we developed scPlantLLM, a Transformer model trained on millions of plant single-cell data points. Using a sequential pretraining strategy incorporating masked language modeling and cell type annotation tasks, scPlantLLM generates robust and interpretable single-cell data embeddings. When applied to Arabidopsis thaliana datasets, scPlantLLM excels in clustering, cell type annotation, and batch integration, achieving an accuracy of up to 0.91 in zero-shot learning scenarios. Furthermore, the model demonstrates an ability to identify biologically meaningful GRNs and subtle cellular subtypes, showcasing its potential to advance plant biology research. Compared to traditional methods, scPlantLLM outperforms in key metrics such as adjusted rand index (ARI), normalized mutual information (NMI), and silhouette score (SIL), highlighting its superior clustering accuracy and biological relevance. scPlantLLM represents a foundation model for exploring plant single-cell expression atlases, offering unprecedented capabilities to resolve cellular heterogeneity and regulatory dynamics across diverse plant systems. The code used in this study is available at https://github.com/compbioNJU/scPlantLLM.

Single-Cell Analysis

CaXML: Chemistry-informed machine learning explains mutual changes between protein conformations and calcium ions in calcium-binding proteins using structural and topological features.

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of CaXML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

Machine Learning

40S Ribosomal protein S6 kinase integrates daylength perception and growth regulation in Arabidopsis thaliana.

Plant growth occurs via the interconnection of cell growth and proliferation in each organ following specific developmental and environmental cues. Therefore, different photoperiods result in distinct growth patterns due to the integration of light and circadian perception with specific Carbon (C) partitioning strategies. In addition, the TARGET OF RAPAMYCIN (TOR) kinase pathway is an ancestral signaling pathway that integrates nutrient information with translational control and growth regulation. Recent findings in Arabidopsis (Arabidopsis thaliana) have shown a mutual connection between the TOR pathway and the circadian clock. However, the mechanistical network underlying this interaction is mostly unknown. Here, we show that the conserved TOR target, the 40S ribosomal protein S6 kinase (S6K) is under circadian and photoperiod regulation both at the transcriptional and post-translational level. Total S6K (S6K1 and S6K2) and TOR-dependent phosphorylated-S6K protein levels were higher during the light period and decreased at dusk especially under short day conditions. Using chemical and genetic approaches, we found that the diel pattern of S6K accumulation results from 26S proteasome-dependent degradation and is altered in mutants lacking the circadian F-box protein ZEITLUPE (ZTL), further strengthening our hypothesis that S6K could incorporate metabolic signals via TOR, which are also under circadian regulation. Moreover, under short days when C/energy levels are limiting, changes in S6K1 protein levels affected starch, sucrose and glucose accumulation and consequently impacted root and rosette growth responses. In summary, we propose that S6K1 constitutes a missing molecular link where day-length perception, nutrient availability and TOR pathway activity converge to coordinate growth responses with environmental conditions.

Arabidopsis

Patient and Public Involvement and Engagement in Pediatric Health Research: A Systematic Review.

BACKGROUND: Patient and public involvement and engagement (PPIE) can increase the relevance and efficiency of research projects. An overview of PPIE approaches and implementation in pediatric research studies is needed to facilitate learning from others' experiences. OBJECTIVE: We aimed to systematically review practices in PPIE across all pediatric health research disciplines regarding characteristics and recruitment of PPIE participants, timepoints and methods used for PPIE, levels of involvement, benefits and barriers of PPIE. SEARCH STRATEGY: We searched Pubmed, EMBASE, Cochrane and PsycInfo using a comprehensive set of terms based on the concepts 'Patient and Public Involvement,' 'Health Research' and 'Pediatrics.' INCLUSION CRITERIA: We included original research articles describing PPIE implementation in pediatric health research published in English or German between 01/2003-10/2024. DATA EXTRACTION AND SYNTHESIS: Data was extracted using predefined categories and synthesized by narrative summary and thematic synthesis. PPIE reporting quality was assessed using the GRIPP2 short form checklist. MAIN RESULTS: Out of 1910 references, we included 37 original research articles, representing 35 studies. PPIE participants were mostly children, adolescents or caregivers involved in all research stages, especially in study design (89%) and recruitment (51%). Key positive impacts of PPIE on research included enhanced recruitment and retention rates and personal benefits for PPIE participants. Barriers to PPIE were financial and time resources required and challenges in recruiting representative PPIE participants. The level of involvement and PPIE reporting quality varied highly between studies. DISCUSSION: Common benefits and barriers of PPIE exist across pediatric research disciplines. Reporting quality varied highly between studies. CONCLUSIONS: PPIE is valuable in pediatric health research. Adherence to guidelines for conducting and reporting PPIE is important to enhance mutual learning. PATIENT OR PUBLIC CONTRIBUTION: PPIE input contributed to the understandability of the lay summary. The findings of this review, together with parent and public input, will inform guidelines for future PPIE activities at the authors' institutions.

Humans

Spatial mutual nearest neighbors for spatial transcriptomics data.

MOTIVATION: Mutual nearest neighbors (MNN) is a widely used computational tool to perform batch correction for single-cell RNA-sequencing data. However, in applications such as spatial transcriptomics, it fails to take into account the 2D spatial information. RESULTS: Here, we present spatialMNN, an algorithm that integrates multiple spatial transcriptomic samples and identifies spatial domains. Our approach begins by building a k-nearest neighbors (kNN) graph based on the spatial coordinates, prunes noisy edges, and identifies niches to act as anchor points for each sample. Next, we construct a MNN graph across the samples to identify similar niches. Finally, the spatialMNN graph can be partitioned using existing algorithms, such as the Louvain algorithm to predict spatial domains across the tissue samples. We demonstrate the performance of spatialMNN using large datasets, including one with N = 31 10x Genomics Visium samples. We also evaluate the computing performance of spatialMNN to other popular spatial clustering methods. AVAILABILITY AND IMPLEMENTATION: Our software package is available on GitHub (https://github.com/Pixel-Dream/spatialMNN). The code is available on Zenodo (https://doi.org/10.5281/zenodo.15073963).

Algorithms

Antibody diversification in cartilaginous fishes: Mechanistic insights from the nurse shark and comparative perspectives across jawed vertebrates.

Antibody diversity in vertebrates arises through the coordinated actions of V(D)J recombination and somatic hypermutation (SHM). Cartilaginous fishes occupy a key phylogenetic position as the sister lineage to bony vertebrates and therefore provide important comparative insights into the evolution of adaptive immunity. This review focuses on the nurse shark (Ginglymostoma cirratum) as a representative model for examining antibody-diversification mechanisms in cartilaginous fishes. Shark immunoglobulin genes exhibit a multicluster organization, while immunoglobulin new antigen receptor (IgNAR), a heavy-chain-only isotype, contains a single variable domain with an extended complementarity-determining region 3 (CDR3) that can be stabilized by non-canonical disulfide bonds. These structural features, together with intracluster multi-D V(D)J recombination and distinctive SHM characterized by single and tandem substitutions and insertions/deletions, contribute to antibody diversification in sharks. By comparing cartilaginous fishes, ray-finned fishes, and mammals, this review highlights lineage-specific combinations of immunoglobulin gene organization, recombination, mutational processing, and affinity maturation. Within the heuristic framework proposed here, shark and mammalian systems are described as emphasizing "breadth-first" repertoire generation and "precision-first" affinity optimization, respectively. These terms indicate relative mechanistic emphases rather than mutually exclusive categories or sequential evolutionary stages, while ray-finned fishes exhibit a distinct combination of genomic organization and mutational features. Investigating antibody diversification in cartilaginous fishes not only advances our understanding of vertebrate immune evolution but also provides structural and mechanistic insights that may inform the development of engineered antibodies based on the IgNAR scaffold.

Antibody diversity

SIVA: diagonal integration of spatial multi-omics data via spatially informed variational autoencoders and anchor guidance.

MOTIVATION: Understanding cellular states and regulatory programs requires integrative analysis of multiple omics layers. Although recent spatial sequencing technologies allow molecular profiling of cells within their tissue context, paired spatial multi-omics assays are still limited by technical complexity and cost. This creates a pressing need for diagonal integration methods that enable joint analysis of unpaired spatial omics datasets. RESULTS: We propose SIVA, a deep generative framework based on Spatially-Informed Variational Autoencoders with Anchor Guidance, for diagonal integration of spatial multi-modal data. SIVA employs modality-specific variational autoencoders (VAEs) with a hybrid latent embedding that integrates Gaussian process and standard Gaussian priors, enabling joint modeling of spatially structured variation and dominant underlying data distributions across modalities. To facilitate cross-modal alignment in the absence of one-to-one cell correspondence, SIVA adopts a dual integration strategy combining global distribution alignment via Maximum Mean Discrepancy and local correspondence guidance using mutual nearest neighbor anchors. Extensive experiments across multiple cross-slice integration scenarios demonstrate that SIVA achieves robust and accurate integration of unpaired spatial omics datasets, consistently outperforming existing methods. AVAILABILITY AND IMPLEMENTATION: The source codes are available at https://github.com/PelenJiang/SIVA.

Autoencoder

Comprehensive Somatic Profiling of Gastroenteropancreatic Neuroendocrine Neoplasms.

BACKGROUND: The incidence of gastroenteropancreatic neuroendocrine neoplasms (GEP-NENs) is rising, yet their biological heterogeneity and variable response to treatments remain poorly understood. Comprehensive genomic characterization may uncover somatic drivers and inform biomarker-driven therapeutic strategies. METHODS: We retrospectively analyzed clinically ordered next-generation sequencing (NGS) results from tumor samples of 111 patients with confirmed GEP-NENs treated at Johns Hopkins Hospital between 2020 and 2022. Pathogenic and likely pathogenic mutations were identified using OncoKB, CHASMplus, and COSMIC databases. Mutational patterns were correlated with clinical characteristics and overall survival using univariate and multivariate analyses. RESULTS: In this retrospective study of 111 patients with gastroenteropancreatic neuroendocrine neoplasms (GEP-NENs), somatic pathogenic or likely pathogenic mutations were identified in 79% of cases. The most frequent alterations involved TP53 (19%), MEN1 (17%), and chromatin remodeling genes such as DAXX (9%) and ATRX (6%). Notably, we also identified a subset of patients (9%) patients with mutations typically associated with hematologic malignancies. Distinct co-mutation and mutual exclusivity patterns were observed between pancreatic and non-pancreatic NENs. Poorly differentiated or high-grade tumors correlated with mutations in TP53, KRAS, and CDKN2A. Mutations in KRAS, DAXX/ATRX, and hematologic malignancy-associated genes were independently associated with worse overall survival. CONCLUSIONS: This study reveals distinct somatic mutation patterns in GEP-NENs associated with tumor differentiation, grade, primary site, and survival. The identification of hematologic malignancy-associated mutations in a subset of GEP-NENs suggests possible shared molecular phenotypes with poor prognostic implications. The presence of KRAS mutations supports exploring pan-RAS inhibitors as potential therapies in select patients. These findings highlight the clinical utility of genomic profiling in GEP-NENs.

Neuroendocrine neoplasms

Genomic Profiling of Epidermal Growth Factor Receptor Mutation-Positive Non-Small Cell Lung Cancer after Progression on First-line Osimertinib: Phase II ORCHARD Study.

PURPOSE: Osimertinib is the standard of care for first-line treatment for epidermal growth factor receptor-mutated (EGFRm) non-small cell lung cancer (NSCLC). Understanding the tumor molecular profile of patients following progression on osimertinib could help inform optimal second-line treatment. PATIENTS AND METHODS: ORCHARD (NCT03944772), a phase II biomarker-directed study, enrolled patients with EGFRm NSCLC who progressed on first-line osimertinib to receive treatment based on their tumor molecular profile after progression. The study comprised three groups into which patients were allocated based on the molecular profile of their tumor, determined via next-generation sequencing (NGS) of a tumor biopsy. We report results from a prespecified, exploratory analysis of baseline tumor tissue and plasma samples to evaluate mechanisms of resistance to first-line osimertinib identified by tissue and plasma NGS. Agreement between tissue and plasma NGS data was also assessed. RESULTS: This study provided a comprehensive dataset exploring tissue (n = 400) and plasma (n = 191) genomics, enabling characterization of the histogenomic landscape after first-line osimertinib treatment. TP53 and MDM2/4 alterations were mutually exclusive and occurred in 86% of tumors. When combining tissue and plasma genomics, resistance alterations were detected in 87% of samples, with multiple resistance alterations in 46%. Alterations in the PI3K pathway, SOX2, and MYC were frequently detected in histologically transformed tumors. Additionally, differential patterns of co-occurring EGFR mutations in tumors with L858R versus exon 19 deletion were observed. CONCLUSIONS: This comprehensive analysis highlights potential heterogeneous resistance to first-line osimertinib treatment, providing a rationale for combining treatments with broad activity to improve patient outcomes. See related commentary by Gupta et al., p. 3718.

Humans