Search PubMedSearch

SEARCH · Search PubMed

Results for “Low-dimensional representation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

12 recordsLinked to original sources

Brain-wide spontaneous neural avalanches: Definition, functional dynamics and cognitive relevance.

Although spontaneous activity is ubiquitous across multiple spatiotemporal scales, its functional organization and cognitive relevance remain poorly understood. Following the classic neuronal avalanche framework, a spontaneous avalanche is defined as consecutively active frames separated by inactive time bins. Hence, multiple distinct avalanches may be considered as one avalanche, thereby ignoring their spatial and temporal distinguishability. Furthermore, group-level power-law fitting of such neural avalanches is often performed to evaluate brain criticality (referring to a system perched between order and disorder) due to the limited recording length of macroscale neuroimaging (such as functional magnetic resonance imaging), and the functional representation of brain-wide neural avalanches is largely unexplored. To address these issues, we proposed large-scale neural avalanches as a single, spatially consecutive cascade pattern and further investigated their functional dynamics, network propagation, and association with task-evoked activity. Compared with the conventional inactive-bin definition, our current approach is more favorable to power-law fitting of avalanche size and duration distributions at the individual level. We also demonstrated that participants whose brain activities were close to the critical point tend to have higher cognitive abilities. Notably, the ratio of neural avalanches that evolved from primary sensory to association networks negatively correlated with cognitive abilities. Moreover, the geometric distance between low-dimensional representations of task-evoked activity and spontaneous avalanches was associated with behavioral performance. This study not only provides a promising avenue for measuring avalanche criticality based on human whole-brain neuroimaging, but also suggests that spontaneous neural avalanches and their low-dimensional representations contribute to human cognitive abilities.

Humans

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

Integration of single cell multiomics data by deep transfer hypergraph neural network.

Multi-omics characterization of individual cells offers remarkable potential for analyzing the dynamics and relationships of gene regulatory states across millions of cells. How to integrate multimodal data is an open problem, existing integration methods struggle with accuracy and modality-specific biological variation retention. In this paper, we present scHyper (scalable, interpretable machine learning for single cell integration), a low-code and data-efficient deep transfer model designed for integrating paired and unpaired single-cell multimodal data. We benchmark scHyper against datasets from different multimodal data. ScHyper learns a low-dimensional representation and aligns the covariance matrices of the measured modalities, achieving high accuracy even with large scale atlas-level datasets with low memory and computational time across different cell lines, shedding light on regulatory relationships between different types of omics. Altogether, we show that scHyper is a versatile and robust tool for cell-type label transfer and integration from multimodal single-cell datasets.

Single-Cell Analysis

jsPCA: fast, scalable, and interpretable identification of spatial domains and variable genes across multi-slice and multi-sample spatial transcriptomics data.

MOTIVATION: Spatial transcriptomics technologies record genome-wide measurements of gene expression with high spatial resolution. These technologies generate large and high-dimensional datasets requiring efficient automated methods for their analysis. We introduce joint spatial PCA (jsPCA), a novel, fast, scalable and interpretable method for the automatic identification of spatial domains and variable genes in multi-slice and multi-sample spatial transcriptomics data. RESULTS: jsPCA relies on a simple mathematical formulation of a spatial covariance defined as the product of the gene expression covariance with the spatial autocorrelation. The principal components of this spatial covariance yield a biologically meaningful low-dimensional representation. From this representation, spatial domains are derived by simple clustering and spatially variable genes are identified directly from the principal component coefficients. A joint representation of multiple slices and samples without spatial alignment is obtained by computing common principal components via joint diagonalization. By leveraging data sparsity and non-convex manifold optimization, jsPCA leads to computing time in the order of seconds to minutes, substantially outperforming state-of-the-art approaches. We benchmarked jsPCA against 10 state-of-the-art methods on two reference databases. Our approach demonstrated excellent performance, comparable or better than state-of-the-art methods, while being much faster, interpretable, and scalable to very large datasets.

Journal Article

X-intNMF: a cross- and intra-omics regularized NMF framework for multi-omics integration.

MOTIVATION: The rapid accumulation of multi-omics data presents a valuable opportunity to advance our understanding of complex diseases and biological systems, driving the development of integrative computational methods. However, the complexity of biological processes, spanning multiple molecular layers and involving intricate regulatory interactions, requires models that can capture both intra- and cross-omics relationships. Most existing integration methods primarily focus on sample-level similarities or intra-omics feature interactions, often neglecting the interactions across different omics layers. This limitation can result in the loss of critical biological information and suboptimal performance. To address this gap, we propose X-intNMF, a network-regularized non-negative matrix factorization (NMF) framework that simultaneously integrates intra- and cross-omics feature interactions into a shared low-dimensional representation (see Fig. 1). By modeling these multi-layered relationships, X-intNMF enhances the representation of biological interactions and improves integration quality and prediction accuracy. RESULTS: For evaluation, we applied X-intNMF to predict breast cancer phenotypes and classify clinical outcomes in lung and ovarian cancers using mRNA expression, microRNA expression, and DNA methylation data from TCGA. The results show that X-intNMF consistently outperforms state-of-the-art methods. Ablation studies confirm that incorporating both cross-omics and intra-omics interactions contributes significantly to the model's improved performance. Additionally, survival analysis on 25 TCGA cancer datasets demonstrates that the integrated multi-omics representation offers strong prognostic value for both overall survival and disease-free status. These findings highlight X-intNMF's ability to effectively model multi-layered molecular interactions while maintaining interpretability, robustness, and scalability within the NMF framework. AVAILABILITY AND IMPLEMENTATION: The source code and datasets supporting this study are publicly available at GitHub (https://github.com/compbiolabucf/X-intNMF) and archived on Zenodo (https://doi.org/10.5281/zenodo.18238385).

Multiomics

IBAS: Interaction-bridged association studies discovering novel genes underlying complex traits.

Genetic contributions to complex traits are often mediated through coordinated gene-gene interaction networks, yet most existing association frameworks focus on marginal single-gene effects and overlook higher-order dependency structures. Direct modeling of interactions remains challenging due to combinatorial complexity and statistical instability. We introduce Interaction-Bridged Association Study (IBAS), a general framework that incorporates pathway-level interaction patterns into genotype-phenotype association analysis without explicitly enumerating interactions. IBAS leverages transcriptomic reference data to construct low-dimensional representations of pathway activity, which guide SNP-weighting and gene-level association testing within a kernel-based framework. In perturbation-based simulations, IBAS demonstrates improved stability and reproducibility compared to conventional TWAS and gene-based methods, while maintaining well-calibrated Type I error under phenotype permutation. Application to the WTCCC datasets identifies both known and novel genes across multiple complex diseases, including candidates with modest marginal effects missed by standard approaches. These findings are supported by replication in an independent cohort, and analyses across multiple reference tissues revealing both shared and tissue-specific signals. Overall, IBAS provides a statistically robust and computationally tractable framework for incorporating interaction effects into association mapping, extending beyond the single-gene paradigm and enabling more comprehensive characterization of complex trait. IBAS is available on GitHub at: https://github.com/QingrunZhangLab/IBAS.

Polymorphism, Single Nucleotide

Dissecting spatial patterning and signaling with directional diffusion in spatial multi-omics.

Spatial multi-omics sequencing enables the simultaneous profiling of transcriptomics, proteomics, and epigenomics at a spatial resolution, offering insights into complex tissue organization and molecular regulation. However, the effective integration of multiple omics modalities in a spatial context remains a major challenge. Here, we present SpaDDM, a spatial multi-omics integration framework based on directional diffusion models (DDMs), which supports spatial pattern identification, cross-omics alignment, and inter-and intracellular signaling flow analysis. SpaDDM employs DDM-based graph networks to learn omics-specific representations by jointly incorporating spatial coordinates and molecular measurements within each modality, followed by an attention mechanism to align features across modalities. We benchmarked SpaDDM on diverse spatial multi-omics datasets, including transcriptomics-epigenomics and transcriptomics-proteomics combinations across multiple tissues and species. SpaDDM consistently outperformed existing methods by more accurately deciphering spatial tissue patterns and effectively reducing the boundary noise between spatial regions. Moreover, the learned low-dimensional coembedded representations of individual cells serve as integral mediators for inferring the signaling flows that underlie spatial patterning. Finally, we demonstrated that SpaDDM alignment of complementary information across multi-omics layers facilitates cross-omics translation and significantly improves the prediction of cell state alignments.

Multiomics

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans

scPOEM: robust co-embedding of peaks and genes revealing peak-gene regulation.

MOTIVATION: Identifying regulatory elements in various chromosomal regions that influence gene expression is a fundamental challenge in epigenomics, with profound implications for understanding gene regulation and disease mechanisms. The advent of paired single-cell RNA sequencing and single-cell ATAC sequencing has created unprecedented opportunities to address this challenge by enabling simultaneous profiling of gene expression and chromatin accessibility at single-cell resolution. However, the inherent signals between them are weak due to the highly sparse and noisy nature of data. RESULTS: This article proposes single-cell meta-Path based Omics Embedding (scPOEM), a novel embedding method that jointly projects chromatin accessibility peaks and expressed genes into a shared low-dimensional space. By integrating the relationships among peak-peak, peak-gene, and gene-gene interactions, scPOEM assigns closer representations in the embedding space to related peak-gene pairs. Our experiments demonstrate that scPOEM generates stable representations of peaks and genes, outperforms existing methods in recovering biologically meaningful peak-gene regulatory relationships and enables new insights in subgroup and differential analysis of gene regulation. These results highlight its potential to uncover gene regulatory mechanisms and enhance the understanding of transcriptional regulation at single-cell resolution. AVAILABILITY AND IMPLEMENTATION: The source code of scPOEM is available at https://github.com/Houyt23/scPOEM. The datasets can be obtained from the 10× Genomics (https://www.10xgenomics.com/datasets/pbmc-from-a-healthy-donor-granulocytes-removed-through-cell-sorting-10-k-1-standard-1-0-0) and GEO database under access codes GSE194122 and GSE239916.

Gene Expression Regulation

One chromatin, many structures: From ensemble contact maps to single-cell 3D organization.

Understanding how chromatin folds in three dimensions remains challenging because most experimental assays capture low-dimensional projections of an underlying, highly heterogeneous polymer. Here, we present an ensemble-based interpretive framework built on the previously introduced Self-Returning Excluded Volume (SR-EV) model, a minimal generator of chromatin conformations using a nucleosome-indexed coarse-grained representation based on stochastic return rules and excluded-volume geometry. Despite its simplicity, SR-EV recapitulates key experimental signatures across scales: heterogeneous nanoscale packing domains resembling ChromEMT and ChromSTEM observations, sparse and highly variable single-configuration contact patterns analogous to single-cell chromosome conformation capture (Hi-C), and robust ensemble-level contact enrichment consistent with topologically associating domains (TADs). In this framework, Hi-C loop and TAD signatures are interpreted as ensemble-level statistical enrichments rather than invariant features of single-cell conformations. SR-EV is explicitly designed to generate large ensembles of complete three-dimensional chromatin configurations that can be projected consistently onto two-dimensional contact maps and one-dimensional genomic profiles. By introducing architectural-protein effects only through ensemble selection rather than explicit forces, SR-EV supports a separation between intrinsic polymer geometry and regulatory bias and suggests that TAD-like features can emerge as statistical enrichments rather than deterministic three-dimensional structures. Coordination number and probe-based accessibility computed directly from SR-EV provide a unified link between three-dimensional packing, two-dimensional contact maps, and one-dimensional genomic profiles. The main contribution of this work is to show, within a single coarse-grained framework, how these multimodal observables arise as linked projections of the same heterogeneous chromatin ensemble through averaging and conditional sampling. Together, these results establish SR-EV as a minimal and geometrically grounded mesoscale reference framework for interpreting how heterogeneous chromatin ensembles give rise to multimodal experimental observables while remaining consistent with the fact that chromatin organization is realized in individual cells.

Chromatin

Systematic background selection with BasCoD enhances contrastive dimension reduction in single cell genomics.

In single-cell experiments spanning diverse conditions, distinguishing variation specific to one condition (e.g., treatment) from shared or background variation (e.g., control) is critical for uncovering treatment-specific molecular responses. However, these studies typically yield ultra-high-dimensional data, necessitating effective dimension reduction for reliable biological interpretation. Contrastive dimension reduction methods address this challenge by identifying low-dimensional features enriched in a target dataset relative to a background dataset that captures shared variation. Despite their growing utility, the success of such methods critically depends on the choice of background, yet no formal criterion exists for evaluating or selecting backgrounds. To address this gap, we introduce BasCoD, a statistical testing framework based on spectral subspace inclusion theory, that enables rigorous evaluation and systematic selection of background datasets. Applying BasCoD across a range of single-cell datasets, we show that it effectively identifies suitable backgrounds, substantially improving the contrast and interpretability of the resulting target representations. We further demonstrate how BasCoD can guide the design of contrastive analyses in large-scale single-cell experiments conducted under heterogeneous conditions and elucidate potential interaction effects in perturbation studies.

Single-Cell Analysis

Systematic evaluation of one-dimensional-to-two-dimensional near-infrared spectroscopy transformations with deep learning for quantifying coconut sap adulteration.

Near-infrared (NIR) spectroscopy have limitations when combined with deep learning (DL) algorithms because they rely on low-dimensional datasets. Therefore, we investigated the potential of transforming one-dimensional (1D) NIR spectra into two-dimensional (2D) spectrograms using synchronous and asynchronous techniques and the continuous wavelet transform (CWT) and their effectiveness by integrating with DL for detecting adulteration in coconut sap. NIR spectra (12,500-4000 cm-1) were collected from binary mixtures (0%-100%;w/w). The performance of all DL (convolutional neural networks-CNN, AlexNet and ResNet) models was compared with that of partial least squares (PLS). The models were ranked in the mentioned order based on their performances: 2D-CWT > 2D-asynchronous > 2D-synchronous > 1D/2D-PLS. The important features of the best model can be explained and visualized using gradient-weighted-class-activation-mapping. The findings highlight that the 1D-to-2D NIR data transformation combined with DL is a highly robust approach because it addresses the feature representation gap in NIR data and effectively captures the spatial-spectral correlations.

Spectroscopy, Near-Infrared