Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Network graphs”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Decoding TnsC Filament Assembly in CRISPR-Associated Transposons Using Interpretable Deep Learning and Molecular Simulations.

CRISPR-associated transposons (CASTs) enable programmable DNA integration, yet how the TnsC regulator forms processive filaments on DNA to coordinate RNA-guided transposition in type V-K CAST systems remains unknown. Here, we integrate large-scale molecular simulations, interpretable deep learning using graph attention networks (GATs), and causal inference analyses to define the molecular determinants of TnsC filament nucleation and elongation. We show that TnsC nucleates by inducing localized DNA deformation that propagates along extended filaments, with Granger causality revealing that TnsC motions precede and predict DNA deformation. Interpretable GAT models demonstrate that elongation is determined during early recognition between incoming and DNA-bound subunits, followed by structural reorganization that regenerates the recruitment interface and enables processive assembly. These results elucidate the molecular mechanism of processive TnsC filament assembly and explain why isolated TnsC filaments preferentially elongate in the 5' → 3' direction, while accessory transposition factors can reshape the interaction landscape and alter filament growth polarity. Together, these findings advance our understanding of CAST function and inform the engineering of programmable DNA integration platforms. Beyond CAST systems, this work introduces an interpretable GAT approach as a general and transferable deep learning strategy for uncovering molecular mechanisms in biological systems, while demonstrating the power of causal inference for dissecting directional relationships in molecular dynamics.

Deep Learning↗

soFusion: facilitating tissue structure identification via spatial multi-omics data fusion.

The rapid advancement of spatial multi-omics technologies has opened new avenues for dissecting tissue architecture with unprecedented resolution. However, inherent disparities across omics modalities, such as differences in biological hierarchy and resolution, pose significant challenges for integrative analysis. To address this, we present soFusion, a method for representation learning on spatial multi-omics data that enables automated identification of tissue compartmentalization. soFusion employs a graph convolutional network (GCN) to extract latent embeddings from spatial omics profiles. To simultaneously capture both cross-modality relationships and modality-specific features, we introduce a novel strategy for intra- and inter-omics feature learning. Moreover, modality-specific decoders are designed to preserve the unique information embedded in each omics type. We evaluated soFusion on multiple datasets including gene expression, protein expression, and epigenetic features. Across all benchmarks, soFusion consistently outperformed existing methods in delineating anatomical structures and identifying spatial domains with improved continuity and reduced noise. Collectively, soFusion offers an effective solution for spatial multi-omics integration, substantially enhancing the robustness of spatial domain identification.

Humans↗

Dual-contrastive learning for spatial domain identification in spatial transcriptomics with STAMGC.

Spatial transcriptomics (STs) have become a valuable approach for understanding the growth and development of organisms. Despite the recent emergence of numerous ST models, accurately identifying spatial domains remains challenging owing to the trade-off between preserving local details and reducing noise. Here, we introduce STAMGC, which is a dual-contrastive learning framework built upon graph convolutional networks. This model leverages regional and topological contrastive learning to jointly optimize the model, effectively reducing the noise in spatial domain identification and enhancing the extraction of detailed features. In this study, Gaussian smoothing, originally developed in the image processing field, is introduced to process ST data, providing a foundation for region-level contrastive learning by mitigating spatial discontinuities of gene expression signals. Experimental results indicate that STAMGC outperforms existing methods across multiple data sets according to comprehensive evaluations. Furthermore, STAMGC not only identifies finer structures in the mouse brain but also brings new discoveries for human breast cancer research.

Journal Article↗

Localized states on comb lattices.

Complex networks and graphs provide a general description of a great variety of inhomogeneous discrete systems. These range from polymers and biomolecules to complex quantum devices, such as arrays of Josephson junctions, microbridges, and quantum wires. We introduce a technique, based on the analysis of the motion of a random walker, that allows us to determine the density of states of a general local Hamiltonian on a graph, when the potential differs from zero on a finite number of sites. We study in detail the case of the comb lattice and we derive an analytic expression for the elements of the resolvent operator of the Hamiltonian, giving its complete spectrum.

Journal Article↗

Mosaic genomes of the six major primate lentivirus lineages revealed by phylogenetic analyses.

To clarify the origin and evolution of the primate lentiviruses (PLVs), which include human immunodeficiency virus types 1 and 2 as well as their simian relatives, simian immunodeficiency viruses (SIVs), isolated from several host species, we investigated the phylogenetic relationships among the six supposedly nonrecombinant PLV lineages for which the full genome sequences are available. Employing bootscanning as an exploratory tool, we located several regions in the PLV genome that seem to have uncertain or conflicting phylogenetic histories. Phylogeny reconstruction based on distance and maximum-likelihood algorithms followed by a number of statistical tests confirms the existence of at least five putative recombinant fragments in the PLV genome with different clustering patterns. Split decomposition analysis also shows that phylogenetic relationships among PLVs may be better represented by network-based graphs, such as the ones produced by SplitsTree. Our findings not only imply that the six so-called pure PLV lineages have in fact mosaic genomes but also make more unlikely the hypothesis of cospeciation of SIVs and their simian hosts.

Algorithms↗

Chromatin structures from integrated AI and polymer physics model.

The physical organization of the genome in three-dimensional space regulates many biological processes, including gene expression and cell differentiation. Three-dimensional characterization of genome structure is critical to understanding these biological processes. Direct experimental measurements of genome structure are challenging; computational models of chromatin structure are therefore necessary. We develop an approach that combines a particle-based chromatin polymer model, molecular simulation, and machine learning to efficiently and accurately estimate chromatin structure from indirect measures of genome structure. More specifically, we introduce a new approach where the interaction parameters of the polymer model are extracted from experimental Hi-C data using a graph neural network (GNN). We train the GNN on simulated data from the underlying polymer model, avoiding the need for large quantities of experimental data. The resulting approach accurately estimates chromatin structures across all chromosomes and across several experimental cell lines despite being trained almost exclusively on simulated data. The proposed approach can be viewed as a general framework for combining physical modeling with machine learning, and it could be extended to integrate additional biological data modalities. Ultimately, we achieve accurate and high-throughput estimations of chromatin structure from Hi-C data, which will be necessary as experimental methodologies, such as single-cell Hi-C, improve.

Chromatin↗

Artificial Intelligence for Natural Products Discovery and Development.

Natural products (NPs) remain a cornerstone of modern drug discovery, offering stereochemical complexity and diverse bioactivities that precisely modulate therapeutic targets, refined through billions of years of evolution. However, their research has long been hindered by inefficient, empirical workflows, high resource consumption, structural complexity, and the "multicomponent, multi-target" nature of their mechanisms. The exponential growth of genomic, metabolomic, and spectral data has overwhelmed conventional analytical methods, exposing critical bottlenecks in handling high-dimensional, heterogeneous datasets that exceed human interpretive capacity. Artificial intelligence (AI) is emerging as a transformative paradigm to address these challenges, integrating multi-omics and chemical data to shift NP research from fragmented empiricism toward mechanism-driven, precision-oriented development. By leveraging deep learning architectures- including graph neural networks, Transformers, and diffusion-based generative models-AI enables systematic decoding of NP biosynthesis, automated structure elucidation, rational target identification, knowledge extraction from vast unstructured scientific literature, and de novo molecular design. This review comprehensively surveys recent advances in AI applications across the full NP discovery and development pipeline, encompassing genome mining, structure-based and ligand-based virtual screening, multimodal structural characterization, lead optimization, and biosynthetic pathway engineering. We further examine the emerging roles of protein-centric, molecule- centric, and multimodal foundation models, as well as large language models, in bridging genotype-to-chemotype gaps and unlocking unstructured scientific knowledge. Finally, we discuss critical challenges including data scarcity, representational limitations for complex stereochemistry, physical plausibility in generative models, and the urgent need for experimental validation, while outlining future directions toward autonomous experimentation, closed-loop optimization, and human-AI collaborative discovery.

Artificial intelligence↗

FREX: a query interface for biological processes with hierarchical and recursive structures.

An intelligent system for signal transduction pathways and other higher order functional knowledge is presented. Molecular mechanisms of biological processes are typically represented as diagrams ("pathways") that have a graph-analogical network structure. However, due to the diversity of topics that pathways cover, their constituent biological entities are highly diverse and range from metal ion to protein to biological processes in general. In addition, the kinds of interactions that connect biological entities are likewise diverse. Consequently, current knowledge about pathways is highly heterogeneous both in the sense of the types of constituents and the granularity of descriptions. To cope with this problem, the proposed system adopts a recursive and hierarchical representation model that enables the annotation and query of pathways or sub-pathways of arbitral granularity. By combining the use of this hierarchical structure and biological ontologies, literature-based information regarding biological mechanisms becomes accessible by computer.

Computational Biology↗

Software systems as complex networks: structure, function, and evolvability of software collaboration graphs.

Software systems emerge from mere keystrokes to form intricate functional networks connecting many collaborating modules, objects, classes, methods, and subroutines. Building on recent advances in the study of complex networks, I have examined software collaboration graphs contained within several open-source software systems, and have found them to reveal scale-free, small-world networks similar to those identified in other technological, sociological, and biological systems. I present several measures of these network topologies, and discuss their relationship to software engineering practices. I also present a simple model of software system evolution based on refactoring processes which captures some of the salient features of the observed systems. Some implications of object-oriented design for questions about network robustness, evolvability, degeneracy, and organization are discussed in the wake of these findings.

Journal Article↗

Network thermodynamics: analysis and synthesis of membrane transport system.

The bond graph expression of network thermodynamics is a useful tool for modeling of biological systems [1, 8, 9, 13]. Epithelial transport systems, such as frog skin and the acinar cell of the salivary gland, have been modeled and simulated using this tool [8, 9]. However, thermodynamic description of a complex system has been considered difficult [1], and doubt has also been cast on the applicability of network thermodynamics to thermal or entropy relations [2]. In this review, I discussed the thermodynamic basis for network thermodynamics using the availability function and assuming local equilibrium [7]. I also showed the applicability of bond graphs to the analysis of nonisothermal systems which include thermal conduction, thermoosmotic and thermoelectric phenomena [13]. Thus, network thermodynamics is suitable for analysis of more complicated systems and for computer simulation.

Animals↗

THEMPO: a knowledge-based system for therapy planning in pediatric oncology.

This article describes the knowledge-based system THEMPO (Therapy Management in Pediatric Oncology), which supports protocol-directed therapy planning and configuration in pediatric oncology. THEMPO provides a semantic network controlled by graph grammars to cover the different types of knowledge relevant in the domain, and offers a suite of acquisition tools for knowledge base authoring. Medical problem solvers, operating on the oncological network, reason about adequate therapeutic and diagnostic timetables for a patient. Furthermore, a corresponding patient record, also based on semantic networks and graph grammars, has been implemented to represent the course of therapy of an oncological patient.

Antineoplastic Combined Chemotherapy Protocols↗

An algorithm for constructing local regions in a phylogenetic network.

The groupings of taxa in a phylogenetic tree cannot represent all the conflicting signals that usually occur among site patterns in aligned homologous genetic sequences. Hence a tree-building program must compromise by reporting a subset of the patterns, using some discriminatory criterion. Thus, in the worst case, out of possibly a large number of equally good trees, only an arbitrarily chosen tree might be reported by the tree-building program as "The Tree." This tree might then be used as a basis for phylogenetic conclusions. One strategy to represent conflicting patterns in the data is to construct a network. The Buneman graph is a theoretically very attractive example of such a network. In particular, a characterization for when this network will be a tree is known. Also the Buneman graph contains each of the most parsimonious trees indicated by the data. In this paper we describe a new method for constructing the Buneman graph that can be used for a generalization of Hadamard conjugation to networks. This new method differs from previous methods by allowing us to focus on local regions of the graph without having to first construct the full graph. The construction is illustrated by an example.

Algorithms↗

The Whitney reduction network: a method for computing autoassociative graphs.

This article introduces a new architecture and associated algorithms ideal for implementing the dimensionality reduction of an m-dimensional manifold initially residing in an n-dimensional Euclidean space where n >> m. Motivated by Whitney's embedding theorem, the network is capable of training the identity mapping employing the idea of the graph of a function. In theory, a reduction to a dimension d that retains the differential structure of the original data may be achieved for some d < or = 2m + 1. To implement this network, we propose the idea of a good-projection, which enhances the generalization capabilities of the network, and an adaptive secant basis algorithm to achieve it. The effect of noise on this procedure is also considered. The approach is illustrated with several examples.

Journal Article↗

Systematic identification of statistically significant network measures.

We present a graph embedding space (i.e., a set of measures on graphs) for performing statistical analyses of networks. Key improvements over existing approaches include discovery of "motif hubs" (multiple overlapping significant subgraphs), computational efficiency relative to subgraph census, and flexibility (the method is easily generalizable to weighted and signed graphs). The embedding space is based on scalars, functionals of the adjacency matrix representing the network. Scalars are global, involving all nodes; although they can be related to subgraph enumeration, there is not a one-to-one mapping between scalars and subgraphs. Improvements in network randomization and significance testing--we learn the distribution rather than assuming Gaussianity--are also presented. The resulting algorithm establishes a systematic approach to the identification of the most significant scalars and suggests machine-learning techniques for network classification.

Algorithms↗

Visualization and characterization of non-covalent networks in molecular crystals: automated assignment of graph-set descriptors for asymmetric molecules.

A method of visualizing intermolecular networks (for example, hydrogen-bonded networks) in the crystalline state has been developed, based on the concept of link atoms, i.e. those atoms deemed to be in contact with each unique molecule or ion in the crystal chemical unit (CCU). Extension of a structure using each of these primary links can be achieved, enabling the generation and investigation of extended networks. Algorithms have been developed for the automatic assignment of graph-set notation for patterns up to second level, i.e. those involving one or two crystallographically independent non-covalent bonds, in the absence of internal crystallographic symmetry in the unique molecules of the CCU. The self, ring, chain and discrete motifs may be displayed by highlighting the atoms and bonds comprising the pattern. These methodologies have been implemented in the Cambridge Structural Database program PLUTO.

Journal Article↗

Neural network simulation of visual inspection of graphs of single-subject interventions.

Judgments of the effectiveness of single-subject behavioral interventions are often based on visual examination of graphs of response data, but previous research indicates that such judgments are often unreliable and flawed. Here it is proposed that artificial neural networks could be developed to simulate the judgments of expert judges. A prototype of such a network was designed and trained in the present study, and its use in novel experiments matched the estimates of the expert whose judgments were simulated significantly better than did a prediction equation developed using a multiple regression approach.

Artificial Intelligence↗

DBRF-MEGN method: an algorithm for deducing minimum equivalent gene networks from large-scale gene expression profiles of gene deletion mutants.

MOTIVATION: Large-scale gene expression profiles measured in gene deletion mutants are invaluable sources for identifying gene regulatory networks. Signed directed graph (SDG) is the most common representation of gene networks in genetics and cell biology. However, no practical procedure that deduces SDGs consistent with such profiles has been developed. RESULTS: We developed the DBRF-MEGN (difference-based regulation finding-minimum equivalent gene network) method in which an algorithm deduces the most parsimonious SDGs consistent with expression profiles of gene deletion mutants. Positive (or negative) directed edges representing positive (or negative) gene regulations are deduced by comparing the gene expression level between the wild-type and mutant. The most parsimonious SDGs are deduced using graph theoretical procedures. Compensation for excess removal of edges by restoring a minimum number of edges makes the method applicable to cyclic gene networks. Use of independent groups of edges greatly reduces the computational cost, thus making the method applicable to large-scale expression profiles. We confirmed the applicability of our method by applying it to the gene expression profiles of 265 Saccharomyces cerevisiae deletion mutants, and we confirmed our method's validity by comparing the pheromone response pathway, general amino acid control system, and copper and iron homeostasis system deduced by our method with those reported in the literature. Interpretation of the gene network deduced from the S. cerevisiae expression profiles by using our method led to the prediction of 132 transcriptional targets and modulators of transcriptional activity of 18 transcriptional regulators. AVAILABILITY: The software is available on request.

Algorithms↗

Co-clustering of biological networks and gene expression data.

MOTIVATION: Large scale gene expression data are often analysed by clustering genes based on gene expression data alone, though a priori knowledge in the form of biological networks is available. The use of this additional information promises to improve exploratory analysis considerably. RESULTS: We propose constructing a distance function which combines information from expression data and biological networks. Based on this function, we compute a joint clustering of genes and vertices of the network. This general approach is elaborated for metabolic networks. We define a graph distance function on such networks and combine it with a correlation-based distance function for gene expression measurements. A hierarchical clustering and an associated statistical measure is computed to arrive at a reasonable number of clusters. Our method is validated using expression data of the yeast diauxic shift. The resulting clusters are easily interpretable in terms of the biochemical network and the gene expression data and suggest that our method is able to automatically identify processes that are relevant under the measured conditions.

Algorithms↗