Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Network graphs”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Degree distributions of growing networks.

The in-degree and out-degree distributions of a growing network model are determined. The in-degree is the number of incoming links to a given node (and vice versa for out-degree). The network is built by (i) creation of new nodes which each immediately attach to a preexisting node, and (ii) creation of new links between preexisting nodes. This process naturally generates correlated in-degree and out-degree distributions. When the node and link creation rates are linear functions of node degree, these distributions exhibit distinct power-law forms. By tuning the parameters in these rates to reasonable values, exponents which agree with those of the web graph are obtained.

Journal Article↗

Disagreement-informed arbitration for gene regulatory network inference: A score-level meta-classifier and a diagnostic typology of inter-method conflict.

Gene regulatory network inference methods routinely disagree about individual edges, and practitioners resolve those conflicts by choosing one method or averaging them all. We ask whether the conflict can instead be arbitrated per edge. A gradient-boosted classifier is trained on the raw scores that ten inference methods-correlation-based, information-theoretic, sparse-regression and tree-ensemble, including GENIE3, GRNBoost2, CLR and ARACNe-assign to each candidate regulator-target pair, so that the weight given to each method varies from edge to edge. Across six single-cell perturbation screens spanning four cell types, arbitration improves on mean ensembling by +0.056 AUROC on Adamson and +0.083 on Shifrut under target-grouped cross-validation. The evaluation protocol turns out to matter more than the model. Edge-level cross-validation, standard in this literature, inflates apparent gains by 0.060 AUROC through target-gene leakage-comparable to the entire honest improvement. The effect is far larger for methods that represent genes implicitly: a supervised graph-attention link predictor trained on identical folds scores AUROC 0.930 under edge-level cross-validation, better than anything else we evaluate, and 0.533 once target genes are held out. Any method that parameterises genes is exposed, which covers most graph- and embedding-based approaches. A five-category typology of inter-method conflict localises where arbitration pays off, with the largest gains on edges where the methods disagree and the smallest where they already agree, while adding nothing as model input; we therefore report it as a diagnostic instrument rather than a modelling contribution. We also characterise what the ground truth measures: most perturbed genes in widely used screens are not transcription factors, and a mediation screen bounds how much of the perturbation response can be direct.

Ensemble methods↗

Comparative evaluation of reverse engineering gene regulatory networks with relevance networks, graphical gaussian models and bayesian networks.

MOTIVATION: An important problem in systems biology is the inference of biochemical pathways and regulatory networks from postgenomic data. Various reverse engineering methods have been proposed in the literature, and it is important to understand their relative merits and shortcomings. In the present paper, we compare the accuracy of reconstructing gene regulatory networks with three different modelling and inference paradigms: (1) Relevance networks (RNs): pairwise association scores independent of the remaining network; (2) graphical Gaussian models (GGMs): undirected graphical models with constraint-based inference, and (3) Bayesian networks (BNs): directed graphical models with score-based inference. The evaluation is carried out on the Raf pathway, a cellular signalling network describing the interaction of 11 phosphorylated proteins and phospholipids in human immune system cells. We use both laboratory data from cytometry experiments as well as data simulated from the gold-standard network. We also compare passive observations with active interventions. RESULTS: On Gaussian observational data, BNs and GGMs were found to outperform RNs. The difference in performance was not significant for the non-linear simulated data and the cytoflow data, though. Also, we did not observe a significant difference between BNs and GGMs on observational data in general. However, for interventional data, BNs outperform GGMs and RNs, especially when taking the edge directions rather than just the skeletons of the graphs into account. This suggests that the higher computational costs of inference with BNs over GGMs and RNs are not justified when using only passive observations, but that active interventions in the form of gene knockouts and over-expressions are required to exploit the full potential of BNs. AVAILABILITY: Data, software and supplementary material are available from http://www.bioss.sari.ac.uk/staff/adriano/research.html

Algorithms↗

Graph theoretical characterization and tracking of the effective neural connectivity during episodes of mesial temporal epileptic seizure.

Via a detailed case study of mesial temporal lobe epilepsy, we show that a method of determining the direction of information flow among signals is able to provide focal localization via the simultaneous analysis of multiple EEG channels. This determination is accomplished by representing information flow direction via directed graphs, where focal electrodes are associated with high observed rates of pertinence to strongly connected subgraphs. Further clinical support to this finding is provided by results for an additional 9 cases of focal epilepsy cases. The graph theoretical approach is a tool for describing and analyzing the effective connectivity dynamics behind epileptic seizures and may provide a common language for studying other complex dynamic relationships between neural structures.

Electroencephalography↗

Pathway logic modeling of protein functional domains in signal transduction.

Protein functional domains (PFDs) are consensus sequences within signaling molecules that recognize and assemble other signaling components into complexes. Here we describe the application of an approach called Pathway Logic to the symbolic modeling signal transduction networks at the level of PFDs. These models are developed using Maude, a symbolic language founded on rewriting logic. Models can be queried (analyzed) using the execution, search and model-checking tools of Maude. We show how signal transduction processes can be modeled using Maude at very different levels of abstraction involving either an overall state of a protein or its PFDs and their interactions. The key insight for the latter is our algebraic representation of binding interactions as a graph.

Computational Biology↗

A graph theoretical approach for analysis of protein flexibility change at protein complex formation.

Hitherto analyses of protein complexes are frequently confined to the changes in the interface of the protein subunits undergoing interaction, while the holistic picture of the protein monomers' structure transformation, or the pervasive rigidity adopted by the newly formed complex are most often than not improperly evaluated in spite of the multiple and deep insights that they can yield about the interaction process itself at the molecular level, or at the higher level of genomic functional analyses for which relevant systems biological information can be obtained. To address this aspect of protein-protein interaction we propose in this work a newly developed algorithm that is based on graph theoretical instances and makes possible the evaluation of the changes in the flexibility of the interacting molecules and the rigidity adopted at complex formation. Since one can also figure out the opposite process, i.e. that in which the complex decomposes into its constituent subunits, each of which may accomplish another vital role in the organism, the methodology proposed here is also able to address such problem. The algorithm we propose performs a rigidity and/or flexibility evaluation of every node (atom) on the network constituted by the entire set of intra and inter-molecular inter-atomic interactions. Comparison of flexible or rigid molecular regions or domains within the complex with those in the respective isolated monomers leads to quantification of the loss (or gain) in the number of degrees of freedom at complex formation and their effects on protein complex formation mechanisms. This index is also valuable in the identification of collective motions within the protein that may play a critical role in the process of complex formation, and the influences they may have in the behavior and function of the complex (as well as the subunits constituting it) within the organism. Furthermore, the methodology, embedded in protein docking algorithms allows the development of a framework for categorizing and ranking decoys output by broadly used grid scoring type algorithms, one of which is the system for protein-protein interaction system MIAX that has been under continuous development in recent years.

Animals↗

2-Acetamido-4-p-tolyl-1,3-thiazole and 2-amino-4-p-tolyl-1,3-thiazolium chloride dihydrate.

The structures of 2-acetamido-4-tolyl-1,3-thiazole, C(12)H(12)N(2)OS, (I), and 2-amino-4-tolyl-1,3-thiazolium chloride dihydrate, C(10)H(11)N(2)S(+).Cl(-).2H(2)O, (II), reveal that both molecules are essentially planar, with the respective dihedral angles between the benzene and thiazole rings being 2.9 (1) and 10.39 (7) degrees . Compound (I) associates via a single N-H...O interaction to form a flat alternate-facing hydrogen-bonded chain [graph-set C(2)(2)(4)]. Compound (II) packs with the hydrogen-bonding associations of the Cl atoms and the water molecules creating a convoluted hydrogen-bonded ribbon made up of five-membered donor-acceptor rings, involving three water O atoms (with associated H atoms) and two Cl atoms. The thiazolium rings form stacked columns, aligned in the same direction as the hydrogen-bonded ribbons, of alternate-facing molecules that are also involved in the hydrogen-bonding network, linking to the Cl atoms and one of the water molecules. Subsequently, each Cl atom is the hydrogen-bond acceptor for five separate O/N-H associations.

Journal Article↗

Dealing with large data sets.

Collection, storage and retrieval of large amounts of data from multiple experiments for subsequent reduction, graphing and statistical analysis need not be a burdensome task. Although turnkey systems may offer significant economies for single well-defined and repetitive tasks, they may not permit sufficient flexibility to achieve the diverse aims required by many research programs. Using popular microcomputers to run one or a few experimental subjects may confront the investigator not only with significant bookkeeping problems, but also with an allocation of labor resources to computer maintenance and support that might be better invested in research effort. By using networked minicomputers, economies of scale emerge both in data collection, transfer, reduction, and analysis, as well as in maintenance, support, and scientific effort.

Data Collection↗

GIDEON: a comprehensive Web-based resource for geographic medicine.

GIDEON (Global Infectious Diseases and Epidemiology Network) is a web-based computer program designed for decision support and informatics in the field of Geographic Medicine. The first of four interactive modules generates a ranked differential diagnosis based on patient signs, symptoms, exposure history and country of disease acquisition. Additional options include syndromic disease surveillance capability and simulation of bioterrorism scenarios. The second module accesses detailed and current information regarding the status of 338 individual diseases in each of 220 countries. Over 50,000 disease images, maps and user-designed graphs may be downloaded for use in teaching and preparation of written materials. The third module is a comprehensive source on the use of 328 anti-infective drugs and vaccines, including a listing of over 9,500 international trade names. The fourth module can be used to characterize or identify any bacterium or yeast, based on laboratory phenotype. GIDEON is an up-to-date and comprehensive resource for Geographic Medicine.

Journal Article↗

Predictive Bayesian neural network models of MHC class II peptide binding.

We used Bayesian regularized neural networks to model data on the MHC class II-binding affinity of peptides. Training data consisted of sequences and binding data for nonamer (nine amino acid) peptides. Independent test data consisted of sequences and binding data for peptides of length </=25. We assumed that MHC class II-binding activity of peptides depends only on the highest ranked embedded nonamer and that reverse sequences of active nonamers are inactive. We also internally validated the models by using 30% of the training data in an internal test set. We obtained robust models, with near identical statistics for multiple training runs. We determined how predictive our models were using statistical tests and area under the Receiver Operating Characteristic (ROC) graphs (A(ROC)). Most models gave training A(ROC) values close to 1.0 and test set A(ROC) values >0.8. We also used both amino acid indicator variables (bin20) and property-based descriptors to generate models for MHC class II-binding of peptides. The property-based descriptors were more parsimonious than the indicator variable descriptors, making them applicable to larger peptides, and their design makes them able to generalize to unknown peptides outside of the training space. None of the external test data sets contained any of the nonamer sequences in the training sets. Consequently, the models attempted to predict the activity of truly unknown peptides not encountered in the training sets. Our models were well able to tackle the difficult problem of correctly predicting the MHC class II-binding activities of a majority of the test set peptides. Exceptions to the assumption that nonamer motif activities were invariant to the peptide in which they were embedded, together with the limited coverage of the test data, and the fuzziness of the classification procedure, are likely explanations for some misclassifications.

Amino Acid Sequence↗

The small world of human language.

Words in human language interact in sentences in non-random ways, and allow humans to construct an astronomic variety of sentences from a limited number of discrete units. This construction process is extremely fast and robust. The co-occurrence of words in sentences reflects language organization in a subtle manner that can be described in terms of a graph of word interactions. Here, we show that such graphs display two important features recently found in a disparate number of complex systems. (i) The so called small-world effect. In particular, the average distance between two words, d (i.e. the average minimum number of links to be crossed from an arbitrary word to another), is shown to be d approximately equal to 2-3, even though the human brain can store many thousands. (ii) A scale-free distribution of degrees. The known pronounced effects of disconnecting the most connected vertices in such networks can be identified in some language disorders. These observations indicate some unexpected features of language organization that might reflect the evolutionary and social history of lexicons and the origins of their flexibility and combinatorial nature.

Biological Evolution↗

Topological and causal structure of the yeast transcriptional regulatory network.

Interpretation of high-throughput biological data requires a knowledge of the design principles underlying the networks that sustain cellular functions. Of particular importance is the genetic network, a set of genes that interact through directed transcriptional regulation. Genes that exert a regulatory role encode dedicated transcription factors (hereafter referred to as regulating proteins) that can bind to specific DNA control regions of regulated genes to activate or inhibit their transcription. Regulated genes may themselves act in a regulatory manner, in which case they participate in a causal pathway. Looping pathways form feedback circuits. Because a gene can have several connections, circuits and pathways may crosslink and thus represent connected components. We have created a graph of 909 genetically or biochemically established interactions among 491 yeast genes. The number of regulating proteins per regulated gene has a narrow distribution with an exponential decay. The number of regulated genes per regulating protein has a broader distribution with a decay resembling a power law. Assuming in computer-generated graphs that gene connections fulfill these distributions but are otherwise random, the local clustering of connections and the number of short feedback circuits are largely underestimated. This deviation from randomness probably reflects functional constraints that include biosynthetic cost, response delay and differentiative and homeostatic regulation.

Gene Expression Regulation, Fungal↗

Annealing by two sets of interactive dynamics.

This work derives the mean field approximation to the mean configuration of a stochastic Hopfield neural network under the Boltzmann assumption. The new approximation is realized by two sets of interactive mean field equations, respectively estimating mean activations subject to mean correlations and mean correlations subject to mean activations. The two sets of interactive dynamics are derived based on two dual mathematical frameworks. Each aims to optimize the objective quantified by a combiation of the Kullback-Leibler (KL) divergence and the correlation strength between any two distinct fluctuated variables subject to fixed mean correlations or activations. The new method is applied to the graph bisection problem. By numerical simulations, we show that the new method effectively improves in both performance and relaxation efficiency against the naive mean field equation

Algorithms↗

A probabilistic methodology for integrating knowledge and experiments on biological networks.

Biological systems are traditionally studied by focusing on a specific subsystem, building an intuitive model for it, and refining the model using results from carefully designed experiments. Modern experimental techniques provide massive data on the global behavior of biological systems, and systematically using these large datasets for refining existing knowledge is a major challenge. Here we introduce an extended computational framework that combines formalization of existing qualitative models, probabilistic modeling, and integration of high-throughput experimental data. Using our methods, it is possible to interpret genomewide measurements in the context of prior knowledge on the system, to assign statistical meaning to the accuracy of such knowledge, and to learn refined models with improved fit to the experiments. Our model is represented as a probabilistic factor graph, and the framework accommodates partial measurements of diverse biological elements. We study the performance of several probabilistic inference algorithms and show that hidden model variables can be reliably inferred even in the presence of feedback loops and complex logic. We show how to refine prior knowledge on combinatorial regulatory relations using hypothesis testing and derive p-values for learned model features. We test our methodology and algorithms on a simulated model and on two real yeast models. In particular, we use our method to explore uncharacterized relations among regulators in the yeast response to hyper-osmotic shock and in the yeast lysine biosynthesis system. Our integrative approach to the analysis of biological regulation is demonstrated to synergistically combine qualitative and quantitative evidence into concrete biological predictions.

Cell Physiological Phenomena↗

A dynamically growing self-organizing tree (DGSOT) for hierarchical clustering gene expression profiles.

MOTIVATION: The increasing use of microarray technologies is generating large amounts of data that must be processed in order to extract useful and rational fundamental patterns of gene expression. Hierarchical clustering technology is one method used to analyze gene expression data, but traditional hierarchical clustering algorithms suffer from several drawbacks (e.g. fixed topology structure; mis-clustered data which cannot be reevaluated). In this paper, we introduce a new hierarchical clustering algorithm that overcomes some of these drawbacks. RESULT: We propose a new tree-structure self-organizing neural network, called dynamically growing self-organizing tree (DGSOT) algorithm for hierarchical clustering. The DGSOT constructs a hierarchy from top to bottom by division. At each hierarchical level, the DGSOT optimizes the number of clusters, from which the proper hierarchical structure of the underlying dataset can be found. In addition, we propose a new cluster validation criterion based on the geometric property of the Voronoi partition of the dataset in order to find the proper number of clusters at each hierarchical level. This criterion uses the Minimum Spanning Tree (MST) concept of graph theory and is computationally inexpensive for large datasets. A K-level up distribution (KLD) mechanism, which increases the scope of data distribution in the hierarchy construction, was used to improve the clustering accuracy. The KLD mechanism allows the data misclustered in the early stages to be reevaluated at a later stage and increases the accuracy of the final clustering result. The clustering result of the DGSOT is easily displayed as a dendrogram for visualization. Based on a yeast cell cycle microarray expression dataset, we found that our algorithm extracts gene expression patterns at different levels. Furthermore, the biological functionality enrichment in the clusters is considerably high and the hierarchical structure of the clusters is more reasonable. AVAILABILITY: DGSOT is available upon request from the authors.

Algorithms↗

Causal circuit tracing reveals distinct computational architectures in single-cell foundation models: inhibitory dominance, biological coherence, and cross-model convergence.

MOTIVATION: Sparse autoencoders (SAEs) decompose foundation-model activations into interpretable features, but the model-internal causal interactions between those features (i.e. what ablating one feature does to the others, as distinct from the biological causal structure of the underlying cells)-and how those model-internal relationships relate to biological structure-are uncharacterized in single-cell foundation models. RESULTS: We introduce model-internal causal circuit tracing-zeroing one SAE feature at a source layer and measuring the resulting change in all downstream SAE features, for each of 120 source features-and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96&#xa0;892 ablation-derived edges, 80&#xa0;191 forward passes). On annotation-selected source features, edges share GO/KEGG/Reactome/STRING/TRRUST ontology terms at 50.9%-68.5%, a 2.9-6.2&#xd7; enrichment over a configuration-preserving permutation null (P<.002); on 20 randomly sampled source features this attenuates to 21.5%-26.3%-still 2.5-3.1&#xd7; above null-quantifying the annotation-selection contribution. Inhibitory dominance (fraction of ablation edges with d<0, i.e. source activation supports downstream target) is 65.5%-89.4%. scGPT produces larger raw per-edge effects (mean |d|=1.40 versus 1.05); after feature-share normalization, Geneformer is stronger (paired gene-pair ratio 0.64 on 33&#xa0;301 shared pairs). Cross-model consensus yields 1142 architecture-invariant domain pairs (ordered pairs of GO biological-process categories "A&#x2192;B" each connected by at least one ablation edge in both models; 10.6&#xd7; enrichment over permutation null; P<.001). Circuit edge magnitude explains <1% of the variance in marginal driver-gene coexpression on the same cells (R2=0.010, n=31&#xa0;176): the graph encodes structure beyond bivariate correlation. Against a matched-cell-type ENCODE ChIP-seq prior, circuit-predicted transcription factor (TF)&#x2192;target pairs are enriched 2.06&#xd7; (Fisher OR 5.84), markedly higher than 1.12&#xd7; against TRRUST; direct ChIP-seq-supported target pairs show 10-30&#xd7; larger CRISPRi sign-bias-corrected excess than indirect pairs. Gene-level CRISPRi validation on Replogle K562 and the noncancer RPE1 arm (and a true primary-T-cell control from Shifrut E, Carnevale J, Tobin V et&#xa0;al. Genome-wide CRISPR screens in primary human T cells reveal key regulators of immune function. Cell 2018; 175: 1958-71.e15) after sign-bias correction shows excess over baseline of +0.03 and +0.35 percentage points on K562 and RPE1, respectively (baseline already 52%-56% from sign marginals); effect-magnitude Spearman correlations &#x3c1;&#x2248;0. Bootstrap and per-cell-type stability (N&#x2208;{50,100,200}; B cell, CD4&#xa0;+ T, macrophage) give Pearson r&#x2265;0.97 on shared edges with 100% sign agreement; edge Jaccard grows monotonically with sample size. The circuit graph is therefore highly reproducible as an effect-size map, cell type specific in edge identity, consistent with coexpression encoding, and weakly but detectably enriched for ChIP-seq-supported direct regulatory edges. AVAILABILITY AND IMPLEMENTATION: https://github.com/Biodyn-AI/bio-sae-circuits (Python). Archival DOI: 10.5281/zenodo.19,633,166 (Zenodo).

Humans↗

Spectral embedding finds meaningful (relevant) structure in image and microarray data.

BACKGROUND: Accurate methods for extraction of meaningful patterns in high dimensional data have become increasingly important with the recent generation of data types containing measurements across thousands of variables. Principal components analysis (PCA) is a linear dimensionality reduction (DR) method that is unsupervised in that it relies only on the data; projections are calculated in Euclidean or a similar linear space and do not use tuning parameters for optimizing the fit to the data. However, relationships within sets of nonlinear data types, such as biological networks or images, are frequently mis-rendered into a low dimensional space by linear methods. Nonlinear methods, in contrast, attempt to model important aspects of the underlying data structure, often requiring parameter(s) fitting to the data type of interest. In many cases, the optimal parameter values vary when different classification algorithms are applied on the same rendered subspace, making the results of such methods highly dependent upon the type of classifier implemented. RESULTS: We present the results of applying the spectral method of Lafon, a nonlinear DR method based on the weighted graph Laplacian, that minimizes the requirements for such parameter optimization for two biological data types. We demonstrate that it is successful in determining implicit ordering of brain slice image data and in classifying separate species in microarray data, as compared to two conventional linear methods and three nonlinear methods (one of which is an alternative spectral method). This spectral implementation is shown to provide more meaningful information, by preserving important relationships, than the methods of DR presented for comparison. Tuning parameter fitting is simple and is a general, rather than data type or experiment specific approach, for the two datasets analyzed here. Tuning parameter optimization is minimized in the DR step to each subsequent classification method, enabling the possibility of valid cross-experiment comparisons. CONCLUSION: Results from the spectral method presented here exhibit the desirable properties of preserving meaningful nonlinear relationships in lower dimensional space and requiring minimal parameter fitting, providing a useful algorithm for purposes of visualization and classification across diverse datasets, a common challenge in systems biology.

Algorithms↗

Data, network, and application: technical description of the Utah RODS Winter Olympic Biosurveillance System.

Given the post September 11th climate of possible bioterrorist attacks and the high profile 2002 Winter Olympics in the Salt Lake City, Utah, we challenged ourselves to deploy a computer-based real-time automated biosurveillance system for Utah, the Utah Real-time Outbreak and Disease Surveillance system (Utah RODS), in six weeks using our existing Real-time Outbreak and Disease Surveillance (RODS) architecture. During the Olympics, Utah RODS received real-time HL-7 admission messages from 10 emergency departments and 20 walk-in clinics. It collected free-text chief complaints, categorized them into one of seven prodromes classes using natural language processing, and provided a web interface for real-time display of time series graphs, geographic information system output, outbreak algorithm alerts, and details of the cases. The system detected two possible outbreaks that were dismissed as the natural result of increasing rates of Influenza. Utah RODS allowed us to further understand the complexities underlying the rapid deployment of a RODS-like system.

Algorithms↗