Search PubMedSearch

SEARCH · Search PubMed

Results for “mutual information”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Mutual Information-based Prognostic Biomarker Discovery in Cancer Genomics: Conceptual Framework and Representative Applications of MI-POG.

Mutual information (MI)-based approaches have increasingly been applied to cancer genomics; however, their use for genome-wide prognostic biomarker discovery remains relatively underexplored. The present article summarizes the conceptual workflow of Mutual Information-based Prognostic Omics Gene (MI-POG) based on previously published applications in breast cancer, lower-grade glioma, and other cancer datasets. The framework consists of clinical endpoint discretization, genome-wide MI-based screening, candidate ranking, and downstream validation using conventional survival-analysis approaches. Previous MI-POG applications identified solute carrier family 20 member 1 (SLC20A1) as a prognostic biomarker in hormone receptor-positive breast cancer. Elevated SLC20A1 expression was associated with unfavorable survival outcomes and was independently validated in the Molecular Taxonomy of Breast Cancer International Consortium (METABRIC) cohort. Methodological analyses demonstrated how survival endpoints can be integrated into an information-theoretic framework through fixed-time outcome discretization, enabling model-independent assessment of molecular-clinical dependencies. Applications across multiple cancer datasets suggested the potential applicability of the framework across biologically distinct tumor types, although further validation will be required to establish its robustness and generalizability. In conclusion, MI-POG can be formalized as an information-theoretic framework for genome-wide identification of prognostic biomarkers by quantifying molecular-clinical dependencies using mutual information. Representative applications from previously published studies suggest that MI-POG may complement conventional survival-analysis approaches and provide a useful strategy for biomarker discovery, although additional benchmarking and prospective validation will be required.

Humans

[The correlation between autoregressive power spectrum of EEG and computerized tomography in tuberous sclerosis (author's transl)].

The autoregressive power spectrum and their component analyses of 36 electroencephalograms of 9 patients with tuberous sclerosis and 21 healthy children were studied. The data were recorded on the analog tapes, and the 20 second artifact free segment of records was digitized at 50 samples/sec. The autoregressive power spectrum and their component were calculated by the methods of Sato (1976) with the minicomputer PDP 11/40. The component consisted of the first and the second order elementary processes. The former showed a transient nonscillatory delta wave, whereas the latter showed damped oscillatory waves of delta, theta, alpha and beta rhythms in the EEG. The characteristics of these component rhythms in the EEG were given by the frequency, the time constant of the nonoscillatory delta, the damping time of the oscillatory component waves (time constant of the envelope of the damped oscillation), the mutual information amounts, etc. Thus, the correlation between these characteristics of component and CT-scan in tuberous sclerosis were examined. The results were as follows: 1. Compared with the characteristics of EEG in normal children, the mutual information amount, the damping time and/3r time constant showed significantly lower value in alpha rhythms of frontal-, central-regions and in theta, delta rhythms of occipital region in the patient with tuberous sclerosis. 2. Multiple subependymal high density areas were found in all of these patients on CT. EMI-number of these high density areas were less than that of calcification in younger children below the age of 3 years, but in older children were equivalent to that of calcification. 3. The correlation between the number of subependymal nodule on CT and the mutual information amount in the EEG showed significantly the negative coefficient in theta, delta rhythms of O1 region in older age group of patients. 4. The correlation between EMI-number of these subependymal high density areas and the mutual information amount in the EEG showed significantly thenegative coefficient in theta & delta rhythms of O1 region in all cases of patients. 5. It is considered that the multiple subependymal nodules may influence the background activities of the central and the occipital regions in the EEG of tuberous sclerosis.

Child

Stress-induced altered expression of hippocampal nuclear and mitochondrial encoded genes in rats and cross-species genetic associations reveal molecular links to depression.

BACKGROUND: Mitochondria play a pivotal role in energy production, and their dysfunction not only hampers cells' ability to meet energy requirements but also contributes to the impairment of neural plasticity, a critical feature of depressive disorders. In this study, mitochondrial cross-omics analysis was carried out in the hippocampus of restraint rats to understand the role of mitochondria in depression pathophysiology. METHODS: The expression profiles of hippocampal mitochondrial and nuclear-encoded genes in mitochondrial fractions from restraint and handled control rats were obtained using high-throughput RNA sequencing. Weighted gene co-expression network analysis (WGCNA) was used to identify the gene co-expression and pathways associated with the restraint phenotype. Mutual Information Network algorithm tools Arance, CLR, and MRNET were additionally used to screen the functional modules and hub genes and their similarity with the WGCNA-based network analysis. Finally, cross-species homology followed by gene association analysis was conducted to obtain SNPs and haplotypes related to depression phenotype. RESULTS: A significant proportion of mitochondrial and nuclear-encoded genes showed differential regulation in the hippocampus of restraint rats. WGCNA and Mutual Information Network analysis yielded distinct functional modules significantly related to restraint phenotype. Further network analysis revealed distinct co-expression patterns associated with differentially expressed genes associated with these modules. Cross-species analysis showed 39 significantly associated SNPs with the depression phenotype, where the most significant SNP, rs10899570, was located within the TENM4 gene. Further, rs1573529 and rs10899570 were distributed into the linkage disequilibrium block where SNPs were highly correlated. Subsequent haplotype analysis showed that rs1573529 and rs10899570 were significantly associated with depressive behavior. CONCLUSIONS: The study demonstrates a significant impact of restraint stress on mitochondrial functions and genetic association, suggesting their critical role in depression pathophysiology.

Animals

A Graph Contrastive Learning Method for Enhancing Genome Recovery in Complex Microbial Communities.

Accurate genome binning is essential for resolving microbial community structure and functional potential from metagenomic data. However, existing approaches-primarily reliant on tetranucleotide frequency (TNF) and abundance profiles-often perform sub-optimally in the face of complex community compositions, low-abundance taxa, and long-read sequencing datasets. To address these limitations, we present MBGCCA, a novel metagenomic binning framework that synergistically integrates graph neural networks (GNNs), contrastive learning, and information-theoretic regularization to enhance binning accuracy, robustness, and biological coherence. MBGCCA operates in two stages: (1) multimodal information integration, where TNF and abundance profiles are fused via a deep neural network trained using a multi-view contrastive loss, and (2) self-supervised graph representation learning, which leverages assembly graph topology to refine contig embeddings. The contrastive learning objective follows the InfoMax principle by maximizing mutual information across augmented views and modalities, encouraging the model to extract globally consistent and high-information representations. By aligning perturbed graph views while preserving topological structure, MBGCCA effectively captures both global genomic characteristics and local contig relationships. Comprehensive evaluations using both synthetic and real-world datasets-including wastewater and soil microbiomes-demonstrate that MBGCCA consistently outperforms state-of-the-art binning methods, particularly in challenging scenarios marked by sparse data and high community complexity. These results highlight the value of entropy-aware, topology-preserving learning for advancing metagenomic genome reconstruction.

canonical correlation analysis

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics

A two-pathway informon theory of conditioning and adaptive pattern recognition.

A neural network theory is proposed which offers an explanation of many of the facts of classical and operant conditioning and adaptive pattern recognition. Interconnected networks of units have been studied and simulated which embody only two rules; firstly, units have inputs from pathways of variable and of fixed conductivity; secondly, the conductivity of a variable pathway is made proportional to the negative of the mutual information function between the signals at its input and output. The signal in a fixed pathway indicates whether the total input to the variable pathways is a member or not of some class. After a learning phase in which the unit, called an informon, receives such labelled inputs, it is able to predict the class of future unlabelled inputs. Such units are stable and their steady state can be calculated.

Conditioning, Psychological

Neurophysiological predictions of a two-pathway informon theory of neural conditioning.

According to the informon theory there must be variable and fixed synapses in a neurone for conditioning to occur. For a variable synapse to behave like an informon pathway its conductivity needs to depend only on the average values of its presynaptic potential and of the internal state of the neurone. Eight predictions are made about the detailed functioning of such a synapse. In a minimal hypothesis all fixed synapses are inhibitory; but sign reversals are considered. Let one unit A in the receptive field of a neurone drive it through a fixed synapse, and all other units, e.g. B, drive variable synapses; then the theory predicts that the conductivity of the B synapse becomes proportional to the mutual information function between the signals at A and B; so inputs which tend to occur with the A signal become connected positively to the neurone. Applied to visual pathways this principle leads to the formation of edge and grating detectors. If X and Y cells excite variable and fixed synapses respectively, simple and complex cells should be driven by both X and Y cells, the latter being inhibitory. The two-pathway theory resolves two apparent conflicts between experimental facts.

Conditioning, Psychological

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping

scPlantLLM: A Foundation Model for Exploring Single-cell Expression Atlases in Plants.

Single-cell RNA sequencing (scRNA-seq) provides unprecedented insights into plant cellular diversity by enabling high-resolution analyses of gene expression at the single-cell level. However, the complexity of scRNA-seq data, including challenges in batch integration, cell type annotation, and gene regulatory network (GRN) inference, demands advanced computational approaches. To address these challenges, we developed scPlantLLM, a Transformer model trained on millions of plant single-cell data points. Using a sequential pretraining strategy incorporating masked language modeling and cell type annotation tasks, scPlantLLM generates robust and interpretable single-cell data embeddings. When applied to Arabidopsis thaliana datasets, scPlantLLM excels in clustering, cell type annotation, and batch integration, achieving an accuracy of up to 0.91 in zero-shot learning scenarios. Furthermore, the model demonstrates an ability to identify biologically meaningful GRNs and subtle cellular subtypes, showcasing its potential to advance plant biology research. Compared to traditional methods, scPlantLLM outperforms in key metrics such as adjusted rand index (ARI), normalized mutual information (NMI), and silhouette score (SIL), highlighting its superior clustering accuracy and biological relevance. scPlantLLM represents a foundation model for exploring plant single-cell expression atlases, offering unprecedented capabilities to resolve cellular heterogeneity and regulatory dynamics across diverse plant systems. The code used in this study is available at https://github.com/compbioNJU/scPlantLLM.

Single-Cell Analysis

Communication and cooperation between professionals in the field of rehabilitation.

Some criteria and characteristics of Rehabilitation as a complex, multidimensional and interdisciplinary field of theory, practice and research are pointed out at the beginning of this paper. It is shown, that such a rehabilitation concept demands the cooperation between professionals in this field (a) in practice within the rehabilitation team, (b) in research for the realization of well coordinated and problem-oriented interdisciplinary projects, (c) in theory for the development of more adequate theoretical concepts and better strategies for teaching and training personnel. Improved mutual information on an international level is regarded as a basis or prerequisite for such cooperation. The multiple dimensions and directions of communication necessary to assure such cooperation are illustrated. In a description of barriers to communication in the field of rehabilitation the following facts are considered as possible causes for the present insufficient cooperation among professionals in this field- (1) specialization and exclusiveness, (2) language and presentation, (3) differences in culture, demographical situations and economic/historical development, (4) geographical, national, legal or administrative barriers, (5) professional competition, patenting and copyright, (6) lack of adequate documentation and dissemination of information, (7) inadequate use of existing information systems and communication channels. The article closes with suggestions on ways to improve communications and cooperation between professionals in rehabilitation.

Communication

A multicenter inquiry into the etiology of pancreatic diseases.

A multicenter study on the etiology and diet of patients with pancreatic diseases has been realized with the collaboration of 36 centers in 19 countries having widely different climatic and racial conditions. 2,478 cases were studied: acute pancreatitis (AP), 222 males, 208 females; calcified chronic pancreatitis (CCP), 801 males, 134 females; non-calcified chronic pancreatitis (NCCP), 525 males, 155 females; pancreatic cancer (PK), 69 males, 14 females; controls, 281 males, 62 females. The analysis of mutual information and the factorial analysis of correspondences have been used. With regard to chronic pancreatitis, the 19 countries could be classified into 4 classes presenting relative similarities. (A) Southern Europe: The diet is rich in carbohydrates, protein and lipids, alcohol intake is primarily in the form of wine and the pathology is dominated by CCP. There are much fewer women than men with chronic pancreatitis. (B) Northern Europe, to which may be added Argentina and Chile, is characterized by a protein- and lipid-rich diet, a beer-based alcohol consumption and a distinct prevalence of AP and NCCP. The prevalence of males with chronic pancreatitis is less marked than in southern Europe. (C) Japan has a lipid-poor diet and a low frequency of CCP and NCCP. (D) A fourth group is mostly composed of tropical countries with mixed races. It may be divided into 2 subclasses: (a) India is the most characteristic country of the first type with low fat, low protein diet, no alcoholism, high frequency of CCP (at an early age); (b) Brasil and South Africa are representative of the second subclass with very high alcohol intake in the form of spirits and a high frequency of CCP.

Adult

[Sports in neurological rehabilitation (author's transl)].

Little has been written about the effects of sports activities on neurological diseases such as cerebral palsy, hemiplegia, brain trauma and paraplegia, and the available reports are mainly related to cardiopulmonary parameters. Compared to the securing of physiological data, it is more difficult to operationally define, or quantify questions concerning the motivation and social background--in a broader sense, sociological background--of the patient. Both data, however, are indispensable to rehabilitation. The experiences gained with a group of children with spina bifida attending a school for the physically handicapped, are used to describe the sequela of the central nervous system defect with which the physical education instructor has to cope, and stresses how important mutual information amongst the team members is for successful rehabilitation. The criteria for sports activities with spina bifida children are: (a) to promote their independence, (b) to give them socialisation stimuli, (c) to enhance their physical performance. It is hoped with this programme that a further personality maturation and stabilisation be achieved.

Adult

Feasibility of computer evaluation of the histopathologic findings in human skeletal muscle disease.

Optical Fourier methods of image processing coupled with mutual information analysis were applied to the separation of human muscle specimens into four diagnostic classes. This process allowed rapid, relatively inexpensive data acquisition, reduction, and classification. The accuracy of the method when compared to that of experienced surgical pathologists is acceptable. Difficulties arise when material outside of the previous experience of the computer is presented.

Computers

CaXML: Chemistry-informed machine learning explains mutual changes between protein conformations and calcium ions in calcium-binding proteins using structural and topological features.

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of CaXML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

Machine Learning

[Early diagnosis of cardiac insufficiency with the spiroergometric method].

Standard and submaximal physical load tests were contrasted in studying the functional state of the cardio-respiratory system in healthy individuals and in patients with coronary atherosclerosis, post-infarction cardiosclerosis, mitral lesions and cardiac-pulmonary insufficiency. The application of physical load tests of different types is shown to be instrumental in obtaining a mutually complementary information about the function of the cardiac-respiratory system and the degree of pulmonary and cardiac decompensation. The authors attach great importance to determining the ratio of an actual oxygen uptake to its proper values and suggest using this indicator, called by them performance capacity index, in quantitative appraisal of physical capacity to perform work. In recognizing early stages of circulatory insufficiency of considerable interest is determining the ratio of the venous blood concentration of lactate, pyruvate and some other biochemical factors to the amount of work done.

Adolescent

The field of possible structures for the chlorophyll a dimer in photosystem I of green plants delineated by polarized photochemistry.

Photoselection experiments with immobilized photosystem I particles have been done to determine the mutual orientation of pigments in the reaction centre. When these particles are excited and interrogated with linearly polarized light, the flash-induced transient absorption changes (mainly from the chlorophyll a dimer) reveal linear dichroism, which yields information on the mutual orientation between the excited and the interrogated transition moments. The interpretation of the data, however, is ambiguous, (1) for reasons of principles inherent in the photoselection technique when applied to complex systems and (2) because of incomplete knowledge about the relative contribution of x- and of y-polarized transitions of chlorophyll a to absorption or to absorption changes at a given wavelength. We find it impossible to attribute any particular structure to the photooxidizable dimer based on photoselection data alone. Instead we present a field of possible structures, imposing constraints on proposed models for the dimer structure.

Chlorophyll