Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “representation learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Comparison of computational models of familiarity discrimination in the perirhinal cortex.

This study compares the efficiency and plausibility of published computational models of familiarity discrimination in the perirhinal cortex. Substantial evidence indicates that the perirhinal cortex is involved in both the familiarity discrimination aspect of recognition memory and in perceptual functions involved with representations of complete stimuli (i.e., object identification). Published models of how the perirhinal cortex may perform familiarity discrimination can be divided into two groups. The first group assumes that a proportion of perirhinal neurons form a network specialised just for familiarity discrimination (these models may be based on Hebbian or anti-Hebbian synaptic plasticity). In contrast, the second group assumes that both familiarity discrimination and learning representations of complete stimuli are performed within a single combined network. This study establishes that when the responses of neurons that provide input to the familiarity discrimination network are correlated (as indicated by experimental data), specialised networks based on anti-Hebbian learning may recognise the previous occurrence of many more stimuli (i.e., have a capacity up to thousands of times larger) than specialised networks based on Hebbian learning. The currently published combined models do not learn an optimal stimulus representation (they do not fully extract statistically independent features), and hence their capacities are even lower than those of the specialised models based on Hebbian learning. Hence, the combined models published thus far are critically less efficient than the specialised models based on anti-Hebbian learning. This study also compares the consistency of the models with experimental observations concerning what is known of synaptic plasticity in the perirhinal cortex and the responses of its neurons. Many theoretically important parameters remain undetermined, and experiments are suggested to provide information critical for refining and distinguishing between the various models. However, the above theoretical arguments and currently published data favour the existence of a separate network specialised for familiarity discrimination.

Animals↗

Teaching Bayesian reasoning: an evaluation of a classroom tutorial for medical students.

How likely is a diagnosis, given a particular medical test result? This probability can be determined by using Bayes's rule; however, previous research has shown that doctors often experience problems with Bayesian inferences. These findings illustrate the need to teach statistical reasoning in medical education. A new method of teaching Bayesian reasoning is representation learning: the key idea is to instruct medical students how to translate probability information into a representation that is easier to process, namely natural frequencies. This approach was implemented in a one-hour classroom tutorial to evaluate its effectiveness in this setting and compared with a traditional rule-learning approach. Evaluation took place two months after training by testing students' ability to correctly solve a Bayesian inference task with information represented as probabilities. While both approaches improved performance, almost three times as many students were able to profit from representation training as opposed to rule training.

Adult↗

GBFN: A gated bimodal fusion network leveraging foundation model embeddings for cancer drug sensitivity prediction.

Despite recent progress in deep learning for cancer drug sensitivity prediction, many existing models still rely on task-specific representation learning or relatively simple multimodal fusion, which may limit their ability to capture complex drug-cell interactions. To address this issue, we developed GBFN, a gated bimodal fusion network for continuous IC50 prediction that integrates pretrained drug and cell-line representations. Specifically, drug embeddings were obtained from SMI-TED, whereas cell-line embeddings were derived from transcriptomic profiles using BulkFormer. These two modalities were then combined through a dimension-wise gated fusion module and used to predict IC50 values in matched drug-cell line pairs. On the CCLE-based benchmark, GBFN outperformed representative neural baselines, including GraphDRP, TGSA, and TransEDRP, and achieved the best overall performance, with an R² of 0.8714 and an RMSE of 0.8938. Moreover, ablation analysis showed that the model using drug features and cell-line expression data with gated fusion performed better than the corresponding model using direct concatenation, indicating that the improvement was associated with the fusion strategy rather than with the input modalities alone. In addition, cell-line expression data were more informative than mutation data in the present setting, and adding mutation data to the model using drug features and expression data did not further improve performance. Across major cancer types, GBFN maintained generally high cell-line-level predictive performance, and perturbation-based attribution identified biologically relevant transcriptomic programs in selected drug-cell line settings. Together, these findings support GBFN as a compact and effective framework for continuous drug response prediction.

Humans↗

Alignment effects on learning multiple, use-relevant classification systems.

People often learn multiple classification systems that are relevant to some goal or use. We compared conditions in which subclassification within a category hierarchy was predicted by values on either the same (alignable) or different (nonalignable) dimensions between category hierarchies. The results indicated that learning in alignable conditions occurred in fewer blocks and with fewer errors than did learning in nonalignable conditions. This facilitation was not the result of differences between conditions in the representations learned by the participants, the number of dimensions needed for subclassification (Experiment 1), or the objective complexity of the learning task (Experiment 2). The facilitated learning in the alignable conditions appears to reflect a commitment on the part of the learner to alignment: the belief that the structure relevant to the use of one category system will also be relevant to the use of a comparable system.

Decision Making↗

In vivo recordings of spontaneous and odor-modulated dynamics in the Limax olfactory lobe.

The major central site of olfactory information processing in the terrestrial slug Limax maximus is the procerebral lobe of the cerebral ganglion, which exhibits oscillatory dynamics of its local field potential and propagates activity waves from its apex to its base, as determined by multisite optical and electrical measurements in vitro. The learning-dependent uptake of Lucifer yellow into procerebral neurons suggests that the procerebral lobe may form learned representations of odors. To determine the role of the procerebral lobe in odor processing and odor learning, we developed procedures to implant fine wire electrodes in the lobe, which allowed recordings of local field potential in freely behaving slugs. The procerebral lobe displays oscillatory dynamics of its local field potential in vivo; however the amplitude and frequency of the local field potential are much more variable in vivo than in vitro. Odor presentation leads to increased frequency and amplitude of the local field potential signal. Several lines of evidence indicate that the variations in the local field potential signal recorded in vivo are not due to movement artifacts or activity in adjacent muscles. Multiple amine, gaseous, and peptide neuromodulators known to be present in the procerebral lobe provide pathways by which activity or coupling of bursting neurons in the procerebral lobe could be altered, resulting in the observed amplitude and frequency modulation of the local field potential.

Animals↗

Reconstruction of vascular networks using three-dimensional models.

Reconstructing vasculature in three dimensions is a challenging problem. Early approaches concentrated on coronary vasculature in X-ray images, recent work uses magnetic resonance imagery of cerebral vasculature. In both cases a priori information has been used, and often the way this is represented has proven limiting to the scope of applications supported. For example, a particular representation may be useful only for X-ray images. This paper addresses two issues: 1) representing a collection of vasculature and 2) the reconstruction of individual vasculature from images. Our representation learns the variations in branching structures and vessel shapes that occur between individuals. It supports a vascular catalogue containing three-dimensional (3-D) anatomical models. The representation is task independent; here we use it to reconstruct vasculature from images. Our algorithm has four features to which we draw attention: 1) it is not premised wholly upon X-ray images (though that is our focus here); 2) it produces several feasible solutions rather than one; 3) it can generalize from the catalogue to reconstruct instances not yet learned; 4) it exhibits polynomial time complexity, reasonable memory consumption, and is reliable. Both our representation and reconstruction algorithm are new and useful approaches. In support of these claims, we present results gathered from X-rays of both simulated and real vasculature.

Algorithms↗

Invariant object recognition in the visual system with novel views of 3D objects.

To form view-invariant representations of objects, neurons in the inferior temporal cortex may associate together different views of an object, which tend to occur close together in time under natural viewing conditions. This can be achieved in neuronal network models of this process by using an associative learning rule with a short-term temporal memory trace. It is postulated that within a view, neurons learn representations that enable them to generalize within variations of that view. When three-dimensional (3D) objects are rotated within small angles (up to, e.g., 30 degrees), their surface features undergo geometric distortion due to the change of perspective. In this article, we show how trace learning could solve the problem of in-depth rotation-invariant object recognition by developing representations of the transforms that features undergo when they are on the surfaces of 3D objects. Moreover, we show that having learned how features on 3D objects transform geometrically as the object is rotated in depth, the network can correctly recognize novel 3D variations within a generic view of an object composed of a new combination of previously learned features. These results are demonstrated in simulations of a hierarchical network model (VisNet) of the visual system that show that it can develop representations useful for the recognition of 3D objects by forming perspective-invariant representations to allow generalization within a generic view.

Neural Networks, Computer↗

How hallucinations may arise from brain mechanisms of learning, attention, and volition.

This article suggests how brain mechanisms of learning, attention, and volition may give rise to hallucinations during schizophrenia and other mental disorders. The article suggests that normal learning and memory are stabilized through the use of learned top-down expectations. These expectations learn prototypes that are capable of focusing attention upon the combinations of features that comprise conscious perceptual experiences. When top-down expectations are active in a priming situation, they can modulate or sensitize their target cells to respond more effectively to matched bottom-up information. They cannot, however, fully activate these target cells. These matching properties are shown to be essential towards stabilizing the memory of learned representations. The modulatory property of top-down expectations is achieved through a balance between top-down excitation and inhibition. The learned prototype is the excitatory on-center in this top-down network. Phasic volitional signals can shift the balance between excitation and inhibition to favor net excitatory activation. Such a volitionally mediated shift enables top-down expectations, in the absence of supportive bottom-up inputs, to cause conscious experiences of imagery and inner speech and thereby to enable fantasy and planning activities to occur. If these volitional signals become tonically hyperactive during a mental disorder, the top-down expectations can give rise to conscious experiences in the absence of bottom-up inputs and volition. These events are compared with data about hallucinations. The article predicts where these top-down expectations and volitional signals may act in the laminar circuits of visual cortex and, by extension, in other sensory and cognitive neocortical areas, and how the level of abstractness of learned prototypes may covary with the abstractness of hallucinatory content. A similar breakdown of volition may lead to delusions of control in the motor system.

Attention↗

A Graph Contrastive Learning Method for Enhancing Genome Recovery in Complex Microbial Communities.

Accurate genome binning is essential for resolving microbial community structure and functional potential from metagenomic data. However, existing approaches-primarily reliant on tetranucleotide frequency (TNF) and abundance profiles-often perform sub-optimally in the face of complex community compositions, low-abundance taxa, and long-read sequencing datasets. To address these limitations, we present MBGCCA, a novel metagenomic binning framework that synergistically integrates graph neural networks (GNNs), contrastive learning, and information-theoretic regularization to enhance binning accuracy, robustness, and biological coherence. MBGCCA operates in two stages: (1) multimodal information integration, where TNF and abundance profiles are fused via a deep neural network trained using a multi-view contrastive loss, and (2) self-supervised graph representation learning, which leverages assembly graph topology to refine contig embeddings. The contrastive learning objective follows the InfoMax principle by maximizing mutual information across augmented views and modalities, encouraging the model to extract globally consistent and high-information representations. By aligning perturbed graph views while preserving topological structure, MBGCCA effectively captures both global genomic characteristics and local contig relationships. Comprehensive evaluations using both synthetic and real-world datasets-including wastewater and soil microbiomes-demonstrate that MBGCCA consistently outperforms state-of-the-art binning methods, particularly in challenging scenarios marked by sparse data and high community complexity. These results highlight the value of entropy-aware, topology-preserving learning for advancing metagenomic genome reconstruction.

canonical correlation analysis↗

PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clustering.

The development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency.

Single-Cell Analysis↗

HAVNET: A New Neural Network Architecture for Pattern Recognition.

A new artificial neural network architecture, specifically designed for two-dimensional binary pattern recognition, is introduced. The network employs a unique similarity metric, based on the Hausdorff distance, to determine the degree of match between an input pattern and a learned representation. Use of this metric in the network leads to behaviour that is more consistent with human performance than that generated by similarity metrics currently in use in other artificial neural networks. A detailed description of the architecture, the learning equations, and the recall equations for the network are presented. An extension of the network is also described in which each class of learned objects is represented by multiple two-dimensional aspects. This extension greatly increases the utility of the network for tasks like character recognition and three-dimensional vision. The network is employed on an example pattern recognition task to demonstrate its application, with very good results. Copyright 1996 Elsevier Science Ltd.

Journal Article↗

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics↗

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models.

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Deep Learning↗

A Knowledge-Enhanced Multimodal Framework with Genomic Reconstruction for DLBCL Drug Response Prediction.

Diffuse large B-cell lymphoma (DLBCL) exhibits substantial biological heterogeneity, leading to pronounced variability in patient response to therapy. Accurate drug response prediction is therefore critical for precision treatment but remains challenging in clinical settings where genomic sequencing, a highly informative modality, is frequently incomplete. Existing methods, often developed from cell-line pharmacogenomic datasets or single-modality data, typically assume fully observed molecular profiles and thus show limited robustness under missing genomic data. To address this limitation, a knowledge-enhanced multimodal framework with genomic reconstruction (KeM-DRP) is proposed for individualized drug response prediction in DLBCL. The framework models the central role of genomics by integrating biological prior knowledge through a gene-pathway-biological process hierarchy, enabling robust representation learning from sparse observations. To compensate for missing genomic measurements, a cross-modal genomic compensation module reconstructs genomically informed latent features from routinely available clinical modalities. Furthermore, a genomics-guided adaptive fusion strategy dynamically integrates heterogeneous modalities conditioned on observed or reconstructed genomic representation. Experiments on a real-world DLBCL cohort demonstrate that KeM-DRP consistently outperforms competitive baselines. The reconstructed genomic representation represents most predictive utility, highlighting the robustness and practical value of the framework under incomplete genomic data.

Journal Article↗

Human theta oscillations related to sensorimotor integration and spatial learning.

oscillations in the rat hippocampus have been implicated in sensorimotor integration (Bland, 1986), especially during exploratory and wayfinding behavior. We propose that human cortical activity coordinates sensory information with a motor plan to guide wayfinding behavior to known goal locations. To test this hypothesis, we analyzed invasive recordings from epileptic patients while they performed a spatially immersive, virtual taxi driver task. Consistent with this hypothesis, we found oscillations during both exploratory search and goal-seeking behavior and, in particular, during virtual movement, when sensory information and motor planning were both in flux, compared with periods of self-initiated stillness. oscillations had different topographic and spectral characteristics during searching than during goal-seeking, suggesting that different cortical networks exhibit depending on which cognitive functions are driving behavior (spatial learning during exploration vs orienting to a learned representation during goal-seeking). In contrast, oscillations in the beta band appeared to be related to simple motor planning, likely a variant of the Rolandic mu rhythm. These findings suggest that human cortical oscillations act to coordinate sensory and motor brain activity in various brain regions to facilitate exploratory learning and navigational planning.

Adolescent↗

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis↗

Associative learning modifies neural representations of odors in the insect brain.

Recording brain activity in vivo during learning is fundamental to understanding how memories are formed. We used functional calcium imaging to track odor representations in the primary chemosensory center of the honeybee, the antennal lobe, while training animals to discriminate a rewarded odor from an unrewarded one. Our results show that associative learning transforms odor representations and decorrelates activity patterns for the rewarded versus the unrewarded odor, making them less similar. Additionally, activity for the rewarded but not for the unrewarded odor is increased. These results indicate that neural representations of the environment may be modified through associative learning.

Animals↗

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans↗