Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “representation learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Representations in learning new faces: evidence from prosopagnosia.

We report the performance of a prosopagnosic patient on face learning tasks under different encoding instructions (i.e., levels of processing manipulations). R.J. performs at chance when given no encoding instructions or when given "shallow" encoding instruction to focus on facial features. By contrast, he performs relatively well with "deep" encoding instructions to rate faces in terms of personality traits or when provided with semantic and name information during the study phase. We propose that the improvement associated with deep encoding instructions may be related to the establishment of distinct visually derived and identity-specific semantic codes. The benefit associated with deep encoding in R.J., however, was found to be restricted to the specific view of the face presented at study and did not generalize to other views of the same face. These observations suggest that deep encoding instructions may enhance memory for concrete or pictorial representations of faces in patients with prosopagnosia, but that these patients cannot compensate for the inability to construct abstract structural codes that normally allow faces to be recognized from different orientations. We postulate further that R.J.'s poor performance on face learning tasks may be attributable to excessive reliance on a feature-based left hemisphere face processing system that operates primarily on view-specific representations.

Aged↗

Dissociating stimulus information from internal representation--a case study in object recognition.

Human object recognition is a function of both internal memory representation(s) and stimulus input information. The role of the latter has been so far largely overlooked, and the nature of the representation is often directly equated with recognition performance. We quantify stimulus information for three classes of objects in order of decreasing object complexity: unconnected balls, balls connected with lines, and balls connected with cylinders. In an object discrimination task, subjects' performance improved with the decreasing object complexity. We show that input information also increases with decreasing object complexity. Therefore, the results could potentially be accounted for either by differences in the object representations learned for each class of objects, or by the increased information about the three-dimensional (3D) structure inherent in images of the less complex objects, or by both. We demonstrate that, when image information is taken into account, by computing efficiencies relative to a set of ideal observers, subjects were more efficient in recognizing the less complex objects. This suggests that differences in subjects' performance for different object classes is at least partly a function of the internal representations learned for the different object classes. We stress that this conclusion cannot be achieved without the quantitative analysis of stimulus input information.

Discrimination, Psychological↗

Resolution-dependent self-supervised transfer in chest radiograph classification.

BACKGROUND: Self-supervised learning (SSL) has improved visual representation learning, but its value in chest radiography remains uncertain. DINOv3 extends earlier SSL models through Gram-anchored self-distillation and explicit high-resolution adaptation. Whether these changes improve transfer learning for chest radiograph classification has not been established. METHODS: We benchmarked DINOv3 against DINOv2 and supervised ImageNet initialization across seven chest radiograph datasets comprising 816,183 radiographs from pediatric and adult cohorts. ViT-B/16 and ConvNeXt-B were evaluated under full fine-tuning at 224 × 224 and 512 × 512 pixels, with targeted 1024 × 1024 experiments on three cohorts. Additional analyses examined parameter-efficient adaptation, synthetic label corruption, external validation, frozen 7B features, and computational efficiency. The primary outcome was the mean area under the receiver operating characteristic curve across labels. RESULTS: In adult cohorts, DINOv3 did not consistently outperform DINOv2 at 224 × 224 pixels, but became the strongest initialization at 512 × 512 pixels, especially with ConvNeXt-B. Gains were greatest for small focal and boundary-dependent abnormalities, whereas large-structure findings changed little. The pediatric cohort showed no significant benefit from DINOv3, higher resolution, or backbone choice. Scaling to 1024 × 1024 rarely improved performance and markedly increased computational cost. ConvNeXt-B remained superior to ViT-B/16 under both full and parameter-efficient adaptation. External validation preserved the 512 × 512 DINOv3 advantage, whereas synthetic label corruption showed that this benefit should not be interpreted simply as superior noise robustness. Frozen DINOv3-7B features underperformed relative to fully adapted 86 to 89M-parameter backbones. CONCLUSIONS: For adult chest radiograph classification, DINOv3 provides its most reliable benefit at 512 × 512 pixels, particularly with ConvNeXt-B. Fully adapted mid-sized models at 512 × 512 pixels provided the best performance-cost trade-off in our benchmark.

Journal Article↗

DPAS-Graph: adaptive spatial-feature relation learning for spatial RNA-to-protein prediction and virtual protein profiling.

Paired spatial multi-omics provides a supervised basis for learning RNA-protein correspondence in situ, but predicting protein abundance from spatial transcriptomic data alone remains challenging across tissue contexts and protein panels. Here, we present DPAS-Graph, an adaptive relation-learning framework for spatial RNA-to-protein prediction. Rather than directly merging spatial proximity and transcriptomic similarity as fixed graph priors, DPAS-Graph represents them as two relation channels on a shared edge support and updates their contributions during representation learning for protein prediction. Its Niche-Coupled Field Encoder combines layer-wise edge-relation modeling, intra-branch relation refinement, and cross-branch residual correction to learn spot representations for protein abundance prediction. In a leave-one-dataset-out benchmark across seven paired spatial multi-omics datasets, DPAS-Graph achieved lower aggregate prediction errors and improved spot-level agreement of protein expression profiles, with gains mainly reflected in error-based metrics and PCC-Spot. Spatial autocorrelation and protein-derived domain agreement analyses were further used to characterize the spatial behavior of the predicted protein maps. When applied to external RNA-only spatial sections, DPAS-Graph generated qualitatively interpretable marker-level virtual protein maps, illustrating its use as a complementary tool for protein-level interpretation of transcriptomics-only spatial data.

RNA↗

soFusion: facilitating tissue structure identification via spatial multi-omics data fusion.

The rapid advancement of spatial multi-omics technologies has opened new avenues for dissecting tissue architecture with unprecedented resolution. However, inherent disparities across omics modalities, such as differences in biological hierarchy and resolution, pose significant challenges for integrative analysis. To address this, we present soFusion, a method for representation learning on spatial multi-omics data that enables automated identification of tissue compartmentalization. soFusion employs a graph convolutional network (GCN) to extract latent embeddings from spatial omics profiles. To simultaneously capture both cross-modality relationships and modality-specific features, we introduce a novel strategy for intra- and inter-omics feature learning. Moreover, modality-specific decoders are designed to preserve the unique information embedded in each omics type. We evaluated soFusion on multiple datasets including gene expression, protein expression, and epigenetic features. Across all benchmarks, soFusion consistently outperformed existing methods in delineating anatomical structures and identifying spatial domains with improved continuity and reduced noise. Collectively, soFusion offers an effective solution for spatial multi-omics integration, substantially enhancing the robustness of spatial domain identification.

Humans↗

[Analysis and cognitive modelling of the analogical process in psychosis].

The disturbances of cognitive processes in psychotic patients are well known: the delusional interpretation is "the inference from a right perception into a wrong concept" (Dromard), "an wrong intuition about the meaning of what is perceived, seen or heard" (H. Ey). Analogy is the very core of any cognitive process: relating a strange thing to some object already part of the experience enables to set up differences, oppositions, connections, classes. Any semantic process (something stands for another thing) originates in analogy. It is the basis of every interpretation and world's knowledge. Its soundness is by no means reliable, but for the inner strength of the analogical network and its power to integrate new objects. It's easy to fall out of the track... A wrong analogy, better, a wrong one that would not be acknowledged as a mistake, would be enough: the gap is quite narrow between interpretations leading either to understanding or to misreading, only filled through the relation of other people providing the necessary clues. The contemporary papers about cognitive process are driving towards two main trends: 1) Neuromimetic models, and the building of neuronal networks, whose emergent properties point out the basically analogical character of representations, learning and memory. 2) Cognitive models, dealing with representations and algorithms, and leading to Artificial Intelligence Programs. We tried to build a model (both cognitive and AI) of the analogical process and its psychotic disturbances. Our model describes how simple analogical problems are solved: If (A) becomes (B), what about (C)? Making up the psychological model and its AI translation led to propound the concept of Universes as sets consisting of ONE likely or relevant link between two objects, and such intrinsic of extrinsic properties of the objects as are involved in this relation. The model uses 3 different universes: Universe U1, made up of one of the possible transformation kinks from (A) to (B) and (A)'s properties involved in this actual transformation. Universe U2, made out of the likeness link between Universe U1 and (C). Universe U3, performing in fact the validation procedure of the result. The analogical reasoning goes through the three universes, along an iterative loop again and again until a nice result is found.(ABSTRACT TRUNCATED AT 400 WORDS)

Artificial Intelligence↗

Multimodal alignment improves generalizability of genomic biomarker prediction in computational pathology.

Computational pathology models that use digitized histopathology whole-slide images have the potential to become a cost-effective and scalable alternative to molecular assays for the prediction of genomic biomarkers, a key task in precision oncology. However, as new genomic biomarkers are discovered or quantified, large, labeled datasets must be prospectively collected to train new models. To address this challenge, we developed multimodal alignment for biomarker learning and generalization (MARBLE), a multimodal contrastive pretraining strategy that integrates structured biomarker knowledge into representation learning of histopathology images. MARBLE aligns histopathology-derived representations with representations of genomic biomarkers generated by a large language model (LLM) and a protein language model (PLM). This biologically informed alignment enables data-efficient generalization to novel, out-of-distribution biomarkers. Using the MSK-IMPACT cohort of over 40,000 patients across multiple biomarker panel versions, we design experiments grounded in real-world data to demonstrate the value of our proposed approach.

CP: computational biology↗

GBFN: A gated bimodal fusion network leveraging foundation model embeddings for cancer drug sensitivity prediction.

Despite recent progress in deep learning for cancer drug sensitivity prediction, many existing models still rely on task-specific representation learning or relatively simple multimodal fusion, which may limit their ability to capture complex drug-cell interactions. To address this issue, we developed GBFN, a gated bimodal fusion network for continuous IC50 prediction that integrates pretrained drug and cell-line representations. Specifically, drug embeddings were obtained from SMI-TED, whereas cell-line embeddings were derived from transcriptomic profiles using BulkFormer. These two modalities were then combined through a dimension-wise gated fusion module and used to predict IC50 values in matched drug-cell line pairs. On the CCLE-based benchmark, GBFN outperformed representative neural baselines, including GraphDRP, TGSA, and TransEDRP, and achieved the best overall performance, with an R² of 0.8714 and an RMSE of 0.8938. Moreover, ablation analysis showed that the model using drug features and cell-line expression data with gated fusion performed better than the corresponding model using direct concatenation, indicating that the improvement was associated with the fusion strategy rather than with the input modalities alone. In addition, cell-line expression data were more informative than mutation data in the present setting, and adding mutation data to the model using drug features and expression data did not further improve performance. Across major cancer types, GBFN maintained generally high cell-line-level predictive performance, and perturbation-based attribution identified biologically relevant transcriptomic programs in selected drug-cell line settings. Together, these findings support GBFN as a compact and effective framework for continuous drug response prediction.

Humans↗

In vivo recordings of spontaneous and odor-modulated dynamics in the Limax olfactory lobe.

The major central site of olfactory information processing in the terrestrial slug Limax maximus is the procerebral lobe of the cerebral ganglion, which exhibits oscillatory dynamics of its local field potential and propagates activity waves from its apex to its base, as determined by multisite optical and electrical measurements in vitro. The learning-dependent uptake of Lucifer yellow into procerebral neurons suggests that the procerebral lobe may form learned representations of odors. To determine the role of the procerebral lobe in odor processing and odor learning, we developed procedures to implant fine wire electrodes in the lobe, which allowed recordings of local field potential in freely behaving slugs. The procerebral lobe displays oscillatory dynamics of its local field potential in vivo; however the amplitude and frequency of the local field potential are much more variable in vivo than in vitro. Odor presentation leads to increased frequency and amplitude of the local field potential signal. Several lines of evidence indicate that the variations in the local field potential signal recorded in vivo are not due to movement artifacts or activity in adjacent muscles. Multiple amine, gaseous, and peptide neuromodulators known to be present in the procerebral lobe provide pathways by which activity or coupling of bursting neurons in the procerebral lobe could be altered, resulting in the observed amplitude and frequency modulation of the local field potential.

Animals↗

Reconstruction of vascular networks using three-dimensional models.

Reconstructing vasculature in three dimensions is a challenging problem. Early approaches concentrated on coronary vasculature in X-ray images, recent work uses magnetic resonance imagery of cerebral vasculature. In both cases a priori information has been used, and often the way this is represented has proven limiting to the scope of applications supported. For example, a particular representation may be useful only for X-ray images. This paper addresses two issues: 1) representing a collection of vasculature and 2) the reconstruction of individual vasculature from images. Our representation learns the variations in branching structures and vessel shapes that occur between individuals. It supports a vascular catalogue containing three-dimensional (3-D) anatomical models. The representation is task independent; here we use it to reconstruct vasculature from images. Our algorithm has four features to which we draw attention: 1) it is not premised wholly upon X-ray images (though that is our focus here); 2) it produces several feasible solutions rather than one; 3) it can generalize from the catalogue to reconstruct instances not yet learned; 4) it exhibits polynomial time complexity, reasonable memory consumption, and is reliable. Both our representation and reconstruction algorithm are new and useful approaches. In support of these claims, we present results gathered from X-rays of both simulated and real vasculature.

Algorithms↗

How hallucinations may arise from brain mechanisms of learning, attention, and volition.

This article suggests how brain mechanisms of learning, attention, and volition may give rise to hallucinations during schizophrenia and other mental disorders. The article suggests that normal learning and memory are stabilized through the use of learned top-down expectations. These expectations learn prototypes that are capable of focusing attention upon the combinations of features that comprise conscious perceptual experiences. When top-down expectations are active in a priming situation, they can modulate or sensitize their target cells to respond more effectively to matched bottom-up information. They cannot, however, fully activate these target cells. These matching properties are shown to be essential towards stabilizing the memory of learned representations. The modulatory property of top-down expectations is achieved through a balance between top-down excitation and inhibition. The learned prototype is the excitatory on-center in this top-down network. Phasic volitional signals can shift the balance between excitation and inhibition to favor net excitatory activation. Such a volitionally mediated shift enables top-down expectations, in the absence of supportive bottom-up inputs, to cause conscious experiences of imagery and inner speech and thereby to enable fantasy and planning activities to occur. If these volitional signals become tonically hyperactive during a mental disorder, the top-down expectations can give rise to conscious experiences in the absence of bottom-up inputs and volition. These events are compared with data about hallucinations. The article predicts where these top-down expectations and volitional signals may act in the laminar circuits of visual cortex and, by extension, in other sensory and cognitive neocortical areas, and how the level of abstractness of learned prototypes may covary with the abstractness of hallucinatory content. A similar breakdown of volition may lead to delusions of control in the motor system.

Attention↗

A Graph Contrastive Learning Method for Enhancing Genome Recovery in Complex Microbial Communities.

Accurate genome binning is essential for resolving microbial community structure and functional potential from metagenomic data. However, existing approaches-primarily reliant on tetranucleotide frequency (TNF) and abundance profiles-often perform sub-optimally in the face of complex community compositions, low-abundance taxa, and long-read sequencing datasets. To address these limitations, we present MBGCCA, a novel metagenomic binning framework that synergistically integrates graph neural networks (GNNs), contrastive learning, and information-theoretic regularization to enhance binning accuracy, robustness, and biological coherence. MBGCCA operates in two stages: (1) multimodal information integration, where TNF and abundance profiles are fused via a deep neural network trained using a multi-view contrastive loss, and (2) self-supervised graph representation learning, which leverages assembly graph topology to refine contig embeddings. The contrastive learning objective follows the InfoMax principle by maximizing mutual information across augmented views and modalities, encouraging the model to extract globally consistent and high-information representations. By aligning perturbed graph views while preserving topological structure, MBGCCA effectively captures both global genomic characteristics and local contig relationships. Comprehensive evaluations using both synthetic and real-world datasets-including wastewater and soil microbiomes-demonstrate that MBGCCA consistently outperforms state-of-the-art binning methods, particularly in challenging scenarios marked by sparse data and high community complexity. These results highlight the value of entropy-aware, topology-preserving learning for advancing metagenomic genome reconstruction.

canonical correlation analysis↗

PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clustering.

The development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency.

Single-Cell Analysis↗

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics↗

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models.

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Deep Learning↗

A Knowledge-Enhanced Multimodal Framework with Genomic Reconstruction for DLBCL Drug Response Prediction.

Diffuse large B-cell lymphoma (DLBCL) exhibits substantial biological heterogeneity, leading to pronounced variability in patient response to therapy. Accurate drug response prediction is therefore critical for precision treatment but remains challenging in clinical settings where genomic sequencing, a highly informative modality, is frequently incomplete. Existing methods, often developed from cell-line pharmacogenomic datasets or single-modality data, typically assume fully observed molecular profiles and thus show limited robustness under missing genomic data. To address this limitation, a knowledge-enhanced multimodal framework with genomic reconstruction (KeM-DRP) is proposed for individualized drug response prediction in DLBCL. The framework models the central role of genomics by integrating biological prior knowledge through a gene-pathway-biological process hierarchy, enabling robust representation learning from sparse observations. To compensate for missing genomic measurements, a cross-modal genomic compensation module reconstructs genomically informed latent features from routinely available clinical modalities. Furthermore, a genomics-guided adaptive fusion strategy dynamically integrates heterogeneous modalities conditioned on observed or reconstructed genomic representation. Experiments on a real-world DLBCL cohort demonstrate that KeM-DRP consistently outperforms competitive baselines. The reconstructed genomic representation represents most predictive utility, highlighting the robustness and practical value of the framework under incomplete genomic data.

Journal Article↗

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis↗