Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

Overcoming Artificial Structures in Resolution-Enhanced Hi-C Data by Signal Decomposition and Multi-Scale Attention.

Computational enhancement is an important strategy for inferring high-resolution features from genome-wide chromosome conformation capture (Hi-C) data, which typically have limited resolution. Deep learning has been highly successful in this task but we show that it creates prevalent artificial structures in the enhanced data due to the need to divide the large contact matrix into small patches. In addition, previous deep learning methods largely focus on local patterns, which cannot fully capture the complexity of Hi-C data. Here we propose Smooth, High-resolution, and Accurate Reconstruction of Patterns (SHARP) for enhancing Hi-C data. It uses the novel approach of decomposing the data into three types of signals, due to one-dimensional proximity, contiguous domains, and other fine structures, respectively, and applies deep learning only to the third type of signals, such that enhancement of the first two is unaffected by the patches. For the deep learning part, SHARP uses both local and global attention mechanisms to capture multi-scale contextual information. We compare SHARP with state-of-the-art methods extensively, including application to data from new samples and another species, and show that SHARP has superior performance in terms of resolution enhancement accuracy, avoiding creation of artificial structures, identifying significant interactions, and enrichment in chromatin states.

Hi‐C↗

CDACHIE: chromatin domain annotation by integrating chromatin interaction and epigenomic data with contrastive learning.

MOTIVATION: Chromatin domain annotation identifies functional genomic regions, such as active and inactive zones, based on epigenomic features like histone modifications, DNA methylation, and chromatin accessibility. While recent methods have utilized both chromatin interaction data (e.g. Hi-C) and epigenomic data, they often overlook the direct relationship between these data types. RESULTS: In this study, we introduce Chromatin Domain Annotation using Contrastive Learning for Hi-C and Epigenomic Data (CDACHIE), a method for identifying chromatin domains from Hi-C and epigenomic data. Our approach leverages contrastive learning to generate aligned representative vectors for both data types at each genomic bin. The concatenated vectors are then clustered using K-means to classify distinct chromatin domain types. CDACHIE achieves superior performance in Variance Explained, evaluated across gene expression, replication timing, and ChIA-PET data. This highlights its robust ability to integrate semantic associations between Hi-C and epigenomic features within the embedding space. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub: https://github.com/maruyama-lab-design/CDACHIE. An archival snapshot of the code used in this study is available on Zenodo: https://doi.org/10.5281/zenodo.15751780.

Chromatin↗

SIGEL: a context-aware genomic representation learning framework for spatial genomics analysis.

Spatial transcriptomics (ST) integrates spatial information into genomics, yet methods for generating spatially-informed gene representations are limited and computationally intensive. We present SIGEL, a cost-effective framework that derives gene manifolds from ST data by exploiting spatial genomic context. The resulting SIGEL-generated gene representations (SGRs) are context-aware, biologically meaningful, and robust across samples, making them highly effective for key downstream tasks, including imputing missing genes, detecting spatial expression patterns, identifying disease-related genes and interactions, and improving spatial clustering. Extensive experiments across diverse ST datasets validate SIGEL's effectiveness and highlight its potential in advancing spatial genomics research.

Genomics↗

MetaChrome: An Open-Source, User-Friendly Tool for Automated Metaphase Chromosome Analysis.

DNA Fluorescence In Situ Hybridization (FISH) is an essential technique to study chromosome biology and genetics, enabling precise visualization of specific genomic loci to study structural abnormalities, gene mapping, and chromosomal rearrangements. High-Throughput Imaging (HTI) can automate the analysis of DNA-FISH chromosome images, but the accurate and automated segmentation of mitotic chromosomes and simultaneous colocalization of FISH signals remains a challenge. While several commercial automated karyotyping tools partially solve these issues, open-source software that effectively combines robust chromosome segmentation with comprehensive colocalization analysis capabilities remains necessary. To address this unmet need, we developed MetaChrome, an open-source software platform built around a graphical user interface and explicitly designed for automated metaphase chromosome analysis. MetaChrome leverages fine-tuned deep learning models to automate metaphase chromosome segmentation, together with colocalization analysis of chromosome-specific FISH probes and immunofluorescent-labeled proteins. Importantly, MetaChrome achieves enhanced segmentation accuracy compared to traditional image processing methods by adopting a Cellpose segmentation model fine-tuned with manually annotated metaphase chromosome datasets. The fine-tuned model ensures precise assignment of DNA-FISH spots to individual chromosomes in an automated manner. This facilitates rapid identification of chromosomal abnormalities, reduces human error, and advances high-throughput chromosome analysis workflows, addressing a key bottleneck in chromosome biology research.

Chromosome segmentation↗

Common to rare transfer learning (CORAL) enables inference and prediction for a quarter million rare Malagasy arthropods.

DNA-based biodiversity surveys result in massive-scale data, including up to millions of species-of which, most are rare. Making the most of such data for inference and prediction requires modeling approaches that can relate species occurrences to environmental and spatial predictors, while incorporating information about their taxonomic or phylogenetic placement. Even if the scalability of joint species distribution models to large communities has greatly advanced, incorporating hundreds of thousands of species has not been feasible to date, leading to compromised analyses. Here we present a 'common to rare transfer learning' (CORAL) approach, based on borrowing information from the common species to enable statistically and computationally efficient modeling of both common and rare species. We illustrate that CORAL leads to much improved prediction and inference in the context of DNA metabarcoding data from Madagascar, comprising 255,188 arthropod species detected in 2,874 samples.

Animals↗

Fairness-aware supervised hierarchical contrastive semantic learning for sexual dimorphism analysis.

MOTIVATION: Sexual dimorphism is a fundamental biological determinant driving systematic differences in disease susceptibility, progression, and clinical outcomes. However, current sex-combined AI-based genomic models often exhibit algorithmic bias and fail to capture these sex-specific mechanisms, creating a critical barrier to unbiased precision medicine. Ensuring fairness in the context of sexual dimorphism requires understanding and addressing the distinct biological mechanisms functioning in each sex, rather than focusing solely on equalizing predictive performance. RESULTS: We propose a fairness-aware supervised hierarchical contrastive learning approach, called FairHICON, to discover unbiased sex-common and sex-specific predictive features. Evaluations on cancer and asthma transcriptomic datasets demonstrate that FairHICON significantly outperforms state-of-the-art benchmarks, improving predictive performance by up to 9% while effectively reducing the performance gap between male and female sexes. Furthermore, prognostic validation confirms that the identified sex-specific pathways stratify patient survival significantly better within their corresponding sex groups. This validates FairHICON to elucidate the molecular heterogeneity of sexual dimorphism, advancing inclusive precision medicine. AVAILABILITY AND IMPLEMENTATION: The source code and data is available at https://github.com/datax-lab/FairHICON.

Sex Characteristics↗

Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction.

Predicting spatial gene expression from Histological images is a fundamental task in understanding tissue organization and molecular phenotypes. However, existing methods often rely on single-model representations or lack effective alignment between image and transcriptomic features. To address these limitations, we propose a unified multimodal learning framework that integrates histological imaging and spatial transcriptomics through a shared latent representation space. Specifically, histological H&E images are encoded by a ResNet50-based convolutional stem and a MobileViT Transformer backbone to extract hierarchical visual representations. Both modalities are projected into a shared latent space via linear-GELU-dropout transformation blocks, enabling cross-modal alignment through a contrastive learning objective that maximizes agreement between the corresponding image and the spot embeddings. Experimental results on the 10x Genomics Visium dataset of human liver tissue demonstrate that MViTGene achieves significantly higher prediction accuracy than existing methods across multiple gene subsets, with improvements of 20%, 33%, and 12% in predicting marker genes, highly expressed genes, and highly variable genes, respectively. The significant improvement in relevance indicates that the model can more accurately capture the true correspondence between tissue morphology and gene expression, therefore enabling more reliable biological interpretation. It provides a computational tool for high-throughput spatial gene expression prediction that balances performance and interpretability.

Humans↗

CSGL: chemical synthesis graph learning for molecule representation.

MOTIVATION: Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. RESULTS: Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-2023/CSGL.

Machine Learning↗

A Meta-learning-driven strategy for adulteration detection in sweet potato starch and vermicelli using Raman spectroscopy.

To address the widespread adulteration of sweet potato starch and its vermicelli with cheaper starches and overcome conventional supervised learning's dependency on large labeled datasets, this study developed a few-shot discrimination method integrating Raman spectroscopy with meta-learning. We constructed a meta-learning framework using cassava- and wheat-adulterated sweet potato starch as the source domain for training, with potato-adulterated sweet potato starch and cassava-adulterated sweet potato vermicelli as two target domains for testing. Raman spectra showed high consistency between sweet potato vermicelli and its raw starch, laying the foundation for cross-domain detection. Testing yielded comprehensive classification accuracies of 95.33% and 98.00% for the two target domains, significantly outperforming SVM, RF, and CNN (max. 85.24%). This approach effectively identifies subtle starch variety differences in complex adulteration, providing novel food quality inspection solutions and verifying the feasibility of raw material-to-finished product cross-domain detection.

Ipomoea batatas↗

EPIC: Event Prototyping via Information Constrained graph learning for personalized cancer driver gene prediction.

MOTIVATION: Precision oncology relies on accurately distinguishing patient-specific driver mutations from the vast background of passenger alterations. While graph-based computational methods have emerged as powerful tools for this task, they often struggle to preserve the distinct genomic context of individual mutations within complex biological networks. Consequently, subtle patient-specific driver signals are frequently obscured by dominant topological patterns, critically impeding the identification of individualized oncogenic events essential for personalized cancer therapy. RESULTS: To address this, we propose EPIC, a novel framework for Event Prototyping via Information Constrained Graph Learning. Unlike traditional node-centric approaches, EPIC redefines driver prediction as a metric learning task in an event embedding space. We introduce an information-constrained learning strategy that imposes explicit geometric constraints on feature variance, effectively preventing feature collapse and ensuring that low-frequency driver signals are distinctively preserved. Experiments on large-scale cancer cohorts demonstrate that EPIC significantly outperforms established baselines. Notably, the model prioritizes low-frequency driver variants typically overlooked by population-based methods, mapping them to critical oncogenic mechanisms associated with drug resistance and metastasis. Furthermore, clinical actionability analysis confirms that EPIC substantially expands the patient population eligible for targeted therapies. EPIC provides a robust and context-aware solution for personalized cancer driver discovery, bridging the gap between genomic data and actionable therapeutic insights. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/spcho-dev/EPIC.

Humans↗

The Continuity Trap in Data Science Health Research.

Secondary use is now the ordinary condition of data science health research rather than an exception to it. Electronic health records collected for clinical care become prediction tools and inputs for generative AI; imaging archives become foundation-model corpora; genomic datasets become resources for polygenic risk scores; and legacy biospecimens become renewable, indefinitely distributable cell lines. Governance has responded by emphasizing verifiable instruments such as provenance logs, repository approvals, broad-consent forms, data-use agreements, model cards, records of processing, and locality-preserving architectures. These instruments are necessary, and they answer real questions about lineage, privacy, institutional responsibility, and accountability, but they are not sufficient to establish that a present use remains ethically justified. We define ethical continuity as the persistence of normatively relevant relationships between the original conditions of data generation or material collection and subsequent downstream uses, such that current uses remain justifiable in light of the expectations, permissions, meanings, and relational obligations present at entrustment. We then define the Continuity Trap as a review-stage governance error in which a salient signal of continuity in one domain is treated as sufficient evidence of ethical continuity overall, causing inquiry into the remaining domains to close prematurely. The trap is not ordinary noncompliance, ethics creep, or a demand for universal rereview; it is a cross-domain inference error that can arise even in careful, good-faith review. We distinguish it from proxy closure, of which it is a continuity-specific subtype, and from Goodhart's and Campbell's laws, which describe how measures degrade once they become targets. We operationalize ethical continuity across 4 domains: provenance, semantics, authorization, and relational standing, developed in our Representational Veracity framework, and we show that these domains can diverge as data are linked, transformed, modeled, and redeployed. We identify the institutional mechanisms-provenance privilege, descriptor sedimentation, authorization fossilization, and community effacement-that cause auditable signals to be overread, and we examine how the US Health Insurance Portability and Accountability Act (HIPAA) of 1996, the General Data Protection Regulation, the European Health Data Space, US Food and Drug Administration guidance, the US National Institute of Standards and Technology (NIST) AI Risk Management Framework, and federated-learning governance can reduce risk while still inducing continuity traps. We apply the framework to consent and nonconsent settings, including public health, immunization, syndromic, and wastewater surveillance, polygenic risk scores, induced pluripotent stem cells, federated learning, and health-related large language models. The policy implication is trigger-based continuity review: rather than rereviewing every reuse, investigators and reviewers should identify the weakest continuity domain at the present data stage and impose a domain-matched safeguard, recorded in a short continuity statement. This reframing is intended for the committees, repositories, funders, and governance bodies that decide whether reuse may proceed, and it matters most in cross-border and low-resource settings. Provenance should begin ethical review; it should not end it.

Data Science↗

PharaCon: a new framework for identifying bacteriophages via conditional representation learning.

MOTIVATION: Identifying bacteriophages (phages) within metagenomic sequences is essential for understanding microbial community dynamics. Transformer-based foundation models have been successfully employed to address various biological challenges. However, these models are typically pre-trained with self-supervised tasks that do not consider label variance in the pre-training data. This presents a challenge for phage identification as pre-training on mixed bacterial and phage data may lead to information bias due to the imbalance between bacterial and phage samples. RESULTS: To overcome this limitation, we proposed a novel conditional BERT framework that incorporates label classes as special tokens during pre-training. Specifically, our conditional BERT model attaches labels directly during tokenization, introducing label constraints into the model's input. Additionally, we introduced a new fine-tuning scheme that enables the conditional BERT to be effectively utilized for classification tasks. This framework allows the BERT model to acquire label-specific contextual representations from mixed sequence data during pre-training and applies the conditional BERT as a classifier during fine-tuning, and we named the fine-tuned model as PharaCon. We evaluated PharaCon against several existing methods on both simulated sequence datasets and real metagenomic contig datasets. The results demonstrate PharaCon's effectiveness and efficiency in phage identification, highlighting the advantages of incorporating label information during both pre-training and fine-tuning. AVAILABILITY AND IMPLEMENTATION: The source code and associated data can be accessed at https://github.com/Celestial-Bai/PharaCon.

Bacteriophages↗

Automated Deep Learning-Based Detection of Early Atherosclerotic Plaques in Carotid Ultrasound Imaging.

BACKGROUND: Carotid plaque presence is associated with cardiovascular risk, even among asymptomatic individuals. While deep learning has shown promise for carotid plaque phenotyping in patients with advanced atherosclerosis, its application in population-based settings of asymptomatic individuals remains unexplored. METHODS: We developed a YOLOv8-based model for plaque detection using carotid ultrasound images from 19,499 participants of the population-based UK Biobank (UKB) and fine-tuned it for external validation in the BiDirect study (N = 2,105). Cox regression was used to estimate the impact of plaque presence and count on major cardiovascular events. To explore the genetic architecture of carotid atherosclerosis, we conducted a genome-wide association study (GWAS) meta-analysis of the UKB and CHARGE cohorts. Mendelian randomization (MR) assessed the effect of genetic predisposition to vascular risk factors on carotid atherosclerosis. RESULTS: Our model demonstrated high performance with accuracy, sensitivity, and specificity exceeding 85%, enabling identification of carotid plaques in 45% of the UKB population (aged 47-83 years). In the external BiDirect cohort, a fine-tuned model achieved 86% accuracy, 78% sensitivity, and 90% specificity. Plaque presence and count were associated with risk of major adverse cardiovascular events (MACE) over a follow-up of up to seven years, improving risk reclassification beyond the Pooled Cohort Equations. A GWAS meta-analysis of carotid plaques uncovered two novel genomic loci, with downstream analyses implicating targets of investigational drugs in advanced clinical development. Observational and MR analyses showed associations between smoking, LDL cholesterol, hypertension, and odds of carotid atherosclerosis. CONCLUSIONS: Our model offers a scalable solution for early carotid plaque detection, potentially enabling automated screening in asymptomatic individuals and improving plaque phenotyping in population-based cohorts. This approach could advance large-scale atherosclerosis research.

atherosclerosis↗

Cross-tissue immune profiling of APOE ε4 reveals early dysregulation in Alzheimer's disease.

INTRODUCTION: Apolipoprotein E (APOE) ε4 is the strongest genetic risk factor for late-onset Alzheimer's disease (AD), but its contribution to disease pathogenesis remains incompletely understood. METHODS: Here, we integrate proteomic profiling of plasma (n = 9028), cerebrospinal fluid (n = 1099), dorsolateral prefrontal cortex (n = 720), and superior temporal gyrus (n = 105) to define the immune phenotype associated with APOE ε4. RESULTS: We identify a conserved, allele dose-dependent pro-inflammatory immune protein signature across peripheral and central tissues independent of AD diagnosis. This signature also emerges in patient-derived cortical organoids prior to amyloid beta and tau pathology, supporting a genotype-driven mechanism. Cross-tissue comparisons reveal shared innate and antiviral responses alongside tissue-specific immune signaling. Notably, a 12-week medical ketogenic diet partially reversed the APOE ε4 immune signature. DISCUSSION: These findings position immune dysregulation as an early and tractable driver of AD risk in APOE ε4 carriers with direct implications for targeted prevention strategies.

Humans↗

Immune-Like Malignant Epithelial Programs Shape Tumor-Immune Interactions and Inform Prognostic Stratification in Lung Adenocarcinoma.

Lung adenocarcinoma (LUAD) is characterized by marked cellular heterogeneity, yet how malignant epithelial states contribute to immune regulation and clinical outcomes remains incompletely defined. We integrated single-cell RNA-sequencing data to map the cellular landscape of LUAD and identify malignant epithelial cells based on inferred copy-number alterations. Epithelial states were further examined through trajectory inference, transcription factor analysis, and cell-cell communication profiling. Single-cell-derived genes were subsequently integrated with TCGA and independent GEO cohorts to construct and validate a machine learning-based prognostic signature. Malignant epithelial cells displayed distinct functional programs, including an immune-like state associated with genomic instability, immune-related transcriptional activity, tumor-immune communication, and patient outcomes. The resulting immune-like malignant epithelial cell signature (IMEC-Sig) consistently stratified survival across multiple cohorts. Low IMEC-Sig scores were accompanied by greater immune infiltration, higher immune checkpoint expression, and increased immunophenoscore, whereas high scores were linked to a comparatively immunosuppressive phenotype. Pan-cancer analyses further identified KRT8 as a gene associated with unfavorable prognosis, and functional experiments showed that KRT8 silencing suppressed proliferation, migration, invasion, and colony formation in LUAD cells. Together, these findings connect malignant epithelial heterogeneity with the immune context and clinical outcomes, support IMEC-Sig as a biologically informed prognostic tool, and nominate KRT8 as a potential therapeutic target in LUAD.

Humans↗

Nanocarrier-Based Gene Delivery Systems: Mechanisms, Clinical Translation, and Future Perspectives.

Gene therapy holds revolutionary potential for managing genetic disorders, cancers and infectious illnesses. However, one of the biggest challenges is delivering DNA or RNA into targeted cells and in the safe and effective way. In this review, nano carrier-based approaches for gene delivery are critically examined, focusing on both viral and non-viral systems. The advancement of CRISPR-Cas genome editing, machine learning-assisted nanocarrier optimization, and biologically inspired delivery systems is being quickly pushed forward in this area. In this review, a comparative analysis of gene delivery systems is being provided, and the key challenges to clinical translation are being pointed out. In addition, expert opinions on future research directions are being offered, with a heavy focus on the development of multifunctional, precisely targeted, and easily scalable delivery systems that can be integrated with next-generation therapeutic technologies.

Humans↗

Accurate prediction of solvent accessibility using neural networks-based regression.

Accurate prediction of relative solvent accessibilities (RSAs) of amino acid residues in proteins may be used to facilitate protein structure prediction and functional annotation. Toward that goal we developed a novel method for improved prediction of RSAs. Contrary to other machine learning-based methods from the literature, we do not impose a classification problem with arbitrary boundaries between the classes. Instead, we seek a continuous approximation of the real-value RSA using nonlinear regression, with several feed forward and recurrent neural networks, which are then combined into a consensus predictor. A set of 860 protein structures derived from the PFAM database was used for training, whereas validation of the results was carefully performed on several nonredundant control sets comprising a total of 603 structures derived from new Protein Data Bank structures and had no homology to proteins included in the training. Two classes of alternative predictors were developed for comparison with the regression-based approach: one based on the standard classification approach and the other based on a semicontinuous approximation with the so-called thermometer encoding. Furthermore, a weighted approximation, with errors being scaled by the observed levels of variability in RSA for equivalent residues in families of homologous structures, was applied in order to improve the results. The effects of including evolutionary profiles and the growth of sequence databases were assessed. In accord with the observed levels of variability in RSA for different ranges of RSA values, the regression accuracy is higher for buried than for exposed residues, with overall 15.3-15.8% mean absolute errors and correlation coefficients between the predicted and experimental values of 0.64-0.67 on different control sets. The new method outperforms classification-based algorithms when the real value predictions are projected onto two-class classification problems with several commonly used thresholds to separate exposed and buried residues. For example, classification accuracy of about 77% is consistently achieved on all control sets with a threshold of 25% RSA. A web server that enables RSA prediction using the new method and provides customizable graphical representation of the results is available at http://sable.cchmc.org.

Artificial Intelligence↗

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8 billion predictions for 1.4 million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein↗