Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

Efficient evidence-based genome annotation with EviAnn.

For many years, machine learning-based ab initio gene finding approaches have been central components of eukaryotic genome annotation pipelines, and they remain so today. The reliance on these approaches was originally sustained by the high cost and low availability of gene expression data, a primary source of evidence for gene annotation along with protein homology. However, innovations in modern sequencing technologies have revolutionized the acquisition of gene expression data, allowing scientists to rely more heavily on this class of evidence. In addition, proteins found in a multitude of well-annotated genomes represent another invaluable resource for gene annotation. Existing annotation packages often underutilize these data sources, which prompted us to develop EviAnn (Evidence-based Annotator), a novel evidence-based eukaryotic gene annotation system. EviAnn takes a strongly data-driven approach, building the exon-intron structure of genes from transcript alignments or protein-sequence homology rather than from purely ab initio gene finding techniques. We show that when provided with the same input data, EviAnn consistently outperforms current state-of-the-art packages including BRAKER3, MAKER2, and FINDER, while utilizing considerably less computer time. Annotation of a mammalian genome can be completed in less than an hour on a single multi-core server. EviAnn is freely available under an open-source license from https://github.com/alekseyzimin/EviAnn_release and from Bioconda as "eviann".

Journal Article↗

Designing smart spatial omics experiments with S2Omics.

Spatial omics technologies have transformed biomedical research by enabling high-resolution molecular profiling while preserving the native tissue architecture. These advances provide unprecedented insights into tissue structure and function. However, the high cost and time-intensive nature of spatial omics experiments necessitate careful experimental design, particularly in selecting regions of interest (ROIs) from large tissue sections. Currently, ROI selection is performed manually, which introduces subjectivity, inconsistency, and a lack of reproducibility. Previous studies have shown strong correlations between spatial molecular patterns and histological features, suggesting that readily available and cost-effective histology images can be leveraged to guide spatial omics experiments. Here, we present S2Omics, an end-to-end workflow that automatically selects ROIs from histology images with the goal of maximizing molecular information content in the ROIs. Through comprehensive evaluations across multiple spatial omics platforms and tissue types, we demonstrate that S2Omics enables systematic and reproducible ROI selection and enhances the robustness and impact of downstream biological discovery.

digital pathology↗

Automated classification of protein crystallization images using support vector machines with scale-invariant texture and Gabor features.

Protein crystallography laboratories are performing an increasing number of experiments to obtain crystals of good diffraction quality. Better automation has enabled researchers to prepare and run more experiments in a shorter time. However, the problem of identifying which experiments are successful remains difficult. In fact, most of this work is still performed manually by humans. Automating this task is therefore an important goal. As part of a project to develop a new and automated high-throughput capillary-based protein crystallography instrument, a new image-classification subsystem has been developed to greatly reduce the number of images that require human viewing. This system must have low rates of false negatives (missed crystals), possibly at the cost of raising the number of false positives. The image-classification system employs a support vector machine (SVM) learning algorithm to classify the blocks making up each image. A new algorithm to find the area within the image that contains the drop is employed. The SVM uses numerical features, based on texture and the Gabor wavelet decomposition, that are calculated for each block. If a block within an image is classified as containing a crystal, then the entire image is classified as containing a crystal. In a study of 375 images, 87 of which contained crystals, a false-negative rate of less than 4% with a false-positive rate of about 40% was consistently achieved.

Algorithms↗

Perceptual adaptive insensitivity for support vector machine image coding.

Support vector machine (SVM) learning has been recently proposed for image compression in the frequency domain using a constant epsilon-insensitivity zone by Robinson and Kecman. However, according to the statistical properties of natural images and the properties of human perception, a constant insensitivity makes sense in the spatial domain but it is certainly not a good option in a frequency domain. In fact, in their approach, they made a fixed low-pass assumption as the number of discrete cosine transform (DCT) coefficients to be used in the training was limited. This paper extends the work of Robinson and Kecman by proposing the use of adaptive insensitivity SVMs [2] for image coding using an appropriate distortion criterion [3], [4] based on a simple visual cortex model. Training the SVM by using an accurate perception model avoids any a priori assumption and improves the rate-distortion performance of the original approach.

Algorithms↗

Global convergence of decomposition learning methods for support vector machines.

Decomposition methods are well-known techniques for solving quadratic programming (QP) problems arising in support vector machines (SVMs). In each iteration of a decomposition method, a small number of variables are selected and a QP problem with only the selected variables is solved. Since large matrix computations are not required, decomposition methods are applicable to large QP problems. In this paper, we will make a rigorous analysis of the global convergence of general decomposition methods for SVMs. We first introduce a relaxed version of the optimality condition for the QP problems and then prove that a decomposition method reaches a solution satisfying this relaxed optimality condition within a finite number of iterations under a very mild condition on how to select variables.

Algorithms↗

Descriptor-based protein remote homology identification.

Here, we report a novel protein sequence descriptor-based remote homology identification method, able to infer fold relationships without the explicit knowledge of structure. In a first phase, we have individually benchmarked 13 different descriptor types in fold identification experiments in a highly diverse set of protein sequences. The relevant descriptors were related to the fold class membership by using simple similarity measures in the descriptor spaces, such as the cosine angle. Our results revealed that the three best-performing sets of descriptors were the sequence-alignment-based descriptor using PSI-BLAST e-values, the descriptors based on the alignment of secondary structural elements (SSEA), and the descriptors based on the occurrence of PROSITE functional motifs. In a second phase, the three top-performing descriptors were combined to obtain a final method with improved performance, which we named DescFold. Class membership was predicted by Support Vector Machine (SVM) learning. In comparison with the individual PSI-BLAST-based descriptor, the rate of remote homology identification increased from 33.7% to 46.3%. We found out that the composite set of descriptors was able to identify the true remote homolog for nearly every sixth sequence at the 95% confidence level, or some 10% more than a single PSI-BLAST search. We have benchmarked the DescFold method against several other state-of-the-art fold recognition algorithms for the 172 LiveBench-8 targets, and we concluded that it was able to add value to the existing techniques by providing a confident hit for at least 10% of the sequences not identifiable by the previously known methods.

Algorithms↗

A systematic approach to standardizing the visual appearance of endometriotic lesions for artificial intelligence recognition.

INTRODUCTION: Numerous studies have shown that the diagnostic performance and reproducibility of visual recognition of endometriosis during laparoscopy are poor. The use of artificial intelligence (AI) seems relevant for exhaustive lesion recognition. Standardization of the visual classification of lesions, in the form of an ontology, is an essential prerequisite to enable medical experts to annotate surgical data consistently and subsequently allow engineers to train and build an artificial intelligence tool for endometriosis recognition. MATERIAL AND METHODS: A systematic search was conducted in the MEDLINE (via PubMed), EMBASE, and the Cochrane Library databases up to May 2022, aiming to identify studies describing the laparoscopic visual appearance of superficial endometriosis, endometriomas, and deep infiltrating endometriosis. The accumulated data in the literature concerning the visual appearance of the different forms of endometriosis were used to create an ontology that could be used for artificial intelligence applications. RESULTS: Out of 932 articles screened, 35 studies were selected based on the inclusion criteria of human subjects with histologically confirmed endometriosis lesions visualized via laparoscopy. The selected studies were reviewed to develop a visual ontology of endometriosis lesions observed via laparoscopy. The lesions were categorized into 4 classes and further subdivided into 11 subclasses: superficial (black, red, white, or subtle), adhesions (dense or filmy), deep (obliteration, retraction, or deformation), and ovarian (endometrioma or chocolate fluid). The positive predictive value (PPV) varied across lesion types: black lesions (PPV 47%-97%), red lesions (PPV 33%-100%), white lesions (PPV 20%-81%), and ovarian endometriosis (PPV 42%-98%). Nonspecific lesions such as adhesions (PPV 16%-50%) and subtle superficial lesions (PPV 0%-67%) presented lower PPVs. Deep endometriosis lesions, often buried within organs, required indirect signs (obliteration, retraction, deformation) for identification. CONCLUSIONS: The visual ontology proposed in this systematic search could facilitate the detection and classification of endometriosis lesions using artificial intelligence. This study highlights the challenges of reaching a consensus on lesion recognition and classification in AI projects due to the diverse visual presentations of endometriosis.

Humans↗

Integrative Multi-Omics Analysis Identifies Thrombosis-Associated Molecular Features Linked to Germline Susceptibility and Immune Cell Communication in Gastric Cancer.

Emerging evidence indicates that coagulation-related molecular programs are associated with thrombosis, tumor progression, and molecular dysregulation in gastric cancer (GC). However, thrombosis-associated molecular features in GC and their potential links to inherited susceptibility remain insufficiently understood. Integrated analyses of transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) datasets were performed to identify thrombosis-associated genes and establish a machine learning-based prognostic signature. Genome-wide association study (GWAS), expression quantitative trait loci (eQTL), transcriptome-wide association study (TWAS), and Mendelian randomization (MR) analyses were conducted to investigate susceptibility-associated transcriptional programs in GC. Functional assays were used to evaluate candidate genes associated with malignant phenotypes. Single-cell RNA sequencing (scRNA-seq) and cell-cell communication analyses were further performed to characterize cell-type-specific expression patterns and potential intercellular interactions. A total of 22 differentially expressed thrombosis-associated genes were identified, and a prognostic signature comprising 14 genes was established. The signature stratified patients into high- and low-risk groups and showed prognostic performance in both the training and validation cohorts. Integrative GWAS, eQTL, and TWAS analyses identified susceptibility-associated transcriptional programs that were positively correlated with the thrombosis-associated risk score. Silencing ACTN2 and CRYAB significantly reduced GC cell migration and invasion. scRNA-seq analysis revealed relatively high CRYAB expression in neutrophils, and CellChat analysis suggested potential neutrophil-B cell interactions involving COLLAGEN-related signaling. This integrative multi-omics study identified a thrombosis-associated molecular signature linked to prognosis and germline susceptibility-associated transcriptional programs in GC. ACTN2 and CRYAB may represent candidate genes associated with GC cell migration and invasion, while single-cell analysis suggested potential immune-related communication features.

Humans↗

Artificial Intelligence Technologies in Nursing Clinical Decision-Making: An Umbrella Review.

AIM: To describe contemporary peer-reviewed literature on artificial intelligence in nurses' clinical decision-making. METHODS: An umbrella review of literature reviews. DATA SOURCES: Four major databases were searched for reviews published between 2019 and 2024. RESULTS: Sixteen literature reviews reported on 965 nursing artificial intelligence primary studies. The studies focused on technology development and emerging performance evaluations, whilst real-world testing or implementation in nursing clinical settings was rare. Rigorous comparative analyses were lacking. While artificial intelligence demonstrates promise in decision-making, challenges such as a lack of controlled studies, algorithmic bias, limited reproducibility and insufficient clinical trials hinder its practical impact. Ethical concerns, transparency and patient data privacy issues pose barriers to AI integration in nursing practice. Ethical and legal guidelines for patient privacy are needed and should be taught along with AI literacy training for nurses. CONCLUSIONS: Artificial intelligence has the potential to enhance clinical nursing decision-making, although evidence is limited by too few examples of nurse participation during development. Underutilisation in administrative nursing functions hinders implementation. Nurses should assume a central role in the design and development of AI applications to ensure that these technologies address the realities of nursing practice. With such improvements, artificial intelligence can transform nursing practice, improve nurses' clinical decision-making and ultimately enhance consumer healthcare outcomes. PATIENT OR PUBLIC INVOLVEMENT: No Patient or Public Involvement. REPORTING METHOD: While there is no reporting checklist for umbrella reviews, the PRISMA guide for systematic reviews was followed.

Artificial Intelligence↗

Artificial Intelligence in Diagnosing Depression Through Behavioural Cues: A Diagnostic Accuracy Systematic Review and Meta-Analysis.

AIM: To synthesise existing evidence concerning the application of AI methods in detecting depression through behavioural cues among adults in healthcare and community settings. DESIGN: This is a diagnostic accuracy systematic review. METHODS: This review included studies examining different AI methods in detecting depression among adults. Two independent reviewers screened, appraised and extracted data. Data were analysed by meta-analysis, narrative synthesis and subgroup analysis. DATA SOURCES: Published studies and grey literature were sought in 11 electronic databases. Hand search was conducted on reference lists and two journals. RESULTS: In total, 30 studies were included in this review. Twenty of which demonstrated that AI models had the potential to detect depression. Speech and facial expression showed better sensitivity, reflecting the ability to detect people with depression. Text and movement had better specificity, indicating the ability to rule out non-depressed individuals. Heterogeneity was initially high. Less heterogeneity was observed within each modality subgroup. CONCLUSIONS: This is the first systematic review examining AI models in detecting depression using all four behavioural cues: speech, texts, movement and facial expressions. IMPLICATIONS: A collaborative effort among healthcare professionals can be initiated to develop an AI-assisted depression detection system in general healthcare or community settings. IMPACT: It is challenging for general healthcare professionals to detect depressive symptoms among people in non-psychiatric settings. Our findings suggested the need for objective screening tools, such as an AI-assisted system, for screening depression. Therefore, people could receive accurate diagnosis and proper treatments for depression. REPORTING METHOD: This review followed the PRISMA checklist. PATIENTS OR PUBLIC CONTRIBUTION: No patients or public contribution.

Humans↗

Evaluating the Antigen and Eplet Accuracy of DQA1 Imputations With the HaploSFHI Two-Field HLA Typing Inference Tool.

Donor/recipient mismatched HLA antigens can lead to the production of Donor-Specific Antibodies by the recipient, which are deleterious to organ transplants. The HLA-DQ locus is the most frequent target, with both the DQ beta and alpha chains involved. For deceased donors in particular, while HLA-DQB1 has been typed in emergencies for a long time, HLA-DQA1 has only recently been included. No imputation algorithmic tool was available to impute HLA-DQA1 until the development of HaploSFHI, trained on 61,393 two-field typings by NGS methods. We evaluated the accuracy of two-field HLA-DQA1 imputation from serological and two-field level HLA-A, B, DRB1, and DQB1 typings. We report a highly accurate two-field HLA-DQA1 prediction using a French test cohort of 7696 individuals, respectively reaching 92.30% and 96.45% accuracy. The average 'False Positive eplet load' stood at 0.19 and 0.07, respectively, and the average 'False Negative eplet load' at 0.18 and 0.08, respectively. A similar performance was obtained on three independent test cohorts of European ancestry (from the USA, the UK, and Portugal). Interestingly, performance was only slightly inferior on five independent test cohorts of other ethnicities (from Hong Kong and the USA) whereas it was significantly lower for two-field DRB1 imputation from its serological level. These results suggest that DQA1 can reliably be imputed even when information is totally missing, with low error risk at both antigen and eplet levels, even if the reference population is not matched. Similar additional initiatives would be welcome to confirm these findings.

Humans↗

Analog circuits for relaxation networks.

Selected examples are presented of recent advances, primarily from the U.S. and Canada, in analog circuits for relaxation networks. Relaxation networks having feedback connections exhibit potentially greater computational power per neuron than feedforward networks. They are also more poorly understood especially with respect to learning algorithms. Examples are described of analog circuits for (i) supervised learning in deterministic Boltzmann machines, (ii) unsupervised competitive learning and feature maps and (iii) networks with resistive grids for vision and audition tasks. We also discuss recent progress on in-circuit learning and synaptic weight storage mechanisms.

Artificial Intelligence↗

Neoadjuvant Immunotherapy Promotes the Formation of Mature Tertiary Lymphoid Structures in a Remodeled Pancreatic Tumor Microenvironment.

Pancreatic ductal adenocarcinoma (PDAC) is a rapidly progressing cancer that responds poorly to immunotherapies. Intratumoral tertiary lymphoid structures (TLS) have been associated with rare long-term PDAC survivors, but the role of TLS in PDAC and their spatial relationships within the context of the broader tumor microenvironment remain unknown. In this study, we report the generation of a spatial multiomic atlas of PDAC tumors and tumor-adjacent lymph nodes from patients treated with combination neoadjuvant immunotherapies. Using machine learning-enabled hematoxylin and eosin image classification models, imaging mass cytometry, and unsupervised gene expression matrix factorization methods for spatial transcriptomics, we characterized cellular states within and adjacent to TLS spanning distinct spatial niches and pathologic responses. Unsupervised learning identified TLS-specific spatial gene expression signatures that are significantly associated with improved survival in patients with PDAC. We identified spatial features of pathologic immune responses, including intratumoral TLS-associated B-cell maturation colocalizing with IgG dissemination and extracellular matrix remodeling. Our findings offer insights into the cellular and molecular landscape of TLS in PDACs during immunotherapy treatment.

Humans↗

New Genetic Loci Implicated in Cardiac Morphology and Function Using Three-Dimensional Population Phenotyping.

BACKGROUND: Cardiac remodeling occurs in the mature heart and is a cascade of adaptations in response to stress, which are primed in early life. A key question remains as to the processes that regulate the geometry and motion of the heart and how it adapts to stress. METHODS: We performed spatially resolved phenotyping using machine learning-based analysis of cardiac magnetic resonance imaging in 47 549 UK Biobank participants. We analyzed 16 left ventricular spatial phenotypes, including regional myocardial wall thickness and systolic strain in both circumferential and radial directions. In up to 40 058 participants, genetic associations across the allele frequency spectrum were assessed using genome-wide association studies with imputed genotype participants, and exome-wide association studies and gene-based burden tests using whole-exome sequencing data. We integrated transcriptomic data from the GTEx project and used pathway enrichment analyses to further interpret the biological relevance of identified loci. To investigate causal relationships, we conducted Mendelian randomization analyses to evaluate the effects of blood pressure on regional cardiac traits and the effects of these traits on cardiomyopathy risk. RESULTS: We found 42 loci associated with cardiac structure and contractility, many of which reveal patterns of spatial organization in the heart. Whole-exome sequencing revealed 3 additional variants not captured by the genome-wide association study, including a missense variant in CSRP3 (minor allele frequency 0.5%). The majority of newly discovered loci are found in cardiomyopathy-associated genes, suggesting that they regulate spatially distinct patterns of remodeling in the left ventricle in an adult population. Our causal analysis also found regional modulation of blood pressure on cardiac wall thickness and strain. CONCLUSIONS: These findings provide a comprehensive description of the pathways that orchestrate heart development and cardiac remodeling. These data highlight the role that cardiomyopathy-associated genes have on the regulation of spatial adaptations in those without known disease.

Humans↗

An SVM-based system for predicting protein subnuclear localizations.

BACKGROUND: The large gap between the number of protein sequences in databases and the number of functionally characterized proteins calls for the development of a fast computational tool for the prediction of subnuclear and subcellular localizations generally applicable to protein sequences. The information on localization may reveal the molecular function of novel proteins, in addition to providing insight on the biological pathways in which they function. The bulk of past work has been focused on protein subcellular localizations. Furthermore, no specific tool has been dedicated to prediction at the subnuclear level, despite its high importance. In order to design a suitable predictive system, the extraction of subtle sequence signals that can discriminate among proteins with different subnuclear localizations is the key. RESULTS: New kernel functions used in a support vector machine (SVM) learning model are introduced for the measurement of sequence similarity. The k-peptide vectors are first mapped by a matrix of high-scored pairs of k-peptides which are measured by BLOSUM62 scores. The kernels, measuring the similarity for sequences, are then defined on the mapped vectors. By combining these new encoding methods, a multi-class classification system for the prediction of protein subnuclear localizations is established for the first time. The performance of the system is evaluated with a set of proteins collected in the Nuclear Protein Database (NPD). The overall accuracy of prediction for 6 localizations is about 50% (vs. random prediction 16.7%) for single localization proteins in the leave-one-out cross-validation; and 65% for an independent set of multi-localization proteins. This integrated system can be accessed at http://array.bioengr.uic.edu/subnuclear.htm. CONCLUSION: The integrated system benefits from the combination of predictions from several SVMs based on selected encoding methods. Finally, the predictive power of the system is expected to improve as more proteins with known subnuclear localizations become available.

Algorithms↗

Genetic targets related to aging for the treatment of coronary artery disease.

BACKGROUND: Coronary Artery Disease (CAD) is the most common cardiovascular disease worldwide, threatening human health, quality of life and longevity. Aging is a dominant risk factor for CAD. This study aims to investigate the potential mechanisms of aging-related genes and CAD, and to make molecular drug predictions that will contribute to the diagnosis and treatment. METHODS: We downloaded the gene expression profile of circulating leukocytes in CAD patients (GSE12288) from Gene Expression Omnibus database, obtained differentially expressed aging genes through "limma" package and GenaCards database, and tested their biological functions. Further screening of aging related characteristic genes (ARCGs) using least absolute shrinkage and selection operator and random forest, generating nomogram charts and ROC curves for evaluating diagnostic efficacy. Immune cells were estimated by ssGSEA, and then combine ARCGs with immune cells and clinical indicators based on Pearson correlation analysis. Unsupervised cluster analysis was used to construct molecular clusters based on ARCGs and to assess functional characteristics between clusters. The DSigDB database was employed to explore the potential targeted drugs of ARCGs, and the molecular docking was carried out through Autodock Vina. Finally, single-cell data (GSE159677) of arterial intima was used to further explore the expression of aging signature genes in different cell subpopulations. RESULTS: We identified 8 ARCGs associated with CAD, in which HIF1A and FGFR3 were up while NOX4, TCF7L2, HK3, CDK18, TFAP4, and ITPK1 were down in CAD patients. Based on this, CAD patients can be divided into two molecular clusters, among which cluster A mainly involves functional pathways such as ECM receptor interaction and focal adhesion; cluster B mainly involves functional pathways such as amimo sugar and nucleotide sugar metabolism and pyrimidine metabolism. In addition, the molecular docking results showed that retinoic acid and resveratrol had good binding affinity with targets genes. Further single-cell analysis results showed that NOX4, TCF7L2, ITPK1, and HIF1A were specifically expressed in different types of cells in atherosclerotic tissues. CONCLUSION: Our study identified several ARCGs that may be involved in the pathogenesis and progression of CAD. Further, retinoic acid and resveratrol were potential candidate molecule drugs for inhibiting these targets.

Humans↗

A three-metabolite microbiota-associated signature for early risk stratification of gestational diabetes mellitus.

BACKGROUND: Gestational diabetes mellitus (GDM) is associated with adverse pregnancy outcomes and long-term metabolic and cardiovascular risk. However, oral glucose tolerance testing at 24-28 gestational weeks limits early risk stratification. Gut microbiota-associated metabolites may reflect early metabolic abnormalities, including those relevant to cardiometabolic health, but robust early-pregnancy biomarkers remain limited. METHODS: We conducted a multicenter nested case-control and prospective study involving 2,693 pregnant women. Untargeted metabolomics and metagenomics were integrated to identify GDM-associated metabolites and gut microbial alterations. Three consistently dysregulated metabolites, 3-hydroxydecanoic acid, γ-Glu-Leu, and propionic acid, were quantified by targeted LC-MS/MS. Candidate algorithms were compared using repeated 10-fold cross-validation, and a final generalized linear model was externally and prospectively validated. RESULTS: Women who later developed GDM showed an adverse early-pregnancy metabolic profile, including higher BMI, triglycerides, and platelet count. Untargeted metabolomics identified 14 persistently altered metabolites enriched in energy, oxidative stress, and amino acid metabolism pathways. Metagenomics revealed taxonomic restructuring and coordinated microbiota-metabolite associations. The three-metabolite model achieved AUCs of 0.838 (95% CI, 0.791-0.885) in training, 0.840 (95% CI, 0.769-0.911) in internal validation, 0.955 (95% CI, 0.925-0.985) and 0.917 (95% CI, 0.875-0.958) in two external cohorts, and 0.969 (95% CI, 0.937-1.000) in the prospective cohort. CONCLUSION: Early microbiota-associated metabolic dysregulation is detectable before routine GDM diagnosis. This compact three-metabolite panel may support early GDM risk stratification and provides metabolic evidence relevant to broader cardiometabolic risk assessment in pregnancy.

Humans↗

Cancer of unknown primary: the evolution of tissue of origin identification in the artificial intelligence era.

Cancer of Unknown Primary (CUP) presents substantial diagnostic and therapeutic challenges owing to its heterogeneous nature and the absence of an identifiable primary tumor site. This review provides a structured search of the pathogenesis, epidemiological characteristics, and limitations of traditional diagnostic and therapeutic approaches for CUP, with an emphasis on the evolution of Tissue of Origin (TOO) identification techniques. Recent advances in precision medicine have accelerated the development of machine learning-based TOO identification tools, representing a paradigm shift in CUP diagnostics. Deep learning (DL) algorithms that integrate multi-omics data (such as genomics and transcriptomics) with clinical features have markedly enhanced the accuracy of tracing tumor origin, and artificial intelligence (AI) driven TOO models are increasingly being incorporated into clinical practice, offering new insights for pathological diagnosis, treatment selection, and prognostic evaluation. Nevertheless, several challenges remain, including issues of data standardization, model generalizability, and interpretability. Ethical considerations related to data privacy, algorithmic fairness, and clinical implementation also warrant careful attention. Future research should focus on establishing standardized multi-center databases, developing more interpretable AI models, and fostering multidisciplinary collaborative strategies for CUP management. Through continued refinement of technical solutions and regulatory guidelines, TOO identification is anticipated to progress from research to routine clinical application, ultimately supporting precise and personalized care for patients with CUP.

Artificial intelligence↗