Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Penalised regression improves imputation of cell-type specific expression using RNA-seq data from mixed cell populations compared to domain-specific methods.

Gene expression studies often use bulk RNA sequencing of mixed cell populations because single cell or sorted cell sequencing may be prohibitively expensive. However, mixed cell studies may miss expression patterns that are restricted to specific cell populations. Computational deconvolution can be used to estimate cell fractions from bulk expression data and infer average cell-type expression in a set of samples (e.g., cases or controls), but imputing sample-level cell-type expression is required for more detailed analyses, such as relating expression to quantitative traits, and is less commonly addressed. Here, we assessed the accuracy of imputing sample-level cell-type expression using a real dataset where mixed peripheral blood mononuclear cells (PBMC) and sorted (CD4, CD8, CD14, CD19) RNA sequencing data were generated from the same subjects (N=158), and pseudobulk datasets synthesised from eQTLgen single cell RNA-seq data. We compared three domain-specific methods, CIBERSORTx, bMIND and debCAM/swCAM, and two cross-domain machine learning methods, multiple response LASSO and ridge, that had not been used for this task before. We also assessed the methods according to their ability to recover differential gene expression (DGE) results. LASSO/ridge showed higher sensitivity but lower specificity for recovering DGE signals seen in observed data compared to deconvolution methods, although LASSO/ridge had higher area under curves than deconvolution methods. Machine learning methods have the potential to outperform domain-specific methods when suitable training data are available.

Humans↗

Proteomics uncovers ICAM2 (CD102) as a novel serum biomarker of proliferative lupus nephritis.

OBJECTIVES: This study aimed to identify novel, non-invasive biomarkers for lupus nephritis (LN) through serum proteomics. METHODS: Serum proteins were detected in patients with LN and healthy control (HC) groups through liquid chromatography-tandem mass spectrometry. The key networks associated with LN were screened out using Cytoscape software, followed by pathway enrichment analysis. The best candidate biomarkers were selected by machine learning models, further validated in a larger independent cohort. Finally, the expression of these candidate markers was verified in kidney tissue samples, and the mechanism was explored by knocking down the expression of intercellular adhesion molecule 2 (ICAM2) through in vitro cell transfection with siRNA. RESULTS: Following the serum proteomic screening of LN, a key network of 20 proteins was identified. Machine learning models were used to select ICAM2 (CD102), metalloproteinase inhibitor 1 (TIMP1) and thrombospondin 1 (THSB1) for validation in independent cohorts. ICAM2 exhibited the highest area under the curve (AUC) value in distinguishing LN from HC (AUC=0.92) and was significantly correlated with activity index, proteinuria, albumin and anti-dsDNA antibody levels. Particularly, ICAM2 was significantly elevated in proliferative LN and was associated with specific pathological attributes, outperforming conventional parameters in distinguishing proliferative LN from non-proliferative LN. ICAM2 expression was also elevated in renal tissue samples from patients with proliferative LN. In vitro, knockdown of ICAM2 expression can inhibit the activation of the PI3K/Akt pathway and alleviate the injury of glomerular endothelial cells. CONCLUSION: ICAM2 (CD102) may serve as a potential serum biomarker for proliferative LN that reflects renal pathology activity, potentially contributing to the progression of LN through the PI3K/Akt pathway.

Humans↗

SSB deficiency-induced R-loop accumulation triggers podocyte inflammation in DKD.

INTRODUCTION: Diabetic kidney disease (DKD) is fundamentally a podocytopathy in which sterile inflammation plays a central pathogenic role, yet the upstream triggers that initiate inflammatory cascades in podocytes remain elusive. R-loops are critical regulators of genomic stability, and their pathological accumulation triggers DNA damage and innate immune activation. Whether R-loop dysregulation contributes to podocyte-driven inflammation in DKD is unknown. METHODS: We integrated single-cell transcriptomic profiling, dual machine learning algorithms, and functional experiments to dissect the R-loop regulatory network in the diabetic kidney. RESULTS: Integrated analysis of human diabetic kidney single-cell RNA-seq data revealed a globally compromised R-loop regulatory network selectively within podocytes. Intersection of podocyte-specific transcriptomic shifts with validated R-loop regulators identified 93 candidate genes, from which dual machine learning algorithms pinpointed SSB (Sjögren syndrome antigen B) as the principal podocyte-selective R-loop resolver and a superior diagnostic biomarker (AUC = 0.983). SSB expression was selectively downregulated in diabetic podocytes and showed the strongest positive correlation with the R-loop resolution module. Mechanistically, SSB loss impaired RNA splicing and stability pathways, leading to aberrant R-loop accumulation that activated the cGAS-dependent inflammatory signaling in podocytes. In two murine DKD models and high glucose-challenged podocytes, SSB was markedly reduced. Remarkably, SSB knockdown in podocytes alone sufficed to trigger R-loop accumulation and pro-inflammatory cytokine expression, whereas both RNase H1-mediated R-loop removal and cGAS co-depletion blunted this response. DISCUSSION: These findings suggest that an SSB-governed R-loop -cGAS -inflammatory signaling axis may link genomic instability to podocyte inflammation and contribute to DKD progression, nominating R-loop homeostasis as a previously unrecognized potential therapeutic target.

Podocytes↗

Identification of NR4A2 as a Potential Predictive Biomarker for Atherosclerosis.

INTRODUCTION/OBJECTIVE: Atherosclerosis, a leading cause of death globally, is characterized by the buildup of immune cells and lipids in medium to large-sized arteries. However, its precise mechanism remains unclear. The purpose of this study is to explore innovative and reliable biomarkers as a viable approach for the identification and management of atherosclerosis. METHODS: The atherosclerosis-related datasets GSE100927 and GSE66360 were retrieved from the Gene Expression Omnibus (GEO) database. The Limma package in the R programming language was utilized, applying the criteria of |logFC| > 1 and P < 0.05. Subsequently, Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed on the 127 identified DEGs using R. Machine learning techniques were then applied to these data to explore and pinpoint potential biomarkers. The diagnostic potential of these markers was assessed via Receiver Operating Characteristic (ROC) curve analysis. Finally, western blot, real-time quantitative PCR (qRT-PCR), and immunohistochemistry (IHC) were employed to confirm the key biomarkers. RESULTS: Our research indicated that a total of 127 DEGs linked to atherosclerosis were successfully identified. Through the application of machine learning methods, eight critical genes were highlighted. Among these, Nuclear Receptor Subfamily 4 Group A Member-2 (NR4A2) emerged as the most promising marker for further investigation. CIBERSORT analysis revealed that NR4A2 expression levels were significantly correlated with multiple immune cell types, including B cells, plasma cells, and macrophages. Additional validation experiments confirmed that NR4A2 expression was indeed elevated in atherosclerotic plaques, supporting its potential as a biomarker for atherosclerosis. CONCLUSION: Our study identified NR4A2 as a potential immune-related biomarker for the diagnosis and treatment of atherosclerosis.

Atherosclerosis↗

CCT2 defines a highly cisplatin-resistant and poor-prognosis subtype of lung adenocarcinoma.

Cisplatin-based chemotherapy is a standard treatment for lung adenocarcinoma (LUAD), yet acquired cisplatin resistance remains a marked cause of treatment failure. The molecular mechanisms driving cisplatin resistance in LUAD have not been fully elucidated. The present study integrated bulk transcriptomic data, genomic mutation profiles and single-cell RNA sequencing data to systematically investigate cisplatin resistance in LUAD. Resistance-associated genes were identified through differential expression, survival analysis and database integration. Unsupervised clustering was used to define cisplatin resistance-associated subtypes. Functional characteristics were explored using pathway enrichment, immune infiltration, tumor mutation burden and weighted gene co-expression network analysis. A machine learning framework incorporating 101 algorithms was applied to identify key genes and construct a prognostic model. Single-cell analyses and in vitro experiments were performed to validate the biological role of the core gene. Molecular docking and molecular dynamics simulations were conducted to identify potential therapeutic compounds. A total of two molecular subtypes with distinct cisplatin resistance levels and prognostic outcomes were identified. The high-resistance subtype exhibited enhanced cell cycle activity, DNA repair signaling and immune heterogeneity. Machine learning analysis revealed a five-gene signature, with chaperonin-containing TCP1 subunit 2 (CCT2) emerging as a key regulator of cisplatin resistance. Single-cell analyses showed that CCT2 was predominantly enriched in resistant epithelial cell subpopulations. Functional experiments demonstrated that CCT2 knockdown significantly inhibited cell proliferation and enhanced cisplatin sensitivity in LUAD cell lines. A number of candidate compounds targeting CCT2 exhibited stable binding in silico. The present findings identified CCT2 as a key mediator of cisplatin resistance in LUAD and provided potential therapeutic strategies to overcome chemotherapy resistance.

chaperonin-containing TCP-1 subunit 2↗

In silico prediction method for plant Nucleotide-binding leucine-rich repeat- and pathogen effector interactions.

Plant Nucleotide-binding leucine-rich repeat (NLR) proteins play a crucial role in effector recognition and activation of Effector triggered immunity following pathogen infection. Genome sequencing advancements have led to the identification of a myriad of NLRs in numerous agriculturally important plant species. However, deciphering which NLRs recognize specific pathogen effectors remains challenging. Predicting NLR-effector interactions in silico will provide a more targeted approach for experimental validation, critical for elucidating function, and advancing our understanding of NLR-triggered immunity. In this study, NLR-effector protein complex structures were predicted using AlphaFold2-Multimer for all experimentally validated NLR-effector interactions reported in literature. Binding affinities- and energies were predicted using 97 machine learning models from Area-Affinity. We show that AlphaFold2-Multimer predicted structures have acceptable accuracy and can be used to investigate NLR-effector interactions in silico. Binding affinities for 58 NLR-effector complexes ranged between -8.5 and -10.6 log(K), and binding energies between -11.8 and -14.4&#x2009;kcal/mol-1, depending on the Area-Affinity model used. For 2427 "forced" NLR-effector complexes, these estimates showed larger variability, enabling identification of novel NLR-effector interactions with 99% accuracy using an Ensemble machine learning model. The narrow range of binding energies- and affinities for "true" interactions suggest a specific change in Gibbs free energy, and thus conformational change, is required for NLR activation. This is the first study to provide a method for predicting NLR-effector interactions, applicable to all pathosystems. Finally, the NLR-Effector Interaction Classification (NEIC) resource can streamline research efforts by identifying NLRs important for plant-pathogen resistance, advancing our understanding of plant immunity.

Plant Proteins↗

Chromatin structures from integrated AI and polymer physics model.

The physical organization of the genome in three-dimensional space regulates many biological processes, including gene expression and cell differentiation. Three-dimensional characterization of genome structure is critical to understanding these biological processes. Direct experimental measurements of genome structure are challenging; computational models of chromatin structure are therefore necessary. We develop an approach that combines a particle-based chromatin polymer model, molecular simulation, and machine learning to efficiently and accurately estimate chromatin structure from indirect measures of genome structure. More specifically, we introduce a new approach where the interaction parameters of the polymer model are extracted from experimental Hi-C data using a graph neural network (GNN). We train the GNN on simulated data from the underlying polymer model, avoiding the need for large quantities of experimental data. The resulting approach accurately estimates chromatin structures across all chromosomes and across several experimental cell lines despite being trained almost exclusively on simulated data. The proposed approach can be viewed as a general framework for combining physical modeling with machine learning, and it could be extended to integrate additional biological data modalities. Ultimately, we achieve accurate and high-throughput estimations of chromatin structure from Hi-C data, which will be necessary as experimental methodologies, such as single-cell Hi-C, improve.

Chromatin↗

Subphenogroups of acute heart failure with preserved ejection fraction: comprehensive proteomics and pathway analysis.

BACKGROUND: Heterogeneity of heart failure with preserved ejection fraction (HFpEF) results in significant challenges for treatment development. Identifying and characterising distinct HFpEF phenogroups may aid in tailoring therapeutic strategies for these patients. The objective of this study was to assess proteomic patterns of HFpEF phenogroups identified through a machine-learning-based clustering model, with the aim of uncovering specific biological pathways associated with each phenogroup. METHODS: This study represents a post-hoc analysis of the ongoing Prospective mUlticenteR obServational stUdy of patIenTs with Heart Failure with preserved Ejection Fraction (PURSUIT-HFpEF) study, which is a multicentre prospective observational study of hospitalised patients with acute decompensated HFpEF. Of the overall cohort (N=1238), this study analysed 198 patients with HFpEF with available proteomics data. These patients were classified into four phenogroups using the machine-learning-based clustering model. The SomaScan assay V.4.1 was used to measure levels of >7000 plasma proteins, and subsequent pathway analysis was conducted to determine the biological differences among the phenogroups. RESULTS: We identified four distinct phenogroups: Phenogroup 1 ('rhythm trouble'), Phenogroup 2 ('ventricular-arterial uncoupling'), Phenogroup 3 ('low output and systemic congestion') and Phenogroup 4 ('systemic failure'). The proteomics revealed distinct protein expression profiles among the phenogroups, with ribonuclease 4, tax1-binding protein 1, regenerating islet-derived protein 3-gamma and alpha-1-antichymotrypsin being the most significant markers to specific identified phenogroups. Pathway analysis suggested differences in immune response, autonomic activation, cellular homeostasis and tissue repair mechanisms across the phenogroups. CONCLUSIONS: Using a comprehensive plasma proteomics approach, our study identified distinct proteomic profiles of HFpEF phenogroups, which in turn suggest specific underlying biological processes. These profiles suggest the involvement of inflammatory activation, tissue injury and regenerative responses, immune modulation and systemic stress signalling as key components of HFpEF pathophysiology. TRIAL REGISTRATION NUMBER: UMIN-CTR ID: UMIN000021831.

Humans↗

Prematurity and Genetic Liability for Autism Spectrum Disorder.

BACKGROUND: Autism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by diverse presentations and a strong genetic component. Environmental factors, such as prematurity, have also been linked to increased liability for ASD, though the interaction between genetic predisposition and prematurity remains unclear. This study aims to investigate the impact of genetic liability and preterm birth on ASD conditions. METHODS: We analyzed phenotype and genetic data from two large ASD cohorts, the Simons Foundation Powering Autism Research for Knowledge (SPARK) and Simons Simplex Collection (SSC), encompassing 78,559 individuals for phenotype analysis, 12,519 individuals with genome sequencing data, and 8,104 individuals with exome sequencing data. Statistical significance of differences in clinical measures was evaluated between individuals with different ASD and preterm status. We assessed the rare variants burden using generalized estimating equations (GEE) models and polygenic load using ASD-associated polygenic risk score (PRS). Furthermore, we developed a machine learning model to predict ASD in preterm children using phenotype and genetic features available at birth. RESULTS: Individuals with both preterm birth and ASD exhibit more severe phenotypic outcomes despite similar levels of genetic liability for ASD across the term and preterm groups. Notably, preterm ASD individuals showed an elevated rate of de novo variants identified in exome sequencing (GEE model, p=0.005) in comparison to the non-ASD preterm group. Additionally, a GEE model showed that a higher ASD PRS, preterm birth, and male sex were positively associated with a higher predicted probability for ASD, reaching a probability close to 90% in SPARK. Lastly, we developed a machine learning model using phenotype and genetic features available at birth with limited predictive power (AUROC = 0.65). CONCLUSIONS: Preterm birth may exacerbate the multimorbidity present in ASD, which was not due to the ASD genetic factors. However, increased genetic factors may elevate the likelihood of a preterm child being diagnosed with ASD. Additionally, a polygenic load of ASD-associated variants had an additive role with preterm birth in the predicted probability for ASD, especially for boys. We propose that incorporating genetic assessment into neonatal care could benefit early ASD identification and intervention for preterm infants.

Autism Spectrum Disorder↗

NRG-P0074 Viral Sample RU1 from Unclassified Mosigvirus Genomic Characterization and Host Range Analysis.

BACKGROUND: Machine learning models for phage-host range prediction and design require comprehensive training data on phage genomes and host ranges to predict phage-host interactions effectively. MATERIALS AND METHODS: This study characterizes phage sample NRG-P0074 viral sample RU1 from unclassified Mosigvirus, originally isolated by the Betty Kutter. The complete genome of NRG-P0074 was sequenced, annotated, and analyzed using various bioinformatic tools. Host range analysis was conducted using the Escherichia coli Reference (ECOR) Library and nine Escherichia coli (E. coli) K12 strains (Keio Knockout Collection) with single nonessential gene deletions. RESULTS: The genome of NRG-P0074 spans 168,357 base pairs with a guanine-cytosine (GC) content of 37.5%. NRG-P0074 exhibited permissiveness in 15.28% of the ECOR isolates and all 9 Keio knockout strains. Comparative genomic analysis revealed that NRG-P0074 is closely related to E. coli phage a20. Its genome is comprised of 270 coding sequences, 153 known genes, 16 terminators, 3 ribosomal-binding sites, 0 tRNAs, and 117 hypothetical proteins. CONCLUSIONS: This research provides valuable data for developing machine learning models to predict phage-host interactions, aiding the development of targeted phage therapies against antibiotic-resistant bacteria.

ECOR Library↗

Multi-criteria decision making and its application to in silico discovery of vaccine candidates for Toxoplasma gondii.

Vaccine discovery against eukaryotic parasites is not trivial and few exist. Reverse vaccinology is an in silico vaccine discovery approach, designed to identify vaccine candidates from the thousands of protein sequences encoded by a target genome. Previously, we produced the Vacceed bioinformatics pipeline for identification of parasite membrane and excreted/secreted proteins that were likely be exposed to the hosts immune system. More recently, we improved upon machine learning as the final decision-making process to identify parasite proteins that induce a protective response in an animal model. Subsequently, we combined Vacceed with metrics on B and T cell epitope types to produce a new in silico discovery workflow. In this study we extend this in silico workflow to the developability of proteins as vaccines by the incorporation of metrics on the physicochemical properties of proteins. To demonstrate this process, every Toxoplasma gondii protein was ranked in its capacity to provide exposure to the immune system (Vacceed exposure score), presence of epitopes and solubility characteristics by several multicriteria decision making (MCDM) tools (such as TOPSIS, VIKOR and MABAC). A consensus rank was subsequently generated from the results of these tools using a variety of aggregate ranking methods. Levels of uncertainty in the aggregate protein rankings was assessed by conformal interval prediction in association with a machine learning model. Several of the top ranked proteins identified by this approach were novel, uncharacterized membrane transporters or proteins associated with RNA metabolism. In conclusion, MCDM automated the decision making using well known algorithms while conformal prediction intervals varied significantly across the 8000+ proteins of T. gondii. Highly ranked proteins (e.g. the top 100) typically generated low prediction intervals, providing high levels of confidence in their ranks.

Toxoplasma↗

Artificial intelligence in treatment prediction for skeletal Class III malocclusion: A systematic review.

In skeletal Class III patients, treatment options range from orthodontics to orthognathic surgery. Choosing the optimal approach requires a comprehensive clinical evaluation, which may be supported by AI tools. The aim of this study was to assess the performance of AI models in predicting the need for orthognathic surgery and in identifying predictors influencing treatment decisions. A PRISMA-guided electronic database search (PubMed, Web of Science; 2009-2024; English/French) was performed to identify studies using machine learning (ML) or deep learning (DL) on cephalometric and clinical data. After screening and assessment for eligibility, 15 studies were critically appraised. Model performance was summarized using accuracy, sensitivity, specificity, and the area under the curve (AUC). ML algorithms (particularly Random Forest and XGBoost) and DL models (ResNet-based convolutional neural networks (CNNs)) achieved high accuracy for predicting surgical need. Frequently selected predictors included Wits appraisal, ANB angle, the maxillomandibular ratio (Mx/Md), overjet, and the divergence of the lower gonial angle. AI methods show promise for assisting treatment decisions in Class III malocclusion, with Random Forest and XGBoost performing well on tabular cephalometric data and CNNs on imaging. Larger, multicentre datasets and external validation are needed to improve reliability, address bias, and support clinical implementation.

Humans↗

Systemic Proteome Profiling to Differentiate Primary Glomerular Diseases.

KEY POINTS: Plasma proteome profiling identified distinct signatures across biopsy-proven primary glomerular disease subtypes. An elastic net model using 93 proteins classified primary glomerular disease subtypes and controls, with external validation. Integrating proteomics with machine learning yields biologically interpretable insights in primary glomerular diseases. BACKGROUND: Primary GN is a heterogeneous group of kidney disorders where understanding of their pathophysiology remains incomplete. Despite the diagnostic potential of high-throughput proteomics, constrained proteomic depth and a reliance on binary comparisons have left the feasibility of using systemic signatures to differentiate multiple GN subtypes largely unexplored. METHODS: To identify protein signatures that noninvasively differentiate major primary glomerular disease subtypes and provide mechanistic insights, we performed large-scale systemic proteome profiling of 5416 plasma proteins via Olink Explore HT in a discovery cohort ( n =147) and an external validation cohort ( n =85) of Korean participants (mean age, 41&#xb1;13 years; 46% female). The study population included patients with four GN subtypes-focal segmental glomerulosclerosis, IgA nephropathy, minimal change disease, and membranous nephropathy-alongside healthy controls. We developed a machine learning (ML) model using logistic regression with elastic net regularization to classify disease groups based on proteomic profiles and evaluated its performance in the independent validation cohort. RESULTS: Plasma proteome profiles were distinct among disease subtypes, emerging as a significant source of data variation independent of conventional markers such as eGFR or proteinuria levels. The ML model performed robustly in both the discovery and validation cohorts, achieving an area under the receiver operating characteristic curve >0.8 for differentiating minimal change disease, membranous nephropathy, and IgA nephropathy. The model, even without clinical information, correctly identified 93% of minimal change disease cases (14 of 15) and 63% of IgA nephropathy cases (20 of 32), but its performance was limited for focal segmental glomerulosclerosis, with only 21% of cases (three of 14) correctly classified. Functional analysis of key proteins highlighted distinct biologic pathways, such as hemostasis in minimal change disease. CONCLUSIONS: We identified distinct systemic proteome signatures for primary glomerular diseases, where disease subtype served as a major determinant of proteomic variance alongside conventional clinical markers. ML models demonstrated robust discriminatory performance for minimal change disease, membranous nephropathy, and IgA nephropathy, underscoring the potential for proteome-based classification.

Humans↗

Computer-derived nuclear "grade" and breast cancer prognosis.

Visual assessments of nuclear grade are subjective yet still prognostically important. Now, computer-based analytical techniques can objectively and accurately measure size, shape and texture features, which constitute nuclear grade. The cell samples used in this study were obtained by fine needle aspiration (FNA) during the diagnosis of 187 consecutive patients with invasive breast cancer. Regions of FNA preparations to be analyzed were digitized and displayed on a computer monitor. Nuclei to be analyzed were roughly outlined by an operator using a mouse. Next, the computer generated a "snake" that precisely enclosed each designated nucleus. Ten nuclear features were then calculated for each nucleus based on these snakes. These results were analyzed statistically and by an inductive machine learning technique that we developed and call "recurrence surface approximation" (RSA). Both the statistical and RSA machine learning analyses demonstrated that computer-derived nuclear features are prognostically more important than are the classic prognostic features, tumor size and lymph node status.

Adult↗

Potential evaluation of SULT1A3 as an early diagnostic marker for nasopharyngeal carcinoma: a study based on serum proteomics screening and ELISA validation.

BACKGROUND: Nasopharyngeal carcinoma (NPC) represents a highly prevalent and aggressive malignancy endemic to Southeast Asia. Early and accurate diagnosis is critical to improving survival outcomes; however, the absence of robust, stage-specific biomarkers remains a key obstacle to clinical implementation of early screening strategies. METHODS: We performed untargeted serum proteomic profiling using mass spectrometry in 15 treatment-na&#xef;ve early-stage NPC patients and 15 VCA-IgA-positive healthy controls. Bioinformatics analyses were conducted to identify differentially expressed proteins (DEPs). Machine learning (random forest combined with recursive feature elimination) was employed to prioritize candidate biomarkers, which were subsequently verified using enzyme-linked immunosorbent assay (ELISA) in independent sample cohorts. RESULTS: In total, 1,428 serum proteins were identified, among which 1,410 were reliably quantified. We observed 31 upregulated and 189 downregulated proteins in NPC patients relative to controls. Spearman correlation analysis revealed significant associations: LTA4H (leukotriene A4 hydrolase) levels correlated with serum cell infiltration (r&#x2009;=&#x2009;0.383, p&#x2009;=&#x2009;0.032) and CD8&#x2009;+&#x2009;T-cell abundance (r&#x2009;=&#x2009;0.408, p&#x2009;=&#x2009;0.021); both SULT1A3 (sulfotransferase family 1&#xa0;A member 3) and FGL1 (fibrinogen-like protein 1) levels were positively associated with M1 macrophage infiltration (r&#x2009;=&#x2009;0.510, p&#x2009;=&#x2009;0.003 and r&#x2009;=&#x2009;0.430, p&#x2009;=&#x2009;0.015, respectively). In a preliminary validation cohort (n&#x2009;=&#x2009;80), ELISA yielded AUC values of 0.631 (95% CI: 0.515-0.736, p&#x2009;=&#x2009;0.04) for LTA4H, 0.787 (95% CI: 0.681-0.871, p&#x2009;<&#x2009;0.001) for SULT1A3, and 0.688 (95% CI: 0.575-0.787, p&#x2009;=&#x2009;0.002) for FGL1. In large-scale independent validation, SULT1A3 achieved an AUC of 0.826 (95% CI: 0.766-0.876; sensitivity&#x2009;=&#x2009;78.89%, specificity&#x2009;=&#x2009;75.47%) in cohort 1 (n&#x2009;=&#x2009;196) and 0.796 (95% CI: 0.723-0.857; sensitivity&#x2009;=&#x2009;76.67%, specificity&#x2009;=&#x2009;76.67%) in cohort 2 (n&#x2009;=&#x2009;150). CONCLUSIONS: Through an integrated workflow combining proteomic screening, machine learning prioritization, and multi-stage ELISA validation, we identified SULT1A3 as a candidate serum-based biomarker for early detection of NPC. Preliminary findings suggest that SULT1A3 may have potential utility in clinical screening, though further validation in independent, multi&#x2011;center cohorts is required.

Humans↗

Machine classification of dental images with visual search.

RATIONALE AND OBJECTIVES: The authors performed this study to assess the performance of a computer-based classification system that uses gaze locations of observers to define the subspace for machine learning. MATERIALS AND METHODS: Thirty-two dental radiographs were classified by an expert viewer into four categories of disease of the periapical region: no disease (normal tooth), mild disease (widened periodontal ligament space), moderate disease (destruction of the lamina dura), and severe disease (resorption of bone in the periapical area). There were eight images in each category. Six observers independently viewed the images while their eye gaze position was recorded. They then classified the images into one of the four categories. A sample of image space was used as input to a machine learning routine to develop a machine classifier. Sample space was determined with three techniques: visual gaze, random selection, and constrained random selection. K analyses were used to compare classification accuracies with the three sampling techniques. RESULTS: With use of the expert classification as a standard of reference, observers classified images with 57% accuracy, and the machine classified images with 84% accuracy by using the same gaze-selected features and image space. Results of kappa analyses revealed mean values of 0.78 for gaze-selected sampling, 0.69 for random sampling, 0.68 for constrained random selection, and 0.44 for observers. The use of sample space selected with the visual gaze technique was superior to that selected with both random-selection techniques and by the observers. CONCLUSION: Machine classification of dental images improves the accuracy of individual observers using gaze-selected image space.

Artificial Intelligence↗

Immunohistochemical analysis and prognostic value of cathepsin D determination in laryngeal squamous cell carcinoma.

Cathepsin D, a protease with the capability of degrading matrix proteins, is implicated in the process of breast and colorectal cancer invasion and metastasis. Biochemical studies in laryngeal cancer have shown a potential prognostic significance of cathepsin D content determination. We studied immunohistochemical positivity of cathepsin D in tumor epithelium and stroma of 61 surgical specimens of squamous cell laryngeal cancer. Immunohistochemical reaction was quantitatively assessed using a PC-based image analysis system SFORM-VAMS. The results were correlated to clinical and morphological parameters and survival. Immunohistochemical positivity was noted in neoplastic cells and tumor stroma. Significant prognostic value for cathepsin D was established separately for epithelial tumor component and tumor stroma using log-rank test, the Cox proportional hazards regression model, and C4.5 machine learning system. In all groups, patients above the median cathepsin D staining showed significantly shorter survival time. C4.5 machine learning system extracted cutoff values for the decision tree that defines the probabilities of patients survival and death with high sensitivity (92.8% alive, 73.6% dead), 100% specificity, and 86.9% accuracy. This makes immunohistochemical cathepsin D estimation an independent prognostic parameter in laryngeal carcinomas within a 5-year period from the time of tumor surgery.

Carcinoma, Squamous Cell↗

Induction of decision trees and Bayesian classification applied to diagnosis of sport injuries.

Machine learning techniques can be used to extract knowledge from data stored in medical databases. In our application, various machine learning algorithms were used to extract diagnostic knowledge which may be used to support the diagnosis of sport injuries. The applied methods include variants of the Assistant algorithm for top-down induction of decision trees, and variants of the Bayesian classifier. The available dataset was insufficient for reliable diagnosis of all sport injuries considered by the system. Consequently, expert-defined diagnostic rules were added and used as pre-classifiers or as generators of additional training instances for diagnoses for which only few training examples were available. Experimental results show that the classification accuracy and the explanation capability of the naive Bayesian classifier with the fuzzy discretization of numerical attributes were superior to other methods and estimated as the most appropriate for practical use.

Artificial Intelligence↗