Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Uncovering the genetic architecture of ME/CFS: a precision approach reveals impact of rare monogenic variation.

BACKGROUND: Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a disabling and heterogeneous disorder lacking validated biomarkers or targeted therapies. Clinical variability and elusive pathophysiology hinder progress toward effective diagnostics and treatment. Core symptoms include persistent fatigue, post-exertional malaise, unrefreshing sleep, cognitive dysfunction, and pain. We tested whether an individualized, “n-of-1” genomic and transcriptomic framework combined with comprehensive, participant-informed phenotyping could reveal molecular signatures unique to each patient. METHODS: Clinical-grade whole-genome sequencing was conducted in 31 affected individuals from 25 families, with RNA-seq performed on a subset (16 affected, 7 unaffected) using blood samples. Machine-learning assisted variant triage, transcript-aware damage prediction, and expert review identified pathogenic or likely pathogenic variants in 8 of 25 probands (32%) and 12 of 31 affected individuals (39%). RESULTS: Findings revealed marked genetic heterogeneity, including large-effect rare and more common variants. Implicated pathways included ATP generation, oxidative phosphorylation, fatty acid oxidation; regulation of glycolysis, amino acid and lipid turnover; ion and solute homeostasis; synaptic signaling, excitability, oxygen transport, and muscle integrity, resilience, and post-exertional recovery; previously implicated processes. Plausible modifiers influencing disease onset, severity, and relapsing–remitting patterns and possibly explaining intrafamilial variability and inconsistent findings across studies, were also identified. Despite gene-level diversity, downstream effects converged on impaired energy production, reduced stress resilience, and vulnerability to post-exertional metabolic failure; disruptions consistent with core ME/CFS symptoms of exertional intolerance, cognitive fog, and fatigue. CONCLUSIONS: Our findings support the hypothesis that at least a subset of ME/CFS cases represent distinct molecular disorders that converge on shared physiological pathways. Validation in larger, more diverse cohorts will be essential to test this hypothesis and establish generalizability, but increase size alone is unlikely to resolve causation in a disorder defined by rarity, heterogeneity, and molecular complexity. We suggest that progress will require experimental designs that integrate individual-level genomic data with deep, participant-informed deep phenotyping, capturing the combined effects of rare and common variants and environmental modifiers on disease expression and progression. We believe that an individualized precision medicine framework will uncover molecular drivers and modifiers of ME/CFS previously obscured by heterogeneity, enabling biologically informed stratification, improved trial design, biomarker discovery, and targeted interventions in this historically neglected condition.

Humans↗

Integrative multi-omics analysis proposes a metabolic classification of gliomas: distinct metabolic states, immune infiltration, and prognosis.

BACKGROUND: The tumor microenvironment (TME) of glioma harbors diverse cell types; however, cell metabolic heterogeneity remains to be explored. This study aims to characterize the metabolic features of different cell types in the TME by integrating multiple datasets, including genomics, bulk and single-cell transcriptomics, and metabolomics. METHODS: Unsupervised machine learning was used to construct an energy metabolic classifier based on the metabolic pathways identified from bulk RNA-seq of gliomas in the TCGA dataset. The classifier was externally validated using multiple datasets, including genomics, bulk RNA-seq, snRNA-seq, and the metabolomics data. Furthermore, metabolic heterogeneity associated with the classifier was further characterized at single-cell resolution. RESULTS: The energy metabolism-based classifier stratified patients into two prognostic clusters: patients in cluster 1 were characterized by high pathway activity of glycolysis, the pentose phosphate pathway (PPP), and fatty acid oxidation (FAO), whereas patients in cluster 2 exhibited higher activity in glutaminolysis. This metabolic classifier revealed both intratumoral and intertumoral metabolic heterogeneity, and the complexity was further validated by the metabolomics profiling and snRNA-seq data from the CPTAC dataset. Notably, OSMR, highly expressed in cluster 1, showed significant co-expression with key glycolytic enzyme genes. The OSM/OSMR/JAK1/STAT3 axis potently drives malignant progression of glioma cells, specially enhancing their invasive and migratory capabilities. Single-cell resolution analyses demonstrated that tumor metabolic heterogeneity is primarily driven by malignant cells rather than non-malignant components, while tumor microenvironment (TME) factors were also found to modulate malignant cell metabolism. Significantly, glycolytic activity in glioma cells increased during the phenotypic transition from PN (proneural) to MES (mesenchymal), with cluster 1 metabolic phenotypes predominating in the tumor core. Compared to cluster 2, cluster 1 patients exhibited higher mRNA expression of immunosuppressive checkpoint genes, which correlated with pronounced immunosuppression in the TME. Furthermore, various immune cells demonstrated distinct metabolic preferences at single-cell resolution. CONCLUSIONS: This study developed an energy metabolic-based classifier for gliomas with prognostic and therapeutic potential. Metabolic reprogramming was linked with the PN-to-MES transition of glioma cells and immunosuppression in the tumor microenvironment. Multi-omics data, especially snRNA-seq, offered insights into metabolism heterogeneity at single-cell resolution, enabling personalized treatment strategies.

Humans↗

Integrative multi-omics and single-cell analysis identifies EGFR pathway activation and metabolic reprogramming as potential synthetic lethal vulnerabilities in resistance to the FGFR inhibitor AZD4547.

BACKGROUND: Although fibroblast growth factor receptor (FGFR) inhibitors (FGFRi) have demonstrated clinical promise, the inevitable emergence of acquired resistance remains a critical bottleneck, severely compromising their long-term clinical efficacy. The pan-cancer molecular landscape and heterogeneous mechanisms driving this resistance, ranging from genetic alterations to dynamic network rewiring, remain poorly understood. METHODS: We integrated large-scale pharmacogenomic profiling of the FGFR inhibitor AZD4547 from the GDSC2 and PRISM databases with single-cell RNA sequencing to dissect the multi-omics landscape of FGFRi resistance across 312 cell lines from 8 cancer types. This multi-omics framework was further extended by machine learning modeling and systematic synthetic lethality screening to uncover actionable therapeutic targets. In vitro viability assays and western blot analysis were subsequently conducted to experimentally evaluate the predicted FGFR-EGFR synthetic lethality. RESULTS: Our dual-database analysis unveiled a multi-dimensional atlas of FGFRi resistance. We identified cancer-specific genomic drivers, such as ELF4 amplification in glioblastoma, alongside key transcriptomic markers including UCP2 and FSCN1, highlighting a shift towards metabolic reprogramming and epithelial-mesenchymal transition (EMT). Single-cell analysis unveiled that resistance is linked to the heterogeneous enrichment of baseline subpopulations characterized by distinct metaprograms, including cell-cycle dysregulation. Furthermore, a random forest model built on a LASSO-derived transcriptomic signature was constructed, demonstrating promising predictive capability for AZD4547 sensitivity (mean test-set AUC = 0.73, 95% CI [0.63, 0.80]); the signature generalized well to erdafitinib but showed limited transferability to some other FGFR inhibitors (e.g. pemigatinib, BGJ398). Most notably, our synthetic lethal screening revealed a convergent reliance on compensatory RTK signaling (specifically EGFR pathway enrichment) and downstream MAPK/PI3K cascades in resistant phenotypes, providing converging computational evidence for EGFR pathway activation as an adaptive bypass mechanism. This predicted synthetic lethality was experimentally supported in two FGFR-dependent cell line models (RT112 and CCLP1), in which combined FGFR-EGFR inhibition produced marked synergistic antiproliferative effects. CONCLUSIONS: This study establishes a comprehensive multi-omics atlas of resistance to the FGFR inhibitor AZD4547, delineating convergent mechanisms of metabolic reprogramming and EGFR-mediated bypass signaling. Our findings characterize the resistance as a dynamic network rewiring and nominate rational combination strategies to overcome this therapeutic bottleneck. While FGFR-EGFR co-inhibition is experimentally supported, metabolic co-targeting remains a computationally derived, hypothesis-generating strategy.

Benzamides↗

Artificial intelligence in healthcare and medicine: clinical applications, therapeutic advances, and future perspectives.

Healthcare systems worldwide face growing challenges, including rising costs, workforce shortages, and disparities in access and quality, particularly in low- and middle-income countries. Artificial intelligence (AI) has emerged as a transformative tool capable of addressing these issues by enhancing diagnostics, treatment planning, patient monitoring, and healthcare efficiency. AI's role in modern medicine spans disease detection, personalized care, drug discovery, predictive analytics, telemedicine, and wearable health technologies. Leveraging machine learning and deep learning, AI can analyze complex data sets, including electronic health records, medical imaging, and genomic profiles, to identify patterns, predict disease progression, and recommend optimized treatment strategies. AI also has the potential to promote equity by enabling cost-effective, resource-efficient solutions in low-resource and remote settings, such as mobile diagnostics, wearable biosensors, and lightweight algorithms. Successful deployment requires addressing critical challenges, including data privacy, algorithmic bias, model interpretability, regulatory oversight, and maintaining human clinical oversight. Emphasizing scalable, ethical, and evidence-driven implementation, key strategies include clinician training in AI literacy, adoption of resource efficient tools, global collaboration, and robust regulatory frameworks to ensure transparency, safety, and accountability. By complementing rather than replacing healthcare professionals, AI can reduce errors, optimize resources, improve patient outcomes, and expand access to quality care. This review emphasizes the responsible integration of AI as a powerful catalyst for innovation, sustainability, and equity in healthcare delivery worldwide.

Humans↗

Predicting extubation outcome in preterm newborns: a comparison of neural networks with clinical expertise and statistical modeling.

Even though ventilator technology and monitoring of premature infants has improved immensely over the past decades, there are still no standards for weaning and determining optimal extubation time for those infants. Approximately 30% of intubated preterm infants will fail attempted extubation, requiring reintubation and resuming of mechanical ventilation. A machine-learning approach using artificial neural networks (ANNs) to aid in extubation decision making is hereby proposed. Using expert opinion, 51 variables were identified as being relevant for the decision of whether to extubate an infant who is on mechanical ventilation. The data on 183 premature infants, born between 1999 and 2002, were collected by review of medical charts. The ANN extubation model was compared with alternative statistical modeling using multivariate logistic regression and also with the clinician's own predictive insight using sensitivity analysis and receiver operating characteristic curves. The optimal ANN model used 13 parameters and achieved an area under the receiver operating characteristic curve of 0.87 (out-of-sample validation), comparing favorably with multivariate logistic regression. It also compared well with the clinician's expertise, which raises the possibility of being useful as an automated alert tool. Because an ANN learns directly from previous data obtained in the institution where it is to be used, this makes it particularly amenable for application to evidence-based medicine. Given the variety of practices and equipment being used in different hospitals, this may be particularly relevant in the context of caring for preterm newborns who are on mechanical ventilation.

Decision Making↗

Exercise Therapy in Down Syndrome: A Systematic Review and Meta-Analysis Focused on Muscle Strength, Redox Balance, and Inflammatory Profile.

OBJECTIVE: This study systematically reviewed and meta-analyzed randomized and quasi-randomized controlled trials investigating the impact of exercise therapy on muscle strength, redox balance, and inflammatory profile in individuals with Down syndrome. DESIGN: Systematic review and meta-analysis. DATA SOURCES: Cochrane Central Register of Controlled Trials, MEDLINE, CINAHL, SPORTDiscus, EMBASE, and PEDro. ELIGIBILITY CRITERIA FOR SELECTING STUDIES: Randomized and quasi-randomized controlled trials exploring exercise therapy effects on muscle strength and redox balance in individuals with Down syndrome. Although no initial restrictions on age, gender, or health condition were applied during the search process, all included studies focused on adult participants (>18 yr old). No language restrictions were applied, and the search covered the period from 1970 to 2021. RESULTS: We assessed the abstract of 1964 studies. Of the 46 studies meeting the inclusion criteria for the period 2004-2021, 32 focused on muscle strength, and 14 examined redox balance and inflammation. A total of 1611 participants with a mean age of 27 yr were included. This review confirmed that different exercise modalities are prone to improve muscle strength (random effect (95% confidence interval): 0.66, 0.54 to 0.78), redox balance and inflammatory profile (random effect (95% confidence interval): -1.04, -1.31 to -0.76) in this population. The multimodel inference suggested that the frequency of training (times per week) might play a significant role in the main effect. Unsupervised machine learning algorithms displayed a pattern-based graphic representation to assess heterogeneity. CONCLUSIONS: Exercise training demonstrated a positive impact on muscle strength in adults with Down syndrome. The review provides valuable insights into the effects of exercise therapy on individuals with Down syndrome, emphasizing the need for tailored training prescriptions.

Humans↗

RNA secondary structure prediction from sequence alignments using a network of k-nearest neighbor classifiers.

We present a machine learning method (a hierarchical network of k-nearest neighbor classifiers) that uses an RNA sequence alignment in order to predict a consensus RNA secondary structure. The input to the network is the mutual information, the fraction of complementary nucleotides, and a novel consensus RNAfold secondary structure prediction of a pair of alignment columns and its nearest neighbors. Given this input, the network computes a prediction as to whether a particular pair of alignment columns corresponds to a base pair. By using a comprehensive test set of 49 RFAM alignments, the program KNetFold achieves an average Matthews correlation coefficient of 0.81. This is a significant improvement compared with the secondary structure prediction methods PFOLD and RNAalifold. By using the example of archaeal RNase P, we show that the program can also predict pseudoknot interactions.

Algorithms↗

Prediction of RNA binding sites in proteins from amino acid sequence.

RNA-protein interactions are vitally important in a wide range of biological processes, including regulation of gene expression, protein synthesis, and replication and assembly of many viruses. We have developed a computational tool for predicting which amino acids of an RNA binding protein participate in RNA-protein interactions, using only the protein sequence as input. RNABindR was developed using machine learning on a validated nonredundant data set of interfaces from known RNA-protein complexes in the Protein Data Bank. It generates a classifier that captures primary sequence signals sufficient for predicting which amino acids in a given protein are located in the RNA-protein interface. In leave-one-out cross-validation experiments, RNABindR identifies interface residues with >85% overall accuracy. It can be calibrated by the user to obtain either high specificity or high sensitivity for interface residues. RNABindR, implementing a Naive Bayes classifier, performs as well as a more complex neural network classifier (to our knowledge, the only previously published sequence-based method for RNA binding site prediction) and offers the advantages of speed, simplicity and interpretability of results. RNABindR predictions on the human telomerase protein hTERT are in good agreement with experimental data. The availability of computational tools for predicting which residues in an RNA binding protein are likely to contact RNA should facilitate design of experiments to directly test RNA binding function and contribute to our understanding of the diversity, mechanisms, and regulation of RNA-protein complexes in biological systems. (RNABindR is available as a Web tool from http://bindr.gdcb.iastate.edu.).

Amino Acid Motifs↗

Weighted sequence motifs as an improved seeding step in microRNA target prediction algorithms.

We present a new microRNA target prediction algorithm called TargetBoost, and show that the algorithm is stable and identifies more true targets than do existing algorithms. TargetBoost uses machine learning on a set of validated microRNA targets in lower organisms to create weighted sequence motifs that capture the binding characteristics between microRNAs and their targets. Existing algorithms require candidates to have (1) near-perfect complementarity between microRNAs' 5' end and their targets; (2) relatively high thermodynamic duplex stability; (3) multiple target sites in the target's 3' UTR; and (4) evolutionary conservation of the target between species. Most algorithms use one of the two first requirements in a seeding step, and use the three others as filters to improve the method's specificity. The initial seeding step determines an algorithm's sensitivity and also influences its specificity. As all algorithms may add filters to increase the specificity, we propose that methods should be compared before such filtering. We show that TargetBoost's weighted sequence motif approach is favorable to using both the duplex stability and the sequence complementarity steps. (TargetBoost is available as a Web tool from http://www.interagon.com/demo/.).

5' Untranslated Regions↗

Data quality in predictive toxicology: identification of chemical structures and calculation of chemical properties.

Every technique for toxicity prediction and for the detection of structure-activity relationships relies on the accurate estimation and representation of chemical and toxicologic properties. In this paper we discuss the potential sources of errors associated with the identification of compounds, the representation of their structures, and the calculation of chemical descriptors. It is based on a case study where machine learning techniques were applied to data from noncongeneric compounds and a complex toxicologic end point (carcinogenicity). We propose methods applicable to the routine quality control of large chemical datasets, but our main intention is to raise awareness about this topic and to open a discussion about quality assurance in predictive toxicology. The accuracy and reproducibility of toxicity data will be reported in another paper.

Data Interpretation, Statistical↗

Personality traits in miners with past occupational elemental mercury exposure.

In this study, we evaluated the impact of long-term occupational exposure to elemental mercury vapor (Hg0) on the personality traits of ex-mercury miners. Study groups included 53 ex-miners previously exposed to Hg0 and 53 age-matched controls. Miners and controls completed the self-reporting Eysenck Personality Questionnaire and the Emotional States Questionnaire. The relationship between the indices of past occupational exposure and the observed personality traits was evaluated using Pearson's correlation coefficient and on a subgroup level by machine learning methods (regression trees). The ex-mercury miners were intermittently exposed to Hg0 for a period of 7-31 years. The means of exposure-cycle urine mercury (U-Hg) concentrations ranged from 20 to 120 microg/L. The results obtained indicate that ex-miners tend to be more introverted and sincere, more depressive, more rigid in expressing their emotions and are likely to have more negative self-concepts than controls, but no correlations were found with the indices of past occupational exposure. Despite certain limitations, results obtained by the regression tree suggest that higher alcohol consumption per se and long-term intermittent, moderate exposure to Hg0 (exposure cycle mean U-Hg concentrations > 38.7 < 53.5 microg/L) in interaction with alcohol remain a plausible explanation for the depression associated with negative self-concept found in subgroups of ex-mercury miners. This could be one of the reason for the higher risk of suicide among miners of the Idrija Mercury Mine in the last 45 years.

Adult↗

Pattern recognition analysis of a set of mutagenic aliphatic N-nitrosamines.

A set of 21 mutagenic aliphatic N-nitrosamines were subjected to a pattern recognition analysis using ADAPT software. Four descriptors based on molecular connectivity, geometry and sigma charge on nitrogen were capable of achieving a 100% classification using the linear learning machine or iterative least squares algorithms. Three descriptors were capable of a 90.5% and two descriptors of a 85.7% overall correct classification. Three of the four descriptors were each capable of classifying 15 of the 16 active chemicals while it required three of the four descriptors to classify correctly two of the five inactive chemicals. These results are in concert with previous observations that molecular connectivity, geometry, and sigma charge on nitrogen are powerful descriptors for separating active from inactive mutagenic and carcinogenic N-nitrosamines.

Mutagens↗

Large-scale mapping and validation of Escherichia coli transcriptional regulation from a compendium of expression profiles.

Machine learning approaches offer the potential to systematically identify transcriptional regulatory interactions from a compendium of microarray expression profiles. However, experimental validation of the performance of these methods at the genome scale has remained elusive. Here we assess the global performance of four existing classes of inference algorithms using 445 Escherichia coli Affymetrix arrays and 3,216 known E. coli regulatory interactions from RegulonDB. We also developed and applied the context likelihood of relatedness (CLR) algorithm, a novel extension of the relevance networks class of algorithms. CLR demonstrates an average precision gain of 36% relative to the next-best performing algorithm. At a 60% true positive rate, CLR identifies 1,079 regulatory interactions, of which 338 were in the previously known network and 741 were novel predictions. We tested the predicted interactions for three transcription factors with chromatin immunoprecipitation, confirming 21 novel interactions and verifying our RegulonDB-based performance estimates. CLR also identified a regulatory link providing central metabolic control of iron transport, which we confirmed with real-time quantitative PCR. The compendium of expression data compiled in this study, coupled with RegulonDB, provides a valuable model system for further improvement of network inference algorithms using experimental data.

Algorithms↗

Landscape of essential growth and fluconazole-resistance genes in the human fungal pathogen Cryptococcus neoformans.

Fungi can cause devastating invasive infections, typically in immunocompromised patients. Treatment is complicated both by the evolutionary similarity between humans and fungi and by the frequent emergence of drug resistance. Studies in fungal pathogens have long been slowed by a lack of high-throughput tools and community resources that are common in model organisms. Here we demonstrate a high-throughput transposon mutagenesis and sequencing (TN-seq) system in Cryptococcus neoformans that enables genome-wide determination of gene essentiality. We employed a random forest machine learning approach to classify the C. neoformans genome as essential or nonessential, predicting 1,465 essential genes, including 302 that lack human orthologs. These genes are ideal targets for new antifungal drug development. TN-seq also enables genome-wide measurement of the fitness contribution of genes to phenotypes of interest. As proof of principle, we demonstrate the genome-wide contribution of genes to growth in fluconazole, a clinically used antifungal. We show a novel role for the well-studied RIM101 pathway in fluconazole susceptibility. We also show that insertions of transposons into the 5' upstream region can drive sensitization of essential genes, enabling screenlike assays of both essential and nonessential components of the genome. Using this approach, we demonstrate a role for mitochondrial function in fluconazole sensitivity, such that tuning down many essential mitochondrial genes via 5' insertions can drive resistance to fluconazole. Our assay system will be valuable in future studies of C. neoformans, particularly in examining the consequences of genotypic diversity.

Cryptococcus neoformans↗

Tumor-immune partitioning and clustering algorithm for identifying tumor-immune cell spatial interaction signatures within the tumor microenvironment.

BACKGROUND: Growing evidence supports the importance of characterizing the organizational patterns of various cellular constituents in the tumor microenvironment in precision oncology. Most existing data on immune cell infiltrates in tumors, which are based on immune cell counts or nearest neighbor-type analyses, have failed to fully capture the cellular organization and heterogeneity. METHODS: We introduce a computational algorithm, termed Tumor-Immune Partitioning and Clustering (TIPC), that jointly measures immune cell partitioning between tumor epithelial and stromal areas and immune cell clustering versus dispersion. As proof-of-principle, we applied TIPC to a prospective cohort incident tumor biobank containing 931 colorectal carcinoma cases. TIPC identified tumor subtypes with unique spatial patterns between tumor cells and T lymphocytes linked to certain molecular pathologic and prognostic features. T lymphocyte identification and phenotyping were achieved using multiplexed (multispectral) immunofluorescence. In a separate hepatocellular carcinoma cohort, we replaced the stromal component with specific immune cell types-CXCR3+CD68+ or CD8+-to profile their spatial relationships with CXCL9+CD68+ cells. RESULTS: Six unsupervised TIPC subtypes based on T lymphocyte distribution patterns were identified, comprising two cold and four hot subtypes. Three of the four hot subtypes were associated with significantly longer colorectal cancer (CRC)-specific survival compared to a reference cold subtype. Our analysis showed that variations in T-cell densities among the TIPC subtypes did not strictly correlate with prognostic benefits, underscoring the prognostic significance of immune cell spatial patterns. Additionally, TIPC revealed two spatially distinct and cell density-specific subtypes among microsatellite instability-high colorectal cancers, indicating its potential to upgrade tumor subtyping. TIPC was also applied to additional immune cell types, eosinophils and neutrophils, identified using morphology and supervised machine learning; here two tumor subtypes with similarly low densities, namely 'cold, tumor-rich' and 'cold, stroma-rich', exhibited differential prognostic associations. Lastly, we validated our methods and results using The Cancer Genome Atlas colon and rectal adenocarcinoma data (n = 570). Moreover, applying TIPC to hepatocellular carcinoma cases (n = 27) highlighted critical cell interactions like CXCL9-CXCR3 and CXCL9-CD8. CONCLUSIONS: Unsupervised discoveries of microgeometric tissue organizational patterns and novel tumor subtypes using the TIPC algorithm can deepen our understanding of the tumor immune microenvironment and likely inform precision cancer immunotherapy.

Humans↗

IgStrand: A universal residue numbering scheme for the immunoglobulin-fold (Ig-fold) to study Ig-proteomes and Ig-interactomes.

The Immunoglobulin fold (Ig-fold) is found in proteins from all domains of life and represents the most populous fold in the human genome, with current estimates ranging from 2 to 3% of protein coding regions. That proportion is much higher in the surfaceome where Ig and Ig-like domains orchestrate cell-cell recognition, adhesion and signaling. The ability of Ig-domains to reliably fold and self-assemble through highly specific interfaces represents a remarkable property of these domains, making them key elements of molecular interaction systems: the immune system, the nervous system, the vascular system and the muscular system. We define a universal residue numbering scheme, common to all domains sharing the Ig-fold in order to study the wide spectrum of Ig-domain variants constituting the Ig-proteome and Ig-Ig interactomes at the heart of these systems. The "IgStrand numbering scheme" enables the identification of Ig structural proteomes and interactomes in and between any species, and comparative structural, functional, and evolutionary analyses. We review how Ig-domains are classified today as topological and structural variants and highlight the "Ig-fold irreducible structural signature" shared by all of them. The IgStrand numbering scheme lays the foundation for the systematic annotation of structural proteomes by detecting and accurately labeling Ig-, Ig-like and Ig-extended domains in proteins, which are poorly annotated in current databases and opens the door to accurate machine learning. Importantly, it sheds light on the robust Ig protein folding algorithm used by nature to form beta sandwich supersecondary structures. The numbering scheme powers an algorithm implemented in the interactive structural analysis software iCn3D to systematically recognize Ig-domains, annotate them and perform detailed analyses comparing any domain sharing the Ig-fold in sequence, topology and structure, regardless of their diverse topologies or origin. The scheme provides a robust fold detection and labeling mechanism that reveals unsuspected structural homologies among protein structures beyond currently identified Ig- and Ig-like domain variants. Indeed, multiple folds classified independently contain a common structural signature, in particular jelly-rolls. Examples of folds that harbor an "Ig-extended" architecture are given. Applications in protein engineering around the Ig-architecture are straightforward based on the universal numbering.

Humans↗

Topologically distinct intratumoral heterogeneity scores for predicting high-risk pathological grades in invasive lung adenocarcinoma: A multicenter study across four institutions.

High-risk subtypes of invasive lung adenocarcinoma (IAC), particularly micropapillary- or solid-predominant patterns, are closely associated with poor prognosis. This multicenter retrospective study developed and validated a predictive model for the preoperative identification of these high-risk subtypes using topologically distinct intratumoral heterogeneity (ITH) scores derived from CT images. The study included 1,051 patients with IAC. Two complementary ITH scores were developed: a two-dimensional ITH score, which integrated local radiomics features with global pixel distribution patterns on the largest cross-sectional CT slice, and a three-dimensional ITH score, which extended this quantification across the entire tumor volume. Clinicoradiological features and ITH scores were incorporated as model inputs to construct six base machine learning classifiers and a final stacking ensemble classifier. Model interpretability and robustness were evaluated using SHapley Additive exPlanations (SHAP)-based ablation analyses. An independent dataset from The Cancer Imaging Archive (TCIA) was used for external validation to investigate associations between ITH scores and pathological characteristics, genomic features, recurrence-free survival, and overall survival. The stacking ensemble classifier achieved the best predictive performance, with an area under the receiver operating characteristic curve of 0.875, outperforming models based solely on radiomics features (0.834) or clinicoradiological features (0.792). SHAP analysis identified the 3D ITH score as the most influential contributor to model output, and TCIA validation showed that higher 3D ITH scores were associated with more aggressive tumor biology and poorer survival outcomes. The topologically distinct 3D ITH score may provide a clinically meaningful imaging biomarker for preoperative risk stratification in IAC.

Journal Article↗

Inhibiting the expression of spindle appendix cooled coil protein 1 can suppress tumor cell growth and metastasis and is associated with cancer immune cells in esophageal squamous cell carcinoma.

Inhibiting the expression of spindle appendix cooled coil protein 1 (SPDL1) can slow down disease progression and is related to poor prognosis in patients with esophageal cancer. However, the specific roles and molecular mechanisms of SPDL1 in esophageal squamous cell carcinoma (ESCC) have not been explored yet. The current study aimed to investigate the expression levels of SPDL1 in ESCC via transcriptome analysis using data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus databases. Moreover, the biological roles, molecular mechanisms, and protein networks involved in SPDL1 were identified using machine learning and bioinformatics. The cell counting kit-8 assay, EdU staining, and transwell assay were used to investigate the effects of inhibiting SPDL1 expression on ESCC cell proliferation, migration, and invasion. Finally, the correlation between the SPDL1 expression and cancer immune infiltrating cells was evaluated by analyzing data from the TCGA database. Results showed that SPDL1 was overexpressed in the ESCC tissues. The SPDL1 expression was related to age in patients with ESCC. The SPDL1 co-expressed genes included those involved in cell division, cell cycle, DNA repair and replication, cell aging, and other processes. The high-risk scores of SPDL1-related long non-coding RNAs were significantly correlated with overall survival and cancer progression in patients with ESCC (P < 0.05). Inhibiting the SPDL1 expression was effective in suppressing the proliferation, migration, and invasion of ESCC TE-1 cells (P < 0.05). The overexpression of SPDL1 was positively correlated with the levels of Th2 and T-helper cells, and was negatively correlated with the levels of plasmacytoid dendritic cells and mast cells. In conclusion, SPDL1 was overexpressed in ESCC and was associated with immune cells. Further, inhibiting the SPDL1 expression could effectively slow down cancer cell growth and migration. SPDL1 is a promising biomarker for treating patients with ESCC.

Humans↗