Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “random forest”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

On the nature of cavities on protein surfaces: application to the identification of drug-binding sites.

In this article we introduce a new method for the identification and the accurate characterization of protein surface cavities. The method is encoded in the program SCREEN (Surface Cavity REcognition and EvaluatioN). As a first test of the utility of our approach we used SCREEN to locate and analyze the surface cavities of a nonredundant set of 99 proteins cocrystallized with drugs. We find that this set of proteins has on average about 14 distinct cavities per protein. In all cases, a drug is bound at one (and sometimes more than one) of these cavities. Using cavity size alone as a criterion for predicting drug-binding sites yields a high balanced error rate of 15.7%, with only 71.7% coverage. Here we characterize each surface cavity by computing a comprehensive set of 408 physicochemical, structural, and geometric attributes. By applying modern machine learning techniques (Random Forests) we were able to develop a classifier that can identify drug-binding cavities with a balanced error rate of 7.2% and coverage of 88.9%. Only 18 of the 408 cavity attributes had a statistically significant role in the prediction. Of these 18 important attributes, almost all involved size and shape rather than physicochemical properties of the surface cavity. The implications of these results are discussed. A SCREEN Web server is available at http://interface.bioc.columbia.edu/screen.

Binding Sites↗

Rapid glycomic analysis of serum EVs reveals altered N-glycosylation patterns in ASD.

Objective laboratory diagnostics for autism spectrum disorder (ASD) are lacking, necessitating rapid clinical screening tools. Because serum extracellular vesicle (EV) N-glycosylation captures critical neurodevelopmental signatures, we developed a fast, biologically interpretable diagnostic strategy. EVs from ASD patients with language impairment and neurotypical controls were isolated using a rapid extra-polyethylene glycol precipitation/filtration (EPF) workflow, benchmarked against ultracentrifugation. Following MALDI-TOF/MS profiling, machine learning was re-evaluated using repeated nested cross-validation to reduce optimistic bias and potential information leakage. Among five classifiers, Random Forest (RF) showed the best overall balance across discrimination, calibration, and classification metrics. RF-based SHAP analysis provided transparent interpretation, highlighting key discriminative glycans, including H4N3S1F1, H5N5S1F1, and H3N5F1. To elucidate molecular mechanisms, we integrated public EV transcriptomic data. This revealed significant dysregulation of N-glycosylation machinery genes (e.g., MAN1A1, NEU1, OSTC, RPN2), whose expression directionally aligned with observed glycan shifts in synaptic pathways. Collectively, this rapid serum EV N-glycomic workflow, combined with leakage-controlled RF-based interpretation, provides a promising foundation for non-invasive ASD biomarker discovery and future multicenter validation.

Humans↗

Integrative analysis and experiment validation of SLC12A8 as a biomarker for the malignant transition from endometriosis to endometriosis associated ovarian cancer.

Endometriosis (EM) is a chronic inflammatory, estrogen‑dependent benign gynecological disorder. A subset of patients with EM may subsequently develop endometriosis‑associated ovarian cancer (EAOC), implying a biological continuum between these two conditions. Nevertheless, the molecular events underlying the progression from benign endometriotic lesions toward EAOC remain incompletely characterized. In this study, transcriptomic datasets retrieved from the GEO database were interrogated through differentially expressed gene screening, functional enrichment analysis, and weighted gene co‑expression network analysis (WGCNA) to identify key genes and pathways relevant to EM and EAOC. Candidate genes were further prioritized by integrating survival analysis via the Kaplan‑Meier Plotter, LASSO regression, random‑forest modeling, and CIBERSORT immune‑infiltration profiling. Loss and gain‑of‑function cellular models were established using siRNA and overexpression plasmids, and in‑vitro functional assays were performed to characterize the phenotypic effects of target genes.We identified several candidate genes associated with EM and EAOC and evaluated their discriminatory performance. Among them, SLC12A8 elevated expression across EM and EAOC tissues and exhibited moderate diagnostic capacity. Higher SLC12A8 expression was also associated with poorer prognosis in EAOC patients. In‑vitro experiments further demonstrated that SLC12A8 modulates proliferation, invasion, and migration in both EM and EAOC cell lines. Collectively, our exploratory research findings support SLC12A8 as a candidate functional mediator and potential biomarker linked to EM‑EAOC pathological progression, thereby extending the mechanistic understanding of these disorders.

Female↗

QSAR analysis of phenolic antioxidants using MOLMAP descriptors of local properties.

Molecular maps of atom-level properties (MOLMAPs) were developed to represent the diversity of chemical bonds existing in a molecule. Chemical reactivity, being related to the ability for bond breaking and bond making, is primarily determined by the properties of bonds available in a molecule. In order to use physicochemical properties of individual bonds for an entire molecule, and at the same time having a fixed-length molecular representation, all the bonds of a molecule are mapped into a fixed-size 2D self-organizing map (MOLMAP). This article illustrates the application of MOLMAP descriptors to QSAR, with a study of the radical scavenging activity of 47 naturally occurring phenolic antioxidants. Counterpropagation neural networks (CPG NNs) were trained with MOLMAP descriptors selected using genetic algorithms to predict antioxidant activity. The model was subsequently validated by the leave-one-out (LOO) procedure obtaining a q(2) of 0.71. Random Forests were grown with the entire set of MOLMAP descriptors giving 70% of correct classifications as potent, active or inactive in a LOO experiment. Interpretations of both models in terms of discriminant variables were concordant and allowed identifying bonds and substructures that are mostly responsible for antioxidant activity. This work shows how MOLMAPs can be used for data mining of structural and biological activity data, leading to the extraction of relationships between local properties and activity.

Algorithms↗

Revealing potential biomarkers and metabolic mechanisms of ovarian aging in hens during late laying period based on machine learning and metabolomics.

Ovarian function decline during the late laying period represents a major bottleneck for the economic efficiency of the global poultry industry. However, the underlying metabolic mechanisms and reliable early-warning biomarkers for ovarian aging remain poorly understood. In this study, we performed the first untargeted LC-MS/MS metabolomics analysis of ovarian tissues from Taihe silky fowls at peak laying (30 weeks) and late laying (50 weeks) stages, and employed an ensemble machine learning strategy integrating LASSO, random forest, and support vector machine (SVM) algorithms to identify high-confidence core biomarkers of ovarian aging. Gene expression analysis was further conducted to validate the potential molecular mechanisms. Our results showed that the metabolic profiles of ovarian tissues differed significantly between the two groups. A total of 6 core biomarkers were identified, 4 of which were long-chain acylcarnitines. Mechanistic analysis revealed that downregulation of key genes in the carnitine shuttle system led to impaired mitochondrial fatty acid β-oxidation, which in turn triggered excessive oxidative stress and compromised ovarian endocrine function. In conclusion, this study identifies long-chain acylcarnitines as potential metabolic biomarkers for ovarian aging in Taihe silky fowls. These findings provide novel insights into the metabolic basis of poultry ovarian aging and lay a theoretical foundation for the precise regulation of reproductive performance in indigenous poultry breeds.

Animals↗

Health-associated key gut microbiota drives the variation in community metabolic interactions in non-human primates.

Gut microbiota often undergo metabolic cross-feeding and resource competition. However, our understanding of global variations in these interactions and their implications for host health remain elusive. By analyzing a microbial genome catalog from 841 fecal metagenomes across 53 primate species worldwide, we identified key microbiota assigned to two taxa, i.e., Bacillota_A and Pseudomonadota, which well predicted the trade-off of community-level interaction types between metabolic competition and cooperation. Specifically, Bacillota_A species were inherently competitive and amino acid auxotrophic and typically found in anaerobic habitats. In contrast, members of Pseudomonadota were inherently cooperative, siderophore producers, and more abundant in aerobic conditions. Random forest models successfully distinguished unhealthy gut samples from healthy samples through the key competitive and cooperative microbiota, suggesting potential links between community metabolic interactions and host health. Together, this study enhances our mechanistic understanding of microbial interaction dynamism within complex gut ecosystems, offering new targets for understanding host health.

Animals↗

CTSG Suppresses Breast Cancer Progression by Inhibiting the EGFR/ERK Signaling Pathway and Enhancing CD8⁺ T Cell Activation.

BACKGROUND: Breast cancer (BC), the most common female malignancy, has metastasis as its main cause of mortality. Cathepsin G (CTSG) is involved in tumorigenesis and immunity. This study explores the role of CTSG in BC progression and CD8 + T cell regulation. METHODS: Differentially expressed genes and proteins (DEGs/DEPs) were analyzed using Limma, and core genes were screened using Random Forest (RF) and Least absolute shrinkage and selection operator (LASSO). CTSG expression was analyzed using GSE36295, the Cancer Genome Atlas (TCGA), reverse transcription-quantitative polymerase chain reaction (RT-qPCR), and western blot. Cell viability, proliferation, cell cycle, migration, and invasion were detected using Cell Counting Kit-8 (CCK8), 5&#x2011;Ethynyl&#x2011;2'&#x2011;deoxyuridine (EdU), flow cytometry, and Transwell assays, respectively. Sphere diameter was analyzed via sphere formation assay. Downstream mechanisms were examined using western blot, CCK8, flow cytometry, and Transwell assays. CD8 + T cell activity was examined using EdU, western blot, and flow cytometry. RESULTS: A total of 177 genes overlapped between GSE36295 DEGs and PDC000173 DEPs. CTSG was the hub gene identified by RF and LASSO. CTSG expression was significantly reduced in BC (P < 0.01). CTSG overexpression suppressed cell viability, proliferation, migration, invasion, sphere formation, and CD44 and CD133 expression (P < 0.01). CTSG up-regulation inhibited epidermal growth factor receptor (EGFR)/extracellular signal-regulated kinase (ERK) signaling axis and reduced cancer cell malignancy (P < 0.01). CTSG overexpression activated CD8 + T cells via EGFR/ERK inhibition, enhancing their cytotoxic effect on cancer cells (P < 0.01). CONCLUSION: CTSG inhibits BC malignancy and enhances CD8 + T cell function via EGFR/ERK inhibition.

Humans↗

In silico screening of anti-atherosclerotic compounds from Morus alba leaves by machine learning and network pharmacology.

OBJECTIVE: This study integrates machine learning with network pharmacology, molecular docking, and molecular dynamics simulations to screen bioactive compounds from Mulberry leaves and elucidate their potential mechanisms against atherosclerosis (AS). METHODS: A training dataset of anti-AS active compounds was compiled and encoded as Morgan fingerprints. Three machine learning classifiers, specifically Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XG-Boost), were constructed and evaluated using multiple performance metrics. Potential active components from Mulberry leaves and AS-related targets were retrieved, followed by protein-protein interaction network construction and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Molecular docking was then performed to evaluate binding affinities between core targets and candidate compounds, and the most stable complex was subjected to molecular dynamics simulations using GROMACS (2025). RESULTS: The RF model achieved superior performance (accuracy= 0.8354, F1 = 0.8408, AUC = 0.9119) with 100% external validation accuracy. Thirteen anti-AS candidates were prioritized from mulberry leaves, four of which have been previously documented. Network pharmacology revealed AKT1 and IL6 as core targets, enriched in pathways such as endocrine resistance. Molecular docking and dynamics simulations confirmed strong binding between oxysanguinarine and AKT1, with the complex exhibiting high stability. CONCLUSION: The RF model provides a reliable computational tool for prioritizing anti-AS compounds from Mulberry leaves. The integrated analysis reveals that Mulberry leaves exert anti-atherosclerotic effects through multi-target (e.g., AKT1, IL6) and multi-pathway (e.g., PI3K-Akt) mechanisms, offering a framework for further experimental validation.

Morus↗

Uncovering hub genes and key pathways responsive to drought stress in rice via meta-analysis of transcriptomic data.

Drought stress presents a formidable threat to global rice cultivation, triggering complex molecular responses that impact plant growth and productivity. To decipher the underlying gene expression dynamics, we performed a comprehensive meta-analysis of transcriptomic datasets derived from drought-tolerant rice genotypes. Via microarray data from three independent studies, we identified a set of consistently expressed differentially expressed genes (DEGs) under drought conditions. Integration of functional annotation tools, including GO and KEGG pathway enrichment, revealed key biological processes and signaling cascades involved in stress mitigation, such as ABA signaling, protein folding, and photosynthesis suppression. Protein-protein interaction (PPI) network construction, followed by hub gene identification via maximal clique centrality (MCC), highlighted pivotal regulators including LEA proteins, dehydrins, HSP70, and several transcription factors. Machine learning approaches further prioritize potential biomarkers, with Random Forest models achieving high classification accuracy and pinpointing key predictive genes. Chromosomal localization analysis provided spatial insights into the distribution of these hub genes, whose expression patterns were further compared against qRT-PCR data from previously published studies. This integrative approach identifies candidate genomic markers and mechanistic insights that may support future breeding strategies for drought-tolerant rice, pending experimental validation.

Cytoscape↗

Diagnosing scrapie in sheep: a classification experiment.

Scrapie is a neuro-degenerative disease in small ruminants. A data set of 3113 records of sheep reported to the Scrapie Notifications Database in Great Britain has been studied. Clinical signs were recorded as present/absent in each animal by veterinary officials (VO) and a post-mortem diagnosis was made. In an attempt to detect healthy animals within the set of suspects using only the clinical signs, 18 classification methods were applied ranging from simple linear classifiers to classifier ensembles such as Bagging, AdaBoost and Random Forests. The results suggest that the clinical classification by the VO was adequate as no further differentiation within the set of suspects was feasible.

Animals↗

Liquid Biopsy-Multiomics Link Adhesion Pathway Dysregulation to Kidney Injury Severity.

INTRODUCTION: Severe acute kidney injury (AKI) is strongly associated with the risk of developing chronic kidney disease; however, little is known about the cell type-specific mechanisms driving kidney injury severity. METHODS: In this multicenter observational study, we used clinically obtained liquid biopsy proteomics and machine learning (ML) to predict severe outcomes in patients with COVID-associated and non-COVID AKI. Further, we orthogonally combined 169 urine proteomics with 437 plasma proteomics samples and 40 urine sediment single-cell transcriptomics samples to identify complementary dysregulated mechanisms. RESULTS: Using a 10-fold cross-validated random forest algorithm, we identified a set of urinary proteins that demonstrate predictive power for both discovery and validation set with AUC of 87% and 76%, respectively. These predictive proteomics features obtained demonstrate that cell adhesion and autophagy-associated pathways are uniquely impacted in severe AKI. Differentially abundant proteins (DAPSs) associated with these pathways are highly expressed in cells of the juxtamedullary nephron, endothelial cells (ECs), and podocytes, indicating that these kidney cell types could be potential targets. Single-cell transcriptomic analysis in the in vitro model of kidney organoids infected with SARS-CoV-2 reveal dysregulation of extracellular matrix (ECM) organization in multiple nephron segments, recapitulating the clinically observed fibrotic response across multiomics datasets. Ligand-receptor interaction analysis of the podocyte and tubule organoid clusters shows significant reduction and loss of interaction between integrins and basement membrane receptors in the infected kidney organoids. CONCLUSION: Collectively, these data suggest that ECM degradation and adhesion-associated mechanisms could be the main driver of severe kidney injury.

AKI↗

Metagenomic insights into ecological risk of antibiotic resistome and mobilome in riverine plastisphere under impact of urbanization.

Microplastics (MPs) are of increasing concern due to their role as reservoirs for antibiotic resistance genes (ARGs) and pathogens. To date, few studies have explored the influence of anthropogenic activities on ARGs and mobile genetic elements (MGEs) within various riverine MPs, in comparison to their natural counterparts. Here an in-situ incubation was conducted along heavily anthropogenically-impacted Houxi River to characterize the geographical pattern of antibiotic resistome, mobilome and pathogens inhabiting MPs- and leaf-biofilms. The metagenomics result showed a clear urbanization-driven profile in the distribution of ARGs, MGEs and pathogens, with their abundances sharply increasing 4.77 to 19.90 times from sparsely to densely populated regions. The significant correlation between human fecal marker crAssphage and ARG (R2&#xa0;=&#xa0;0.67, P=0.003) indicated the influence of anthropogenic activity on ARG proliferation in plastisphere and natural leaf surfaces. And mantel tests and random forest analysis revealed the impact of 17 socio-environmental factors, e.g., population density, antibiotic concentrations, and pore volume of materials, on the dissemination of ARGs. Partial least squares-path modeling further unveiled that intensifying human activities not only directly boosted ARGs abundance but also exerted a comparable indirect impact on ARGs propagation. Furthermore, the polyvinylchloride plastisphere created a pathogen-friendly habitat, harboring higher abundances of ARGs and MGEs, while polylactic acid are not likely to serve as vectors for pathogens in river, with a lower resistome risk score than that in leaf-biofilms. This study highlights the diverse ecological risks associated with the dissemination of ARGs and pathogens in varied MPs, offering insights for the policymaking of usage and control of plastics within urbanization.

Urbanization↗

Temporal shifts in gyrA mutation types and sublineage replacement in ST11 Salmonella enterica&#xa0;serovar Enteritidis over a decade (2014-2023): A genomic epidemiological study in Guangxi, China.

The overuse or abuse of antibiotics drives the global health threat of antimicrobial resistance. Although bans on certain veterinary antibiotics, such as colistin, have proven effective, the impact of fluoroquinolone stewardship on the evolution of the foodborne pathogen Salmonella enterica serovar Enteritidis (S. Enteritidis) remains unclear. Here, we conducted a decade-long (2014-2023) retrospective longitudinal genomic epidemiological analysis of 441&#xa0;ST11 S. Enteritidis isolates from Guangxi, China, alongside a global reference dataset of 4297 genomes. Our aim was to elucidate the effect of real-world antibiotic stewardship on the shift of gyrA point mutations and lineage distribution. Surveillance identified three global epidemic clade sublineages (GEC-L2, L3, L4), with the multidrug-resistant GEC-L4 (i.e., GC-c or MMC2), characterized by the gyrA mutation with amino acid substitution D87Y, being domestically dominant (70.07%, 309/441). Following China's 2016 ban on the veterinary use of critical fluoroquinolones, the proportion of the highly resistant GEC-L4 sublineage decreased continuously (from 86.84% in 2017 to 56.00% in 2023), while the less resistant GEC-L3 sublineage (i.e., GC-b or MMC1), mainly characterized by gyrA D87G, increased simultaneously (from 13.16% to 44.00%). This phenomenon might be attributed to the fact that the GEC-L4 sublineage exhibited a higher fitness cost compared with the GEC-L3 sublineage, as confirmed by the competition assay. A Random Forest Model validated that the gyrA mutation with amino acid substitution&#xa0;D87Y was the paramount feature for these sublineages' identification. In contrast, global data showed a continuous increase in gyrA mutations (from 8.63% in 2006 to 68.85% in 2024), primarily D87Y (from 1.44% to 31.15%) and D87N (from 4.32% to 22.95%), correlating with rising average fluoroquinolone consumption. This study provides direct genomic evidence that national-level antibiotic stewardship can drive the replacement of highly resistant sublineages with moderately resistant ones. These findings offer crucial scientific evidence for evaluating the impact of antibiotic management policies and inform strategies for the rational use of antimicrobials.

China↗

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

PURPOSE: The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results. We performed multiple modeling experiments integrating clinical and demographic data from electronic health records with genetic data to understand which decisions may affect performance. METHODS: Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from 2 large independent health systems, and polygenic risk scores (PRS) were generated across all patients of European ancestry with genetic data in the corresponding biobanks. Crohn's disease was studied based on its substantial genetic component, established electronic health records-based definition, and sufficient prevalence for training and testing. We investigated the impact of choices regarding the PRS integration method, training sample, model complexity, and performance metrics. RESULTS: Overall, our results showed that including PRS resulted in higher performance, but this gain was only robust in situations with limited clinical information. We found consistent performance increases from more compute-intensive models, such as random forest, but the impact of other decisions varied by site. CONCLUSION: This work highlights the importance of considering methodological decision points in interpreting the impact of PRS on prediction performance in clinical models.

Humans↗

Development and validation of a machine learning prognostic model based on an epigenomic signature in patients with pancreatic ductal adenocarcinoma.

BACKGROUND: In Pancreatic Ductal Adenocarcinoma (PDAC), current prognostic scores are unable to fully capture the biological heterogeneity of the disease. While some approaches investigating the role of multi-omics in PDAC are emerging, the analysis of methylation data is under exploited. MATERIALS AND METHODS: We analyzed CpG sites from two publicly available datasets, the TCGA-PAAD used as discovery set and the CPTAC-PDA as external test set. Single mutations and co-mutation of KRAS and TP53 genes were identified as targets, and differentially methylated CpG sites (DMC) were detected accordingly. We trained and validated Random Forest (RF) models to predict each target. Area Under the Receiver Operating Characteristic curve (AUROC) and Area Under the Precision-Recall curve (AUPRC) were used as performance metrics. Then, we performed consensus clustering from the DMCs to identify novel patients' profiles. Finally, we trained and validated a combination of eXtreme Gradient Boosting (XGB) and tree models to select an epigenomic prognostic determinant. RESULTS: From 598 DMCs extracted, an RF model predicted KRAS and TP53 co-mutation on the external test set with AUROC of 0.77 and AUPRC of 0.87. The consensus clustering allowed us to identify 4 clusters (C1, C2, C3, and C4) of patients. The C4 cluster captured a subgroup of patients with favorable Overall Survival (OS) with respect to others. The XGB model perfectly predicted C4 vs other clusters on the discovery set. In both cohorts, patients were stratified into two risk groups according to methylation levels of cg16854533, individuated as the most important CpG site. CONCLUSION: We analyzed methylation data to develop a classifier for the TP53 and KRAS mutational status. Four prognostic clusters were pointed out and a prognostic model using a CpG site was validated in an independent cohort. Our results evidence that the proposed use of methylation data facilitates risk stratification for PDAC.

Humans↗

Machine learning-assisted plasma PEA proteomics enables differential diagnosis of melancholic depression and bipolar disorder.

Differentiating bipolar disorder (BD) from major depressive disorder (MDD) remains a critical unmet need in psychiatry due to overlapping clinical presentations and the absence of reliable biological markers. In this study, we assessed the capacity of multivariate machine learning models to accurately differentiate BD from MDD with melancholic features using plasma proteomic profiles obtained via Proximity Extension Assay (PEA) technology. A total of 67 participants were included (23 BD, 20 MDD, and 24 HC), and plasma protein expression was assessed using the Olink Target 96 Neurology panel. Differential proteomic analysis revealed distinct disorder-specific expression patterns, identifying 21 differentially expressed proteins in BD versus MDD, 18 in BD versus healthy controls, and 7 in MDD versus healthy controls. Using a stepwise feature reduction strategy, machine learning models were trained on three feature sets comprising all proteins, the top 20 most informative proteins, and the top 5 most beneficial proteins, and evaluated across BD-MDD, BD-HC, and MDD-HC classification tasks using five algorithms. For BD-MDD discrimination, the Random Forest model achieved the highest performance when trained on the top 5 protein set (LXN, HAGH, MATN3, PLXNB1, and CTSC), yielding an AUC of 0.905, with similarly strong performance observed using the top 20 protein set. Feature importance analysis highlighted proteins involved in neurodevelopmental processes, immune regulation, and extracellular matrix organization. Overall, these findings demonstrate that integrating plasma proteomics with machine learning enables robust differentiation between BD and MDD with melancholic features, supporting the development of scalable and biologically informed diagnostic tools for precision psychiatry.

Bipolar disorder↗

Mining the structural genomics pipeline: identification of protein properties that affect high-throughput experimental analysis.

Structural genomics projects represent major undertakings that will change our understanding of proteins. They generate unique datasets that, for the first time, present a standardized view of proteins in terms of their physical and chemical properties. By analyzing these datasets here, we are able to discover correlations between a protein's characteristics and its progress through each stage of the structural genomics pipeline, from cloning, expression, purification, and ultimately to structural determination. First, we use tree-based analyses (decision trees and random forest algorithms) to discover the most significant protein features that influence a protein's amenability to high-throughput experimentation. Based on this, we identify potential bottlenecks in various stages of the structural genomics process through specialized "pipeline schematics". We find that the properties of a protein that are most significant are: (i.) whether it is conserved across many organisms; (ii). the percentage composition of charged residues; (iii). the occurrence of hydrophobic patches; (iv). the number of binding partners it has; and (v). its length. Conversely, a number of other properties that might have been thought to be important, such as nuclear localization signals, are not significant. Thus, using our tree-based analyses, we are able to identify combinations of features that best differentiate the small group of proteins for which a structure has been determined from all the currently selected targets. This information may prove useful in optimizing high-throughput experimentation. Further information is available from http://mining.nesg.org/.

Algorithms↗

Integrated Multi-Omics Analyses Reveal Lipid Metabolic Signature in Osteoarthritis.

Osteoarthritis (OA) is the most common degenerative joint disease and the second leading cause of disability worldwide. Single-omics analyses are far from elucidating the complex mechanisms of lipid metabolic dysfunction in OA. This study identified a shared lipid metabolic signature of OA by integrating metabolomics, single-cell and bulk RNA-seq, as well as metagenomics. Compared to the normal counterparts, cartilagesin OA patients exhibited significant depletion of homeostatic chondrocytes (HomCs) (P&#xa0;=&#xa0;0.03) and showed lipid metabolic disorders in linoleic acid metabolism and glycerophospholipid metabolism which was consistent with our findings obtained from plasma metabolomics. Through high-dimensional weighted gene co-expression network analysis (hdWGCNA), weidentified PLA2G2A as a hub gene associated with lipid metabolic disorders in HomCs. And an OA-associated subtype of HomCs, namely HomC1 (marked by PLA2G2A, MT-CO1, MT-CO2, and MT-CO3) was identified, which also exhibited abnormal activation of lipid metabolic pathways. This suggests the involvement of HomC1 in OA progression through the shared lipid metabolism aberrancies, which were further validated via bulk RNA-Seq analysis. Metagenomic profiling identified specific gut microbial species significantly associated with the key lipid metabolism disorders, including Bacteroides uniformis (P&#xa0;<&#xa0;0.001, R&#xa0;=&#xa0;-0.52), Klebsiella pneumonia (P&#xa0;=&#xa0;0.003, R&#xa0;=&#xa0;0.42), Intestinibacter_bartlettii (P&#xa0;=&#xa0;0.009, R&#xa0;=&#xa0;0.38), and Streptococcus anginosus (P&#xa0;=&#xa0;0.009, R&#xa0;=&#xa0;0.38). By integrating the multi-omics features, a random forest diagnostic model with outstanding performance was developed (AUC&#xa0;=&#xa0;0.97). In summary, this study deciphered the crucial role of a integrated lipid metabolic signature in OA pathogenesis, and established a regulatory axis of gut microbiota-metabolites-cell-gene, providing new insights into the gut-joint axis and precision therapy for OA.

Humans↗