Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “random forest”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Information assessment on predicting protein-protein interactions.

BACKGROUND: Identifying protein-protein interactions is fundamental for understanding the molecular machinery of the cell. Proteome-wide studies of protein-protein interactions are of significant value, but the high-throughput experimental technologies suffer from high rates of both false positive and false negative predictions. In addition to high-throughput experimental data, many diverse types of genomic data can help predict protein-protein interactions, such as mRNA expression, localization, essentiality, and functional annotation. Evaluations of the information contributions from different evidences help to establish more parsimonious models with comparable or better prediction accuracy, and to obtain biological insights of the relationships between protein-protein interactions and other genomic information. RESULTS: Our assessment is based on the genomic features used in a Bayesian network approach to predict protein-protein interactions genome-wide in yeast. In the special case, when one does not have any missing information about any of the features, our analysis shows that there is a larger information contribution from the functional-classification than from expression correlations or essentiality. We also show that in this case alternative models, such as logistic regression and random forest, may be more effective than Bayesian networks for predicting interactions. CONCLUSIONS: In the restricted problem posed by the complete-information subset, we identified that the MIPS and Gene Ontology (GO) functional similarity datasets as the dominating information contributors for predicting the protein-protein interactions under the framework proposed by Jansen et al. Random forests based on the MIPS and GO information alone can give highly accurate classifications. In this particular subset of complete information, adding other genomic data does little for improving predictions. We also found that the data discretizations used in the Bayesian methods decreased classification performance.

Artificial Intelligence↗

Metabolomics Reveals Metabolic Characteristics of Functional Cure in Chronic Hepatitis B Treated With Entecavir Combined With Pegylated Interferon Alpha.

BACKGROUND: Entecavir (ETV) combined with pegylated interferon alpha (PEG-IFNα) improves chronic hepatitis B (CHB) functional cure rates, but therapeutic heterogeneity and underlying metabolic mechanisms remain unclear. This study used untargeted metabolomics to identify metabolic signatures, mechanisms, and predictive biomarkers of functional cure with ETV-PEG-IFNα. METHODS: Thirty-eight CHB patients were grouped into ETV monotherapy (Group E, n = 12) and ETV-PEG-IFNα combination therapy (Group Z, n = 26); Group Z was subdivided into cured (Group A, n = 13) and noncured (Group B, n = 13). Serum metabolomic profiling, multivariate statistics, and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis identified differential metabolites. A random forest model was built using key metabolites. RESULTS: Three hundred eighty-eight metabolites were identified. Four differential metabolites distinguished Group A and B (upregulated guanidinoacetic acid, uracil 5-carboxylate; downregulated L-methionine S-oxide, oleamide), enriching amino acid metabolism pathways. Nine differential metabolites between Group E and Z implicated amino acid, immune, and fatty acid pathways. The random forest model based on the four Group A/B metabolites showed 88.5% cross-validation accuracy (AUC = 0.920), with L-methionine S-oxide and oleamide as key predictors. CONCLUSIONS: This study reveals metabolic rewiring in CHB functional cure via ETV-PEG-IFNα therapy, involving energy metabolism, oxidative stress, and immunomodulation, based on which we propose a tentative metabolism-immunity synergy model to guide future research. Key metabolites, especially L-methionine S-oxide and oleamide, show exploratory predictive potential for functional cure that warrants further validation in independent cohorts.

Humans↗

Deciphering microbial and metabolic influences in gastrointestinal diseases-unveiling their roles in gastric cancer, colorectal cancer, and inflammatory bowel disease.

INTRODUCTION: Gastrointestinal disorders (GIDs) affect nearly 40% of the global population, with gut microbiome-metabolome interactions playing a crucial role in gastric cancer (GC), colorectal cancer (CRC), and inflammatory bowel disease (IBD). This study aims to investigate how microbial and metabolic alterations contribute to disease development and assess whether biomarkers identified in one disease could potentially be used to predict another, highlighting cross-disease applicability. METHODS: Microbiome and metabolome datasets from Erawijantari et al. (GC: n = 42, Healthy: n = 54), Franzosa et al. (IBD: n = 164, Healthy: n = 56), and Yachida et al. (CRC: n = 150, Healthy: n = 127) were subjected to three machine learning algorithms, eXtreme gradient boosting (XGBoost), Random Forest, and Least Absolute Shrinkage and Selection Operator (LASSO). Feature selection identified microbial and metabolite biomarkers unique to each disease and shared across conditions. A microbial community (MICOM) model simulated gut microbial growth and metabolite fluxes, revealing metabolic differences between healthy and diseased states. Finally, network analysis uncovered metabolite clusters associated with disease traits. RESULTS: Combined machine learning models demonstrated strong predictive performance, with Random Forest achieving the highest Area Under the Curve(AUC) scores for GC(0.94[0.83-1.00]), CRC (0.75[0.62-0.86]), and IBD (0.93[0.86-0.98]). These models were then employed for cross-disease analysis, revealing that models trained on GC data successfully predicted IBD biomarkers, while CRC models predicted GC biomarkers with optimal performance scores. CONCLUSION: These findings emphasize the potential of microbial and metabolic profiling in cross-disease characterization particularly for GIDs, advancing biomarker discovery for improved diagnostics and targeted therapies.

Humans↗

Chemoinformatics-based classification of prohibited substances employed for doping in sport.

Representative molecules from 10 classes of prohibited substances were taken from the World Anti-Doping Agency (WADA) list, augmented by molecules from corresponding activity classes found in the MDDR database. Together with some explicitly allowed compounds, these formed a set of 5245 molecules. Five types of fingerprints were calculated for these substances. The random forest classification method was used to predict membership of each prohibited class on the basis of each type of fingerprint, using 5-fold cross-validation. We also used a k-nearest neighbors (kNN) approach, which worked well for the smallest values of k. The most successful classifiers are based on Unity 2D fingerprints and give very similar Matthews correlation coefficients of 0.836 (kNN) and 0.829 (random forest). The kNN classifiers tend to give a higher recall of positives at the expense of lower precision. A naïve Bayesian classifier, however, lies much further toward the extreme of high recall and low precision. Our results suggest that it will be possible to produce a reliable and quantitative assignment of membership or otherwise of each class of prohibited substances. This should aid the fight against the use of bioactive novel compounds as doping agents, while also protecting athletes against unjust disqualification.

Algorithms↗

Simple statistical models predict C-to-U edited sites in plant mitochondrial RNA.

BACKGROUND: RNA editing is the process whereby an RNA sequence is modified from the sequence of the corresponding DNA template. In the mitochondria of land plants, some cytidines are converted to uridines before translation. Despite substantial study, the molecular biological mechanism by which C-to-U RNA editing proceeds remains relatively obscure, although several experimental studies have implicated a role for cis-recognition. A highly non-random distribution of nucleotides is observed in the immediate vicinity of edited sites (within 20 nucleotides 5' and 3'), but no precise consensus motif has been identified. RESULTS: Data for analysis were derived from the the complete mitochondrial genomes of Arabidopsis thaliana, Brassica napus, and Oryza sativa; additionally, a combined data set of observations across all three genomes was generated. We selected datasets based on the 20 nucleotides 5' and the 20 nucleotides 3' of edited sites and an equivalently sized and appropriately constructed null-set of non-edited sites. We used tree-based statistical methods and random forests to generate models of C-to-U RNA editing based on the nucleotides surrounding the edited/non-edited sites and on the estimated folding energies of those regions. Tree-based statistical methods based on primary sequence data surrounding edited/non-edited sites and estimates of free energy of folding yield models with optimistic re-substitution-based estimates of approximately 0.71 accuracy, approximately 0.64 sensitivity, and approximately 0.88 specificity. Random forest analysis yielded better models and more exact performance estimates with approximately 0.74 accuracy, approximately 0.72 sensitivity, and approximately 0.81 specificity for the combined observations. CONCLUSIONS: Simple models do moderately well in predicting which cytidines will be edited to uridines, and provide the first quantitative predictive models for RNA edited sites in plant mitochondria. Our analysis shows that the identity of the nucleotide -1 to the edited C and the estimated free energy of folding for a 41 nt region surrounding the edited C are the most important variables that distinguish most edited from non-edited sites. However, the results suggest that primary sequence data and simple free energy of folding calculations alone are insufficient to make highly accurate predictions.

Arabidopsis↗

A comparative study of discriminating human heart failure etiology using gene expression profiles.

BACKGROUND: Human heart failure is a complex disease that manifests from multiple genetic and environmental factors. Although ischemic and non-ischemic heart disease present clinically with many similar decreases in ventricular function, emerging work suggests that they are distinct diseases with different responses to therapy. The ability to distinguish between ischemic and non-ischemic heart failure may be essential to guide appropriate therapy and determine prognosis for successful treatment. In this paper we consider discriminating the etiologies of heart failure using gene expression libraries from two separate institutions. RESULTS: We apply five new statistical methods, including partial least squares, penalized partial least squares, LASSO, nearest shrunken centroids and random forest, to two real datasets and compare their performance for multiclass classification. It is found that the five statistical methods perform similarly on each of the two datasets: it is difficult to correctly distinguish the etiologies of heart failure in one dataset whereas it is easy for the other one. In a simulation study, it is confirmed that the five methods tend to have close performance, though the random forest seems to have a slight edge. CONCLUSIONS: For some gene expression data, several recently developed discriminant methods may perform similarly. More importantly, one must remain cautious when assessing the discriminating performance using gene expression profiles based on a small dataset; our analysis suggests the importance of utilizing multiple or larger datasets.

Data Interpretation, Statistical↗

Radiogenomic MRI biomarkers for noninvasive prediction of GPC3 expression and tumor microenvironment in hepatocellular carcinoma.

BACKGROUND: Glypican-3 (GPC3) is frequently overexpressed in hepatocellular carcinoma (HCC) and plays a key role in immune and metabolic remodeling of the tumor microenvironment. Reliable noninvasive biomarkers for predicting GPC3 status could improve patient stratification and support precision immunotherapy. METHODS: This multicenter retrospective study included 274 patients with pathologically confirmed hepatocellular carcinoma from three institutions, 34 external cases with MRI from The Cancer Imaging Archive, and 363 transcriptomic profiles from The Cancer Genome Atlas. Contrast-enhanced T1-weighted imaging and diffusion-weighted imaging were analyzed. Tumor and peritumoral regions were segmented manually and radiomic features extracted using PyRadiomics. Feature selection was performed with correlation filtering and least absolute shrinkage and selection operator regression. Machine learning classifiers including logistic regression, random forest, support vector machine, k-nearest neighbor, and decision tree were trained with 10-fold cross-validation and tested on independent external cohorts. A radiomics score was calculated for each patient. Radiogenomic analysis correlated radiomics scores with transcriptomic data using weighted gene co-expression network analysis. Hub genes and enriched pathways were identified, and immune infiltration and predicted immunotherapy response were assessed using computational methods. RESULTS: The random forest model using contrast-enhanced T1-weighted imaging achieved an area under the curve of 0.966 in training and 0.935 in internal validation. The integrated contrast-enhanced T1-weighted imaging plus diffusion-weighted imaging model reached an internal validation area under the curve of 0.979. In external testing, the best performance was obtained with a support vector machine model (area under the curve 0.756). Radiomics scores were significantly correlated with GPC3 expression (R&#x2009;=&#x2009;0.78, p&#x2009;<&#x2009;0.05). Transcriptomic analysis identified a 10-gene signature enriched in hypoxia and lipid metabolism pathways that stratified patients into prognostic subgroups (concordance index 0.720, hazard ratio 4.07, p&#x2009;<&#x2009;0.0001). High-risk patients had greater immune infiltration and a lower predicted immune evasion score, suggesting a potential benefit from immunotherapy. CONCLUSIONS: MRI-based radiomics models can noninvasively predict GPC3 expression in hepatocellular carcinoma. Radiomics scores reflect underlying hypoxia and lipid metabolism pathways and stratify patients by prognosis and predicted immunotherapy response. These findings support radiogenomics as a translational approach to imaging-guided precision treatment in hepatocellular carcinoma.

Humans↗

Identification of Immune Response-Related Proteomic Biomarkers in Moyamoya Disease Using Serum Olink Proteomics.

Moyamoya disease, a rare chronic cerebrovascular disorder, requires invasive digital subtraction angiography (DSA) for diagnosis. This study employed high-throughput proteomics to identify plasma biomarkers for Moyamoya disease diagnosis. We conducted immunopanel analysis using the Olink platform to evaluate 92 immune-related proteins in plasma samples from 88 Moyamoya disease patients and 88 healthy controls. Key proteins were identified through differential expression analysis, GO, and KEGG enrichment analysis. A diagnostic model was constructed using LASSO regression, Boruta algorithm, and machine learning models including random forest and XGBoost. Validation of these proteins was performed using GEO external data sets, followed by prediction of potential therapeutic drugs and molecular docking validation through pharmacogenomic databases. A total of 44 differentially expressed proteins were identified through the Olink immunopanel, with 12 downregulated and 32 upregulated. GO and KEGG analyses revealed significant enrichment of these proteins in innate immune responses and signaling pathways such as NF-kB and MAPK. Through LASSO, random forest, and protein under-area analysis, four potential biomarkers for Moyamoya disease (MGMT, SIT1, PRDX1, TRAF2) were identified. A diagnostic model using these proteins showed the highest AUC value with the XGBoost model. Additionally, TRAF2 and PRDX1 exhibited significant expression differences in Moyamoya disease patients within the GEO data set. Our study revealed the immune landscape of Moyamoya disease, identified four biomarkers, and established a variety of diagnostic models.

Humans↗

Genomic selection in timothy (Phleum pratense L.): a comprehensive evaluation of prediction models, multi-trait strategies, and forward validation across Norwegian environments.

This study presents a comprehensive evaluation of genomic selection (GS) in timothy (Phleum pratense L.), comparing nine prediction models across yield and quality traits at two Norwegian locations. Forward validation with independent full-sib (FS2) families revealed a substantial generalization gap, highlighting the need for realistic accuracy assessment in polyploid forage breeding. Timothy (Phleum pratense L.) is the most important forage grass in Northern Europe, yet genomic selection has not been systematically evaluated in this hexaploid species. We assessed 889 FS2-families originating from biparental crosses among 49 cultivars/populations. The FS2-families were genotyped with 30,698 SNP markers derived from genotyping-by-sequencing (GBS) and field tested for three harvest years at a highland and a lowland continental location in Southern Norway. Nine genomic prediction models were compared for six yield traits (dry matter yield per cut and total) and six quality traits (protein, digestibility, and fiber fractions) across three cuts/year. Within-training cross-validation accuracies were moderate to high (mean r = 0.62), with Random Forest and SVR consistently outperforming GBLUP. However, forward validation using 213 independent FS2-families revealed dramatically lower accuracies (mean r = 0.16), with only 16 of 30 trait-dataset combinations reaching statistical significance (p < 0.05). Genomic heritabilities (GREML), estimated across environments, ranged from near zero for the quality traits to 0.55 for the yield traits. Multi-trait models improved accuracy by 3-5% over single-trait approaches, while FS2 families-by-environment interaction models with Random Forest achieved the highest within-training accuracy (mean r = 0.71). Marker density analysis showed accuracy plateauing at approximately 15000 SNPs. Genetic correlations among the yield component traits were estimated by multi-trait REML; correlations among the quality traits could not be estimated reliably because their genomic heritabilities were low. A multi-trait selection index identified top-performing FS2-families for further crossing recommendations. These results provide a benchmark for GS implementation in hexaploid timothy and emphasize that cross-validation substantially overestimates prediction accuracy for truly independent material.

Norway↗

Integrating Genomic and Nongenomic Data to Stratify the Risk of Contralateral Breast Cancer After Radiation Therapy.

PURPOSE: Women treated with radiation therapy (RT) for breast cancer have an increased risk of developing radiation-associated contralateral breast cancer (CBC). Predicting CBC events is challenging because of the complex interplay of genomic, treatment, personal, and clinical factors. This study investigated computational methods that integrate genome-wide single-nucleotide polymorphisms and nongenomic data to develop a risk stratification model for developing CBC in women treated with RT for their first primary breast cancer. METHODS AND MATERIALS: This study used a subset of the population-based Women's Environmental Cancer and Radiation Epidemiology study that included 633 CBC cases and 1253 individually matched unilateral breast cancer controls who were treated with RT and had single-nucleotide polymorphism data available from a genome-wide association study. The study population was split into training, validation, and test sets for rigorous modeling and validation. Three data integration methods were compared in terms of their ability to stratify CBC risk: (1) naive integration; (2) sequential integration; and (3) sequential iterative integration. A biological analysis of the final model was performed using gene set enrichment analysis and protein-protein interaction analysis with gene annotation information informed by the model. RESULTS: The best-performing integration method was the sequential iterative integration equipped with the mixed-effect random forest algorithm. This approach achieved an area under the curve of 0.64 to stratify CBC risk in the test set, representing moderate predictive power. Calibration analysis showed good agreement between the lowest and highest risk bins stratified using sorted predicted values in the test set, resulting in an odds ratio of 3.27 for both predicted and observed CBC occurrence. Gene set enrichment analysis and protein-protein interaction analysis revealed that genes with high importance scores were associated with pathways relevant to lipid and fatty acid metabolism as well as breast cancer sensitivity to tamoxifen. CONCLUSIONS: The mixed-effect random forest approach demonstrated the potential for integrating high-dimensional genomic and low-dimensional nongenomic data to stratify CBC risk.

Humans↗

CD4+CD8+ double-positive T cells are associated with severity of tuberculosis.

BACKGROUND: Tuberculosis (TB) remains a global public health burden, and how immune cell subsets regulate host anti-TB immunity and disease progression remains incompletely understood. While previous studies have focused on single-positive (SP) T cells (CD4+ or CD8+) in TB pathogenesis, the association between CD4+CD8+ double-positive (DP) T cells and TB susceptibility, severity, and treatment outcomes have not been fully elucidated. This study aimed to investigate the relationship between DP T cells and other immune cell subsets with TB, and to explore the potential diagnostic and prognostic value of DP T cells in active TB. METHODS: A Genome-Wide Association Study (GWAS) was conducted to analyze 731 immune cell traits and a dataset encompassing 895 patients with TB. Subsequently, a cohort including 647 patients with active TB and 632 healthy controls was used to verify the findings of Mendelian randomization (MR). The correlation between the percentage of DP T cells in lymphocytes and TB severity, treatment efficacy, and Mycobacterium tuberculosis (Mtb)-specific IFN-&#x3b3; production was evaluated. Finally, a random forest model incorporating the percentage of DP T cells in leukocytes and other peripheral blood parameters was constructed to distinguish severe from mild active TB. RESULTS: MR analysis suggested potential causal links between the percentage of DP T cells among peripheral leukocytes and TB status. Clinical sample validation showed that the percentage of peripheral DP T cell among leukocytes was significantly lower in patients with active TB than in healthy controls (P < 0.001), and was inversely correlated with disease severity. Additionally, the percentage of DP T cells in leukocytes was positively correlated with Mtb-specific antigen-stimulated IFN-&#x3b3; production. Flow cytometric analysis demonstrated that DP T cells had a significantly higher frequency of IFN-&#x3b3;-expressing cells compared to CD8+ SP T cells (P < 0.001). The constructed random forest model effectively distinguished severe from mild TB, with good diagnostic performance (AUC&#xa0;=&#xa0;0.985). CONCLUSIONS: Our findings indicate that DP T cells are closely associated with TB severity, and are positively associated with Mtb-specific IFN-&#x3b3; response. The percentage of peripheral DP T cells in leukocytes could serve as a potential non-invasive biomarker for TB severity stratification.

Humans↗

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-&#x3ba;B, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-&#x3ba;B, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4↗

Immunogenetic risk and protective factors for juvenile dermatomyositis in Caucasians.

OBJECTIVE: To define the relative importance (RI) of class II major histocompatibility complex (MHC) alleles and peptide binding motifs as risk or protective factors for juvenile dermatomyositis (DM), and to compare these with HLA associations in adult DM. METHODS: DRB1 and DQA1 typing was performed in 142 Caucasian patients with juvenile DM, and the results were compared with HLA typing data from 193 patients with adult DM and 797 race-matched controls. Random Forests classification and multiple logistic regression were used to assess the RI of the HLA associations. RESULTS: The HLA-DRB1*0301 allele was a primary risk factor (odds ratio [OR] 3.9), while DQA1*0301 (OR 2.8), DQA1*0501 (OR 2.1), and homozygosity for DQA1*0501 (OR 3.2) were additional risk factors for juvenile DM. These risk factors were not present in patients with adult DM without defined autoantibodies. DQA1 alleles *0201 (OR 0.37), *0101 (OR 0.38), and *0102 (OR 0.51) were identified as novel protective factors for juvenile DM, the latter 2 also being protective factors in adult DM. The peptide binding motif DRB1 (9)EYSTS(13) was a risk factor, and DQA1 motifs F(25), S(26), and (45)(V/A)W(R/K)(47) were protective. Random Forests classification analysis revealed that among the identified risk factors for juvenile DM, DRB1*0301 had a higher RI (100%) than DQA1*0301 (RI 57%), DQA1*0501 (RI 42%), or the peptide binding motifs. In a logistic regression model, DRB1*0301 and DQA1*0201 were the strongest risk and protective factors, respectively, for juvenile DM. CONCLUSION: DRB1*0301 is ranked higher in RI than DQA1*0501 as a risk factor for juvenile DM. DQA1*0301 is a newly identified HLA risk factor for juvenile DM, while 3 of the DQA1 alleles studied are newly identified protective factors for juvenile DM.

Adolescent↗

Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategy.

BACKGROUND: Understanding how amino acid substitutions affect protein functions is critical for the study of proteins and their implications in diseases. Although methods have been developed for predicting potential effects of amino acid substitutions using sequence, three-dimensional structural, and evolutionary properties of proteins, the applications are limited by the complication of the features and the availability of protein structural information. Another limitation is that the prediction results are hard to be interpreted with physicochemical principles and biological knowledge. RESULTS: To overcome these limitations, we proposed a novel feature set using physicochemical properties of amino acids, evolutionary profiles of proteins, and protein sequence information. We applied the support vector machine and the random forest with the feature set to experimental amino acid substitutions occurring in the E. coli lac repressor and the bacteriophage T4 lysozyme, as well as to annotated amino acid substitutions occurring in a wide range of human proteins. The results showed that the proposed feature set was superior to the existing ones. To explore physicochemical principles behind amino acid substitutions, we designed a simulated annealing bump hunting strategy to automatically extract interpretable rules for amino acid substitutions. We applied the strategy to annotated human amino acid substitutions and successfully extracted several rules which were either consistent with current biological knowledge or providing new insights for the understanding of amino acid substitutions. When applied to unclassified data, these rules could cover a large portion of samples, and most of the covered samples showed good agreement with predictions made by either the support vector machine or the random forest. CONCLUSION: The prediction methods using the proposed feature set can achieve larger AUC (the area under the ROC curve), smaller BER (the balanced error rate), and larger MCC (the Matthews' correlation coefficient) than those using the published feature sets, suggesting that our feature set is superior to the existing ones. The rules extracted by the simulated annealing bump hunting strategy have comparable coverage and accuracy but much better interpretability as those extracted by the patient rule induction method (PRIM), revealing that the strategy is more effective in inducing interpretable rules.

Amino Acid Sequence↗

Transcriptome Analysis and Experimental Validation of Palmitoylation- Related Biomarkers in Atherosclerosis.

INTRODUCTION: Protein palmitoylation contributes to membrane localisation, signal transduction, and cell-fate regulation. It is closely associated with lipid metabolic dysfunction, immune inflammation, and vascular remodelling in atherosclerosis (AS). However, key palmitoylation-related transcriptomic markers and their potential causal associations with AS remain incompletely defined. METHODS: The Gene Expression Omnibus (GEO) dataset GSE100927 was used as the training cohort, and GSE43292 was used as an external validation cohort. Differentially expressed genes were identified using limma and intersected with palmitoylation-related genes to obtain palmitoylation-related differentially expressed genes (PRDEGs). Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were then performed using clusterProfiler. Two-sample Mendelian randomisation was used to evaluate potential causal relationships between characteristic genes and AS. Feature selection was conducted using random forest and support vector machine recursive feature elimination (SVM-RFE), and the overlapping genes selected by both methods were retained. Receiver operating characteristic (ROC) curves were used to assess diagnostic performance. A five-gene nomogram was constructed, and its clinical utility was evaluated using calibration curves and decision curve analysis (DCA). Gene set variation analysis (GSVA) was applied to compare pathway activity between high- and low-expression groups for each core gene. Single-cell analysis using Seurat and expression-based cell-cell communication analysis using CellChat were conducted with GSE159677, and upstream transcription factors were predicted using NetworkAnalyst. For in vivo validation, an AS model was established in ApoE&#x2078;/&#x2078; mice fed a high-fat diet, and aortic gene and protein expression were assessed by RT-qPCR and western blotting. RESULTS: In GSE100927, 51 PRDEGs were identified. GO and KEGG enrichment analyses highlighted pathways associated with regulation of monoatomic ion transport, sarcomere and myofibril organisation, and immune inflammation. Mendelian randomisation suggested a potential protective causal association between SLC7A7 and AS. By integrating MR with random forest and SVM-RFE feature selection, we prioritised five core genes: PLCB2, GMIP, NEXN, PLN, and SLC7A7. These genes showed good diagnostic performance in GSE43292. The resulting nomogram was well calibrated and demonstrated stable net benefit in decision curve and clinical impact curve analyses. Single-gene GSVA identified consistently activated pathways across multiple genes, including innate and adaptive immune recognition, calcium signalling and myocardial contraction/cardiomyopathy, extracellular matrix-receptor interaction, cell junction pathways, autophagy-lysosome pathways, and several metabolic programmes. At the single-cell level, PLCB2 and GMIP were predominantly expressed in T cells and macrophages, NEXN and PLN were enriched in vascular smooth muscle cells, and SLC7A7 was mainly expressed in macrophages. CellChat analysis indicated increased signals for immune-related ligand-receptor interactions. In ApoE&#x2078;/&#x2078; mice fed a high-fat diet, PLCB2, GMIP, and SLC7A7 were upregulated, whereas NEXN and PLN were downregulated; protein-level changes were concordant with the transcriptomic trends. DISCUSSION: These findings indicate that palmitoylation-related dysregulation in AS converges on immune inflammation, calcium signalling/contractile programmes, ECM remodelling, and autophagy-linked metabolism. The five-gene panel is supported by external validation, single-cell localisation to immune and vascular compartments, and concordant results in ApoE&#x2078;/&#x2078; mice. CONCLUSION: This study identified and validated five palmitoylation-related genes associated with AS. SLC7A7 showed a potential protective causal signal in MR analysis. The enriched pathway patterns linked these genes to immune inflammation, calcium signalling-contraction coupling, ECM remodelling, cell adhesion, and autophagy- associated metabolic reprogramming. The five-gene nomogram showed potential utility for diagnostic classification and decision support, nominating candidate biomarkers and pathway targets for AS molecular subtyping, diagnosis, and mechanistic investigation.

Atherosclerosis (AS)↗

Cell and tumor classification using gene expression data: construction of forests.

The advent of gene chips has led to a promising technology for cell, tumor, and cancer classification. We exploit and expand the methodology of recursive partitioning trees for tumor and cell classification from microarray gene expression data. To improve classification and prediction accuracy, we introduce a deterministic procedure to form forests of classification trees and compare their performance with extant alternatives. When two published and commonly used data sets are used, we find that the deterministic forests perform similarly to the random forests in terms of the error rate obtained from the leave-one-out procedure, and all of the forests are far better than the single trees. In addition, we provide graphical presentations to facilitate interpretation of complex forests and compare our findings with the current biological literature. In addition to numerical improvement, the main advantage of deterministic forests is reproducibility and scientific interpretability of all steps in tree construction.

Cells↗

Integrative machine learning and transcriptomic analysis reveals molecular mechanisms underlying low survival rate in larval Chinese Bahaba (Bahaba taipingensis).

Chinese Bahaba (Bahaba taipingensis) is a Class I protected marine fish endemic to China. Low larvae survival during artificial breeding severely hinder population recovery. To investigate the molecular mechanism of high mortality in larval fish, this study performed RNA-seq on liver from naturally deceased (ND) and mass-dead (MD) individuals, combined with least absolute shrinkage and selection operator (LASSO) regression and random forest (RF) algorithms to screen for core signature genes. A total of 873 differentially expressed genes (DEGs) were identified, including 112 upregulated and 761 downregulated genes. GO and KEGG enrichment analyses revealed significant enrichment in amino acid metabolism disorders, one&#x2011;carbon folate pool impairment, PPAR signaling abnormalities, ECM-receptor interaction, focal adhesion pathway, indicating widespread metabolic suppression accompanied by extracellular matrix remodeling and signaling disturbances in the livers of MD fish. MAD pre-filtering combined with dual machine learning algorithms yielded 18 robust core signature genes, among which SLC38A4, MMP1, FADD, FKBP5, and APOB were consistently identified as high-frequency core genes by both algorithms. SLC38A4 exhibited the highest importance score in the RF model and was significantly downregulated, making it the primary molecule distinguishing ND from MD phenotypes. ROC curve analysis showed that both models achieved an AUC of 1.000 (95% CI lower bound: 0.610), confirming the precise discriminatory ability of the core genes. GSEA further demonstrated significant enrichment of this core gene set in ND samples. This study provides the first systematic elucidation of the molecular mechanisms underlying liver dysfunction in low survival rate B. taipingensis, characterized by amino acid transport impairment, metabolic reprogramming, and structural remodeling, offering theoretical foundations for health assessment, early mortality risk warning, and artificial breeding conservation of this species.

Animals↗

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans↗