Search PubMedSearch

SEARCH · Search PubMed

Results for “Feature selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Development and validation of a serum peptidomic signature for early detection of asymptomatic ovarian cancer: A multi-center prospective study.

Early detection of asymptomatic ovarian cancer (asym-OC) remains a critical challenge, the failure of which underlies its high mortality. Performing serum peptidomic profiling of 843 participants in the cohort SOCFCP, we distill 1,081 initial features into a 7-marker panel for asym-OC detection via a biology-informed machine-learning (ML)-based feature selection strategy. Three markers significantly revert toward non-OC levels after surgery. Integrating the panel with age, CA125, and HE4, we develop and externally validate (n = 159) a LightGBM model, ProMS+. For early-stage OC detection, ProMS+ shows a specificity of 92.6% at 95.0% sensitivity, outperforming CA125 (44.7%), HE4 (11.2%), and Risk of Ovarian Malignancy Algorithm (ROMA) (24.0%), with an area under the curve (AUC) of 0.993. In a simulated high-risk population (n = 100,000; OC prevalence = 1%), ProMS+ yields a high AUC (0.983) and a higher positive predictive value than CA125, HE4, and Age + CA125 + HE4 combined model (0.201 vs. 0.027, 0.090, and 0.064). ProMS+ offers a promising, non-invasive, and interpretable approach for the early detection of asym-OC.

Humans

Diffusion coefficients of hemoglobin by intensity fluctuation spectroscopy: effects of varying pH and ionic strength.

Measurements of the mutual diffusion coefficients (D) of the liganded human hemoglobins (Hb) oxy-HbA and oxy-HbS were performed as a function of Hb concentration (CHb), pH, and ionic strength (tau) by intensity fluctuation spectroscopy (IFS). Average diffusion coefficients, (D), and normalized variances, ((D/(D) - 1)2), were recorded. Results are reported and select features are discussed quantitatively. (a) for tau = 0.15 M, the shape of the (d) vs. CHb curve is found to vary with pH. We developed a precise description of this effect in the form of an algebraic relationship between (D), CHb, and Z, the titration charge. (b) only slight differences between the (D) values of oxy-HbS and oxy-HbA are observed, at tau = 0.15 M, for CHb Less Than or Equal To 10 g%. These differences are explained by the theory of part a. (c) No evidence of aggregation is found in solutions of oxy-HbA or oxy-HbS, at tau = 0.15 M, for CHb Less Than or Equal To 10 g%. (d) Indications of aggregation appear in oxy-HbA solutions at very low concentrations of salt. An estimate is made of the extent of aggregation, and the average radius of a cluster is determined.

Diffusion

A statewide characterization of hospital infection control practices and practitioners.

Selected features of infection control programs among the 163 general hospitals in Tennessee were surveyed in 1976 and 1979. Each hospital but one had a designated infection control practitioner. Three-fourths of the hospitals had fewer than 200 beds and most were in rural areas. The practitioners in these small hospitals worked in an isolated professional milieu: few (4%) had attended a basic training course or were members of a national (11%) or local (16%) infection control association. They also had significantly less access to standard infection control resource publications than did practitioners in large hospitals. Use of aqueous quaternary ammonium compounds for disinfection was reported by 37% of all hospitals in 1979; 68% of hospitals routinely performed bacteriologic cultures of personnel or the environment. In contrast, only 3% of hospitals did not have a policy specifying the use of sterile closed-system drainage of indwelling bladder catheters. Although these practices varied somewhat by hospital size, the differences were not statistically significant. Modest improvement in each parameter was noted since 1976. Pathology was the most common medical specialty (34%) among chairman of infection control committees; internal medicine and pediatrics accounted for only 13%. The practice of routine microbiologic monitoring was significantly more common among hospitals with chairmen who were pathologists. The implications of these findings for national priorities in hospital infection control are discussed.

Bacteria

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)─a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC ≥ 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median ρ ∼ 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals

Dual-Matrix Platform for Highly Specific Multi-Omics Profiling of Renal Cell Carcinoma.

Multiomics interrogation provides complementary information beyond single-omics approaches for improved disease characterization. To enable such multilayer profiling, we expanded the rapid functionalized mesoporous nanoparticle-coupled laser desorption/ionization mass spectrometry (fMNPLDI-MS) platform by designing two structurally homologous but functionally tailored fMNPs. This design enables efficient acquisition of both serum metabolic and peptide fingerprints from a total of only 2.05 μL of serum, with an LDI MS analysis time of approximately 90 s per sample, while addressing the limitation of single-matrix systems in simultaneously optimizing analytical performance for different biomolecular species. Through statistical analysis and machine learning-based feature selection, an integrated multiomics biomarker panel was established, comprising 5 peptides and 4 metabolites. Notably, this integrated panel outperformed both single-omics panels across all evaluation metrics in the validation set, improving the area under curve from 0.985 to 1.000 and increasing the classification accuracy from 0.947 (metabolites) and 0.930 (peptides) to 0.965, while showing consistent improvements in F1-score, precision, and recall. Collectively, these results demonstrate the robust performance of the dual-matrix design and multiomics integration for renal cell carcinoma classification, with potential relevance for broader applications in complex disease profiling.

Carcinoma, Renal Cell

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine

Isolation and characterization of distinct domains of sarcolemma and T-tubules from rat skeletal muscle.

1. Several cell-surface domains of sarcolemma and T-tubule from skeletal-muscle fibre were isolated and characterized. 2. A protocol of subcellular fractionation was set up that involved the sequential low- and high-speed homogenization of rat skeletal muscle followed by KCl washing, Ca2+ loading and sucrose-density-gradient centrifugation. This protocol led to the separation of cell-surface membranes from membranes enriched in sarcoplasmic reticulum and intracellular GLUT4-containing vesicles. 3. Agglutination of cell-surface membranes using wheat-germ agglutinin allowed the isolation of three distinct cell-surface membrane domains: sarcolemmal fraction 1 (SM1), sarcolemmal fraction 2 (SM2) and a T-tubule fraction enriched in protein tt28 and the alpha 2-component of dihydropyridine receptor. 4. Fractions SM1 and SM2 represented distinct sarcolemmal subcompartments based on different compositions of biochemical markers: SM2 was characterized by high levels of beta 1-integrin and dystrophin, and SM1 was enriched in beta 1-integrin but lacked dystrophin. 5. The caveolae-associated molecule caveolin was very abundant in SM1, SM2 and T-tubules, suggesting the presence of caveolae or caveolin-rich domains in these cell-surface membrane domains. In contrast, clathrin heavy chain was abundant in SM1 and T-tubules, but only trace levels were detected in SM2. 6. Immunoadsorption of T-tubule vesicles with antibodies against protein tt28 and against GLUT4 revealed the presence of GLUT4 in T-tubules under basal conditions and it also allowed the identification of two distinct pools of T-tubules showing different contents of tt28 and dihydropyridine receptors. 7. Our data on distribution of clathrin and dystrophin reveal the existence of subcompartments in sarcolemma from muscle fibre, featuring selective mutually exclusive components. T-tubules contain caveolin and clathrin suggesting that they contain caveolin- and clathrin-rich domains. Furthermore, evidence for the heterogeneous distribution of membrane proteins in T-tubules is also presented.

Animals

Opposing effects of estradiol and progesterone on oxytocin receptors in rabbit uterus.

Estradiol-17beta administration to young (10- to 12-week-old) rabbits to produce the "estrogen-dominated" uterus increased the uterine contractile response to both oxytocin and methacholine in vitro. In "progesterone-dominated" uteri, obtained from rabbits that received progesterone for 4 days after estrogen pretreatment, the contractile response to oxytocin in vitro was selectively abolished; the response to methacholine was unaffected. Parallel changes were observed in the concentration (but not affinity) of specific sites in uterine microsomal membranes that bind [(3)H]oxytocin with selectivity features expected for oxytocin receptors. Thus, estrogen-dominated uteri have an increased number of specific [(3)H]oxytocin binding sites per mg of membrane protein relative to untreated controls, whereas specific oxytocin binding sites are reduced to barely detectable levels in the progesterone-dominated uterus. Similar results are obtained when binding sites are measured in membranes from the myometrium of estrogen- or progesterone-dominated uteri. Short-term (24-hr) progesterone administration to estrogen-pretreated rabbits decreased, but did not abolish, specific [(3)H]oxytocin binding; the concentration of specific [(3)H]oxytocin binding sites was reduced without influence on the affinity of these sites. A sublethal dose of actinomycin D, administered over a 24-hr period to rabbits pretreated with estradiol for 4 days, likewise reduced specific oxytocin binding; additive effects were not observed when progesterone and actinomycin D were administered together. These results suggest that the regulatory effects of estrogens and progesterone upon the rabbit uterine contractile response to oxytocin are achieved, at least in part, by the opposing actions of these steroids in regulating the number of oxytocin receptors in smooth muscle cells. Estradiol increased the concentration of uterine oxytocin receptors; the maintenance of high receptor levels appears to depend upon the continuous de novo synthesis of oxytocin receptors. In contrast, progesterone, like actinomycin D, appears to act at the nuclear locus to repress synthesis of oxytocin receptors.

Animals

Assessing individual genetic susceptibility to metabolic syndrome: interpretable machine learning method.

BACKGROUND: Genome-wide association studies have provided profound insights into the genetic aetiology of metabolic syndrome (MetS). However, there is a lack of machine-learning (ML)-based predictive models to assess individual genetic susceptibility to MetS. This study utilized single-nucleotide polymorphisms (SNPs) as variables and employed ML-based genetic risk score (GRS) models to predict the occurrence of MetS, bringing it closer to clinical application. METHODS: Feature selection was performed using Least Absolute Shrinkage and Selection Operator. Six ML algorithms were employed to construct GRS models. A fivefold cross-validation was utilized to aid in the internal validation of models. The receiver operating characteristic (ROC) curve was used to select the better-performing GRS model. The SHapley Additive exPlanations (SHAP) was then applied to interpret the model. After extracting GRS, stratified analysis of BMI, age and gender was performed. Finally, these conventional risk factors and GRS were integrated through multivariate logistic regression to establish a combined model. RESULTS: A total of 17 SNPs were selected for analysis. Among the GRS models, the extreme gradient boosting (XGBoost) model demonstrated superior discriminative performance (AUC = 0.837). The XGBoost's optimal robustness was also validated through five-fold cross-validation (mean ROC-AUC = 0.706). The XGBoost-based SHAP algorithm not only elucidated the global effects of 17 SNPs across all samples, but also described the interaction between SNPs, providing a visual representation of how SNPs impact the prediction of MetS in an individual. There was a strong correlation between GRS and MetS risk, particularly observed among young individuals, males and overweight individuals. Furthermore, the model combining conventional risk factors and GRS exhibited excellent discriminative performance (AUC = 0.962) and outstanding robustness (mean ROC-AUC = 0.959). CONCLUSION: This study established a reliable XGBoost-based GRS model and a GRS prediction platform (https://metabolicsyndromeapps.shinyapps.io/geneticriskscore/) to assess individual genetic susceptibility to MetS. This model has high interpretability and can provide personalized reference for determining the necessity of primary prevention measures for MetS. Additionally, there may be interactions between traditional risk factors and GRS, and the integration of both in a comprehensive model is useful in the prediction of MetS occurrence.

Humans

Stage-specific ROMO1 in rheumatoid arthritis: predictive immune insights into the MIF pathway and HLA-DR/IL2RA axis via integrated GWAS, transcriptomic, single-cell, and spatial profiling.

Emerging evidence links reactive oxygen species modulator 1 (ROMO1), a key mitochondrial ROS regulator, to rheumatoid arthritis (RA) pathogenesis. However, its exact mechanism remains elusive given the conflicting evidence about its specific function. We used a four-level integrative framework combining multi-omics data and literature‑supported mechanistic inference. At the genetic level, Mendelian randomization (MR) was performed to explore potential causal relationships between ROMO1, IL2RA, HLA-DR, MIF, and RA risk, followed by differential expression analysis and machine learning-based feature selection to identify key mROS genes. The temporal expression dynamics of ROMO1 were assessed in RA progression. At the cellular and tissue levels, we integrated single-cell RNA sequencing and spatial transcriptomics to map cell-type-specific expression and synovial localization of ROMO1-related immune cells and pathways. Finally, our multi-omics findings were contextualized with literature-supported mechanistic inference. (1) MR results were consistent with a potential protective effect of ROMO1 on RA (OR = 0.52) and its potential regulation of risk factors IL2RA (OR = 0.46) and HLA-DR (OR = 0.40). Conversely, IL2RA (OR = 1.42), HLA-DR (OR = 1.88), and MIF (OR = 1.17) were positively associated with RA risk. Additionally, ROMO1 was identified as a top candidate diagnostic predictor with stage-specific dynamics: downregulated in the early but upregulated in the late/remission stages. (2) Single-cell RNA sequencing showed ROMO1's cell-specific expression in CD14+ HLA-DR+ CD74+ monocytes and CD4+ IL2RA+ T cells. Cell communication analysis further suggested that these cells may participate in MIF pathway regulation. Spatial transcriptomics subsequently identified that ROMO1-related cells localized to synovial pathological regions, with MIF pathway changes correlated with RA progression. (3) Finally, literature-supported mechanistic inference suggests that ROMO1 may modulate mROS levels to promote anti-inflammatory M2 macrophage polarization, which could theoretically contribute to reduced systemic inflammation and the alleviation of multi-organ decline in RA. This integrated multi-omics investigation, supported by literature-based mechanistic inference, suggests ROMO1 as a stage-dependent biomarker candidate and potential immune regulator in RA.

Humans

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans

CpGene: a web application for epigenetic signature identification from DNA methylation arrays.

MOTIVATION: DNA methylation (DNAme) is the best studied epigenetic mechanism that plays pivotal role in tissue differentiation and epigenetic disruption has been correlated to diverse disease types (e.g. cancer, metabolic disorders). While various DNAme array platforms have been discovered, data analysis remains a challenging task which often requires in-depth bioinformatic expertise. Here, we developed a user-friendly web-based application for data analysis and visualization that accommodates users ranging from early-career basic/translational researchers to experienced bioinformaticians. RESULTS: CpGene is a web application for analyzing DNA methylation array data. It supports Illumina 450K, EPIC, and EPICv2 methylation array platforms and processes .idat files with integrated preprocessing, normalization, and quality control. Biomarker discovery is available through either classic differential methylation point analysis or machine learning-based feature selection as well as gene enrichment analysis. Results are summarized with clear visualizations, to aid interpretation. By combining these functions in a unified interface, CpGene streamlines methylation analysis and helps identify CpG sites and genes with biological and clinical relevance. AVAILABILITY AND IMPLEMENTATION: CpGene is openly accessible as a web service through http://cpgene.duckdns.org:8001/ and it's source code is available on https://github.com/kostaslazaros/cpgenene.

DNA Methylation

Model-based multifacet clustering with high-dimensional omics applications.

High-dimensional omics data often contain intricate and multifaceted information, resulting in the coexistence of multiple plausible sample partitions based on different subsets of selected features. Conventional clustering methods typically yield only one clustering solution, limiting their capacity to fully capture all facets of cluster structures in high-dimensional data. To address this challenge, we propose a model-based multifacet clustering (MFClust) method based on a mixture of Gaussian mixture models, where the former mixture achieves facet assignment for gene features and the latter mixture determines cluster assignment of samples. We demonstrate superior facet and cluster assignment accuracy of MFClust through simulation studies. The proposed method is applied to three transcriptomic applications from postmortem brain and lung disease studies. The result captures multifacet clustering structures associated with critical clinical variables and provides intriguing biological insights for further hypothesis generation and discovery.

Humans

Dietary cadmium, zinc and copper: effects on chick lung morphology and elastin cross-linking.

Day-old White Leghorn cockerels were divided into seven dietary groups and fed one of the following diets: 1) a casein-based basal diet; 2) a casein-based diet supplemented with 10 mg/kg cadmium, 3) 100 mg/kg cadmium, 4) or 800 mg/kg zinc; 5) a casein-based diet pair-fed to the 100 mg/kg Cd group; 6) a spray-dried nonfat milk-based diet with no added copper, or 7) a spray-dried nonfat milk-based diet supplemented with 5 mg/kg copper. At termination (5 weeks), the birds were killed, and the effects of the diets on selected features of lung composition and morphology were assessed. Body weights were reduced in the 100 mg/kg Cd, pair-fed, and Cu-deficient groups when compared to their controls (casein-based or milk-based copper-supplemented diets). There were no differences in lung weights (expressed relative to metabolic body size) among the groups, although copper deficiency did result in a slight decrease in the dry to wet weight ratio of lung. Lung elastin content and the desmosine content in elastin were significantly lower in the Cu-deficient group and tended to be lower in the group fed 800 ppm Zn. Significant alterations (enlargement of the tertiary bronchial lumen) in morphology were also observed in lungs from both the 100 mg/kg Cd and Cu-deficient groups. Alteration in lung morphology observed in the 100 mg/kg Cd group could not be explained by changes in the elastin content of lung.

Animals

Proteomic Immune Signatures of Severe HIV-Associated Tuberculosis in Sub-Saharan Africa: A Prospective, Multicenter Analysis From Uganda.

OBJECTIVES: Severe tuberculosis (TB) is a major cause of critical illness and death in people living with HIV (PLWH) worldwide. Despite this, the immunopathology of severe HIV-associated TB (HIV/TB) is poorly understood. We aimed to identify an immunopathologic signature of severe HIV/TB in sub-Saharan Africa. DESIGN AND SETTING: We analyzed proteomic data from two prospective observational cohorts of adults hospitalized with severe undifferentiated infection in Uganda: an urban discovery cohort (Entebbe, n = 241) and a rural validation cohort (Tororo, n = 253). PATIENTS: Adults (age &#x2265; 18 yr) hospitalized with severe febrile illness. INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: Across both cohorts, severe HIV/TB was common, affecting 18% of participants in the discovery cohort and 21% in the validation cohort. Overall mortality was significant (30-d mortality of 22% in the discovery cohort and 60-d mortality of 26% in the validation cohort). Participants were stratified into three HIV/TB phenotypes: HIV-negative without TB, PLWH without TB, and PLWH with microbiologically diagnosed TB. We applied ordinal random forest models in the discovery cohort as a supervised feature-selection approach to identify proteins associated with progressive HIV/TB phenotype. In both cohorts, PLWH with microbiologically diagnosed TB were at highest risk of critical illness and death (30-d mortality of 42% in the discovery cohort and 60-d mortality of 52% in the validation cohort). An eight-protein signature reliably distinguished this phenotype, reflecting mediators of macrophage/dendritic cell activation (lysosome-associated membrane glycoprotein 3), natural killer cell and T-cell stimulation and cytotoxicity (cluster of differentiation 70, class I-restricted T-cell-associated molecule), B-cell activation (immunoglobulin lambda constant 2), protease-mediated tissue injury (protease, serine 2 [trypsin-2]), dysregulated coagulation (serpin peptidase inhibitor, clade A [alpha-1 antitrypsin], member 5), extracellular matrix remodeling (epidermal growth factor-containing fibulin-like extracellular matrix protein 1), and growth hormone/insulin-like growth factor axis dysregulation (insulin-like growth factor binding protein 3). CONCLUSIONS: We identified an immunologic signature of severe HIV/TB defined by mediators of macrophage/dendritic cell and cytotoxic lymphocyte activation, extracellular matrix remodeling, and dysregulated coagulation. These findings offer new insight into HIV/TB pathobiology and highlight potential targets for host-directed therapies in this high-risk population.

Humans

Estimating population structure using epigenome-wide methylation data.

INTRODUCTION: In epigenome-wide association analysis (EWAS), unaddressed population stratification often leads to inflation. We aimed to compute methylation population scores (MPSs) that predict genetic principal components (GPCs) using a feature selection and regression approach. METHODS: We used multi-ethnic methylation data (Illumina 450K/EPIC array) from unrelated MESA (n=929), CARDIA (n=1123), JHS (n=1365), ARIC (n=2338), and HCHS/SOL (n=1475) individuals, randomly assigning 85% of participants from each cohort to a training dataset and the remaining 15% to a test dataset. First, we estimated the associations of GPCs with each available CpG methylation site using linear regression within each cohort, adjusting for age, sex, smoking status, race/ethnic background (as a proxy for background information associated with lifestyle and other environmental exposures that may impact methylation), alcohol use status, body mass index, and cell type proportions. We meta-analyzed the associations across cohorts and selected CpG sites with association FDR-adjusted q-value <0.05. We next aggregated individuallevel data across the cohort-specific training datasets, and applied two-stage weighted least squares Lasso regression, with the GPCs as the outcomes and the selected CpG sites as penalized predictors, adjusting for the aforementioned covariates. The developed MPSs are the weighted sum of selected CpG sites from the Lasso. To evaluate the developed MPSs, we constructed them in the test dataset, and compared them with GPCs, and with MPSs constructed based on a previously-published paper. Comparison was based on correlation analysis and data visualization. We demonstrate the use of the MPSs in EWAS. RESULTS: In the test dataset, the MPSs were highly correlated with GPCs, with correlation decreasing, though not monotonically, for later components. Specifically, MPS1 and GPC1 had R2= 0.99, while MPS7 and GPC7 had R2=0.27 (the lowest observed correlation). In data visualization, MPSs had similar patterns as GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups, while outperforming MPC constructed using alternative published methods. MPSs showed comparable performance to GPCs in reducing some of the inflation in EWAS. CONCLUSIONS: Methylation-based population scores provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent. Unlike previous methods based on unsupervised methylation PCA, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations. The weights for each GPCs derived in our study can be applied to generate MPSs in other studies.

Journal Article

Knowledge-driven interpretable neural networks provide mechanistic insight.

Analyzing omics data in the context of pathway knowledge is critical for understanding the molecular mechanisms underlying pathological changes. However, current pathway analysis methods do not model the detailed mechanistic nature of biological interactions, limiting the understanding of pathway behavior to a relatively shallow level. To address this issue, we present a knowledge-driven machine learning framework that embeds features into pathway graphs and models reactions analytically, producing interpretable feature hierarchies and subnetworks in which functional associations are estimated to model biological interactions. The approach is agnostic to feature selection, enabling the use of full omics data sets without discarding weak signals. Applications to breast cancer microRNA-gene regulation data and COVID-19 metabolomic data highlight immune and metabolic pathways relevant to disease progression. This framework bridges predictive modeling with mechanistic interpretation and offers a foundation for integrative pathway analysis.

Humans

Large-Scale Plasma Proteomics Identifies Early Molecular Deviations and Improves Risk Prediction for Heart Failure Among Individuals With Obesity.

AIMS: Heart failure (HF) is a major global public health challenge, with obesity being one of its key risk factors. Although several HF risk prediction models have been developed in the general population, few are specifically tailored to individuals with obesity. This underscores the urgent need for precise biomarkers to improve individual risk stratification and enable personalized prevention strategies. We aimed to develop and validate a plasma proteomics-based protein risk score (PRS) to predict incident HF among individuals with obesity. MATERIALS AND METHODS: We analysed 9831 participants with obesity (BMI &#x2265;&#x2009;30&#x2009;kg/m2) from the UK Biobank with baseline measurements of 2911 circulating proteins and up to 16&#x2009;years of follow-up. Multivariable Cox regression identified proteins associated with incident HF after comprehensive covariate adjustment. A PRS was constructed using LASSO regression and evaluated in a held-out test set. Protein trajectories before HF onset were reconstructed using LOESS modelling. To enhance clinical feasibility, a minimal protein panel was identified using LightGBM with forward feature selection. RESULTS: A total of 727 participants developed HF during follow-up. Multivariable cox analyses identified 578 proteins significantly associated with HF. LASSO regression further selected 81 proteins to build the PRS, which showed a strong association with HF risk in both training (HR 3.57; 95% CI 3.19-4.00) and test cohorts (HR 2.45; 95% CI 2.20-2.74). Adding the PRS improved prediction beyond age and sex (&#x394;C&#x2009;=&#x2009;0.091) and beyond the Pooled Cohort Equations to Prevent Heart Failure (PCP-HF) model (&#x394;C&#x2009;=&#x2009;0.052), with consistent gains in NRI and IDI. Proteomic deviations were detectable up to 16&#x2009;years before diagnosis. A four-protein panel (GDF15, NT-proBNP, TNFRSF10B, CTHRC1) achieved robust discrimination (AUC 0.789), outperforming NT-proBNP alone (AUC 0.695) and complementing the PCP-HF model (combined AUC 0.803). DISCUSSION: Large-scale plasma proteomics substantially improves HF risk prediction in individuals with obesity and reveals long-standing molecular alterations preceding clinical onset. A simplified four-protein panel maintains robust predictive accuracy and provides a practical approach for the early detection and targeted prevention of obesity-related HF.

Humans