Search PubMedSearch

SEARCH · Search PubMed

Results for “large-scale data analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

TET2 promotes monocyte inflammatory activation in asthma via ALKBH5-m6A regulation and PI3K signaling: evidence from m6A-SNP and single-cell analyses.

Asthma is a complex inflammatory airway disease with strong genetic determinants, yet the functional relevance of most asthma-associated non-coding variants remains unclear. Emerging evidence suggests that N6-methyladenosine (m6A) modification may serve as a critical epitranscriptomic link between genetic variation and immune regulation. In this study, we aimed to systematically identify functionally relevant m6A-regulated genes in asthma by integrating large-scale GWAS data, m6A-SNP annotations, and single-cell transcriptomic analyses, and to investigate their roles in monocyte-driven airway inflammation. We identified TET2 as a key m6A-regulated gene associated with both asthma and lung function, which was selectively upregulated in monocytes during asthma and accompanied by activation of inflammatory and PI3K signaling pathways. Mechanistic experiments further demonstrated that inflammatory stimulation induced ALKBH5 expression, reduced m6A modification of TET2 mRNA, and increased TET2 protein levels, thereby promoting PI3K/AKT signaling and pro-inflammatory cytokine production, whereas inhibition of TET2 or ALKBH5 attenuated these effects. Collectively, these findings demonstrate that ALKBH5-mediated m6A regulation of TET2 enhances PI3K/AKT signaling in monocytes, thereby promoting inflammatory responses in asthma. Our study establishes TET2 as a key m6A-regulated gene linking genetic susceptibility to monocyte-driven inflammation, and highlights the ALKBH5-m6A-TET2 axis as a potential therapeutic target for modulating aberrant immune responses in asthma.

Humans

Disentangling the association between chronic pain and sarcopenia-related traits: A bidirectional Mendelian randomization study.

ObjectiveThis study aimed to investigate the potential causal relationships between chronic pain and three key sarcopenia-related quantitative traits: (a) hand grip strength; (b) usual walking pace; and (c) appendicular lean mass, using bidirectional two-sample Mendelian randomization.MethodsWe conducted bidirectional two-sample Mendelian randomization using summary-level data from large-scale genome-wide association studies to assess the genetically predicted associations between chronic pain, including multisite chronic pain and chronic widespread musculoskeletal pain, and the aforementioned sarcopenia-related traits.ResultsMendelian randomization revealed that multisite chronic pain was significantly associated with an increased risk of low hand grip strength (odds ratio = 1.70; p&#x2009;<&#x2009;0.001) and decreased usual walking pace (odds ratio = 0.81; p&#x2009;<&#x2009;0.001); chronic widespread musculoskeletal pain was also significantly associated with decreased usual walking pace (odds ratio = 0.15; p&#x2009;<&#x2009;0.001). Additionally, higher left hand grip strength was significantly associated with a lower risk of multisite chronic pain (odds ratio = 0.90; p<&#x2009;0.001) and chronic widespread musculoskeletal pain (odds ratio = 0.99; p&#x2009;=&#x2009;0.002); higher right hand grip strength was significantly associated with a lower risk of multisite chronic pain (odds ratio = 0.91; p&#x2009;=&#x2009;0.002); and higher usual walking pace was significantly associated with a lower risk of multisite chronic pain (odds ratio = 0.49; p&#x2009;<&#x2009;0.001) and chronic widespread musculoskeletal pain (odds ratio = 0.92; p&#x2009;<&#x2009;0.001). No significant causal associations were detected for appendicular lean mass in either direction (all p&#x2009;>&#x2009;0.05).ConclusionThis study provides genetic evidence supporting potential causal links between chronic pain and key phenotypic components of sarcopenia.

Humans

Effectiveness of Wearable Digital Therapeutics in Improving Sleep Outcomes Among Individuals With Insomnia: Systematic Review and Meta-Analysis of Randomized Controlled Trials.

BACKGROUND: Wearable devices are increasingly used for sleep monitoring and as adjunctive treatment. Existing meta-analyses mostly pool composite digital therapies and rarely isolate stand-alone wearables or distinguish between objective and subjective end points. Whether stand-alone wearable interventions improve sleep outcomes in adults with insomnia, and which factors moderate treatment heterogeneity, remains unclear. OBJECTIVE: This study aims to evaluate the effectiveness of wearable digital interventions on sleep outcomes in adults with insomnia versus control strategies and explore moderators of effectiveness, including device-wearing position, intervention duration, and control type, using meta-regression. METHODS: This systematic review and meta-analysis was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta&#x2011;Analyses) 2020 statement and the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta&#x2011;Analyses Literature Search Extension) guideline. Five electronic databases and clinical trial registries were searched from inception to May 18, 2026. Eligible studies were randomized controlled trials (RCTs) evaluating wearable digital interventions in adults with insomnia compared with sham, waitlist, usual care, or active control conditions and had an intervention duration of at least 1 week. Study screening, data extraction, and risk-of-bias assessment were carried out independently by 2 reviewers. Pooled estimates were calculated using a restricted maximum likelihood random-effects model with the Hartung-Knapp-Sidik-Jonkman correction. Heterogeneity was assessed using the I&#xb2; statistic, and 95% prediction intervals (PIs) were calculated for the primary analyses. The certainty of evidence was rated using the GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) approach. RESULTS: Sixteen RCTs (N=910) were included. Wearable digital interventions were associated with a significant reduction in objective sleep-onset latency (SOL; mean difference [MD] -4.52, 95% CI -8.38 to -0.67, PI -9.52 to 0.47 min) and a significant improvement in subjective sleep efficiency (SE; MD 2.00%, 95% CI 1.90%-2.11%, PI 1.85%-2.15%). Subjective total sleep time (TST) also showed a significant increase (MD 19.11, 95% CI 2.98-35.24, PI -16.20 to 54.43 minutes). Meta-regression showed that control type, intervention duration, and device location did not explain the heterogeneity of the insomnia severity index (ISI) (R&#xb2;=0). Sensitivity analysis confirmed the robustness of pooled ISI estimates, and an Egger test indicated no small-study effects (P=.07). Certainty of evidence ranged from moderate to high. CONCLUSIONS: Wearable digital interventions provide selective benefits for objective SOL, subjective SE, and subjective TST in adults with insomnia, with no improvement in overall ISI. Despite statistically significant effects on several sleep parameters, wide PIs, substantial heterogeneity, and limited study numbers indicate preliminary, nonconclusive findings. Wearables should be viewed as affordable adjunctive tools requiring further validation, not substitutes for first-line cognitive behavioral therapy for insomnia. Large-scale, long-term RCTs with standardized protocols and patient-level external validation are required to consolidate the evidence base.

Humans

Predicting the First Onset of Suicidal Thoughts and Behaviors in Adolescents Using Multimodal Risk Factors: A 4-Year Longitudinal Study.

OBJECTIVE: Suicide is one of the leading causes of death among youth worldwide, yet existing studies that aimed to predict the first onset of suicidal thoughts and behaviors (STB) included a limited number of data modalities and/or focused on adult populations. This study aimed to prospectively predict first-onset STB across 4-year follow-ups in adolescents using an existing STB history classification model that was previously applied to baseline data and a new machine learning model with 195 biopsychosocial features. METHOD: Participants were 7,503 unrelated adolescents (54.5% female, ages 9-11 years at baseline) from the multisite, longitudinal Adolescent Brain Cognitive Development (ABCD) Study. An existing baseline STB history classification model was applied to predict longitudinal first-onset STB in adolescents compared with healthy controls and clinical controls (individuals with a mental health disorder but no STB). A new elastic net logistic regression model with 195 features was trained on data from 14 sites (n = 5,220), and the resulting top 15 features were validated at 7 independent sites (n = 2,283). RESULTS: The previously developed model to classify STB lifetime history also prospectively predicted first-onset STB in adolescents with an area under the curve (AUC) [95% CI] of 0.73 [0.70, 0.75], p < .001, compared with healthy controls and AUC [95% CI] of 0.63 [0.60, 0.66], p < .001, compared with clinical controls. The newly trained model with top 15 features performed similarly with AUC [95% CI] of 0.73 [0.71, 0.76], p < .001, and AUC [95% CI] of 0.64 [0.60, 0.66], p < .001, for the same comparison groups. The most consistent predictors across models included female sex, sleep disturbances, and maladaptive home and school environments. CONCLUSION: The models predicted first-onset STB in adolescents with moderate accuracy. This study also confirmed the roles of well-established psychological risk factors for STB and identified several novel neurocognitive and brain imaging risk factors. Future studies should validate these models in large-scale diverse samples before clinical translation. PLAIN LANGUAGE SUMMARY: This study followed over 7,500 adolescents for 4 years and tested 2 machine learning models using psychological, social, and brain data to identify those at risk of experiencing suicidal thoughts or behaviors. Both models predicted first-time suicidal thoughts or behaviors with moderate accuracy. Key risk factors that were identified included being female, experiencing sleep problems, and negative home and school environments. DIVERSITY & INCLUSION STATEMENT: We worked to ensure sex and gender balance in the recruitment of human participants. We worked to ensure race, ethnic, and/or other types of diversity in the recruitment of human participants. We worked to ensure that the study questionnaires were prepared in an inclusive way. Diverse cell lines and/or genomic datasets were not available. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented racial and/or ethnic groups in science. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented sexual and/or gender groups in science. We actively worked to promote sex and gender balance in our author group. One or more of the authors of this paper received support from a program designed to increase minority representation in science. We actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our author group. While citing references scientifically relevant for this work, we also actively worked to promote sex and gender balance in our reference list. While citing references scientifically relevant for this work, we also actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our reference list. The author list of this paper includes contributors from the location and/or community where the research was conducted who participated in the data collection, design, analysis, and/or interpretation of the work.

Adolescent

Effects of time-restricted eating on markers of glucose metabolism and regulation in individuals with prediabetes or type 2 diabetes: a systematic review and meta-analysis of randomised controlled trials.

AIMS/HYPOTHESIS: This systematic review and meta-analysis aimed to investigate the effects of time-restricted eating (TRE) on glucose metabolism and regulation in individuals with prediabetes (fasting blood glucose of 5.6-6.9 mmol/l or HbA1c of 39-47 mmol/mol [5.7-6.4%]) or type 2 diabetes (fasting blood glucose &#x2265;7 mmol/l or HbA1c &#x2265;48 mmol/mol [6.5%]). METHODS: A literature search was performed in MEDLINE, Embase and CENTRAL from inception to 5 August 2025. Moreover, forward and backward citation searches were performed. Eligible studies were RCTs in adults with prediabetes or type 2 diabetes, lasting &#x2265;2 weeks, reporting markers of glucose metabolism and regulation, comparing TRE (&#x2264;12 h eating window) with a non-time-restricted control diet. Studies involving pregnancy, other fasting regimens, or non-peer-reviewed publications were excluded. Data were pooled as weighted mean differences with 95% CIs using random-effects generic inverse variance models in Cochrane Review Manager Web, and results are presented as forest plots. The certainty of evidence was defined using Grading of Recommendations, Assessment, Development and Evaluations methodology, and risk of bias was estimated by using the Revised Cochrane risk-of-bias tool for randomised trials (RoB 2). RESULTS: Out of 2043 records identified through the database search, as well as 1249 from forward and backward citation searches, ten RCTs including 599 participants were included. The mean length of the studies was 4 months, and the eating windows ranged from 4 to 10 h per day. The pooled meta-analysis showed no overall effect of TRE on HbA1c (-3.33 mmol/mol; 95% CI -6.87, 0.20 (-0.30% points; -0.63, 0.02); p=0.06, moderate certainty). Nevertheless, following stratification by subgroups, TRE resulted in a reduction in HbA1c of 0.93 mmol/mol (-1.70, -0.17 [-0.09% points; -0.16, -0.02]; p=0.02) in individuals with prediabetes but not in individuals with type 2 diabetes (-4.68 mmol/mol; -10.08, 0.72 (-0.43% points; -0.92, 0.07); p=0.09). TRE reduced fasting blood glucose in the pooled analysis (-0.30 mmol/l; -0.53, -0.07; p<0.01, moderate certainty) as well as in the subgroup analyses in individuals with prediabetes (-0.14 mmol/l; -0.27, -0.01; p=0.03) and with type 2 diabetes (-0.48 mmol/l; -0.78, -0.17; p<0.01). Moreover, TRE lowered body weight by 1.6 kg (-2.2, -1.0; p<0.001) in the pooled analysis. The evidence was limited by imprecision arising from wide confidence intervals in some of the included studies, which may be due to small sample sizes. Lastly, the effects of TRE on markers of insulin sensitivity, beta cell function and continuous glucose monitoring measurements were inconclusive. CONCLUSIONS/INTERPRETATION: Moderate-certainty evidence indicates that TRE reduces fasting blood glucose but not HbA1c. The subgroup analyses revealed that TRE improved HbA1c and fasting glucose in individuals with prediabetes and improved fasting glucose in individuals with type 2 diabetes. Future large-scale studies should investigate long-term effects of TRE in prevention and treatment of type 2 diabetes. TRIAL REGISTRATION: PROSPERO CRD42024523591 FUNDING: This research received no specific grant from any funding agency in the public, commercial or not-for-profit sectors. Three authors (JS, A-DT, THA) are employed at Steno Diabetes Center Copenhagen, a public hospital and research institution under the Capital Region of Denmark, partly funded by a grant from the Novo Nordisk Foundation.

Humans

Genetic evidence for repurposing GLP-1 receptor agonists in chronic kidney disease and IgA nephropathy: Metabolic and anti-inflammatory pathways beyond glycaemic control.

AIMS: Despite observational links between glucagon-like peptide-1 receptor agonists (GLP-1RAs) and kidney benefits, causal mechanisms remain unclear. This study aims to dissect genetic causality and mediation pathways underlying the effects of GLP-1RAs on chronic kidney disease (CKD) and related renal outcomes. MATERIALS AND METHODS: Using large-scale Genome - Wide Association Study (GWAS) data, we applied two-sample Mendelian randomisation (MR) to estimate the causal effects of GLP-1RAs on CKD, estimated glomerular filtration rate (eGFR) and subtypes (IgA nephropathy, membranous nephropathy, nephrotic syndrome and chronic glomerulonephritis), with sensitivity analyses. The glycaemic markers (glycated haemoglobin [HbA1c] and blood glucose), type 2 diabetes mellitus (T2DM) and diabetic nephropathy (DN) served as positive controls. Mediation MR assessed body mass index (BMI), lipids, glycaemic markers and inflammatory proteins. Data were sourced from MRC Integrative Epidemiology Unit Open Genome - Wide Association Studies OpenGWAS, FinnGen, GWAS Catalogue and cohort-specific studies. RESULTS: Positive control analyses revealed that genetically predicted GLP-1R activation was associated with reduced levels of HbA1c (p&#x2009;=&#x2009;4.93E-15) and blood glucose (p&#x2009;=&#x2009;9.73E-5), as well as a decreased risk of T2DM (p&#x2009;=&#x2009;2.45E-4) and DN (p&#x2009;=&#x2009;6.35E-4), fully validating the reliability of the genetic instruments. Genetic proxies for GLP-1R activation lowered risks of CKD (odds ratio [OR]&#x2009;=&#x2009;0.83, p&#x2009;=&#x2009;9.22E-9), immunoglobulin A nephropathy (IgAN) (OR&#x2009;=&#x2009;0.70, p&#x2009;=&#x2009;2.11E-3) and kidney function preservation (&#x3b2;&#x2009;=&#x2009;0.01, p&#x2009;=&#x2009;9.11E-3), but showed null effects on other CKD subtypes. Mediation analyses indicated that fibroblast growth factor 23 (FGF23) suppression mediated 26.57% of the effect on eGFR and 13.50% of CKD protection, whereas metabolic traits (BMI: 2.08% for CKD, 5.51% for eGFR; high-density lipoprotein: 0.79% for CKD, 2.34% for eGFR; HbA1c: 8.25% for eGFR) partially explained the benefits on CKD and eGFR. Only BMI exhibited a mediation effect on IgAN. Sensitivity analyses confirmed minimal pleiotropy. CONCLUSIONS: This study provides robust genetic evidence for repurposing GLP-1RAs in CKD and IgAN through anti-inflammatory (FGF23) and metabolic pathways, extending their utility beyond glucose control. While European ancestry data limit generalisability, our framework prioritises FGF23 and metabolic modulation as key targets for clinical trials in renal protection.

Humans

Genome-wide gene-sleep interaction study identifies novel lipid loci in 732,564 participants.

BACKGROUND AND AIMS: Deviations from the population mean in sleep duration have been associated with increased risk for developing dyslipidemia and atherosclerotic cardiovascular disease, but the mechanism of effect is poorly characterized. We performed large-scale genome-wide gene-sleep interaction analyses of lipid levels to identify genetic variants underpinning the biomolecular pathways of sleep-associated lipid disturbances and to suggest possible druggable targets. METHODS: We collected data from 55 cohorts with a combined sample size of 732,564 participants (87&#xa0;% European ancestry) with data on lipid traits (high-density lipoprotein [HDL-c] and low-density lipoprotein [LDL-c] cholesterol and triglycerides [TG]). Short (STST) and long (LTST) total sleep time were defined by the extreme 20&#xa0;% of the age- and sex-standardized values within each cohort. Based on cohort-level summary statistics data, we performed meta-analyses for one-degree of freedom tests of interaction and two-degree of freedom joint tests of the SNP-main and -interaction effect on lipid levels. RESULTS: The one-degree of freedom variant-sleep interaction test identified 10 novel loci (Pint<5.0e-9), and we additionally identify 7 loci within the two-degree of freedom analyses (Pjoint<5.0e-9 in combination with Pint<6.6e-6). Multiple loci, including those mapped to APSH (target for aspartic and succinic acid) and SLC8A1 showed biological plausibility and druggability potential based on literature. CONCLUSIONS: Collectively, the 17 (9 with short and 8 with long sleep) loci provided evidence into the biomolecular mechanisms underlying sleep-associated lipid changes, including potential involvement of the vitamin D receptor pathway. Collectively, these findings may contribute developing novel interventions for treating dyslipidemia in people with sleep disturbances.

Humans

Application of SPI-guided analgesia in laparoscopic gynecologic surgery: a randomized controlled trial evaluating the remifentanil-sparing effect and predictive value of time-weighted SPI.

This study aimed to achieve two primary objectives: (1) to evaluate the opioid-sparing effect of Surgical Pleth Index (SPI)-directed analgesia during surgery via a randomized controlled trial (RCT), and (2) to propose and preliminarily assess a novel dynamic metric, Threshold-based Time-Weighted SPI (Tb-TW-SPI), which integrates stimulus intensity and duration, for its predictive efficacy regarding postoperative moderate-to-severe pain. Employing an RCT combined with exploratory analysis, 61 patients undergoing elective laparoscopic gynecologic surgery were randomized into an SPI-directed analgesia group or a conventional analgesia group. The primary outcome was total intraoperative remifentanil consumption. Postoperatively, an exploratory analysis of the control group data evaluated the correlation between Tb-TW-SPI and Numeric Rating Scale (NRS) pain scores in the post-anesthesia care unit (PACU), calculating its predictive value for moderate-to-severe pain (NRS&#x2009;&#x2265;&#x2009;4). Results: The SPI-directed group required significantly less intraoperative remifentanil than the conventional group [median (IQR): 5.84(5.02,6.62)vs. 6.96(5.81,8.19)&#xb5;g/kg/h; P&#x2009;=&#x2009;0.016]. Postoperative pain scores did not differ significantly between groups (P&#x2009;>&#x2009;0.05). Exploratory analysis of the conventional analgesia group revealed that Tb-TW-SPI values were significantly higher in patients with moderate-to-severe postoperative pain (NRS&#x2009;&#x2265;&#x2009;4) compared to those without (P&#x2009;=&#x2009;0.0417).The area under the ROC curve for Tb-TW-SPI predicting this pain was 0.74 (95% CI: 0.52-0.96), with 67% sensitivity and 76% specificity at an optimal cutoff of 1210. This RCT suggests that SPI-directed analgesia can safely and moderately reduce intraoperative remifentanil consumption. Furthermore, the proposed Tb-TW-SPI metric, in this exploratory analysis, suggests potential for predicting postoperative pain, though this finding requires validation in larger cohorts with higher-frequency SPI sampling, offering a new direction for SPI interpretation. Large-scale, multicenter trials are warranted to validate the predictive utility of Tb-TW-SPI. Clinical Trial Registration, China Clinical Trial Registry: ChiCTR2400088444.

Humans

Steroid hormone biosynthesis and dietary related metabolites associated with excessive daytime sleepiness.

BACKGROUND: Excessive daytime sleepiness (EDS) is a complex sleep problem that affects approximately 33% of the United States population. Although EDS usually occurs in conjunction with insufficient sleep and other sleep and circadian disorders, recent studies have shown unique genetic markers and metabolic pathways underlying EDS. Here, we aimed to further elucidate the biological profile of EDS using large-scale single- and pathway-level metabolomics analyses. METHODS: Metabolomics data were available for 877 metabolites in 6071 individuals from the Hispanic Community Health Study/Study of Latinos (HCHS/SOL). EDS was assessed using the Epworth Sleepiness Scale (ESS) questionnaire. We performed linear regression for each metabolite on the continuous ESS score, adjusting for demographic, lifestyle, and physiological confounders, and in sex specific groups. Subsequently, gaussian graphical modelling was performed coupled with pathway and enrichment analyses to generate a holistic interactive network of the metabolomic profile of EDS associations. FINDINGS: We identified seven metabolites belonging to steroids, sphingomyelin, and long-chain fatty acids sub-pathways in the primary model associated with EDS, and an additional three metabolites in the male-specific analysis. INTERPRETATION: Our findings indicate that an EDS metabolomic profile is characterised by endogenous and dietary metabolites within the steroid hormone biosynthesis pathway, with some pathways that differ by sex. These pathways may be useful for understanding the causes or consequences of EDS and related sleep disorders. FUNDING: Details regarding funding supporting this work and all studies involved are provided in the acknowledgements section.

Humans

Scalable, generalizable and uncertainty-aware integration of spatial multiomics across diverse modalities and platforms with SCIGMA.

Recent advances in spatial omics technologies have enabled simultaneous profiling of transcriptomic, proteomic, epigenomic, metabolomic and imaging data at high spatial resolution, offering unprecedented opportunities to dissect tissue complexity. However, integrating these diverse and large-scale spatial multimodal datasets remains a major computational challenge. We present SCIGMA, a scalable and generalizable deep learning framework for spatial multiomics integration. SCIGMA introduces an uncertainty-aware contrastive learning objective and multiview graph neural networks to preserve modality-specific signals while learning biologically meaningful joint representations. Unlike previous methods, SCIGMA provides spatially resolved uncertainty estimates, interpretably identifying regions of biological or technical heterogeneity. SCIGMA supports integration of up to five modalities, and its modular framework is extensible to future technologies with even more modalities. It also scales to more than 1 million spatial locations, enabling analysis of high-resolution datasets such as Visium HD and Xenium Prime. We evaluated SCIGMA across 19 datasets spanning 8 modalities, 10 tissues and 9 platforms. On benchmarkable datasets, SCIGMA outperformed other methods in spatial domain detection, modality preservation, feature reconstruction and reproducibility. SCIGMA identifies biologically meaningful structures, refined spatial domains and modality-specific regulatory programs, providing a robust, flexible and future-ready solution for scalable spatial multimodal integration.

Multiomics

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348&#xa0;handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

Sparse polygenic risk score inference with the spike-and-slab LASSO.

MOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs.

Multifactorial Inheritance

PLAID: ultrafast single-sample gene set enrichment scoring.

SUMMARY: In recent years, computational methods have emerged that calculate enrichment of gene signatures within individual samples. These signatures offer critical insights into the coordinated activity of functionally related genes, proteins or metabolites, enabling the identification of unique molecular profiles in individual cells and patients. This strategy is pivotal for patient stratification and advancement of personalized medicine. However, the rise of large-scale datasets, including single-cell profiles and population biobanks, has exposed significant computational inefficiencies in existing methods. Current methods often demand excessive runtime and memory resources, becoming impractical for large datasets. Overcoming these limitations is a focus of current efforts by bioinformatics teams in academia and the pharmaceutical industry, as essential to support basic and clinical biomedical research. To address this critical need, we developed PLAID (Pathway Level Average Intensity Detection), an ultrafast and memory optimized single sample gene set enrichment algorithm that utilizes sparse matrix computation. PLAID delivers highly accurate gene set scoring and surpasses the performance of current methods in single-cell and bulk transcriptomics, and proteomics data. PLAID uniquely integrates the most widely used gene set scoring algorithms, enabling researchers to apply multiple methods for cross-validation with outstanding runtime efficiency and minimal memory requirement. AVAILABILITY AND IMPLEMENTATION: PLAID is implemented in the R language for statistical computing. PLAID source code and installation instructions are available with no restrictions at https://github.com/bigomics/plaid.

Algorithms

Secure bioinformatics: privacy-preserving federated analytics using homomorphic encryption.

MOTIVATION: Large-scale bioinformatics analyses increasingly require collaboration across multiple cohorts and institutions, yet existing workflows often rely on data co-localization, which is slow, difficult to scale, and raises privacy concerns. We present a privacy-preserving federated analytics framework that enables secure statistical analysis across distributed datasets without transferring raw data, by performing all computations on encrypted data via cryptographic methods. RESULTS: We evaluate the framework by validating polygenic risk scores and conducting meta-analyses on two real-world cohorts. The proposed solution achieves over 99.9% accuracy relative to plaintext analyses, while maintaining scalable runtime performance with increasing data size and number of participating sites. These results demonstrate the feasibility of secure federated analytics for practical bioinformatics applications involving sensitive data.

Computational Biology

ADAMIXTURE: adaptive first-order optimization for biobank-scale genetic clustering.

MOTIVATION: Estimating genetic clusters from sequencing data is a fundamental task in population and medical genetics, enabling demographic inference and adjustment for population structure in association studies. ADMIXTURE, a widely used model-based clustering method, employs an accelerated Expectation-Maximization (EM) algorithm to infer population parameters; however, its computational demands scale poorly, limiting its usefulness for modern biobank-sized datasets. While recent EM acceleration strategies employing second-order quasi-Newton schemes preserve accuracy, they remain computationally intensive. Conversely, EM-free approaches that prioritize speed often compromise solution quality. RESULTS: We introduce ADAMIXTURE, a novel optimization framework that integrates the EM algorithm with Adaptive Moment Estimation (Adam). Unlike traditional acceleration methods, ADAMIXTURE utilizes first-order gradients with adaptive learning rates derived from raw and squared moments to approximate curvature information, bypassing the computational overhead of Hessian approximations. This approach surpasses the convergence efficiency of second-order methods while maintaining the low computational complexity of first-order updates. Across simulated and large-scale empirical datasets, ADAMIXTURE demonstrates substantial reductions in wall-clock runtime and enhanced scalability compared to state-of-the-art methods, while maintaining comparable or improved inference accuracy. Its GPU implementation runs in under 2&#xa0;h on half a million samples and variants, a two order of magnitude speedup over current state-of-the-art. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://github.com/AI-sandbox/ADAMIXTURE.

Clustering Algorithms

ECHO: a nanopore sequencing-based workflow for (epi)genetic profiling of the human repeatome.

SUMMARY: The human genome is dominated by repetitive DNA, whose genetic and epigenetic variation plays a key role in gene regulation, genome stability, and disease. Recent advances in long-read sequencing now enable large-scale, haplotype-resolved, and DNA methylation-informative analysis of the human genome, including on previously inaccessible complex and repetitive regions. However, the comprehensive, simultaneous characterisation of the "human repeatome" remains challenging, largely due to the lack of comprehensive tools integrated in a single pipeline that can capture the full spectrum of variation across diverse types of DNA repeats. Here, we present ECHO, a user-friendly, Snakemake-based pipeline for the "(Epi)genomic Characterisation of Human Repetitive Elements using Oxford Nanopore Sequencing." ECHO provides a reproducible and scalable framework for end-to-end analysis of whole-genome nanopore sequencing data, enabling integrative but also tailored (epi)genetic analyses of the human repeatome. AVAILABILITY AND IMPLEMENTATION: ECHO is freely available at Github: https://github.com/leenput/ECHO-pipeline, with the archived version at Zenodo: https://zenodo.org/records/19068468.

Humans

CoSAG-nf: A Scalable Nextflow Pipeline for Co-assembly, Optimization, and Interactive Visualization of High-Throughput Single-Cell Genomes.

MOTIVATION: Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. RESULTS: We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. AVAILABILITY: CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Journal Article

Association of Vitamin D Polygenic Risk Scores and Disease Outcome in People With Multiple Sclerosis.

BACKGROUND AND OBJECTIVES: Observational studies suggest low levels of 25-hydroxyvitamin D (25[OH]D) may be associated with increased disease activity in people with multiple sclerosis (PwMS). Large-scale genome-wide association studies (GWAS) suggest 25(OH)D levels are partly genetically determined. The resultant polygenic scores (PGSs) could serve as a proxy for 25(OH)D levels, minimizing potential confounding and reverse causation in analyses with outcomes. Herein, we assess the association of genetically determined 25(OH)D and disease outcomes in MS. METHODS: We generated 25(OH)D PGS for 1,924 PwMS with available genotyping data pooled from 3 studies: the CombiRx trial (n = 575), Johns Hopkins MS Center (n = 1,152), and Immune-Mediated Inflammatory Diseases study (n = 197). 25(OH)D-PGS were derived using summary statistics (p < 5 &#xd7; 10-8) from a large GWAS including 485,762 individuals with circulating 25(OH)D levels measured. We included clinical and imaging outcomes: Expanded disability status scale (EDSS), timed 25-foot walk (T25FW), nine-hole peg test (9HPT), radiologic activity, and optical coherence tomography-derived ganglion cell inner plexiform layer (GCIPL) thickness. A subset (n = 935) had measured circulating 25(OH)D levels. We fitted multivariable models based on the outcome of interest and pooled results across studies using random effects meta-analysis. Sensitivity analyses included a modified p value threshold for inclusion in the PGS (5 &#xd7; 10-5) and applying Mendelian randomization (MR) rather than using PGS. RESULTS: Initial analyses demonstrated a positive association between generated 25(OH)D-PGS and circulating 25(OH)D levels (per 1SD increase in 25[OH]D PGS: 3.08%, 95% CI: 1.77%, 4.42%; p = 4.33e-06; R2 = 2.24%). In analyses with outcomes, we did not observe an association between 25(OH)D-PGS and relapse rate (per 1SD increase in 25[OH]D-PGS: 0.98; 95% CI: 0.87-1.10), EDSS worsening (per 1SD: 1.05; 95% CI: 0.87-1.28), change in T25FW (per 1SD: 0.07%; 95% CI: -0.34 to 0.49), or change in 9HPT (per 1SD: 0.09%; 95% CI: -0.15 to 0.33). 25(OH)D-PGS was not associated with new lesion accrual, lesion volume or other imaging-based outcomes (whole brain, gray, white matter volume loss or GCIPL thinning). The results were similarly null in analyses using other p value thresholds or those applying MR. DISCUSSION: Genetically determined lower 25(OH)D levels were not associated with worse disease outcomes in PwMS and raises questions about the plausibility of a treatment effect of vitamin D in established MS.

Humans