Search PubMedSearch

SEARCH · Search PubMed

Results for “sparse learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

29 records · Page 2Linked to original sources

PATTY corrects open chromatin bias for improved bulk and single-cell CUT&Tag profiling.

Precise profiling of epigenomes is essential for better understanding chromatin biology and gene regulation. Cleavage Under Targets & Tagmentation (CUT&Tag) is an efficient epigenomic profiling technique that can be performed on a low number of cells and at the single-cell level. With its growing adoption, CUT&Tag datasets spanning diverse biological systems are rapidly accumulating in the field. CUT&Tag assays use the hyperactive transposase Tn5 for DNA tagmentation. Tn5's preference toward accessible chromatin alters CUT&Tag sequence read distributions in the genome and introduces open chromatin bias that can confound downstream analysis, an issue more substantial in sparse single-cell data. We show that open chromatin bias extensively exists in published CUT&Tag datasets, including those generated with recently optimized high-salt protocols. To address this challenge, we present PATTY (Propensity Analyzer for Tn5 Transposase Yielded bias), a comprehensive computational method that corrects open chromatin bias in CUT&Tag data by leveraging accompanying ATAC-seq. By integrating transcriptomic and epigenomic data using machine learning and integrative modeling, we demonstrate that PATTY enables accurate and robust detection of occupancy sites for both active and repressive histone modifications, including H3K27ac, H3K27me3, and H3K9me3, with experimental validation. We further develop a single-cell CUT&Tag analysis framework built on PATTY and show improved cell clustering when using bias-corrected single-cell CUT&Tag data compared to using uncorrected data. Beyond CUT&Tag, PATTY sets a foundation for further development of bias correction methods for improving data analysis for all Tn5-based high-throughput assays.

Journal Article

PATTY corrects open-chromatin bias for improved bulk and single-cell CUT&Tag profiling.

Precise profiling of epigenomes is essential for better understanding chromatin biology and gene regulation. Cleavage Under Targets & Tagmentation (CUT&Tag) is an efficient epigenomic profiling technique that can be performed on a low number of cells and at the single-cell level. With its growing adoption, CUT&Tag datasets spanning diverse biological systems are rapidly accumulating in the field. CUT&Tag assays use the hyperactive transposase Tn5 for DNA tagmentation. Tn5's preference toward accessible chromatin alters CUT&Tag sequence read distributions in the genome and introduces open-chromatin bias that can confound downstream analysis, an issue more substantial in sparse single-cell data. We show that open-chromatin bias extensively exists in published CUT&Tag datasets, including those generated with recently optimized high-salt protocols. To address this challenge, we present PATTY (Propensity Analyzer for Tn5 Transposase Yielded bias), a comprehensive computational method that corrects open-chromatin bias in CUT&Tag data by leveraging accompanying ATAC-seq. By integrating transcriptomic and epigenomic data using machine learning and integrative modeling, we demonstrate that PATTY enables accurate and robust detection of occupancy sites for both active and repressive histone modifications, including H3K27ac, H3K27me3, and H3K9me3, with experimental validation. We further develop a single-cell CUT&Tag analysis framework built on PATTY and show improved cell clustering when using bias-corrected single-cell CUT&Tag data compared to using uncorrected data. Beyond CUT&Tag, PATTY sets a foundation for further development of bias correction methods for improving data analysis for all Tn5-based high-throughput assays.

Journal Article

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Integrative multi-omics profiling deciphers tumor microenvironment heterogeneity and immunotherapy vulnerabilities in lung neuroendocrine carcinomas.

INTRODUCTION: Lung neuroendocrine carcinomas (Lu-NECs) are rare, highly aggressive lung tumors with poor prognosis and limited therapeutic options. Understanding the tumor immune microenvironment (TIME) is crucial towards personalized therapeutic strategies. OBJECTIVES: This study aims to systematically characterize the heterogeneity and complexity of the TIME in Lu-NECs by integrating proteomic, transcriptomic, and genomic data. METHODS: We performed comprehensive immune-proteomic profiling of 76 Lu-NECs across diverse histopathological subtypes to elucidate intra-tumoral TIME heterogeneity at the proteomic level. Validation was conducted in multiple independent cohorts, including 112 Lu-NECs using immunohistochemistry, 147 Lu-NECs, and 17 small cell lung carcinoma samples using transcriptomics. We integrated proteomic, transcriptomic, genomic, and clinical data to assess molecular, immunological, and clinical features, as well as therapeutic vulnerabilities across different immune subtypes. RESULTS: We delineated the immuno-proteomic landscape of Lu-NECs and identified two major immuno-proteomic clusters with distinct immunological, molecular, and clinical characteristics. IPC1 was characterized by high immune cell infiltration, while IPC2 exhibited sparse immune cell presence. Genomic analysis revealed distinct mutational patterns, with IPC1 showing a higher incidence of APOBEC-associated mutation signatures and IPC2 being enriched for mutations associated with defective DNA mismatch repair and tobacco-related mutagens. Functional analyses indicated that IPC1 was related to immune and oncogenic signaling activity, whereas IPC2 was associated with cancer stemness and proliferation-related features. Furthermore, IPC1 and IPC2 demonstrated histological subtype-specific clinical benefits from postoperative chemotherapy. Finally, we developed a machine learning model (iPROM) to predict Lu-NECs immune classification and improve risk stratification, which was validated across multiple independent cohorts. CONCLUSIONS: This study advances the understanding of the tumor immune microenvironment in Lu-NECs through multi-omics characterization and highlights potential personalized therapeutic vulnerabilities tailored to the specific immune landscapes of Lu-NECs.

Humans

A novel high-dimensional model for identifying regional DNA methylation QTLs.

Varying coefficient models offer the flexibility to learn the dynamic changes of regression coefficients. Despite their good interpretability and diverse applications, in high-dimensional settings, existing estimation methods for such models have important limitations. For example, we routinely encounter the need for variable selection when faced with a large collection of covariates with nonlinear/varying effects on outcomes, and no ideal solutions exist. One illustration of this situation could be identifying a subset of genetic variants with local influence on methylation levels in a regulatory region. To address this problem, we propose a composite sparse penalty that encourages both sparsity and smoothness for the varying coefficients. We present an efficient proximal gradient descent algorithm that scales to high-dimensional predictor spaces, providing sparse solutions for the varying coefficients. A comprehensive simulation study has been conducted to evaluate the performance of our approach in terms of estimation, prediction and selection accuracy. We show that the inclusion of smoothness control yields much better results over sparsity-only approaches. An adaptive version of the penalty offers additional performance gains. We further demonstrate the utility of our method in identifying regional mQTLs from asymptomatic samples in the CARTaGENE cohort. The methodology is implemented in the R package sparseSOMNiBUS, available on GitHub.

Humans

Profile and prevalence of the brain fag syndrome: psychiatric morbidity in school populations in Africa.

The profile and prevalence of a syndrome of somaticised anxiety associated with education in Africa was explored by survey of 2040 senior secondary school students in different types of school: rural, urban, and elite. Response to two different screening methods, an open question to elicit symptoms spontaneously, and the SRQ-24, was compared. Symptom prevalence was higher in rural schools, 34%, than periurban, 22%, and elite, 6%, but the central urban school serving a shanty town was also high at 35%. Three categories of the culturally relevant symptoms were identified--somatic, cognitive and 'spiritual'--with affective symptoms sparsely represented in the cultural idiom. The SRQ-24 items screening for psychosis were associated with a range of spontaneous symptoms representing anxiety. This 'spiritual' expression of neurosis reflects the world views and beliefs of the culture. Intensification under stress could produce the picture of transient reactive psychosis.

Adolescent

Connections between the retrosplenial cortex and the hippocampal formation in the rat: a review.

The retrosplenial cortex is situated at the crossroads between the hippocampal formation and many areas of the neocortex, but few studies have examined the connections between the hippocampal formation and the retrosplenial cortex in detail. Each subdivision of the retrosplenial cortex projects to a discrete terminal field in the hippocampal formation. The retrosplenial dysgranular cortex (Rdg) projects to the postsubiculum, caudal parts of parasubiculum, caudal and lateral parts of the entorhinal cortex, and the perirhinal cortex. The retrosplenial granular b cortex (Rgb) projects only to the postsubiculum, but the retrosplenial granular a cortex (Rga) projects to the postsubiculu, rostral presubiculum, parasubiculum, and caudal medial entorhinal cortex. Reciprocating projections from the hippocampal formation to Rdg originate in septal parts of CA1, postsubiculum, and caudal parts of the entorhinal cortex, but these are only sparse projections. In contrast, Rgb and Rga receive dense projections from the hippocampal formation. The hippocampal projection to Rgb originates in area CA1, dorsal (septal) subiculum, and post-subiculum. Conversely, Rga is innervated by ventral (temporal) subiculum and postsubiculum. Further, the connections between the retrosplenial cortex and the hippocampal formation are topographically organized. Rostral retrosplenial cortex is connected primarily to the septal (rostrodorsal) hippocampal formation, while caudal parts of the retrosplenial cortex are connected with temporal (caudoventral) areas of the hippocampal formation. Together, the elaborate connections between the retrosplenial cortex and the hippocampal formation suggest that this projection provides an important pathway by which the hippocampus affects learning, memory, and emotional behavior.

Afferent Pathways

CROP: a feature-independent context-aware method for CRISPR-Cas9 frameshift prediction.

MOTIVATION: The CRISPR-Cas9 complex has revolutionized genome-editing technologies. By designing a 20 nt-long guide RNA, a Cas9 nuclease can be guided to cleave almost any genomic target site (followed by NGG). The cleavage induces double-stranded DNA breaks, which are then repaired by cellular pathways. Accurate CRISPR-Cas9 repair-outcome prediction is essential for designing guide RNAs with desired genomic effects, such as gene knockout. A central challenge is quantifying the rate of frameshifts, i.e. repair-outcomes that lead to a change in the local length that is not a multiple of three. Previous methods for frameshift-rate prediction were trained on only a few experimental or cellular contexts, mostly relied on manually defined microhomology features, and were limited by sparse features and class labels. RESULTS: We developed CROP, a feature-independent context-aware repair-outcome prediction method. By aggregating specific repair outcomes as Δlength classes, CROP overcomes class sparsity. We designed CROP to work with variable input sequence lengths and output classes to utilize multiple datasets simultaneously. We benchmarked CROP against state-of-the-art repair-outcome prediction methods over 18 datasets, which we curated and standardized from various studies. Across all datasets, CROP outperformed all competing methods in frameshift-rate prediction. We performed cross-experiment and cross-cellular frameshift-rate predictions to investigate the generalizability of repair mechanisms. Finally, we show that CROP learned microhomology principles from raw sequences without explicit feature engineering, establishing an end-to-end architecture for CRISPR-Cas9 repair-outcome prediction that learns from multiple datasets. AVAILABILITY AND IMPLEMENTATION: CROP is available at https://github.com/OrensteinLab/CROP.

CRISPR-Cas Systems

Density-dependent accumulation of basic fibroblast growth factor in the subendothelial matrix.

Recent evidence indicates that basic fibroblast growth factor (bFGF), which lacks a conventional signal recognition sequence, is a component of the subendothelial matrix. However, the molecular mechanisms regulating its cellular release and subsequent matrix deposition remain equivocal. To examine the cellular and subcellular mechanisms regulating bFGF release and subendothelial sequestration, we generated polyclonal antibodies against a chemically cross-linked bFGF. We then used anti-bFGF IgG in conjunction with 3T3 cell [3H]thymidine incorporation assays, enzyme immunoassays and immunofluorescence to learn whether bFGF accumulation in the subendothelial matrix is dependent upon endothelial cell (EC)-cell contact, which coincides with growth arrest. In contrast to subconfluent cultures, which lacked any detectable extracellular matrix bFGF localization, bovine aortic and microvascular EC plated at confluent densities displayed a punctate extracellular staining pattern that was abolished when EC were pretreated with 10 micrograms/ml cycloheximide. Additionally, when EC were treated with either 1 mM beta-D xyloside, an inhibitor of proteoglycan assembly, or 100 micrograms/ml heparin, there was a 40% reduction in matrix-associated bFGF (quantified by image analysis of antibody stained cultures). 3T3 [3H]thymidine incorporation assays indicated that the beta-D xyloside-induced reduction of matrix-associated bFGF coincided with a significant increase in bFGF activity in the conditioned media. Neither sparsely-plated nor confluent EC cultures possessed specific bFGF localization of the nuclear compartment when cells were fixed using cold methanol; however, when EC were fixed in formaldehyde and lysed in isotonic buffers containing 0.1% Triton X-100 or absolute acetone, there was a marked decrease in anti-bFGF staining of the postconfluent extracellular matrix and a concomitant increase in nuclear fluorescence. Because bFGF-stimulated vascular cell growth has been implicated in controlling neointimal cell proliferation, we screened normal and atherosclerotic coronary blood vessels for bFGF, but we were unable to detect it either in lesioned or normal intima. In contrast, significant bFGF levels were observed in association with the EC and mesangial cells of the renal corpuscle, where heparan sulfate accumulates within the glomerular basement membrane. Our in vitro results suggest that bFGF accumulates within the proteoglycan-containing subendothelial matrix concomitant with the formation of cell-cell contacts. In situ, the composition of the microvascular matrix and the cellular phenotype may facilitate the selective accumulation of bFGF that we observed. This, in turn, may influence vascular morphogenesis and remodeling during angiogenesis.

3T3 Cells

ARISE: RNA-anchored shared-edge topology and hierarchical fusion for spatial multi-omics integration.

MOTIVATION: Spatial multi-omics technologies jointly profile transcriptomes, proteins and chromatin accessibility in situ, enabling integrative analysis of tissue organization across molecular layers. However, most existing graph-based integration methods rely on independently constructed modality-specific k-nearest-neighbor graphs. When auxiliary modalities are sparse or noisy, these graphs can become topologically discordant, propagate spurious edges, weaken cross-modal alignment, and reduce spatial domain resolution. RESULTS: We present Anchored RNA for Integrated Spatial Embedding (ARISE), an RNA expression anchored framework for spatial multi-omics integration. ARISE defines a shared-edge topology by intersecting RNA feature-similarity and spatial-proximity graphs, encodes auxiliary modalities on this common scaffold, and integrates them through inside-out hierarchical fusion. We further show theoretically that graph intersection minimizes false-positive edges within a broad class of k-of-r graph fusion rules, providing a principled basis for topology anchoring. Across various spatial multi-omics benchmarks spanning simulated and real datasets in bi-modal and tri-modal settings, ARISE improves spatial domain identification, cross-modal consistency, and preservation of tissue structure relative to existing methods. Furthermore, the learned representation supports biologically meaningful downstream analyses, including marker-based domain annotation, pathway enrichment, and cis-regulatory inference, indicating that ARISE yields a robust and interpretable framework for spatial multi-omics integration. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/XiangxiangWang-code/ARISE. The archived version used in this study is available at https://doi.org/10.6084/m9.figshare.32686137.v2.

Multiomics

Disagreement-informed arbitration for gene regulatory network inference: A score-level meta-classifier and a diagnostic typology of inter-method conflict.

Gene regulatory network inference methods routinely disagree about individual edges, and practitioners resolve those conflicts by choosing one method or averaging them all. We ask whether the conflict can instead be arbitrated per edge. A gradient-boosted classifier is trained on the raw scores that ten inference methods-correlation-based, information-theoretic, sparse-regression and tree-ensemble, including GENIE3, GRNBoost2, CLR and ARACNe-assign to each candidate regulator-target pair, so that the weight given to each method varies from edge to edge. Across six single-cell perturbation screens spanning four cell types, arbitration improves on mean ensembling by +0.056 AUROC on Adamson and +0.083 on Shifrut under target-grouped cross-validation. The evaluation protocol turns out to matter more than the model. Edge-level cross-validation, standard in this literature, inflates apparent gains by 0.060 AUROC through target-gene leakage-comparable to the entire honest improvement. The effect is far larger for methods that represent genes implicitly: a supervised graph-attention link predictor trained on identical folds scores AUROC 0.930 under edge-level cross-validation, better than anything else we evaluate, and 0.533 once target genes are held out. Any method that parameterises genes is exposed, which covers most graph- and embedding-based approaches. A five-category typology of inter-method conflict localises where arbitration pays off, with the largest gains on edges where the methods disagree and the smallest where they already agree, while adding nothing as model input; we therefore report it as a diagnostic instrument rather than a modelling contribution. We also characterise what the ground truth measures: most perturbed genes in widely used screens are not transcription factors, and a mediation screen bounds how much of the perturbation response can be direct.

Ensemble methods