Search PubMedSearch

SEARCH · Search PubMed

Results for “computational frameworks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

MPAC: a computational framework for inferring pathway activities from multi-omic data.

Fully capturing cellular state requires examining genomic, epigenomic, transcriptomic, proteomic, and other assays for a biological sample and comprehensive computational modeling to reason with the complex and sometimes conflicting measurements. Modeling these so-called multi-omic data is especially beneficial in disease analysis, where observations across omic data types may reveal unexpected patient groupings and inform clinical outcomes and treatments. We present Multi-omic Pathway Analysis of Cells (MPAC), a computational framework that interprets multi-omic data through prior knowledge from biological pathways. MPAC leverages network relationships encoded in pathways through a factor graph to infer consensus activity levels for proteins and associated pathway entities from multi-omic data, runs permutation testing to eliminate spurious activity predictions, and groups biological samples by pathway activities to allow identifying and prioritizing proteins with potential clinical relevance, e.g., associated with patient prognosis. Using DNA copy number alteration and RNA-seq data from head and neck squamous cell carcinoma patients from The Cancer Genome Atlas as an example, we demonstrate that MPAC predicts a patient subgroup related to immune responses not identified by analysis with either input omic data type alone. Key proteins identified via this subgroup have pathway activities related to clinical outcome as well as immune cell compositions. Our MPAC R package, available at https://bioconductor.org/packages/MPAC, enables similar multi-omic analyses on new datasets.

Journal Article

MPAC: a computational framework for inferring pathway activities from multi-omic data.

MOTIVATION: Fully capturing cellular state requires examining genomic, epigenomic, transcriptomic, proteomic, and other assays for a biological sample and comprehensive computational modeling to reason with the complex and sometimes conflicting measurements. Modeling these so-called multi-omic data is especially beneficial in disease analysis, where observations across omic data types may reveal unexpected patient groupings and inform clinical outcomes and treatments. RESULTS: We present Multi-omic Pathway Analysis of Cells (MPAC), a computational framework that interprets multi-omic data through prior knowledge from biological pathways. MPAC leverages network relationships encoded in pathways through a factor graph to infer consensus activity levels for proteins and associated pathway entities from multi-omic data, runs permutation testing to eliminate spurious activity predictions, and groups biological samples by pathway activities to allow identifying and prioritizing proteins with potential clinical relevance, e.g. associated with patient prognosis. Using DNA copy number alteration and RNA-seq data from head and neck squamous cell carcinoma patients from The Cancer Genome Atlas as an example, we demonstrate that MPAC predicts a patient subgroup related to immune responses not identified by analysis with either input omic data type alone. Key proteins identified via this subgroup have pathway activities related to clinical outcome as well as immune cell composition. Our MPAC R package enables similar multi-omic analyses on new datasets. AVAILABILITY AND IMPLEMENTATION: The MPAC package is available at Bioconductor https://bioconductor.org/packages/MPAC.

Humans

An end-to-end computational framework for "Record-seq" transcriptional recording data.

MOTIVATION: Record-seq captures cumulative transcriptional activity over time in engineered Escherichia coli by integrating cellular RNA-derived spacer sequences into clustered regularly interspaced short palindromic repeats (CRISPR) arrays, which are read out by sequencing. Unlike the approximately uniform transcript sampling of RNA-seq, Record-seq records biological signal as spacers sampled by the CRISPR spacer acquisition machinery. Consequently, standard RNA-seq analysis strategies are not directly applicable, limiting sensitivity and interpretability. Our previous pipeline addressed these challenges only partially, retained inherited RNA-seq assumptions, and had limited algorithmic efficiency. RESULTS: Here, we present an end-to-end computational framework for Record-seq data. To address the primary computational bottleneck of spacer sequence extraction, we implemented a wavefront alignment approach for efficient quasi-local pattern matching, achieving an approximately 30-fold speedup. We introduce transcription unit-based feature counting as an alternative to gene-body quantification to better represent prokaryotic transcription and increase statistical power by capturing signal from untranslated regions, which are spacer acquisition hotspots. For downstream analyses, we incorporate multiple normalization strategies and a nonparametric differential expression testing framework designed for sparse datasets. Further, we analyze spacer acquisition patterns and train sequence-based neural models that predict acquisition propensity from genomic sequence and annotations, providing a framework for assessing whether acquisition rules generalize as Record-seq is extended to new microbial hosts. AVAILABILITY AND IMPLEMENTATION: The primary analysis workflow, the recoRdseq package, acquisition modeling repository, and relevant data are all linked at https://github.com/plattlab/Record-seq-Framework. Acquisition models and training data are on Zenodo at https://doi.org/10.5281/zenodo.18891434.

Escherichia coli

ProgModule: A novel computational framework to identify mutation driver modules for predicting cancer prognosis and immunotherapy response.

BACKGROUND: Cancer originates from dysregulated cell proliferation driven by driver gene mutations. Despite numerous algorithms developed to identify genomic mutational signatures, they often suffer from high computational complexity and limited clinical applicability. METHODS: Here, we presented ProgModule, an advanced computational framework designed to identify mutation driver modules for cancer prognosis and immunotherapy response prediction. In ProgModule, we introduced the Prognosis-Related Mutually Exclusive Mutation (PRMEM) score, which optimizes the balance between exclusive mutation coverage and the incorporation of mutation combination mechanisms critical for cancer prognosis. RESULTS: Applying to BLCA and HNSC cohorts, ProgModule successfully identified driver modules that stratify patients into distinct prognostic subgroups, and the combination of these modules could serve as an effective prognostic biomarker. Extending our method to diverse cancers, ProgModule presented robust prognostic performance and stability across model parameters, including stopping criteria and network topology. Moreover, our analysis suggested that driver modules can predict immunotherapeutic benefit more effectively than existing signatures. Further analyses based on published CRISPR data indicated that genes within these modules may serve as potential therapeutic targets. CONCLUSIONS: Altogether, ProgModule emerges as a powerful tool for identifying mutation driver modules as prognostic and immunotherapy response biomarkers, and genes within these modules may be used as potential therapeutic targets for cancer, offering new insights into precision oncology.

Humans

Multimodal computational framework resolves B cell maturation in autoimmunity and ageing.

Identification of the origin of pathogenic immune cells is crucial for therapeutic interventions and diagnosis but pseudotime methods struggle to trace immune cells accurately. Current trajectory inference methods for B cell development and response in health and disease either ignore or underutilize antigen receptor sequence information, limiting their ability to resolve developmental pathways, particularly for pathogenic populations. Widely used methods such as Monocle 3 reconstruct developmental paths from transcriptomic similarity alone, discarding the features from immune receptors. Dandelion has combined the immune receptor features with transcriptomics but it struggles to simulate the trajectory path of B cells. Here we present ClonoTrace, a computational framework that integrates BCR sequence features with transcriptomic trajectory inference through gated fusion of multimodal embeddings. In fetal B cell development and germinal centre development, ClonoTrace demonstrates closer concordance with the canonical reference ordering than Monocle 3 and Dandelion. Applied to systemic lupus erythematosus, ClonoTrace indicates a memory B cell extrafollicular maturation route alongside the naïve B cell route, accompanied by induction of ZEB2 with a concomitant decline of BACH2 along the trajectory, as a candidate alternative route to pathogenic double negative 2 B cells (DN2) in systemic lupus erythematosus (SLE) patients. In healthy ageing, ClonoTrace resolved three candidate age-related B cell maturation routes, from naïve, IgM+ memory and switched-memory B cells, each passing through a DN2-associated transcriptional state that is ordered before age-associated B cells along the inferred trajectory. ClonoTrace's fate probability algorithm indicated that IgM+ memory B cell to ABC transition as the leading candidate age-associated transition, which may be distinct from SLE DN2 maturation. ClonoTrace provides a generalizable framework for receptor-informed trajectory inference, describing candidate developmental routes of pathogenic B cell populations in autoimmunity and ageing.

Humans

FusionTarget: Computational framework for drug repurposing against modeled fusion protein structures from genomic breakpoints.

Many fusion genes have been recognized as biomarkers and therapeutic targets. However, the lack of knowledge on protein structures and targeting approaches made it challenging to develop effective targeting therapeutics. To fill this, we developed a computational pipeline, FusionTarget, which annotates the genomic DNA breakage to RNA and protein sequences, predicts the 3D structures of fusion proteins, and performs comparative virtual screening, comparative molecular dynamics simulation, and quantitative analyses to identify the fusion protein-selective small molecules by selecting drugs with consistent high-fold binding affinity between fusion and wild-type proteins in multiple isoforms. We applied our pipeline to EWSR1::FLI1 in Ewing sarcoma and KMT2A::AFF1 in infant acute lymphoblastic leukemia. Further cell assay experiments confirmed that cells expressing individual fusion genes were more sensitive to the suggested drugs, and the key downstream genes were affected by our drugs. FusionTarget provides a unique foundation for developing therapeutics targeting fusion proteins.

applied computing in medical science

A python based automated computational framework to classify and comparative genomics analysis of the global diversity of chili leaf curl virus (ChiLCV) strains to understand virus host interactions.

Chili leaf curl virus (ChiLCV) is a Begomovirus chillicapsici that is one of the most devastating viruses impacted on the production of chili in the world, especially in South Asia. In the present study, we combined high-throughput computational genomics with experimental analysis of global diversity. A workflow was created using automated Python scripts to download, curate and process ChiLCV genomes from public database. About 410 complete ChiLCV genomes download from public databases. Using a phylogenetic approach, these isolates were subdivided into 34 strains, belonging to 10 major clades, showing significant genetic diversity. Geographic analysis revealed that Pakistan (207 isolates) and India (148 isolates) were the main sources of ChiLCV diversity and the remainder of the isolates were from Oman, Bangladesh, Iran, Saudi Arabia and Sri Lanka. Recombination was observed as a major evolutionary force as more than twenty recombination events were detected. Analysis of cis-regulatory elements showed a complex structure of the viral promoter, including multiple binding sites for transcription factors, hormone-response elements, light-responsive elements, and stress-responsive elements, indicating a high number of interactions between viral regulatory elements and host signaling pathways. Pangenome analysis showed the presence of a highly dynamic open pangenome made up of strain-specific orthologous groups (species-specific orthogroups). Experimental inoculation of chili plants was also carried out to assess the biological effects of infection, along with phytochemical, FTIR, HPLC, and qPCR analyses.

Begomovirus

A computational ontology framework for the synthesis of multi-level pathology reports from brain MRI scans.

BackgroundConvolutional neural network (CNN) based volumetry of MRI data can help differentiate Alzheimer's disease (AD) and the behavioral variant of frontotemporal dementia (bvFTD) as causes of cognitive decline and dementia. However, existing CNN-based MRI volumetry tools lack a structured hierarchical representation of brain anatomy, which would allow for aggregating regional pathological information and automated computational inference.ObjectiveDevelop a computational ontology pipeline for quantifying hierarchical pathological abnormalities and visualize summary charts for brain atrophy findings, aiding differential diagnosis.MethodsUsing FastSurfer, we segmented brain regions and measured volume and cortical thickness from MRI scans pooled across multiple cohorts (N = 3433; ADNI, AIBL, DELCODE, DESCRIBE, EDSD, and NIFD), including healthy controls, prodromal and clinical AD cases, and bvFTD cases. Employing the Web Ontology Language (OWL), we built a semantic model encoding hierarchical anatomical information. Additionally, we created summary visualizations based on sunburst plots for visual inspection of the information stored in the ontology.ResultsOur computational framework dynamically estimated and aggregated regional pathological deviations across different levels of neuroanatomy abstraction. The disease similarity index derived from the volumetric and cortical thickness deviations achieved an AUC of 0.88 for separating AD and bvFTD, which was also reflected by distinct atrophy profile visualizations.ConclusionsThe proposed automated pipeline facilitates visual comparison of atrophy profiles across various disease types and stages. It provides a generalizable computational framework for summarizing pathologic findings, potentially enhancing the physicians' ability to evaluate brain pathologies robustly and interpretably.

Humans

A reproducible computational transcriptomic framework for cell-type-resolved fibroinflammatory-AKT remodeling in human heart failure.

BACKGROUND: Human heart failure involves multicellular transcriptional remodeling, but public transcriptomic studies often remain disconnected from cell-type localization and perturbational interpretation. METHODS: We developed a reproducible computational workflow integrating human left-ventricular bulk transcriptomes, donor-level cell-type pseudobulk results from a human heart-failure single-cell/single-nucleus atlas, external snRNA-seq support, curated module scoring, focused ligand-receptor prioritization and LINCS/L1000 perturbational matching. RESULTS: Cross-cohort analysis identified 14,358 same-direction HF-associated genes, including 1633 replicated HF-up and 785 replicated HF-down genes. Donor-level pseudobulk analysis localized disease remodeling to cardiomyocyte, fibroblast and myeloid compartments. Activated fibroblast and inflammatory myeloid programs defined a fibroinflammatory remodeling axis connected to context-dependent AKT-associated transcriptional shifts. External snRNA-seq support was strongest for fibroblast activation and AKT-associated remodeling, with etiology-dependent heterogeneity across validation resources. L1000FWD screening prioritized safety-aware perturbational hypotheses, including glimepiride and simvastatin as interpretable candidates requiring experimental validation. CONCLUSIONS: This study provides a computational transcriptomic framework linking reproducible human HF signatures, cell-type-resolved fibroinflammatory remodeling and perturbational genomic prioritization without claiming drug efficacy or AKT causality.

Humans

Integrating structure and experimental data annotations with computational modeling framework for predicting micro-nanoplastics toxicities.

The wide use of plastic materials leads to increased emissions of micro-nanoplastics (MNPs) into the environment, raising significant concerns about their impact on human health. Traditional experimental approaches for assessing MNPs toxicity are costly, time-consuming, and there are no experimental protocols that are universally acceptable. Computational modeling using machine learning (ML) approaches provides an efficient alternative to MNP toxicity assessment. However, most modeling studies of MNPs are limited due to the lack of high-quality data and there are few previous modeling studies considering complex structures of MNPs for model training. To address this challenge, we constructed three MNP datasets with popular toxicity endpoints from various resources and used nanostructure annotation techniques to create virtual MNPs (vMNPs) for all MNP structures. The MNP structures were digitalized from annotated vMNPs, and geometrical descriptors were calculated using the Delaunay Tessellation approach. Moreover, important experimental information, such as concentrations and cell lines, were transformed into extra training variables. Partial least squares regression (PLSR) models were built using both experimental and geometrical descriptors and validated through a leave-one-out cross validation procedure. The resulting models showed reasonable performance in predicting toxicity potentials of MNPs for the three endpoints in the present datasets. Moreover, an additional library of vMNPs with their predicted properties and bioactivities was constructed, directing further research of new MNPs. This study provides three novel ML models for MNPs by integrating geometrical and experimental descriptors, which have the potential to assess new MNPs for their toxicity. The modeling strategy developed in this study can be easily expanded to model other MNP toxicity endpoints and create promising new models for MNP toxicity assessments.

Data annotation

Gut microbiota-derived metabolites target C5AR1/KDM2A/HCAR3 axis in inflammatory bowel disease: a multi-machine learning algorithms and molecular docking study.

BACKGROUND: Inflammatory bowel disease (IBD) is a chronic recurrent disorder. Gut microbiota-derived metabolites regulate intestinal homeostasis, but their molecular mechanisms in IBD remain unclear. Current studies lack systematic "microbiota-metabolite-target" network mining with multi-method validation. This study integrates network pharmacology, three machine learning algorithms, and molecular docking to construct this regulatory network in IBD. METHODS: Transcriptome data were obtained from the Gene Expression Omnibus (GEO) database. Differentially expressed genes (DEGs) were identified using limma (p < 0.05, |log2FC| > 0.5). Weighted gene co-expression network analysis (WGCNA) with an optimal soft threshold of &#x3b2; = 7 was performed to identify key module genes. Candidate genes were obtained by intersecting DEGs, gut microbiota-associated genes from the gutMGene database, and WGCNA module genes. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were conducted to explore the functional roles of candidate genes. Core genes were identified using three machine learning algorithms (LASSO, Boruta, and SVM-RFE), followed by protein-protein interaction (PPI) network analysis. Molecular docking was performed to assess the binding affinities between hub proteins and gut microbiota-derived metabolites. RESULTS: A total of 885 DEGs were identified between the IBD and control groups, including 463 upregulated and 422 downregulated genes. WGCNA identified 280 key module genes from the purple and yellow modules. The intersection of DEGs, gut microbiota-associated genes, and WGCNA module genes yielded 19 core candidate genes. PPI network analysis combined with three machine learning algorithms jointly identified C5AR1, KDM2A, and HCAR3 as core hub genes. ROC curve analysis demonstrated that all three hub genes achieved AUC values greater than 0.7 in both the training and validation sets, indicating excellent diagnostic performance for IBD. Enrichment analysis revealed significant associations with the TNF, NF-&#x3ba;B, and IL-17 signaling pathways. Molecular docking confirmed stable binding of C5AR1 with 1,3-Diphenylpropan-2-Ol (-7.87 &#xb1; 0.83 kcal&#xb7;mol-&#xb9;) and HCAR3 with 3-Indolepropionic Acid (-6.35 &#xb1; 0.70 kcal&#xb7;mol-&#xb9;), both below -5.0 kcal&#xb7;mol-&#xb9;. CONCLUSION: This study first constructs a "gut microbiota-metabolite-hub gene" axis in IBD, providing a computational framework for microbiota-targeted precision therapy, and identifying C5AR1/KDM2A/HCAR3 as computationally predicted diagnostic biomarkers and 1,3-Diphenylpropan-2-Ol/3-Indolepropionic Acid as candidate intervention molecules that warrant further experimental validation.

Molecular Docking Simulation

Decoding spatiotemporal fibrotic and cellular immunosuppression of therapeutic T cells in live pancreatic ductal adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDA) is profoundly immunosuppressive. To help define this behavior, we present integrated experimental and computational frameworks to elucidate therapeutic T cell dynamics. Through the development of TME-CARTographer (TME-CART), a computational pipeline integrating high-dimensional data, graph theory, behavior analysis, and deep learning (DL), we present quantitative insights on 4D T cell-TME interactions in live PDA tumors. Mapping physical immunosuppression demonstrates that collagen fiber architectures direct migration while concomitantly limiting off-axis movement, creating immune exclusion zones. Expanding these findings, we establish that the collagen matrix harbors and spatially organizes immunosuppressive myeloid cells to serve as cooperative co-modulators of T cell behaviors, including migration, sampling, repulsion, and sequestration. Consistent with these findings, DL defines both linear and nonlinear collagen matrix and cellular neighborhood interactions as drivers of T cell behavior. The TME-CART DL framework also accurately predicts shifts in immunosuppression following depletion of myeloid cells. Overall, we identify synergistic barriers impeding anti-tumor T cell behaviors and present TME-CART as a discovery platform for interpreting complex 4D data to enhance the understanding and design of immunotherapies.

Journal Article

Iterative, multimodal, and scalable single-cell profiling for discovery and characterization of signaling regulators.

Cell signaling plays a critical role in regulating cellular state, yet uncovering regulators of signaling pathways and understanding their molecular consequences remains challenging. Here, we present an iterative experimental and computational framework to identify and characterize regulators of signaling proteins, using the mTOR marker phosphorylated RPS6 (pRPS6) as a case study. We present a customized workflow that uses the 10x Flex assay to jointly profile intracellular protein levels, transcriptomes, and CRISPR perturbations in single cells. We use this to generate a "glossary" dataset of paired protein-RNA measurements across targeted perturbations, which we leverage to train a predictive model of pRPS6 levels based solely on transcriptomic data. Applying this model to a genome-wide Perturb-seq dataset enables in silico screening for pRPS6 and nominates novel regulators of mTOR signaling. Experimental validation confirms these predictions and reveals mechanistic diversity among hits, including changes in signaling output driven by anabolic activity, cellular proliferation and multiple stress pathways. Our work demonstrates how integrated experimental and computational approaches provide a scalable framework for multimodal phenotyping and discovery.

Journal Article

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model.

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a "promoter" or "non-promoter," which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model's ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Promoter Regions, Genetic

Accelerated long-read variant calling with Clair3 for whole-genome sequencing.

SUMMARY: The rapid growth of genomic data and increasing adoption of long-read sequencing technologies have rendered variant calling one of the most computationally demanding tasks in genomic analysis. Although deep learning-based methods currently outperform conventional approaches in distinguishing true variants from complex sequencing noise, they impose prohibitive computational and time requirements. To address this limitation, we present a computational framework based on Clair3 that integrates parallelized feature generation, enhanced variant phasing, in-memory read haplotagging, and GPU-accelerated neural network inference to accelerate variant calling. By dynamically optimizing the use of both GPU and CPU resources, our method achieves substantial runtime improvements without compromising accuracy. We evaluated our framework across a range of sequencing depths, diverse samples, and multiple hardware configurations. Our results demonstrate that the optimized pipeline completes variant calling for a 30&#xd7; whole-genome sequence in 12-20&#x2009;minutes using standard computational resources (32 CPU threads and one NVIDIA GPU), and in 12-15&#x2009;minutes on an Apple Mac Studio (32 threads), which is &#x223c;10-20-fold speedup compared with its initial release. In addition to exceptional efficiency, our method maintains state-of-the-art accuracy, achieving SNP F1-scores of 99.32% and 99.70% on 30&#xd7; ONT and PacBio GIAB HG003 datasets, respectively. This work introduces a rapid, accurate, and scalable variant calling framework that effectively supports large-cohort genomic studies and time-sensitive clinical applications. AVAILABILITY AND IMPLEMENTATION: The accelerated implementation of Clair3 is open source and available at: https://github.com/HKU-BAL/Clair3/tree/gpu.

Whole Genome Sequencing

Virtual Tumors Enable Prediction of Personalized Therapeutic Combinations for Non-Small Cell Lung Cancer.

UNLABELLED: The disease burden from non-small cell lung cancer (NSCLC) adenocarcinoma is substantial, with a million new cases diagnosed globally each year and a 5-year survival rate of less than 20%. The lack of therapeutic options personalized to individual patients leads to high variation in survival. The combination of patient stratification with personalized treatment has the potential to improve outcomes; however, the variation in mutations found in patients with NSCLC adenocarcinoma makes experimentally determining treatment combinations time-consuming and expensive. In this study, we developed an interpretable mechanistic model to decipher complex signaling interplay and guide personalized therapy in NSCLC adenocarcinoma. This "virtual tumor" model encompassed key tumor-intrinsic oncogenic signaling pathways for efficiently predicting rational drug-drug and drug-radiotherapy combination therapies in NSCLC. Diverse genetic profiles were simulated for testing more than 10,000 therapeutic strategies to identify optimal approaches to overcome resistance mechanisms specific to genetic profiles and p53 status. The virtual tumor model reproduced drug additivity screens, predicted radiosensitizing genes validated in a CRISPR screen, and identified 53BP1 as a potential drug target that improved the therapeutic window during radiotherapy. A 19-gene signature derived from the virtual tumor framework stratified patients most likely to benefit from radiotherapy, which was validated using The Cancer Genome Atlas (TCGA) data. These results show the utility of virtual tumors to predict effective therapeutic combinations and present a computational resource for large-scale screening of personalized therapies to guide clinical decision-making in patients with NSCLC. SIGNIFICANCE: A computational framework that simulates thousands of personalized treatment strategies offers a scalable, cost-effective way to tailor therapies and improve outcomes for patients with genetically diverse NSCLC.

Humans

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348&#xa0;handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

PEARL: integrative multi-omics classification and omics feature discovery via deep graph learning.

MOTIVATION: Integrating multi-omics data provides valuable insights into biological processes by capturing information across multiple molecular layers, enabling a comprehensive understanding of complex diseases and driving advancements in precision medicine. However, existing computational methods for multi-omics integration face significant challenges, such as low reliability and poor generalizability, due to the high dimensionality and low sample size nature of omics data. RESULTS: To address these challenges, we present PEARL (Pearson-Enhanced spectrAl gRaph convoLutional networks), a novel deep graph learning method for biomedical classification and functional important omics features identification. PEARL leverages a simple yet effective learning architecture to achieve superior and robust performance in high-dimensional, low-sample-size multi-omics settings. Our results demonstrate that PEARL significantly outperforms existing state-of-the-art methods on both synthetic and real biomedical datasets. Furthermore, applied to Alzheimer's disease (AD) brain multi-omics data, features prioritized by PEARL lead to functionally important genes that demonstrate significant enrichment in AD-related pathways. These findings highlight PEARL's practical utility in biomedical research and its potential to enhance biological interpretability in multi-omics studies. AVAILABILITY AND IMPLEMENTATION: The source code of our computational framework is available at https://github.com/zqq121017/PEARL.

Multiomics