Search PubMedSearch

SEARCH · Search PubMed

Results for “Computational approach”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Characterization of Tumor Antigens from Multi-omics Data: Computational Approaches and Resources.

Tumor-specific antigens, also known as neoantigens, have potential utility in anti-cancer immunotherapy, including immune checkpoint blockade (ICB), neoantigen-specific T cell receptor-engineered T (TCR-T), chimeric antigen receptor T (CAR-T), and therapeutic cancer vaccines (TCVs). After recognizing presented neoantigens, the immune system becomes activated and triggers the death of tumor cells. Neoantigens may be derived from multiple origins, including somatic mutations (single nucleotide variants, insertions/deletions, and gene fusions), circular RNAs, alternative splicing, RNA editing, and polymorphic microbiomes. An increasing amount of bioinformatics tools and algorithms are being developed to predict tumor neoantigens derived from different sources, which may require inputs from different multi-omics data. In addition, calculating the peptide-major histocompatibility complex (MHC) affinity can aid in selecting putative neoantigens, as high binding affinities facilitate antigen presentation. Based on these approaches and previous experiments, many resources have been developed to reveal the landscape of tumor neoantigens across multiple cancer types. Herein, we summarize these tools, algorithms, and resources to provide an overview of computational analysis for neoantigen discovery and prioritization, as well as the future development of potential clinical utilities in this field.

Humans

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology

In silico identification of DNMT1 inhibitors from the PlantCyc database through computational approach to assess the anti-cancer potential of nutraceutical compounds in breast cancer.

Breast cancer accounts for a disproportionate share of global cancer-related deaths, with 670,000 fatalities and 2.3 million new diagnoses recorded in women during 2022 alone. Existing treatment modalities carry considerable toxicity burdens, and resistance to available agents remains an unresolved clinical problem. DNA methyltransferase 1 (DNMT1), the enzyme chiefly responsible for maintaining genome-wide methylation patterns during DNA replication, has been mapped out as a high-value target in breast cancer because its dysregulation silences tumour suppressor genes through promoter hypermethylation. The present work involves hierarchical in silico workflow to screen 4549 plant-derived compounds from the PlantCyc database (v16.0.3) against the human DNMT1 catalytic domain (PDB ID: 4WXX). Ten top-scoring compounds were taken forward for molecular docking via AutoDock Vina; Quercetin and Kaempferol both recorded the highest binding affinities at -9.5 kcal/mol, Wogonin (-9.3 kcal/mol) and Xanthohumol (-8.1 kcal/mol) also emerged as strong binders. Pharmacokinetic evaluation using ADMET-AI confirmed that all 10 compounds met Lipinski's rule of five, with human intestinal absorption values at or above 0.98. Wogonin and Xanthohumol were selected for a 100 ns all-atom molecular dynamics (MD) simulation in GROMACS due to their well-rounded ADMET profiles and limited existing data on their specific interactions with DNMT1 in breast cancer. Across all measured trajectory metrics, backbone RMSD, residue fluctuation, radius of gyration, solvent-accessible surface area, and intermolecular hydrogen bond count, Wogonin formed a more stable, compact complex. These findings suggest that Wogonin and Xanthohumol are non-toxic nutraceutical candidates suitable for DNMT1 targeted epigenetic therapy, with computational foundation strong enough to facilitate future in vitro and in vivo validation work.

Humans

A computerized tomography-computer graphics approach to stereotaxic localization.

This paper describes the use of three-dimensional computer graphics to display structures visible in computerized tomography (CT) scans and to accurately determine optimum stereotaxic probe placement relative to those structures. A prototype Lucite stereotaxic frame designed for use in CT body scanners was fitted with a phantom consisting of small Lucite spheres representing intracranial tumors with diameters of 6 to 19 mm. A series of CT scans was obtained of the frame and phantom together. Edge outlines of the spheres were extracted from each scan in the series. A three-dimensional visual representation of the spheres was obtained by displaying their outlines from all CT scans. A superimposed visual simulation of the stereotaxic frame was adjusted interactively using several analog dials in order to simulate trajectory choice and probe depth before actual surgery. The frame settings and probe depth calculated by the graphical computer were then applied to the actual prototype frame in order to assess the accuracy of this combined CT-computer graphics approach to stereotaxic localization. The results of 22 such experiments and the implications for future clinical use of this new and precise localization technique are discussed.

Brain

A failure to improve radiologists' performances in diagnosing pulmonary lesions by a computer-aided approach.

A traditional radiological evaluation of five lung diseases and their computer-aided diagnosis with Bayes' approach, were based on the plain chest film. Parameters of diagnostic performance for 12 radiologists, who made 1224 diagnoses by each method, were compared. Under the conditions of this study the use of a computer-aided method did not improve radiologists' performances independently of their diagnostic ability and experience.

Bronchial Neoplasms

Modeling and computer stimulation approach to the mechanism of foot-and-mouth disease virus neutralization assays.

Block neutralization data since 1949 for the foot-and-mouth disease virus system have been analyzed in terms of a unified mass-action theory for computing the amounts of infectious complexes. Proof that infectious complexes contribute considerably to the assays was obtained by demonstrating a reduction in titer after an additional reaction with anti-Ig antibody before the assay. In the suckling-mice assay with intraperitoneal inoculation, both the data of others and our own on several types indicate that for IgG probably three of the unknown total number of critical sites on the virion must be available for infectivity and death. For IgM, just one of an unknown different total number of critical sites on the virion must be available. In tissue culture infectivity assays the minimum number is two or three, whereas in the bovine tongue assay it could be one or two, but probably two. The difficulties in establishing the at present unknown total numbers of neutralization sites to both IgG and IgM are considerable. However, by the simplest interpretation of the data, the number is estimated to be between 5 and 10 for IgG and perhaps just 1 for IgM. A speculation, consistent with the known virion architecture, is that just 1 of the 12 vertices is uniquely involved in infectivity and death, at least in the suckling-mice assay.

Animals

Bioinformatic approaches for accurate assessment of A-to-I editing in complete transcriptomes.

A-to-I RNA editing is an RNA modification that alters the RNA sequence relative to the its genomic blueprint. It is catalyzed by double-stranded RNA-specific adenosine deaminase (ADAR) enzymes, and contributes to the complexity and diversification of the proteome. Advancement in the study of A-to-I RNA editing has been facilitated by computational approaches for accurate mapping and quantification of A-to-I RNA editing based on sequencing data. In this chapter we review some of the main computational approaches currently used, describe potential hurdles, challenges and pitfalls, and discuss possible ways to mitigate them.

RNA Editing

Dock & design: engineering specificity for an alternative pimaradiene outcome with the ent-kaurene synthase from Bradyrhizobium japonicum.

The complexity of the reactions catalyzed by terpene synthases has hindered enzymatic engineering. In most cases such efforts result in non-specific product outcome, with the targeted compound being produced alongside others, hindering further use. Previous work with the structurally characterized ent-kaurene synthase from Bradyrhizobium japonicum (BjKS) identified a serine for alanine substitution (A167S) that led to premature deprotonation, yielding a pair of ent-pimaradiene double-bond isomers, with retrospective analysis by the TerDockin computational approach indicating that the introduced hydroxyl acts as a catalytic base for both. Here this route to 'short-circuiting' the BjKS catalyzed reaction for ent-pimaradiene production was further explored, with prospective application of TerDockin, via design-build-test cycles, enabling specific production of a novel pimaradiene isomer via introduction of a water molecule as the catalytic base. The resulting mutants, BjKS:F72S and particularly BjKS:F72Y/Y280S specifically yield the targeted ent-pimara-8,15-diene with reasonable catalytic efficiency, demonstrating the applicability of this computationally inexpensive approach to engineering terpene synthase product outcomes.

Journal Article

Reversal of cancer gene expression identifies repurposed drugs for diffuse intrinsic pontine glioma.

Diffuse intrinsic pontine glioma (DIPG) is an aggressive incurable brainstem tumor that targets young children. Complete resection is not possible, and chemotherapy and radiotherapy are currently only palliative. This study aimed to identify potential therapeutic agents using a computational pipeline to perform an in silico screen for novel drugs. We then tested the identified drugs against a panel of patient-derived DIPG cell lines. Using a systematic computational approach with publicly available databases of gene signature in DIPG patients and cancer cell lines treated with a library of clinically available drugs, we identified drug hits with the ability to reverse a DIPG gene signature to one that matches normal tissue background. The biological and molecular effects of drug treatment was analyzed by cell viability assay and RNA sequence. In vivo DIPG mouse model survival studies were also conducted. As a result, two of three identified drugs showed potency against the DIPG cell lines Triptolide and mycophenolate mofetil (MMF) demonstrated significant inhibition of cell viability in DIPG cell lines. Guanosine rescued reduced cell viability induced by MMF. In vivo, MMF treatment significantly inhibited tumor growth in subcutaneous xenograft mice models. In conclusion, we identified clinically available drugs with the ability to reverse DIPG gene signatures and anti-DIPG activity in vitro and in vivo. This novel approach can repurpose drugs and significantly decrease the cost and time normally required in drug discovery.

Humans

Panaln: indexing pangenome for read alignment.

MOTIVATION: Pangenome indexing is a critical supporting technology in biological sequence analysis such as read alignment applications. The need to accurately identify billions of small sequencing fragments carrying sequencing errors and genomic variants drives the development of scalable and efficient pangenome indexing approach. RESULTS: We propose a new wavelet tree-based approach, called Panaln, for indexing pangenome and introduce a batch computation approach for fast count query over Panaln. We present a simple and effective seeding strategy and develop a pangenome program that uses the seed-and-extend paradigm for read alignment. Experimental results on simulated and real data demonstrate that Panaln uses significantly less space for the compared pangenome methods with generally higher accuracy. We provide a scalable index construction by representing pangenome with a linear model. Additionally, Panaln brings enhanced accuracy compared to the popular single reference methods. AVAILABILITY AND IMPLEMENTATION: Package: https://anaconda.org/bioconda/panaln and source code: https://github.com/Lilu-guo/Panaln.

Software

Disease candidate genes prediction using positive labeled and unlabeled instances.

Identifying disease genes and understanding their performance is critical in producing drugs for genetic diseases. Nowadays, laboratory approaches are not only used for disease gene identification but also using computational approaches like machine learning are becoming considerable for this purpose. In machine learning methods, researchers can only use two data types (disease genes and unknown genes) to predict disease candidate genes. Notably, there is no source for the negative data set. The proposed method is a two-step process: The first step is the extraction of reliable negative genes from a set of unlabeled genes by one-class learning and a filter based on distance indicators from known disease genes; this step is performed separately for each disease. The second step is the learning of a binary model using causing genes of each disease as a positive learning set and the reliable negative genes extracted from that disease. Each gene in the unlabeled gene's production and ranking step is assigned a normalized score using two filters and a learned model. Consequently, disease genes are predicted and ranked. The proposed method evaluation of various six diseases and Cancer class indicates better results than other studies.

Humans

Competing subclones and fitness diversity shape tumor evolution across cancer types.

MOTIVATION: Intratumor heterogeneity arises from ongoing somatic evolution and complicates cancer diagnosis, prognosis, and treatment. Reconstructing evolutionary dynamics typically requires spatiotemporal samples, which are often unavailable in clinical settings. Computational approaches that can infer tumor evolutionary history from single-timepoint bulk sequencing data remain limited. RESULTS: We present estimating evolutionary events through single-timepoint sequencing (TEATIME), a novel computational framework that models tumors as mixtures of two competing cell populations: an ancestral clone with baseline fitness and a derived subclone with elevated fitness. Using cross-sectional bulk sequencing data, TEATIME estimates mutation rates, timing of subclone emergence, relative fitness, and number of generations of growth. To quantify intratumor fitness asymmetries, we introduce a novel metric-fitness diversity-which captures the imbalance between competing cell populations and serves as a measure of functional intratumor heterogeneity. Applying TEATIME to 33 tumor types from The Cancer Genome Atlas, we revealed divergent as well as convergent evolutionary patterns. Notably, we found that immune-hot microenvironments constraint subclonal expansion and limit fitness diversity. Moreover, we detected temporal dependencies in mutation acquisition, where early driver mutations in ancestral clones epistatically shape the fitness landscape, predisposing specific subclones to selective advantages. These findings underscore the importance of intratumor competition and tumor-microenvironment interactions in shaping evolutionary trajectories, driving intratumor heterogeneity. Lastly, we demonstrate that TEATIME-derived evolutionary parameters and fitness diversity offer novel prognostic insights across multiple cancer types. AVAILABILITY AND IMPLEMENTATION: R implementation of TEATIME is available on GitHub (https://github.com/liliulab/TEATIME) and Zenodo (https://zenodo.org/records/17422174).

Neoplasms

FIERCE: reconstructing dynamic trajectories from the differentiation potency of single cells.

MOTIVATION: Since the introduction of single-cell RNA sequencing (scRNA-seq), numerous computational approaches have been developed to reconstruct dynamic cellular processes from static transcriptional profiles. These methods order cells along continuous trajectories by assessing their similarity in the gene-expression space. However, they rely on several assumptions, such as prior knowledge of the structure and directionality of the expected genealogy. These assumptions can limit their application to complex cellular systems with poorly understood developmental paths. RESULTS: To address this challenge, we introduce FIERCE (Framework for InfERence of the veloCity of Entropy), a novel computational pipeline designed to predict the changes in the differentiation potency of single cells during dynamic processes. Through a fully unsupervised approach, FIERCE enables the inference of cell lineages directly on the differentiation landscape of the biological system, thus eliminating the need for prior specification of developmental parameters. We demonstrate the efficacy of FIERCE by reconstructing three well-known mouse differentiation systems and by quantifying its accuracy on simulated data. AVAILABILITY AND IMPLEMENTATION: The FIERCE R package is available on GitHub at https://github.com/bicciatolab/FIERCE.

Cell Differentiation

Recent Advances in Multi-Omics of Systemic Lupus Erythematosus.

This comprehensive narrative review examines recent advances in multi-omics research for Systemic Lupus Erythematosus (SLE), emphasizing integrated approaches over single-omics studies. The review critically evaluates technological advancements, methodological innovations, and clinical applications while identifying current limitations and future research directions. We conducted a comprehensive narrative review following SANRA guidelines, searching PubMed, Web of Science, Scopus, and Embase, covering publications from January 2018 to June 2025. The review focuses on studies integrating two or more omics layers in SLE research, with emphasis on computational methods, biomarker validation, and clinical applications. Multi-omics integration has revealed critical insights into SLE pathogenesis, including immune cell heterogeneity, gene-environment interactions, and metabolic dysregulation. However, significant challenges remain in data integration methodologies, small sample sizes, and biomarker reproducibility. Current computational approaches include early integration (concatenation), intermediate integration (joint dimensionality reduction), and late integration (ensemble methods). While multi-omics approaches offer unprecedented insights into SLE complexity, standardized integration protocols and robust validation frameworks are urgently needed. Small sample sizes and heterogeneity issues limit reproducibility, particularly affecting biomarker discovery and clinical translation. Multi-omics integration represents a paradigm shift toward precision medicine in SLE, but realizing this potential requires addressing current methodological limitations, standardizing validation processes, and developing robust computational frameworks for reliable clinical applications.

Humans

Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.

MOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998.

Promoter Regions, Genetic

Annotation of RxLR Effectors in Oomycete Genomes.

Pathogens have evolved effector proteins to suppress host immunity and facilitate plant infections. RxLR effectors are small, secreted effector proteins with conserved RxLR and dEER amino acid motifs at the N terminus and highly variable C termini and are commonly found in oomycete species. We provide computational approaches to annotate RxLR candidate effector genes in a genome assembly in FASTA format with an available GFF file. Hidden Markov Modeling (HHM) is used in combination with regular expressions to search for RxLR and EER amino acid patterns.

Oomycetes

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (Ψ) represents one of the most abundant and conserved RNA modifications. Ψ provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of Ψ sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel Ψ site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA Ψ-site prediction. The Ψ modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA Ψ-site prediction. Meta-PseU offers a new framework for robust Ψ-site identification by using long sequences.

Pseudouridine