Search PubMedSearch

SEARCH · Search PubMed

Results for “Computational methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Prediction and Evaluation of Protein Aggregation with Computational Methods.

Protein and peptide aggregation has recently become one of the most studied biomedical problems due to its central role in several neurodegenerative disorders and of biotechnological importance. Multiple in silico methods, databases, tools, and algorithms have been developed to predict aggregation of proteins and peptides to better understand fundamental mechanisms of various aggregation diseases. Here, we attempt to provide a brief overview of bioinformatic methods and tools to better understand molecular mechanisms of aggregation disorders. Furthermore, through a better understanding of protein aggregation mechanisms, it might be possible to design novel therapeutic agents to treat and hopefully prevent protein aggregation diseases.

Computational Biology

Development of a multi-copy integration platform in Kluyveromyces marxianus enabled by a computational method for genome-wide identification of multi-copy integration loci.

Multi-copy integration is a core strategy for redirecting metabolic flux toward target compounds. However, its application has been hampered by the absence of methods for systematically identifying native multi-copy genomic loci. To overcome this, we developed a computational procedure for genome-wide identification of such loci. Theoretically, this method is potentially applicable to any genome-sequenced species as it only requires the genomic assembly of the target species as input. Applying the procedure to Kluyveromyces marxianus, we identified four groups of loci (KmCS1-4). Combining these loci-KmCS1-4 and the traditional 26S rDNA-with 14 markers with graded selection strengths, we established a versatile multi-copy integration toolkit comprising 70 plasmids. Each plasmid exhibits a unique integration pattern, collectively forming an integration profile. This profile serves as a manual, enabling users to select appropriate tools tailored to the expression requirements of rate-limiting enzymes in their pathways. Applying representative plasmids exhibiting low-, medium-, and high-copy integration patterns to lycopene biosynthesis modules resulted in lycopene titers of 3.5, 6.8 and 40.5 mg/L, corresponding to 2, 6 and 9 genomic copies, respectively, demonstrating a positive correlation between lycopene titers, genomic copy numbers and integration patterns, which highlights the versatility of the toolkit and its supporting manual. Our study not only provides a broadly applicable methodology for genome-wide identification of multi-copy loci, but also an efficient integration platform for K. marxianus.

Kluyveromyces marxianus

Absolute copy number aware CNV calling of sub-megabase segments in ultra-low coverage single-cell DNA sequencing data.

Recent advances in ultra-low coverage whole-genome sequencing (WGS) of single cells have enabled detailed analysis of copy number variation at a throughput approaching that of single-cell RNA sequencing. However, downstream computational methods have not seen comparable advances and are largely adaptations of deep sequencing methodology with reduced precision. Here, we present ASCENT, a computational method built to take full advantage of modern direct tagmentation-based WGS at ultra-low depth. Using joint segmentation with high-resolution bins, we accurately detect small segments, achieving accurate copy number profiles even at 100 000 reads per cell. ASCENT implements true absolute copy state inference for single cells, based on statistical modeling of coverage rather than comparison to a reference, while taking variable segment copy state into account. Further, ASCENT implements per-segment copy-neutral loss of heterozygosity (LOH) calling without the need for non-tumor or bulk WGS reference. When applied to a pediatric B-ALL sample, ASCENT finds copy-neutral LOH in a small segment and a minor subclone defined by breakpoints missed in bulk WGS. Thus, by applying appropriate computational methods, single-cell WGS provides clear advantages over bulk, even at a relatively low cell number and sequencing depth.

DNA Copy Number Variations

Accurate prediction of toxicity peptide and its function using multi-view tensor learning and latent semantic learning framework.

MOTIVATION: Therapeutic peptide is an important ingredient in the treatment of various diseases and drug discovery. The toxicity of peptides is one of the major challenges in peptide drug therapy. With the abundance of therapeutic peptides generated in the post-genomics era, it is a challenge to promptly identify toxicity peptides using computational methods. Although several efforts have been made, few algorithms are designed to identify whether a query peptide exhibits toxicity. Considering the varied levels of biological activities, the toxicity peptides should be further classified into multi-functional peptides. RESULTS: This study introduces a two-level predictor, ToxPre-2L, developed using the multi-view tensor learning and latent semantic learning framework. The proposed method utilized multi-label learning with feature induced labels to avoid the redundancy of information from each view. Then the multi-view tensor learning was employed to establish the latent semantic information among different views, while low-rank constraint learning was leveraged to exploit the correlation information among multi-labels. Finally, we constructed an updated toxicity peptide benchmark dataset to assess the effectiveness of the proposed method. Experimental results demonstrated that ToxPre-2L achieves a better performance than alternative computational methods in the prediction of toxicity peptides and their multi-functional types. AVAILABILITY AND IMPLEMENTATION: The source code and data of ToxPre-2L can be accessed at http://bliulab.net/ToxPre-2L.

Peptides

A multi-omic analysis of MCF10A cells provides a resource for integrative assessment of ligand-mediated molecular and phenotypic responses.

The phenotype of a cell and its underlying molecular state is strongly influenced by extracellular signals, including growth factors, hormones, and extracellular matrix proteins. While these signals are normally tightly controlled, their dysregulation leads to phenotypic and molecular states associated with diverse diseases. To develop a detailed understanding of the linkage between molecular and phenotypic changes, we generated a comprehensive dataset that catalogs the transcriptional, proteomic, epigenomic and phenotypic responses of MCF10A mammary epithelial cells after exposure to the ligands EGF, HGF, OSM, IFNG, TGFB and BMP2. Systematic assessment of the molecular and cellular phenotypes induced by these ligands comprise the LINCS Microenvironment (ME) perturbation dataset, which has been curated and made publicly available for community-wide analysis and development of novel computational methods ( synapse.org/LINCS_MCF10A ). In illustrative analyses, we demonstrate how this dataset can be used to discover functionally related molecular features linked to specific cellular phenotypes. Beyond these analyses, this dataset will serve as a resource for the broader scientific community to mine for biological insights, to compare signals carried across distinct molecular modalities, and to develop new computational methods for integrative data analysis.

Epidermal Growth Factor

Generative AI Models in Time-Varying Biomedical Data: Scoping Review.

BACKGROUND: Trajectory modeling is a long-standing challenge in the application of computational methods to health care. In the age of big data, traditional statistical and machine learning methods do not achieve satisfactory results as they often fail to capture the complex underlying distributions of multimodal health data and long-term dependencies throughout medical histories. Recent advances in generative artificial intelligence (AI) have provided powerful tools to represent complex distributions and patterns with minimal underlying assumptions, with major impact in fields such as finance and environmental sciences, prompting researchers to apply these methods for disease modeling in health care. OBJECTIVE: While AI methods have proven powerful, their application in clinical practice remains limited due to their highly complex nature. The proliferation of AI algorithms also poses a significant challenge for nondevelopers to track and incorporate these advances into clinical research and application. In this paper, we introduce basic concepts in generative AI and discuss current algorithms and how they can be applied to health care for practitioners with little background in computer science. METHODS: We surveyed peer-reviewed papers on generative AI models with specific applications to time-series health data. Our search included single- and multimodal generative AI models that operated over structured and unstructured data, physiological waveforms, medical imaging, and multi-omics data. We introduce current generative AI methods, review their applications, and discuss their limitations and future directions in each data modality. RESULTS: We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and reviewed 155 articles on generative AI applications to time-series health care data across modalities. Furthermore, we offer a systematic framework for clinicians to easily identify suitable AI methods for their data and task at hand. CONCLUSIONS: We reviewed and critiqued existing applications of generative AI to time-series health data with the aim of bridging the gap between computational methods and clinical application. We also identified the shortcomings of existing approaches and highlighted recent advances in generative AI that represent promising directions for health care modeling.

Artificial Intelligence

A systematic review of human avoidance learning: Cognition, computation, and methods.

Avoidance behaviour is fundamental for survival but can become maladaptive in clinical conditions. A large body of literature has accumulated on the dynamics of human avoidance learning. However, current theories and overviews do not provide an exhaustive account of this evidence. In this systematic review, we identify N = 116 studies on human avoidance learning. We analyse these studies with the goal of distilling robust empirical phenomena as a basis for theory-building, and examine their diagnostic value in differentiating between competing theories. We find that the evidence is difficult to reconcile with foundational two-factor and classical safety-signal accounts, and most strongly supports expectancy- and inference-based views, in which avoidance responses are selected with respect to represented consequences. At the same time, no current framework provides a complete account of the evidence: several findings point to an additional role for operant valuation, Pavlovian influences, and contextual or latent-state control over the expression of avoidance. Methodologically, we observe that the problem setting in the most common experimental paradigms is radically simpler than real-world avoidance and therefore unlikely to expose the limits of inferential or reflective mechanisms. Consequently, we argue that paradigms with greater computational demands and more realistic action affordances are required to identify the mechanisms underlying avoidance learning. Collectively, these insights provide a foundation for theoretical refinement, computational modelling, and methodological innovation, with implications for advancing interventions targeting maladaptive avoidance.

Humans

Likelihood-based optimization enables accurate copy number estimation for paralogous genes using exome data.

MOTIVATION: Exome sequencing is widely used for genetic studies; however, accurate detection of copy number variants (CNV) in paralogous genes is challenging due to short-read mapping ambiguity and extensive copy-number variation. The human genome contains several hundred paralogous genes, many of which are known to harbor disease-associated CNVs. Existing exome CNV callers are primarily designed for rare CNV detection in uniquely mappable regions and are not well-suited for paralogous genes. METHODS: We describe a computational method (EdgeCopy) for copy number profiling of paralogous genes using whole-exome sequence data. EdgeCopy aggregates reads mapped to all copies of paralogous genes and relates observed read depth to copy number for multiple exome samples using an approximate composite likelihood function. The likelihood function is optimized using numerical optimization to obtain gene-level fractional copy number estimates that are discretized and refined using a Hidden Markov Model to obtain exon-level copy number estimates. RESULTS: Benchmarking of Edgecopy using experimental copy number data showed high concordance (mean = 0.973) for six disease-associated paralogous genes. We evaluated performance using whole-exome data from approximately 2400 samples across five continental populations from the 1000 Genomes Project. EdgeCopy shows robust concordance with whole-genome sequencing based estimates (0.974-0.982) across populations and 130 paralogous genes spanning a wide range of copy-number variation. In comparison, copy number analysis using a state-of-the-art exome CNV caller failed to estimate copy number for paralogous genes with very high mapping ambiguity and showed much lower concordance (0.565) for CNV events compared to EdgeCopy (0.908). AVAILABILITY: EdgeCopy is freely available at https://github.com/vibansal-lab/edgecopy.

Humans

Drug repurposing in status epilepticus.

The treatment of status epilepticus (SE) has changed little in the last 20 years, largely because of the high risks and costs of new drug development for SE. Moreover, SE poses specific challenges to drug development, such as patient diversity, logistical hurdles, and the need for acute treatment strategies that differ from chronic seizure prevention. This has reduced the appetite of industry to develop new drugs in this area. Drug repurposing is an attractive approach to address this unmet need. It offers significant advantages, including reduced development time, lower costs, and higher success rates, compared to novel drug development. Here I demonstrate how novel methods integrating biological knowledge and computational methods can be applied to drug repurposing in status epilepticus. Biological approaches focus on addressing mechanisms underlying drug resistance in SE (using for example ketamine, tacrolimus and safinamide) and longer-term consequences (using for example omaveloxolone, celecoxib and losartan). Additionally, artificial intelligence platforms, such as ChatGPT, can rapidly generate promising drug lists, while in silico methods can analyze gene expression changes to predict molecular targets. Combining AI and in silico approaches has identified several candidate drugs, including metformin, sirolimus and riluzole, for SE treatment. Despite the promise of repurposing, challenges remain, such as intellectual property issues and regulatory barriers. Nonetheless, drug repurposing presents a viable solution to the high costs and slow progress of traditional drug development for SE. This paper is based on a presentation made at the 9th London-Innsbruck Colloquium on Status Epilepticus and Acute Seizures, in April 2024.

Animals

Decoding the Functional Interactome of Non-Model Organisms with PHILHARMONIC.

Despite the widespread availability of genome sequencing pipelines, many genes remain part of the genome's "dark matter," where existing inference tools cannot even begin to guess the biological function of their proteins from sequence alone. This challenge is especially pronounced in organisms that are highly evolutionarily distant from well-studied models, where homology-based methods break down. Here, we describe PHILHARMONIC, a computational method that combines deep learning-based de novo protein interaction network inference with robust unsupervised spectral clustering and remote homology to illuminate functional organization in any non-model organism. From only a sequenced proteome, we show PHILHARMONIC predicts protein functions, functional communities, and higher-order network structure with high accuracy. We validate its performance using experimental gene expression and pathway data in D. melanogaster, and we demonstrate its broad utility by analyzing temperature sensing and stress response pathways in the reef-building coral P. damicornis and its algal symbiont C. goreaui. PHILHARMONIC provides a general-purpose engine for functional discovery and biological hypothesis generation in non-model organisms, enabling systems-level insights across the full diversity of life.

Journal Article

PLAID: ultrafast single-sample gene set enrichment scoring.

SUMMARY: In recent years, computational methods have emerged that calculate enrichment of gene signatures within individual samples. These signatures offer critical insights into the coordinated activity of functionally related genes, proteins or metabolites, enabling the identification of unique molecular profiles in individual cells and patients. This strategy is pivotal for patient stratification and advancement of personalized medicine. However, the rise of large-scale datasets, including single-cell profiles and population biobanks, has exposed significant computational inefficiencies in existing methods. Current methods often demand excessive runtime and memory resources, becoming impractical for large datasets. Overcoming these limitations is a focus of current efforts by bioinformatics teams in academia and the pharmaceutical industry, as essential to support basic and clinical biomedical research. To address this critical need, we developed PLAID (Pathway Level Average Intensity Detection), an ultrafast and memory optimized single sample gene set enrichment algorithm that utilizes sparse matrix computation. PLAID delivers highly accurate gene set scoring and surpasses the performance of current methods in single-cell and bulk transcriptomics, and proteomics data. PLAID uniquely integrates the most widely used gene set scoring algorithms, enabling researchers to apply multiple methods for cross-validation with outstanding runtime efficiency and minimal memory requirement. AVAILABILITY AND IMPLEMENTATION: PLAID is implemented in the R language for statistical computing. PLAID source code and installation instructions are available with no restrictions at https://github.com/bigomics/plaid.

Algorithms

pKAKA: a protein language model for prioritizing kinase-disrupting variants in diseases.

Protein kinases are pivotal regulators of cellular signaling, and their genetic variations are frequently implicated in diseases. Although numerous kinase mutations have been identified as drivers of altered activity, with a few successfully targeted therapeutically, the functional impact of most variants remains uncharacterized. To bridge this gap, we curate a comprehensive dataset that contains 2553 experimentally validated kinase activity-related key alterations (KAKAs) from the literature. While many mutations outside canonical functional regions are known to affect kinase activity, systematic methods to predict their functional consequences are lacking. Consequently, we develop a computational method to predict potential KAKAs, leveraging transfer learning on the pre-trained protein language model ProtBert. Our model, termed pKAKA, achieves an impressive AUC score of 0.9593 and outperforms the AlphaMissense benchmark in comparative testing. Systematic analysis of kinase missense mutations underscores the critical role of KAKAs in pathogenesis, with highlights including JAK2 V617F in atherosclerotic cardiovascular disease, LRRK2 G2385R in Parkinson's disease, EGFR L858R in lung adenocarcinoma, and EGFR G598V in glioma. Overall, this study significantly advances our understanding of how mutations that influence kinase activity contribute to disease mechanisms.

Humans

Diagnosing scientific replicability through probabilistic distinguishability.

MOTIVATION: Despite the widely recognized importance of replicability in biological research, computational methods to quantify irreplicability and identify irreplicable instances remain underdeveloped. This article presents an efficient and robust computational framework to address this gap. RESULTS: To tackle the challenge of defining an acceptable level of intrinsic heterogeneity among replicable studies, we introduce a distinguishability criterion, ensuring that replicable effects, while potentially heterogeneous, can be distinguished from zero effects and maintain consistent directions with high probability. We implement a Bayesian model criticism approach, reporting a Bayesian P-value to identify potential irreplicable instances. Through numerical experiments, we demonstrate the efficacy of the proposed methods in detecting batch effects in high-throughput experiments and identifying instances of the publication bias. Finally, we apply the framework to multi-tissue eQTL data from the GTEx consortium, uncovering tissue-specific eQTLs that represent biological heterogeneity across tissues. AVAILABILITY AND IMPLEMENTATION: An R package DiscRep implementing our method is available on GitHub (https://github.com/PengWang96/DiscRep).

Bayes Theorem

HiCPotts: An R/Bioconductor package to identify significant interactions in chromosome conformation capture data and model sources of bias.

MOTIVATION: Chromosome Conformation Capture methods, including Hi-C, micro-C or Capture-C, are used to map chromatin interactions genome-wide. Most of the existing computational methods do not account for sources of bias (such as DNA accessibility, GC content or TE content) in the data. RESULTS: We previously developed ZipHiC, a Bayesian method based on the hidden Markov random field (HMRF) model and the Approximate Bayesian Computation (ABC), that uses zero-inflated Poisson distribution to model the noise, signal and false signal of the data and showed that this approach was able to detect bias from DNA accessibility, GC content and TE content in both Hi-C and micro-C data. Here, we present HiCPotts, another Bayesian method based on the HMRF model and the ABC that uses a zero-inflated Negative Binomial distribution instead to model the noise and signal of the data. We systematically show that HiCPotts reduces false positives and increases recovery of true interactions compared to ZipHiC, but also compared to other methods such as FastHiC, Juicer and HiCExplorer. Most importantly, we provide an R/Bioconductor package that allows modelling the noise, signal and false signal using various distributions such as the zero-inflated Negative Binomial (ZINB) and the zero-inflated Poisson distribution (ZIP). AVAILABILITY AND IMPLEMENTATION: https://bioconductor.org/packages/HiCPotts/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Approximate Bayesian Computation

Computational strategies for copy number variation detection, disease association, and beyond.

Copy number variations (CNVs) are key structural variations that contribute to human genetic diversity, evolution, and disease susceptibility. Advances in sequencing technologies and computational methods have improved CNV detection, yet association studies remain challenged by methodological limitations and a lack of standardisation. This review provides an overview of computational strategies for germline CNV detection and disease association. We highlight the value of CNV analysis for uncovering genetic contributions to complex traits and disease risk and outline an analysis workflow including key benchmarking methods. We also discuss current challenges and future directions for advancing CNV detection and association analysis.

Humans

Identifying fate-determining transcription factors with single-cell omics.

Single-cell sequencing enables the systematic discovery of cell fate-determining transcription factors (TFs), or key TFs, that define cellular identity or drive cell state transitions. A wide range of computational methods have been developed for this goal, but they differ substantially in the input data and the biological questions they address. In this article, we systematically review computational approaches for key TF identification and organize them from three perspectives: whether they identify TFs defining cell state identity or driving state transitions, whether transitions are modeled as discrete or continuous processes, and whether TFs act individually or combinatorially. We summarize key features and application scenarios of relevant methods to guide tool selection and discuss emerging trends in this field toward programmable and active control of cell fate.

Transcription Factors

Recent Advances in Multi-Omics of Systemic Lupus Erythematosus.

This comprehensive narrative review examines recent advances in multi-omics research for Systemic Lupus Erythematosus (SLE), emphasizing integrated approaches over single-omics studies. The review critically evaluates technological advancements, methodological innovations, and clinical applications while identifying current limitations and future research directions. We conducted a comprehensive narrative review following SANRA guidelines, searching PubMed, Web of Science, Scopus, and Embase, covering publications from January 2018 to June 2025. The review focuses on studies integrating two or more omics layers in SLE research, with emphasis on computational methods, biomarker validation, and clinical applications. Multi-omics integration has revealed critical insights into SLE pathogenesis, including immune cell heterogeneity, gene-environment interactions, and metabolic dysregulation. However, significant challenges remain in data integration methodologies, small sample sizes, and biomarker reproducibility. Current computational approaches include early integration (concatenation), intermediate integration (joint dimensionality reduction), and late integration (ensemble methods). While multi-omics approaches offer unprecedented insights into SLE complexity, standardized integration protocols and robust validation frameworks are urgently needed. Small sample sizes and heterogeneity issues limit reproducibility, particularly affecting biomarker discovery and clinical translation. Multi-omics integration represents a paradigm shift toward precision medicine in SLE, but realizing this potential requires addressing current methodological limitations, standardizing validation processes, and developing robust computational frameworks for reliable clinical applications.

Humans

Next-Generation Disease Profiling by Integrating Histopathology with Spatial Multi-Omics Data.

The field of pathology has experienced several transformative changes in recent years with the advent of digital pathology and spatial multi-omics. These technologies have enhanced every aspect of pathology practice, from streamlining daily workflows to generating high-fidelity multi-omics data that provide pathologists with novel tools to refine disease profiling and clinical diagnosis. Each layer of multimodal data (genomic, metabolomic, proteomic, or transcriptomic) has uncovered a distinct facet of disease pathologies, and combined with machine learning/artificial intelligence-based data analysis and pattern recognition models, has provided holistic understanding of regulatory mechanisms underpinning them. However, high-dimensional data have far exceeded the volume, scale, and complexity of immunostaining methods implemented by pathologists and, thus, have generated significant challenges related to deconvolution, interpretation, and clinical translation. Furthermore, these multimodal studies have predominantly relied on computational methods to process data and extract disease-relevant insights, thus raising questions around relevance or role of a pathologist in this new era of multi-omics. This review will provide a perspective on the evolving fields of molecular histopathology and spatial -omics, leveraging them to approach disease profiling, and redefining the role of a pathologist during this process.

Humans