Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatics workflow”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

myGrid: personalised bioinformatics on the information grid.

MOTIVATION: The (my)Grid project aims to exploit Grid technology, with an emphasis on the Information Grid, and provide middleware layers that make it appropriate for the needs of bioinformatics. (my)Grid is building high level services for data and application integration such as resource discovery, workflow enactment and distributed query processing. Additional services are provided to support the scientific method and best practice found at the bench but often neglected at the workstation, notably provenance management, change notification and personalisation. RESULTS: We give an overview of these services and their metadata. In particular, semantically rich metadata expressed using ontologies necessary to discover, select and compose services into dynamic workflows.

Computational Biology↗

Portable metagenomics for preventive surveillance and outbreak control in livestock and poultry: Pathogen detection, resistome profiling, and antimicrobial stewardship.

Conventional diagnostics for livestock and poultry outbreaks commonly rely on culture or targeted PCR panels, which may be too slow or too narrow to guide early control decisions. Portable metagenomics, particularly real-time nanopore sequencing, offers a route to broad pathogen detection, antimicrobial-resistance gene profiling, and outbreak investigation within an integrated workflow. This implementation-focused review evaluates how near-point-of-care metagenomics may support preventive veterinary medicine through earlier detection, surveillance, cohorting, biosecurity decisions, and antimicrobial stewardship. We synthesize sample-to-answer workflows for enteric and respiratory disease in food-producing animals, including sampling, nucleic-acid extraction, host depletion or target enrichment, library preparation, sequencing, bioinformatics, quality control, and interpretation. Applications in calf diarrhea, bovine respiratory disease, poultry outbreaks, mastitis, and resistome monitoring are considered alongside the central limitation that detection alone does not establish causation. Pathogen and resistance-gene signals must therefore be interpreted with clinical signs, lesions, epidemiology, controls, and confirmatory testing. We also propose a minimum reporting checklist, intended as a practical framework rather than a validated consensus standard. Portable metagenomics is not a replacement for conventional diagnostics, but appropriately validated workflows can reduce uncertainty during time-sensitive outbreaks and support more judicious antimicrobial use.

Animals↗

GNARE: automated system for high-throughput genome analysis with grid computational backend.

Recent progress in genomics and experimental biology has brought exponential growth of the biological information available for computational analysis in public genomics databases. However, applying the potentially enormous scientific value of this information to the understanding of biological systems requires computing and data storage technology of an unprecedented scale. The Grid, with its aggregated and distributed computational and storage infrastructure, offers an ideal platform for high-throughput bioinformatics analysis. To leverage this we have developed the Genome Analysis Research Environment (GNARE)--a scalable computational system for the high-throughput analysis of genomes, which provides an integrated database and computational backend for data-driven bioinformatics applications. GNARE efficiently automates the major steps of genome analysis including acquisition of data from multiple genomic databases; data analysis by a diverse set of bioinformatics tools; and storage of results and annotations. High-throughput computations in GNARE are performed using distributed heterogeneous Grid computing resources such as Grid2003, TeraGrid, and the DOE Science Grid. Multi-step genome analysis workflows involving massive data processing, the use of application-specific tools and algorithms and updating of an integrated database to provide interactive web access to results are all expressed and controlled by a "virtual data" model which transparently maps computational workflows to distributed Grid resources. This paper describes how Grid technologies such as Globus, Condor, and the Gryphyn Virtual Data System were applied in the development of GNARE. It focuses on our approach to Grid resource allocation and to the use of GNARE as a computational framework for the development of bioinformatics applications.

Computational Biology↗

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus↗

TriosCompass: a snakemake workflow for integrated detection of SNVs, indels, STRs, and structural de novo variants in parent-child trios.

MOTIVATION: The accurate and sensitive identification of de novo variants, which are unique to an individual and not found in the parents' germlines, is critical for understanding the genetic basis of rare diseases, developmental disorders, and evolutionary processes. Existing de novo variant detection pipelines often lack the flexibility to handle multiple variant types, struggle with speed and reproducibility across computational environments, demand extensive manual configuration, or require bioinformatics expertise for downstream curation and analysis, limiting their scalability and usability for large genomic studies. Accordingly, there is a pressing need to better address these challenges. RESULTS: We introduce TriosCompass, an open-source Snakemake workflow that addresses these challenges by providing a modular, accelerated, and environmentally-configurable end-to-end solution for comprehensive de novo variant discovery. It integrates state-of-the-art tools into a reproducible framework, empowering researchers to discover novel genetic insights with greater efficiency and reliability. AVAILABILITY: TriosCompass is implemented as a Snakemake workflow and is freely available at https://github.com/NCI-CGR/TriosCompass_v2 or on Zenodo (10.5281/zenodo.17981062). SUPPLEMENTARY INFORMATION: Supplementary data is available on GitHub at https://github.com/NCI-CGR/TriosCompass_v2/tree/manuscript/report_dashboards. Supplementary methods on DeepTrio benchmark runs can be viewed at: https://github.com/NCI-CGR/TriosCompass_v2/blob/manuscript/TriosCompass_Supp_Methods_deeptrio_benchmark.md.

Software↗

Bioconductor: an open source framework for bioinformatics and computational biology.

This chapter describes the Bioconductor project and details of its open source facilities for analysis of microarray and other high-throughput biological experiments. Particular attention is paid to concepts of container and workflow design, connections of biological metadata to statistical analysis products, support for statistical quality assessment, and calibration of inference uncertainty measures when tens of thousands of simultaneous statistical tests are performed.

Animals↗

Workflow for Long-Read Amplicon Sequencing of Chikungunya Virus Using Oxford Nanopore Technology.

This protocol provides a comprehensive, step-by-step workflow for whole-genome sequencing of Chikungunya virus (CHIKV) using an amplicon-based strategy optimized for Oxford Nanopore Technologies (ONT) platforms. The procedure includes detailed instructions for sample handling, viral RNA extraction, quality control, cDNA synthesis, multiplex PCR amplification, library preparation, sequencing, and primary bioinformatic processing. The protocol is designed to maximize reproducibility across laboratories and is suitable for genomic surveillance applications, including outbreak investigation and molecular epidemiology, even when working with low-to-moderate viral loads.

Chikungunya virus↗

Systematic performance evaluation and application validation of an end-to-end NGS workstation.

Next-generation sequencing (NGS) library preparation is a core component of precision genomics, but it is commonly constrained by inefficiency, variability, and low throughput of manual protocols. To address these limitations, we developed and systematically evaluated a fully automated NGS workstations and further validated its performance across representative application scenarios. The automated system reduced total processing time from 8 to 10 to 4–6 h. At the same time, it maintained similar performance in pre-library metric, including DNA yield and fragment size, as well as post-capture sequencing metrics (Q30 > 90%, mapping rates > 95%, on-target rates 85–90%). The duplication rate was reduced to 5–8%, compared with 10–15% for manual methods, indicating increased library complexity. Bioinformatic evaluation of inter-species read mapping showed minimal cross-contamination, with a maximum contamination ratio of 0.0003%, indicating effective sample isolation in the automated workflow. High concordance in variant detection was observed between automated and manual workflows. Overall, this automated workstation provides a standardized and reproducible workflow that supports scalable precision genomics applications.

High-Throughput Nucleotide Sequencing↗

Modelling biological processes using workflow and Petri Net models.

MOTIVATION: Biological processes can be considered at many levels of detail, ranging from atomic mechanism to general processes such as cell division, cell adhesion or cell invasion. The experimental study of protein function and gene regulation typically provides information at many levels. The representation of hierarchical process knowledge in biology is therefore a major challenge for bioinformatics. To represent high-level processes in the context of their component functions, we have developed a graphical knowledge model for biological processes that supports methods for qualitative reasoning. RESULTS: We assessed eleven diverse models that were developed in the fields of software engineering, business, and biology, to evaluate their suitability for representing and simulating biological processes. Based on this assessment, we combined the best aspects of two models: Workflow/Petri Net and a biological concept model. The Workflow model can represent nesting and ordering of processes, the structural components that participate in the processes, and the roles that they play. It also maps to Petri Nets, which allow verification of formal properties and qualitative simulation. The biological concept model, TAMBIS, provides a framework for describing biological entities that can be mapped to the workflow model. We tested our model by representing malaria parasites invading host erythrocytes, and composed queries, in five general classes, to discover relationships among processes and structural components. We used reachability analysis to answer queries about the dynamic aspects of the model. AVAILABILITY: The model is available at http://smi.stanford.edu/projects/helix/pubs/process-model/.

Animals↗

A web services choreography scenario for interoperating bioinformatics applications.

BACKGROUND: Very often genome-wide data analysis requires the interoperation of multiple databases and analytic tools. A large number of genome databases and bioinformatics applications are available through the web, but it is difficult to automate interoperation because: 1) the platforms on which the applications run are heterogeneous, 2) their web interface is not machine-friendly, 3) they use a non-standard format for data input and output, 4) they do not exploit standards to define application interface and message exchange, and 5) existing protocols for remote messaging are often not firewall-friendly. To overcome these issues, web services have emerged as a standard XML-based model for message exchange between heterogeneous applications. Web services engines have been developed to manage the configuration and execution of a web services workflow. RESULTS: To demonstrate the benefit of using web services over traditional web interfaces, we compare the two implementations of HAPI, a gene expression analysis utility developed by the University of California San Diego (UCSD) that allows visual characterization of groups or clusters of genes based on the biomedical literature. This utility takes a set of microarray spot IDs as input and outputs a hierarchy of MeSH Keywords that correlates to the input and is grouped by Medical Subject Heading (MeSH) category. While the HTML output is easy for humans to visualize, it is difficult for computer applications to interpret semantically. To facilitate the capability of machine processing, we have created a workflow of three web services that replicates the HAPI functionality. These web services use document-style messages, which means that messages are encoded in an XML-based format. We compared three approaches to the implementation of an XML-based workflow: a hard coded Java application, Collaxa BPEL Server and Taverna Workbench. The Java program functions as a web services engine and interoperates with these web services using a web services choreography language (BPEL4WS). CONCLUSION: While it is relatively straightforward to implement and publish web services, the use of web services choreography engines is still in its infancy. However, industry-wide support and push for web services standards is quickly increasing the chance of success in using web services to unify heterogeneous bioinformatics applications. Due to the immaturity of currently available web services engines, it is still most practical to implement a simple, ad-hoc XML-based workflow by hard coding the workflow as a Java application. For advanced web service users the Collaxa BPEL engine facilitates a configuration and management environment that can fully handle XML-based workflow.

Computational Biology↗

Machine learning approaches for cancer prognosis and diagnosis via non-coding RNA: a comprehensive review.

Non-coding RNAs (ncRNAs), once considered genomic dark matter, are now established as key regulators of gene expression with widespread roles in cellular homeostasis and disease. In cancer, ncRNA expression is frequently and systematically dysregulated, and many of these molecules circulate in stable, protected form within biofluids, offering a compelling basis for non-invasive or minimally invasive diagnostic strategies. However, their clinical translation remains substantially hindered to date due to biological complexity, technical noise, and high dimensionality inherent to ncRNA expression datasets. In this context, machine learning (ML) has emerged as a powerful analytical tool to address these challenges, enabling the identification of subtle, reproducible ncRNA signatures predictive of diverse malignancies. This review critically evaluates ML-driven frameworks for cancer diagnosis and prognosis across four ncRNA subclasses, namely miRNAs, lncRNAs, circRNAs, and piRNAs, while also acknowledging the biophysical and thermodynamic models that reinforce ncRNA bioinformatics. Despite substantial methodological progress in ML-based cancer diagnosis and prognosis, key challenges persist, including tumor biological heterogeneity, limited multicenter validation, and the lack of widely adopted standardized protocols for preprocessing, normalization, and reporting workflows. Furthermore, many current ML models lack interpretability in biological or clinical context, constraining their translational utility. By synthesizing recent advances and identifying unresolved barriers, this review charts a roadmap for developing a robust, clinically actionable ncRNA biomarker platform for cancer detection. With global cancer incidence projected to exceed 35 million annual cases by 2050, validated ncRNA-ML-driven frameworks hold potential to revolutionize early-stage detection and personalized therapeutic strategies, thereby reducing the escalating socio-economic burden of cancer worldwide.

Humans↗

Evolution of web services in bioinformatics.

Bioinformaticians have developed large collections of tools to make sense of the rapidly growing pool of molecular biological data. Biological systems tend to be complex and in order to understand them, it is often necessary to link many data sets and use more than one tool. Therefore, bioinformaticians have experimented with several strategies to try to integrate data sets and tools. Owing to the lack of standards for data sets and the interfaces of the tools this is not a trivial task. Over the past few years building services with web-based interfaces has become a popular way of sharing the data and tools that have resulted from many bioinformatics projects. This paper discusses the interoperability problem and how web services are being used to try to solve it, resulting in the evolution of tools with web interfaces from HTML/web form-based tools not suited for automatic workflow generation to a dynamic network of XML-based web services that can easily be used to create pipelines.

Computational Biology↗

Fedflow: cloud orchestration for federated learning with the FeatureCloud platform.

MOTIVATION: Federated learning (FL) enables collaborative model training on geographically distributed genomic and clinical datasets while complying with data privacy laws and regulatory constraints. FeatureCloud is an existing platform for FL that provides an accessible web-based interface and a large repository of implemented methods. However, due to its graphical interface, FeatureCloud requires manual interaction of all participants, limiting automation, iteration, and reproducibility. RESULTS: We introduce fedflow, a Python-based command-line tool for headless orchestration of FL tasks with FeatureCloud. This tool uses distributed computing resources such as virtual machines or cloud instances to automate such workflows. This allows for scalable federated computing either in local simulations or deployed in a trusted environment. Further, we demonstrate how fedflow can be used to integrate FeatureCloud in reproducible Snakemake workflows. For this, we reanalyse a metagenomic dataset with two federated algorithms and compare the results to the centralized approach with pooled data. Overall, fedflow enables automation of multi-client FL tasks, facilitates embedding of FeatureCloud in standard bioinformatics pipelines and thereby helps increase reproducibility. AVAILABILITY: Fedflow is open-source and available at https://github.com/W-L/fedflow.

Journal Article↗

A scalable HPC framework for bioinformatics in resource-limited settings: design principles, implementation, and sustainability from the UVRI experience.

MOTIVATION: Building and sustaining High-Performance Computing (HPC) infrastructure for bioinformatics research in resource-limited settings presents significant technical, financial and operational challenges. Institutions in low-and middle-income regions often face constraints such as limited technical expertise, unstable infrastructure and restricted funding which can hinder the deployment of large-scale computational platforms necessary for modern genomics and bioinformatics analyses. RESULTS: We present a scalable and modular HPC framework developed at the Uganda Virus Research Institute (UVRI) to support large-scale genomics and other omics data analyses in resource-limited settings. The framework integrates open-source HPC management tools, infrastructure automation, and reproducible configuration management to enable reliable deployment and maintenance. Optimized storage and networking configurations combined with a phased capacity-building strategy support high-throughput genomic workflows while strengthening local technical expertise. From our implementation experience, we derive ten practical design and operational rules that provide a transferable methodology for establishing and sustaining in-house HPC infrastructure. These rules emphasize strategic investment in human capacity, structured planning, leveraging collaborations, adoption of open-source technologies and service management practices to improve operational resilience and long-term sustainability. AVAILABILITY: The design principles, automation strategies and implementation guidelines described in this work are applicable to institutions seeking to establish sustainable HPC resources for bioinformatics research in resource-constrained environments.

Computational Biology↗

MACS3: A Peak-calling Platform for Bulk and Single-cell Regulatory Genomics.

Since the original publication of Model-based Analysis for ChIP-Seq (MACS), the software has been widely used to identify enriched genomic regions in ChIP-seq, ATAC-seq, CUT&RUN, DNase-seq, and related regulatory genomics assays. Over the years, MACS has evolved substantially, with MACS version 3 (MACS3) now serving as the actively maintained implementation. MACS3 preserves the core MACS framework for fragment pileup, dynamic local background noise, statistical enrichment testing, and peak refinement, while adding functionality needed for contemporary bulk and single-cell workflows. It supports conventional bulk peak calling, paired-end and fragment-based file formats, modular signal processing, direct analysis of single-cell ATAC-seq fragment files, barcode-restricted pseudobulk and cluster-level peak calling, specialized ATAC-seq and variant-calling modules, as well as command-line and programmatic interfaces. MACS3 is distributed through standard software channels and supported by continuous testing across operating systems, Python versions, and CPU architectures. Here we describe the architecture, current capabilities, and recommended use of MACS3, providing an updated reference for applying the MACS framework in contemporary bulk and single-cell regulatory genomics workflows. MACS3 is open-source software available at https://github.com/macs3-project/MACS.

Bioinformatics software↗

Bioinformatics for medical diagnostics: assessment of microarray data in the context of clinical databases.

MOTIVATION: To identify genes suitable for medical diagnostics microarray data is assessed in the context of clinical databases, which store complex information about the patient phenotype. The wealth of data and lacking standards make it difficult to analyse this kind of data. RESULTS: We present a workflow for exploratory analysis of microarray data together with clinical data consisting of four steps: definition of clinically meaningful research questions in a masterfile, generation of analysis files, selection and characterization of differentially expressed genes, and estimation of classification accuracy. We applied this workflow to large data sets from the field of cardiology and oncology (n~500 patients). Systematic data management of microarray data and clinical data helps to make results more transparent and comparable.

Cardiology↗

ChromBERT-tools: a versatile toolkit for context-specific regulatory representations of transcription regulators across different cell types.

SUMMARY: Representations that encode the genome-wide regulatory behavior of transcription regulators provide a foundation for flexible transcription modeling and in silico regulatory analysis. Existing regulator representations are commonly derived from gene co-expression, motif annotations, or static protein features, which capture useful but limited aspects of regulator identity but do not directly model how regulators participate in region-specific regulatory programs across the genome. ChromBERT addresses this gap by learning context-aware regulatory representations from large-scale ChIP-seq data. However, routine bioinformatics applications require lightweight, accessible, and modular tools for generating, adapting, and interpreting these representations in user-defined biological contexts. Here, we present ChromBERT-tools, a user-oriented toolkit built upon ChromBERT that converts its regulatory representation framework into practical workflows for customizable analysis across cellular contexts. ChromBERT-tools provides command-line interfaces and Python APIs organized into three functional layers: representation generation, predictive modeling, and regulatory interpretation. The representation generation layer produces representations of genomic regions and transcription regulators. The predictive modeling layer fine-tunes ChromBERT for genome-wide regulatory activity prediction through classification or regression tasks, with optimized implementation to reduce running time and computational resource requirements. The regulatory interpretation layer supports inference of the context-specific roles of cis-regulatory elements and transcription regulators. These modules can be used independently or integrated into end-to-end workflows, enabling flexible analyses across diverse datasets. ChromBERT-tools lowers the barrier to applying context-specific regulatory representations in routine genomic analyses. AVAILABILITY AND IMPLEMENTATION: ChromBERT-tools is freely available at https://github.com/TongjiZhanglab/ChromBERT-tools, with documentation at https://chrombert-tools.readthedocs.io/en/latest/. A frozen archival snapshot is available on Zenodo under DOI: 10.5281/zenodo.20094206.

Software↗

Designing and executing scientific workflows with a programmable integrator.

MOTIVATION: As in many other fields of science, computational methods in molecular biology need to intersperse information access and algorithm execution in a computational workflow. Users often find difficulties when transferring data between data sources and applications. In most cases there is no standard solution for workflow design and execution and tailored scripting mechanisms are implemented in a case by case basis. RESULTS: In this paper, we present a general purpose 'programmable integrator' that can access information from a variety of sources in a coordinated manner. Its usefulness in complex bioinformatics applications is claimed and supported by some application examples. AVAILABILITY: Tools are freely available to non-profit educations and research institutions. Usage by commercial organizations requires a license agreement. Software requirements: Java v1.3 (http://java.sun.com), Xerces XML Parser (http://xml.apache.org/xerces-j) and Kweelt implementation of XQuery (http://kweelt.sourceforge.net/).

Algorithms↗