Search PubMedSearch

SEARCH · Search PubMed

Results for “automated workflows”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Trustworthy Agentic AI in Bioinformatics: From Workflow Automation to Traceable and Validated Biological Inference.

Agentic artificial intelligence is extending bioinformatics beyond conversational assistance by enabling systems to select tools, execute code, revise analytical plans, and interpret biological data. These capabilities may accelerate research, but they also redistribute decisions that determine whether biological conclusions are valid. We conducted a targeted, structured PubMed search in July 2026 and identified 11 peer-reviewed agentic bioinformatics systems for descriptive review based on predefined eligibility criteria for analytical decision-making, tool or code execution, iterative evaluation, or coordinated agent activity. The evidence base covered single-cell transcriptomics, microbial genomics, cancer genomics, and omics applications, together with methodological literature on reproducibility and biological validation. We examined how current systems report delegated authority, provenance, validation, evidence, abstention, and human oversight. Existing platforms implement safeguards such as sandboxed execution, restricted commands, interaction logs, evidence identifiers, automated checks, critic agents, quality scores, and expert assessment. However, published reports rarely provide a connected account linking the original biological question to samples, reference resources, analytical decisions, computational actions, statistical results, supporting evidence, validation outcomes, and final claims. We distinguish inherited bioinformatics errors, errors amplified through autonomous action, and emergent failures arising from memory, retrieval, tool interaction, or agent coordination. We further propose a multidimensional decision-rights profile, consequence-sensitive validation gates, and a claim-to-evidence provenance architecture organized through the Traceable History of Research Evidence, Agent Actions, and Decisions in Bioinformatics (THREAD-Bio) framework. Illustrative cases show that technically successful execution may still support misleading inference. Trustworthy agentic bioinformatics therefore requires claims to remain reconstructible, challengeable, validated, and proportionate to the evidence.

accountable autonomy

Systematic performance evaluation and application validation of an end-to-end NGS workstation.

Next-generation sequencing (NGS) library preparation is a core component of precision genomics, but it is commonly constrained by inefficiency, variability, and low throughput of manual protocols. To address these limitations, we developed and systematically evaluated a fully automated NGS workstations and further validated its performance across representative application scenarios. The automated system reduced total processing time from 8 to 10 to 4–6 h. At the same time, it maintained similar performance in pre-library metric, including DNA yield and fragment size, as well as post-capture sequencing metrics (Q30 > 90%, mapping rates > 95%, on-target rates 85–90%). The duplication rate was reduced to 5–8%, compared with 10–15% for manual methods, indicating increased library complexity. Bioinformatic evaluation of inter-species read mapping showed minimal cross-contamination, with a maximum contamination ratio of 0.0003%, indicating effective sample isolation in the automated workflow. High concordance in variant detection was observed between automated and manual workflows. Overall, this automated workstation provides a standardized and reproducible workflow that supports scalable precision genomics applications.

High-Throughput Nucleotide Sequencing

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus

A liquid handling platform for standardized quantification of cell-free enzymatic activity encoded by antimicrobial resistance genes.

Antimicrobial resistance (AMR) is a growing global threat to human health, and rapid methods for characterizing emerging antimicrobial resistance genes (ARGs) are needed. Here, we develop a semi-automated workflow using cell-free gene expression systems to measure the activity of two ARGs encoded on plasmid DNA that produce rifampicin-inactivating and gentamicin-inactivating enzymes. We validated the use of a small benchtop Myra liquid handling system compared to manual pipetting, with no statistical differences observed. After optimizing the pre-incubation time of ARGs and dispensing protocol, expression of aac(3)-IIa increased the half-maximal inhibition concentration (IC50) of gentamicin by over 150-fold, whilst arr-3 increased the IC50 of rifampicin by ~20-fold compared to controls. This methodology for rapid, semi-automated ARG characterization offers a strategy to combat AMR by assessing novel ARGs identified through genomic surveillance or profiling activity of new or derivative antibiotics.

aminoglycoside resistance

DeepPlaque: a scalable multimodal platform for Aβ pathology and cell analysis in Alzheimer's disease.

Histological analysis is essential for understanding disease pathology and the microenvironment, particularly in Alzheimer's disease (AD), characterized by beta-amyloid (Aβ) plaques that exist as diffuse, fibrillar, and core species, with distinct toxicity levels. However, accurate classification of Aβ plaque types in postmortem brain tissues and profiling of surrounding cells present significant challenges. To address these challenges, we developed "DeepPlaque", an integrated system featuring "PlaqueNet", a deep learning model for automated classification of Aβ plaque species from diverse imaging platforms. DeepPlaque includes automated workflows for cellular phenotyping and proteomic profiling through targeted laser microdissection. PlaqueNet achieves expert-level accuracy (AUC > 90%) in classifying the 3 major Aβ plaque species, supporting consistent and large-scale annotation. By integrating spatial cellular phenotyping with laser microdissection, DeepPlaque enables high-throughput proteomic analysis of Aβ plaque niches, revealing that microglia are more abundant around core and fibrillar Aβ plaques, with increased expression of apolipoprotein E and amyloid precursor protein in core Aβ plaques. This customizable platform enhances the molecular and cellular characterization of Aβ plaque-associated environments, providing critical insights into AD pathology.

Alzheimer Disease

Mass spectrometry-based ligand binding assays in biomedical research.

INTRODUCTION: Ligand binding assays combining immunoaffinity enrichment steps with mass spectrometry (MS) readout have gained attention as a highly specific and sensitive tool for protein quantification. These techniques typically combine enzymatic fragmentation of the sample or enriched protein with capture on the protein or peptide-level for quantification. Antibodies ensure specific target recognition, while MS offers quantitative accuracy with isotopically labeled internal standards. This dual approach supports a broad dynamic range, enabling protein measurements from picomolar to nanomolar levels. These methods have diverse applications, from quantifying signaling proteins in basic research to biomarker monitoring in clinical trials and analyzing the pharmacokinetics of therapeutic proteins. AREAS COVERED: This review delves into the diverse workflows of immunoaffinity-MS, shedding light on the innovative strategies employed, their practical applications, efficacy, and inherent limitations in the realm of protein quantification. EXPERT OPINION: Immunoaffinity-MS has transformed protein analysis, but widespread adoption is hindered by complex workflows, high instrument costs, and limited capture molecule availability. Efforts to enhance automation, standardize workflows, and advance technological innovation aim to overcome these barriers. Improvements in mass spectrometer sensitivity, advances in recombinant capture technologies, and support from public initiatives are poised to further improve the reliability and accessibility of this method.

Mass Spectrometry

Integrative evidence-knowledge marker selection enhances LLM-based cell type annotation in single-cell RNA-seq analysis.

BACKGROUND: Cell type annotation is essential for gaining biological insight from single-cell RNA sequencing data, yet manual labeling remains time-consuming and difficult to reproduce. Various computational approaches have been developed to automate this process, and recent studies suggest that large language models can infer cell types with promising accuracy in single-cell analysis. However, most workflows still rely on cluster-specific markers derived from gene expression alone or manual curation. As a result, marker selection can be sensitive to statistical criteria and dataset-dependent bias, which may lead to the selection of less informative genes or missing important markers, while providing limited biological context. RESULTS: To address this limitation, we introduce CELLIA, an LLM-based workflow for automated and robust cell type annotation. CELLIA employs an integrative evidence-knowledge marker selection strategy that combines statistical differential expression criteria with curated tissue-specific marker resources to identify informative marker genes. In benchmarking analyses of 102 cell types, this approach improved agreement with manual annotations. In addition, CELLIA achieved higher agreement in subtype-level analyses of closely related immune populations and was further evaluated in a non-immune stromal subtype setting, covering 25 cell types in total. CONCLUSION: By integrating evidence-knowledge from gene expression with curated biological prior knowledge, CELLIA provides a more stable marker selection and improves the reliability of LLM-cell type annotation.

Cell type annotation

Fedflow: cloud orchestration for federated learning with the FeatureCloud platform.

MOTIVATION: Federated learning (FL) enables collaborative model training on geographically distributed genomic and clinical datasets while complying with data privacy laws and regulatory constraints. FeatureCloud is an existing platform for FL that provides an accessible web-based interface and a large repository of implemented methods. However, due to its graphical interface, FeatureCloud requires manual interaction of all participants, limiting automation, iteration, and reproducibility. RESULTS: We introduce fedflow, a Python-based command-line tool for headless orchestration of FL tasks with FeatureCloud. This tool uses distributed computing resources such as virtual machines or cloud instances to automate such workflows. This allows for scalable federated computing either in local simulations or deployed in a trusted environment. Further, we demonstrate how fedflow can be used to integrate FeatureCloud in reproducible Snakemake workflows. For this, we reanalyse a metagenomic dataset with two federated algorithms and compare the results to the centralized approach with pooled data. Overall, fedflow enables automation of multi-client FL tasks, facilitates embedding of FeatureCloud in standard bioinformatics pipelines and thereby helps increase reproducibility. AVAILABILITY: Fedflow is open-source and available at https://github.com/W-L/fedflow.

Journal Article

Colora: a Snakemake workflow for complete chromosome-scale de novo genome assembly.

MOTIVATION: De novo assembly creates reference genomes that underpin many modern biodiversity and conservation studies. Large numbers of new genomes are being assembled by labs around the world. To avoid duplication of efforts and variable data quality, we desire a best-practice assembly process, implemented as an automated portable workflow. RESULTS: Here, we present Colora, a Snakemake workflow that produces chromosome-scale de novo primary or phased genome assemblies complete with organelles using Pacific Biosciences HiFi, Hi-C, and optionally Oxford Nanopore Technologies reads as input. Colora is a user-friendly, versatile, and reproducible pipeline that is ready to use by researchers looking for an automated way to obtain high-quality de novo genome assemblies. AVAILABILITY AND IMPLEMENTATION: The source code of Colora is available on GitHub (https://github.com/LiaOb21/colora) and has been deposited in Zenodo under DOI https://doi.org/10.5281/zenodo.13321576. Colora is also available at the Snakemake Workflow Catalog (https://snakemake.github.io/snakemake-workflow-catalog/? usage=LiaOb21%2Fcolora).

Software

Tractor workflow: a scalable Nextflow framework for local ancestry-aware genome-wide association studies.

MOTIVATION: The routine exclusion of admixed individuals from traditional genome-wide association studies (GWAS) due to concerns about spurious associations has limited multi-ancestry genetic discovery. Tractor addresses this issue by incorporating local ancestry into association testing, enabling the identification of ancestry-enriched signals and generating ancestry-specific summary statistics. However, adoption has been constrained by the complexity of prerequisite steps, including phasing and local ancestry inference, which require substantial bioinformatics expertise and introduce key analytical decision points. RESULTS: We developed a scalable, automated Nextflow workflow that integrates phasing, local ancestry inference, and Tractor association testing into a reproducible end-to-end pipeline. To demonstrate its utility, we applied the workflow to 32 blood biomarkers in 6245 two-way African-European admixed individuals from the UK Biobank. This pipeline performed efficiently at scale, replicating known associations and uncovering key ancestry-specific loci. These associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously masked genetic signals. AVAILABILITY AND IMPLEMENTATION: The workflow is modular, customizable, and compatible with commonly used phasing and local ancestry tools, minimizing manual intervention while preserving analytical flexibility. By lowering technical barriers to implementation, this framework facilitates broader adoption of local ancestry-aware GWAS, paving the way for expanded genetic discovery.

Humans

Automated chromatin profiling with spa-ChIP-seq uncovers the impacts of condition variations.

Chromatin immunoprecipitation followed by sequencing (ChIP-seq) is widely used to study the genomic localization of DNA-associated proteins. However, conventional protocols include multiple manual steps that can introduce inconsistency and limit scalability, thereby restricting the inclusion of appropriate replicates and controls. Although the introduction of liquid handling platforms has improved reproducibility, most existing efforts have automated only a subset of the workflow, and extending automation to efficiently map non-histone proteins, such as chromatin regulators, remains challenging. Here, we present a fully automated implementation of our previously developed single-pot ChIP-seq protocol (Texari et al. 2021), named spa-ChIP-seq, which enables scalable processing of 8 to 96 ChIP-seq samples from crosslinked cells to sequencing-ready library in approximately three days with an estimated cost of $70 per sample. Benchmarking spa-ChIP-seq against manual ChIP-seq performed in parallel demonstrates comparable signal-to-noise ratio between the two workflows. Using spa-ChIP-seq, we systematically evaluate multiple parameters including shearing and crosslinking conditions, buffer compositions, and the ratio of antibody to cell-number. We find, for the first time to our knowledge, that weaker genomic localization signals are sensitive to changing the antibody to cell-number ratio, whereas the stronger signals remain unaffected. This finding underscores the importance of maintaining consistent antibody-to-cell-number ratio for comparative studies, such as treatment responses or chromatin-QTL mapping. The spa-ChIP-seq protocol is publicly available, including deck setups, operational parameters, and scripts. We envision that this robust, cost-efficient protocol will facilitate high-throughput, reproducible ChIP-seq analyses, supporting large-scale studies of antibody validation, compound screening, population genomics, and diagnostic frameworks.

Journal Article

PopGLen-a Snakemake pipeline for performing population genomic analyses using genotype likelihood-based methods.

SUMMARY: PopGLen is a Snakemake workflow for performing population genomic analyses within a genotype-likelihood framework, integrating steps for raw sequence processing of both historical and modern DNA, quality control, multiple filtering schemes, and population genomic analysis. Currently, the population genomic analyses included allow for estimating linkage disequilibrium, kinship, genetic diversity, genetic differentiation, population structure, inbreeding, and allele frequencies. Through Snakemake, it is highly scalable, and all steps of the workflow are automated, with results compiled into an HTML report. PopGLen provides an efficient, customizable, and reproducible option for analyzing population genomic datasets across a wide variety of organisms. AVAILABILITY AND IMPLEMENTATION: PopGLen is available under GPLv3 with code, documentation, and a tutorial at https://github.com/zjnolen/PopGLen. An example HTML report using the tutorial dataset is included in the Supplementary Material.

Software

Evaluation of sequencing reads at scale using rdeval.

MOTIVATION: Large sequencing datasets are being produced and deposited into public archives at unprecedented rates. The availability of tools that can reliably and efficiently generate and store sequencing read summary statistics has become critical. RESULTS: As part of the effort by the Vertebrate Genomes Project (VGP) to generate high-quality reference genomes at scale, we sought to address the community's need for efficient sequence data evaluation by developing rdeval, a standalone tool to quickly compute and interactively display sequencing read metrics. Rdeval can either run on the fly or store key sequence data metrics in tiny read 'snapshot' files. Statistics can then be efficiently recalled from snapshots for additional processing. Rdeval can convert fa*[.gz] files to and from other popular formats including BAM and CRAM for better compression. Overall, while CRAM achieves the best compression, the gain compared to BAM is marginal, and BAM achieves the best compromise between data compression and access speed. Rdeval also generates a detailed visual report with multiple data analytics that can be exported in various formats. We showcase rdeval's functionalities using long-read data from different sequencing platforms and species, including human. For PacBio long-read sequencing, our analysis shows dramatic improvements in both read length and quality over time, as well as the benefit of increased coverage for genome assembly, though the magnitude varies by taxa. AVAILABILITY AND IMPLEMENTATION: Rdeval is implemented in C++ for data processing and in R for data visualization. Precompiled releases (Linux, MacOS, Windows) and commented source code for rdeval are available under MIT license at https://github.com/vgl-hub/rdeval. Documentation is available on ReadTheDocs (https://rdeval-documentation.readthedocs.io). Rdeval is also available in Bioconda and in Galaxy (https://usegalaxy.org). An automated test workflow ensures the consistency of software updates.

Software

Scalable, open-access and multidisciplinary data integration pipeline for climate-sensitive diseases.

Climate-sensitive infectious diseases pose an important challenge for human, animal and environmental health and it has been estimated that over half of known human pathogenic diseases can be aggravated by climate change. While climatic and weather conditions are important drivers of transmission of vector-borne diseases, socio-economic, behavioural, and land-use factors as well as the interactions among them impact transmission dynamics. Analysis of drivers of climate-sensitive diseases require rapid integration of interdisciplinary data to be jointly analysed with epidemiological (including genomic and clinical) data. Current tools for the integration of multiple data sources are often limited to one data type or rely on proprietary data and software. To address this gap, we develop a scalable and open-access pipeline for the integration of multiple spatio-temporal datasets that requires only the declaration of the country and temporal range and resolution of the study. The tool is locally deployable and can easily be integrated into existing climate-disease-modelling applications. We demonstrate the utility of the tool for dengue modelling in Vietnam where epidemiological data are legally required to remain local. We include a pipeline for bias correction of climate data to enhance their quality for downstream modelling tasks. The Dengue Advanced Readiness Tools-Pipeline empowers users by simplifying complex download, correction, and aggregation steps, fostering data-driven discovery of relationships between infectious diseases and their drivers in space and time, and enhancing reproducibility in research. Additional modules and datasets can be added to the existing ones to make the pipeline extendable to use cases other than the ones presented here.

automated workflows

Hard to Halt: Automation Bias in Agent-Driven Sequencing Prior Authorization Workflows.

PURPOSE: Prior authorization (PA) for exome or genome sequencing is a time-consuming process that impedes timely rare disease diagnosis. Large language model-based browser agents offer potential for automating these workflows, but their clinical reliability remain uncharacterized. METHODS: We developed a sandbox compromising a simulated ES/GS PA submission payer portal and a synthetic EHR containing 836 patient records spanning compliant profiles and deficient profiles with different types of issues. Gemini 3 Pro, Gemini 3 Flash, and Claude Opus 4.5 were evaluated on task completion rate, form completion accuracy, and appropriate withholding for deficient profiles. RESULTS: Larger models achieved much higher task completion rates (Gemini 3 Pro 95.45%, Claude Opus 4.5 93.67%) compared to Gemini 3 Flash (56.05%), but nearly universally failed to withhold submission for deficient profiles whereas Gemini 3 Flash ironically demonstrated superior withholding performance (17.33%). In a non-agentic setting, Gemini 3 Pro correctly identified 91% of the issues in deficient profiles, indicating that withholding failure is attributable to the browser interaction rather than the model's reasoning limitations. CONCLUSION: Current LLM-based browser agents exhibit a systematic bias towards form submission that poses risks in PA workflows. A modular, multi-agent architecture with human supervision is necessary for a safe clinical deployment.

Journal Article

Hide and seek: de novo identification in sugar beet reveals impact of non-autonomous LTR retrotransposons.

Plant genomes are filled with retrotransposons and their derivatives, constantly undergoing sequence diversification and structural rearrangement. Among them, short, non-autonomous retrotransposons lack full coding capacity and often form subfamilies. As a result, non-autonomous retrotransposons are incompletely identified in most to all genome assemblies.Here, we capitalize on our comprehensive understanding of the transposable element (TE) landscape in sugar beet (Beta vulgaris) to assess the extent of the blind spot for non-autonomous long terminal repeat (LTR) retrotransposons. This use case serves to answer if all of these sequences are derivatives of easier-to-identify full-length elements or if there is more variability that is currently overlooked.For this we applied a semi-automated structural discovery workflow followed by in-depth manual verification to characterize non-autonomous LTR retrotransposons in sugar beet. We retrieve more than 100 non-autonomous LTR retrotransposon families that lack complete autonomous coding capacity, including canonical terminal-repeat retrotransposons in miniature (TRIMs), elongated non-coding derivatives and families retaining fragmented coding remnants. The identified families span a broad range, including elements exceeding 15,000 bp in length and display evidence for reshuffling and modular evolution. Only a subset of families could be confidently linked to autonomous retrotransposons, showing sequence diversification within the non-autonomous LTR retrotransposon fraction beyond the autonomous genomic templates.We highlight that a large fraction of non-autonomous LTR retrotransposons is incompletely recovered with the current TE identification workflows, even if the output is well-curated and condensed into TE libraries and suggest procedures to remedy this gap. This study gives a genome-wide view into the non-autonomous LTR retrotransposon landscape of a single plant genome and highlights the importance of structure-based approaches for their identification and classification.

LTR retrotransposons

Accelerating natural product discovery, characterization and engineering by biofoundries.

Covering: From early developments to the presentNatural product (NP) discovery is increasingly constrained by low-throughput screening, repeated rediscovery, and challenges in scaling genome mining-guided validation workflows. This highlight examines how automated biofoundries are accelerating NP discovery, characterization, and engineering through integrated design-build-test-learn (DBTL) cycles. We discuss recent advances in phenotype-first and genome-first discovery strategies enabled by robotics, high-throughput pathway reconstitution, and automated screening platforms. We further highlight emerging technologies, including cell-free biosynthesis, automated culturomics, programmable chassis engineering, and AI-assisted workflow orchestration, that may enable increasingly autonomous biofoundries for scalable exploration of NP chemical space and therapeutic discovery.

Journal Article

A python based automated computational framework to classify and comparative genomics analysis of the global diversity of chili leaf curl virus (ChiLCV) strains to understand virus host interactions.

Chili leaf curl virus (ChiLCV) is a Begomovirus chillicapsici that is one of the most devastating viruses impacted on the production of chili in the world, especially in South Asia. In the present study, we combined high-throughput computational genomics with experimental analysis of global diversity. A workflow was created using automated Python scripts to download, curate and process ChiLCV genomes from public database. About 410 complete ChiLCV genomes download from public databases. Using a phylogenetic approach, these isolates were subdivided into 34 strains, belonging to 10 major clades, showing significant genetic diversity. Geographic analysis revealed that Pakistan (207 isolates) and India (148 isolates) were the main sources of ChiLCV diversity and the remainder of the isolates were from Oman, Bangladesh, Iran, Saudi Arabia and Sri Lanka. Recombination was observed as a major evolutionary force as more than twenty recombination events were detected. Analysis of cis-regulatory elements showed a complex structure of the viral promoter, including multiple binding sites for transcription factors, hormone-response elements, light-responsive elements, and stress-responsive elements, indicating a high number of interactions between viral regulatory elements and host signaling pathways. Pangenome analysis showed the presence of a highly dynamic open pangenome made up of strain-specific orthologous groups (species-specific orthogroups). Experimental inoculation of chili plants was also carried out to assess the biological effects of infection, along with phytochemical, FTIR, HPLC, and qPCR analyses.

Begomovirus