Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatics workflow”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Beyond the genome--SMi Conference 27-28 January 2003, London, UK.

Over a decade of astonishing developments, genomics and proteomics have promised a fundamentally new approach to drug discovery. Although there has been an undeniable increase in the range of potential targets available, this has not led to an increased output of the drug discovery pipeline into the clinic. With tighter markets and increasing competition, the major pharmaceutical companies are under intense pressure to achieve rapid, concrete delivery of those early promises, but there remain acute problems in the genes-to-drugs pipeline. This meeting showcased a range of novel approaches from proteomics and bioinformatics to address these problems. A common theme in the range of proteomics offerings was the prioritization of potential novel targets on the basis of their accessibility to drugs and their functional link to disease phenotypes. Informatics and in silico offerings also concentrated on fast, accurate, drug-focused workflows built on large integrative databases and novel data-mined algorithms.

Animals↗

Syndromic cholera diagnosis masks diverse causes of diarrhoeal disease in Burundi revealed by portable metagenomics.

BACKGROUND: Cholera outbreaks remain a major public-health challenge in sub-Saharan Africa, where diagnostic capacity is limited and clinical case definitions are non-specific and re ly heavily on syndromic diagnosis. Rapid identification of Vibrio cholerae is critical, yet cholera-suspected diarrhoea can have multiple infectious causes not captured by targeted diagnostics. METHODS: We evaluated a mobile, culture-independent metagenomic sequencing workflow for on-site detection of gastrointestinal pathogens directly from faecal samples in Burundi. The offline workflow combined long-read Oxford Nanopore Technologies (ONT) sequencing with rapid, laptop-based taxonomic and antimicrobial resistance (AMR) screening and was deployed across a health centre, a district hospital, and a refugee transit camp. The frontline and real-time results were verified using both conventional culturing and in-depth bioinformatic analyses. RESULTS: V. cholerae signals were only detected in a subset of suspected cholera cases, while many samples were dominated by alternative bacterial taxa, most frequently Escherichia coli. V. cholerae abundance correlated strongly with detection of the C holera T oxin P hage CTXφ, supporting differentiation between toxigenic signal and background exposure. AMR genes were detected across samples, providing early situational insight into resistance determinants among gastrointestinal bacteria. CONCLUSIONS: Mobile, offline metagenomic sequencing enables rapid frontline characterization of gastrointestinal disease, especially cholera-suspected, in resource-limited settings and complements existing diagnostics by improving etiological resolution and outbreak response.

Humans↗

Fantastic microbes and where to find them: evaluating learning-by-doing outcomes in a crowdfunded metagenomics workshop.

Metagenomics offers a powerful framework for authentic, interdisciplinary learning, yet it remains underrepresented in undergraduate education due to technical and infrastructural barriers. We hypothesized that a research-based, learning-by-doing metagenomics workshop supported by accessible bioinformatics tools could enhance students' perceived skills, self-efficacy, and conceptual understanding of metagenomic analysis. To test this hypothesis, we designed and evaluated a hybrid hands-on workshop in which undergraduate and postgraduate students analyzed real environmental shotgun metagenomic datasets generated from soil samples collected during a citizen science initiative. Using the graphical workflow platform KBase, participants completed an end-to-end metagenomic analysis, from quality control and assembly to genome reconstruction, taxonomic classification, functional annotation, and scientific presentation of results. Educational outcomes were assessed through validated retrospective pre-post questionnaires, self-efficacy scales, and an open-ended conceptual understanding task. Participants showed significant increases in perceived metagenomic skills and confidence in performing metagenomic analyses, while gains in perceived learning showed a positive trend. Conceptual understanding improved across educational levels, particularly among participants with limited prior experience. Together, these findings demonstrate that authentic, data-driven metagenomics activities can effectively lower barriers to computational biology and foster meaningful learning through hands-on research experiences.

Metagenomics↗

A bioinformatics pipeline for high-throughput microbial multilocus sequence typing (MLST) analyses.

Multilocus sequence typing (MLST) analysis for semi-routine applications is hindered by the downstream, manually intensive steps of processing the raw sequence data files. This report describes the development of an MLST pipeline that automates DNA sequence editing and analysis in order to significantly reduce the time required for processing data. Validation using a pneumococcal dataset revealed complete agreement between the results generated by manual and automated workflows. The MLST pipeline was developed for both double-strand and single-strand sequencing.

Bacteria↗

The MIGenAS integrated bioinformatics toolkit for web-based sequence analysis.

We describe a versatile and extensible integrated bioinformatics toolkit for the analysis of biological sequences over the Internet. The web portal offers convenient interactive access to a growing pool of chainable bioinformatics software tools and databases that are centrally installed and maintained by the RZG. Currently, supported tasks comprise sequence similarity searches in public or user-supplied databases, computation and validation of multiple sequence alignments, phylogenetic analysis and protein-structure prediction. Individual tools can be seamlessly chained into pipelines allowing the user to conveniently process complex workflows without the necessity to take care of any format conversions or tedious parsing of intermediate results. The toolkit is part of the Max-Planck Integrated Gene Analysis System (MIGenAS) of the Max Planck Society available at www.migenas.org (click 'Start Toolkit').

Animals↗

VirDetector: a bioinformatic pipeline for virus surveillance using nanopore sequencing.

SUMMARY: Virus surveillance programmes are designed to counter the growing threat of viral outbreaks to human health. Nanopore sequencing, in particular, has proven to be suitable for this purpose, as it is readily available and provides rapid results. However, as special bioinformatic programs are required to extract the relevant information from the sequencing data, applications are needed that allow users without extensive bioinformatics knowledge to carry out the relevant analysis steps. We present VirDetector, a bioinformatic pipeline for virus surveillance using nanopore sequencing. The pipeline automatically installs all required programs and databases and allows all its steps to be executed with a single console command. After preprocessing the samples, including the possibility for basecalling, the pipeline classifies each sample taxonomically and reconstructs the viral consensus genomes, which are then used in phylogenetic analyses. This streamlined workflow provides a user-friendly and efficient solution for monitoring viral pathogens. AVAILABILITY AND IMPLEMENTATION: VirDetector is freely available at https://github.com/NLKaiser/VirDetector and https://zenodo.org/records/14637302 (10.5281/zenodo.14637302).

Nanopore Sequencing↗

DNA methylation and machine learning: challenges and perspective toward enhanced clinical diagnostics.

DNA methylation is an epigenetic modification that regulates gene expression by adding methyl groups to DNA, affecting cellular function and disease development. Machine learning, a subset of artificial intelligence, analyzes large datasets to identify patterns and make predictions. Over the past two decades, advances in bioinformatics technologies for arrays and sequencing have generated vast amounts of data, leading to the widespread adoption of machine learning methods for analyzing complex biological information for medical problems. This review explores recent advancements in DNA methylation studies that leverage emerging machine learning techniques for more precise, comprehensive, and rapid patient diagnostics based on DNA methylation markers. We present a general workflow for researchers, from clinical research questions to result interpretation and monitoring. Additionally, we showcase successful examples in diagnosing cancer, neurodevelopmental disorders, and multifactorial diseases. Some of these studies have led to the development of diagnostic platforms that have entered the global healthcare market, highlighting the promising future of this field.

Humans↗

Targeted Modulation of Abundant Proteins Enhances Proteomic Profiling of Ovarian Cancer Ascites: A Pilot Technical Workflow Comparison.

Ascites from ovarian cancer patients are increasingly recognized as a valuable biofluid for cancer research, as its protein composition reflects the disease state and may reveal biomarkers of treatment sensitivity and response. However, the detection of low-abundance proteins is hindered by the presence of highly abundant proteins such as albumin. In this study, we evaluated five protein preparation methods for their effectiveness in depleting high-abundance or enriching low-abundance proteins in ovarian cancer ascites. The Norgen (Nor), Minutes (Min), and Perchloric acid (PerCA) methods were based on abundant protein depletion, while the Urine (Uri) and Nanomics (Nano) kits focused on low-abundance protein enrichment. Processed samples were analyzed using label-free quantitative bottom-up proteomics by LC-MS/MS, followed by a bioinformatics assessment. Compared with undepleted ascites (UnD), Min, Nor, Nano, and PerCA increased protein identifications, whereas Uri produced profiles similar to those of UnD. Notably, PerCA and Nano enabled the identification of distinct protein subsets associated with cancer-related pathways, including immune responses and autophagy. PerCA enriched transmembrane and secreted immunomodulatory glycoproteins, whereas Nano enrichment primarily captured secreted, nuclear, and cytoplasmic soluble proteins. Overall, our results show that both high-abundance protein depletion and low-abundance enrichment improve ascites proteome coverage, each offering distinct advantages in identifying biologically relevant low-abundance proteins.

Female↗

Clinical proteomics in inborn errors of metabolism: from biomarker discovery to implementation.

INTRODUCTION: Inborn errors of metabolism (IEMs) are rare, heterogeneous disorders traditionally diagnosed through genetic testing, enzyme assays, and metabolite measurements. However, these tools often do not fully explain phenotypic variability, organ involvement, disease progression, or treatment response. Clinical proteomics provides a complementary functional layer by capturing changes in protein abundance, proteoforms, post-translational modifications (PTM), and biological pathways, offering insights beyond genotype- and metabolite-based approaches. AREAS COVERED: This review examines the role of high-resolution mass spectrometry and computational proteomics in biomarker discovery and clinical decision-making for IEMs. It focuses on their contribution to diagnosis, variant interpretation, patient stratification, and treatment monitoring. Disease-specific applications are discussed, with the strongest evidence in lysosomal storage disorders, mitochondrial diseases, congenital disorders of glycosylation, and selected neurodegenerative or renal metabolic conditions. The literature search was performed in PubMed, Scopus, Web of Science, and Google Scholar, covering peer-reviewed articles available up to 2026, with emphasis on methodological advances and translational applications in clinical proteomics for IEMs. EXPERT OPINION: Proteomics will not replace established diagnostic tools, but it can help address clinically actionable questions in selected contexts. Translation into clinical practice will require standardized workflows, multicenter validation, clinically anchored endpoints, and integration with other omics approaches.

Humans↗

Multi-criteria decision making and its application to in silico discovery of vaccine candidates for Toxoplasma gondii.

Vaccine discovery against eukaryotic parasites is not trivial and few exist. Reverse vaccinology is an in silico vaccine discovery approach, designed to identify vaccine candidates from the thousands of protein sequences encoded by a target genome. Previously, we produced the Vacceed bioinformatics pipeline for identification of parasite membrane and excreted/secreted proteins that were likely be exposed to the hosts immune system. More recently, we improved upon machine learning as the final decision-making process to identify parasite proteins that induce a protective response in an animal model. Subsequently, we combined Vacceed with metrics on B and T cell epitope types to produce a new in silico discovery workflow. In this study we extend this in silico workflow to the developability of proteins as vaccines by the incorporation of metrics on the physicochemical properties of proteins. To demonstrate this process, every Toxoplasma gondii protein was ranked in its capacity to provide exposure to the immune system (Vacceed exposure score), presence of epitopes and solubility characteristics by several multicriteria decision making (MCDM) tools (such as TOPSIS, VIKOR and MABAC). A consensus rank was subsequently generated from the results of these tools using a variety of aggregate ranking methods. Levels of uncertainty in the aggregate protein rankings was assessed by conformal interval prediction in association with a machine learning model. Several of the top ranked proteins identified by this approach were novel, uncharacterized membrane transporters or proteins associated with RNA metabolism. In conclusion, MCDM automated the decision making using well known algorithms while conformal prediction intervals varied significantly across the 8000+ proteins of T. gondii. Highly ranked proteins (e.g. the top 100) typically generated low prediction intervals, providing high levels of confidence in their ranks.

Toxoplasma↗

Expanding and improving analyses of nucleotide recoding RNA-seq experiments with the EZbakR suite.

Nucleotide recoding RNA sequencing methods (NR-seq; TimeLapse-seq, SLAM-seq, TUC-seq, etc.) are powerful approaches for assaying transcript population dynamics. In addition, these methods have been extended to probe a host of regulated steps in the RNA life cycle. Current bioinformatic tools significantly constrain analyses of NR-seq data. To address this limitation, we developed EZbakR (https://github.com/isaacvock/EZbakR), an R package to facilitate a more comprehensive set of NR-seq analyses, and fastq2EZbakR (https://github.com/isaacvock/fastq2EZbakR), a Snakemake pipeline for flexible preprocessing of NR-seq datasets, collectively referred to as the EZbakR suite. Together, these tools generalize many aspects of the NR-seq analysis workflow. The fastq2EZbakR pipeline can assign reads to a diverse set of genomic features (e.g., genes, exons, splice junctions), and EZbakR can perform analyses on any combination of these features. EZbakR extends standard NR-seq mutational modeling to support multi-label analyses (e.g., s4U and s6G dual labeling), and implements an improved hierarchical model to better account for transcript-to-transcript variance in metabolic label incorporation. EZbakR also generalizes dynamical systems modeling of NR-seq data to support analyses of premature mRNA processing and flow between subcellular compartments. Finally, EZbakR implements flexible and well-powered comparative analyses of all estimated parameters via design matrix-specified generalized linear modeling. The EZbakR suite will thus allow researchers to make full, effective use of NR-seq data.

Software↗

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology↗

Clinical impact of 16S rRNA RC-PCR NGS on infectious disease management.

16S rRNA metagenomics provides a culture-independent method for diagnosing infections with fastidious or uncultivable organisms, guiding targeted therapy, and detecting polymicrobial communities. This study utilizes reverse complement (RC)-PCR next-generation sequencing (NGS) to accurately identify bacterial pathogens from clinical specimens and assess its impact on clinical decision-making, setting it apart from conventional 16S sequencing approaches. A retrospective analysis of an ISO 15189 accredited 16S RC-PCR NGS diagnostic workflow targeting the V1-6 and V9 regions of the 16S rRNA gene was conducted over a 2-year period, including 390 clinical specimens from 316 patients. 16S RC-PCR NGS results were discussed in a multidisciplinary consultation and subsequently reported to the clinic. In total, 1,283 RC-PCR results were analyzed, of which 517 were from clinical specimens, 284 were negative controls, 66 were positive controls, and 416 were from wet lab and bioinformatic pipeline validation. 16S RC-PCR NGS assay detected bacterial taxa in 179/390 (45.9%) of clinical specimens, while 201/390 (51.5%) were negative, and 10/390 (2.6%) yielded uninterpretable results. The specimen types pus, pleural fluid, and heart valves exhibited the highest positivity rate (68% to 70%). Overall, 16S RC-PCR NGS influenced diagnostic decision making in 145/282 (51.4%) clinical cases and guided therapeutic management in 77/282 (27.3%) cases. Results providing definite evidence for either the presence or absence of bacterial infection were considered clinically valuable. Integration of 16S RC-PCR NGS pathogen detection with multidisciplinary consultation markedly improved clinical management, directly impacting diagnosis and treatment of complex clinical cases in a tertiary care setting. The effect was most pronounced in brain abscess patients, where RC-PCR results guided treatment decisions in 9/13 (69.2%) of cases.IMPORTANCETimely and accurate diagnosis is essential for managing serious infections, yet clinicians often face situations where routine laboratory tests do not provide clear answers. This study demonstrates that next-generation sequencing (NGS) of the bacterial 16S rRNA gene can decisively resolve these uncertainties. By revealing whether bacteria are present in clinical specimens, this approach influenced clinical reasoning and supported treatment decisions across a variety of challenging cases. 16S reverse-complement PCR was especially powerful for brain abscesses and infections where the causative microorganism was unclear, providing clarity that directly improved patient care. These findings show that integrating advanced sequencing with expert clinical interpretation can enhance the management of complex infections and support more confident, evidence-based therapy.

Humans↗

Pangenomes aid accurate detection of large insertions and deletions from targeted sequencing: the case of cardiomyopathies.

BACKGROUND: Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it often fails to detect many larger variants. Recent studies have recommended the adoption of pangenome references (as opposed to linear reference genomes like GRCh38) to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. METHODS: Here, we analyze 1969 cardiomyopathy cases and 1805 controls sequenced with the Illumina Trusight Cardio panel using a pangenome-based workflow (GRAF) and five conventional orthogonal methodologies (GATK HaplotypeCaller, GATK-gCNV, ExomeDepth, Manta and Lumpy-SV) to detect variants ≥ 20 bp in size. RESULTS: Following lab-based variant validation by means of PCR and Sanger sequencing, we show that GRAF conjugates higher precision and recall (F1 score 0.86) compared with other methods (F1 0-0.57) in detecting potentially pathogenic variants ≥ 20 bp from short-read panel data. Results were complemented by a comparison of the tools' performance in detecting ground truth variants on reference sample HG002 from Genome In A Bottle, which confirmed GRAF to outperform other tools also on exome sequencing (F1 0.97 vs. 0-0.94). Notably, in the HG002 benchmark dataset, GRAF also showed slightly improved performance compared to GATK HaplotypeCaller in the identification of small variants (1-19 bp; F1 0.975 vs. 0.968). CONCLUSIONS: Our results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.

Humans↗

VisPan: real-time visualisation of multiplex amplicon-based sequencing panels for rapid syndromic surveillance and pathogen detection.

MOTIVATION: Infectious diseases persist as a major global public health challenge. Diverse factors, including climate change, globalization, deforestation, human-animal interactions, lifestyle choices, and various biological factors, can contribute to their emergence and reemergence. Rapid detection and characterization of (re)emerging pathogens are therefore critical for effective outbreak management and for enhancing our understanding of epidemics by monitoring the transmission, spread, evolution, and genomics of pathogens. In this context, next-generation sequencing technologies (NGS), particularly long-read platforms such as Oxford Nanopore Technologies (ONT), have opened new avenues for real-time pathogen monitoring. However, the bioinformatics bottleneck remains a challenge, emphasizing the need for efficient, accessible, and user-friendly analysis tools. RESULTS: Here, we present a tool adapted from the RAMPART software that enables real-time data visualisation of multiplex PCR syndromic panels combined with Oxford Nanopore sequencing. This real-time analysis enables rapid pathogen detection, from raw data acquisition to taxonomic assignment, within minutes. The interface offers dynamic visual tracking of the sequencing run and amplicon coverage, facilitating immediate insights during diagnostic workflows. Validation experiments confirmed the system's reliability, accurately identifying all pathogens present in complex clinical or environmental samples. This tool provides an integrated, user-friendly solution for genomic pathogen surveillance in field or clinical settings.

Software↗

Lectin capture strategies combined with mass spectrometry for the discovery of serum glycoprotein biomarkers.

The application of mass spectrometry to identify disease biomarkers in clinical fluids like serum using high throughput protein expression profiling continues to evolve as technology development, clinical study design, and bioinformatics improve. Previous protein expression profiling studies have offered needed insight into issues of technical reproducibility, instrument calibration, sample preparation, study design, and supervised bioinformatic data analysis. In this overview, new strategies to increase the utility of protein expression profiling for clinical biomarker assay development are discussed with an emphasis on utilizing differential lectin-based glycoprotein capture and targeted immunoassays. The carbohydrate binding specificities of different lectins offer a biological affinity approach that complements existing mass spectrometer capabilities and retains automated throughput options. Specific examples using serum samples from prostate cancer and hepatocellular carcinoma subjects are provided along with suggested experimental strategies for integration of lectin-based methods into clinical fluid expression profiling strategies. Our example workflow incorporates the necessity of early validation in biomarker discovery using an immunoaffinity-based targeted analytical approach that integrates well with upstream discovery technologies.

Amino Acid Sequence↗

Open source tools and toolkits for bioinformatics: significance, and where are we?

This review summarizes important work in open-source bioinformatics software that has occurred over the past couple of years. The survey is intended to illustrate how programs and toolkits whose source code has been developed or released under an Open Source license have changed informatics-heavy areas of life science research. Rather than creating a comprehensive list of all tools developed over the last 2-3 years, we use a few selected projects encompassing toolkit libraries, analysis tools, data analysis environments and interoperability standards to show how freely available and modifiable open-source software can serve as the foundation for building important applications, analysis workflows and resources.

Algorithms↗

Dual RNA isolation from blood: an optimized protocol for host and bacterial RNA purification for dual RNA-sequencing analysis in whole blood sepsis samples.

Dual RNA-sequencing (dual RNA-seq) holds significant promise for deciphering bacterial virulence mechanisms during systemic infections. However, its application in sepsis research is hindered by technical challenges, including a low bacterial burden in blood and limited sample volumes and RNA yield from vulnerable populations, such as neonates. We developed an optimized protocol [dual RNA isolation from blood (DRIB)] for simultaneous stabilization, isolation and purification of high-quality host leukocyte and bacterial RNA from low-volume whole blood samples (0.5 ml). This protocol is compatible with clinical sample collection workflows and high-throughput RNA sequencing. The feasibility of DRIB for dual RNA-seq was validated using a pilot cohort of clinical adult sepsis samples, enabling the investigation of host-bacterial gene expression during sepsis. The DRIB protocol yielded 2.10-6.91 µg of total RNA per clinical sample in our pilot cohort. Dual-species ribosomal RNA (rRNA) depletion and RNA-seq generated 16.6-24.8 million filtered reads per sample, with 63±7% of reads uniquely mapped to host or bacterial sequences. Host genes accounted for 51-68% (8.4-10.9 million) reads, while 0.5-6.7% (79,496-789,808 reads) mapped to bacterial genomes. Bioinformatic analysis revealed that both shared and individual transcriptional patterns were identified in host and bacterial responses, including pathways related to immune metabolism and metal-ion binding. Our optimized DRIB protocol and RNA-seq pipeline effectively captured both host and bacterial RNA transcription in clinical sepsis samples. Expanding this approach to larger cohorts and varying disease timepoints will provide crucial new insights into host-bacterial gene co-expression dynamics in sepsis progression and outcomes.

Humans↗