Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatics workflow”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Formalization of mouse embryo anatomy.

MOTIVATION: The Edinburgh Mouse Atlas and Gene Expression Database project has developed a digital atlas of mouse development to provide a spatio-temporal framework for spatially mapped data such as in situ gene expression and cell lineage. As part of this database, a mouse embryo anatomy ontology has been created. A formalization of this anatomy is required to document its precise semantics and how it is used in the context of the Mouse Atlas. RESULTS: The paper describes the existing anatomy ontology and formalizes aspects of it using a predicate logic based approach. It therefore provides a guide for users of the current version of the ontology, as well as the basis for a description of the anatomy using an ontology language, such as OWL, thus enabling future work on reasoning about the Mouse Atlas in the context of an intelligent gene expression bioinformatics workflow system. The logic has been implemented in a Prolog prototype. AVAILABILITY: The Mouse Atlas is available on-line at http://genex.hgu.mrc.ac.uk

Algorithms↗

Global maintenance of histone post-translational modifications during the transition into anoxia in embryos of the annual killifish Austrofundulus limnaeus.

Many organisms have adapted to survive anoxic or hypoxic environments, but the epigenetic responses involved in this successful stress response are not well described in most species. Embryos of the annual killifish Austrofundulus limnaeus have the greatest tolerance to anoxia of all vertebrates, making them a powerful model to study the cellular mechanisms necessary for anoxia tolerance. However, the global histone landscape of this species has never been quantified or explored in relation to stress tolerance. Liquid chromatography-mass spectrometry and a Python bioinformatics workflow were used to identify histones and their post-translational modifications. This pipeline resulted in the detection of 252 unique biologically relevant histone post-translational modifications (hPTMs) (unimod + residue). These PTMs represent 16 types of biologically relevant hPTMs present during both anoxia and normoxia in Wourms' stage 36 embryos. This hPTM library presents an exciting opportunity to study histone modifications across development and in response to environmental stressors. No significant changes in PTM or histone abundance were observed between anoxic and normoxic embryos, suggesting that 24 h of anoxia is not sufficient to induce epigenetic or histone isoform changes at the organismal level. This result is inconsistent with data presented for similar stresses in mammalian cells and thus stabilization of the hPTM landscape may be an adaptation that supports anoxia tolerance.

anoxia↗

Computational framework for the prediction of transcription factor binding sites by multiple data integration.

Control of gene expression is essential to the establishment and maintenance of all cell types, and its dysregulation is involved in pathogenesis of several diseases. Accurate computational predictions of transcription factor regulation may thus help in understanding complex diseases, including mental disorders in which dysregulation of neural gene expression is thought to play a key role. However, biological mechanisms underlying the regulation of gene expression are not completely understood, and predictions via bioinformatics tools are typically poorly specific. We developed a bioinformatics workflow for the prediction of transcription factor binding sites from several independent datasets. We show the advantages of integrating information based on evolutionary conservation and gene expression, when tackling the problem of binding site prediction. Consistent results were obtained on a large simulated dataset consisting of 13050 in silico promoter sequences, on a set of 161 human gene promoters for which binding sites are known, and on a smaller set of promoters of Myc target genes. Our computational framework for binding site prediction can integrate multiple sources of data, and its performance was tested on different datasets. Our results show that integrating information from multiple data sources, such as genomic sequence of genes' promoters, conservation over multiple species, and gene expression data, indeed improves the accuracy of computational predictions.

Animals↗

XML schemas for common bioinformatic data types and their application in workflow systems.

BACKGROUND: Today, there is a growing need in bioinformatics to combine available software tools into chains, thus building complex applications from existing single-task tools. To create such workflows, the tools involved have to be able to work with each other's data--therefore, a common set of well-defined data formats is needed. Unfortunately, current bioinformatic tools use a great variety of heterogeneous formats. RESULTS: Acknowledging the need for common formats, the Helmholtz Open BioInformatics Technology network (HOBIT) identified several basic data types used in bioinformatics and developed appropriate format descriptions, formally defined by XML schemas, and incorporated them in a Java library (BioDOM). These schemas currently cover sequence, sequence alignment, RNA secondary structure and RNA secondary structure alignment formats in a form that is independent of any specific program, thus enabling seamless interoperation of different tools. All XML formats are available at http://bioschemas.sourceforge.net, the BioDOM library can be obtained at http://biodom.sourceforge.net. CONCLUSION: The HOBIT XML schemas and the BioDOM library simplify adding XML support to newly created and existing bioinformatic tools, enabling these tools to interoperate seamlessly in workflow scenarios.

Algorithms↗

ProGenGrid: a grid-enabled platform for bioinformatics.

In this paper we describe the ProGenGrid (Proteomics and Genomics Grid) system, developed at the CACT/ISUFI of the University of Lecce which aims at providing a virtual laboratory where e-scientists can simulate biological experiments, composing existing analysis and visualization tools, monitoring their execution, storing the intermediate and final output and finally, if needed, saving the model of the experiment for updating or reproducing it. The tools that we are considering are software components wrapped as Web Services and composed through a workflow. Since bioinformatics applications need to use high performance machines or a high number of workstations to reduce the computational time, we are exploiting a Grid infrastructure for interconnecting wide-spread tools and hardware resources. As an example, we are considering some algorithms and tools needed for drug design, providing them as services, through easy to use interfaces such as the Web and Web service interfaces built using the open source gSOAP Toolkit, whereas as Grid middleware we are using the Globus Toolkit 3.2, exploiting some protocols such as GSI and GridFTP.

Computational Biology↗

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome↗

Parameter-space screening: a powerful tool for high-throughput crystal structure determination.

The determination of protein structures on a genomic scale requires both computing capacity and efficiency increases at many stages along the complex process. By combining bioinformatics workflow-management techniques, cluster-based computing and popular crystallographic structure-determination software packages, an efficient and powerful new tool for structural biology/genomics has been developed. Using the workflow manager and a simple web interface, the researcher can, in a few easy steps, set up hundreds of structure-determination jobs, each using a slightly different set of program input parameters, thus efficiently screening parameter space for the optimal input-parameter combination, i.e. a set of parameters that leads to a successful structure determination. Upon completion, results from the programs are harvested, analyzed, sorted based on success and presented to the user via the web interface. This approach has been applied with success in more than 30 cases. Examples of successful structure determinations based on single-wavelength scattering (SAS) are described and include cases where the 'rational' crystallographer-based selection of input parameters values had failed.

Computational Biology↗

Misdetection of frameshifts in SARS-CoV-2 genomes: need for additional harmonisation and efficient monitoring of data workflows.

Five years after the outbreak of the SARS-CoV-2 pandemic in 2020, diagnostic laboratories have moved from massive sequencing of thousands of samples to routine surveillance of SARS-CoV-2 cases, as with all other respiratory viruses. Surveillance remains of paramount importance to prevent a further SARS-CoV-2 surge, as the virus has been shown to mutate rapidly and can render available drugs and vaccines ineffective. During the pandemic, several bioinformatics pipelines and workflows have been developed to streamline analysis, shorten turnaround time and ensure reproducibility. As the number of samples decreases, laboratories are moving towards more flexible sequencing strategies and optimizing the cost per sample. However, workflow redesigns, even if individual steps have proven successful time and time again, can lead to challenges when changes in a bioinformatics pipeline are introduced (e.g. version updates, implementation of new features, etc.), a new combination of viral mutations emerge or a change in wet-lab procedures leads to unpredictable results. Here, we present a report of misidentified frameshift mutations in the consensus sequence of SARS-CoV-2, which led to an incorrect assumption of mutations in the spike and nucleocapsid viral proteins with the potential to affect PCR detection or even antigen testing. This investigation exemplifies the need for better awareness of the challenges that can occur even when using routinely applied protocols and analytical workflows and highlights the need for cooperation between experts of NGS, bioinformaticians and decision-makers towards more harmonized data workflows.

SARS-CoV-2↗

Wildfire: distributed, Grid-enabled workflow construction and execution.

BACKGROUND: We observe two trends in bioinformatics: (i) analyses are increasing in complexity, often requiring several applications to be run as a workflow; and (ii) multiple CPU clusters and Grids are available to more scientists. The traditional solution to the problem of running workflows across multiple CPUs required programming, often in a scripting language such as perl. Programming places such solutions beyond the reach of many bioinformatics consumers. RESULTS: We present Wildfire, a graphical user interface for constructing and running workflows. Wildfire borrows user interface features from Jemboss and adds a drag-and-drop interface allowing the user to compose EMBOSS (and other) programs into workflows. For execution, Wildfire uses GEL, the underlying workflow execution engine, which can exploit available parallelism on multiple CPU machines including Beowulf-class clusters and Grids. CONCLUSION: Wildfire simplifies the tasks of constructing and executing bioinformatics workflows.

Algorithms↗

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational↗

DNA sequencing for microbial surveillance in cystic fibrosis airways: advances, challenges, and clinical translation.

SUMMARYDNA sequencing has revolutionized microbial surveillance in cystic fibrosis (CF), transforming pathogen identification from culture-dependent to total microbial community identification using molecular-based approaches. Techniques such as 16S rRNA gene sequencing have uncovered the complexity of the CF airway microbiome, while shotgun metagenomics, metatranscriptomics, and viromics now provide strain-level, functional, and viral insights beyond bacterial identification. Despite these advances, key technical and logistical challenges remain, including the processing of high-viscosity sputum samples, overwhelming host DNA contamination, managing large data sets, and the integration of complex bioinformatic outputs into clinical workflows. Emerging innovations such as host DNA depletion protocols, targeted enrichment panels, and adaptive sampling on Oxford Nanopore platforms are helping to overcome these barriers, improving microbial recovery and sequencing efficiency. As cystic fibrosis transmembrane conductance regulator (CFTR) modulator therapies are changing the lives of people with cystic fibrosis (pwCF), sequencing offers an unprecedented opportunity to track potential microbial adaptation in response. This review investigates current advances, limitations, and translational opportunities in DNA sequencing for CF airway microbiome surveillance, highlighting how these technologies can help reshape research and clinical microbiology in the post-modulator era.

Cystic Fibrosis↗

High performance workflow implementation for protein surface characterization using grid technology.

BACKGROUND: This study concerns the development of a high performance workflow that, using grid technology, correlates different kinds of Bioinformatics data, starting from the base pairs of the nucleotide sequence to the exposed residues of the protein surface. The implementation of this workflow is based on the Italian Grid.it project infrastructure, that is a network of several computational resources and storage facilities distributed at different grid sites. METHODS: Workflows are very common in Bioinformatics because they allow to process large quantities of data by delegating the management of resources to the information streaming. Grid technology optimizes the computational load during the different workflow steps, dividing the more expensive tasks into a set of small jobs. RESULTS: Grid technology allows efficient database management, a crucial problem for obtaining good results in Bioinformatics applications. The proposed workflow is implemented to integrate huge amounts of data and the results themselves must be stored into a relational database, which results as the added value to the global knowledge. CONCLUSION: A web interface has been developed to make this technology accessible to grid users. Once the workflow has started, by means of the simplified interface, it is possible to follow all the different steps throughout the data processing. Eventually, when the workflow has been terminated, the different features of the protein, like the amino acids exposed on the protein surface, can be compared with the data present in the output database.

Automation↗

High performance GRID based implementation for genomics and protein analysis.

Starting from the genomic and proteomic sequence data, a complex computational infrastructure as been established with the objective to develop a GRID based system to to automate the analysis, prediction and annotation processes of genomic DNA. To support of this type of analysis, several algorithms as been used to recognize biological signals involved in the identification of genes and proteins. The system implemented can be use to analyse the content of the large number of genomic sequences. For this reason, the system realized is capable of using a computational architecture specifically designed for intensive computing based on GRID technologies developed throughout the BIOINFOGRID European project. We developed a GRID based workflow to correlate different kind of Bioinformatics data, going from the Genomics Nucleotide to the Protein Sequence. The first step in the workflow consists of submitting a nucleotide sequence that is elaborated by a specific software for gene prediction. In particular this tool performs a search in the nucleotide sequence to find out the key components of gene. The predicted gene is then translated in the corresponding protein sequence. Based on protein sequence is then possible to identify the domains that characterize the protein functionality using specific tools of domain prediction. Protein domains classification are very important in the analysis of the macromolecular functionality. To analyze a whole protein family from large genome of various organism means to elaborate a large amount of data that requires huge computational resources. To analyze all this data we suggest the use of a high performance platform based on grid technology. We have implemented our applications on a wide area grid platform for scientific applications [http://www.grid.it and http://grid-it.cnaf.infn.it] composed of about 1000 CPU's. The grid infrastructure consists in a collection of computing elements and storage elements that jointly concur to define a platform for high performance elaboration. In this study a grid based application is presented to compute the protein domain analysis in a distributed way. This approach has high performance because the protein domains are checked with different software in parallel in different grid sites.

Computational Biology↗

A qualitative study of the implementation of a bioinformatics tool in a biological research laboratory.

OBJECTIVE: To explore how the implementation of a comprehensive new bioinformatics analysis system would affect workflow, collaboration and information management in a small genetic research lab. DESIGN: This was a longitudinal qualitative study of seven individuals involved in genomic and proteomic research. The study data were gathered using the illuminative/responsive approach of immersion in the environment. Additional qualitative data were gathered using informal semi-structured interviews, participant observation in lab meetings, and direct observation of lab researchers engaged in specific tasks. MEASUREMENTS: Interview, observation and field note data were coded and analyzed based on three analysis perspectives. A subset of the data was independently evaluated by an external researcher to enhance the trustworthiness of results. RESULTS: Three reoccurring themes were observed in the study. (1) Satisfaction and acceptance of software tools tended to be role and goal specific. (2) The system was seen primarily as a measurement system rather than a "total laboratory analysis system". (3) Lab meetings deemphasized the system, preferring more traditional data analysis techniques. These themes support the observations that the system was not used to its full potential in the lab. CONCLUSION: Themes identified in this study suggest that sophisticated genetic researchers face similar problems of technology implementation as do professionals in other fields. We recommend that leadership support and on-going training and evolution of academic curricula can improve chances of bioinformatics analysis systems becoming used more effectively.

Computational Biology↗

Is There a Fly in My Soup? To What Extent Do Metabarcoding and Individual Barcoding Tell the Same Story?

Metabarcoding has become the method of choice for characterizing complex arthropod communities. The extent to which metabarcoded bulk samples will recover the same community composition as individual sequencing of all individuals in the sample remains poorly quantified. Biases such as unequal extraction of DNA from different taxa, primer mismatches and non-random PCR may cause the selective drop-out of species from metabarcoding data. At the same time, DNA metabarcoding may reveal arthropod taxa present not as individuals, but as DNA residues on the surface or in the gut of insects. To quantify the consistency in sample contents established by different means, we metabarcoded 45 bulk insect samples, then extracted all arthropods and sequenced them individually. Metabarcoding targeted 418 bp at the 3' end of the Folmer barcoding region, while individual barcodes captured the entire 658 bp Folmer region. The metabarcoding workflow, including PCR amplification, sequencing and bioinformatics, was performed in three replicates from three separate lysate aliquots per sample. For the main analyses, sequences were assigned to Barcode Index Numbers (BINs) as identical taxonomic categories across data types, thereby allowing the detection of even rare but biologically true taxa. Since such reference-based validation will be unavailable to any researcher dealing with metabarcoding data alone, we validated our key findings through an alternative workflow, i.e., de novo clustering of sequences. We found that metabarcoding is replicable, as different replicates of the same sample recover similar species richness and composition. Individual barcoding and metabarcoding provide similar impressions of relative differences in community structure: species-rich vs. species-poor samples rank similarly among data types (Spearman's ⍴ = 0.88-0.99) as do differences in relative dissimilarity between sample pairs (Spearman's ⍴ = 0.55-0.90). Dissimilarity between data types varies with BIN richness in the sample, but this relationship reflects nestedness rather than turnover: metabarcoding recovers the same set of core species as individual barcoding but adds hundreds of species on top. Any BIN recovered as an individual occurred with high probability in the metabarcoding data, and any BIN found in high read abundances by metabarcoding was likely found as an individual (p > 0.8). In terms of abundances, the number of individual insects per BIN was well predicted by the number of metabarcoding reads (R2 > 0.68 for a model including taxonomy as a random effect). Our analysis suggests that metabarcoding data will be informative of the sample contents in terms of arthropod species richness, composition and taxon-specific abundances. Taxa recovered in low copy numbers in metabarcoding sequence data will likely represent DNA left as residues from past biotic interactions. Barring sequencing errors, both types of data yield biologically relevant insights into the taxa present in the source community.

Animals↗

Backtracking Cell Phylogenies in the Human Brain with Somatic Mosaic Variants.

Somatic mosaic variants, and especially somatic single nucleotide variants (sSNVs), occur in progenitor cells in the developing human brain frequently enough to provide permanent, unique, and cumulative markers of cell divisions and clones. Here, we describe an experimental workflow to perform lineage studies in the human brain using somatic variants. The workflow consists in two major steps: (1) sSNV calling through whole-genome sequencing (WGS) of bulk (non-single-cell) DNA extracted from human fresh-frozen tissue biopsies, and (2) sSNV validation and cell phylogeny deciphering through single nuclei whole-genome amplification (WGA) followed by targeted sequencing of sSNV loci.

Humans↗

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software↗

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning↗