Search PubMedSearch

SEARCH · Search PubMed

Results for “Nextflow”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

24 records · Page 2Linked to original sources

CERTOMICS: trusted single-cell multiomics pipeline for high-resolution profiling of adoptive cellular immunotherapies.

SUMMARY: Adoptive cellular immunontherapies, such as chimeric antigen receptor (CAR) T cell therapy, have transformed cancer treatment, yet challenges such as resistance, relapse, and high costs limit their efficacy and accessibility. A comprehensive understanding of cellular heterogeneity and molecular profiles is essential to improve these therapies. Advanced single-cell multiomics technologies have the power to analyze the complex interactions between CAR-engineered cells, immune cells, and tumor cells. However, standardized single-cell multiomics computational pipelines specifically tailored to CAR-engineered cell products are lacking. Due to the synthetic nature of CAR transgenes, additional steps for reliable identification and characterization of CAR-positive cells are required but not included in existing data-processing workflows. To address this, we present CERTOMICS, a Nextflow-based, CAR-aware pipeline offering enhanced CERTainty in immunophenotyping and data interpretation, tailored for single-cell multiOMICSprofiling of adoptive cellular immunotherapies. The pipeline standardizes processing 10x Genomics single-cell multiomics data and integrates CAR-specific identification and quality control. Additionally, a curated repository of CAR construct sequences and annotation data is provided, serving as an extensible resource to support the analysis and development of CAR T cell therapies. AVAILABILITY AND IMPLEMENTATION: Detailed documentation of this pipeline, along with a resource on latest FDA-approved CAR therapies is available on our website: https://fraunhofer-izi.github.io/Living-Drugs-Wiki/. The data underlying this article are available on GitHub at https://github.com/fraunhofer-izi/CERTOMICS. The code is also published on Zenodo at https://doi.org/10.5281/zenodo.18709693.

Multiomics

TaxTriage: an open-source metagenomic sequencing data analysis pipeline enabling putative pathogen detection.

MOTIVATION: TaxTriage is a comprehensive pathogen identification workflow designed for both short- and long-read untargeted DNA and RNA sequencing data. Combining read classification, mapping, and de novo assembly approaches, putative pathogens are identified through comparisons to curated pathogens and abundance expectations from healthy cohort data. Flexible installation options are enabled using Nextflow™ (NF), including cloud deployment via NF Tower (Seqera Platform) and local installation on a variety of systems, including standalone installations without external internet access. Final analysis summaries are compiled into an Organism Discovery Report, which lists likely pathogens and supporting data, including a custom confidence score. RESULTS: Evaluation of published in silico, clinical, and outbreak datasets identified performance comparable to alternative cloud-based processing pipelines for expected pathogen and co-infection detection with similar sensitivity and increased specificity. To support both public health and veterinary diagnostics communities, customization options have been incorporated to enable improved performance for host species of interest. AVAILABILITY AND IMPLEMENTATION: Source code for TaxTriage is freely available at https://github.com/jhuapl-bio/taxtriage. TaxTriage v2.1.1 has been archived on Zenodo at https://zenodo.org/records/17081354 to permit reproducible analysis as described in this manuscript.

Software

nf-core/magmap: Map metatranscriptomes to large collections of genomes.

SUMMARY: The lack of publicly available reference genomes has forced annotation of metatranscriptomes to either use direct alignment of sequence reads to reference databases or de novo assembly. As more and more natural environments are covered by metagenomic surveys, this is rapidly changing. This opens up the possibility of genome-resolved studies of prokaryotic metatranscriptomes by mapping to genomes from public repositories or metagenome-assembled genomes derived from the same environment. Here, we present the nf-core/magmap pipeline that provides a reproducible, easy-to-access, and well-documented workflow for selecting reference genomes, mapping to them, and quantifying features. Genomes can be drawn from public sources or originate from private collections. The pipeline is primarily aimed at prokaryotic communities but can, together with collections of reference mature gene sequences, also be applied to eukaryotes. AVAILABILITY AND IMPLEMENTATION: The nf-core/magmap pipeline is implemented in Nextflow and part of the nf-core collaboration. The pipeline is available at the nf-core website (https://nf-co.re/magmap) and GitHub (https://github.com/nf-core/magmap).

Software

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software

Semiparametric efficient estimation of small genetic effects in large-scale population cohorts.

Population genetics seeks to quantify DNA variant associations with traits or diseases, as well as interactions among variants and with environmental factors. Computing millions of estimates in large cohorts in which small effect sizes and tight confidence intervals are expected, necessitates minimizing model-misspecification bias to increase power and control false discoveries. We present TarGene, a unified statistical workflow for the semi-parametric efficient and double robust estimation of genetic effects including $ k $-point interactions among categorical variables in the presence of confounding and weak population dependence. $ k $-point interactions, or Average Interaction Effects (AIEs), are a direct generalization of the usual average treatment effect (ATE). We estimate genetic effects with cross-validated and/or weighted versions of Targeted Minimum Loss-based Estimators (TMLE) and One-Step Estimators (OSE). The effect of dependence among data units on variance estimates is corrected by using sieve plateau variance estimators based on genetic relatedness across the units. We present extensive realistic simulations to demonstrate power, coverage, and control of type I error. Our motivating application is the targeted estimation of genetic effects on trait, including two-point and higher-order gene-gene and gene-environment interactions, in large-scale genomic databases such as UK Biobank and All of Us. All cross-validated and/or weighted TMLE and OSE for the AIE $ k $-point interaction, as well as ATEs, conditional ATEs and functions thereof, are implemented in the general purpose Julia package TMLE.jl. For high-throughput applications in population genomics, we provide the open-source Nextflow pipeline and software TarGene which integrates seamlessly with modern high-performance and cloud computing platforms.

Humans

Whole Genome Methylation Sequencing via Enzymatic Conversion (EM-seq): Protocol, Data Processing, and Analysis.

Whole genome bisulfite sequencing (WGBS) has been the gold standard technique for base resolution analysis of DNA methylation for the last 15 years. It has been, however, associated with technical biases, which lead to overall overestimation of global and regional methylation values, and significant artifacts in extreme cytosine-rich DNA sequence contexts. Enzymatic conversion of cytosine is the newest approach, set to replace entirely the use of the damaging bisulfite conversion of DNA. The EM-seq technique utilizes TET2, T4-BGT, and APOBEC in a two-step conversion process, where the modified cytosines are first protected by oxidation and glucosylation, followed by deamination of all unmodified cytosines to uracil. As a result, EM-seq is degradation-free and bias-free, requires low DNA input, and produces high library yields with longer reads, little batch variation, less duplication, uniform genomic coverage, accurate methylation over a larger number of captured CpGs, and no sequence-specific artifacts.

DNA Methylation