Search PubMedSearch

SEARCH · Search PubMed

Results for “python”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans

WinPCA: a package for windowed principal component analysis.

SUMMARY: With chromosomal reference genomes and population-scale whole genome-sequencing becoming increasingly accessible, contemporary studies often include characterizations of the genomic landscape as it varies along chromosomes, commonly termed genome scans. While traditional summary statistics like FST and dXY between pre-assigned populations remain integral to characterizing the genomic divergence profile, PCA differs by providing single-sample resolution, thereby supporting the identification of polymorphic inversions, introgression and other types of divergent sequence that may not be fully aligned with global population structure. Here, we introduce WinPCA, a user-friendly package to compute, polarize and visualize genetic principal components in windows along the genome. To accommodate low-coverage whole genome-sequencing datasets, WinPCA can optionally make use of PCAngsd methods to compute principal components in a genotype likelihood framework. WinPCA accepts variant data in either VCF or BEAGLE format and can generate rich plots for interactive data exploration and downstream presentation. AVAILABILITY AND IMPLEMENTATION: WinPCA is implemented in Python and freely available at https://github.com/MoritzBlumer/winpca and https://doi.org/10.5281/zenodo.15614979.

Software

DeNoFo: a file format and toolkit for standardized, comparable de novo gene annotation.

MOTIVATION: De novo genes emerge from previously non-coding regions of the genome, challenging the traditional view that new genes primarily arise through duplication and adaptation of existing ones. Characterized by their rapid evolution and their novel structural properties or functional roles, de novo genes represent a young area of research. Therefore, the field currently lacks established standards and methodologies, leading to inconsistent terminology and challenges in comparing and reproducing results. RESULTS: This work presents a standardized annotation format to document the methodology of de novo gene datasets in a reproducible way. We developed DeNoFo, a toolkit to provide easy access to this format that simplifies annotation of datasets and facilitates comparison across studies. Unifying the different protocols and methods in one standardized format, while providing integration into established file formats, such as fasta or gff, ensures comparability of studies and advances new insights in this rapidly evolving field. AVAILABILITY AND IMPLEMENTATION: DeNoFo is available through the official Python Package Index (PyPI) and at https://github.com/EDohmen/denofo. All tools have a graphical user interface and a command line interface. The toolkit is implemented in Python3, available for all major platforms and installable with pip and uv.

Software

GRUMB: a genome-resolved metagenomic framework for monitoring urban microbiomes and diagnosing pathogen risk.

SUMMARY: Urban infrastructure hosts dynamic microbial communities that complicate biosurveillance and AMR monitoring. Existing tools rarely combine genome-resolved reconstruction with ecological modeling and batch-aware analytics tailored to infrastructure-scale studies. We present GRUMB (Genome-Resolved Urban Microbiome Biosurveillance), an open-source, SLURM-compatible pipeline that reconstructs high-quality metagenome-assembled genomes (MAGs) from shotgun sequencing reads and integrates taxonomic/functional annotation (CARD, VFDB), batch-aware normalization, ecological diagnostics and machine learning classification of environment types with uncertainty and risk scoring. GRUMB accepts either SRA project accessions or paired-end FASTQ files with metadata, and produces assemblies, MAGs, taxonomic and functional profiles, ecological outputs and risk-informed classification. Its modular design enables reproducible, infrastructure-scale biosurveillance across diverse environments. AVAILABILITY AND IMPLEMENTATION: GRUMB is freely available under the MIT License at: https://github.com/SuleimanAminu/genome-resolved-urban-microbiome-biosurveillance; Zenodo DOI: https://doi.org/10.5281/zenodo.15505402. Requirements: Linux (Ubuntu 20.04+), Python 3.11, R 4.2+, SLURM. Issues and feature requests are tracked on GitHub.

Microbiota

Synteny plot quality control with SyntenyQC.

SUMMARY: SyntenyQC is a data pre-processing tool for the construction of synteny plots. It supports genomic data collection, annotation and dereplication to facilitate (and in some cases fundamentally enable) the construction of informative synteny plots. AVAILABILITY AND IMPLEMENTATION: SyntenyQC is a command line app developed using Python version 3.10 and tested using pytest. SyntenyQC is available on PyPI (https://pypi.org/project/SyntenyQC) under the MIT License, along with a detailed user tutorial. Package tests can be viewed at https://github.com/Tim-Kirkwood/SyntenyQC.

Synteny

SISTEM: simulation of tumor evolution, metastasis, and DNA-seq data under genotype-driven selection.

SUMMARY: SISTEM is a software package and mathematical framework for simulating tumor evolution and cell migrations at single-cell resolution. Unlike existing frameworks which simulate cancer cell populations under the neutral coalescent or using simple birth-death models, SISTEM simulates tumor populations under somatic clonal selection using an agent-based framework. SISTEM can generate mutation profiles, read counts, and DNA sequencing reads along with ground truth cell lineages and migration graphs under a number of easily customizable mutation and selection models. For improved realism, SISTEM allows for cell fitness to be driven by genomic events of various scales including single nucleotide variants, segmental gains and losses, whole-chromosomal and chromosome-arm aberrations, and whole-genome duplications. SISTEM also includes numerous migration models to simulate metastatic cancers, facilitating the exploration and evaluation of diverse migration patterns. AVAILABILITY AND IMPLEMENTATION: SISTEM is written in Python and is freely available open-source under GNU GPLv3 from: https://github.com/samsonweiner/sistem.

Software

Profiler: an open web platform for multi-omics analysis.

MOTIVATION: High-throughput multi-omics technologies produce increasingly large and heterogeneous datasets that are difficult to analyze without advanced computational expertise. Existing bioinformatics tools are often fragmented or limited to specific omics types, hindering reproducibility and accessibility. There is a critical need for an integrated, user-friendly, and scalable platform capable of supporting multi-omics analyses across different data modalities. RESULTS: We present Profiler, an open-source, modular platform that unifies data import, quality control, preprocessing, statistical testing, machine and deep learning, biomarker discovery, pathway and drug-target enrichment, and survival modeling within a single reproducible environment. Built in Python with Streamlit, Profiler is available as both a web-based platform deployed on high-performance computing and a desktop version for local execution, enabling flexible usage across computational infrastructures. Profiler supports diverse omics modalities, including proteomics, transcriptomics, lipidomics, and electroencephalogram data. Through applications to glioblastoma proteomic, pancancer, and multi-omics datasets, Profiler reproduced known molecular subtypes, revealed potential therapeutic targets, and generated fully traceable analysis reports within minutes. By integrating advanced analytics behind an intuitive interface, Profiler democratizes multi-omics analysis and provides a robust, scalable foundation for systems biology and precision medicine research. AVAILABILITY AND IMPLEMENTATION: Profiler is open-source and freely available via its web platform (https://prism-profiler.univ-lille.fr) and GitHub (web version: https://github.com/yanisZirem/Profiler_v1_requests_datatests, desktop version: https://github.com/yanisZirem/prism-profiler), and archived on Zenodo (DOI: https://doi.org/10.5281/zenodo.17478158).

Software

Chrom-Sig: de-noising 1D genomic profiles by signal processing methods.

MOTIVATION: Modern genomic research is driven by next-generation sequencing experiments such as ChIP-seq, CUT&Tag, and CUT&RUN that generate coverage files for transcription factor binding, as well as ATAC-seq that yield coverage files for chromatin accessibility. Due to the inherent technical noise present in the experimental protocols, researchers need statistically rigorous and computationally efficient methods to extract true biological signal from a mixture of signal and noise. However, existing approaches are often computationally demanding or require input or spike-in controls. RESULTS: We developed Chrom-Sig, a Python package to quickly de-noise 1D genomic coverage tracks by computing the empirical null distribution without prior assumptions or experimental controls. When tested on 19 ChIP-seq, CUT&RUN, ATAC-seq, and snATAC-seq datasets, Chrom-Sig can effectively decompose the data into signal and noise components. Notably, Chrom-Sig performs de-noising and peak calling in 1-2 h using around 20 GB of memory. The de-noised signal corroborates with biologically meaningful results: CTCF CUT&RUN data retained a high percentage of peaks overlapping CTCF binding motifs, while ATAC-seq and RNA Polymerase II data were enriched in enhancers and promoters. We envision Chrom-Sig to be a versatile and general tool for current and future genomic technologies. AVAILABILITY AND IMPLEMENTATION: Chrom-Sig is publicly available on GitHub (https://github.com/minjikimlab/chromsig) and Zenodo (doi: 10.5281/zenodo.17488772) under the MIT licence.

Genomics

MHASS: Microbiome HiFi Amplicon Sequencing Simulator.

SUMMARY: Microbiome HiFi Amplicon Sequence Simulator (MHASS) creates realistic synthetic PacBio HiFi amplicon sequencing datasets for microbiome studies, by integrating genome-aware abundance modeling, realistic dual-barcoding strategies, and empirically derived pass-number distributions from actual sequencing runs. MHASS generates datasets tailored for rigorous benchmarking and validation of long-read microbiome analysis workflows, including ASV clustering and taxonomic assignment. AVAILABILITY AND IMPLEMENTATION: Implemented in Python with automated dependency management, the source code for MHASS is freely available at https://github.com/rhowardstone/MHASS along with installation instructions. Our code is also published on Zenodo at https://doi.org/10.5281/zenodo.17486364. The data underlying this article are available on GitHub at https://github.com/rhowardstone/MHASS_evaluation/.

Software

esloco: simulation-based estimation of local coverage in long-read DNA sequencing.

SUMMARY: Long-read DNA sequencing is increasingly applied for whole-genome studies, yet experimental planning often lacks reliable estimates of target region coverage, leading to costly and time-consuming pilot studies and replicates. We present esloco, a Monte Carlo-based simulation framework for estimating local coverage in long-read sequencing experiments, including scenarios with unknown target regions (e.g. viral integration, CRISPR-Cas9) or PCR-free designs (e.g. base modifications). By modeling coverage as a function of sequencing depth and read length distribution, esloco enables informed predictions of local sequencing outcomes. Benchmarking across a 45-gene panel demonstrated close agreement with empirical data, underscoring the framework's reliability. AVAILABILITY AND IMPLEMENTATION: esloco is a Python package available on PyPI (https://pypi.org/project/esloco/), GitHub (https://github.com/aweich/esloco), and Zenodo (https://doi.org/10.5281/zenodo.17776161).

Sequence Analysis, DNA

diffMONT: predicting methylation-specific PCR biomarkers based on nanopore sequencing data for clinical application.

MOTIVATION: DNA methylation serves as a key biomarker in clinical diagnostics, especially in cancer detection. With methylation-specific PCR (MSP), a widely used approach, patient samples can be screened fast and efficiently for differential methylation. During MSP, methylated regions are selectively amplified with specific primers. With nanopore sequencing, knowledge about DNA methylation is generated during direct DNA sequencing without needing pretreatment of the DNA. Multiple methods, mainly developed for whole-genome bisulfite sequencing (WGBS) data, exist to predict differentially methylated regions (DMRs) in the genome. However, the predicted DMRs are often very large and not sufficiently discriminating to generate meaningful results in MSP, creating a gap between theoretical cancer marker research and practical application, as no tool currently provides methylation difference predictions tailored for PCR-based diagnostics. RESULTS: Here, we present diffMONT, a tool that predicts differentially methylated regions specifically suited for MSP primer design, enabling rapid translation into practical applications. diffMONT takes into account (i) the specific length of primer and amplicon regions, (ii) the fact that one condition should be unmethylated, and (iii) a minimal required amount of differentially methylated cytosines within the primer regions. We compared the results of diffMONT to metilene and DSS based on a publicly available nanopore sequencing dataset and show that the regions predicted by diffMONT are more specific toward hypermethylated regions. diffMONT accelerates the design of methylation-specific diagnostic assays, bridging the gap between theoretical research and clinical application. AVAILABILITY AND IMPLEMENTATION: The source code for diffMONT, an open-source Python-based tool, is available at https://github.com/rnajena/diffMONT/, with an archived release under https://zenodo.org/records/17641031.

DNA Methylation

HXMS: a standardized file format for HX-MS data.

MOTIVATION: Hydrogen/deuterium exchange-mass spectrometry (HX-MS) is a rapidly expanding technique used to investigate protein conformational ensembles. The growing popularity and utility of HX-MS has driven the development of diverse instrumentation and software, resulting in inconsistent, non-standardized data analysis and representation. Most HX-MS data formats also employ only mean deuteration representations of the data rather than full isotopic mass spectra, which reduces the information content of the data and limits downstream quantitative analysis. RESULTS: Inspired by reliable protein structure and genomics data formats, we present HXMS, a unified, lightweight, scalable, and human-readable file format for HX-MS data. The HXMS format preserves the isotopic mass envelopes for all peptides, captures the full experimental time-course including fully deuterated control samples, and contains all other key information. It supports multimodal distributions, post-translational modifications (PTMs), and experimental replicates. To promote compatibility with existing HX-MS workflows, we also developed PFLink, a Python package that converts exported data files from commonly used HX-MS software to the HXMS format. PFLink and the HXMS format will enable quantitative, higher-resolution data processing, improved data sharing and storage among HX-MS practitioners, future machine learning applications, and further developments in HX-MS analysis. AVAILABILITY AND IMPLEMENTATION: PFLink is publicly available to install locally on HuggingFace, alongside documentation, or use online at HuggingFace (https://huggingface.co/spaces/glasgow-lab/PFlink). The supplementary information includes sample input files, sample HXMS files, and a generic unfilled PFlink custom CSV file that users may populate with key experimental conditions and results, which can then be read and converted into the HXMS format.

Software

igv-reports: embedding interactive genomic visualizations in HTML reports to aid variant review.

SUMMARY: We present igv-reports, a command-line tool to create standalone HTML pages embedding interactive genomic visualizations of read alignments and associated annotations to support variant inspection workflows. The reports contain all data and code required for visualization of the variant sites, with no dependencies on the input data files. AVAILABILITY AND IMPLEMENTATION: igv-reports is a command-line application written in Python. It is freely available at https://github.com/igvteam/igv-reports under an MIT license.

Software

NCBoost v2: a classifier for non-coding single-nucleotide variants in Mendelian diseases.

MOTIVATION: The current diagnostic rate of rare diseases through whole-genome sequencing has stabilized at around 30% on average, highlighting the need for improved computational scores to identify pathogenic variants. In 2019, we developed NCBoost, a supervised-learning approach that mined a comprehensive set of sequence constraint features and proved particularly well suited to identifying high-effect pathogenic non-coding variants in genetic diseases. Since its first release, the substantial increase in the number of variants available for training, as well as the enhanced capacity to detect purifying selection signals from large-scale genome sequencing projects, motivated an update of NCBoost. RESULTS: We implemented NCBoost v2, a pathogenicity score for non-coding single-nucleotide variants, trained on the largest set of curated pathogenic variants in monogenic Mendelian diseases available to date. It leverages conservation features computed from recent large-scale genomic consortia such as Zoonomia and gnomAD, and incorporates recent splice-altering predictive scores. NCBoost v2 outperformed alternative state-of-the-art methods in a variety of scenarii, providing more consistent scores across non-coding genomic regions and fine-tuning the scoring of pathogenic splice-altering variants in Mendelian disease genes. AVAILABILITY AND IMPLEMENTATION: NCBoost v2 software is implemented in Python 3.10 and is freely available under the GNU General Public License Version 3 at https://doi.org/10.5281/zenodo.16029049 and https://github.com/RausellLab/NCBoost-2, together with precomputed scores for the human genome assembly GRCh38.

Polymorphism, Single Nucleotide

gMISpy: integration of complex regulatory networks and genome scale metabolic models.

MOTIVATION: Genome-scale metabolic models lack explicit regulatory mechanisms, limiting their predictive accuracy for genetic interventions. Current methods for computing genetic Minimal Cut Sets either ignore regulatory networks entirely or use simplified acyclic representations that cannot capture regulatory feedback loops, ubiquitous features critical in cellular modeling. RESULTS: We developed gMISpy, a Python package that that enables efficient computation of genetic Minimal Intervention Sets (gMISs) in integrated genome-scale metabolic and regulatory networks. gMISpy incorporates cyclic regulatory logic into our previous computational framework using layered Boolean networks and BoNesis framework, resulting in a more accurate modeling of how regulatory interactions affect metabolic genes. Benchmarking across four different regulatory networks with Human-GEM showed consistent improvements in prediction accuracy, with Matthews correlation coefficient gains ranging from 2.50% to 14.42%. Validation against cancer data from DepMap and Project Score confirmed that cyclic integration reduces false positives and better captures biological vulnerabilities compared to acyclic approaches. AVAILABILITY AND IMPLEMENTATION: https://github.com/PlanesLab/cyclic-gMISpy.

Software

QCatch: a framework for quality control assessment and analysis of single-cell sequencing data.

MOTIVATION: Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstream results. While tools like alevin-fry and simpleaf (and augmented execution context for the alevin-fry), offer flexibility and computational efficiency to process single-cell data, this ecosystem will further benefit from a standardized QC reporting tailored for its outputs. RESULTS: We introduce QCatch, a Python-based command-line tool that generates comprehensive and interactive HTML QC reports designed specifically for single-cell quantification results. Taking the output directory of alevin-fry or simpleaf as the input, QCatch is able to perform essential processing steps, like cell calling, and generate detailed QC reports that contain informative visualizations and statistics, including unique molecular identifier (UMI) count distributions, sequencing saturation estimates, and splicing status information, for QC assurance. Built for seamless integration into downstream analysis workflows, QCatch exports the processed results in a richly-annotated H5AD format file, a widely used data format common among many downstream single-cell data analysis tools. AVAILABILITY AND IMPLEMENTATION: The source code and documentation of QCatch are available on GitHub at https://github.com/COMBINE-lab/QCatch. QCatch can be installed via both Bioconda and PyPI.

Single-Cell Analysis

Segzoo: a turnkey system that summarizes genome annotations.

MOTIVATION: Segmentation and automated genome annotation (SAGA) techniques, such as Segway and ChromHMM, assign labels to every part of the genome, identifying similar patterns across multiple genomic input signals. Inferring biological meaning in these patterns remains challenging. Doing so requires a time-consuming process of manually downloading reference data, running multiple analysis methods, and interpreting many individual results. RESULTS: To simplify these tasks, we developed the turnkey system Segzoo. As input, Segzoo only requires a genome annotation file in browser extensible data (BED) format. It automatically downloads the rest of the data required for comparisons. Segzoo performs analyses using these data and summarizes results in a single visualization. AVAILABILITY AND IMPLEMENTATION: The source code for Python ≥ 3.7 on Linux is freely available for download at https://github.com/hoffmangroup/segzoo under the GNU General Public License (GPL) version 2. Segzoo is also available in the Bioconda package segzoo: https://anaconda.org/bioconda/segzoo. We have deposited in Zenodo the version of the Segzoo source which produced the results in this article (https://doi.org/10.5281/zenodo.10988775), other code and data used to produce the results (https://doi.org/10.5281/zenodo.10477083), and the results (https://doi.org/10.5281/zenodo.10477106).

Software

LAML-Pro: joint maximum likelihood inference of cell genotypes and cell lineage trees.

MOTIVATION: Recent dynamic lineage tracing technologies use genome editing to induce heritable mutations, or edits, that accumulate across successive cell divisions. These edits are measured using single-cell sequencing or imaging, providing data to reconstruct cell lineages at single-cell resolution. Current computational approaches to infer cell lineage trees, or phylogenies, from these data perform two separate steps: (i) Identify each cell's edits (genotype) from the raw sequencing or imaging data; (ii) Infer a cell lineage tree from the cell genotypes. However, genotyping cells is an inexact process and genotype errors can yield an inaccurate lineage tree. For example, using fluorescence based-imaging to measure edits results in a high fraction (≈25%-50%) of uncertain or erroneous genotypes. RESULTS: We introduce Lineage Analysis via Maximum Likelihood with PRobabilistic Observations (LAML-Pro), an algorithm that jointly infers cell genotypes and a cell lineage tree. LAML-Pro is based on the Probabilistic Mixed-type Missing Observation (PMMO) model, which we derive to describe both the genome editing and genotype observation processes. LAML-Pro constructs lineage trees from thousands of cells in under an hour by leveraging the sparsity of transitions under the PMMO model. On simulated data, we demonstrate that LAML-Pro corrects genotype errors and infers substantially more accurate trees than existing methods which are vulnerable to genotype errors. Applied to data from two recent imaging-based lineage tracing systems, LAML-Pro reduces genotype errors by 5-fold and produces more spatially coherent lineage trees compared to existing methods. AVAILABILITY AND IMPLEMENTATION: LAML-Pro is implemented in C++ and is available as both a command-line interface and as a Python library at: github.com/raphael-group/LAML-Pro.

Cell Lineage