Search PubMedSearch

SEARCH · Search PubMed

Results for “Python”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

AEGIS: an annotation extraction and genomic integration resource.

MOTIVATION: Genome annotation files (GFF3/GTF) are the standard for storing genomic feature data, yet their flexibility often results in formatting inconsistencies that create bottlenecks for downstream bioinformatics analyses. A robust, unified framework is required to parse, standardise, and validate these files to ensure interoperability and facilitate complex comparative genomic tasks. RESULTS: We present AEGIS (Annotation Extraction and Genomic Integration Suite), a comprehensive toolkit designed to parse, correct, and standardise genome annotations. Beyond quality control, AEGIS provides advanced modules for flexible feature extraction (e.g., coding sequences, promoters) and comparative genomic analysis. Uniquely, it integrates multiple lines of evidence, including sequence homology, synteny, and coordinate-based lift-overs, to assess gene model correspondence and infer orthology. We demonstrate the utility of AEGIS by quantifying complex structural changes between Arabidopsis annotation versions and identifying high-confidence orthologues across diverse plant genomes. AVAILABILITY: AEGIS is implemented in Python. Source code and documentation are freely available under the GPL-3 license at https://github.com/Tomsbiolab/aegis and as a Docker container at https://hub.docker.com/r/tomsbiolab/aegis. The package is also available on PyPI (pip install aegis-bio).

Software

Causal circuit tracing reveals distinct computational architectures in single-cell foundation models: inhibitory dominance, biological coherence, and cross-model convergence.

MOTIVATION: Sparse autoencoders (SAEs) decompose foundation-model activations into interpretable features, but the model-internal causal interactions between those features (i.e. what ablating one feature does to the others, as distinct from the biological causal structure of the underlying cells)-and how those model-internal relationships relate to biological structure-are uncharacterized in single-cell foundation models. RESULTS: We introduce model-internal causal circuit tracing-zeroing one SAE feature at a source layer and measuring the resulting change in all downstream SAE features, for each of 120 source features-and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96&#xa0;892 ablation-derived edges, 80&#xa0;191 forward passes). On annotation-selected source features, edges share GO/KEGG/Reactome/STRING/TRRUST ontology terms at 50.9%-68.5%, a 2.9-6.2&#xd7; enrichment over a configuration-preserving permutation null (P<.002); on 20 randomly sampled source features this attenuates to 21.5%-26.3%-still 2.5-3.1&#xd7; above null-quantifying the annotation-selection contribution. Inhibitory dominance (fraction of ablation edges with d<0, i.e. source activation supports downstream target) is 65.5%-89.4%. scGPT produces larger raw per-edge effects (mean |d|=1.40 versus 1.05); after feature-share normalization, Geneformer is stronger (paired gene-pair ratio 0.64 on 33&#xa0;301 shared pairs). Cross-model consensus yields 1142 architecture-invariant domain pairs (ordered pairs of GO biological-process categories "A&#x2192;B" each connected by at least one ablation edge in both models; 10.6&#xd7; enrichment over permutation null; P<.001). Circuit edge magnitude explains <1% of the variance in marginal driver-gene coexpression on the same cells (R2=0.010, n=31&#xa0;176): the graph encodes structure beyond bivariate correlation. Against a matched-cell-type ENCODE ChIP-seq prior, circuit-predicted transcription factor (TF)&#x2192;target pairs are enriched 2.06&#xd7; (Fisher OR 5.84), markedly higher than 1.12&#xd7; against TRRUST; direct ChIP-seq-supported target pairs show 10-30&#xd7; larger CRISPRi sign-bias-corrected excess than indirect pairs. Gene-level CRISPRi validation on Replogle K562 and the noncancer RPE1 arm (and a true primary-T-cell control from Shifrut E, Carnevale J, Tobin V et&#xa0;al. Genome-wide CRISPR screens in primary human T cells reveal key regulators of immune function. Cell 2018; 175: 1958-71.e15) after sign-bias correction shows excess over baseline of +0.03 and +0.35 percentage points on K562 and RPE1, respectively (baseline already 52%-56% from sign marginals); effect-magnitude Spearman correlations &#x3c1;&#x2248;0. Bootstrap and per-cell-type stability (N&#x2208;{50,100,200}; B cell, CD4&#xa0;+ T, macrophage) give Pearson r&#x2265;0.97 on shared edges with 100% sign agreement; edge Jaccard grows monotonically with sample size. The circuit graph is therefore highly reproducible as an effect-size map, cell type specific in edge identity, consistent with coexpression encoding, and weakly but detectably enriched for ChIP-seq-supported direct regulatory edges. AVAILABILITY AND IMPLEMENTATION: https://github.com/Biodyn-AI/bio-sae-circuits (Python). Archival DOI: 10.5281/zenodo.19,633,166 (Zenodo).

Humans

VIJB: a companion of the JBROWSE genome browser for the visually impaired people.

MOTIVATION: The availability of touch-sensitive and haptic devices has been a keystone development for the inclusion of visually impaired people (VIPs) in modern, highly digitized work environments. Braille displays have proven efficient and versatile enough to parse large and complex text files, making bioinformatics and text-heavy programming accessible to VIPs. However, the complex graphical objects -combining numerous datasets- typically generated during data integration remain challenging, even with the aid of descriptive AI. This is particularly true in functional genomics. Here, we present VIJB, a simple application that displays the multilayered output of the JBROWSE genome browser on a Braille reader, enabling VIPs to fully participate in data integration in functional genomics. AVAILABILITY AND IMPLEMENTATION: VIJB is programmed in Python and relies on the scientific library NumPy, the braillegraph and pyBigWig libraries, and the TABIX software. The architecture is summarized in Supplementary Material 1, available as supplementary data at Bioinformatics online. VIJB is available for download at the GitHub repository https://GitHub.com/NiBuMNHN/VIJB and is licenced under the GPL 3.0.

Persons with Visual Disabilities

ChromBERT-tools: a versatile toolkit for context-specific regulatory representations of transcription regulators across different cell types.

SUMMARY: Representations that encode the genome-wide regulatory behavior of transcription regulators provide a foundation for flexible transcription modeling and in silico regulatory analysis. Existing regulator representations are commonly derived from gene co-expression, motif annotations, or static protein features, which capture useful but limited aspects of regulator identity but do not directly model how regulators participate in region-specific regulatory programs across the genome. ChromBERT addresses this gap by learning context-aware regulatory representations from large-scale ChIP-seq data. However, routine bioinformatics applications require lightweight, accessible, and modular tools for generating, adapting, and interpreting these representations in user-defined biological contexts. Here, we present ChromBERT-tools, a user-oriented toolkit built upon ChromBERT that converts its regulatory representation framework into practical workflows for customizable analysis across cellular contexts. ChromBERT-tools provides command-line interfaces and Python APIs organized into three functional layers: representation generation, predictive modeling, and regulatory interpretation. The representation generation layer produces representations of genomic regions and transcription regulators. The predictive modeling layer fine-tunes ChromBERT for genome-wide regulatory activity prediction through classification or regression tasks, with optimized implementation to reduce running time and computational resource requirements. The regulatory interpretation layer supports inference of the context-specific roles of cis-regulatory elements and transcription regulators. These modules can be used independently or integrated into end-to-end workflows, enabling flexible analyses across diverse datasets. ChromBERT-tools lowers the barrier to applying context-specific regulatory representations in routine genomic analyses. AVAILABILITY AND IMPLEMENTATION: ChromBERT-tools is freely available at https://github.com/TongjiZhanglab/ChromBERT-tools, with documentation at https://chrombert-tools.readthedocs.io/en/latest/. A frozen archival snapshot is available on Zenodo under DOI: 10.5281/zenodo.20094206.

Software

Interactive exploration of biobank-scale ancestral recombination graphs with Lorax.

MOTIVATION: Ancestral Recombination Graphs (ARGs) provide a comprehensive representation of genetic ancestry and underpin analyses of natural selection, disease association, and population history. However, existing visualization tools are limited in scalability and interactivity, making ARGs difficult to explore at biobank scale. RESULTS: We introduce Lorax, a GPU-accelerated, web-native platform for real-time visualization of population-scale ARGs. Lorax integrates genomic position, coalescent time, local genealogy, and metadata, enabling interactive exploration of ancestry and variant inheritance in biobank-scale datasets. AVAILABILITY AND IMPLEMENTATION: Lorax is freely available as a live demo at https://lorax.ucsc.edu/ and as a Python package "lorax-arg" on PyPI. The source code and documentation are available on GitHub at https://github.com/pratikkatte/lorax.

Software

An interpretable deep learning framework uncovers features governing CRISPR-Cas9 genome-editing efficiency.

MOTIVATION: CRISPR-Cas9 genome-editing efficiency is strongly influenced by the sequence composition and positional context of single-guide RNAs (sgRNAs). Although numerous deep learning-based models have been developed to predict Cas9 efficiency from sgRNA sequences, most operate as black boxes, offering limited insight into the sequence determinants underlying Cas9 activity. In addition, previous studies often overlook how the positional context of sequence motifs within sgRNAs influences their effects on Cas9 binding or cleavage. RESULTS: We introduce DeepCC9, an interpretable machine learning framework that combines explicit sequence feature extraction with a residual block-based deep architecture to improve interpretability and identify composition- and position-based motifs governing Cas9 genome-editing efficiency. We applied this method to multiple Cas9 variant datasets, achieving superior predictive performance compared with existing methods while enabling direct interpretation of sequence motifs and their positional effects. Our analysis uncovered 74 sequence motifs enriched or depleted at specific positions within sgRNAs and strongly associated with Cas9 efficiency, providing mechanistic insight into sequence features that influence guide performance. Together, these results establish DeepCC9 as a generalizable and interpretable framework for modeling sequence-function relationships and advancing the understanding of the sequence determinants underlying CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The authors have implemented their algorithm in the Python programming language (version 3.X), which is accessible using (https://zenodo.org/records/20073890).

Deep Learning

ReGAIN: a bioinformatics platform for assessing probabilistic co-occurrence between resistance genes in bacterial pathogens.

MOTIVATION: Multidrug-resistant bacterial pathogens continue to rise globally, yet scalable methods are needed to infer how resistance determinants co-occur across pathogen populations and to quantify conditional dependencies underlying co-occurrence and shared genetic context. RESULTS: We present ReGAIN (Resistance Gene Association and Inference Network), an open-source platform that applies Bayesian network structure learning to infer probabilistic, conditional dependency relationships among antibiotic resistance, heavy metal tolerance, stress response, and virulence determinants in bacteria. In contrast to pairwise co-occurrence analyses, ReGAIN reports conditional probabilities, relative risks, and absolute risk differences with confidence intervals to prioritize candidate relationships for downstream prioritization. Applied across ESKAPEE pathogens, ReGAIN recapitulated established resistance gene relationships and identified additional candidate patterns consistent with co-selection and shared genetic context. Together, these results support scalable, reproducible population-wide analysis of resistance networks for surveillance, comparative genomics and epidemiology. AVAILABILITY: ReGAIN analyses are performed using Python v3.11.5 and R v4.4.1 and is available as open-source software through Bioconda at {https://anaconda.org/bioconda/regain-cli}. Source code and documentation can be found at {https://github.com/ERBringHorvath/regain_CLI}. All genomes used in this publication were downloaded from the National Center for Biotechnology Information database. Large supplementary tables and results data from the ESKAPEE pathogen example network analyses can be downloaded from https://figshare.com/articles/dataset/ReGAIN_command_line_software_and_supplemental_figures_/28959431.

Computational Biology

WILDkCAT: extract, retrieve, and predict enzyme turnover numbers of constraint-based metabolic models.

SUMMARY: Accurate enzyme turnover numbers are essential for building enzyme-constrained genome-scale metabolic models. However, collecting and curating these parameters remains a major bottleneck. Indeed, kcat values are scattered across multiple databases, reported under varying experimental conditions, and often missing for many enzymes. To address this challenge, we present WILDkCAT, a Python-based pipeline that enables the retrieval of kcat values from wild-type enzyme measured under user-specified pH and temperature ranges for a given metabolic model. The application to Escherichia coli (iML1515) and Homo sapiens (Human-GEM) models demonstrated the ability of WILDkCAT to retrieve substantial kcat coverage and its applicability across diverse genome-scale models. AVAILABILITY AND IMPLEMENTATION: WILDkCAT is available at https://github.com/sysbiolux/WILDkCAT and from PyPI. WILDkCAT works on all major operating systems and computer architectures. The documentation is available at https://sysbiolux.github.io/WILDkCAT.

Software

Making multi-axis Gaussian graphical models scalable to millions of cells.

MOTIVATION: Networks underlie the generation and interpretation of many biological datasets: gene networks shed light on the regulatory structure of the genome, and cell networks can capture structure of the tumor micro-environment. However, most methods that learn such networks make the faulty "independence assumption"; to learn the gene network, they assume that no cell network exists. "Multi-axis" methods, which do not make this assumption, fail to scale beyond a few thousand cells or genes. This limits their applicability to only the smallest datasets. RESULTS: We develop a multi-axis method, which learns conditional dependency networks, capable of processing million-cell datasets within minutes. This was previously impossible, and unlocks the use of such methods on modern scRNA-seq datasets, as well as more complex datasets. We apply the method to a new scRNA-seq dataset for neuronal cell development, and compare the result to an existing state of the art method, hdWGCNA. We demonstrate that the new method yields gene networks that have a more focused biological interpretation and that the simultaneously learned cell network has advantages over a conventional kNN-based clustering. Further, our method yields novel biological insights by identifying long non-coding RNAs that potentially have a role in neuronal development. AVAILABILITY AND IMPLEMENTATION: Our methodology is available as a Python package GmGM on PyPI (https://pypi.org/project/GmGM/0.5.3/). The code for all experiments performed in this article is available on GitHub (https://github.com/BaileyAndrew/GmGM-Bioinformatics) and Zenodo (10.5281/zenodo.20384566).

Gene Regulatory Networks

Quantifying uncertainty of predictions from cancer progression models.

MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).

Humans

WxS-QC-a quality control pipeline for human germline short-variant Whole-Genome and Whole-Exome cohorts for population-scale analyses.

SUMMARY: Whole-exome (WES) and whole-genome (WGS) sequencing are rapidly becoming preferred methods for population-scale analysis of the human genetic landscape. However, there are currently no standardized quality control (QC) pipelines for human WES and WGS datasets. In this paper, we present WxS-QC, a powerful, scalable, and convenient pipeline for the QC of human germline short-variant WGS and WES cohorts for population-scale analyses. Our pipeline is suitable for both rare-variant discovery and common-variant association studies. It is based on deeply refactored gnomAD v3 and v4 quality control pipelines, contains several methods we have developed de novo, and is aligned with current best practices in WGS/WES germline cohort QC. We provide all methods in a single codebase, aligned to work together and controlled via a single YAML config, with automatic export of resulting graphs and summary tables, excellent performance and scalability, and comprehensive documentation. The pipeline can run in any UNIX-like environment and can efficiently process cohorts of up to 200&#x2009;000 whole-exome samples, with the potential to handle bigger datasets. AVAILABILITY AND IMPLEMENTATION: The pipeline code is written in Python using the Hail library and is freely available under the BSD-3 license here: https://github.com/wtsi-hgi/wxs-qc. The detailed description of the pipeline is available in the pipeline documentation: https://github.com/wtsi-hgi/wxs-qc/blob/main/README.md. We also provide an open dataset with all required metadata, which is available at https://wxs-qc-data.cog.sanger.ac.uk/wxs-qc_public_dataset_v3.tar. An example of test dataset analysis is available in the supplementary materials.

Humans

pLAST-a tool for rapid comparison and classification of bacterial plasmid sequences.

MOTIVATION: The increasing number of fully sequenced bacterial plasmids being annotated and catalogued has prompted the development of computational tools for comparing and classifying them. Existing approaches typically compare full-length DNA sequences (e.g. Mash, BLASTn, and ANI-based methods) or translated open reading frames (ORFs) (e.g. DIAMOND), with plasmid-level scores obtained by aggregating ORF-to-ORF similarities; however, they are either restricted to closely related plasmids or become computationally demanding in large-scale analyses. RESULTS: We describe pLAST (plasmid Language Analysis and Search Tool), a plasmid-search tool built using word2vec representations of protein-family content informed by local genomic context. Benchmarks indicate that pLAST outperforms nucleotide-based methods and performs comparably to DIAMOND in identifying functionally similar plasmids and compared with the widely used Mash, it achieves 26% and 24% improvements in detecting shared mating-pair formation system type and relaxase type, respectively. This performance scales to database searches across hundreds of thousands of sequences, as demonstrated using the precomputed PlasmidScope collection of &#x223c;750&#xa0;000 plasmids. Beyond global similarity, pLAST also returns per-ORF plasmid-plasmid alignments, enabling detection of shared functional modules. AVAILABILITY AND IMPLEMENTATION: pLAST is freely accessible as a web server at&#x202f;https://plast.lbs.cent.uw.edu.pl/ or https://plast.lbs.biol.uw.edu.pl/ and available as a Python module along with a precomputed database at&#x202f;https://github.com/labstructbioinf/pLAST for customized analysis.

Plasmids

Zone equalisation normalisation for improved alignment of epigenetic signal.

MOTIVATION: High-throughput genomic technologies have transformed our understanding of biological systems, yet direct comparison and visualisation of these complex datasets remains challenging. Existing normalisation methods often fail to align genomic signal across samples due to sensitivity to sequencing depth differences and localised high-signal artefacts, leading to inconsistent replicate behaviour and increased downstream variability. RESULTS: We introduce Zone Equalisation Normalisation (ZEN), a novel approach designed to improve cross-sample signal alignment of genomic data. ZEN rescales genomic signal based on variance estimated within biologically enriched regions, reducing the influence of extreme outliers while preserving underlying biological structure. Using a diverse collection of data and our new genome-wide benchmarking approach, we reveal that ZEN improves biological and technical replicate alignment across the majority of tested conditions and experimental platforms. We further show that this improved signal comparability is associated with fewer differential accessibility calls between technical replicates and a more conservative set of biological differences. Together, these results demonstrate that ZEN provides a complementary framework to improve the accuracy and reliability of genomic data analysis and that normalisation choice can affect downstream analyses and biological interpretation. AVAILABILITY AND IMPLEMENTATION: ZEN is available as an open-source Python package via conda and PyPI. Source code, documentation, tutorials, and code to reproduce the analyses are available at https://github.com/Genome-Function-Initiative-Oxford/Zone-Equalisation-Normalisation and Zenodo (https://doi.org/10.5281/zenodo.21067751).

Epigenesis, Genetic

Pesci: fast and user-friendly software to compare single-cell gene expression across species.

SUMMARY: Recent technological advances have propelled comparative functional genomics into the single-cell era, spurring a rapid development of methods to analyse these complex datasets. However, comparing single-cell gene expression across species to quantify expression similarity and ultimately identify homologous cell types remains an open problem. The ICC algorithm (Iterative Correlation of Coexpression) has been recently proposed as an attractive approach to tackle this challenge, but, to date, no software implementation is available. Here, we introduce Pesci (Pretty Easy Single-cell Comparisons using ICC), an efficient and user-friendly implementation of the ICC algorithm applied to pairwise comparisons of single-cell gene expression atlases across species. AVAILABILITY: Pesci is implemented in Python 3 (&#x2265;3.7). It is available for download on Linux, macOS and Windows via pip, conda and GitHub at https://github.com/eparey/pesci. The source code is permanently archived on Zenodo (https://doi.org/10.5281/zenodo.21477543).

Software

Global maintenance of histone post-translational modifications during the transition into anoxia in embryos of the annual killifish Austrofundulus limnaeus.

Many organisms have adapted to survive anoxic or hypoxic environments, but the epigenetic responses involved in this successful stress response are not well described in most species. Embryos of the annual killifish Austrofundulus limnaeus have the greatest tolerance to anoxia of all vertebrates, making them a powerful model to study the cellular mechanisms necessary for anoxia tolerance. However, the global histone landscape of this species has never been quantified or explored in relation to stress tolerance. Liquid chromatography-mass spectrometry and a Python bioinformatics workflow were used to identify histones and their post-translational modifications. This pipeline resulted in the detection of 252 unique biologically relevant histone post-translational modifications (hPTMs) (unimod&#xa0;+&#xa0;residue). These PTMs represent 16 types of biologically relevant hPTMs present during both anoxia and normoxia in Wourms' stage 36 embryos. This hPTM library presents an exciting opportunity to study histone modifications across development and in response to environmental stressors. No significant changes in PTM or histone abundance were observed between anoxic and normoxic embryos, suggesting that 24 h of anoxia is not sufficient to induce epigenetic or histone isoform changes at the organismal level. This result is inconsistent with data presented for similar stresses in mammalian cells and thus stabilization of the hPTM landscape may be an adaptation that supports anoxia tolerance.

anoxia

scATAnno: Automated Cell Type Annotation for Single-cell ATAC-seq Data.

Recent advances in single-cell epigenomic techniques have increased the demand for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) analysis. One key analytical task is to determine cell type identity based on epigenetic data. Here, we introduce scATAnno, a Python package designed to automatically annotate scATAC-seq data using large-scale scATAC-seq reference atlases. This workflow generates reference atlases from publicly available datasets, enabling accurate cell type annotation by integrating query data with reference atlases without the use of single-cell RNA sequencing (scRNA-seq) data. To enhance annotation accuracy, we incorporated k-nearest neighbors (KNN)-based and weighted distance-based uncertainty scores to effectively detect cell populations within the query data that are distinct from all cell types in the reference data. We compared and benchmarked scATAnno against five other published cell annotation approaches, demonstrating its superior performance across multiple datasets and metrics. We further showcased the utility of scATAnno across multiple datasets, including peripheral blood mononuclear cells (PBMCs), triple-negative breast cancer (TNBC), and basal cell carcinoma (BCC), and demonstrated that scATAnno accurately annotates cell types across diverse biological conditions. Overall, scATAnno is a useful tool for scATAC-seq reference atlas construction and cell type annotation and can facilitate the interpretation of new scATAC-seq datasets in complex biological systems. scATAnno is publicly available at https://scatanno-main.readthedocs.io/.

Single-Cell Analysis

MACS3: A Peak-calling Platform for Bulk and Single-cell Regulatory Genomics.

Since the original publication of Model-based Analysis for ChIP-Seq (MACS), the software has been widely used to identify enriched genomic regions in ChIP-seq, ATAC-seq, CUT&RUN, DNase-seq, and related regulatory genomics assays. Over the years, MACS has evolved substantially, with MACS version 3 (MACS3) now serving as the actively maintained implementation. MACS3 preserves the core MACS framework for fragment pileup, dynamic local background noise, statistical enrichment testing, and peak refinement, while adding functionality needed for contemporary bulk and single-cell workflows. It supports conventional bulk peak calling, paired-end and fragment-based file formats, modular signal processing, direct analysis of single-cell ATAC-seq fragment files, barcode-restricted pseudobulk and cluster-level peak calling, specialized ATAC-seq and variant-calling modules, as well as command-line and programmatic interfaces. MACS3 is distributed through standard software channels and supported by continuous testing across operating systems, Python versions, and CPU architectures. Here we describe the architecture, current capabilities, and recommended use of MACS3, providing an updated reference for applying the MACS framework in contemporary bulk and single-cell regulatory genomics workflows. MACS3 is open-source software available at https://github.com/macs3-project/MACS.

Bioinformatics software

CoDIAC: A comprehensive approach for interaction analysis reveals novel insights into SH2 domain function and regulation.

Protein domains are conserved structural and functional units that serve as building blocks of proteins. Through evolutionary expansion, domain families are represented by multiple members in diverse configurations with other domains, evolving new specificities for their interacting partners. Here, we develop a structure-based interface analysis to comprehensively map domain interfaces from experimental and predicted structures, including interfaces with macromolecules and intraprotein interfaces. We hypothesized that comprehensive contact mapping of domains could yield new insights into domain selectivity, conservation of domain-domain interfaces across proteins, and identify conserved post-translational modifications (PTMs), relative to interaction interfaces, allowing for the inference of specific effects due to PTMs or mutations. We applied this approach to the human SH2 domain family, a modular unit central to phosphotyrosine-mediated signaling, identifying a novel approach to understanding binding selectivity and evidence of coordinated regulation of SH2 domain binding interfaces by tyrosine and serine/threonine phosphorylation and acetylation. These findings suggest multiple signaling systems can regulate protein activity and SH2 domain interactions in a coordinated manner. We provide the extensive features of the human SH2 domain family and this modular approach as an open source Python package for COmprehensive Domain Interface Analysis of Contacts (CoDIAC).

SH2 domains