Search PubMedSearch

PubMed · 42635226

An end-to-end computational framework for "Record-seq" transcriptional recording data.

Abstract

MOTIVATION: Record-seq captures cumulative transcriptional activity over time in engineered Escherichia coli by integrating cellular RNA-derived spacer sequences into clustered regularly interspaced short palindromic repeats (CRISPR) arrays, which are read out by sequencing. Unlike the approximately uniform transcript sampling of RNA-seq, Record-seq records biological signal as spacers sampled by the CRISPR spacer acquisition machinery. Consequently, standard RNA-seq analysis strategies are not directly applicable, limiting sensitivity and interpretability. Our previous pipeline addressed these challenges only partially, retained inherited RNA-seq assumptions, and had limited algorithmic efficiency. RESULTS: Here, we present an end-to-end computational framework for Record-seq data. To address the primary computational bottleneck of spacer sequence extraction, we implemented a wavefront alignment approach for efficient quasi-local pattern matching, achieving an approximately 30-fold speedup. We introduce transcription unit-based feature counting as an alternative to gene-body quantification to better represent prokaryotic transcription and increase statistical power by capturing signal from untranslated regions, which are spacer acquisition hotspots. For downstream analyses, we incorporate multiple normalization strategies and a nonparametric differential expression testing framework designed for sparse datasets. Further, we analyze spacer acquisition patterns and train sequence-based neural models that predict acquisition propensity from genomic sequence and annotations, providing a framework for assessing whether acquisition rules generalize as Record-seq is extended to new microbial hosts. AVAILABILITY AND IMPLEMENTATION: The primary analysis workflow, the recoRdseq package, acquisition modeling repository, and relevant data are all linked at https://github.com/plattlab/Record-seq-Framework. Acquisition models and training data are on Zenodo at https://doi.org/10.5281/zenodo.18891434.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Florian Hugi, Tanmay Tanna, Randall J Platt. 2026-08-01. An end-to-end computational framework for "Record-seq" transcriptional recording data.. https://doi.org/10.1093/bioinformatics%2Fbtag479

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Whole-genome surveillance supports hazard profiling of Escherichia coli lineages in recycled water treatment systems.

UNLABELLED: The use of treated wastewater is increasingly important for sustainable water management under a changing climate, yet conventional monitoring based on Escherichia coli enumeration provides limited insight into strain diversity and associated public health hazards. Here, we applied longitudinal whole-genome sequencing (WGS) to 180 E. coli isolates collected across the treatment continuum of a recycled water facility, from influent to final effluent. Genomic analysis revealed extensive strain-level heterogeneity, comprising 88 sequence types across eight phylogroups, with greater diversity in influent than in treated effluent. Phylogenetic comparisons with contextual Australian genomes indicated clustering with strains associated with companion animals, wild birds, humans, and livestock, suggesting multiple potential source reservoirs rather than a single dominant origin, although source contributions were not definitive. Despite a >90% reduction in total E. coli loads, isolates recovered from upstream and downstream stages exhibited broadly comparable virulence factor and antimicrobial resistance gene (ARG) profiles, suggesting that, within the cultured isolate collection, reductions in abundance exceeded shifts in genomic composition. To assess operational relevance, we prototyped a genomics-informed hazard framework integrating virulence determinants, ARGs, plasmid-associated mobility, and reuse-specific exposure context. Using this framework, 92.8% of isolates were classified as low hazard, and 7.2% as moderate hazard, with no isolates meeting criteria for high or critical hazard classifications. These findings demonstrate that genomic profiling of indicator organisms can reveal population structure and hazard heterogeneity not captured by conventional enumeration alone, and can provide a practical basis for incorporating genomic information into hazard-informed monitoring of recycled water systems. IMPORTANCE: Routine recycled water monitoring relies largely on culture-based E. coli counts, which indicate regulatory compliance but provide limited insight into strain diversity, persistence, and genomic characteristics relevant to public health. Using longitudinal whole-genome sequencing, we show that genetically distinct E. coli lineages, including isolates carrying combinations of virulence and antimicrobial resistance determinants, can persist through advanced treatment despite substantial reductions in overall E. coli loads. While most isolates were classified as low genomic hazard and no high- or critical-hazard isolates were detected, these findings demonstrate that conventional enumeration alone cannot distinguish between genetically diverse lineages with differing hazard potential in highly treated systems. By integrating genomic data into a hazard classification framework, this study demonstrates an applied approach to contextualize E. coli detections and distinguish low-risk background populations from isolates with elevated genomic hazard profiles. This work supports the use of genomic profiling of indicator organisms to improve surveillance, inform treatment performance assessment, and enable more risk-based management of recycled water systems.

Escherichia coli

Global Diffusion of IncC Plasmid Harboring blaNDM-1in the High-Risk Escherichia coli ST131 Clone.

AIMS: The global expansion of quinolone-resistant Escherichia coli (QR-EC) is increasingly associated with β-lactam resistance and mobile genetic elements that facilitate resistance dissemination. This study investigated the molecular mechanisms underlying fluoroquinolone and β-lactam resistance in clinical QR-EC isolates and explored the plasmid type associated. METHODS AND RESULTS: A total of 123 non-duplicate QR-EC clinical isolates responsible mainly for gastrointestinal colonization were collected between 2019 and 2021. Plasmid-mediated quinolone resistance (PMQR), extended-spectrum β-lactamase (ESBL), and carbapenemase genes were screened by PCR. Mutations in the quinolone resistance-determining regions (QRDR) of gyrA and parC were analyzed using sequencing and mismatch amplification mutation assay PCR (MAMA-PCR). Selected isolates underwent multilocus sequence typing (MLST). Whole-genome sequencing (WGS) of a representative extensively drug-resistant strain carrying multiple quinolone resistance determinants, ESBL genes, and carbapenemase genes, was performed. PMQR genes were prevalent among QR-EC, dominated by aac(6')-Ib-cr (60.9% of isolates). ESBL genes were identified in 93.5% of isolates, predominantly blaCTX-M (95.7%). Among ertapenem-resistant isolates (QCR-EC) (n=18), blaNDM-1 and blaOXA-48 were detected in 13 and 11 isolates, respectively. QRDR mutations were highly frequent, particularly gyrA83 (98.4%) and parC80 (30.9%). Major QCR-EC genotypes belonged to sequence types ST167 (n=2), ST1196, ST469, and ST410. High-risk E. coli ST131 clone harboring IncC plasmid encoding blaNDM-1 was described for the first time in Africa, following its emergence, in two continents, Asia and America. Despite the very rare description of these strains worldwide, their description in three continents sign their global diffusion. CONCLUSIONS: This finding highlights the ongoing spread of carbapenem resistance and underscores the urgent need for strengthened genomic surveillance.

Escherichia coli

A Draft Map of E. coli Proteoforms.

Top-down proteomics (TDP) enables direct characterization of intact proteoforms, providing protein-level insights into molecular diversity arising from post-translational modifications and sequence variations. Despite this advantage, proteome coverage in TDP remains limited relative to bottom-up proteomics (BUP). To expand coverage, we developed an integrated multidimensional approach combining sequential protein extraction, size-exclusion chromatography (SEC) fractionation, and capillary zone electrophoresis (CZE)-tandem mass spectrometry (MS/MS) and reversed-phase liquid chromatography (RPLC)-MS/MS. This approach identified 743 proteoform families and 10,613 proteoforms from E. coli cells through hundreds of MS runs. By incorporating previous E. coli TDP data sets from our group, we identified 14,932 proteoforms from 985 proteoform families, covering 43% of the E. coli proteome. The data represent the highest proteome coverage of cells by MS-based TDP, creating a draft map of E. coli proteoforms. The results offer strong evidence that MS-based TDP can reach high proteome coverage.

Escherichia coli