Search PubMedSearch

SEARCH · Search PubMed

Results for “single-molecule genomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A practical guide to studying genome function using single-molecule genomics.

Single-molecule genomics (SMG) has transformed our ability to study the mechanisms that regulate the genome by enabling profiling of the activity of regulatory factors on individual DNA molecules genome-wide. SMG is able to quantify molecular heterogeneity and the co-occurrence of regulatory events, including epigenetic modifications, transcription factor binding and chromatin organization on single DNA molecules. SMG reveals dynamics of chromatin interactions that cannot be measured by conventional genomics assays. Therefore, SMG offers a unique platform to study how regulatory events combine to control genome activity. In this Expert Recommendation article, we provide a practical guide for adopting SMG and outline best practices.

Journal Article

nf-core/pacvar: a pipeline for analyzing long-read PacBio whole genome and repeat expansion sequencing data.

MOTIVATION: Pacific Biosciences (PacBio) single-molecule, long-read sequencing enables whole genome annotation and the characterization of 20 complex repetitive repeat regions, especially relevant to neurodegenerative diseases, through their PureTarget panel. Long-read whole-genome sequencing (WGS) also allows for the detection of structural variants that would be difficult to detect with traditional short-read sequencing. However, the raw unaligned Binary Alignment Map data need to be processed before analysis. There is a need for an intuitive comprehensive bioinformatic pipeline that can analyze these data. RESULTS: We present nf-core/pacvar, a comprehensive pipeline for analyzing both PacBio single-molecule PureTarget and WGS data that demultiplexes and parallelizes pre-processing, variant calling and repeat characterization. nf-core/pacvar is compatible with little configuration and has few dependencies. This pipeline enables rapid end-to-end, parallel processing of PacBio single-molecule whole genome and targeted repeat expansion sequencing. AVAILABILITY AND IMPLEMENTATION: nf-core/pacvar is available on nf-core website (https://nf-co.re/pacvar/) and on github (https://github.com/nf-core/pacvar) under MIT License (DOI: 10.5281/zenodo.14813048).

Software

NusG-Spt5 Transcription Factors: Universal, Dynamic Modulators of Gene Expression.

The accurate and efficient biogenesis of RNA by cellular RNA polymerase (RNAP) requires accessory factors that regulate the initiation, elongation, and termination of transcription. Of the many discovered to date, the elongation regulator NusG-Spt5 is the only universally conserved transcription factor. With orthologs and paralogs found in all three domains of life, this ubiquity underscores their ancient and essential regulatory functions. NusG-Spt5 proteins evolved to maintain a similar binding interface to RNAP through contacts of the NusG N-terminal domain (NGN) that bridge the main DNA-binding cleft. We propose that varying strength of these contacts, modulated by tethering interactions, either decrease transcriptional pausing by smoothing the rugged thermodynamic landscape of transcript elongation or enhance pausing, depending on which conformation of RNAP is stabilized by NGN contacts. NusG-Spt5 contains one (in bacteria and archaea) or more (in eukaryotes) C-terminal domains that use a KOW fold to contact diverse targets, tether the NGN, and control RNA biogenesis. Recent work highlights these diverse functions in different organisms. Some bacteria contain multiple specialized NusG paralogs that regulate subsets of operons via sequence-specific targeting, controlling production of antibiotics, toxins, or capsule proteins. Despite their common origin, NusG orthologs can differ in their target selection, interacting partners, and effects on RNA synthesis. We describe the current understanding of NusG-Spt5 structure, interactions with RNAP and other regulators, and cellular functions including significant recent progress from genome-wide analyses, single-molecule visualization, and cryo-EM. The recent findings highlight the remarkable diversity of function among these structurally conserved proteins.

Archaea

Molecular co-accessibility identifies coordinated regulation between distant cis-regulatory elements.

In metazoans, gene expression is typically regulated by a cis-regulatory landscape (CRL) composed of a promoter and multiple enhancers. How these cis-regulatory elements (CREs) coordinate their function across large genomic distances remains unclear. For example, is the simultaneous activation of multiple enhancers required to promote transcription? Here, we combined single-molecule footprinting with long-read sequencing to quantify how often chromatin accessibility and transcription factor binding co-occur across entire CRLs in the Drosophila genome. Analysis of thousands of individual DNA molecules at each locus revealed that CREs form a specific network with shared single-molecule chromatin accessibility profiles. Co-accessibility is not limited to adjacent CREs and is frequently observed between CREs brought into proximity by chromatin looping. Co-accessible CREs exhibit strong coordination in their cell-type-specific accessibility, linking enhancer activity with transcriptional activation. Our data uncover dependencies between CREs genome-wide and suggest that coordinated enhancer activation is a widespread mechanism regulating gene expression.

Animals

Whole-Genome Analysis of Bacillus Licheniformis Ali5 and Synthesis of Lichenysin via Genome Shuffling.

Whole-genome sequencing of Bacillus licheniformis Ali5 was performed via MGI-seq PE150 and Nanopore single-molecule real-time sequencing. The strain has a 4,114,664 bp circular genome encoding 4030 protein-coding genes. Functional annotation across NR, COG, GO, KEGG, CARD, BacMet, and CAZy databases identified 4025, 2812, 988, 1242, 72, 69, and 94 corresponding genes, respectively, and antiSMASH 6.0 revealed multiple antimicrobial biosynthetic gene clusters, including intact lichenysin and lichenicidin VK21 A1/A2 gene clusters. Three rounds of recursive protoplast fusion-based genome shuffling, paired with a dual-index screening system, significantly improved strain growth and lichenysin biosynthesis. Recombinants exhibited shortened lag phase, enhanced proliferation, improved stationary-phase stability, and higher diauxic peak biomass. PP3-176 and PP3-186 showed 4.6%-8.1% higher 12-h shake-flask titer and 3.1%-4.0% higher maximum titer than the parental average, with excellent fermentation stability. 1-L bioreactor validation confirmed strong scale-up potential. PP3-186 achieved 27.2% and 31.6% titer increases at 12 h and 20 h, while PP3-176 yielded 20.4% and 14.6% improvements with robust metabolic performance. This study validates genome shuffling as an effective strategy for enhancing lichenysin production, providing candidate strains and technical support for industrial application.

Bacillus licheniformis

Promises and pitfalls of long-read sequencing for resolving microbial complexity.

Long-read sequencing (LRS) has driven a transition in microbial genomics, overcoming the assembly fragmentation inherent to short-read sequencing. This review elucidates the impact of LRS across isolate genomics, metagenomics, and multi-omics domains. By spanning extensive repetitive regions, LRS facilitates the reconstruction of circular chromosomes and precisely resolves mobile genetic elements (MGEs). In metagenomics, LRS enables strain-level resolution, the recovery of circular metagenome-assembled genomes, and the precise localization of MGEs within host replicons. Furthermore, the single-molecule, amplification-free properties of LRS provide enhanced resolution of native epigenetic modifications and full-length transcriptomes. Despite these advancements, widespread implementation remains constrained by multidimensional challenges, including stringent high-molecular-weight DNA requirements, depth deficits, and computational overhead. Nevertheless, LRS is increasingly becoming the method of choice for isolate genomics and metagenomics. As detection technologies and algorithms progress, LRS will further improve our ability to decipher the structural and functional diversity of microbial ecosystems.

Metagenomics

Nanopore Sequencing for Chikungunya Virus: Principles and Application.

Nanopore sequencing is transforming viral genomics through real-time, portable, long-read analysis of RNA and DNA. Unlike traditional short-read platforms, it detects nucleotide sequences by measuring ionic current changes as nucleic acids pass through nanoscale pores, enabling direct single-molecule sequencing and base modification detection. Its simplicity, flexibility, and capacity for ultra-long reads make it ideal for resolving complex genomic regions, structural variants, and full viral genomes. These advantages have accelerated its use in pathogen surveillance and outbreak response, especially in resource-limited settings. For chikungunya virus (CHIKV), nanopore sequencing allows rapid, culture-independent recovery of complete genomes from clinical and vector samples, enabling real-time tracking of viral diversity, evolution, and spread. Experiences from Ebola, Zika, and COVID-19 have demonstrated the power of portable sequencing, now applied to CHIKV monitoring. Advances in tools such as Guppy, Dorado, Minimap2, and Medaka enhance read quality, consensus accuracy, and downstream analyses. Despite challenges in basecalling and error correction, robust quality control pipelines ensure reliable results. Ongoing improvements in chemistry, flow cell design, and machine learning will further enhance fidelity and throughput, establishing nanopore sequencing as a cornerstone of CHIKV genomic surveillance and epidemic preparedness.

Chikungunya virus

Multivalent cations stabilize DNA duplexes beyond charge neutralization.

Multivalent cations are abundant in cells and play essential roles in DNA duplex stability, genome packaging, and DNA-protein interactions. They can also condense DNA, making it challenging to determine their influence on DNA duplex stability. To overcome this challenge, we studied DNA unpeeling at equilibrium under high tension using magnetic tweezers, thereby preventing condensation. Experiments show that DNA duplex stability first increases and then decreases as cation concentration increases and the maximum DNA duplex stability increases with cation valence. The maximum free energy change of DNA was 3.33 k B T/bp for Na+ and increased to 3.98 k B T/bp for protamine, which is a small arginine-rich protein with a highly positive charge (≈21 for salmon sperm), corresponding to a relative increase of 19.5%. Consistently, all-atom molecular dynamics simulations show that higher-valent cations preferentially embed in the minor groove of DNA and clamp the minor groove, in contrast to the major-groove clamping reported for RNA, thereby stabilizing the helix more efficiently. These findings establish a single-molecule framework for quantifying DNA thermodynamics in complex ionic environments, which contributes to understanding ionic control of genome stability and to designing ion-tunable DNA-based nanostructures and delivery systems.

Journal Article

An ATP-Driven N Protein-DDX21 Molecular Switch Dynamically Controls SARS-CoV-2 RNA G-Quadruplex Heterogeneity.

The SARS-CoV-2 RNA genome functions as a highly structured regulatory scaffold. Although bioinformatic analyses predict widespread RNA G-quadruplexes (G4s) across the viral genome, their structural diversity and regulatory mechanisms remain poorly understood. Here, we report a diverse landscape of viral G4s encompassing parallel and non-canonical topologies with remarkable thermostability. Unlike typical eukaryotic G4s, these two-tetrad viral G4s exhibit a hierarchical ion-dependent mechanism, in which K+ establishes the core fold, and Mg2 + acts as a secondary regulator promoting conformational compaction. Single-molecule FRET analysis further distinguishes rigid, long-lived G4 folds from highly dynamic, metastable species, defining a continuum of conformational states along the viral genome. Functionally, we identify a synergistic yet competitive interplay between the viral nucleocapsid (N) protein and host helicase DDX21. While the N protein acts as a molecular chaperone to promote G4 folding, DDX21 selectively resolves these structures in an ATP-dependent manner. Strikingly, N and DDX21 jointly constitute a finely tuned, ATP-driven molecular switch, where ATP availability dictates the equilibrium between G4-stabilized and resolved states. Our findings establish a mechanistic framework for the active regulation of SARS-CoV-2 RNA architecture and reveal a multilayered host-virus regulatory axis that modulates viral genome heterogeneity.

DEAD‐box helicases

Somatic mutations and genome mosaicism in aging and disease.

Age-related genome mosaicism is an inherent feature of multicellularity and genomic instability. It occurs because of DNA mutations, the accumulation of which leads to diverse genomic landscapes across different tissues. DNA mutations in the genome are consequences of DNA damage, changes in the chemical structure of DNA, such as strand breaks or loss of bases. DNA damage is very frequent and normally repaired quickly. However, errors intrinsic to DNA repair or replication can give rise to permanent changes in genome sequence information. Such DNA mutations are diverse and include single-nucleotide variants, small insertions and deletions, and larger genome structural variants. Since the 1950s, somatic mutations have been proposed to be a major cause of aging. Indeed, somatic mutations are the cause of cancer, the risk of which increases exponentially with age, and possibly other age-related diseases, such as neurodegenerative diseases and cardiomyopathies. Somatic mutations vary from cell to cell owing to the innate stochasticity of their occurrence, from error-prone processing of randomly inflicted DNA damage. With the emergence of single-cell and single-molecule sequencing, it has become possible to quantitatively analyze somatic mutations in human cells and tissues. Here, we discuss a possible causal relationship between mutation-driven mosaicism of the somatic genome and aging-related functional decline and disease by exploring several predictions of the somatic mutation theory of aging.

Humans

Bridging-driven condensation by eukaryotic SMC complexes is a conserved feature of genome organization.

The Structural Maintenance of Chromosome (SMC) protein family plays a central role in higher-order genome organization through ATP-dependent DNA loop extrusion by cohesin and condensin and other processes. Whether these activities fully account for the complexity of chromosome architecture remains unknown. Here, we uncover a conserved ATP-independent mechanism of chromatin condensation by SMC complexes, occurring via biomolecular condensation. Using single-molecule fluorescence imaging, we show that a variety of SMCs form dynamic DNA-bound condensates that exhibit key features of biomolecular condensates, including droplet coalescence, fluorescence recovery after photobleaching, and rapid exchange with free SMC complexes. Atomic force microscopy analysis of human cohesin-DNA assemblies reveals DNA-length-dependent clustering, providing evidence for bridging-driven condensation. Analyses of in vivo super-resolution imaging and high-throughput chromosome conformation capture (Hi-C) data indicate that these condensates form chromatin-associated clusters with multi-loop structures. Together, our results establish that SMC complexes employ ATP-independent phase condensation as well as ATP-dependent activities to shape genome architecture. This work reveals a broadly conserved principle of chromosomal organization across eukaryotes.

Chromosomal Proteins, Non-Histone

A protein-dependent riboswitch activates ribosomal frameshifting in cardioviruses.

Programmed -1 ribosomal frameshifting (PRF) is a translational control mechanism used by RNA viruses to regulate the relative abundance of proteins encoded in different reading frames. Cardioviruses exhibit the highest known PRF efficiency, with ∼85% of ribosomes shifting into the -1 frame. This unusual event requires an interaction between the viral 2A protein and a stimulatory element in the RNA genome, but the basis for protein dependence is unclear. To address this, here we investigate the structure and dynamics of the PRF signal in Theiler's murine encephalitis virus (TMEV). By combining X-ray crystallography, small-angle X-ray scattering (SAXS), and single-molecule fluorescence resonance energy transfer (smFRET), we show that 2A binding switches the RNA from a stem-loop conformation into a pseudoknot, and we demonstrate that pseudoknot formation is essential for efficient PRF in vitro and in cells. Together, these findings illustrate how the cardiovirus PRF element behaves as a protein-dependent riboswitch, defining the molecular mechanism by which frameshifting is conditionally activated.

Frameshifting, Ribosomal

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis

The multi-functional Smc5/6 complex in genome protection and disease.

Structural maintenance of chromosomes (SMC) complexes are ubiquitous genome regulators with a wide range of functions. Among the three types of SMC complexes in eukaryotes, cohesin and condensin fold the genome into different domains and structures, while Smc5/6 plays direct roles in promoting chromosomal replication and repair and in restraining pathogenic viral extra-chromosomal DNA. The importance of Smc5/6 for growth, genotoxin resistance and host defense across species is highlighted by its involvement in disease prevention in plants and animals. Accelerated progress in recent years, including structural and single-molecule studies, has begun to provide greater insights into the mechanisms underlying Smc5/6 functions. Here we integrate a broad range of recent studies on Smc5/6 to identify emerging features of this unique SMC complex and to explain its diverse cellular functions and roles in disease pathogenesis. We also highlight many key areas requiring further investigation for achieving coherent views of Smc5/6-driven mechanisms.

Animals

Yeast growth is controlled by the proportional scaling of mRNA and ribosome concentrations.

Despite growth being fundamental to all aspects of cell biology, we do not yet know its organizing principles in eukaryotic cells. Classic models derived from the bacteria E. coli posit that protein-synthesis rates are set by mass-action collisions between charged tRNAs produced by metabolic enzymes and mRNA-bound ribosomes. These models show that faster growth is achieved by simultaneously raising both ribosome content and peptide elongation speed. Here, we test if these models are valid for eukaryotes by combining single-molecule tracking, spike-in RNA sequencing, and proteomics in 15 carbon- and nitrogen-limited conditions using the budding yeast S. cerevisiae. Ribosome concentration increases linearly with growth rate, as in bacteria, but the peptide elongation speed remains constant (~9 amino acids/s) and charged tRNAs are not limiting. Total mRNA concentration rises in direct proportion to ribosomes, driven by enhanced RNA polymerase II occupancy of the genome. We show that a simple kinetic model of mRNA-ribosome binding predicts both the fraction of active ribosomes, the growth rate, and responses to transcriptional perturbations. Yeast accelerate growth by coordinately and proportionally co-up-regulating total mRNA and ribosome concentrations, not by speeding elongation. Taken together, our work establishes a new framework for eukaryotic growth control and resource allocation.

Journal Article

Mapping Protein Occupancy on DNA with an Unnatural Cytosine Modification.

The epigenome provides a dynamic layer of gene regulatory control above the static genetic sequence. DNA base modifications are key epigenetic regulators, predominantly found within CpG contexts in mammalian genomes. Working in tandem with these DNA modifications, chromatin-associated proteins and transcription factors further control gene expression. Given the interplay of these factors, concurrent mapping of DNA base modifications with protein-DNA occupancy can greatly aid in interpreting the epigenome. Existing multimodal mapping methods include the use of DNA methyltransferases to mark accessible, protein-unbound DNA in non-CpG contexts. However, such approaches can either confound readouts with native DNA modifications or constrain users to third-generation sequencing approaches. To circumvent these limitations, we explored the possibility of introducing an unnatural DNA base modification, 5-carboxymethylcytosine, as an alternative label for protein occupancy. Here, we report our efforts to rationally engineer non-CpG-specific DNA methyltransferases to take on neomorphic DNA carboxymethyltransferase (CxMTase) activities. We find that DNA carboxymethylation of cytosines in GpC contexts shows broad compatibility with the most widely used epigenetic detection methods and can be used to reliably report on protein occupancy states. Using this approach, we reveal the single-molecule binding patterns of LexA, a master repressor in the bacterial DNA damage (SOS) response, at its self-regulated and endogenously methylated promoter. We thus show that unnatural DNA modifications can uncover novel biological insights and potentiate new approaches to multimodal epigenetic profiling.

DNA

Structure and operating principles of a monkeypox virus replisome.

Poxviruses are double-stranded DNA viruses with large genomes. Among them, monkeypox virus (MPXV) has been responsible for two recent public health emergencies as declared by the World Health Organization1. The MPXV polymerase comprises three subunits-a catalytic subunit (F8) and a heterodimeric processivity factor (A22 and E4). The viral polymerase must coordinate activities with the hexameric helicase-primase (E5) to initiate replication of the viral genome2. Although structures of MPXV E5 (refs. 3,4) and the polymerase5-7 in isolation are available, how they assemble into a functional replisome remains unclear. In isolation, E5 is in an autoinhibited conformation and has very weak helicase activity3,4, and the mechanism for helicase activation is unclear. Here we used cryo-electron microscopy to determine the structures of DNA-bound MPXV replisomes comprising the polymerase holoenzyme (F8, A22 and E4) and the E5 helicase hexamer. We show that, during replisome assembly, E5 undergoes large-scale conformational changes that allow two of its primase domains to interact with the polymerase F8 thumb and A22 subunit. Biochemical assays and single-molecule experiments reveal that this E5 conformational change is coupled to helicase activation and enhances primase activity. Taken together, these findings identify fundamental mechanisms governing coordinated helicase and polymerase activities during DNA replication for an important class of viral pathogens.

Journal Article