Search PubMedSearch

SEARCH · Search PubMed

Results for “genome assembly validation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75 Gb with a contig N50 of 35.0 Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5 Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals

Protocol for telomere-to-telomere assembly of Borrelia genomes using a hybrid method.

Borrelia has a linear chromosome and linear plasmids capped by hairpin telomeres that short-read sequencing cannot resolve. Here, we present a protocol for telomere-to-telomere assembly of Borrelia genomes. We describe steps for spanning B. burgdorferi culture, DNA extraction, and sequencing through hybrid genome assembly to generate complete Borrelia genomes. The pipeline integrates Oxford Nanopore long reads and Illumina short reads to assemble hairpin telomeres, resolve paralogous linear and circular plasmids, and annotate and validate the assembled complete Borrelia genome. For complete details on the use and execution of this protocol, please refer to Amin et al.1.

Bioinformatics

Highly Contiguous Is Not Chromosomally Accurate: Integrated Cytogenetic and Genomic Mapping in Two Turtle Genome.

High-quality genome assemblies are essential for robust research across biological and medical fields. Assembly errors can have far-reaching consequences for downstream analyses, including gene annotation and the inference of synteny. In contrast to the rapid growth of genomic data volume, there is a notable lag in the integration of chromosome-level assemblies with cytogenetic data. We conducted the first direct genome-to-genome comparison, integrating comparative chromosome painting, the alignment of chromosome-specific probes to available genome assemblies, and synteny-based comparison of independent chromosome-level assemblies of the loggerhead sea turtle (Caretta caretta, 2n = 56) and the red-eared slider (Trachemys scripta elegans, 2n = 50). Using two independent sets of flow-sorted chromosome-specific probes in cross-species hybridizations, together with the sequencing and mapping of chromosome-derived DNA libraries, we assigned assembled scaffolds to all physical chromosomes of both species. In C. caretta, chromosomal assignments and genome-wide synteny were fully consistent with the published assembly, except for the reduced sizes of two microchromosome scaffolds, which we attribute to under-representation of repetitive DNA. In contrast, in T. s. elegans, cytogenetic validation of the assemblies revealed a false rearrangement compared to a missed one. Our results show that even highly contiguous vertebrate genome assemblies can misrepresent chromosome structure. When cytogenetic analyses reveal such inaccuracies, updated reference genomes should be generated for widely studied species to enable accurate inference of karyotype evolution and downstream comparative genomic analyses.

FISH

Germline stem cell isolation, lineage tracing, and aging in a protochordate.

Germline stem cells (GSCs), the source of gametes, are the only stem cells capable of passing genes to future generations and are therefore considered units of natural selection. Yet, the factors that influence GSC fitness, and thus govern GSC competition, which exist in both protochordates and mammals, remain poorly understood. We studied how aging affects GSC fitness in the protochordate Botryllus schlosseri, an evolutionary crosspoint between invertebrates and vertebrates. GSCs were isolated and distinguished from developing and mature gametes using flow cytometry and scRNA-Seq, facilitated by a new PacBio genome assembly. Moreover, their function was validated through a novel lineage tracing approach that combines membrane-labeled GSC transplantation with scRNA-Seq. Leveraging our method to isolate them, single-cell transcriptomics showed significant age-related changes between young and old GSCs. Spermatids and sperm, however, showed minimal changes, suggesting that reproductive aging is governed by GSCs rather than by gametes. Reduced expressions of markers like DDX4 and PIWIL1 in aged GSCs mirrored trends in mammalian datasets, pointing to a conserved GSC-driven aging mechanism across chordate evolution. This study provides new techniques that lay the foundation to investigate further drivers of GSC fitness and highlights fertility-related genes as promising targets for therapies to preserve reproductive health.

Journal Article

Diversification of Cellulose Synthase (CESA) Genes in Mosses Suggests Both Ancient and Recent Gene duplications.

Cellulose is an important polysaccharide that constitutes all plant cell walls, giving them strength and stability. The plant cellulose synthase (CESA) gene family, which encodes the catalytic subunits of cellulose synthesis complexes (CSCs), has diversified independently in several plant lineages, providing an interesting model for understanding selection for gene duplication. Here we quantified the presence of CESA genes across mosses to understand how the process of gene family diversification occurred in this group and how it parallels diversification in other groups. We first examined the CESA gene family in eight species of mosses across seven families for which whole genome assemblies were available. We then identified CESA genes from additional species, for which only short-read sequence data was available, by using BLAST searches and targeted gene assemblies. We validated this approach by comparing the assembled paralogs from the short-read data to the genes identified from whole genome assemblies in the eight reference species. This approach allowed us to identify paralogs directly from short-read data and greatly expand our sample set. Results from the combined empirical data support the hypothesis that CESA genes diversified within the moss lineage at least as early as the mesozoic period, during or possibly even prior to the onset of moss diversification, but also continue to diversify within modern species. In addition, we found evidence for purifying selection as the dominant force shaping these genes and observed that different lineages experienced different levels of evolutionary constraint. Lastly, our approach to assemble paralogs has the potential to allow researchers to improve analyses of gene duplication events.

Physcomitrium patens

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals

ClinGen recuration of hearing loss-associated genes demonstrates significant changes in gene-disease validity over time.

PURPOSE: The Clinical Genome Resource (ClinGen) Hearing Loss Gene Curation Expert Panel was assembled in 2016 and has since curated 174 gene-disease relationships (GDRs) using ClinGen's semiquantitative framework. ClinGen mandates the timely recuration of all GDRs classified as Disputed, Limited, Moderate, and Strong every 2 to 3 years. METHODS: Thirty-five GDRs met the criteria for recuration within 2 years of original curation. Previous evidence was reevaluated using the latest curation guidelines, and a comprehensive literature review was performed to obtain new evidence. Recurations were approved by the Gene Curation Expert Panel and published on the ClinGen website (www.clinicalgenome.org). RESULTS: Eight of 35 GDRs (22%) changed their classification. Two Moderate and 5 Strong GDRs were upgraded to Definitive because of new case evidence. One Strong was subsumed under another Definitive GDR after evaluation of the lumping/splitting of disease entities. Twenty-seven of 35 patients remained unchanged, with little to no new evidence reported. CONCLUSION: Genes classified as Moderate and Strong were likely to build evidence and change their classification over time, whereas Limited were unlikely to gain evidence. These findings highlight the critical role of recuration in ensuring that genetic tests and research studies incorporate the most recent evidence into their efforts.

Humans

Comparative genomics reveals hidden biosynthetic diversity in Streptomyces spp. and metal-dependent regulatory features associated with untapped specialized metabolites.

The genus Streptomyces is one of the richest sources of bioactive natural products; however, a substantial proportion of its biosynthetic gene clusters (BGCs) remain cryptic and their metabolic products are unresolved. Advances in genome mining and computational prediction now enable comprehensive exploration of this hidden biosynthetic repertoire. In this study, whole-genome sequencing and comparative genomic analyses were performed on three three newly isolated Streptomyces strains to evaluate their specialized metabolic potential. Genome assemblies were annotated and systematically analyzed using antiSMASH, DeepBGC, GECCO, and PRISM to identify, cross-validate, and functionally characterize BGCs while predicting their associated secondary metabolite scaffolds. Taxonomic analyses based on Average Nucleotide Identity (ANI), phylogenomics, and BLAST identified the isolates as Streptomyces thinghirensis, Streptomyces novocaesareae, and Streptomyces griseorubens. Applying the consensus framework across the three Streptomyces genomes yielded 43 cryptic BGCs, lacking close similarity to reference BGCs in the MIBiG database, of which 26 were classified as HIGH, 10 as MEDIUM, and 7 as LOW confidence. Notably, numerous BGCs exhibited low abundance to characterized reference clusters, indicating a high potential for previously undescribed biosynthetic pathways and novel metabolite scaffolds. Comparative analyses further revealed strain-specific biosynthetic architectures together with putative metal-responsive regulatory systems; Fur, Zur, and Nur, which were frequently associated with specialized metabolite biosynthetic loci. Collectively, these findings demonstrate the effectiveness of integrated genome-mining strategies for prioritizing cryptic biosynthetic gene clusters and highlight the remarkable biosynthetic potential of newly identified Streptomyces isolates as a source of novel natural products.

comparative genomics

Genome assembly and annotation of the parasitoid jewel wasp Nasonia oneida.

The jewel wasp, Nasonia (Hymenoptera: Pteromalidae), is a well-established model system for evolutionary genetics and host-microbial interactions. Here, we present the genome of N. oneida, a species lacking prior genomic characterization, using 10× Genomics linked-read (400× coverage), Illumina short-read (120× coverage), and transcriptome data (30× coverage). The assembled genome size is 267 Mb, comprising 4,675 scaffolds, with a scaffold N50 of 1 Mb and 98.40% Benchmarking Universal Single-Copy Orthologues (BUSCOs) completeness score. Annotation revealed 32.29% (86.46 Mb) of repetitive sequences and 14,221 protein-coding genes. Comparative genomics of N. oneida with 15 other hymenopteran species validated the presence of 5,939 gene families shared among them, including 3643 single-copy and 2296 multicopy gene families. This study provides the first de novo assembly of N. oneida, providing a significant addition to the growing repertoire of molecular tools for comparative genomics and functional studies to understand the evolution of closely related species as well as the evolution of parasitic wasps.

Animals

Transcriptomic and Metabolomic Profiling Identifies a Core Gene-Metabolite Axis Driving African Swine Fever Virus Replication in the Soft Tick Ornithodoros lahorensis.

African swine fever virus (ASFV) causes an incurable swine disease with nearly 100% mortality, posing a catastrophic threat to global pig production. The soft tick Ornithodoros lahorensis acts as a critical biological vector that sustains persistent ASFV replication and mediates long-distance viral transmission, yet the molecular mechanisms governing ASFV-tick interplay remain poorly understood. Here, we integrated transcriptomics and metabolomics to systematically dissect molecular changes in O.&#xa0;lahorensis across three infection stages: Uninfected control, early infection (7&#x2009;days post-infection, dpi), and late persistent infection (21 dpi). Multi-omics integration revealed that ASFV extensively remodels tick host metabolism, predominantly activating purine/pyrimidine metabolism, lipid biosynthesis, and energy metabolism. We further characterized a conserved regulatory module consisting of 12 core genes and 8 signature metabolites that collectively support ASFV genome replication and virion assembly. Three hub metabolic genes (TK1, ATP5F1B, and IMPDH) were selected for functional validation via siRNA silencing in ticks; individual gene silencing suppressed ASFV loads by 89.2%, 91.5%, and 87.8%, respectively (p&#x2009;<&#x2009;0.001***). This work represents the first comprehensive multi-omics investigation of ASFV infection in O. lahorensis. We identified tick-specific molecular targets to block vector-mediated ASFV spread and established a standardized multi-omics analytical pipeline for tick-virus interaction research. Our findings elucidate the mechanistic basis of long-term ASFV persistence in soft ticks and deliver novel actionable clues for developing vector-targeted ASF intervention strategies.

Animals

Fantastic microbes and where to find them: evaluating learning-by-doing outcomes in a crowdfunded metagenomics workshop.

Metagenomics offers a powerful framework for authentic, interdisciplinary learning, yet it remains underrepresented in undergraduate education due to technical and infrastructural barriers. We hypothesized that a research-based, learning-by-doing metagenomics workshop supported by accessible bioinformatics tools could enhance students' perceived skills, self-efficacy, and conceptual understanding of metagenomic analysis. To test this hypothesis, we designed and evaluated a hybrid hands-on workshop in which undergraduate and postgraduate students analyzed real environmental shotgun metagenomic datasets generated from soil samples collected during a citizen science initiative. Using the graphical workflow platform KBase, participants completed an end-to-end metagenomic analysis, from quality control and assembly to genome reconstruction, taxonomic classification, functional annotation, and scientific presentation of results. Educational outcomes were assessed through validated retrospective pre-post questionnaires, self-efficacy scales, and an open-ended conceptual understanding task. Participants showed significant increases in perceived metagenomic skills and confidence in performing metagenomic analyses, while gains in perceived learning showed a positive trend. Conceptual understanding improved across educational levels, particularly among participants with limited prior experience. Together, these findings demonstrate that authentic, data-driven metagenomics activities can effectively lower barriers to computational biology and foster meaningful learning through hands-on research experiences.

Metagenomics

Role of RNA G-Quadruplexes in the Japanese Encephalitis Virus Genome and Their Recognition as Prospective Antiviral Targets.

G-quadruplexes (GQs) have been primarily studied in the context of cancer and neurodegenerative pathologies. However, recent research has shifted focus to their existence and functional roles in viral genomes, revealing GQ-regulated key pathways in various human pathogenic viruses. While GQ structures have been reported in the genomes of emerging and re-emerging viruses, RNA viruses have been understudied compared to DNA viruses, including notable examples such as human immunodeficiency virus-1, hepatitis C virus, Ebola virus, Nipah virus, Zika virus, and SARS-CoV-2. The flavivirus family, comprising the Japanese encephalitis virus (JEV), poses a significant global threat due to recurring outbreaks yet lacks approved antivirals. In this study, we identified and characterized eight putative G-quadruplex-forming motifs within essential genes involved in genome replication, assembly, and internalization in the host cell, conserved across different JEV isolates. The formation and stability of these motifs were validated through a multitude of biophysical and cell-based assays. The interaction and binding affinity of these motifs with the known GQ-binding ligand BRACO-19 were supported by biophysical assays, confirming the capability of these motifs to form GQ structures. Notably, BRACO-19 also exerted antiviral properties through reduction of viral replication and infectious virus titers as well as inhibition of viral protein expression, as evaluated by the cell-based assays. This comprehensive molecular characterization of G-quadruplex structures within the JEV genome highlights their potential as promising antiviral targets for intervention strategies against JEV infection through GQ-specific ligands.

G-Quadruplexes

Comparative genomic analysis of Streptococcus parasuis and Streptococcus suis reveals mobile element-associated enrichment of antimicrobial resistance and lack of detectable same-MGE colocalization with virulence-associated genes within stable species boundaries.

Streptococcus suis is a major porcine pathogen and a zoonotic agent that causes meningitis and septicemia in humans. Streptococcus parasuis, a recently recognized close relative, remains poorly characterized with regard to its clinical significance and genomic features. In this study, we generated a single-contig closed genome assembly with genome-wide DNA methylation profiles for S. parasuis strain A1, isolated from a diseased pig in Xinjiang, China, and complemented in silico genomic predictions with isolate-level experimental validation of antimicrobial resistance (AMR) genotypes, virulence genotypes, and phenotypic susceptibility for this reference strain. Using this high-quality genome as a reference anchor, we performed comparative genomic analyses across 195 streptococcal genomes, comprising 15 S. parasuis and 180 S. suis strains, to distinguish genome-level co-occurrence of resistance and virulence determinants from their physical colocalization on the same mobile genetic element (MGE).Species boundaries remained clearly delineated at the genomic level, with a median interspecies average nucleotide identity (ANI) of approximately 86.0%, compared with intraspecies ANI medians of 97.5% for S. parasuis and 96.2% for S. suis. Pangenome analysis identified 12,693 gene clusters, of which 1086 were core clusters, and functional annotation revealed significant differences in accessory gene repertoires between the two species. Within this stable genomic framework, S. parasuis genomes carried a higher AMR gene burden; strain A1 harbored 10 AMR genes, multiple virulence-associated genes, three genomic islands, and eight prophage regions. For strain A1, PCR validation confirmed six AMR genes and six virulence genes, and disk diffusion testing demonstrated a multidrug-resistant phenotype consistent with the genotypic profile.Among 235 predicted mobile elements, 19 harbored AMR genes and seven carried Virulence Factor Database (VFDB) homologs, but none carried both categories simultaneously. This finding reflects a lack of detectable same-MGE colocalization under the applied annotation and assembly framework; it should not be interpreted as evidence of biological physical decoupling. Under a random-placement model, the expected number of co-carrying regions was only 0.57, and the probability of observing zero co-carrying regions was P&#x202f;=&#x202f;0.55. This negative result should be interpreted with caution, given the limited number of cargo-bearing regions and the predominantly draft status of most genomes. Furthermore, the A1 genome contained multiple restriction-modification systems, showed depletion of several methylation motif families in mobile regions, and had limited CRISPR spacer matching evidence, suggesting prior exposure to the relevant sequence space. None of the genomes met our predefined criteria for whole-genome convergence.Collectively, our results support a model in which S. parasuis accumulates AMR-related genes in a modular fashion via mobile elements within stable species boundaries, with no detectable same-MGE colocalization of AMR and virulence determinants under our analytical pipeline. These findings imply that AMR surveillance strategies for this species should prioritize tracking mobile genetic elements rather than inferring wholesale genomic convergence toward S. suis.

Streptococcus suis

Integrated Genome Mining and Bioactivity-Guided Isolation of Antimicrobial Peptides from Bacillus amyloliquefaciens BS4.

Bacterial resistance remains a critical global health challenge, driving the continuous search for novel antimicrobial agents. Bacillus amyloliquefaciens is a recognized repository of bioactive metabolites; however, its full biosynthetic potential requires integrated genomic and experimental validation. This study characterized the antimicrobial profile of B. amyloliquefaciens BS4 through a hybrid pipeline. Genome sequencing and de novo assembly revealed a 3.9&#xa0;Mb chromosome with a G&#x2009;+&#x2009;C content of 46.14%. Functional annotation identified 3,887 coding sequences, including pathways for siderophore biosynthesis and a complete bacilysin biosynthetic cluster. BGC analysis using antiSMASH v7.1.0 and BAGEL4 identified 18 biosynthetic gene clusters, while similarity network analysis via BiG-SCAPE highlighted unique singleton BGCs, indicating untapped biosynthetic diversity. Although in silico screening via Macrel predicted two putative cationic antimicrobial peptides (AMPs), bioactivity-guided purification utilizing sequential RP-HPLC, and de novo sequencing revealed a distinct set of four active peptides. Notably, three of these sequences were identified as fragments derived from the BclA exosporium protein family, highlighting the structural proteome as a non-canonical source of antimicrobials. The purified fractions exhibited activity against M. luteus and E. coli, while displaying no significant hemolytic activity or cytotoxicity, even above the MIC values. Molecular docking further supported the interaction of these candidates with bacterial targets. Overall, this hybrid strategy effectively uncovers the antimicrobial complexity of BS4, revealing 'cryptic' peptide candidates with therapeutic potential.

Bacillus amyloliquefaciens BS4

Unveiling novel antimicrobial peptides from the ruminant gastrointestinal microbiomes: A deep learning-driven approach yields an anti-MRSA candidate.

INTRODUCTION: Antimicrobial peptides (AMPs) present a promising avenue to combat the growing threat of antibiotic resistance. The ruminant gastrointestinal microbiome serves as a unique ecosystem that offers untapped potential for AMP discovery. OBJECTIVES: The aims of this study are to develop an effective methodology for the identification of novel AMPs from ruminant gastrointestinal microbiomes, followed by evaluating their antimicrobial efficacy and elucidating the mechanisms underlying their activity. METHODS: We developed a deep learning-based model to identify AMP candidates from a dataset comprising 120 metagenomes and 10,373 metagenome-assembled genomes derived from the ruminant gastrointestinal tract. Both in vivo and in vitro experiments were performed to examine and validate the antimicrobial activities of the AMP candidates that were selected through bioinformatic analysis and subsequently synthesized chemically. Additionally, molecular dynamics simulations were conducted to explore the action mechanism of the most potent AMP candidate. RESULTS: The deep learning model identified 27,192 potential secretory AMP candidates. Following bioinformatic analysis, 39 candidates were synthesized and tested. Remarkably, all synthesized peptides demonstrated antimicrobial activity against Staphylococcus aureus, with 79.5% showing effectiveness against multiple pathogens. Notably, Peptide 4, which exhibited the highest antimicrobial activity against methicillin-resistant Staphylococcus aureus (MRSA), confirmed this effect in a mouse model with wound infection, exhibiting a low propensity for resistance development and minimal cytotoxicity and hemolysis towards mammalian cells. Molecular dynamics simulations provided insights into the mechanism of Peptide 4, primarily its ability to disrupt bacterial cell membranes, leading to cell death. CONCLUSION: This study highlights the power of combining deep learning with microbiome research to uncover novel therapeutic candidates, paving the way for the development of next-generation antimicrobials like Peptide 4 to combat the growing threat of MRSA would infections. It also underscores the value of utilizing ruminant microbial resources.

Animals

Clinical and genomic features of mitis group streptococcal bacteremia in patients with febrile neutropenia.

BACKGROUND: Viridans group streptococci (VGS) can cause the life-threatening viridans streptococcal shock syndrome (VSSS) in patients with febrile neutropenia (FN). The Mitis group, a major subgroup of VGS, is frequently implicated in these severe infections, but its specific clinical and genomic characteristics remain incompletely characterized, particularly in patients with FN. This study aimed to systematically describe these features in this population. METHODS: In this single-center retrospective study, we compared the clinical data and whole-genome sequencing (WGS) results of Mitis group streptococcal isolates from patients with and without FN. Virulence-associated and antimicrobial resistance genes were initially screened using a reference-based approach, followed by assembly-based reanalysis and manual sequence validation. RESULTS: Compared with the non-FN cohort (n&#x2009;=&#x2009;34), the FN cohort (n&#x2009;=&#x2009;61) was significantly younger, had a higher prevalence of hematologic malignancy, and more frequently presented with primary bacteremia. VSSS occurred exclusively in the FN group (11.5%) and was associated with high mortality (14-day mortality, 42.9%), which did not correlate with in vitro antimicrobial susceptibility. Genomic analyses revealed marked diversity among isolates. Initial screening suggested variable detection of several virulence-associated loci, including pavA, slrA, and rfb-related loci; however, subsequent assembly-based analyses indicated that many apparent absences were attributable to extreme allelic divergence rather than true gene loss. No single virulence determinant clearly segregated with clinical severity. CONCLUSIONS: Mitis group bacteremia in patients with FN appears to be characterized by distinct clinical features and marked genomic diversity. Our findings suggest that the development of severe disease, including VSSS, may not be explained by microbial factors alone and potentially reflects complex host-pathogen interactions. CLINICAL TRIAL: Not applicable.

Humans

Turbo-charging crop improvement: harnessing multiplex editing for polygenic trait engineering and beyond.

Multiplex CRISPR editing has emerged as a transformative platform for plant genome engineering, enabling the simultaneous targeting of multiple genes, regulatory elements, or chromosomal regions. This approach is effective for dissecting gene family functions, addressing genetic redundancy, engineering polygenic traits, and accelerating trait stacking and de novo domestication. Its applications now extend beyond standard gene knockouts to include epigenetic and transcriptional regulation, chromosomal engineering, and transgene-free editing. These capabilities are advancing crop improvement not only in annual species but also in more complex systems such as polyploids, undomesticated wild relatives, and species with long generation times. At the same time, multiplex editing presents technical challenges, including complex construct design and the need for robust, scalable mutation detection. We discuss current toolkits and recent innovations in vector architecture, such as promoter and scaffold engineering, that streamline workflows and enhance editing efficiency. High-throughput sequencing technologies, including long-read platforms, are improving the resolution of complex editing outcomes such as structural rearrangements-often missed by standard genotyping-when targeting repetitive or tandemly spaced loci. To fully realize the potential of multiplex genome engineering, there is growing demand for user-friendly, synthetic biology-compatible, and scalable computational workflows for gRNA design, construct assembly, and mutation analysis. Experimentally validated inducible or tissue-specific promoters are also highly desirable for achieving spatiotemporal control. As these tools continue to evolve, multiplex CRISPR editing is poised to become a foundational technology of next-generation crop improvement to address challenges in agriculture, sustainability, and climate resilience.

Gene Editing

The reference genome of the human diploid cell line RPE-1.

Recent technological advances have facilitated the assembly of telomere-to-telomere (T2T) genomes. The current T2T CHM13 showcases the complete architecture of the human genome, yet its use in functional experiments is limited by discrepancies with the actual genome of the specific biological system under study. Access to reference assemblies for experimentally relevant cell lines is therefore essential in advancing sequencing-based analyses and precise manipulation, particularly in highly variable regions such as centromeres. Here, we present RPE1v1.1, the near-complete diploid genome assembly of the hTERT RPE-1 cell line, a non-cancerous human retinal epithelial model with a stable karyotype. Using high-coverage Pacific Biosciences and Oxford Nanopore Technologies long-read sequencing, we generate a high-quality de novo assembly, validate it through multiple methods, and phase it by integrating high-throughput chromosome conformation capture (Hi-C) data. Our assembly includes chromosome-level scaffolds that span centromeres for all chromosomes. Comparing both haplotypes with the CHM13 genome, we detect haplotype-specific genomic variations, including the translocation between chromosome 10 and chromosome X t(X;10)(Xq28;10q21.2) characteristic of RPE-1 cells, and divergence peaking at centromeres. Altogether, the RPE1v1.1 genome provides a reference-quality diploid assembly of a widely used cell line, supporting high-precision genetic and epigenetic studies in this model system.

Humans