Search PubMedSearch

SEARCH · Search PubMed

Results for “WGS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Graph-KIR: graph-based KIR copy number estimation and allele calling using short-read sequencing data.

MOTIVATION: The Killer-cell Immunoglobulin-like Receptor (KIR) is a highly polymorphic region in the human genome, associated with autoimmune diseases and organ transplantation. The sequences of KIR genes are highly similar among star alleles as well as in between individual genes, with the copy number of each KIR gene typically ranging from 0 to 4. In this study, we introduce Graph-KIR, a tool designed to estimate gene copy numbers and predict full-resolution (7-digit, encompassing both coding and non-coding sequence variations) from a whole genome sequencing (WGS) sample. RESULTS: Graph-KIR is capable of independently typing KIR alleles per sample with no reliance on the distribution of any framework gene in a cohort. In a set of 100 simulated samples, Graph-KIR demonstrated 99.2% accuracy in copy number estimation and high F1-score of allele typing: 91.79% at 7-digit resolution, 97.37% at 5-digit resolution, and 97.11% at 3-digit resolution. Graph-KIR outperforms existing tools such as Geny (96.39% F1-score), PING's WGS version (92.77% F1-score), and T1K (90.44% F1-score) at 5-digit resolution. By analyzing the results on 44 HPRC samples, Graph-KIR achieves better F1-score than Geny and PING at 7-digit resolution. The release of Graph-KIR adds another valuable tool to assist users in accurately estimating copy numbers and calling alleles of KIR genes from WGS samples. AVAILABILITY AND IMPLEMENTATION: The Graph-KIR and paper-related pipeline codes are available at https://github.com/linnil1/KIR_graph.

Receptors, KIR

WxS-QC-a quality control pipeline for human germline short-variant Whole-Genome and Whole-Exome cohorts for population-scale analyses.

SUMMARY: Whole-exome (WES) and whole-genome (WGS) sequencing are rapidly becoming preferred methods for population-scale analysis of the human genetic landscape. However, there are currently no standardized quality control (QC) pipelines for human WES and WGS datasets. In this paper, we present WxS-QC, a powerful, scalable, and convenient pipeline for the QC of human germline short-variant WGS and WES cohorts for population-scale analyses. Our pipeline is suitable for both rare-variant discovery and common-variant association studies. It is based on deeply refactored gnomAD v3 and v4 quality control pipelines, contains several methods we have developed de novo, and is aligned with current best practices in WGS/WES germline cohort QC. We provide all methods in a single codebase, aligned to work together and controlled via a single YAML config, with automatic export of resulting graphs and summary tables, excellent performance and scalability, and comprehensive documentation. The pipeline can run in any UNIX-like environment and can efficiently process cohorts of up to 200 000 whole-exome samples, with the potential to handle bigger datasets. AVAILABILITY AND IMPLEMENTATION: The pipeline code is written in Python using the Hail library and is freely available under the BSD-3 license here: https://github.com/wtsi-hgi/wxs-qc. The detailed description of the pipeline is available in the pipeline documentation: https://github.com/wtsi-hgi/wxs-qc/blob/main/README.md. We also provide an open dataset with all required metadata, which is available at https://wxs-qc-data.cog.sanger.ac.uk/wxs-qc_public_dataset_v3.tar. An example of test dataset analysis is available in the supplementary materials.

Humans

Global spread of Streptococcus pyogenes A genomics-supported narrative review.

Group A Streptococcus (GAS) has recently reemerged as a leading cause of both mild and severe invasive infections worldwide, with recent upsurges in invasive disease among children and adults. Notwithstanding a partial synchronicity with the COVID-19 pandemic, this rapid global dissemination of more virulent GAS lineages has been promptly detected, as well as the molecular shifts underlying the observed changes in clinical patterns. Whole-genome sequencing (WGS)-based genomic epidemiology allowed us to gain relevant insights into this upsurge as it was happening. This review integrates the canonical research publication-based approach with genomic data and metadata and identifies a subset of genomic clusters playing a major role in invasive GAS (iGAS) infections worldwide, which were named as Global Pathogenic Lineages (GPLs). The four GPLs broadly coincide with five sequence types (STs): GPL1 with ST28, GPL2 with ST15 and ST315, GPL3 with ST52, and GPL4 with ST39. While non-GPLs clusters maintain a baseline reservoir of antimicrobial-resistance and virulence genes, GPLs show varying but noteworthy resistance profiles and are frequent causes of iGAS. The integration of WGS into routine diagnostics procedures is a forthcoming improvement, aimed not only at informing tailored therapy and implementing infection control strategies, but also to perform continuous surveillance. Ongoing WGS in clinical microbiology, as a matter of fact, will provide unparalleled insights into lineage emergence, transmission dynamics, and the geographic clustering of virulence and resistance determinants.

Streptococcus pyogenes

Accurate identification of abnormal ploidy using an artificial intelligence model in preimplantation genetic testing.

STUDY QUESTION: Can ultra-low-coverage whole-genome sequencing (ulc-WGS) accurately identify abnormal ploidy during preimplantation genetic testing (PGT)? SUMMARY ANSWER: The artificial intelligence (AI)-based PGT-Plus model demonstrates high accuracy in ploidy detection, offering a cost-effective solution that enhances clinical utility of PGT. WHAT IS KNOWN ALREADY: The predominant PGT for aneuploidy can identify chromosomal aneuploidies but cannot determine ploidy status. Transferring embryos with ploidy abnormalities can result in miscarriage and molar pregnancy. On the other hand, in ART, fertilization is assessed by morphological pronuclear assessment at the zygote stage. However, it has a low specificity in the prediction of abnormal ploidy status and embryos deemed abnormally fertilized can yield healthy pregnancies. Accurately identified abnormal ploidy in PGT-A can resolve current limitations and expand the utility range of PGT-A. Several studies have identified ploidy abnormalities; however, they were mainly based on single-nucleotide polymorphism (SNP) arrays or needed to combine additional targeted-next-generation sequencing (NGS) information. Studies based on ulc-WGS remain scarce. STUDY DESIGN SIZE DURATION: The study consisted of two stages: methodology establishment and validation. An AI model, named PGT-Plus, was developed using 653 samples with known ploidy status, which was further validated using 792 different ploidy status samples. In the clinical application stage, the approach was used to analyse the ploidy status of 19&#x2009;103 normally fertilized PGT blastocysts and 140 single pronucleus (1PN)-derived blastocysts collected between May 2022 and December 2023. All blastocysts were tested using trophectoderm biopsy and NGS. PARTICIPANTS/MATERIALS SETTING METHODS: The methodology is based on the ulc-WGS data. First, based on samples with known ploidy status: the heterozygosity rate of high-frequency biallelic SNPs, the likelihood ratio (LLR) of alleles was calculated under different assumptions ('both parental homologs' [BPH] from a single parent, 'single parental homolog' [SPH] from each parent, disomy, and monosomy) by leveraging allele frequencies and linkage disequilibrium (LD) measured in the 1000 genomes project database. Twenty-three continuous candidate features derived from heterozygosity rates and LLRs of chromosomes or selected windows were included to establish the ploidy prediction AI model. Gini importance analysis and multicollinearity mitigation was performed for feature selection, then the performance of Random Forest (RF), Support Vector Machine (SVM), and Logistic Regression for modelling was compared. Subsequently, the parameter optimization was performed based on the RF model. Ploidy constitution concordance was evaluated in known ploidy status samples. The frequency of abnormal ploidy in normal fertilized PGT blastocysts and 1PN-derived blastocysts (including conventional IVF and ICSI) was evaluated. MAIN RESULTS AND THE ROLE OF CHANCE: Eleven features were collected for model architecture compared to SVM and Logistic Regression; RF achieved superior performance for ploidy detection. The AI model achieved an AUC of 1 for genome-wide-uniparental diploidy (GW-UPD), 1 for triploidy, and 0.99 for diploidy. For the 792 validation samples, 99.5% of samples were successfully detected using the AI model, and the model showed 100% accuracy for ploidy classification. In the clinical application stage, out of 19&#x2009;103 PGT samples, 19&#x2009;069 were successfully analysed using the model, with 110 (0.57%) identified as having abnormal ploidy embryos. Among these, 12.7% (14/110) were identified as GW-UPD, and 87.3% (96/110) were triploid. Among 5563 diploid blastocysts transferred, 3478 clinical pregnancies were achieved. Subsequent ploidy analysis was performed for 217 spontaneous abortion and 935 prenatal diagnostic samples, and no abnormal ploidy was identified. Furthermore, of the 140 1PN embryos tested, 40 (28.6%) exhibited GW-UPD, 3 (2.1%) exhibited triploidy, and 97 (69.3%) were determined to be biparental and normally fertilized. Among the 97 biparental embryos, 46 were diploid, 11 were mosaic, and 40 were aneuploid. In terms of the insemination pattern, the percentage of abnormal ploidy in ICSI was significantly higher than in conventional IVF (P&#x2009;<&#x2009;0.01, 37.1% vs. 2.9%, respectively). With full informed consent, 20 patients without euploidy from normal fertilization chose 1PN-derived biparental and diploid blastocysts to transfer, resulting in 10 clinical pregnancies and 9 ongoing pregnancies. LARGE-SCALE DATA: N/A. LIMITATIONS REASONS FOR CAUTION: Some rare ploidy abnormalities, such as polyploidy with an equal number of identical sets of chromosomes and ploidy mosaicism cannot be accurately identified. Moreover, the origin of abnormal ploidy was not identified due to the unavailability of DNA from both parents. WIDER IMPLICATIONS OF THE FINDINGS: The PGT-Plus AI model provides a ploidy evaluation method based on the conventional PGT-A data and integrates directly into standard PGT-A workflows. Clinical utility results suggest that the model is a valuable tool for identifying embryos with abnormal ploidy in PGT-A and rescuing normal diploid embryos from abnormally fertilized embryos. These findings demonstrate that PGT-Plus significantly enhances the diagnostic accuracy of PGT. STUDY FUNDING/COMPETING INTERESTS: This study was supported by grants from Major Scientific Program of CITIC Group (No. 2023ZXKYB34100, to Ge.L.), Hunan Provincial Grant for Innovative Province Construction (2019SK4012), Hunan Xiangjiang New District (Changsha High-tech Zone) key core technology research project in 2023, and Science Foundation of Hunan Province (Grant 2023JJ30422). All authors declared no conflicts of interest..

artificial intelligence

Integrating genetic predictors into subsequent breast cancer risk prediction in survivors of childhood cancer.

PURPOSE: Female survivors of childhood cancer are at high risk for developing breast cancer. The contributions of most general population primary breast cancer genetic predictors to this risk have not been explored. METHODS: Analyses included females who survived &#x2265;5 years after their childhood cancer diagnosis with available array (N&#x2009;=&#x2009;2096, subsequent breast cancer [SBC]=218) or whole-genome sequencing (WGS; N&#x2009;=&#x2009;3292, SBC=101) data from the Childhood Cancer Survivor Study and St. Jude Lifetime Cohort. We computed 99 externally-validated primary breast cancer polygenic risk scores (PRS). Using deep-coverage WGS, ClinVar-annotated pathogenic/likely pathogenic (P/LP) variants in breast cancer susceptibility genes were identified. Cox proportional hazards models assessed associations with SBC risk, adjusting for treatments and genetic ancestry. RESULTS: Among 5388 female survivors (genetic ancestry, European: N&#x2009;=&#x2009;4,752; African: N&#x2009;=&#x2009;444; East Asian: N&#x2009;=&#x2009;192), 319 developed SBC. Most (90.9%) PRSs were nominally associated with SBC risk (P&#x2009;<&#x2009;0.05), but effect sizes varied substantially. PRSs with superior discriminatory ability had greater genome-wide coverage (e.g., 6.4 million-variant PRS, HR per SD&#x2009;=&#x2009;1.71, 95% CI&#x2009;=&#x2009;1.43 to 2.05; P&#x2009;=&#x2009;4.2x10-9) and 7.7-fold higher odds (P&#x2009;=&#x2009;7.0x10-4) of including variants in multiple DNA damage repair pathways compared with PRSs with weaker risk associations. Among survivors with WGS, 1.6% carried P/LP variants in clinical testing panel genes, which was associated with a 7.4-fold greater risk (95% CI&#x2009;=&#x2009;3.16 to 17.19). Including genetic factors improved SBC risk prediction by age 40 (P&#x2009;<&#x2009;0.001) compared to treatment exposures alone. CONCLUSIONS: Externally-validated primary breast cancer genetic susceptibility predictors are relevant for SBC risk prediction and should be prioritized for risk stratification in survivors.

Journal Article

Target Capture of Ancient Shell DNA Enables Phylogenetic Reconstruction of Deep-Sea Molluscs.

Target capture is widely used to enrich endogenous DNA from calcium phosphate skeletal material in vertebrates, but its performance on calcium carbonate hard parts widely produced by invertebrates remains poorly understood. Here, we compared DNA recovery from four fresh and 12 ancient (eight radiocarbon-dated to 1671-1135&#x2009;years old before present) deep-sea vesicomyid clam shells, including species Archivesica marissinica, A. nanshaensis and A. okutanii, using whole-genome sequencing (WGS) or target capture of ultraconserved elements (UCEs). WGS achieved 16.65% on-target read recovery of UCEs from fresh soft tissue, but <&#x2009;1% from shell specimens. By contrast, UCE capture in the same specimen increased on-target reads by up to 155-fold, reaching 29.84% in fresh shells and up to 72-fold, reaching 19.89% in ancient shells. Target capture of UCEs recovered 142-1001 loci per sample compared to 0-230 with WGS alone. Ancient shells of A. marissinica and A. okutanii, based on reads mapped with bwa-mem2 and bbmap, exhibited characteristic post-mortem DNA damage signals, with average 5'-end C-to-T misincorporation rates of 3.46% and 15.97%, respectively, exceeding the levels observed in fresh A. marissinica shells (maximum 1.24%). UCE-based phylogenetic reconstructions incorporating shell ancient DNA recovered two major clades within Pliocardiinae, consistent with published phylogenomic trees. Together, these findings demonstrate that target-capture enrichment enables effective recovery of highly degraded DNA from ancient mollusc shells and supports robust phylogenetic inference at the intrageneric scale, expanding the utility of shells-one of the most abundant invertebrate remains-for evolutionary, biogeographic and conservation studies.

Animals

FTIR typing of the emerging NDM-14-producing Klebsiella pneumoniae ST147 clone.

UNLABELLED: The emergence and rapid dissemination of NDM-14-producing Klebsiella pneumoniae ST147 represents a major challenge for infection control, requiring timely and reliable outbreak detection tools. In this study, we evaluated Fourier-transform infrared (FTIR) spectroscopy as a rapid typing method for outbreak investigation and compared its performance with whole-genome sequencing (WGS). A collection of 64 carbapenemase-producing K. pneumoniae isolates, including 30 NDM-14-producing ST147 isolates associated with a regional outbreak in the Canary Islands, was analyzed using FTIR spectroscopy and WGS. FTIR-based clustering was optimized using the polysaccharide spectral region and a customized distance cutoff. Genomic relatedness was assessed using multilocus sequence typing, core-genome single-nucleotide polymorphism (SNP) analysis at multiple thresholds, and clustering agreement indices. FTIR identified a dominant spectral cluster comprising 31 isolates, capturing all outbreak-related isolates with 100% sensitivity and 97% specificity. FTIR clustering showed concordance with genomic outbreak definitions at stringent SNP thresholds (10-18 SNPs), with accuracy exceeding 98%. Pairwise distance analysis revealed low FTIR dissimilarity among closely related isolates, whereas increased dispersion occurred at intermediate genomic distances (15-30 SNPs). Agreement indices showed improved concordance as genomic stringency increased, with the Modified Adjusted Rand Index values reaching 94.75 at the 10-SNP threshold. Importantly, FTIR identified an NDM-14-producing isolate from a distinct clonal background. Overall, FTIR spectroscopy provides a rapid and reliable first-line screening tool for identifying homogeneous outbreak clusters. However, due to lineage-dependent behavior and limited resolution at intermediate genomic distances, WGS remains essential for confirmatory analysis and precise delineation of transmission events. IMPORTANCE: The rapid spread of multidrug-resistant Klebsiella pneumoniae poses a major challenge for infection control, particularly during hospital outbreaks where timely identification of transmission is essential. In this study, we evaluate Fourier-transform infrared (FTIR) spectroscopy as a rapid typing approach and compare its performance with whole-genome sequencing in the context of an outbreak caused by NDM-14-producing K. pneumoniae ST147. Our results show that FTIR can reliably identify highly related isolates within a clonal outbreak, supporting early outbreak recognition. However, its performance is influenced by the underlying genomic structure of the population and may require dataset-specific optimization. These findings highlight the potential of FTIR as a first-line screening tool while emphasizing the need for cautious interpretation and integration with genomic methods for accurate outbreak delineation.

Klebsiella pneumoniae

Effectiveness of mass spectrometry and genomic analysis in the surveillance of nontuberculous Mycobacterium in Taiwan.

Nontuberculous mycobacteria (NTM) are diverse, and species-level identification remains challenging in routine diagnostics. We analyzed NTM isolates collected at three regional centers of the National Taiwan University Hospital (NTUH) from 2019 to 2024 to assess geographic variation and identification performance after implementation of matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS). Among 3,188 cases meeting the microbiological criteria for probable pulmonary NTM disease, the species distribution differed by region: Mycobacterium avium complex predominated in central Taiwan (Yunlin, 47.3%), whereas M. abscessus complex (Taipei, 26.5%) and M. kansasii (Hsinchu, 12.4%) were more common in northern Taiwan. In 2019, 14.5% of isolates were reported to be unidentified by MALDI-TOF MS; with workflow optimization and database updates, this percentage decreased but plateaued at 4.5-4.8%. Whole-genome sequencing (WGS) of 61 randomly selected persistently unidentified isolates revealed eight average nucleotide identity (ANI)-defined clusters; 55 isolates (90.2%) could not be assigned to known species using current reference databases. Two clusters detected only in Hsinchu were phylogenetically closest to M. kyorinense, with ANI values below the species demarcation threshold. Overall, we observed marked regional heterogeneity of NTM in Taiwan and a persistent identification gap that remained after MALDI-TOF MS optimization and follow-up WGS.IMPORTANCEThis study characterized regional differences in the NTM species distribution across Taiwan, and the results highlight the limitations of current identification approaches. MALDI-TOF MS identifies most isolates, but locally circulating lineages represent a persistent gap in global reference libraries. Even with whole-genome sequencing (WGS), 90.2% (55/61) of persistently unresolved isolates could not be assigned to known species in the current reference databases despite the formation of clear ANI- and phylogeny-defined clusters. These findings show that both proteomic and genomic reference resources for clinical NTM remain incomplete. Expanding regionally representative databases and performing WGS for isolates that remain unresolved by MALDI-TOF MS will be necessary to improve species-level resolution for surveillance and clinical interpretation.

Taiwan

Whole-genome sequencing-based pathogen characterization for streptococcal infection directly from positive blood culture samples.

Clinical laboratories are increasingly using diagnostic tests directly on positive blood cultures, which may lead to fewer attempts to recover bacterial isolates. Consequently, public health laboratories can benefit from assays that directly process blood culture samples without requiring submission of clinical isolates to determine additional pathogen features not identified by clinical tests, such as vaccine serotype and bacterial genomic relatedness, for surveillance and outbreak response purposes. In partnership with the Minnesota Active Bacterial Core surveillance (ABCs) site, we identified blood culture samples positive for ABCs streptococcal pathogens and characterized them by a direct whole-genome sequencing from blood culture (dWGS) assay. The dWGS results were compared with the results of a reference method (WGS of isolates from the same cultures) to evaluate concordance in pathogen features and genome assemblies. Of the 97 eligible blood culture samples, 83 (86%) passed dWGS quality control criteria and were subjected to a total of 655 dWGS-based tests, which yielded 651 (99.3%) evaluable results. The percent agreement with reference results was 100% (83/83) for M protein gene (emm)/capsular types and 100% (81/81) for multilocus sequencing types. For genotypic antimicrobial susceptibility testing prediction, the percent prediction agreement was 100% (487/487), false resistant prediction rate was 0% (0/417), and the false susceptible prediction rate was 0% (0/66). Assemblies of pathogen genomes from the same patient differed by 1.08 &#xb1; 1.68 (mean &#xb1; SD) sites per genome. The dWGS assay can extract high-quality, important streptococcal strain characteristics directly from positive blood culture samples to support evolving public health needs.IMPORTANCEWhole-genome sequencing (WGS) technologies have emerged as a transformative toolkit used by public health microbiology laboratories to detect and characterize pathogens. The surveillance of bacterial diseases often relies on clinical laboratories to submit pathogen isolates to regional or national public health laboratories, which have the capacity to routinely conduct WGS-based strain characterization. Clinical laboratories are increasingly using diagnostic tests directly on positive blood cultures, which may lead to fewer attempts to recover bacterial isolates. The study evaluated a direct whole-genome sequencing from blood culture (dWGS) assay that directly processes blood culture samples. The dWGS assay recovered high quality, important streptococcal strain characteristics, including vaccine serotypes and whole-genome assemblies, without requiring submission of clinical isolates. Thus, the dWGS assay represents a promising tool for addressing the evolving needs of public health laboratories in the metagenomics era.

Humans

Genomic insights into low-level rifampicin resistance mediated by borderline rpoB mutations in Mycobacterium tuberculosis: prevalence and phylogeny in Northeast China.

The emergence of low-level rifampicin (RIF) resistance in Mycobacterium tuberculosis poses a challenge to tuberculosis (TB) control, as it often leads to discordance between genotypic resistance detected by molecular assays (e.g., Xpert MTB/RIF) and phenotypic susceptibility in conventional drug susceptibility testing (DST). In this study, we performed whole-genome sequencing (WGS) on 17 clinical isolates from Changchun, Northeast China, which exhibited such discordance. All isolates harbored functional borderline mutations in the rpoB RRDR region, predominantly Leu452Pro and Leu430Pro (29% each), followed by His445Asn (18%). RIF minimum inhibitory concentration (MIC) values ranged from &#x2264;0.25 to 1.0 mg/L, confirming low-level resistance. Notably, 53% (9/17) of the isolates were co-resistant to fluoroquinolones and 24% (4/17) to isoniazid (INH). According to WHO classification, 59% (10/17) were pre-extensively drug-resistant TB (Pre-XDR-TB) or multidrug-resistant TB (MDR-TB). Phylogenetic analysis revealed that 94% (16/17) belonged to the East Asian Beijing lineage (Lineage 2.2.1), with no evidence of recent local transmission. These findings underscore the complexity of low-level RIF resistance and its frequent association with broader drug resistance in a dominant lineage, highlighting the need for integrating MIC and WGS into diagnostic algorithms to guide appropriate treatment and surveillance.IMPORTANCEThe accurate detection of RIF resistance is critical for the management of TB, yet standard phenotypic methods often fail to identify strains with low-level resistance conferred by borderline rpoB mutations. This study provides the first genomic characterization of such discordant isolates in Northeast China, revealing a high prevalence of co-resistance to other key drugs and a strong association with the locally dominant Beijing lineage. The findings emphasize that reliance on phenotypic DST alone may lead to underestimation of drug resistance and inappropriate treatment, potentially contributing to the emergence and spread of Pre-XDR-TB and MDR-TB. Incorporating MIC determination and WGS into routine diagnostics could enhance detection, inform tailored therapy, and improve surveillance of these clinically significant strains.

Mycobacterium tuberculosis

Performance of the IR Biotyper, Nanopore, and Illumina sequencing to discriminate Escherichia coli strains originating from poultry.

UNLABELLED: Escherichia coli is a highly diverse bacterial species that includes avian pathogenic E. coli (APEC), one of the most prevalent causative agents of disease in poultry worldwide. Rapid and accurate discrimination of E. coli strains is essential for outbreak management, antimicrobial resistance surveillance, and vaccine development. In this study, we compared the performance of Fourier Transform Infrared (FTIR) spectroscopy using the IR Biotyper system with Nanopore and Illumina whole-genome sequencing (WGS) for typing 200 E. coli isolates, originating from four poultry rearing farms in the Netherlands. From each farm, we sampled 10 one-day-old meat type rearing chicks, and from every chick, we isolated 5 E. coli strains. FTIR clustering showed strong concordance with WGS-based classifications, particularly serotyping and core-genome similarity determined by PopPUNK analysis (Adjusted Rand Index 0.75-0.92). While Nanopore and Illumina sequencing provided the highest genetic resolution, FTIR offered a faster (max 6 vs 12-28 days for 200 isolates) and more cost-effective alternative for assessing clonality. Across all methods, multiple strains were detected per farm, whereas most birds carried a single dominant E. coli strain. Our findings demonstrate that FTIR provides a reliable and scalable phenotypic method for rapid strain discrimination in E. coli, complementing WGS in diagnostic, surveillance, and epidemiological settings where speed and throughput are critical. IMPORTANCE: Escherichia coli is a major pathogen in poultry and a potential zoonotic risk for humans. Rapid and accurate discrimination of avian pathogenic E. coli (APEC) strains is critical for outbreak management, antimicrobial resistance surveillance, and the design of effective autogenous vaccines. In this study, we compared Fourier Transform Infrared (FTIR) spectroscopy with Nanopore and Illumina whole-genome sequencing for strain typing of E. coli isolates originating from poultry. The results show that FTIR provides comparable clustering accuracy to genomic approaches at a fraction of the time and costs. This work demonstrates that FTIR can serve as a practical, high-throughput alternative for routine monitoring of E. coli in veterinary diagnostics and food safety of poultry meat, enabling faster decision-making and more targeted interventions across the poultry production chain.

Animals

No phenotypic resistance observed for most group-3 and -4 variants in Mycobacterium tuberculosis genes related to bedaquiline, clofazimine, delamanid, and pretomanid in a Central and West African context.

The interpretation of genetic variants' association (or not) with phenotypic resistance to newly introduced and repurposed antituberculosis drugs remains challenging, as many mutations detected by whole-genome sequencing (WGS) are classified as of uncertain significance (group 3) or not associated with resistance-interim (group 4) by the World Health Organization (WHO) mutation catalog v2. We evaluated the phenotypic impact of such variants on minimum inhibitory concentrations (MICs) for bedaquiline (BDQ), clofazimine (CFZ), delamanid (DLM), and pretomanid (PA) in Mycobacterium tuberculosis complex isolates from the multi-country DIAMA cohort in sub-Saharan Africa (SSA), which recruited RR/RS-TB patients na&#xef;ve to these drugs. Among 1,475 isolates with available WGS data, 163 variants met eligibility criteria; due to viable strain unavailability, 89 isolates carrying 29 unique BDQ/CFZ-related and 60 unique DLM/PA-related variants were tested for MIC determination using broth microdilution. Additional structural modeling was performed to explore potential effects of amino-acid substitutions on protein stability. Among BDQ/CFZ-related variants, MICs above the critical concentrations (CCs) were consistently associated with mmpR5 variants, whereas variants in atpE, pepQ, and Rv1979c were not. DLM/PA variants (ddn, fbiA-D, and fgd1) were frequently detected as non-fixed populations, yet rarely yielding MIC values above the CC. Predicted structural destabilization showed no consistent association with MIC values or variant fixation status. Under the conditions tested, phenotypic resistance was not detected for most group 3 and 4 variants detected by WGS. Our data provide evidence from SSA to support improved interpretation of resistance-associated mutations for new and repurposed antituberculosis drugs.IMPORTANCEWhole-genome sequencing increasingly detects Mycobacterium tuberculosis complex mutations classified by the World Health Organization (WHO) mutation catalog v2 as group 3 variants of uncertain significance or group 4 variants not associated with resistance-interim, limiting reliable prediction of resistance to new and repurposed antituberculosis drugs. By generating minimum inhibitory concentration (MIC) data for such variants identified in a multi-country sub-Saharan African cohort, this study provides phenotypic evidence to support future refinement and expansion of the WHO mutation catalog v2. Notably, mmpR5 variants associated with elevated bedaquiline/clofazimine MICs were identified in eight isolates, suggesting that some patients in this cohort may have harbored pre-existing resistance-associated variants yet remained potentially eligible for bedaquiline-containing regimens. These findings contribute to improving the interpretation of genomic resistance data and strengthening surveillance of resistance to bedaquiline, clofazimine, delamanid, and pretomanid.

Mycobacterium tuberculosis

Whole genome sequencing and phylogenetic classification accelerate the implementation of respiratory syncytial virus genomic surveillance in Canada: a pilot study.

UNLABELLED: Whole genome sequencing (WGS) has emerged as a powerful tool to facilitate the study of existing and emerging infectious diseases. WGS-based genomic surveillance provides information on the genetic diversity and tracks the evolution of important viral pathogens, including respiratory syncytial virus (RSV). Multiplex tiling polymerase chain reaction (PCR) assays have been used to facilitate sequencing of a variety of pathogens in support of genomics-based surveillance initiatives. We developed, optimized, and implemented multiplex tiling PCR assays for RSVA and RSVB capable of generating near-complete genomes in the majority of contemporaneous specimens tested. A pilot data set comprising 52 RSVA and 37 RSVB genomes derived from Canadian clinical specimens during the 2022-2023 respiratory virus season was used to perform phylogenetic analyses using both near-complete genome and glycoprotein (G) sequences. Overall, the RSV phylogenetic tree built with whole genomes showed identical lineage clusters as compared to the G gene but was more discriminatory. Moreover, the availability of complete genomes enables the identification of a broader range of mutations. For instance, mutations identified in the fusion protein among Canadian isolates tested here, including S377N, K272M, S276N, S211N, S206I, and S209Q, could affect the efficacy of current vaccines or antiviral-based therapeutics. In conclusion, our work reinforces other recent studies demonstrating the utility of multiplex tiling PCR assays to facilitate high-throughput WGS of RSV, which is capable of supporting enhanced genomic surveillance initiatives, as well as the more comprehensive genomic analyses required to inform public health strategies for the development and usage of vaccines and antiviral drugs. IMPORTANCE: We present assays to efficiently sequence genomes of RSVA and RSVB. This enables researchers and public health agencies to acquire high-quality genomic data using rapid and cost-effective approaches. Genomic data-based comparative analysis can be used to conduct surveillance and monitor circulating isolates for efficacy of vaccines and antiviral therapeutics.

Humans

Mobilization of blaVIM genes via the Tn6292 transposon among carbapenem-resistant Enterobacter cloacae complex isolates from colonized patients in a Spanish hospital.

UNLABELLED: The aim of this study was to perform molecular characterization of the carbapenem-resistant Enterobacter cloacae complex (ECC) isolates from colonized patients in a hospital using whole-genome sequencing (WGS) technology. As part of routine surveillance for multidrug-resistant bacterial colonization, 21 ECC isolates were recovered from patients at San Carlos Hospital in Madrid (Spain) between December 2020 and November 2024. WGS was used to determine their genetic relatedness. Furthermore, species identification, sequence type (ST), resistome, plasmid content, and flanking mobile genetic elements (MGEs) of the carbapenemase genes were derived from the WGS data. The most prevalent carbapenemase gene identified was blaVIM-1 (n = 18, 85.7%), with other notable genes including blaKPC-2 (n = 1, 4.8%), blaKPC-3 (n = 1, 4.8%), and blaOXA-48 (n = 1, 4.8%). Several blaACT and blaESBL variants were also found among the carbapenem-resistant ECC isolates. All of them carried at least one blaACT gene, with blaACT-7 (11/21) and blaTEM-type (14/21) genes being the most common AmpC and ESBL-encoding genes, respectively. Additionally, two isolates exhibited the presence of the mcr-9 gene. Overall, E. hormaechei subsp. steigerwaltii (ST93), followed by E. hormaechei subsp. hoffmanii (ST78 and ST50), were the predominant species and STs circulating among the carbapenem-resistant ECC strains. The blaVIM-1 gene was part of class 1 integrons located within a Tn3-family transposon, Tn6292. blaKPC and blaOXA-48 were linked to Tn4401 and Tn1999 transposons, respectively. In conclusion, the presence of the blaVIM within a transposon Tn6292 enhances its mobility across bacterial genomes, underscoring the value of high-throughput sequencing in monitoring the spread of carbapenem-resistant ECC isolates. IMPORTANCE: This study highlights why monitoring the spread of antibiotic-resistant bacteria in hospitals is critical. By analyzing the complete DNA of carbapenem-resistant bacteria, antibiotics were considered a last line of treatment. We found that the resistance genes are not isolated. Instead, they are embedded within mobile elements called transposons. This means that they can "jump" between different bacteria, accelerating the spread of resistance. These findings emphasize the importance of high-resolution genomic technologies to track and control the spread of these dangerous bacteria in clinical settings, helping preserve the effectiveness of life-saving treatments.

Humans

Mutagenic Impact and Evolutionary Influence of Chemoradiotherapy in Hematologic Malignancies.

UNLABELLED: Ionizing radiotherapy (RT) is a widely used treatment strategy for malignancies. In solid tumors, RT-induced double-strand breaks lead to the accumulation of insertion-deletions (indels; ID), and their repair by nonhomologous end joining has been linked to the ID8 mutational signature in surviving cells. However, the extent of RT-induced mutagenesis in hematologic malignancies and its impact on their mutational profiles and interplay with commonly used chemotherapies has not yet been explored. In this study, we interrogated 580 whole-genome sequence (WGS) samples from patients with large B-cell lymphoma, multiple myeloma, and myeloid neoplasms and identified ID8 only in relapsed disease. Yet ID8 was detected after exposure to both RT and mutagenic chemotherapy (i.e., platinum and melphalan). Using WGS of single-cell colonies derived from treated lymphoma cells, we revealed a dose-response relationship between RT and platinum and ID8. Finally, using ID8 as a genomic barcode, we demonstrate that a single RT-surviving cell may seed distant relapse. SIGNIFICANCE: RT and the ID8 indel signature are related, but their genomic impact on hematologic malignancies is unclear. Leveraging WGS, we linked ID8 to both RT and mutagenic chemotherapy and validated that platinum can induce ID8. We used ID8 as a genomic barcode to reveal that RT-resistant cells may seed systemic relapse.

Humans

Analysis of genetic differences underlying chilling stress tolerance using whole genome Re-Sequencing in walnut (Juglans regia L.).

Walnut (Juglans regia L.) is prized worldwide for both its nutritional value and economic importance, yet it remains vulnerable to cold stress, with significant differences in tolerance among varieties. This study combined physiological analyses with whole-genome resequencing (WGS) to evaluate the cold stress responses of two varieties, &#x2018;Qingxiang&#x2019; and &#x2018;Liaoning No.8&#x2019;. Under chilling stress (0&#xa0;&#xb0;C), we measured electrolyte leakage and antioxidant enzyme activity, applying both exogenous methyl jasmonate (MeJA) and the jasmonate inhibitor DIECA. Genomic variations were analyzed using WGS. Results showed that &#x2018;Liaoning No.8&#x2019; exhibited superior cold tolerance. Application of MeJA reduced electrolyte leakage by 37% and MDA accumulation by 52% on average, whereas DIECA exacerbated stress-related damage. WGS achieved 16.24&#x2013;16.26&#xd7; coverage and identified 2.73&#x2013;2.78&#xa0;million SNPs, 378&#x2013;382k InDels, 25&#x2013;26k SVs, and 7.2&#x2013;7.9k CNVs. Twenty genes containing sequence variants showed transcriptional responses under cold stress that were significantly correlated with mutation density (r&#x2009;=&#x2009;0.62, P&#x2009;<&#x2009;0.01). One gene, XM_018985465.2, which lacked SNPs in &#x2018;Liaoning No.8&#x2019;, was expressed 4.2 times more in this variety, suggesting cis-regulatory influence. These findings highlight the role of jasmonic acid signaling in enhancing cold tolerance in walnut and offer genomic insights into its underlying adaptive mechanisms.

Juglans

Analysis of targeted and whole genome sequencing of PacBio HiFi reads for a comprehensive genotyping of gene-proximal and phenotype-associated Variable Number Tandem Repeats.

Variable Number Tandem repeats (VNTRs) refer to repeating motifs of size greater than five bp. VNTRs are an important source of genetic variation, and have been associated with multiple Mendelian and complex phenotypes. However, the highly repetitive structures require reads to span the region for accurate genotyping. Pacific Biosciences HiFi sequencing spans large regions and is highly accurate but relatively expensive. Therefore, targeted sequencing approaches coupled with long-read sequencing have been proposed to improve efficiency and throughput. In this paper, we systematically explored the trade-off between targeted and whole genome HiFi sequencing for genotyping VNTRs. We curated a set of 10&#xa0;,&#xa0;787 gene-proximal (G-)VNTRs, and 48 phenotype-associated (P-)VNTRs of interest. Illumina reads only spanned 46% of the G-VNTRs and 71% of P-VNTRs, motivating the use of HiFi sequencing. We performed targeted sequencing with hybridization by designing custom probes for 9,999 VNTRs and sequenced 8 samples using HiFi and Illumina sequencing, followed by adVNTR genotyping. We compared these results against HiFi whole genome sequencing (WGS) data from 28 samples in the Human Pangenome Reference Consortium (HPRC). With the targeted approach only 4,091 (41%) G-VNTRs and only 4 (8%) of P-VNTRs were spanned with at least 15 reads. A smaller subset of 3,579 (36%) G-VNTRs had higher median coverage of at least 63 spanning reads. The spanning behavior was consistent across all 8 samples. Among 5,638 VNTRs with low-coverage (&#xa0;<&#xa0;15), 67% were located within GC-rich regions (&#xa0;>&#xa0;60%). In contrast, the 40X WGS HiFi dataset spanned 98% of all VNTRs and 49 (98%) of P-VNTRs with at least 15 spanning reads, albeit with lower coverage. Spanning reads were sufficient for accurate genotyping in both cases. Our findings demonstrate that targeted sequencing provides consistently high coverage for a small subset of low-GC VNTRs, but WGS is more effective for broad and sufficient sampling of a large number of VNTRs.

Minisatellite Repeats

An alignment-free strategy for circulating tumor DNA detection and tumor fraction estimation from whole-genome sequencing data.

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed alignment-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth&#x2009;~&#x2009;28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmer's tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

Circulating Tumor DNA