Search PubMedSearch

SEARCH · Search PubMed

Results for “whole-genome- sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Performance of MALDI-TOF MS for human Capnocytophaga identification verified by whole-genome sequencing.

OBJECTIVE: This study aims to evaluate the performance of matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) for species identification of human Capnocytophaga and to confirm results by whole-genome sequencing. METHODS: Six reference strains, representing human Capnocytophaga species and one taxon, and a total of 126 clinical strains, selected based on their biochemical profiles from a large collection of preliminarily identified Capnocytophaga isolates, were analyzed. RESULTS: Of those, 125 strains (94%) were identified at least at the genus level (log score variation of 1.7-1.999), while 52 strains (39%) were identified at the species level with a cut-off score of &#x2265;2.0. Eight strains (6%) remained unidentified with a log score of <1.69. C. leadbetteri and Capnocytophaga genospecies AHN8471 strains were accurately identified at the genus level. Minor identification errors were observed in three cases: C. leadbetteri (n=1), C. ochracea (n=2), and Capnocytophaga genospecies AHN8471 (n=38). MALDI-TOF MS was unable to distinguish between C. sputigena and Capnocytophaga genospecies AHN8471 at the species level but clustered them together in the Main Spectra Profile (MSP) dendrogram. CONCLUSIONS: MALDI-TOF MS shows promise as a diagnostic tool for identifying human Capnocytophaga species when correct taxonomy and sufficient reference strains are available in the database. Based on the close phenotypic, ribosomal, and genotypic structures, we propose to establish the term "C. sputigena group" encompassing C. sputigena, Capnocytophaga genospecies AHN8471, and other related Capnocytophaga variants. Nevertheless, updating and expanding the MALDI-TOF MS reference database is essential to improve identification accuracy.

Capnocytophaga spp.

Whole-genome Sequence Analysis Revealed Novel Subjective Cognitive Decline-associated Genes in 10,763 Chinese.

Subjective cognitive decline (SCD) is widely regarded as a potential preclinical stage of Alzheimer's disease (AD), yet its genetic basis remains poorly understood. To address this gap, we investigated genetic biomarkers associated with SCD using whole-genome sequencing (WGS) in 10,763 Chinese participants from the Healthy Zhejiang One Million People Cohort (HOPE Cohort). The discovery stage included 9284 samples, with 1479 samples used for validation. Using a two-stage design, we systematically investigated both common and rare variants associated with SCD. In rare variant analyses, we identified and replicated an association between the upstream region of SEPHS2 and SCD. SEPHS2 is involved in selenophosphate synthesis, and a Mendelian randomization analysis reveals that its expression levels in both blood and brain cerebellum are associated with AD. Additionally, we identified CLVS2, which encodes a protein primarily expressed in neuronal cells, as a potential regulator for SCD based on missense rare variants. Multi-omics evidence suggests that both SEPHS2 and CLVS2 may play roles in neurodegenerative diseases. For common variants, we validated 8 known loci related to cognitive decline, 3 of which originated from the only existing SCD genetic study conducted under a migraine background. Overall, our WGS-based study fills the gap in SCD research by providing vital genetic evidence from an East Asian population and offers insights into the pathogenic mechanisms of SCD.

Aged

Large-scale low-coverage whole-genome sequencing reveals the genetic architecture of wool and growth traits in fine-wool sheep.

Breeding sheep with superior growth performance and wool quality is essential for the sustainability of the fine-wool sheep industry. In this study, we perform low-coverage whole-genome sequencing (lcWGS) on 3842 individuals from 5 sheep breeds (4 fine-wool and 1 semi-fine wool) and generate a large genomic dataset. By comparing these breeds with coarse-wool sheep, we characterize the genomic landscape and selection signatures of fine-wool sheep. We identify several known functional genes associated with hair follicle development and skin morphology, including EGFR, KRT74, EDAR, EREG, and GLI2. Furthermore, GWAS of 19 traits identifies 156 candidate genes significantly associated with growth and wool characteristics, including LCORL for body size, EGFR for clean wool yield, and PRDM1 for fiber diameter. Notably, EGFR is detected in both GWAS and selection signature analyses, indicating its important role in phenotype formation and historical selection. Overall, our findings reveal the genetic basis of growth and wool traits in fine-wool and semi-fine wool sheep, highlight EGFR, LCORL, and PRDM1 as candidate genes, and provide valuable genomic resources and candidate markers for future functional validation and molecular breeding.

Body size

A dedicated caller for DUX4 rearrangements from whole-genome sequencing data.

Rearrangements involving the DUX4 gene (DUX4-r) define a subtype of paediatric and adult acute lymphoblastic leukaemia (ALL) with a favourable outcome. Currently, there is no 'standard of care' diagnostic method for their confident identification. Here, we present an open-source software tool designed to detect DUX4-r from short-read, whole-genome sequencing (WGS) data. Evaluation on a cohort of 210 paediatric ALL cases showed that our method detects all known, as well as previously unidentified, cases of IGH::DUX4 and rearrangements with other partner genes. These findings demonstrate the possibility of robustly detecting DUX4-r using WGS in the routine clinical setting.

Humans

Detection of House Dust Mite-derived DNA in Human Lung Tumors by Whole-Genome Sequencing.

Lung cancer in never-smokers (LCINS) accounts for an increasing proportion of lung cancer cases, yet its risk factors remain poorly understood. House dust mites (HDM) are common aeroallergens that induce airway inflammation, but their potential contribution to lung cancer is unknown. We analyzed unmapped whole-genome sequencing reads from 783 lung cancers from the Sherlock-Lung (n = 621 never-smokers) and EAGLE (n = 162 smokers) cohorts, including 328 matched adjacent normal lung tissues. After removal of human sequences, reads were aligned to reference genomes from the two major HDM species and confirmed by BLAST. Samples with top BLAST matches were classified as HDM-detected. Associations between HDM detection and genomic, microbiome, and bulk RNA-seq-derived immune features were evaluated. HDM-derived DNA was detected at low abundance in a subset of tumors and adjacent normal tissues, with higher detection frequencies in tumors than matched normal tissues and in smokers than never-smokers. In LCINS tumors, HDM detection was not associated with tumor mutational burden or recurrent driver alterations but was associated with modest differences in immune cell composition and a limited but reproducible bacterial co-detection pattern. These findings provide a foundation for investigating aeroallergen-derived DNA signatures and their potential relationship to the lung tumor microenvironment.

Environmental exposure

Whole-Genome Sequencing Reveals Co-Infection with Bovine Viral Diarrhea Virus, Bovine Enterovirus, and Caprine Parainfluenza Virus Type 3 in a Calf from a Cattle Herd in Xizang, China.

Although mixed viral infections are increasingly recognized as contributors to bovine diarrhea syndrome, diagnosing such co-infections remains challenging, particularly in high-altitude regions where surveillance is limited. In July 2024, a calf presenting with severe diarrhea and respiratory distress was identified on a cattle farm in Linzhi, Xizang, China. Using unbiased whole-genome sequencing (WGS) of the fecal sample, we assembled near-complete genomes of three distinct RNA viruses: two bovine viral diarrhea virus type 1 (BVDV-1) strains (subtypes 1v and 1q, designated BVDV-1/XZ87 and XZ87), one bovine enterovirus (genotype EV-E, designated BEV/XZ87), and one caprine parainfluenza virus type 3 (CPIV3/XZ87). The CPIV3/XZ87 genome exhibited 99.9% nucleotide identity to the goat-derived GS2017-2 strain from Jiangsu, China, raising the possibility of viral spread through livestock trade. Quantitative real-time PCR (RT-qPCR) confirmed the presence of all three pathogens (Ct values: 24.78 for BEV, 25.98 for CPIV3, and 31.28 for BVDV). This study provides the genomic evidence of a triple co-infection involving BVDV-1, BEV, and CPIV3 in Xizang. It illustrates the potential of WGS for unbiased pathogen detection in complex clinical specimens. The near-complete genomes generated here fill critical gaps in the virological surveillance of this epidemiologically under-sampled high-altitude region.

bovine enterovirus

Genetic Heterogeneity of Inborn Errors of Immunity Revealed by Whole-Genome Sequencing: Insights from a Russian Patient Cohort.

Identifying genetic cause(s) is a key step for management and treatment of patients with inborn errors of immunity (IEI). Here, in an observational cross-sectional genomic study, we analyzed whole-genome sequencing (WGS) data of 72 IEI patients from Saint Petersburg and Northwestern Russia: 42 patients with common variable immunodeficiency (CVID)-like phenotypes, 6 patients with clinically diagnosed X-linked agammaglobulinemia (XLA or Bruton's disease), and 24 patients with other forms of IEI. Causative pathogenic and likely pathogenic variants in BTK, CYBB, CHD7, AIRE, ATM, SBDS, NFKB1, and CTLA4 genes were identified in 14 (19%) patients. Variants of uncertain significance that could be linked to observed clinical phenotypes were detected in 6 patients. These included a BTK variant in a patient with Bruton's disease, variants in SH2D1A, SOCS1, and IKBKB in patients with CVID, and variants in CARD11 and CD40LG in patients with other forms of IEI. Additional rare variants that were mostly unique to individual patients were found in multiple IEI genes from the International Union of Immunological Societies (IUIS) Expert Committee 2024 list. In the CVID-like subcohort, pathway-level analysis of these rare variants revealed patterns associated with clinical manifestations. Taken together, our results expand the genetic characterization of an understudied regional IEI cohort, particularly of patients with CVID-like phenotypes, and identify genetic factors that are implicated in or may contribute to the disease.

Humans

Whole-genome sequencing links a Salmonella Newport ST164 outbreak on Fernando de Noronha to prior circulation in the Brazilian poultry supply chain.

Foodborne outbreaks at geographically isolated tourist destinations pose distinctive One Health challenges, combining limited local surveillance capacity, complex intercontinental supply chains, and high visitor turnover. In May 2021, a diarrheal outbreak linked to a gastronomic festival in Fernando de Noronha, that is a remote UNESCO World Heritage island off northeastern Brazil, was attributed to Salmonella enterica serovar Newport ST164. We applied an integrated genomic approach and epidemiological investigation to propose a transmission chain contextualizing and refining case definition of the S. Newport epidemic clone within national and international diversity. Whole-genome sequencing (WGS), SNP-based phylogenomic, pangenome analysis, Salmonella pathogenicity island (SPI) profiling, and resistome characterization was performed on 17 epidemiologically attributed outbreak isolates and 68 contextual genomes from Brazil, France, the United Kingdom, and the United States. The SNP analysis identified a 13 genome clonal core with less than 20 different SNPs demonstrating the possible connection between 9 patient isolates, 2 food isolates, and 2 food handler isolates, consistent with the involvement of colonised kitchen staff in cross-contamination of the ready-to-eat mussel dish. Three poultry isolates in 2020 from a mainland producer, &#x223c;2180&#xa0;km from Fernando de Noronha, differed only 13 to 17 Core-SNPs from the outbreak core, suggesting prior lineage circulation in the supply chain. Pangenome analysis also supports this evidence revealing near-complete genomic overlap of 4544 shared genes within the 5745 gene clusters (99.9%) between outbreak and non-outbreak backgrounds that mostly differentiate by a defense/prophage-associated accessory module. The resistome comprised intrinsic efflux determinants without acquired resistance and showed 35.3% of intermediate ciprofloxacin susceptibility. This One Health based study provides a WGS genomic reconstruction of a S. Newport ST164 outbreak at a remote tourist island, supporting the possibility of circulation from poultry-associated mainland reservoirs and findings consistent with cross-contamination at a gastronomic seafood festival.

Brazil

Whole-Genome Sequencing of Feline Uropathogens Reveals Multidrug Resistance and Zoonotic Potential in Domestic Cats in Tunisia.

BACKGROUND: Urinary tract infections (UTIs) in cats are increasingly recognized as clinically relevant conditions frequently associated with multidrug-resistant (MDR) bacteria of potential zoonotic origin, yet genomic data on feline uropathogens remain scarce in Tunisia. METHODS: We used whole-genome sequencing to characterize seven bacterial isolates recovered from six cats with clinical signs of UTI: Mammaliicoccus lentus (n = 2), Staphylococcus schleiferi (n = 1), Mammaliicoccus sciuri (n = 1), Enterococcus faecalis (n = 1), Enterococcus casseliflavus (n = 1), and Klebsiella aerogenes (n = 1). RESULTS: Resistome analysis revealed determinants conferring resistance to &#x3b2;-lactams (blaZ, blaCMY-132), methicillin (mecC-type), macrolides (erm(43), ermB), tetracyclines (tet(M), tet(45), tetB), fosfomycins (fosI, fosB, fosA5), and aminoglycosides (aac(6'), aph(3')-IIIa, aph(6)-Id), alongside efflux pump genes (efrA, sepA, sdrM, oqxA, KpnE/F/G), vancomycin-operon genes (vanT, vanY, vanC, vanG), and biofilm/biocide-tolerance genes (salB, qacG). Notably, M. lentus S104 carried mecC-type elements, the first such report in Tunisia, while K. aerogenes displayed an extensive MDR profile, including blaCMY-132 and fosA5. Multilocus sequence typing/ribosomal multilocus sequence typing (MLST/rMLST) identified diverse lineages, including the internationally distributed E. faecalis ST19 and the rarely reported K. aerogenes ST242. Plasmids were absent in all isolates; a Tn916/1545-type transposon occurred in E. casseliflavus, and clustered regularly interspaced short palindromic repeats (CRISPR)-Cas systems were unevenly distributed. CONCLUSIONS: These findings highlight companion animals as reservoirs of clinically important resistance genes, reinforcing the need for One Health AMR surveillance.

Animals

Accelerated long-read variant calling with Clair3 for whole-genome sequencing.

SUMMARY: The rapid growth of genomic data and increasing adoption of long-read sequencing technologies have rendered variant calling one of the most computationally demanding tasks in genomic analysis. Although deep learning-based methods currently outperform conventional approaches in distinguishing true variants from complex sequencing noise, they impose prohibitive computational and time requirements. To address this limitation, we present a computational framework based on Clair3 that integrates parallelized feature generation, enhanced variant phasing, in-memory read haplotagging, and GPU-accelerated neural network inference to accelerate variant calling. By dynamically optimizing the use of both GPU and CPU resources, our method achieves substantial runtime improvements without compromising accuracy. We evaluated our framework across a range of sequencing depths, diverse samples, and multiple hardware configurations. Our results demonstrate that the optimized pipeline completes variant calling for a 30&#xd7; whole-genome sequence in 12-20&#x2009;minutes using standard computational resources (32 CPU threads and one NVIDIA GPU), and in 12-15&#x2009;minutes on an Apple Mac Studio (32 threads), which is &#x223c;10-20-fold speedup compared with its initial release. In addition to exceptional efficiency, our method maintains state-of-the-art accuracy, achieving SNP F1-scores of 99.32% and 99.70% on 30&#xd7; ONT and PacBio GIAB HG003 datasets, respectively. This work introduces a rapid, accurate, and scalable variant calling framework that effectively supports large-cohort genomic studies and time-sensitive clinical applications. AVAILABILITY AND IMPLEMENTATION: The accelerated implementation of Clair3 is open source and available at: https://github.com/HKU-BAL/Clair3/tree/gpu.

Whole Genome Sequencing

Evaluating 12 automated, whole-genome sequencing analysis pipelines for Mycobacterium tuberculosis complex: a comparative study.

BACKGROUND: Reliance on complex, custom-built bioinformatics pipelines is a barrier to the implementation of whole-genome sequencing (WGS) of Mycobacterium tuberculosis in high-burden settings in some low-income and middle-income countries (LMICs). Automated analysis pipelines could address this inequity in access to WGS-based diagnostics and surveillance. This study aimed to systematically evaluate the performance and usability of publicly available WGS pipelines for M tuberculosis. METHODS: We identified automated M tuberculosis WGS analysis pipelines through searches of PubMed and GitHub from database inception up to Aug 31, 2024. Accuracy, cost, accessibility, and scalability were assessed for each pipeline. We evaluated the accuracy of genotypic drug susceptibility testing (gDST) using publicly available sequences with phenotypic susceptibility data for 12 antituberculosis drugs. We estimated pooled sensitivity and specificity for each pipeline, across all drugs, by conducting a bivariate meta-analysis, with random effects representing between-drug variability. Lineage classifications were compared, and a previously epidemiologically well-characterised dataset was used to compare measures of genomic relatedness. FINDINGS: Among 28 candidate pipelines, 16 were excluded as they were unmaintained and inexecutable. 12 pipelines (11 compatible with Illumina and four compatible with Nanopore), all free to use, were included for evaluation. Six pipelines processed and stored data remotely, but for five of these six, scalability was limited by the need to upload sequences through web portals. For local processing pipelines, scalability was dependent on substantial local computational resources, data storage capacity, and command-line interfaces that limited user-friendliness. Only one of six remote-processing pipelines removed human DNA sequences before server upload. gDST was similarly accurate across ten of 11 Illumina-compatible pipelines and three of four Nanopore-compatible pipelines. All pipelines classified the main lineages consistently, although there were differences at sublineage resolution. Outputs from three of four pipelines reporting genomic relatedness were compatible with commonly cited single nucleotide polymorphism difference thresholds. INTERPRETATION: Numerous automated analysis pipelines capable of enhancing equity in M tuberculosis WGS are available. Given the overall similarities between the pipelines evaluated in this study in terms of gDST performance, lineage classification, and genomic relatedness inference, non-functional attributes such as availability, accessibility, scalability, and privacy could represent the point of difference for prospective users in LMICs with a high burden of tuberculosis. FUNDING: The Rhodes Trust, Wellcome, Ellison Institute of Technology, and the UK National Institute for Health and Care Research Oxford Biomedical Research Centre.

Mycobacterium tuberculosis

Novel Genetic Loci in Early-Onset Gout Derived From Whole-Genome Sequencing of an Adolescent Gout Cohort.

OBJECTIVE: Mechanisms underlying the adolescent-onset and early-onset gout are unclear. This study aimed to discover variants associated with early-onset gout. METHODS: We conducted whole-genome sequencing in a discovery adolescent-onset gout cohort of 905 individuals (gout onset 12 to 19 years) to discover common and low-frequency single-nucleotide variants (SNVs) associated with gout. Candidate common SNVs were genotyped in an early-onset gout cohort of 2,834 individuals (gout onset &#x2264;30 years old), and meta-analysis was performed with the discovery and replication cohorts to identify loci associated with early-onset gout. Transcriptome and epigenomic analyses, quantitative real-time polymerase chain reaction and RNA sequencing in human peripheral blood leukocytes, and knock-down experiments in human THP-1 macrophage cells investigated the regulation and function of candidate gene RCOR1. RESULTS: In addition to ABCG2, a urate transporter previously linked to pediatric-onset and early-onset gout, we identified two novel loci (Pmeta < 5.0 &#xd7; 10-8): rs12887440 (RCOR1) and rs35213808 (FSTL5-MIR4454). Additionally, we found associations at ABCG2 and SLC22A12 that were driven by low-frequency SNVs. SNVs in RCOR1 were linked to elevated blood leukocyte messenger RNA levels. THP-1 macrophage culture studies revealed the potential of decreased RCOR1 to suppress gouty inflammation. CONCLUSION: This is the first comprehensive genetic characterization of adolescent-onset gout. The identified risk loci of early-onset gout mediate inflammatory responsiveness to crystals that could mediate gouty arthritis. This study will contribute to risk prediction and therapeutic interventions to prevent adolescent-onset gout.

Humans

An alignment-free strategy for circulating tumor DNA detection and tumor fraction estimation from whole-genome sequencing data.

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed alignment-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth&#x2009;~&#x2009;28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmer's tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

Circulating Tumor DNA

Whole-genome sequencing and analysis of the endophytic fungus Alternaria alternata Y-2 from Leymus chinensis.

To explore the genetic basis and functional potential of beneficial symbiosis between the endophytic fungus Alternaria alternata Y-2 and its host Leymus chinensis, we performed Illumina-based draft whole-genome sequencing and systematic bioinformatic analysis. Although this assembly does not reach telomere-to-telomere completeness, it provides high-quality gene-level information for gene prediction, functional annotation, carbohydrate-active enzyme (CAZyme) identification, and secondary metabolite biosynthetic gene cluster analysis. The final genome size of A. alternata Y-2 was 34,383,676&#xa0;bp with a GC content of 51.0%, containing 12,724 predicted protein-coding genes, 90 tRNAs, and 12 rRNAs. BUSCO assessment showed 98.9% completeness, supporting the high quality of this draft genome. A total of 12,627 genes were successfully annotated in the NCBI NR database, and 17,183 genes were functionally categorized using GO terms. In total, 448 CAZyme genes and 21 secondary metabolite biosynthetic gene clusters were identified, which are potentially involved in lignocellulose degradation, cellular redox homeostasis and biosynthesis of bioactive metabolites. Based on ITS sequence alignment, NR annotation, and phylogenetic analysis of single-copy orthologous genes, the strain was confidently identified as A. alternata. This study firstly reports the draft genome of an endophytic A. alternata strain derived from L. chinensis and provides valuable genetic resources for exploring the endophytic lifestyle, stress tolerance, and bioactive metabolite potential of this fungus.

Alternaria

Whole-genome sequencing reveals hidden antimicrobial resistance genes in phenotypically susceptible probiotic candidate lactic acid bacteria.

Phenotypic assays commonly used to evaluate probiotic safety may fail to detect clinically relevant antimicrobial resistance (AMR), potentially allowing genetically concerning strains to appear acceptable based on MIC testing alone. To explore this issue, we applied whole-genome sequencing (WGS) to three lactic acid bacteria (LAB) isolates previously identified as probiotic candidates based on acid and bile tolerance, antagonism against enteric pathogens, and biofilm formation in vitro: Lactiplantibacillus plantarum L25F and L22F (from pigs) and Ligilactobacillus salivarius AF2319 (from a chicken). Genome annotation identified extensive repertoires of probiotic-associated genes (46-47 per strain) linked to stress tolerance, adhesion, immunomodulation, and quorum sensing, supporting functional potential. The two L. plantarum strains exhibited broader predicted metabolic capacities than L. salivarius AF2319. However, genomic analysis revealed acquired AMR genes with complex genotype-phenotype relationships not fully apparent from phenotypic testing. The L. plantarum strains harbored lnu(A) (99.79% identity) on extrachromosomal DNA, conferring the L-phenotype (lincomycin resistance, clindamycin susceptibility); clindamycin MICs (1&#xa0;mg/L) were concordant with this genotype, though lincomycin MICs were not determined. L. salivarius AF2319 carried tet(M), tet(L), and erm(C) (99.48%, 99.49%, and 99.45% identity by ResFinder, respectively) on extrachromosomal DNA; notably, the erythromycin MIC (1&#xa0;mg/L) was precisely at the EFSA breakpoint (&#x2264;&#x2009;1&#xa0;mg/L), representing borderline genotype-phenotype discordance potentially due to silent gene expression. Under current EFSA QPS criteria, these acquired ARGs would preclude all three strains from approval as probiotic feed additives despite favorable functional profiles, underscoring the indispensable role of WGS-based AMR gene detection in modern probiotic safety evaluation.

Probiotics

Whole-genome sequencing, strain composition, and predicted antimicrobial resistance of Streptococcus pneumoniae causing invasive disease in England in 2017-20: a prospective national surveillance study.

BACKGROUND: Surveillance of the invasive disease burden caused by Streptococcus pneumoniae in England is performed by the UK Health Security Agency (UKHSA). In 2017, UKHSA switched from phenotypic methods to whole-genome sequencing (WGS) approaches for pneumococcal surveillance. Here, we present the first results of national WGS surveillance, up to the start of the COVID-19 pandemic, with the aim of describing the population genomics of this important pathogen. METHODS: We examined prospective national surveillance data from England, using bacterial isolates from cases of invasive pneumococcal disease (IPD) submitted to the national reference laboratory at UKHSA. A bioinformatic pipeline was developed to quality control WGS data and routinely report species and serotype. We assembled isolate data, assigned global pneumococcal sequencing clusters (GPSCs), and predicted antimicrobial resistance (AMR) profiles for isolates that passed further quality control. We collected additional data on patient outcomes and characteristics using enhanced surveillance questionnaires completed by patients' general practitioners. We used logistic regression analysis to assess the effects of various genomic and patient characteristics on the outcomes of IPD. FINDINGS: In England, between July 1, 2017, and Feb 29, 2020, there were 15&#x2009;400 cases of IPD. From these cases, 13&#x2009;749 (89&#xb7;3%) isolates were sequenced, passed quality control, and were included in analyses. Serotype diversity was high during the study period, with 2751 (20%) isolates serotyped as 13-valent pneumococcal conjugate vaccine (PCV13) types, whereas serotype 8 was the most prevalent serotype (n=3074 [22&#xb7;4%]) overall. There were 157 GPSCs within the collection, with GSPC3 the most common, encompassing 98&#xb7;7% (3033 of 3074) of serotype 8 isolates. Most isolates (n=10&#x2009;198 [74&#xb7;2%]) did not contain AMR-associated genes. Resistance to co-trimoxazole was the most frequently predicted resistance (n=2331 [17%]), followed by resistance to tetracycline (n=1199 [8&#xb7;7%]) and &#x3b2;-lactams (n=1149 [8&#xb7;4%]). Logistic regression analysis found the presence of AMR-associated genes significantly increased the odds of patient death (odds ratio 1&#xb7;18, 95% CI 1&#xb7;01-1&#xb7;38). Some GPSCs were also associated with a significant increase in the odds of patient death, such as GPSC12 (1&#xb7;88, 1&#xb7;48-2&#xb7;38). Isolates from 2018 were associated with a significant increase in the odds of patient death (1&#xb7;12, 1&#xb7;00-1&#xb7;25), whereas younger patient age was significantly associated with a reduction in the odds of patient death compared with being aged 85 years or older. INTERPRETATION: WGS-based surveillance has allowed us to interrogate country-wide population dynamics driving changes in pneumococcal serotype frequency. Here, we observe a stable but diverse population before the COVID-19 pandemic restrictions were enforced in England, with low rates of AMR. These findings will provide the baseline for pandemic and post-pandemic data, to collectively inform implementation and development of the vaccination programme within the country. FUNDING: None.

Streptococcus pneumoniae

Large-scale whole-genome sequencing reveals the landscape and health implications of de novo mutations.

De novo mutations (DNMs) are an important source of congenital diseases. With delayed parenthood and assisted reproductive technology (ART) use increasing, it is essential to elucidate how these reproductive factors influence DNMs and whether resulting mutations influence offspring health. Here we performed whole-genome sequencing of 24,030 individuals from 7,851 parent-offspring families, identifying 390,924 de novo single-nucleotide variants (dnSNVs). Paternal and maternal aging exhibited distinct mutational patterns, with maternal DNM accumulation accelerating at advanced ages. Increased paternal dnSNVs partially accounted for the association between advanced parental age and shorter gestational duration. Moreover, ART showed age-independent, procedure-specific effects: intracytoplasmic sperm injection (ICSI) and ovarian stimulation were associated with increased paternal and maternal dnSNVs, respectively, and ICSI-associated paternal dnSNVs also partially accounted for the association between ICSI and shorter gestational duration. In vitro embryo manipulation was associated with increased early post-zygotic mosaic mutations, particularly C&#x2009;>&#x2009;A substitutions linked to delayed neurocognitive development at 1&#x2009;year. Collectively, these findings advance understanding of the determinants and consequences of de novo mutagenesis.

Journal Article