Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “WGS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Diagnostic and clinical utility of exome sequencing and chromosomal microarray in children with GDD/iD: a meta-analysis.

BACKGROUND: Global developmental delay/intellectual disability (GDD/ID) is among the most common neurodevelopmental disorders, with up to half of cases are attributed to genetic factors. Chromosome microarray (CMA) has traditionally been the primary genetic test for idiopathic GDD/ID. However, whole exome sequencing (WES) and whole genome sequencing (WGS) have recently emerged, substantially increasing diagnostic yields in these populations. METHODS: We conducted a comprehensive literature search of PubMed, Scopus, EMBASE, and the Cochrane Library from inception to April 29, 2025. Studies reporting the diagnostic utility of these tests in children with GDD/ID were included and analyzed. RESULTS: A total of 102 studies, comprising 55,752 children, were reviewed. The pooled diagnostic yield of WES was 0.37 (95% CI: 0.33-0.41; I2 = 93%), significantly higher than that of CMA at 0.19 (95% CI: 0.16-0.21; I2 = 95%). Subgroup analyses showed that WES yielded significantly higher diagnostic rates than CMA in both same-sample comparisons (OR = 2.27, 95% CI: 1.08-4.78) and different-sample comparisons (OR = 1.65, 95% CI: 1.15-2.37). Only one study evaluated WGS, reporting a diagnostic yield of 0.27. Meta-regression revealed a significant association between CMA diagnostic yield and the proportion of male participants (p&#x2009;<&#x2009;0.01), but not with WES. No significant difference in diagnostic utility was observed between isolated GDD/ID and GDD/ID with comorbidities. CONCLUSION: In children with unexplained GDD/ID, WES demonstrates superior diagnostic and clinical utility compared to CMA. Incorporating WES as a first-line investigation in the diagnostic evaluation of GDD/ID may be warranted.

Humans↗

Responding to a protracted tuberculosis outbreak: lessons from multiple rounds of investigation in a Chinese boarding school.

PURPOSE: This study analysed a multi-semester pulmonary tuberculosis (PTB) cluster outbreak in a Chinese boarding school to provide evidence for future epidemic control. METHODS: Contacts were screened via symptoms, infection tests and chest radiography. Screening expanded progressively from close contacts to same-floor contacts, then all students and staff. Whole-genome sequencing (WGS) with single nucleotide polymorphism (SNP) and bioinformatics analysis was used for lineage classification, transmission clustering (&#x2264;12 SNPs defining a cluster) and drug resistance prediction. RESULTS: From 2020 to 2022, 20 students were diagnosed with PTB, half laboratory-confirmed. Most cases clustered in class 16 and were epidemiologically linked to the primary case (case 0), who had household PTB exposure. Case 0 and case 1 had diagnostic delays exceeding 3 and 6&#xa0;months, respectively. WGS of five isolates (case 1, 3, 4, 9 and 10) collected over three semesters showed all belonged to lineage 2 and differed by &#x2264;12 SNPs, confirming the same transmission chain. The infection rate in class 16 (46.34%) was significantly higher than other case classes (19.05%) and classes without cases (8.27%) (&#x3c7;2&#xa0;=&#xa0;61.169, p&#xa0;<&#xa0;0.001). No new cases were detected during a one-year follow-up of students involved in the outbreak after the final round of screening, nor among household contacts of all cases followed up to the present. CONCLUSIONS: Lack of entry health examinations facilitated the outbreak. Delayed diagnosis, incomplete contact screening and absence of preventive treatment led to cross-semester persistence. The infection rate disparity confirms class 16 as the outbreak epicentre. Improving community case management, extending contact follow-up and enhancing cluster outbreak measures are recommended to prevent future outbreaks.

Humans↗

A preprocessor for shotgun assembly of large genomes.

The whole-genome shotgun (WGS) assembly technique has been remarkably successful in efforts to determine the sequence of bases that make up a genome. WGS assembly begins with a large collection of short fragments that have been selected at random from a genome. The sequence of bases at each end of the fragment is determined, albeit imprecisely, resulting in a sequence of letters called a "read." Each letter in a read is assigned a quality value, which estimates the probability that a sequencing error occurred in determining that letter. Reads are typically cut off after about 500 letters, where sequencing errors become endemic. We report on a set of procedures that (1) corrects most of the sequencing errors, (2) changes quality values accordingly, and (3) produces a list of "overlaps," i.e., pairs of reads that plausibly come from overlapping parts of the genome. Our procedures, which we call collectively the "UMD Overlapper," can be run iteratively and as a preprocessor for other assemblers. We tested the UMD Overlapper on Celera's Drosophila reads. When we replaced Celera's overlap procedures in the front end of their assembler, it was able to produce a significantly improved genome.

Animals↗

Efficient storage and regression computation for population-scale genome sequencing studies.

MOTIVATION: The growing availability of large-scale population biobanks has the potential to significantly advance our understanding of human health and disease. However, the massive computational and storage demands of whole genome sequencing (WGS) data pose serious challenges, particularly for underfunded institutions or researchers in developing countries. This disparity in resources can limit equitable access to cutting-edge genetic research. RESULTS: We present novel algorithms and regression methods that dramatically reduce both computation time and storage requirements for WGS studies, with particular attention to rare variant representation. By integrating these approaches into PLINK 2.0, we demonstrate substantial gains in efficiency without compromising analytical accuracy. In an exome-wide association analysis of 19.4 million variants for the body mass index phenotype in 125&#xa0;077 individuals (AllofUs project data), we reduced runtime from 695.35&#x2009;min (11.5&#x2009;h) on a single machine to 1.57&#x2009;min with 30 GB of memory and 50 threads (or 8.67&#x2009;min with 4 threads). Additionally, the framework supports multi-phenotype analyses, further enhancing its flexibility. AVAILABILITY AND IMPLEMENTATION: Our optimized methods are fully integrated into PLINK 2.0 and can be accessed at: https://www.cog-genomics.org/plink/2.0/.

Humans↗

nf-core/pacvar: a pipeline for analyzing long-read PacBio whole genome and repeat expansion sequencing data.

MOTIVATION: Pacific Biosciences (PacBio) single-molecule, long-read sequencing enables whole genome annotation and the characterization of 20 complex repetitive repeat regions, especially relevant to neurodegenerative diseases, through their PureTarget panel. Long-read whole-genome sequencing (WGS) also allows for the detection of structural variants that would be difficult to detect with traditional short-read sequencing. However, the raw unaligned Binary Alignment Map data need to be processed before analysis. There is a need for an intuitive comprehensive bioinformatic pipeline that can analyze these data. RESULTS: We present nf-core/pacvar, a comprehensive pipeline for analyzing both PacBio single-molecule PureTarget and WGS data that demultiplexes and parallelizes pre-processing, variant calling and repeat characterization. nf-core/pacvar is compatible with little configuration and has few dependencies. This pipeline enables rapid end-to-end, parallel processing of PacBio single-molecule whole genome and targeted repeat expansion sequencing. AVAILABILITY AND IMPLEMENTATION: nf-core/pacvar is available on nf-core website (https://nf-co.re/pacvar/) and on github (https://github.com/nf-core/pacvar) under MIT License (DOI: 10.5281/zenodo.14813048).

Software↗

Large-scale simulation of coverage and error rate tradeoffs for cancer detection in cell-free DNA whole-genome sequencing.

MOTIVATION: Cell-free DNA (cfDNA) whole-genome sequencing (WGS) is a promising approach for detecting cancer recurrence. It enables cancer detection by identifying all tumor-derived cfDNA (ctDNA) molecules carrying somatic single nucleotide variants (sSNVs). While ideally, a sequencing platform should be highly accurate for reliable ctDNA detection, in reality, all sequencing platforms introduce sequencing errors that generate false positives indistinguishable from true SNVs. Understanding how sequencing parameters influence ctDNA detection sensitivity at low tumor fractions (TFs) in cfDNA samples is essential for guiding sequencing strategies in clinical contexts. To model cfDNA sequencing for tumor detection, which contains asymmetric noise and multiple interacting parameters, analytical modeling is intractable, motivating large-scale parallelized simulation. RESULTS: We developed a simulation framework to generate in silico cfDNA data across 10 cancer types. In total, 480 million cfDNA samples were simulated from tumor WGS profiles. Overall, the lowest detectable TF differs substantially between cancer types under identical sequencing conditions due to variations in mutational load. For cancers with high mutational load, 3&#xd7; coverage with low-error techniques reliably detects TFs below 0.1%. In contrast, cancers with low mutational load require at least six-fold higher coverage to achieve comparable detection thresholds. Increasing sequencing quality scores from Q30 to Q55 at 30&#xd7; coverage further enhances sensitivity, enabling detection of TFs as low as 1&#x2009;&#xd7;&#x2009;10-5. This study provides a comprehensive framework for optimizing sequencing parameters, offering valuable guidance for tailoring future technology development for specific cancer types and clinical applications. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/UMCUGenetics/cfdetect/tree/main.

Whole Genome Sequencing↗

Whole-genome Sequence Analysis Revealed Novel Subjective Cognitive Decline-associated Genes in 10,763 Chinese.

Subjective cognitive decline (SCD) is widely regarded as a potential preclinical stage of Alzheimer's disease (AD), yet its genetic basis remains poorly understood. To address this gap, we investigated genetic biomarkers associated with SCD using whole-genome sequencing (WGS) in 10,763 Chinese participants from the Healthy Zhejiang One Million People Cohort (HOPE Cohort). The discovery stage included 9284 samples, with 1479 samples used for validation. Using a two-stage design, we systematically investigated both common and rare variants associated with SCD. In rare variant analyses, we identified and replicated an association between the upstream region of SEPHS2 and SCD. SEPHS2 is involved in selenophosphate synthesis, and a Mendelian randomization analysis reveals that its expression levels in both blood and brain cerebellum are associated with AD. Additionally, we identified CLVS2, which encodes a protein primarily expressed in neuronal cells, as a potential regulator for SCD based on missense rare variants. Multi-omics evidence suggests that both SEPHS2 and CLVS2 may play roles in neurodegenerative diseases. For common variants, we validated 8 known loci related to cognitive decline, 3 of which originated from the only existing SCD genetic study conducted under a migraine background. Overall, our WGS-based study fills the gap in SCD research by providing vital genetic evidence from an East Asian population and offers insights into the pathogenic mechanisms of SCD.

Aged↗

Characterization of carbapenem-resistant Pseudomonas aeruginosa in Canadian hospitals: 6&#x2005;years of the CANWARD study (2018-23).

OBJECTIVES: Antimicrobial resistance in Pseudomonas aeruginosa is of increasing concern in Canada, leading to limited treatment options and poor clinical outcomes. Herein we characterized carbapenem-resistant P. aeruginosa identified through the Canadian national surveillance program CANWARD. METHODS: Antimicrobial susceptibility for 1725 P. aeruginosa isolates was assessed using broth microdilution and 2024 CLSI breakpoints. WGS of carbapenem-resistant isolates was used to identify STs, resistance and virulence markers. Genetic relatedness was further assessed using cgMLST for select STs. RESULTS: From 2018 to 2023, CANWARD collected 1725 P. aeruginosa isolates, of which 371 (21.5%) were carbapenem-resistant. The majority of carbapenem-resistant P. aeruginosa were isolated from respiratory specimens of male patients aged 18-65&#x2005;years living in central Canada. Only 0.8% (n&#x200a;=&#x200a;3) of the carbapenem-resistant isolates harboured a carbapenemase gene. WGS identified mutations associated with OprD dysfunction, MexAB-OprM efflux and AmpC overexpression in 73.6%, 1.1% and 4.9% of isolates, respectively. Most isolates (98.1%) harboured at least one of the following class D &#x3b2;-lactamase genes: OXA-2, OXA-5, OXA-10 or OXA-50-like subfamily. Wide genetic diversity was observed with 151 different STs identified. The most common STs were ST17 (4.6%), ST27 (4.6%) and high-risk clones ST235 (4.3%), ST244 (3.5%), ST253 (6.4%) and ST357 (2.7%). cgMLST clusters were identified amongst 34.9% of the high-risk clones, suggesting clonal dissemination. CONCLUSIONS: Currently, >20% of clinical isolates of P. aeruginosa in Canada are carbapenem-resistant. Genetic evidence indicates that clonal dissemination of high-risk clones is occurring in Canada. High-risk clones are virulent and often MDR. Continued surveillance of P. aeruginosa is important.

Pseudomonas aeruginosa↗

Genomic structure of class 1 and 2 integrons in non-typhoidal Salmonella isolated from food animals and related meat products in the USA.

OBJECTIVES: Integrons facilitate the capture and expression of exogenous genes, including antimicrobial resistance (AMR) genes. This study aimed to detect the presence of integrons, examine their genomic structure and location, and analyse integron-associated AMR, virulence and stress response genes in Salmonella using WGS. METHODS: WGS data from 193 Salmonella strains, representing 38 serotypes isolated from food animals and related meat products (2001-2019), were analysed using bioinformatic tools to assess integron presence and characterize their genomic architectures. RESULTS: Of 193 isolates, 116 (60.1%) harboured class 1 and/or class 2 integrons. Class 1 integrons alone were detected in 105 isolates, with some containing multiple copies. One S. Infantis isolate harboured only class 2 integrons, whereas 10 others contained both classes. No class 3-5 integrons were found. Twenty-seven class 1 integrons were chromosomal; the rest were plasmid-associated, linked to various plasmid incompatibility (Inc) types. Sixty-nine distinct AMR genes conferring resistance to 11 antimicrobial classes were found in integron cassettes or integron-associated plasmids. Genes linked to resistance to quaternary ammonium compounds and heavy metals, as well as ISs and transposons, were also identified. Significant virulence and stress response genes and proteins such as groES-groEL, LysR and EAL (glutamate, alanine and leucine) were common in integron cassettes. CONCLUSIONS: Class 1 integrons are prevalent in MDR Salmonella isolates from food animals and related meat products and are linked to diverse plasmid types. Their association with AMR, virulence and stress response genes underscores their role in AMR dissemination, and bacterial adaptation and pathogenicity.

Integrons↗

Phenotypic and genotypic profiles of clinical isolates of various Nocardia species to carbapenems and fluoroquinolones.

OBJECTIVES: To establish patterns of the antimicrobial susceptibility of Nocardia species to carbapenems and fluoroquinolones and analysis of phenotypic-genotypic correlations. METHODS: Isolates were identified to the species using 16S rRNA, secA1, or rpoB gene sequencing analysis. The antimicrobial susceptibility testing was performed using the broth microdilution method, and WGS was employed to analyse the presence of resistance genes and/or mutations of Nocardia species against carbapenems and fluoroquinolones. RESULTS: Among 143 Nocardia isolates, N. farcinica (27.27%, 39/143) and N. cyriacigeorgica (25.17%, 36/143) were the most common species, followed by N. abscessus Complex (18.88%, 27/143). The MIC90s of the seven carbapenems were 8&#x2005;mg/L for doripenem, 8&#x2005;mg/L for meropenem, 16&#x2005;mg/L for ertapenem, 16&#x2005;mg/L for biapenem, 64&#x2005;mg/L for imipenem, 64&#x2005;mg/L for faropenem and 128&#x2005;mg/L for tebipenem, respectively. The susceptibility rates to imipenem were 76.9% and 88.9% for N. farcinica and N. cyriacigeorgica, respectively, but only 14.3% and 0% for N. otitidiscavarium and N. brasiliensis, respectively. Further, 90% of N. brasiliensis and 50% of N. otitidiscaviarum isolates were susceptible and intermediate to meropenem. WGS identified blaFAR-1 gene in N. farcinica and blaAST-1 gene in N. cyriacigeorgica, respectively. The MIC90s of the four fluoroquinolones were 1&#x2005;mg/L for sitafloxacin, 4&#x2005;mg/L for nemonoxacin, 4&#x2005;mg/L for moxifloxacin and 16&#x2005;mg/L for ciprofloxacin, respectively. The susceptibility rate of Nocardia species to ciprofloxacin was low except for N. farcinica. The resistance to fluoroquinolones arise from mutations in the gyrA gene. CONCLUSIONS: Nocardia spp. exhibited varying patterns of susceptibility to carbapenems and fluoroquinolones respectively. Importantly, different Nocardia spp. exhibited different patterns of susceptibility to carbapenems and fluoroquinolones, respectively.

Nocardia↗

Co-existence of the oxazolidinone resistance genes cfr and optrA on a novel multiresistance plasmid from a methicillin-resistant Macrococcoides bohemicum strain.

OBJECTIVES: To identify and characterize the oxazolidinone resistance genes cfr and optrA from a methicillin-resistant Macrococcoides bohemicum strain of chicken origin. METHODS: The presence of mobile oxazolidinone resistance genes was detected by PCR. Antimicrobial susceptibility testing was conducted by broth microdilution. Transfer experiments were carried out to evaluate horizontal transferability of the plasmid. WGS was performed using a combination of Illumina NovaSeq/Oxford Nanopore PromethION platforms. RESULTS: The M. bohemicum strain HLJ23 exhibited an MDR phenotype and was positive for both cfr and optrA genes. WGS revealed that the genes cfr and optrA co-exist on the novel MDR plasmid pHLJ23-71kb. Although conjugation experiments were unsuccessful, plasmid pHLJ23-71kb could be transferred to Staphylococcus aureus RN4220 by electrotransformation. Genetic context analysis showed that the cfr and optrA together with another four antimicrobial resistance genes are located in an MDR region on plasmid pHLJ23-71kb. Sequence analysis suggested that this MDR region possibly originated from Mammaliicoccus or Staphylococcus spp. CONCLUSIONS: To the best of our knowledge, this study represents the first report of the oxazolidinone resistance genes cfr and optrA in the genus Macrococcoides. Furthermore, attention should be paid to the exchange of resistance determinants between members of the genera Staphylococcus, Mammaliicoccus and Macrococcoides.

Plasmids↗

Whole-genome automated assembly pipeline for Chlamydia trachomatis strains from reference, in vitro and clinical samples using the integrated CtGAP pipeline.

Whole genome sequencing (WGS) is pivotal for the molecular characterization of Chlamydia trachomatis (Ct)-the leading bacterial cause of sexually transmitted infections and infectious blindness worldwide. Ct WGS can inform epidemiologic, public health and outbreak investigations of these human-restricted pathogens. However, challenges persist in generating high-quality genomes for downstream analyses given its obligate intracellular nature and difficulty with in vitro propagation. No single tool exists for the entirety of Ct genome assembly, necessitating the adaptation of multiple programs with varying success. Compounding this issue is the absence of reliable Ct reference strain genomes. We, therefore, developed CtGAP-Chlamydia trachomatisGenome Assembly Pipeline-as an integrated 'one-stop-shop' pipeline for assembly and characterization of Ct genome sequencing data from various sources including isolates, in vitro samples, clinical swabs and urine. CtGAP, written in Snakemake, enables read quality statistics output, adapter and quality trimming, host read removal, de novo and reference-guided assembly, contig scaffolding, selective ompA, multi-locus-sequence and plasmid typing, phylogenetic tree construction, and recombinant genome identification. Twenty Ct reference genomes were also generated. Successfully validated on a diverse collection of 363 samples containing Ct, CtGAP represents a novel pipeline requiring minimal bioinformatics expertise with easy adaptation for use with other bacterial species.

Chlamydia trachomatis↗

Amplicon-based analyses of single-nucleotide polymorphisms reveal the genetic structure of a forest insect baculovirus.

Amplicon-based next-generation sequencing (aNGS) is a powerful tool in diagnostics and genetic studies. We developed an aNGS approach to study the population structure of the Lymantria dispar multiple nucleopolyhedrovirus (LdMNPV), a specific pathogen of the spongy moth Lymantria dispar, a devastating lepidopteran pest in European, Asian, and American deciduous forests. Naturally occurring pathogens, such as LdMNPV, are frequently reported to cause epizootics and a rapid decline of insect pest populations. DNA samples of pooled LdMNPV-infected larvae from forest regions in Northern Bavaria (Germany) were subjected to whole genome sequencing (WGS) and aNGS optimization. Then, five marker regions were identified in the genome of LdMNPV for PCR amplification, covering 21 highly specific single-nucleotide polymorphism (SNP) positions that enabled comprehensive analysis at the intra- and intersample levels. These markers were used in aNGS analyses of 70 single larvae collected in 12 forest sites, followed by SNP-based hierarchical clustering on principal components (HCPC). This approach identified three LdMNPV population clusters consisting of homogenous (pure) and heterogeneous (mixed) LdMNPV samples. To explain the genetic variability within each sample, a model based on linear optimization was developed and validated by comparing the predictions from aNGS and WGS data. The analyses showed that LdMNPV from Bavarian forests carried genetic variants highly similar to those present in the commercial product Gypchek&#xae;, developed for biocontrol. The distribution of genetic characteristics showed some trends of geographic and temporal prevalence, which are indicative of short-distance and long-distance transmission. The aNGS approach offers a fast, cost-effective, and comprehensive insight into the natural population structure of LdMNPV.

insects↗

Whole -genome survival analysis of 144&#x200a;286 people from the UK Biobank identifies novel loci associated with blood pressure.

This study utilized UK Biobank data from 144&#x200a;286 participants and employed whole-genome sequencing (WGS) data and time-to-event data over a 12-year follow-up period to identify susceptibility in genetic variants associated with hypertension. Following genotype quality control, 6&#x200a;319&#x200a;822 single nucleotide polymorphisms underwent analysis, revealing 31 significant variant-level associations. Among these, 29 were novel - 15 in Fibrillin-2 ( FBN2 ) and 4 in Junctophilin-2 ( JPH2 ). Mendelian randomization utilizing two identified variants (rs17677724 and rs1014754) suggested that a genetically induced decrease in heart FBN2 expression and an increase in adrenal gland JPH2 expression were causally linked to hypertension. Phenome-wide association (PheWAS) analysis using the FinnGen dataset confirmed positive associations of rs17677724 and rs1014754 with hypertension, assessed across 2727 traits in 377&#x200a;277 individuals. Lastly, rs1014754 positively associated with kallistatin, whereas rs17677724 negatively associated with renin in the Fenland study, suggesting a counterregulatory response to high blood pressure. This study, employing WGS data, identified novel genetic loci and potential therapeutic targets for hypertension.

Humans↗

Multiplex PCR assay for the rapid detection of Klebsiella pneumoniae pathotypes.

Introduction. Klebsiella pneumoniae (Kp) is a major cause of nosocomial infections, with its evolving pathotypes including multidrug-resistant, hypervirulent (hvKp) and convergent strains posing significant diagnostic and treatment challenges due to combined antimicrobial resistance and virulence.Gap Statement. While there is a pressing requirement for thorough detection of Kp pathotypes, current assays in resource-limited environments are unable to effectively focus on essential carbapenemase and hypervirulence genes with the necessary reliability and precision.Aim. To develop and validate a multiplex PCR (m-PCR) assay capable of simultaneously detecting Kp isolates including those carrying partial or full virulence markers, alongside antimicrobial resistance.Methodology. In this study, an m-PCR assay was designed and optimized for the simultaneous detection of key biomarkers associated with hypervirulent (rmpA, rmpA2, iucA, peg344 and iroB), carbapenem-resistant (bla NDM, bla OXA-48-like and bla KPC) and convergent Kp pathotypes in clinical isolates. The assay was evaluated on clinical isolates and validated against whole-genome sequencing (WGS) data for accuracy, specificity and sensitivity.Results. The developed m-PCR assay exhibited 100% specificity when compared to WGS data, successfully detecting all target genes without cross-amplification in ATCC control strains. The assay demonstrated high sensitivity, efficiently amplifying bacterial genomes from minimal DNA input as low as 1&#x2009;ng &#xb5;l-1. Additionally, validation through sequencing confirmed the accuracy of detected amplicons.Conclusion. This m-PCR assay offers a rapid, sensitive and specific diagnostic tool for differentiating Kp pathotypes in clinical settings, aiding in timely intervention and improved infection control measures.

Klebsiella pneumoniae↗

Multi-strain carriage and intrahost diversity of Staphylococcus aureus among Indigenous adults in the USA.

Staphylococcus aureus (SA) is an opportunistic pathogen and human commensal that is frequently present in the upper respiratory tract, gastrointestinal tract and skin. While SA can cause diseases ranging from minor skin infections to life-threatening bacteraemia, it can also be carried asymptomatically. Indigenous individuals in the Southwest USA experience high rates of invasive SA disease. As carriage is the most significant risk factor for disease, understanding the dynamics of SA carriage, and in particular co-carriage of multiple strains, is important to develop strategies to prevent transmission in vulnerable communities. Here, we investigated SA co-carriage and intrahost evolution by sampling several colonies from multiple anatomical sites and whole-genome sequencing (WGS) on 310 SA isolates collected from 60 Indigenous adults participating in a cross-sectional carriage study. We assessed the richness and diversity of SA isolates via differences in multilocus sequence type, core-genome SNPs and genome content. Using WGS data, we identified 95 distinct SA intra-subject lineages (ISLs) among 60 participants; co-carriage was detected in 42% (25/60). Notably, two participants each carried four distinct SA ISLs. Variation in antibiotic resistance determinants among carried strains was identified among 42% (25/60) of participants. Lastly, we found unequal distribution of clonal complex by body site, suggesting that certain lineages may be adapted to specific anatomical sites. Together, these findings suggest that co-carriage may occur more frequently than previously appreciated and further our understanding of SA intrahost diversity during carriage, which has implications for surveillance activities and epidemiological investigations.

Humans↗

Development and extensive sequencing of a broadly-consented Genome in a Bottle matched tumor-normal pair.

The Genome in a Bottle Consortium (GIAB), hosted by the National Institute of Standards and Technology (NIST), is developing new matched tumor-normal samples, the first to be explicitly consented for public dissemination of genomic data and cell lines. Here, we describe a comprehensive genomic dataset from the first individual, HG008, including DNA from an adherent, epithelial-like pancreatic ductal adenocarcinoma (PDAC) tumor cell line and matched normal cells from duodenal and pancreatic tissues. Data for the tumor-normal matched samples comes from seventeen distinct state-of-the-art whole genome measurement technologies, including high depth short and long-read bulk whole genome sequencing (WGS), single cell WGS, and Hi-C, and karyotyping. In future publications, these data will be used by the GIAB Consortium to develop matched tumor-normal benchmarks for somatic variant detection. We expect these data to facilitate innovation for whole genome measurement technologies, de novo assembly of tumor and normal genomes, and bioinformatic tools to identify small and structural somatic mutations. This first-of-its-kind broadly consented open-access resource will facilitate further understanding of sequencing methods used for cancer biology.

Journal Article↗

Whole-Genome Sequencing Uncovers Chromosomal and Plasmid-Borne Multidrug Resistance and Virulence Genes in Poultry-Associated Escherichia coli from Nigeria.

BACKGROUND: Broad and unregulated antibiotic use in livestock production, particularly poultry farming, has increased the development and persistence of multidrug-resistant (MDR) bacterial strains in animals. These resistant pathogens and their antibiotic resistance genes (ARGs) can spread to humans through environmental exposure and the food chain, posing serious public health risks. Whole-genome sequencing (WGS), alongside phenotypic antimicrobial susceptibility testing (AST), enables a comprehensive understanding of resistance mechanisms and informs antimicrobial stewardship strategies, particularly in resource-limited settings. AIM: This study aimed to characterize the phenotypic and genotypic antimicrobial resistance profiles, plasmid content, and virulence factors of an MDR E. coli strain (S3) isolated from a poultry farm in Enugu State, Nigeria, to elucidate potential risks to public health and the role of poultry as a reservoir for resistance determinants. METHODS: E. coli strain S3 was isolated from chicken droppings using standard microbiological methods and confirmed by MALDI-TOF mass spectrometry. AST was assessed using disc diffusion and broth microdilution to determine minimum inhibitory concentrations (MICs) for ten antibiotics across multiple classes. WGS was performed with a hybrid approach combining Illumina and Nanopore platforms, followed by genome assembly and annotation. ARGs, plasmid replicons, and virulence factors were identified in silico using AMRFinderPlus, starAMR, RGI/CARD, PlasmidFinder, MOB-suite, and the Virulence Factor Database (VFDB). RESULTS: Phenotypic testing revealed extensive resistance, with complete resistance to six of seven tested antibiotics (cefotaxime, ampicillin, erythromycin, gentamicin, ciprofloxacin, and doxycycline). MICs exceeded clinical breakpoints for multiple classes, confirming an MDR phenotype. Genome analysis indicated a 5.33 Mb genome distributed across five contigs, including one chromosome and four plasmid-associated contigs. The strain harboured numerous ARGs, including bla CTX-M-15, bla OXA-1, bla TEM-1, aac(6')-Ib-cr, aadA5, aph(3")-Ib, sul1/sul2, tet(A), dfrA17, and mph(A), co-localized on plasmids indicative of horizontal gene transfer (HGT) potential. Plasmid types included Col156, IncF, and two rep clusters. Virulence profiling revealed genes associated with adhesion (pap cluster, ECP), iron acquisition (enterobactin, yersiniabactin, aerobactin, heme uptake), and toxins (sat, senB), highlighting the isolate's potential for urinary tract and intestinal infections. CONCLUSION: This study highlights the significant role of poultry-associated bacteria as reservoirs of AMR genes, particularly those harboured on mobile plasmids with potential for HGT. E. coli strain S3 exhibits extensive multidrug resistance and carries a complex plasmid repertoire facilitating horizontal transfer of ARGs. Coupled with a rich virulence gene profile, this strain underscores the public health risk posed by poultry-associated E. coli in Nigeria. These findings demonstrate the urgent need for stringent antimicrobial stewardship, regulatory oversight, and genomic surveillance in poultry production milieus to mitigate the dissemination of MDR pathogens.

Escherichia coli↗