Search PubMedSearch

SEARCH · Search PubMed

Results for “whole-genome- sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Whole-Genome Sequencing Reveals Population Structure, Genetic Diversity, and Selection Signatures in Kazakh Dromedary and Bactrian Camels.

Understanding the genomic basis of environmental adaptation is essential for the conservation and genetic improvement of domestic camels. In this study, we investigated the population structure, genetic diversity, and genomic variation potentially associated with environmental adaptation of Kazakh dromedary and Bactrian camels using whole-genome sequencing. Whole-genome sequencing data were generated for Kazakh camels (15 dromedaries and 16 Bactrian camels) and integrated with 131 publicly available genomes representing camel populations from the Arabian Peninsula, Iran, Xinjiang, Inner Mongolia, and Mongolian wild camels. Population structure, genetic diversity, and genome-wide selection were evaluated using principal component analysis, ADMIXTURE, nucleotide diversity, linkage disequilibrium, runs of homozygosity, genomic inbreeding (FROH), and selection scans based on FST, θπ ratio, and XP-EHH. Population genomic analyses revealed clear differentiation between dromedary and Bactrian camels, whereas Kazakh camel populations exhibited higher nucleotide diversity (θπ = 1.307-1.551 × 10-3), and lower genomic inbreeding (median FROH: 0.037-0.056) than Arabian populations. Genome-wide selection analyses identified MC4R as the prominent candidate gene in Kazakh dromedaries and RYR1 as a prominent candidate gene in Kazakh Bactrian camels. Functional enrichment analyses highlighted pathways related to energy metabolism, thermogenesis, calcium signaling, skeletal muscle function, mitochondrial activity, and oxidative stress response. These findings provide new insights into genomic variation potentially associated with environmental adaptation in Kazakh camels and offer valuable genomic resources for future conservation, breeding, and evolutionary studies.

MC4R

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

Routine methods misidentify Serratia spp.: Limitations of MALDI-TOF MS revealed by whole-genome sequencing.

Accurate species-level identification within the genus Serratia remains challenging due to extensive phenotypic overlap and high genomic relatedness among closely related and recently described taxa. This study presents an evaluation of routine and genome-based identification approaches applied to clinical Serratia isolates, integrating phenotypic assays, MALDI-TOF MS (Bruker Daltonics), 16S rRNA gene sequencing, and Whole-Genome Sequencing (WGS). A total of 103 isolates collected from a teaching hospital were analyzed. WGS was performed on a subset of isolates. Conventional biochemical methods classified all isolates as Serratia marcescens, whereas MALDI-TOF MS identified 60.1% as S. marcescens, 11.6% as S. ureilytica, and 28.1% just at the genus level. Peak analysis from MALDI-TOF MS revealed specific peaks associated with S. marcescens and S. ureilytica, but limited discriminatory power. WGS of six isolates initially identified as S. ureilytica by MALDI-TOF MS revealed reclassification as Serratia sarumanii (n = 5) and Serratia montpellierensis (n = 1), supported by Average Nucleotide Identity (ANI), Average Amino Acid Identity (AAI), and Digital DNA-DNA Hybridization (dDDH) thresholds. In contrast, 16S rRNA analysis showed limited species-level resolution. Phylogenomic and SNP-based analyses confirmed these classifications with strong support. Overall, this study underscores the critical role of high-resolution genomic approaches for precise species identification and highlights the need for continuous expansion and curation of MALDI-TOF MS reference databases to support reliable clinical diagnostics and epidemiological surveillance of emerging Serratia species.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans

A tiled amplicon protocol for culture-free whole-genome sequencing of M. tuberculosis from clinical specimens.

Whole-genome sequencing of Mycobacterium tuberculosis can be a valuable tool for TB surveillance and treatment, providing insights into transmission patterns and comprehensive drug susceptibility testing. However, the slow growth of M. tuberculosis means traditional culture-based sequencing methods can take weeks to return results, which has limited the widespread adoption of these techniques and limited their use in clinical decision-making. Tiled amplicon sequencing is a fast, reliable, and cost-effective method of whole-genome sequencing that can be done directly on clinical specimens and has been implemented at scale in academic and public health laboratories across the world; it was the cornerstone of SARS-CoV-2 sequencing and has been adapted for a wide range of viral pathogens. However, similar methods are not yet available for far larger bacterial genomes. Extending this approach to M. tuberculosis would significantly reduce the cost, labor, and turnaround time for whole-genome sequencing. We designed a tiled amplicon panel consisting of 5,128 primers that covers the entire M. tuberculosis genome, the largest tiled amplicon sequencing panel we are aware of to date. Applying our amplicon panels to clinical samples of sputum, we show the ability to recover whole-genome bacterial sequences without the need for culture. The resulting sequence data can be used to determine M. tuberculosis lineage and reliably identify markers of drug resistance. Using this approach in clinical settings could reduce the time needed for comprehensive drug susceptibility testing from weeks to days and enable genomic epidemiology to be performed at scale, even in resource-limited settings.IMPORTANCEWe have developed and tested an amplicon panel, TB-seq, for the priority pathogen Mycobacterium tuberculosis, demonstrating recovery of near-full genomes directly from patient sputum, including mixed and low-concentration samples. This approach significantly reduces the turnaround time for this slow-growing bacterium while maintaining high accuracy in detecting clinically relevant mutations, including those associated with drug resistance. Given the global burden of tuberculosis and the critical need for faster diagnostic solutions, we believe our method has the potential to improve clinical decision-making and public health strategies.

Mycobacterium tuberculosis

Whole-genome sequencing of 490,640 UK Biobank participants.

Whole-genome sequencing provides an unbiased and complete view of the human genome and enables the discovery of genetic variation without the technical limitations of other genotyping technologies. Here we report on whole-genome sequencing of 490,640 UK Biobank participants, building on previous genotyping effort1. This advance deepens our understanding of how genetics associates with disease biology and further enhances the value of this open resource for the study of human biology and health. Coupling this dataset with rich phenotypic data, we surveyed within- and cross-ancestry genomic associations and identified novel genetic and clinical insights. Although most associations with disease traits were primarily observed in individuals of European ancestries, strong or novel signals were also identified in individuals of African and Asian ancestries. With the improved ability to accurately genotype structural variants and exonic variation in both coding and UTR sequences, we strengthened and revealed novel insights relative to whole-exome sequencing2,3 analyses. This dataset, representing a large collection of whole-genome sequencing&#xa0;data that is available to the UK Biobank research community, will enable advances of our understanding of the human genome, facilitate the discovery of diagnostics and&#xa0;therapeutics with higher efficacy and improved safety profile, and enable precision medicine strategies with the potential to improve global health.

Humans

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans

First-line drug-resistant tuberculosis among children under 15 years in Ethiopia: insights from phenotypic and whole-genome sequencing approaches.

BACKGROUND: Childhood drug-resistant tuberculosis is often underdiagnosed and inadequately characterized due to the paucibacillary nature of the disease. This study aimed to assess resistance to first-line anti-tuberculosis drugs in children using phenotypic drug susceptibility testing and whole-genome sequencing. METHODS: A retrospective-prospective study was conducted on culture-confirmed childhood tuberculosis cases in Ethiopia (2017&#x2013;2023). Phenotypic drug susceptibility testing was performed on 110 Mycobacterium tuberculosis complex isolates. Whole-genome sequencing was completed for 85 of these isolates, which were analyzed using the TB-Profiler and MTBSeq pipelines. We assessed the sensitivity, specificity, predictive values, and kappa agreement of whole-genome sequencing compared with phenotypic drug susceptibility testing. RESULTS: Phenotypic resistance to at least one first-line anti-TB drug was observed in 26/110 (23.6%) of the examined isolates, with isoniazid resistance being the most frequent, 23/110 (20.9%), followed by rifampicin resistance, 18/110 (16.4%). TB-Profiler showed almost perfect agreement with phenotypic drug susceptibility testing for rifampicin (sensitivity 94.4%, kappa&#x2009;=&#x2009;0.96) and isoniazid (sensitivity 91.3%, kappa&#x2009;=&#x2009;0.91), whereas MTBSeq showed slightly lower performance. Both pipelines demonstrated moderate to weak agreement with phenotypic drug susceptibility testing for detecting resistance to ethambutol, pyrazinamide, and streptomycin. The most frequently observed resistance mutations among phenotypically resistant isolates were rpoB (Ser450Leu), katG (S315Thr), embB (Met306Ile), and pncA (C-11&#xa0;A&#x2009;>&#x2009;G) for rifampicin, isoniazid, ethambutol, and pyrazinamide, respectively. Discrepancies between genotypic and phenotypic drug susceptibility testing were observed across all first-line anti-TB drug-resistant isolates, particularly for ethambutol and pyrazinamide. CONCLUSION: We found a high prevalence of isoniazid resistance, along with rifampicin resistance, underscoring the need for early detection in vulnerable groups. Whole-genome sequencing showed good accuracy for these drugs, with TB-Profiler performing best. CLINICAL TRIAL NUMBER: Not applicable.

Humans

X-linked spondyloepiphyseal dysplasia tarda misdiagnosed as growth hormone deficiency: identification of a novel intronic TRAPPC2 variant by whole-genome sequencing.

BACKGROUND: X-linked spondyloepiphyseal dysplasia tarda (SEDT) is a rare skeletal dysplasia caused by pathogenic variants in TRAPPC2 and typically presents in late childhood or adolescence with short-trunk disproportion and vertebral dysplasia. CASE PRESENTATION: We describe a family series centered on an adolescent male initially diagnosed with GHD due to reduced height velocity and subnormal GH stimulation results, who received recombinant human GH (rhGH) therapy for three years with negligible improvement. During puberty, he developed progressive short-trunk disproportion and characteristic radiographic features, including platyspondyly and posterior hump-shaped vertebral endplates, suggestive of SEDT. Whole-exome sequencing (WES) was nondiagnostic, whereas whole-genome sequencing (WGS) identified a novel intronic TRAPPC2 variant, c.239-20_239-12delinsAATGAA, initially classified as a variant of uncertain significance (VUS). Segregation analysis across the family enabled reclassification of the variant to likely pathogenic, confirming X-linked SEDT. The proband's younger brother exhibited earlier radiologic abnormalities and, notably, a favorable response to rhGH, whereas the younger sister-an asymptomatic heterozygous carrier-showed normal spinal morphology, consistent with expected female carrier phenotypes. CONCLUSIONS: This family-based report underscores the generally limited therapeutic effect of rhGH in SEDT while highlighting potential interindividual variability, as evidenced by the younger male sibling's response. It further emphasizes the diagnostic utility of WGS for detecting deep intronic variants missed by WES and the importance of segregation analysis in resolving VUS in rare skeletal dysplasias.

Humans

Genetic diversity and drug resistance profiles of Mycobacterium tuberculosis among Ethiopian children as determined by whole-genome sequencing.

UNLABELLED: Ethiopia ranks 30th among the tuberculosis (TB) burden countries, with children representing a significant yet understudied population group. This study aims to investigate the genetic diversity and drug-resistant profile among Ethiopian children. We included children under 15 years of age diagnosed with culture-confirmed pulmonary TB/drug-resistant TB between January 2017 and June 2023. Phenotypic drug susceptibility testing and whole-genome sequencing were conducted for 85 Mycobacterium tuberculosis (MTB) isolates. Demographic data were combined with genomic information. Lineage 4 was the most dominant (77.6%), while lineage 2 was less common (1%). Within lineage 4, several sub-lineages were identified, with lineage 4.2.2.2 being notably the most predominant (48%). Most of these cases were from Oromia (58%), including the hotspot areas for lineage 4 that were identified at a 99% confidence level. Among 17 MDR/pre-XDR-TB isolates, lineages 3 and 4.2.2.2 were the dominantly observed lineages/sub-lineages, with proportions of 29% and 65%, respectively. Of the 85 cases, 30.5% were drug-resistant TB to at least one of the five first-line anti-TB drugs tested by phenotypic drug susceptibility testing. Of these 26 drug-resistant TB cases, 23 were concordant with whole-genome sequencing characterization. The most frequent resistance mutations to rifampicin were found in the rpoB gene, specifically p.Ser450Leu (88%), followed by isoniazid in the katG gene, p.Ser315Thr (86%). Multidrug-resistant TB was strongly associated with MTB lineages (P = 0.007). This study identified high genetic diversity of M. tuberculosis and related drug-resistance mutations, with a strong concordance between whole-genome sequencing-based predictions and phenotypic drug susceptibility testing. IMPORTANCE: Our findings revealed a high genetic diversity of Mycobacterium tuberculosis among Ethiopian children, with the most common lineage being lineage 4, specifically lineage 4.2.2.2, in which a higher frequency of multidrug-resistant tuberculosis (TB) was observed. Additionally, we identified regional hotspots, suggesting ongoing community transmission. Moreover, whole-genome sequencing demonstrated high concordance with phenotypic drug susceptibility testing and identified mutation genes associated with first- and second-line anti-TB drugs, highlighting its usefulness in providing comprehensive results for resistance detection in children. Thus, it is essential for integrating genomic surveillance into childhood TB and drug resistance control.

Humans

A deep intronic IFT172 variant causing pseudoexon inclusion identified by whole-genome sequencing in nephronophthisis.

Nephronophthisis is an autosomal recessive ciliopathy and a major genetic cause of end-stage kidney disease in children and young adults. Although next-generation sequencing panels have improved diagnostic yield, some patients remain genetically unresolved, partly due to deep intronic variants that disrupt pre-mRNA splicing and are not captured by exon-focused approaches. We report a 13-year-old boy who presented with advanced kidney dysfunction, small renal cysts, and kidney histopathology consistent with nephronophthisis. Targeted gene panel sequencing failed to identify causative pathogenic variants beyond a missense variant of uncertain significance. Whole-genome sequencing subsequently revealed compound heterozygous variants in IFT172 (NM_015662.3): a missense variant (c.4696C > T, p.Arg1566Cys) and a deep intronic variant (c.4915-94A > G). In silico analysis predicted activation of cryptic splice sites leading to inclusion of an 86-bp pseudoexon, which was confirmed by a minigene splicing assay. These findings established a molecular diagnosis of IFT172-related nephronophthisis. To our knowledge, this is the first report demonstrating pseudoexon inclusion in IFT172, thereby expanding its mutational spectrum. Our case underscores the importance of evaluating deep intronic regions using whole-genome sequencing and functional validation in genetically unresolved nephronophthisis.

Humans

Whole-genome sequencing of adenovirus 41 directly from wastewater using nested overlapping PCR and MinION.

Human adenovirus F41 (HAdV-F41) is one of the leading causes of children's acute gastroenteritis and was recently linked to an outbreak of severe acute hepatitis of unknown etiology among children during 2021 to 2022. While most evidence is based on clinical data, wastewater-based epidemiology offers a community-level approach to monitoring circulating strains and enhancing outbreak preparedness. In this study, we developed an overlapping amplicon-based whole-genome sequencing approach to directly detect HAdV-F41 from archived wastewater samples, using nested PCR with 13 primer sets. Archived wastewater samples were collected between 2021 and 2022 from three treatment plants in Seattle, USA. The viral load ranged from 1.2 &#xd7; 103 to 8.4 &#xd7; 103 genome copies per liter. The Oxford Nanopore platform was used for whole-genome sequencing. Complete or partial (>84%) HAdV-F41 genomes were recovered from wastewater samples, with mean coverage depths ranging from 10&#xb3; to 10&#x2075;. The consensus sequences showed more than 99% similarity to reference genomes in the NCBI database. The phylogenetic analysis revealed that 2 sequences clustered within lineage 2a and 11 within lineage 2b, reflecting that at least two sub-lineages were circulating in the community at that time. Our results demonstrate that the overlapping amplicon-based whole-genome sequencing approach using the Oxford Nanopore platform reliably recovers HAdV-F41 genomes from wastewater. This method offers high-resolution genomic surveillance of circulating, clinically relevant HAdV-F41, supporting wastewater-based epidemiology as a valuable tool for detecting emerging variants and strengthening the early warning system for future disease outbreaks.IMPORTANCEHuman adenovirus F41 is a primary cause of childhood gastroenteritis and has been linked to recent outbreaks of severe acute hepatitis in children, yet community-level genomic surveillance of this virus remains limited. This study shows that wastewater can be used to recover nearly complete HAdV-F41 genomes through a targeted overlapping-amplicon sequencing strategy on the Oxford Nanopore platform. By applying this method to archived wastewater samples, we detected the simultaneous circulation of multiple viral lineages in a large city. These findings extend wastewater-based epidemiology beyond SARS-CoV-2 and emphasize its importance for monitoring clinically significant enteric viruses. The method described here offers a scalable tool for tracking viral evolution in communities and enhancing early warning systems for future outbreaks.

Wastewater

Third-generation whole-genome sequencing reveals the role of CNTNAP2 as a tumor suppressor gene in high-risk neuroblastomas.

BACKGROUND: Neuroblastoma is a common and aggressive pediatric sympathetic nervous system tumor. Genomic structural variants (SVs) contribute substantially to neuroblastoma, yet remain under-characterized in high-risk neuroblastomas. We aimed to elucidate neuroblastoma pathogenesis using third-generation whole-genome sequence high-risk cases to identify driver aberrations and explore potential therapeutic strategies. METHODS: We analyzed third-generation whole-genome sequencing data of 20 high-risk neuroblastoma samples and combined the findings with those obtained from the analysis of clinical samples, in vitro models, and public datasets. RESULTS: The contactin-associated protein-like 2 (CNTNAP2) gene was observed to be frequently aberrated because of structural variants in high-risk neuroblastoma samples. CNTNAP2 expression was significantly correlated with favorable histology and could be used to predict prognosis using clinical samples and neuroblastoma datasets. Overexpression and knockdown experiments and transcriptomic analysis revealed that CNTNAP2 was primarily involved in neuronal differentiation and axon guidance pathways; moreover, CNTNAP2 was required for neuroblastoma differentiation and affected cancer stemness. Immunoprecipitation and mass spectrometry revealed that CNTNAP2 interacted with cytoskeletal proteins like drebrin 1 (DBN1) and myosin-heavy chain 9 (MYH9). CNTNAP2 dynamically reorganises actin and microtubules for DBN1-mediated neuronal differentiation. CNTNAP2 also reduces CTNNB1 transcription and &#x3b2;-catenin pathway activation by inhibiting MYH9 nuclear translocation. CNTNAP2 overexpression in neuroblastoma cell lines resulted in cell cycle arrest, decreased cell proliferation and metastasis. CONCLUSIONS: The recurrent loss of CNTNAP2 in neuroblastoma contributes to an aggressive phenotype by impairing neuronal differentiation and increasing cancer stemness. These findings may serve as a foundation for developing therapeutic strategies to overcome barriers to differentiation.

Humans

Large-scale simulation of coverage and error rate tradeoffs for cancer detection in cell-free DNA whole-genome sequencing.

MOTIVATION: Cell-free DNA (cfDNA) whole-genome sequencing (WGS) is a promising approach for detecting cancer recurrence. It enables cancer detection by identifying all tumor-derived cfDNA (ctDNA) molecules carrying somatic single nucleotide variants (sSNVs). While ideally, a sequencing platform should be highly accurate for reliable ctDNA detection, in reality, all sequencing platforms introduce sequencing errors that generate false positives indistinguishable from true SNVs. Understanding how sequencing parameters influence ctDNA detection sensitivity at low tumor fractions (TFs) in cfDNA samples is essential for guiding sequencing strategies in clinical contexts. To model cfDNA sequencing for tumor detection, which contains asymmetric noise and multiple interacting parameters, analytical modeling is intractable, motivating large-scale parallelized simulation. RESULTS: We developed a simulation framework to generate in silico cfDNA data across 10 cancer types. In total, 480 million cfDNA samples were simulated from tumor WGS profiles. Overall, the lowest detectable TF differs substantially between cancer types under identical sequencing conditions due to variations in mutational load. For cancers with high mutational load, 3&#xd7; coverage with low-error techniques reliably detects TFs below 0.1%. In contrast, cancers with low mutational load require at least six-fold higher coverage to achieve comparable detection thresholds. Increasing sequencing quality scores from Q30 to Q55 at 30&#xd7; coverage further enhances sensitivity, enabling detection of TFs as low as 1&#x2009;&#xd7;&#x2009;10-5. This study provides a comprehensive framework for optimizing sequencing parameters, offering valuable guidance for tailoring future technology development for specific cancer types and clinical applications. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/UMCUGenetics/cfdetect/tree/main.

Whole Genome Sequencing

Whole-Genome Sequencing Reveals Virulence and Antimicrobial Resistance Determinants of Lactococcus garvieae Causing Lactococcosis in Cage-Cultured Nile Tilapia (Oreochromis niloticus) in Thailand.

Lactococcosis is an important bacterial disease affecting farmed fish worldwide and is primarily associated with Lactococcus garvieae, Lactococcus petauri, and Lactococcus formosensis. In Thailand, information on L. garvieae infection in tilapia remains limited, particularly regarding genome-based identification, virulence determinants, and antimicrobial resistance profiles. This study characterized two L. garvieae isolates, AAHM-LG2501 and AAHM-LG2509, recovered from a lactococcosis outbreak in cage-cultured Nile tilapia (Oreochromis niloticus) in Ubon Ratchathani province, Thailand. Both isolates exhibited typical phenotypic characteristics of L. garvieae, including Gram-positive cocci, alpha hemolysis, positive capsule staining, and positive carbohydrate fermentation. Whole-genome sequencing confirmed both isolates as L. garvieae, with genome sizes of approximately 1.95 Mb and a G + C content of 38.9%. Genome-based taxonomic analysis supported species identification based on dDDH and ANI values, and both isolates were assigned to sequence type ST95 and serotype I. Virulence factor analysis identified 288 virulence-associated genes representing 97 virulence factors across 14 functional categories. Capsule-associated genes were prominent, together with genes involved in heme uptake, adhesion, hemolysis, stress survival, biofilm formation, and host adaptation. Ten capsule biosynthesis genes, including cpsABCFGKO, cps4A, and cps4I, as well as LPxTG cell wall anchor protein genes, were detected. Antimicrobial susceptibility testing showed resistance to nalidixic acid, oxolinic acid, and oxacillin, while reduced inhibition zones were observed for enrofloxacin and sulfamethoxazole-trimethoprim. Genome analysis identified predicted antimicrobial resistance determinants, including lsaD, vanT, vanY, and mdtA. Resistance-associated protein variants were detected in gyrA and gyrB, suggesting that target alteration may contribute to fluoroquinolone resistance. Overall, this study provides genome-level evidence of virulence and antimicrobial resistance determinants in L. garvieae from Thai tilapia and highlights the importance of whole-genome sequencing for accurate diagnosis, epidemiological surveillance, and disease management in aquaculture.

Animals

Whole-genome sequencing in 333,100 individuals reveals rare non-coding single variant and aggregate associations with height.

The role of rare non-coding variation in complex human phenotypes is still largely unknown. To elucidate the impact of rare variants in regulatory elements, we performed a whole-genome sequencing association analysis for height using 333,100 individuals from three datasets: UK Biobank (N&#x2009;=&#x2009;200,003), TOPMed (N&#x2009;=&#x2009;87,652) and All of Us (N&#x2009;=&#x2009;45,445). We performed rare (&#x2009;<&#x2009;0.1% minor-allele-frequency) single-variant and aggregate testing of non-coding variants in regulatory regions based on proximal-regulatory, intergenic-regulatory and deep-intronic annotation. We observed 29 independent variants associated with height at P&#x2009;<&#x2009;after conditioning on previously reported variants, with effect sizes ranging from -7cm to +4.7&#x2009;cm. We also identified and replicated non-coding aggregate-based associations proximal to HMGA1 containing variants associated with a 5&#x2009;cm taller height and of highly-conserved variants in MIR497HG on chromosome 17. We have developed an approach for identifying non-coding rare variants in regulatory regions with large effects from whole-genome sequencing data associated with complex traits.

Humans

Applicability of Nanopore-only whole-genome sequencing for Pseudomonas aeruginosa outbreak investigation in the ICU setting: a multicentric study.

UNLABELLED: Pseudomonas aeruginosa outbreaks frequently occur in intensive care units (ICUs). In particular, ICU patients requiring mechanical ventilation are vulnerable to P. aeruginosa ventilator-associated pneumonia, which is associated with high morbidity and mortality. Fast and accurate genotyping during the early stage is crucial to document and manage P. aeruginosa outbreaks at the ICU. In this study, we have evaluated the applicability of Oxford Nanopore whole-genome sequencing (WGS) for outbreak investigation and antimicrobial resistance (AMR) prediction. To evaluate whether a Nanopore-only WGS workflow was able to reproduce Illumina-confirmed transmission clusters, 19 P. aeruginosa isolates from ICUs at UZ Brussels (Belgium) that were previously sequenced with Illumina were sequenced using a Nanopore-only workflow based on the latest V14 chemistry, followed by bioinformatic analysis via BugSeq and MBioSEQ Ridom Typer. Although both bioinformatic platforms showed high concordance between Illumina and Nanopore data, MBioSEQ Ridom Typer yielded the lowest allelic distance (maximum one cgMLST allele), confirming all outbreak clusters. When applying the Nanopore-only workflow to longitudinally collected isolates, low genetic heterogeneity (maximum three cgMLST alleles) was observed between isolates from the same patient. WGS and subsequent outbreak analysis of 65 respiratory P. aeruginosa isolates collected from 38 different ICU patients across six Belgian hospitals during a 9-month period showed no intra- or inter-hospital transmission. When the Nanopore-only WGS data were used to predict AMR, there was high categorical agreement (95%) between AMR genotype and phenotype. These findings highlight the potential of Nanopore WGS as a rapid and accurate tool for outbreak investigation of P. aeruginosa. IMPORTANCE: In recent years, Nanopore sequencing has found its way to clinical laboratories because of its affordability, scalability, and, most importantly, its ability to obtain sequencing results in near-real time. However, despite improved raw read accuracies with the latest generation R10.4.1 flow cells, the question remains whether the achieved accuracy is sufficient for accurate bacterial outbreak investigation, particularly in high-risk settings such as intensive care units (ICUs). In this study, we show that Nanopore-only whole-genome sequencing (WGS) is able to match Illumina-only WGS in terms of accuracy for Pseudomonas aeruginosa outbreak investigation in the ICU setting, although important sequence type-dependent and even strain-specific methylation issues need to be resolved in order to guarantee this accuracy. By providing a fast and accurate workflow for reliable P. aeruginosa outbreak investigation, this study could pave the way for large-scale implementation of Nanopore-only WGS, leading to faster outbreak response times.

Humans