Search PubMedSearch

SEARCH · Search PubMed

Results for “Illumina”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Performance of the IR Biotyper, Nanopore, and Illumina sequencing to discriminate Escherichia coli strains originating from poultry.

UNLABELLED: Escherichia coli is a highly diverse bacterial species that includes avian pathogenic E. coli (APEC), one of the most prevalent causative agents of disease in poultry worldwide. Rapid and accurate discrimination of E. coli strains is essential for outbreak management, antimicrobial resistance surveillance, and vaccine development. In this study, we compared the performance of Fourier Transform Infrared (FTIR) spectroscopy using the IR Biotyper system with Nanopore and Illumina whole-genome sequencing (WGS) for typing 200 E. coli isolates, originating from four poultry rearing farms in the Netherlands. From each farm, we sampled 10 one-day-old meat type rearing chicks, and from every chick, we isolated 5 E. coli strains. FTIR clustering showed strong concordance with WGS-based classifications, particularly serotyping and core-genome similarity determined by PopPUNK analysis (Adjusted Rand Index 0.75-0.92). While Nanopore and Illumina sequencing provided the highest genetic resolution, FTIR offered a faster (max 6 vs 12-28 days for 200 isolates) and more cost-effective alternative for assessing clonality. Across all methods, multiple strains were detected per farm, whereas most birds carried a single dominant E. coli strain. Our findings demonstrate that FTIR provides a reliable and scalable phenotypic method for rapid strain discrimination in E. coli, complementing WGS in diagnostic, surveillance, and epidemiological settings where speed and throughput are critical. IMPORTANCE: Escherichia coli is a major pathogen in poultry and a potential zoonotic risk for humans. Rapid and accurate discrimination of avian pathogenic E. coli (APEC) strains is critical for outbreak management, antimicrobial resistance surveillance, and the design of effective autogenous vaccines. In this study, we compared Fourier Transform Infrared (FTIR) spectroscopy with Nanopore and Illumina whole-genome sequencing for strain typing of E. coli isolates originating from poultry. The results show that FTIR provides comparable clustering accuracy to genomic approaches at a fraction of the time and costs. This work demonstrates that FTIR can serve as a practical, high-throughput alternative for routine monitoring of E. coli in veterinary diagnostics and food safety of poultry meat, enabling faster decision-making and more targeted interventions across the poultry production chain.

Animals

Comparative metagenomic assessment of Illumina-compatible library preparation methods, short-read lengths, and PacBio HiFi sequencing reveals differences in microbial and functional diversity recovery from a complex environmental sample.

UNLABELLED: Metagenomics enables comprehensive exploration of microbial communities but is influenced by library preparation and sequencing technologies, affecting recovery of microbial genomes and proteins. Here, we benchmarked six Illumina-compatible short-read library preparation conditions in triplicate at 2 × 150 bp and 2 × 250 bp read lengths alongside PacBio HiFi long-read sequencing using a composite environmental sample of marine mangrove sediment and terrestrial palm tree soil. Longer short reads (2 × 250 bp) combined with optimal library preparation approaches improved assembly quality, protein detection, and metagenome-assembled genome (MAG) recovery, achieving results approaching those of long-read sequencing. TruSeq libraries at 2 × 250 bp recovered more than sevenfold more unique proteins than the same kit at 2 × 150 bp (811,701 vs 110,108) using the same number of sequencing reads, while recovering a comparable number of high-quality MAGs to PacBio HiFi long-read sequencing (11 vs 18) and surpassing it in protein discovery by almost 10-fold (811,701 vs 87,745) at less than half of the sequencing cost. Furthermore, biosynthetic gene cluster analysis identified 46 biosynthetic gene clusters in TruSeq-250PE assemblies compared to 38 in PacBio HiFi, with several showing no close match in the MIBiG database. Although long reads yield more contiguity and complete genomes, longer short reads offer a cost-effective, scalable alternative for uncovering microbial and functional diversity. These findings provide critical guidance for metagenomic experimental design, demonstrating that strategic selection of library preparation chemistry and sequencing parameters can reveal more unknown microbial information in complex biomes without requiring additional sequencing depth. IMPORTANCE: Metagenomic outcomes are strongly influenced by library preparation and sequencing strategies, yet their combined effects in complex environmental samples remain poorly defined. Here, we provide the first direct comparison of Illumina NovaSeq short-read metagenomic sequencing at 2 × 150 bp and 2 × 250 bp across multiple library preparation kits, alongside PacBio HiFi long-read sequencing. We show that sequencing read length and library preparation critically shape assembly quality, protein recovery, and metagenome-assembled genome (MAG) reconstruction. These findings demonstrate that short-read sequencing at 2 × 250 bp, with appropriate library preparation, can match long-read technologies in MAG recovery while substantially surpassing them in protein discovery. With less than half of the sequencing price and a 3.5-fold reduction in cost per gigabase of usable data, this method facilitates more accessible large-scale metagenomic analysis within complex environmental systems.

Metagenomics

HTSinfer: inferring metadata from bulk Illumina RNA-Seq libraries.

SUMMARY: The Sequencing Read Archive is one of the largest and fastest-growing repositories of sequencing data, containing tens of petabytes of sequenced reads. Its data is used by a wide scientific community, often beyond the primary study that generated them. Such analyses rely on accurate metadata concerning the type of experiment and library, as well as the organism from which the sequenced reads were derived. These metadata are typically entered manually by contributors in an error-prone process, and are frequently incomplete. In addition, easy-to-use computational tools that verify the consistency and completeness of metadata describing the libraries to facilitate data reuse, are largely unavailable. Here, we introduce HTSinfer, a Python-based tool to infer metadata directly and solely from bulk RNA-sequencing data generated on Illumina platforms. HTSinfer leverages genome sequence information and diagnostic genes to rapidly and accurately infer the library source and library type, as well as the relative read orientation, 3' adapter sequence and read length statistics. HTSinfer is written in a modular manner, published under a permissible free and open-source license and encourages contributions by the community, enabling easy addition of new functionalities, e.g. for the inference of additional metrics, or the support of different experiment types or sequencing platforms. AVAILABILITY AND IMPLEMENTATION: HTSinfer is released under the Apache License 2.0. Latest code is available via GitHub at https://github.com/zavolanlab/htsinfer, while releases are published on Bioconda. A snapshot of the HTSinfer version described in this article was deposited at Zenodo at 10.5281/zenodo.13985958.

Metadata

Reflective Evaluation of Next-Generation Sequencing Data during Early Phase Detection of the Delta Variant.

During the SARS-CoV-2 pandemic, next-generation sequencing (NGS) technologies like the Ion Torrent S5 and Illumina MiSeq, alongside advanced software, improved genomic surveillance in South Africa. This study analysed anonymized samples from the Eastern Cape using Genome Detective and NextClade, showing Ion Torrent S5 and Illumina MiSeq success rates of 96% and 94%, respectively. The study focused on genomic coverage (above 80%) and mutation detection (below 100), with the Ion Torrent S5 achieving 99% coverage compared to Illumina MiSeq's 80%, likely due to different primers used in amplification. The Ion Torrent S5 was more effective in sequencing varied viral loads, whereas Illumina MiSeq had difficulties with lower loads. Both platforms were adept at identifying clades, successfully differentiating between Beta (<45%) and Delta variants (<30%), despite minor discrepancies in assignments due to Illumina MiSeq's lower coverage, leading to a failure rate of up to 6%. Manual library preparation showed similar sample processing and clade identification capabilities for both platforms. However, differences in sequencing duration (3.5 vs. 36 hours), automation level, genomic coverage (80% vs. 99%), and viral load compatibility were noted, highlighting each platform's unique advantages and challenges in SARS-CoV-2 genomic surveillance. In conclusion, the Illumina MiSeq and Ion Torrent S5 platforms are both efficacious in executing whole-genome sequencing (WGS) via amplicons, facilitating precise, accurate, and high-throughput examinations of SARS-CoV-2 viral genomes. However, it is important to note the existence of disparities in the quality of data produced by each platform. Each system offers unique benefits and limitations, rendering them viable choices for the genomic surveillance of SARS-CoV-2.

Illumina MiSeq

Towards genomic medicine: a tailored next-generation sequencing panel for hydroxyurea pharmacogenomics in Tanzania.

BACKGROUND: Pharmacogenomics of hydroxyurea is an important aspect in the management of sickle cell disease (SCD), especially in the era of genomic medicine. Genetic variations in loci associated with HbF induction and drug metabolism are prime targets for hydroxyurea (HU) pharmacogenomics,&#xa0;as these can significantly impact the therapeutic efficacy and safety of HU in SCD patients. METHODS: This study involved designing of a custom panel targeting BCL11A, ARG2, HBB, HBG1, WAC, HBG2, HAO2, MYB, SAR1A, KLF10, CYP2C9, CYP2E1 and NOS1 as potential HU pharmacogenomics targets.&#xa0;These genes were selected based on their known roles in HbF induction and HU metabolism. The panel was designed using the Illumina Design Studio (Illumina, San Diego, CA, USA) and achieved a total coverage of 96% of all genomic targets over a span of 51.6 kilobases (kb). This custom panel was then sequenced using the Illumina MiSeq platform to ensure high coverage and accuracy. RESULTS: We are reporting a successfully designed Illumina (MiSeq) HU pharmacogenomics custom panel encompassing 51.6 kilobases. The designed panel achieved greater than 1000x amplicon coverage which is sufficient for genomic analysis. CONCLUSIONS: This study provides a valuable tool for research in HU pharmacogenomics, especially in Africa where SCD is highly prevalent, and personalized medicine approaches are crucial for improving patient outcomes.&#xa0;The custom-designed Illumina (MiSeq) panel, with its extensive coverage and high sequencing depth, provides a robust platform for studying genetic variations associated with HU response. This panel can contribute to the development of tailored therapeutic strategies, ultimately enhancing the management of SCD through more effective and safer use of hydroxyurea.

Hydroxyurea

Applicability of Nanopore-only whole-genome sequencing for Pseudomonas aeruginosa outbreak investigation in the ICU setting: a multicentric study.

UNLABELLED: Pseudomonas aeruginosa outbreaks frequently occur in intensive care units (ICUs). In particular, ICU patients requiring mechanical ventilation are vulnerable to P. aeruginosa ventilator-associated pneumonia, which is associated with high morbidity and mortality. Fast and accurate genotyping during the early stage is crucial to document and manage P. aeruginosa outbreaks at the ICU. In this study, we have evaluated the applicability of Oxford Nanopore whole-genome sequencing (WGS) for outbreak investigation and antimicrobial resistance (AMR) prediction. To evaluate whether a Nanopore-only WGS workflow was able to reproduce Illumina-confirmed transmission clusters, 19 P. aeruginosa isolates from ICUs at UZ Brussels (Belgium) that were previously sequenced with Illumina were sequenced using a Nanopore-only workflow based on the latest V14 chemistry, followed by bioinformatic analysis via BugSeq and MBioSEQ Ridom Typer. Although both bioinformatic platforms showed high concordance between Illumina and Nanopore data, MBioSEQ Ridom Typer yielded the lowest allelic distance (maximum one cgMLST allele), confirming all outbreak clusters. When applying the Nanopore-only workflow to longitudinally collected isolates, low genetic heterogeneity (maximum three cgMLST alleles) was observed between isolates from the same patient. WGS and subsequent outbreak analysis of 65 respiratory P. aeruginosa isolates collected from 38 different ICU patients across six Belgian hospitals during a 9-month period showed no intra- or inter-hospital transmission. When the Nanopore-only WGS data were used to predict AMR, there was high categorical agreement (95%) between AMR genotype and phenotype. These findings highlight the potential of Nanopore WGS as a rapid and accurate tool for outbreak investigation of P. aeruginosa. IMPORTANCE: In recent years, Nanopore sequencing has found its way to clinical laboratories because of its affordability, scalability, and, most importantly, its ability to obtain sequencing results in near-real time. However, despite improved raw read accuracies with the latest generation R10.4.1 flow cells, the question remains whether the achieved accuracy is sufficient for accurate bacterial outbreak investigation, particularly in high-risk settings such as intensive care units (ICUs). In this study, we show that Nanopore-only whole-genome sequencing (WGS) is able to match Illumina-only WGS in terms of accuracy for Pseudomonas aeruginosa outbreak investigation in the ICU setting, although important sequence type-dependent and even strain-specific methylation issues need to be resolved in order to guarantee this accuracy. By providing a fast and accurate workflow for reliable P. aeruginosa outbreak investigation, this study could pave the way for large-scale implementation of Nanopore-only WGS, leading to faster outbreak response times.

Humans

Long-read sequencing reveals putatively mobilizable resistance genes and multi-drug resistance plasmids underestimated by short-read metagenomics.

While shotgun metagenomics is often used to profile antibiotic resistome in gut microbial communities, few studies have investigated if the choice of sequencing platform and assembly strategy affect what mobile genetic elements and antimicrobial resistance genes are recovered. In this study, we compared three platforms (Illumina, Oxford Nanopore, and PacBio HiFi) and seven assembly strategies on gut metagenomes from cattle, pig, and human as case studies. Long-read assemblies recovered 5- to 7-fold more plasmid sequence than Illumina in cattle and pig (mean 17.0 Mb vs. 3.1 Mb), while Illumina performed comparably in the less diverse human gut where high per-species coverage enabled effective short-read plasmid assembly. Long reads also detected more resistance genes on plasmid contigs. Hybrid assembly results depended on the algorithm: scaffolding-based OPERA-MS preserved long-read contiguity and recovered more plasmid-borne resistance genes, while the short-read-centric metaSPAdes hybrid mode produced fragmented assemblies. After collapsing haplotype redundancy, PacBio HiFi identified 2 and 49 unique multi-drug resistance plasmid lineages in cattle and pig, respectively. On the other hand, only 2 and 4 were identified from Illumina. Long reads also placed far more ARGs in a putative mobilization context (50-73%) compared to 14-21% for short reads. Platform and assembly strategy are thus key variables in mobilome and resistome characterization and should be accounted for in antimicrobial resistance surveillance.

Animals

Next-Generation Sequencing Methods for Sensitive Hepatitis B Viral Genome Analysis: A European Study.

This multicentre study investigated the utility of next-generation sequencing (NGS) to detect and generate hepatitis B virus (HBV) genomes in samples of low viral load (from 0.2 to 6207 IU/mL). 23 HBV DNA-positive plasma samples of genotypes A-E and one HBV-negative control sample were assayed blindly via 9 established NGS methods from 6 European laboratories. Methods included untargeted metagenomics, pre-enrichment by probe-capture followed by Illumina sequencing, and HBV-specific PCR pre-amplification followed by sequencing with Nanopore or Illumina. Full HBV genomes were obtained only from samples with viral loads >&#x2009;1000 IU/mL using probe-capture methods, >&#x2009;200 IU/mL using PCR-Illumina methods, >&#x2009;10 IU/mL using PCR-Nanopore methods, and in no samples using metagenomic methods. Contamination was observed in the negative control and samples with very low viral loads in PCR-based methods. Probe-capture and metagenomic methods detected additional viruses not routinely screened in blood donations, including polyomaviruses and herpesviruses; positive results were confirmed by PCR. In conclusion, NGS may delineate whole-genome sequences at low viral loads if supported by a PCR pre-amplification step. Probe-capture methods also reliably detect HBV without pre-amplification but show limited genome coverage for samples with low viral loads; they may additionally detect a wide range of blood-borne viruses.

Humans

Assembling genomes of non-model plants: A case study with evolutionary insights from Ranunculus (Ranunculaceae).

Whereas genome sequencing and assembly technologies are improving, cost can still be prohibitive for plant species with large, complex genomes. As a consequence, genomics work on some taxa in evolutionarily pivotal positions in the vascular plant tree of life has been hampered. The species-rich genus Ranunculus (Ranunculaceae) is an important angiosperm group for the study of polyploidy, apomixis, and reticulate evolution. However, neither mitochondrial nor high-quality nuclear genome sequences are available. This limits phylogenomic, functional, and taxonomic analyses thus far. Here, we tested Illumina short-read, Oxford Nanopore Technology (ONT) and PacBio (HiFi) long-read, and hybrid-read assembly strategies. We sequenced the diploid progenitor species R. cassubicifolius (R.&#x2009;auricomus species complex) and selected the best assemblies in terms of completeness, contiguity, and quality scores. We first assembled the plastome (156&#x2009;kbp, 85 genes) and mitogenome (1.18&#x2009;Mbp, 40 genes) sequences using Illumina and Illumina-PacBio-hybrid strategies, respectively. We also present an updated plastome and the first mitogenome phylogeny of Ranunculaceae, including studies of gene loss (e.g., infA, ycf15, or rps) with evolutionary implications. For the nuclear genome sequence, we favored a PacBio-based assembly polished three times with filtered short reads and subsequently scaffolded into eight pseudochromosomes by chromatin conformation data (Hi-C). We obtained a haploid genome sequence of 2.69&#x2009;Gbp, with 94.1% complete BUSCO genes found and 35&#x2009;482 annotated genes, and inferred ancient gene duplications compared to existing Ranunculales genomes. The genomic information presented here will enable advanced evolutionary-functional analyses for the species complex, but also for the genus and beyond Ranunculaceae.

Ranunculus

Evaluation of cross-platform compatibility of a DNA methylation-based glucocorticoid response biomarker.

BACKGROUND: Identifying blood-based DNA methylation patterns is a minimally invasive way to detect biomarkers in predicting age, characteristics of certain diseases and conditions, as well as responses to immunotherapies. As microarray platforms continue to evolve and increase the scope of CpGs measured, new discoveries based on the most recent platform version and how they compare to available data from the previous versions of the platform are unknown. The neutrophil dexamethasone methylation index (NDMI 850) is a blood-based DNA methylation biomarker built on the Illumina MethylationEPIC (850K) array that measures epigenetic responses to dexamethasone (DEX), a synthetic glucocorticoid often administered for inflammation. Here, we compare the NDMI 850 to one we built using data from the Illumina Methylation 450K (NDMI 450). RESULTS: The NDMI 450 consisted of 22 loci, 15 of which were present on the NDMI 850. In adult whole blood samples, the linear composite scores from NDMI 450 and NDMI 850 were highly correlated and had equivalent predictive accuracy for detecting DEX exposure among adult glioma patients and non-glioma adult controls. However, the NDMI 450 scores of newborn cord blood were significantly lower than NDMI 850 in samples measured with both assays. CONCLUSIONS: We developed an algorithm that reproduces the DNA methylation glucocorticoid response score using 450K data, increasing the accessibility for researchers to assess this biomarker in archived or publicly available datasets that use the 450K version of the Illumina BeadChip array. However, the NDMI850 and NDMI450 do not give similar results in cord blood, and due to data availability limitations, results from sample types of newborn cord blood should be interpreted with care.

Adult

Reference-Free Microsatellite Instability Detection from Tumor Sequencing Using Intrasample Variability Modeling.

Microsatellite instability (MSI) is a predictive biomarker in several tumor types. However, many next-generation sequencing-based callers require matched normal samples, reference panels, or pretrained models, limiting their portability across assays and sequencing centers. We developed PROMIS (PROfiling of Microsatellite InStability), a tumor-only, reference-free pipeline that uses a discrete mixture model to characterize intrasample repeat-length distributions at predefined microsatellite loci. Locus-level classifications are then aggregated into a continuous MSI score. We benchmarked PROMIS in colorectal (CRC), endometrial (UCEC), and gastric (STAD) cancers from The Cancer Genome Atlas. PROMIS achieved an overall area under the receiver operating characteristic curve (AUC) of 0.995 and cohort-specific AUCs of 1.00 in CRC and stomach adenocarcinoma and 0.999 in uterine corpus endometrial carcinoma, comparable to established tools despite not using matched normals or pretrained models. Subsampling demonstrated robust performance with substantially fewer loci. In silico dilution showed progressively reduced MSI-microsatellite-stable discrimination, with the pooled AUC declining from 0.83 at 10% tumor fraction to 0.53 at 1%. At low tumor fractions, tumor-type-specific baseline microsatellite variability increasingly influenced PROMIS scores. Finally, in prostate and CRC cell-free DNA cohorts, including Illumina TSO500 data and an 18-gene panel, PROMIS yielded MSI scores concordant with orthogonal tissue- and panel-based classifications across the evaluated Illumina-based sequencing contexts. Accordingly, the present validation should be considered limited to Illumina-based sequencing platforms. PROMIS is intended to complement existing genomic profiling workflows by enabling MSI assessment from sequencing data already generated for broader molecular analyses. Prospective clinical validation remains necessary before clinical implementation.

Journal Article

Evaluating 12 automated, whole-genome sequencing analysis pipelines for Mycobacterium tuberculosis complex: a comparative study.

BACKGROUND: Reliance on complex, custom-built bioinformatics pipelines is a barrier to the implementation of whole-genome sequencing (WGS) of Mycobacterium tuberculosis in high-burden settings in some low-income and middle-income countries (LMICs). Automated analysis pipelines could address this inequity in access to WGS-based diagnostics and surveillance. This study aimed to systematically evaluate the performance and usability of publicly available WGS pipelines for M tuberculosis. METHODS: We identified automated M tuberculosis WGS analysis pipelines through searches of PubMed and GitHub from database inception up to Aug 31, 2024. Accuracy, cost, accessibility, and scalability were assessed for each pipeline. We evaluated the accuracy of genotypic drug susceptibility testing (gDST) using publicly available sequences with phenotypic susceptibility data for 12 antituberculosis drugs. We estimated pooled sensitivity and specificity for each pipeline, across all drugs, by conducting a bivariate meta-analysis, with random effects representing between-drug variability. Lineage classifications were compared, and a previously epidemiologically well-characterised dataset was used to compare measures of genomic relatedness. FINDINGS: Among 28 candidate pipelines, 16 were excluded as they were unmaintained and inexecutable. 12 pipelines (11 compatible with Illumina and four compatible with Nanopore), all free to use, were included for evaluation. Six pipelines processed and stored data remotely, but for five of these six, scalability was limited by the need to upload sequences through web portals. For local processing pipelines, scalability was dependent on substantial local computational resources, data storage capacity, and command-line interfaces that limited user-friendliness. Only one of six remote-processing pipelines removed human DNA sequences before server upload. gDST was similarly accurate across ten of 11 Illumina-compatible pipelines and three of four Nanopore-compatible pipelines. All pipelines classified the main lineages consistently, although there were differences at sublineage resolution. Outputs from three of four pipelines reporting genomic relatedness were compatible with commonly cited single nucleotide polymorphism difference thresholds. INTERPRETATION: Numerous automated analysis pipelines capable of enhancing equity in M tuberculosis WGS are available. Given the overall similarities between the pipelines evaluated in this study in terms of gDST performance, lineage classification, and genomic relatedness inference, non-functional attributes such as availability, accessibility, scalability, and privacy could represent the point of difference for prospective users in LMICs with a high burden of tuberculosis. FUNDING: The Rhodes Trust, Wellcome, Ellison Institute of Technology, and the UK National Institute for Health and Care Research Oxford Biomedical Research Centre.

Mycobacterium tuberculosis

Genomic signatures of cold adaptation in a Himalayan drosophilid.

Drosophila nepalensis is a cold-adapted drosophilid endemic to the Himalayan region. Its ability to survive in harsh, cold conditions makes it a valuable Drosophila model for investigating how adaptation to thermal extremes may influence species persistence under future climate change. Here, we report the first de novo genome assembly of D. nepalensis, based on a hybrid sequencing strategy that combines Illumina short reads and Oxford Nanopore long reads. Illumina sequencing generated 49.88 million 150&#x2005;bp paired-end reads (&#x223c;14.96&#x2005;Gbp), while Nanopore sequencing produced 1.35 million long reads totaling &#x223c;0.76&#x2005;Gbp. The assembled genome spanned &#x223c;178&#x2005;Mb with an N50 of 83.6&#x2005;kb and 98% BUSCO completeness, comparable to other well-annotated Drosophila genomes. Annotation identified 10,560 protein-coding genes, including transcription factor-rich and stress-related domains such as zinc fingers, WD40 repeats, and ankyrin motifs. Comparative orthology analysis across 6 Drosophila species identified 14,168 orthologous clusters, of which 9,173 were shared among all 6 species, indicating a conserved core genomic set across the sampled taxa. D. nepalensis showed 83 unique orthogroups and 50 singletons, suggesting some lineage-specific gene expansions associated with cold adaptation and endemicity, including families encoding caspase-family apoptotic regulators, chromatin remodeling proteins (HMGB/protamine-like), and SNARE-domain vesicle trafficking factors. Gene family evolution analysis revealed the highest expansions in the cold-tolerant Himalayan drosophilid, D. nepalensis, including significant expansions in serine protease, chaperone, and neurotransmitter transporter families, alongside dramatic contractions of core histone gene families, suggesting lineage-specific chromatin remodeling and ecological specialization.

Drosophila nepalensis

AVITI sequencing of a four-generation CEPH/Utah pedigree confirms low mutation rates at homopolymer loci despite their low sequence complexity.

BACKGROUND: Short tandem repeats (STRs) and homopolymers are among the most mutable loci in the human genome. Despite their presumed mutability owing to replication slippage, homopolymer loci exhibit lower mutation rates and minimal paternal age effects compared to other STRs. This paradox questions if technical limitations, rather than biological mechanisms, explain these observations. RESULTS: We used the Element Biosciences AVITI platform to sequence the genomes of a 48-member, four-generation CEPH/Utah pedigree. As the AVITI platform reduces error rates at repetitive sequences compared to Illumina, this design enabled accurate mutation discovery at 90% of assayed homopolymers and a 1.7-fold increase in discoverable mutations compared to Illumina. We identified a median of 35 de novo homopolymer mutations per trio and a mutation rate of 5.28 &#xd7; 10-5 DNMs per locus per generation, confirming a lower rate than dinucleotides (1.94 &#xd7; 10-4). Most DNMs were single base-pair expansions or contractions. Despite comprising <1% of homopolymer loci, G/C homopolymers showed 18-fold higher mutation rates than A/T homopolymers; in contrast, the high dinucleotide mutation rate is not driven by a particular motif class. Parent-of-origin analysis revealed 78% of homopolymer mutations are paternal in origin, but no significant paternal age effect was observed. CONCLUSIONS: This study confirms that homopolymers exhibit lower mutation rates and lack strong paternal age effects compared to other STRs, likely owing to the combination of a lower propensity to form slippage-causing secondary structures and more efficient mismatch repair. Our set of high-quality mutations suggest these phenomena are biological rather than technical in nature. Finally, we demonstrate that AVITI sequencing unlocks previously intractable regions of the genome and will be a powerful tool for continued investigation of repeat mutation.

AVITI

IBDV-SSA, a novel molecular approach for the recovery of infectious bursal disease virus whole genomes from FTA cards.

Infectious bursal disease (IBD), a highly contagious viral disease in young chickens, poses significant economic losses due to high mortality and immunosuppression. While IBD virus (IBDV) virulence is influenced by multiple genes, whole-genome sequencing (WGS) of IBDV is crucial for defining the strain pathotype and clinical profile. Flinders Technology Associates (FTA) cards are convenient for field sample collection, but their filter paper matrix can hinder nucleic acid recovery, impacting sequencing efficiency. This study evaluated two enrichment strategies, single primer amplification (SPA) and IBDV segment-specific amplification (SSA), coupled with short-read (Illumina) and long-read (Oxford Nanopore Technologies, ONT) sequencing platforms, to optimize IBDV whole-genome recovery from FTA cards. Illumina sequencing produced comparable raw read counts for both methods, yet IBDV-SSA samples achieved significantly higher genome mapping rates (76%) than IBDV-SPA (12%). Genome coverage analysis revealed that IBDV-SSA provided uniform read distribution across both genomic segments, ensuring complete coverage, while IBDV-SPA exhibited significant bias, with most reads mapping to segment B, and limited coverage of segment A. Importantly, IBDV-SSA also proved compatible with ONT long-read sequencing, providing complete genome coverage. Notably, IBDV-SSA coupled with short-read sequencing successfully characterized coinfections in two samples. This optimized approach using IBDV-SSA enables efficient and comprehensive WGS of IBDV from FTA cards, facilitating strain characterization, virulence prediction, and epidemiological investigations.IMPORTANCEThis research tackles a significant problem for poultry farmers: a virus called infectious bursal disease virus (IBDV) that harms young chickens, causing high death rates and economic losses. To fight it effectively, scientists need to analyze its complete genetic makeup. Traditionally, collecting and preserving IBDV field samples was challenging. Flinders Technology Associates (FTA) cards have simplified this process, but getting usable genetic material from them has been difficult. This study introduces a new genome enrichment method, IBDV segment-specific amplification (IBDV-SSA), which successfully allows for IBDV complete genome recovery from FTA cards. By using this improved approach, scientists can accurately identify virus strains, assess how harmful they are, and monitor their spread. This, in turn, helps to improve vaccines and protect flocks. IBDV-SSA is a powerful tool for outbreak surveillance, supporting the poultry industry and ensuring a stable food supply.

Infectious bursal disease virus

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score &#x2265; 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S

Analysis of targeted and whole genome sequencing of PacBio HiFi reads for a comprehensive genotyping of gene-proximal and phenotype-associated Variable Number Tandem Repeats.

Variable Number Tandem repeats (VNTRs) refer to repeating motifs of size greater than five bp. VNTRs are an important source of genetic variation, and have been associated with multiple Mendelian and complex phenotypes. However, the highly repetitive structures require reads to span the region for accurate genotyping. Pacific Biosciences HiFi sequencing spans large regions and is highly accurate but relatively expensive. Therefore, targeted sequencing approaches coupled with long-read sequencing have been proposed to improve efficiency and throughput. In this paper, we systematically explored the trade-off between targeted and whole genome HiFi sequencing for genotyping VNTRs. We curated a set of 10&#xa0;,&#xa0;787 gene-proximal (G-)VNTRs, and 48 phenotype-associated (P-)VNTRs of interest. Illumina reads only spanned 46% of the G-VNTRs and 71% of P-VNTRs, motivating the use of HiFi sequencing. We performed targeted sequencing with hybridization by designing custom probes for 9,999 VNTRs and sequenced 8 samples using HiFi and Illumina sequencing, followed by adVNTR genotyping. We compared these results against HiFi whole genome sequencing (WGS) data from 28 samples in the Human Pangenome Reference Consortium (HPRC). With the targeted approach only 4,091 (41%) G-VNTRs and only 4 (8%) of P-VNTRs were spanned with at least 15 reads. A smaller subset of 3,579 (36%) G-VNTRs had higher median coverage of at least 63 spanning reads. The spanning behavior was consistent across all 8 samples. Among 5,638 VNTRs with low-coverage (&#xa0;<&#xa0;15), 67% were located within GC-rich regions (&#xa0;>&#xa0;60%). In contrast, the 40X WGS HiFi dataset spanned 98% of all VNTRs and 49 (98%) of P-VNTRs with at least 15 spanning reads, albeit with lower coverage. Spanning reads were sufficient for accurate genotyping in both cases. Our findings demonstrate that targeted sequencing provides consistently high coverage for a small subset of low-GC VNTRs, but WGS is more effective for broad and sufficient sampling of a large number of VNTRs.

Minisatellite Repeats