Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Metagenomes”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Rapid pan-microbial metagenomics for pathogen detection and personalised therapy in the intensive care unit: a single-centre prospective observational study.

BACKGROUND: Most clinical metagenomic studies do not provide rapid results, detect pathogens from all microbial kingdoms, or measure clinical impacts. We aimed to evaluate the feasibility, performance, and clinical impacts of a rapid pan-microbial respiratory metagenomic service for patients admitted to intensive care units (ICUs). METHODS: This was a single-centre observational study of a rapid metagenomics service that tests respiratory samples from ICU patients at Guy's and St Thomas' hospitals, London, UK, between Dec 5, 2023, and April 12, 2024. Testing used a previously published pan-microbial metagenomics workflow, which simultaneously detects bacteria, fungi, and DNA and RNA viruses; provides same-day preliminary results after 2 h; and provides final results after 24 h. Patients were included if they were aged 18 years or older, admitted to the ICU, had confirmed respiratory failure requiring supplemental oxygen or advanced airway support, and had at least one of the following: (1) clinical suspicion of lower respiratory tract infection based on clinical, biochemical, or radiological findings, (2) sepsis of unknown origin, and (3) concern from an intensive care physician regarding inflammatory pathology. Patients with a suspected or confirmed containment level three organism were excluded. The outcome was performance characteristics of the metagenomic test compared with routine diagnostic testing, detection of additional pathogens by metagenomics, change in antimicrobial prescribing within 24 h of testing, and initiation of immunomodulation. FINDINGS: We processed 114 samples (1-5 per day) from 74 patients (39 [53%] female and 35 [47%] male). 107 (94%) of 114 samples passed quality control, of which 101 (94%) provided same-day preliminary results. Bacteria were detected in 45 (43%) of 104 tested specimens, fungal organisms in 17 (16%) of 104 tested specimens, and viruses in 28 (34%) of 83 tested specimens. Sensitivity in lower respiratory tract samples after 24 h was 97% (95% CI 87-100) for bacteria, 89% (65-99) for fungi, and 89% (71-98) for viruses, with only one false positive for bacteria. Metagenomics identified 42 pathogens not detected by other tests in 32 (30%) of 107 samples. Antimicrobial therapy was changed after metagenomic results from 30 (28%) of 107 samples: 22 (21%) were de-escalated and eight (7%) were escalated. Metagenomics contributed to the initiation of immunomodulation in 15 (20%) of 74 patients for a range of inflammatory conditions. Pathogens with clinical significance to local infection control or national public health were found in ten (14%) of 74 patients, including three invasive Group A streptococci, two parvovirus B19, and one each of HIV-1, measles virus, Mycobacterium tuberculosis, Neisseria meningitidis, and Mycoplasma pneumoniae. INTERPRETATION: Respiratory metagenomics for ICU patients showed good performance and turnaround time, and diverse clinical and public health benefits. This ability to inform both personalised patient therapy and infectious disease surveillance needs evaluation in multicentre studies. FUNDING: None.

Humans↗

Computed tomography-guided precision biopsy combined with metagenomic next-generation sequencing for etiological diagnosis in patients with blood culture-negative systemic infections.

ObjectiveTo evaluate the diagnostic efficacy of computed tomography-guided percutaneous biopsy combined with metagenomic next-generation sequencing in patients with blood culture-negative systemic infections and to assess the clinical impact of using this combined strategy for etiological confirmation and guidance of targeted antimicrobial therapy.MethodsThis single-center retrospective observational cohort study enrolled 78 patients who met the Sepsis-3 consensus criteria for suspected systemic infection and had negative conventional microbiological work-ups (at least two sets of blood cultures) between April 2022 and March 2025. All patients underwent computed tomography-guided biopsy of radiologically identified infectious foci, with specimens processed concurrently for conventional culture and metagenomic next-generation sequencing. Diagnostic performance was benchmarked against the final comprehensive clinical diagnosis, and the influence of metagenomic next-generation sequencing findings on antimicrobial therapy modification was analyzed. Sample size calculation, based on a prior study estimating an metagenomic next-generation sequencing detection rate of 85% (&#x3b1;&#x2009;=&#x2009;0.05, &#x3b2;&#x2009;=&#x2009;0.2), indicated a minimum of 68 cases; accordingly, 78 patients were enrolled.ResultsComputed tomography-guided biopsy was technically successful in all 78 patients (100%). The pathogen detection rate of metagenomic next-generation sequencing (91.0%, 71/78) was significantly higher than that of conventional culture (55.1%, 43/78; p&#x2009;<&#x2009;0.001). Using the final clinical diagnosis as the reference standard, metagenomic next-generation sequencing achieved a sensitivity of 94.7% (95% confidence interval: 86.9-98.5), specificity of 100.0% (95% confidence interval: 29.2-100.0), positive predictive value of 100.0% (95% confidence interval: 94.9-100.0), and negative predictive value of 42.9% (95% confidence interval: 9.9-81.6). Among the 35 culture-negative specimens, metagenomic next-generation sequencing established a definitive microbiological diagnosis in 28 cases (80.0%) and detected polymicrobial infections in 11 cases (14.1% of the cohort). Antimicrobial therapy was rationally adjusted based on metagenomic next-generation sequencing results in 69.2% (54/78) of the patients.ConclusionsThe integration of computed tomography-guided precision biopsy with metagenomic next-generation sequencing offers a highly effective diagnostic approach for blood culture-negative systemic infections. This synergistic strategy improves etiological diagnosis by providing high-yield target specimens that enable comprehensive, unbiased pathogen screening, facilitates differentiation between infectious and non-infectious etiologies, and supplies critical evidence for guiding precision antimicrobial therapy. These findings highlight the growing role of interventional radiology in the contemporary framework of precision infectious disease management.

Humans↗

Fantastic microbes and where to find them: evaluating learning-by-doing outcomes in a crowdfunded metagenomics workshop.

Metagenomics offers a powerful framework for authentic, interdisciplinary learning, yet it remains underrepresented in undergraduate education due to technical and infrastructural barriers. We hypothesized that a research-based, learning-by-doing metagenomics workshop supported by accessible bioinformatics tools could enhance students' perceived skills, self-efficacy, and conceptual understanding of metagenomic analysis. To test this hypothesis, we designed and evaluated a hybrid hands-on workshop in which undergraduate and postgraduate students analyzed real environmental shotgun metagenomic datasets generated from soil samples collected during a citizen science initiative. Using the graphical workflow platform KBase, participants completed an end-to-end metagenomic analysis, from quality control and assembly to genome reconstruction, taxonomic classification, functional annotation, and scientific presentation of results. Educational outcomes were assessed through validated retrospective pre-post questionnaires, self-efficacy scales, and an open-ended conceptual understanding task. Participants showed significant increases in perceived metagenomic skills and confidence in performing metagenomic analyses, while gains in perceived learning showed a positive trend. Conceptual understanding improved across educational levels, particularly among participants with limited prior experience. Together, these findings demonstrate that authentic, data-driven metagenomics activities can effectively lower barriers to computational biology and foster meaningful learning through hands-on research experiences.

Metagenomics↗

Optimizing a culture-enriched hybrid metagenomics pipeline to assess the AMR footprint of livestock manure in anaerobic digestate.

The role of environmental samples from livestock production systems, including manure and anaerobic digestate, as reservoirs of antimicrobial resistance genes (ARGs) is likely underestimated because conventional metagenomic approaches can overlook low-abundance ARGs and often lack the resolution to associate these genes with their microbial hosts and co-localized mobile genetic elements (MGEs). We evaluated whether culture-enriched metagenomics (CEMG), with and without antibiotic selection, enhances ARG detection in anaerobic digestate and improves the resolution of ARG-MGE-host associations using hybrid short- and long-read metagenomic assembly. CEMG increased ARG recovery; mean ARG abundance rose from 15.4 counts per million (CPM) in metagenomic fresh digestate (FD) to 124 CPM in CEMG without antibiotics and 160 CPM in antibiotic-selective CEMG. In FD, only 9 unique ARGs were detected, whereas CEMG recovered 112, including ARGs of clinical importance, such as glycopeptide resistance, beta-lactamase genes, and the cfr 23S rRNA methyltransferase conferring cross-resistance to multiple antibiotic classes. Antibiotic selection induced targeted, class-specific shifts in ARG profiles, with ARGs associated with tetracycline resistance consistently enriched across treatments. Hybrid metagenomic assembly resolved the genomic context of 784 ARGs, of which 59.3% were co-localized with at least one class of MGEs, predominantly plasmids and integrative conjugative elements/integrative mobilizable elements. Biocide and metal resistance genes frequently co-occurred with ARGs on the same contigs. Together, these findings demonstrate that antibiotic-selective culture enrichment enhances resistome surveillance by improving detection of low-abundance ARGs, while hybrid assembly provides critical genomic context for assessing their mobility and host associations.IMPORTANCELivestock manure and its byproducts, such as anaerobic digestate, are recognized as important environmental reservoirs of antimicrobial resistance genes (ARGs) and resistant bacteria, yet current metagenomic approaches may underestimate this risk by failing to detect low-abundance but clinically relevant ARGs. Here, we show that integrating culture enrichment with hybrid metagenomics improves ARG recovery and reveals ARG co-localization with mobile genetic elements and putative bacterial hosts. This approach captures a cultivable and condition-responsive fraction of the resistome that is not readily accessible through direct metagenomic sequencing alone, providing a more informative framework for environmental AMR surveillance.

anaerobic digestion↗

Multilocus sequence typing breathes life into a microbial metagenome.

Shot-gun sequencing of DNA isolated from the environment and the assembly of metagenomes from the resulting data has considerably advanced the study of microbial diversity. However, the subsequent matching of these hypothetical metagenomes to cultivable microorganisms is a limitation of such cultivation-independent methods of population analysis. Using a nucleotide sequence-based genetic typing method, multilocus sequence typing, we were able for the first time to match clonal cultivable isolates to a published and controversial bacterial metagenome, Burkholderia SAR-1, which derived from analysis of the Sargasso Sea. The matching cultivable isolates were all associated with infection and geographically widely distributed; taxonomic analysis demonstrated they were members of Burkholderia cepacia complex Group K. Comparison of the Burkholderia SAR-1 metagenome to closely related B. cepacia complex genomes indicated that it was greater than 98% intact in terms of conserved genes, and it also shared complete sequence identity with the cultivable isolates at random loci beyond the genes sampled by the multilocus sequence typing. Two features of the extant cultivable clones support the argument that the Burkholderia SAR-1 sequence may have been a contaminant in the original metagenomic survey: (i) their growth in conditions reflective of sea water was poor, suggesting the ocean was not their preferred habitat, and (ii) several of the matching isolates were epidemiologically linked to outbreaks of infection that resulted from contaminated medical devices or products, indicating an adaptive fitness of this bacterial strain towards contamination-associated environments. The ability to match identical cultivable strains of bacteria to a hypothetical metagenome is a unique feature of nucleotide sequence-based microbial typing methods; such matching would not have been possible with more traditional methods of genetic typing, such as those based on pattern matching of genomic restriction fragments or amplified DNA fragments. Overall, we have taken the first steps in moving the status of the Burkholderia SAR-1 metagenome from a hypothetical entity towards the basis for life of cultivable strains that may now be analysed in conjunction with the assembled metagenomic sequence data by the wider scientific community.

Bacterial Typing Techniques↗

Comparative metagenomic assessment of Illumina-compatible library preparation methods, short-read lengths, and PacBio HiFi sequencing reveals differences in microbial and functional diversity recovery from a complex environmental sample.

UNLABELLED: Metagenomics enables comprehensive exploration of microbial communities but is influenced by library preparation and sequencing technologies, affecting recovery of microbial genomes and proteins. Here, we benchmarked six Illumina-compatible short-read library preparation conditions in triplicate at 2 &#xd7; 150 bp and 2 &#xd7; 250 bp read lengths alongside PacBio HiFi long-read sequencing using a composite environmental sample of marine mangrove sediment and terrestrial palm tree soil. Longer short reads (2 &#xd7; 250 bp) combined with optimal library preparation approaches improved assembly quality, protein detection, and metagenome-assembled genome (MAG) recovery, achieving results approaching those of long-read sequencing. TruSeq libraries at 2 &#xd7; 250 bp recovered more than sevenfold more unique proteins than the same kit at 2 &#xd7; 150 bp (811,701 vs 110,108) using the same number of sequencing reads, while recovering a comparable number of high-quality MAGs to PacBio HiFi long-read sequencing (11 vs 18) and surpassing it in protein discovery by almost 10-fold (811,701 vs 87,745) at less than half of the sequencing cost. Furthermore, biosynthetic gene cluster analysis identified 46 biosynthetic gene clusters in TruSeq-250PE assemblies compared to 38 in PacBio HiFi, with several showing no close match in the MIBiG database. Although long reads yield more contiguity and complete genomes, longer short reads offer a cost-effective, scalable alternative for uncovering microbial and functional diversity. These findings provide critical guidance for metagenomic experimental design, demonstrating that strategic selection of library preparation chemistry and sequencing parameters can reveal more unknown microbial information in complex biomes without requiring additional sequencing depth. IMPORTANCE: Metagenomic outcomes are strongly influenced by library preparation and sequencing strategies, yet their combined effects in complex environmental samples remain poorly defined. Here, we provide the first direct comparison of Illumina NovaSeq short-read metagenomic sequencing at 2 &#xd7; 150 bp and 2 &#xd7; 250 bp across multiple library preparation kits, alongside PacBio HiFi long-read sequencing. We show that sequencing read length and library preparation critically shape assembly quality, protein recovery, and metagenome-assembled genome (MAG) reconstruction. These findings demonstrate that short-read sequencing at 2 &#xd7; 250 bp, with appropriate library preparation, can match long-read technologies in MAG recovery while substantially surpassing them in protein discovery. With less than half of the sequencing price and a 3.5-fold reduction in cost per gigabase of usable data, this method facilitates more accessible large-scale metagenomic analysis within complex environmental systems.

Metagenomics↗

Benchmarking DNA extraction protocols across use cases for culture-independent Nanopore metagenomics.

Oxford Nanopore Technologies (ONT) sequencing offers several advantages for metagenomics, including long reads, rapid turnaround, low upfront cost, scalability and portability. However, for ONT metagenomics, DNA yield, quality and integrity are important considerations when selecting an extraction method. Many metagenomic extraction methods use harsh lysis conditions to extract a wide range of species and provide an accurate community composition, but these conditions can compromise DNA fragment length. Therefore, extraction methods for ONT metagenomics must balance DNA shearing and recovery with representative community lysis. We systematically evaluated DNA extraction methods for ONT metagenomic sequencing using a use case-oriented framework. Among nearly 50 extraction methods screened, 7 were selected for detailed comparison based on suitability for metagenomics, variation in methodology, availability, cost and processing time: Norgen BioTek Corp's Stool DNA Isolation (NG), Zymo Research's ZymoBIOMICS Quick-DNA HMW MagBead (ZMG), Qiagen's DNeasy Blood and Tissue (QBT), Macherey-Nagel's NucleoMag DNA Microbiome (MN), Zymo Research's ZymoBIOMICS DNA Mini Prep (ZMI), Qiagen's DNeasy PowerSoil/QIAamp PowerFecal Pro (PS) and Qiagen's QIAamp Fast DNA Stool Mini (QIA). Methods were tested using Zymo Research's ZymoBIOMICS Microbial Community Standard (MCS), a matrix-free mock community with known composition. DNA extracts were sequenced on an ONT PromethION using the Rapid Barcoding Kit, except QIA due to insufficient DNA yield. Metrics for the method, DNA extracts, sequencing and genomes were evaluated, revealing trade-offs between methods. The two magnetic bead methods, MN and ZMG, produced the highest mean read length N50 values (13.9 and 16.5&#x2009;kb, respectively) but showed apparent community compositions skewed towards Gram-negative bacteria. In contrast, ZMI and PS maintained a community composition close to expected, with reduced mean read length N50 values (4.5 vs. 7.5&#x2009;kb). Performance across various metrics is presented in the context of the following use cases: maximizing genome coverage and assembly completeness, preserving composition accuracy, targeting specific species and limiting required resources (equipment, time or budget). The metrics and use case considerations presented offer practical guidance for informed selection of DNA extraction methods for ONT metagenomics. For accurate community composition, ZMI or PS are recommended, while PS and ZMG perform best at maximizing genome coverage and assembly completeness. NG and QBT may be the most economical options, though performance trade-offs were observed. Finally, PS may be the preferred method for time-sensitive diagnostic or field applications.

Metagenomics↗

High-resolution metagenome assembly for modern long reads with myloasm.

Long-read metagenome assembly promises complete genomic recovery from microbiomes. However, the complexity of metagenomes poses challenges. We present myloasm, a metagenome assembler for PacBio HiFi and Oxford Nanopore Technologies (ONT) R10.4 long reads. Myloasm uses polymorphic k-mers to construct a high-resolution string graph and then leverages differential abundance for graph simplification. On real-world ONT metagenomes, myloasm assembled three times more complete circular contigs than the next-best assembler. Myloasm can make ONT and HiFi comparable for assembly: for a jointly sequenced gut metagenome, myloasm with ONT assembled more complete circular genomes than any assembler with HiFi. Myloasm recovers previously inaccessible within-species diversity; we recovered six complete Prevotella copri single-contig genomes from a gut metagenome and eight complete TM7 (Saccharibacteria) contigs with > 93% similarity from an oral metagenome. With this improved resolution, we resolved two 98% similar ermF antibiotic resistance genes spreading through distinct strain-specific mobile genetic elements in a human gut.

Journal Article↗

Intracellular screen to identify metagenomic clones that induce or inhibit a quorum-sensing biosensor.

The goal of this study was to design and evaluate a rapid screen to identify metagenomic clones that produce biologically active small molecules. We built metagenomic libraries with DNA from soil on the floodplain of the Tanana River in Alaska. We extracted DNA directly from the soil and cloned it into fosmid and bacterial artificial chromosome vectors, constructing eight metagenomic libraries that contain 53,000 clones with inserts ranging from 1 to 190 kb. To identify clones of interest, we designed a high throughput "intracellular" screen, designated METREX, in which metagenomic DNA is in a host cell containing a biosensor for compounds that induce bacterial quorum sensing. If the metagenomic clone produces a quorum-sensing inducer, the cell produces green fluorescent protein (GFP) and can be identified by fluorescence microscopy or captured by fluorescence-activated cell sorting. Our initial screen identified 11 clones that induce and two that inhibit expression of GFP. The intracellular screen detected quorum-sensing inducers among metagenomic clones that a traditional overlay screen would not. One inducing clone carries a LuxI homologue that directs the synthesis of an N-acyl homoserine lactone quorum-sensing signal molecule. The LuxI homologue has 62% amino acid sequence identity to its closest match in GenBank, AmfI from Pseudomonas fluorescens, and is on a 78-kb insert that contains 67 open reading frames. Another inducing clone carries a gene with homology to homocitrate synthase. Our results demonstrate the power of an intracellular screen to identify functionally active clones and biologically active small molecules in metagenomic libraries.

Alaska↗

ZILA-SRM: a probabilistic framework with zero-inflated latent models for robust strain reconstruction from metagenomes.

UNLABELLED: Resolving bacterial strain diversity from shotgun metagenomic data is fundamental to understanding intra-host evolution, transmission dynamics, and phenotypic heterogeneity. However, current probabilistic approaches face a severe "identifiability limit" when disentangling highly similar genomes. Under high-noise conditions, sequencing errors, coverage overdispersion, and collinearity confound standard expectation-maximization algorithms, resulting in overfitting and spurious "ghost" strains. Here, we introduce zero-inflated latent allocation for strain reconstruction from metagenomes with adaptive sparsity regularization (ZILA-SRM) to overcome this barrier through three innovations. First, we integrate a zero-inflated Poisson mixture model to decouple "structural zeros" (true strain absence) from "sampling zeros" (stochastic dropout), addressing overdispersion in standard Poisson-based tools. Second, we impose a convex adaptive sparsity regularization penalty that leverages biological sparsity priors to shrink noise artifacts dynamically. Third, we implement a graph-theoretic refinement step using maximal clique enumeration to resolve haplotype collinearity. Benchmarking against StrainFinder and MixtureS on 702 synthetic data sets shows that ZILA-SRM achieves a 20% improvement in precision in high-complexity scenarios while maintaining over 80% recall for minor variants at 0.5% abundance. Re-analysis of deep-sequencing data from 195 Mycobacterium tuberculosis clinical samples reveals cryptic low-abundance drug-resistant variants in 12% of patients, including a minor clone carrying the rpoB S450L mutation. Furthermore, application to skin microbiome data sets further reveals a strong negative correlation between dominant Staphylococcus aureus and Staphylococcus epidermidis strains, providing genomic evidence for competitive exclusion. These findings establish ZILA-SRM as a robust tool for resolving strain-level diversity in complex metagenomes. IMPORTANCE: Understanding microbial communities at the strain level is critical because closely related strains can differ dramatically in traits such as drug resistance, virulence, and ecological interactions. However, resolving individual strains from metagenomic sequencing data remains difficult, especially when strains are highly similar or present at low abundance. As a result, biologically meaningful diversity is often obscured or misinterpreted as noise. In this study, we introduce a new framework that improves the reliability of strain reconstruction from complex metagenomic data. By reducing false-positive strain detection while preserving sensitivity to rare variants, our approach enables more accurate characterization of microbial populations. This improved resolution reveals previously hidden subpopulations in clinical and microbiome datasets, providing clearer insights into microbial evolution, competition, and the emergence of clinically relevant traits such as antibiotic resistance.

Metagenomics↗

Development of metagenomic DNA shuffling for the construction of a xenobiotic gene.

We describe a metagenomic DNA shuffling process by combining protein engineering process mutation generator and the high potential diversity of metagenomic DNA derived from the environment. Numerous previous shuffling processes attempted to recombine more or less related parental sequences. At the same time, metagenomic approaches unveiled a huge diversity of DNA sequences and genomes, which have not yet been identified to date. In this study, we attempted to combine these two approaches in order to regenerate a novel gene. Here, we present the possibility that DNA fragments from an entire microbial community (metagenome) might be available for the creation of novel genes capable of degrading pollutants. Metagenomic DNA extracted from non-polluted soil was shuffled in vitro to recreate the linA gene responsible for the first steps of lindane degradation. In this work, 74% of the ORF came from separate subsets of the metagenomic pool from a lindane-free and linA-free soil. Our results demonstrate that microbial community genetic diversity can serve as a source for novel gene construction during in vitro manipulation. This in vitro gene construction might also simulate the mosaic nature of novel genes. This demonstration might lead to other attempts to mimic bacterial adaptation and to construct degradative genes for novel compounds not yet released into the environment.

Bacteria↗

Bacterial diversity of metagenomic and PCR libraries from the Delaware River.

To determine whether metagenomic libraries sample adequately the dominant bacteria in aquatic environments, we examined the phylogenetic make-up of a large insert metagenomic library constructed with bacterial DNA from the Delaware River, a polymerase chain reaction (PCR) library of 16S rRNA genes, and community structure determined by fluorescence in situ hybridization (FISH). The composition of the libraries and community structure determined by FISH differed for the major bacterial groups in the river, which included Actinobacteria, beta-proteobacteria and Cytophaga-like bacteria. Beta-proteobacteria were underrepresented in the metagenomic library compared with the PCR library and FISH, while Cytophaga-like bacteria were more abundant in the metagenomic library than in the PCR library and in the actual community according to FISH. The Delaware River libraries contained bacteria belonging to several widespread freshwater clusters, including clusters of Polynucleobacter necessarius, Rhodoferax sp. Bal47 and LD28 beta-proteobacteria, the ACK-m1 and STA2-30 clusters of Actinobacteria, and the PRD01a001B Cytophaga-like bacteria cluster. Coverage of bacteria with > 97% sequence identity was 65% and 50% for the metagenomic and PCR libraries respectively. Rarefaction analysis of replicate PCR libraries and of a library constructed with re-conditioned amplicons indicated that heteroduplex formation did not substantially impact the composition of the PCR library. This study suggests that although it may miss some bacterial groups, the metagenomic approach can sample other groups (e.g. Cytophaga-like bacteria) that are potentially underrepresented by other culture-independent approaches.

Actinobacteria↗

An application of statistics to comparative metagenomics.

BACKGROUND: Metagenomics, sequence analyses of genomic DNA isolated directly from the environments, can be used to identify organisms and model community dynamics of a particular ecosystem. Metagenomics also has the potential to identify significantly different metabolic potential in different environments. RESULTS: Here we use a statistical method to compare curated subsystems, to predict the physiology, metabolism, and ecology from metagenomes. This approach can be used to identify those subsystems that are significantly different between metagenome sequences. Subsystems that were overrepresented in the Sargasso Sea and Acid Mine Drainage metagenome when compared to non-redundant databases were identified. CONCLUSION: The methodology described herein applies statistics to the comparisons of metabolic potential in metagenomes. This analysis reveals those subsystems that are more, or less, represented in the different environments that are compared. These differences in metabolic potential lead to several testable hypotheses about physiology and metabolism of microbes from these ecosystems.

Algorithms↗

MetaflowX: a scalable and resource-efficient workflow for multi-strategy metagenomic analysis.

Microbiomes play crucial roles in diverse ecosystems, spanning environmental, agricultural, and human health domains. However, in-depth metagenomic data analysis presents significant technical and resource challenges, particularly at scale. Existing computational pipelines are typically limited to either reference-based or reference-free approaches and exhibit inefficiencies in process large datasets. Here, we introduce MetaflowX (https://github.com/01life/MetaflowX), an open-resource workflow integrating both analytical paradigms for enhanced metagenomic investigations. This modular framework encompasses short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality metagenome-assembled genome (MAG) identification, as well as bin refinement and reassembly. Benchmarking tests showed that MetaflowX completed full metagenomic analyses up to 14-fold faster and with 38% less disk usage than existing workflows. It also recovered the highest number of high-quality and taxonomically diverse MAGs. A dedicated reassembly module further improved MAG quality, increasing completeness by 5.6% and reducing contamination by 53% on average. Functional annotation modules enable detection of key features, including virulence and antibiotic resistance genes. Designed for extensibility, MetaflowX provides an efficient solution addressing current and emerging demands in large-scale metagenomic research.

Metagenomics↗

VicMAG, an open-source tool for visualizing circular metagenome-assembled genomes highlighting bacterial virulence and antimicrobial resistance.

Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.

Metagenome↗

De novo assembly and authentication of ancient DNA metagenomes with nf-core/mag.

Ancient DNA provides a direct window into the evolutionary processes that have shaped living microbial species today, as well as their now extinct relatives. Advances in both sequencing methods and de novo assembly techniques have not only resulted in a flood of modern metagenomic sequencing data, but they have also allowed palaeogenomicists to retrieve vast amounts of ancient DNA from past microorganisms, including species and strains without modern reference genomes. However, the degraded nature of ancient DNA means that the standard techniques of genome assembly developed for modern DNA are unlikely to perform effectively, unless heavily modified. This hinders the incorporation of ancient data into broader metagenomic studies that would otherwise benefit from having deep time information on the evolution of different microbial species. In this primer and protocol paper, we provide guidance on ways to adapt existing metagenomic de novo assembly processes, including data input, tools, and settings, in order to perform more robustly and effectively on ancient DNA. After assembly, we then further describe how ancient DNA contigs can be identified and validated. The key steps of ancient metagenomic assembly are now integrated in a dedicated ancient DNA mode in the established pipeline nf-core/mag. By introducing support for ancient DNA data in nf-core/mag, we aim to improve the ability of researchers to more regularly integrate de novo assembled ancient microbial data into broader metagenomics studies of microbial ecology and evolution.

DNA, Ancient↗

MADCAP: isolation of novel nAb-na&#xef;ve AAV capsids from metagenomic data.

UNLABELLED: Gene therapy using adeno-associated virus (AAV) vectors offers promising treatment for genetic disorders, but significant limitations restrict clinical application. Current AAV serotypes exhibit strong liver tropism and require high doses for extra-hepatic targeting, and pre-existing antibodies (NAbs) exclude up to 50% of potential patients. Evolutionarily distant isolates can evade neutralization but typically transduce human tissues poorly and require extensive engineering. We developed MADCAP (Metagenomic AAV Discovery and Capsid Annotation Pipeline) to systematically mine metagenomic data for functional, clinically relevant AAV capsids. We hypothesized that these sources might contain capsids that do not circulate widely in humans, can transduce human cells, and avoid neutralization. We screened 4.2 million metagenomic samples and identified 139 novel AAV capsid isolates which were tested for viral capsid assembly, viability, neutralization evasion, and tissue transduction in non-human primates. While natural serotypes (AAV1, AAV2, AAV9) were neutralized at low dilutions of pooled human immunoglobulin (IVIG), 68% of tested MADCAP capsids exhibited minimal to undetectable neutralization even at supra-physiological IVIG concentrations. Systemically delivered MADCAP capsids effectively transduced multiple clinically relevant tissues in non-human primates. Two capsids, MC46 and MC55, demonstrated improved CNS tropism compared to AAV9 while maintaining comparable production yields. In passive transfer studies, MC46 retained full transduction efficiency in the presence of human antibodies, while AAV9 transduction was completely lost. This work establishes metagenomic mining as a powerful tool for accelerating AAV capsid discovery, identifying isolates with favorable tissue tropisms and resistance to broadly neutralizing antibodies. IMPORTANCE: This work provides proof of concept that potentially clinically relevant AAVs can be isolated from metagenomic data. Our findings lay the groundwork for accelerated discovery of AAV capsids which could potentially increase the accessibility and effectiveness of AAV gene therapy.

AAV↗

Metagenomic insights and biosynthetic potential of Candidatus Entotheonella symbiont associated with Halichondria marine sponges.

Korea, being surrounded by the sea, provides a rich habitat for marine sponges, which have been a prolific source of bioactive natural products. Although a diverse array of structurally novel natural products has been isolated from Korean marine sponges, their biosynthetic origins remain largely unknown. To explore the biosynthetic potential of Korean marine sponges, we conducted metagenomic analyses of sponges inhabiting the East Sea of Korea. This analysis revealed a symbiotic association of Candidatus Entotheonella bacteria with Halichondria sponges. Here, we report a new chemically rich Entotheonella variant, which we named Ca. Entotheonella halido. Remarkably, this symbiont makes up 69% of the microbial community in the sponge Halichondira dokdoensis. Genome-resolved metagenomics enabled us to obtain a high-quality Ca. E. halido genome, which represents the largest (12 Mb) and highest quality among previously reported Entotheonella genomes. We also identified the biosynthetic gene cluster (BGC) of the known sponge-derived Halicylindramides from the Ca. E. halido genome, enabling us to determine their biosynthetic origin. This new symbiotic association expands the host diversity and biosynthetic potential of metabolically talented bacterial genus Ca. Entotheonella symbionts.IMPORTANCEOur study reports the discovery of a new bacterial symbiont Ca. Entotheonella halido associated with the Korean marine sponge Halichondria dokdoensis. Using genome-resolved metagenomics, we recovered a high-quality Ca. E. halido MAG (Metagenome-Assembled Genome), which represents the largest and most complete Ca. Entotheonella MAG reported to date. Pangenome and BGC network analyses revealed a remarkably high BGC diversity within the Ca. Entotheonella pangenome, with almost no overlapping BGCs between different MAGs. The cryptic and genetically unique BGCs present in the Ca. Entotheonella pangenome represents a promising source of new bioactive natural products.

Animals↗