Search PubMedSearch

SEARCH · Search PubMed

Results for “Genome, Microbial”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

CompareM2 is a genomes-to-report pipeline for comparing microbial genomes.

SUMMARY: Here, we present CompareM2, a genomes-to-report pipeline for comparative analysis of bacterial and archaeal genomes derived from isolates and metagenomic assemblies. CompareM2 is easy to install and operate, designed in such a way that the user can install the complete software in one step and launch all analyses on a set of microbial genomes (bacterial and archaeal) in a single action. The central results generated via the CompareM2 workflow are emphasized in a portable dynamic report document. AVAILABILITY AND IMPLEMENTATION: CompareM2 is a free software that is scalable to a range of project sizes, and welcomes modifications and pull requests from the community on its Git repository at https://github.com/cmkobel/comparem2.

Software

Seqwin: ultrafast identification of signature sequences in microbial genomes.

MOTIVATION: Polymerase chain reaction (PCR) enables rapid, cost-effective diagnostics but requires prior identification of genomic regions that allow sensitive and specific detection of target microbial groups, herein referred to as microbial signature sequences. We introduce Seqwin, an open-source framework designed to automate microbial genome signature discovery. Tens of thousands of microbial genomes are now available for a single species, limiting the application of existing manual and automated approaches for identifying signatures. Modern approaches that are capable of leveraging all available microbial genomes will ensure sensitive and accurate DNA signature identification and enable robust pathogen detection for clinical, environmental, and public health applications. RESULTS: Seqwin builds weighted pan-genome minimizer graphs and uses a traversal algorithm to identify signature sequences that occur frequently in target genomes but remain rare in non-targets. Unlike earlier tools that depend on strict presence or absence of sequences, Seqwin accommodates natural sequence variation and scales to very large genome collections. When applied to genomes from C. difficile, M. tuberculosis, and S. enterica, Seqwin recovered more high-quality signatures than alternative methods with lower computational burden. Seqwin's analysis of nearly 15 000 S. enterica genomes yielded over 200 candidate signatures in three minutes. Seqwin provides an open-source solution for the long-standing need for scalable microbial signature discovery and diagnostic assay design. AVAILABILITY AND IMPLEMENTATION: Seqwin is available on GitHub (https://github.com/treangenlab/Seqwin) and can be installed via Bioconda (https://bioconda.github.io/recipes/seqwin/README.html). Benchmarking datasets, outputs, and scripts are available on Zenodo (https://doi.org/10.5281/zenodo.19874011).

Software

De novo discovery of conserved gene clusters in microbial genomes with Spacedust.

Metagenomics has revolutionized environmental and human-associated microbiome studies. However, the limited fraction of proteins with known biological processes and molecular functions presents a major bottleneck. In prokaryotes and viruses, evolution favors keeping genes participating in the same biological processes colocalized as conserved gene clusters. Conversely, conservation of gene neighborhood indicates functional association. Here we present Spacedust, a tool for systematic, de novo discovery of conserved gene clusters. To find homologous protein matches, Spacedust uses fast and sensitive structure comparison with Foldseek. Partially conserved clusters are detected using novel clustering and order conservation P values. We demonstrate Spacedust's sensitivity with an all-versus-all analysis of 1,308 bacterial genomes, identifying 72,843 conserved gene clusters containing 58% of the 4.2 million genes. It recovered 95% of antiviral defense system clusters annotated by the specialized tool PADLOC. Spacedust's high sensitivity and speed will facilitate the annotation of large numbers of sequenced bacterial, archaeal and viral genomes.

Metagenomics

The role of microbial genomics in delivering the UK's national action plan for confronting antimicrobial resistance 2024-29.

Antimicrobial resistance (AMR) is a major threat to human and animal health, in addition to environmental resilience. Countries set the agenda on their national action against AMR in the form of National Action Plans (NAPs), with the UK's latest NAP released in May, 2024. Advances in genomics have strengthened our ability to work towards NAP priorities; however, to date, no mapping of the role genomics plays in contributing to specific goals within the NAP has been undertaken. The UK Research and Innovation-funded Transdisciplinary Antimicrobial Resistance Genomics Network brought together a range of stakeholders to discuss the role of genomics for action on AMR and to deliver policy priority-led research, as outlined in the UK NAP 2024-29. We report our discussions in this Personal View, with key roles for genomics, including informing targeted stewardship in health-care settings, supporting AMR literacy, and supporting effective antimicrobial innovation. However, changes in infrastructure, communication, and cross-sector coordination are needed to support implementation.

United Kingdom

Microbial genomic database of the Yangtze River, the third-longest river on Earth.

Microbes play an important role in mediating the nutrient cycling in the river ecosystem as a hotspot for biogeochemical processes. Due to scattered sampling efforts, however, there is a lack of a systematic study of the diversity of prokaryotic genomes in the Yangtze River, the third longest river on Earth. Here, we collected 602 metagenomic datasets of water, sediment and riparian soil samples spanning the Upper, Middle, and Lower basins of the Yangtze River over a 6,300 km continuum. We reconstructed 8,110 qualified genomes represented by 927 species-level genomes at the 95% ANI threshold, spanning 31 bacterial and five archaeal phyla. We further showed that more than half of these species (61.3% ~ 82.4%) were novel according to the genomic comparison against the curated databases, greatly expanding the known diversity of river prokaryotes. This dataset depicts an overview of microbial genomic diversity in the Yangtze River and provides a resource for in-depth investigation of metabolic potential, ecology, and evolution of riverine microbiomes.

Rivers

Gut microbial genomes with paired isolates from China illustrate probiotic and cardiometabolic effects.

The gut microbiome displays genetic differences among populations, and characterization of the genomic landscape of the gut microbiome in China remains limited. Here, we present the Chinese Gut Microbial Reference (CGMR) set, comprising 101,060 high-quality metagenomic assembled genomes (MAGs) of 3,707 nonredundant species from 3,234 fecal samples across primarily rural Chinese locations, 1,376 live isolates mainly from lactic acid bacteria, and 987 novel species relative to worldwide databases. We observed region-specific coexisting MAGs and MAGs with probiotic and cardiometabolic functionalities. Preliminary mouse experiments suggest a probiotic effect of two Faecalibacillus intestinalis isolates in alleviating constipation, cardiometabolic influences of three Bacteroides fragilis_A isolates in obesity, and isolates from the genera Parabacteroides and Lactobacillus in host lipid metabolism. Our study expands the current microbial genomes with paired isolates and demonstrates potential host effects, contributing to the mechanistic understanding of host-microbe interactions.

Probiotics

Bioprospecting microbial genomes to expand the biocatalytic toolbox of rubber oxygenases.

A set of rubber oxygenases was discovered through phylogenetic analysis and AI-based structural modeling of complexes of the putative enzymes with a substrate mimicking cis-1,4-polyisoprene. Sixteen candidate proteins were selected from thermophilic microorganisms, all sequence-related to the Latex clearing protein from Streptomyces sp. K30 (LcpK30). Sequence truncation and solubility tags were then evaluated to enhance protein expression, with the SUMO tag proving to be the most effective. Including LcpK30, nine heme-containing oxygenases were successfully expressed in E. coli NEB 10-beta cells, purified (35-157 mg L-1 yield) and characterized. Steady-state kinetics revealed significant rubber latex-degrading properties for six of them, with the truncated SUMO-fused LcpK30 (SUMO-LcpK30T) showing activity in agreement with literature. Notably, the catalytic efficiencies of all the expressed homologs lay within one order of magnitude and the oxygenase from Thermomonospora echinospora was found to be particularly promising in terms of activity, especially at high latex concentrations (more than 1% w/v). The analysis of reaction mixtures by both HPLC and HPLC-MS confirmed the oxidation of cis-1,4-polyisoprene to form the expected isoprenoid oligomers (n = 2-12), whose distribution was consistent with the usual endo-type cleavage pattern in all but one case. This bioprospecting effort afforded a platform of new rubber-degrading enzymes with diverse efficiencies and product profiles, capable of adapting to targeted applications.

Oxygenases

The planktonic microbiome of the Great Barrier Reef.

Large genome databases have markedly improved our understanding of marine microorganisms1-5. Although these resources have focused on prokaryotes, genomes from many dominant marine lineages, such as Pelagibacter and Prochlorococcus, are conspicuously underrepresented. Here we present the Great Barrier Reef Microbial Genomes Database (GBR-MGD), comprising 5,283 prokaryotic genomes obtained from Great Barrier Reef seawater samples using Nanopore and Illumina sequencing, including a collection of high-quality genomes of underrepresented groups. We show that standard short-read assemblies miss these populations owing to a combination of strain heterogeneity and low-GC-percentage sequencing bias. The GBR-MGD also comprises 20 chromosome-level picoeukaryote and 808,585 viral genomes, including a newly described clade of marine Crassvirales. We demonstrate the utility of the GBR-MGD to identify indicator taxa that can reliably predict the effects of reef management practices, such as the establishment of marine protected zones.

Bacteria

Toxoplasma gondii infection disrupts secondary bile acid transformation in feline gut microbiota.

UNLABELLED: Bile acid (BA) transformation relies on gut microbiota and is vulnerable to Toxoplasma gondii infection, yet feline microbial BA-transforming capacity upon toxoplasmosis remains unclear. Here, we constructed a catalog of 2,474 nonredundant feline gut microbial genomes and integrated serum metabolomic data to verify BA transformation alterations. The results revealed that the feline gut microbiome harbored widespread genetic potential for BA transformation but lacked a complete 7α-dehydroxylation pathway due to the absence of the key gene baiE. The BA transformation-related genomes (2,045 in total) were predominantly from the phyla Bacillota_A and Actinomycetota, among which only 37 encoded baiB, all belonging to Bacillota_A. The distribution of BA transformation-related genes varied across intestinal regions: genes encoding 7α-HSDH were primarily enriched in the small intestine, whereas genes encoding 3α-HSDH, baiCD, and baiH were more abundant in the large intestine. Additionally, the abundance of genes encoding BSH and 3α-HSDH increased significantly in the small intestine on day 3 post-infection, accompanied by increases in the phylum Bacillota_C and genera such as Blautia_A, Enterococcus_E, and Ligilactobacillus. Serum metabolomics revealed a significant increase in cholesterol levels post-infection, supporting the impact of T. gondii infection on intestinal BA transformation. These findings illustrated that the feline gut microbiota played an important role in BA transformation and that T. gondii infection disrupted the microbial potential for secondary BA transformation. This study provided new insights into gut microbiota-associated metabolic perturbations during feline toxoplasmosis. IMPORTANCE: Bile acid (BA) transformation plays a critical role in host metabolism and immune regulation. Although studies on BA transformation are increasing, the capacity for BA transformation within the feline gut microbiota and the impact of Toxoplasma gondii infection on this capacity remain unclear. To bridge this gap, we constructed a catalog of 2,474 nonredundant feline gut microbial genomes and integrated serum metabolomic data to verify BA transformation alterations. Our findings revealed that the feline gut microbiome lacked a complete 7α-dehydroxylation pathway, and the specific functions involved in BA transformation may differ between the small and large intestines. Furthermore, integrated metagenomic and serum metabolomic analyses suggested that T. gondii infection disrupted BA transformation capacity in the small intestine. This study provided new insights into gut microbiota-associated metabolic perturbations during feline toxoplasmosis.

Toxoplasma gondii

Characterization of microbial dark matter at scale with MetaSBT and taxonomy-aware Sequence Bloom Trees.

Metagenomics has become a powerful tool for studying microbial communities, allowing researchers to investigate microbial diversity within complex environmental samples. Recent advances in sequencing technology have enabled the recovery of near-complete microbial genomes directly from metagenomic samples, also known as metagenome-assembled genomes (MAGs). However, accurately characterizing these genomes remains a significant challenge due to the presence of sequencing errors, incomplete assembly, and contamination. Here we present MetaSBT, a new tool for organizing, indexing, and characterizing microbial reference genomes and MAGs. It is able to identify clusters of genomes at all seven taxonomic levels, from the kingdom all the way down to the species level, using the Sequence Bloom Tree (SBT) data structure that relies on Bloom Filters (BFs) to index massive amounts of genomes based on their k-mers composition. We have built an initial set of databases composed of over 190 thousand viral genomes from NCBI GenBank and public sources grouped into sequence consistent clusters at different taxonomic levels, making it the first software solution for the classification of viruses at different ranks, including still unknown ones. This results in the definition of over 40 thousand species clusters where ~80% do not match with any known viral species in reference databases to date. Furthermore, we show how our databases can be used as a new basis for existing quantitative metagenomic profilers to unlock the detection of unknown microbes and the estimation of their abundance in metagenomic samples. Finally, the framework is released open-source and, along with its public databases, is fully integrated into the Galaxy Platform enabling broad accessibility.

metagenome-assembled genomes

A genomic catalog of Earth's bacterial and archaeal symbionts.

Microbial symbiosis drives the functional and phylogenomic diversification of life on Earth yet remains underexplored because of culturing challenges. This study used machine learning (ML) to predict symbiotic lifestyles in more than a hundred thousand microbial genomes from diverse environmental metagenome samples and reference genomes. Predictions were performed using symclatron, an ML framework developed to identify genomic signatures of symbionts. Predictions were deposited in a catalog we established called Symbiont Genomes (SymGs). The results indicate that 15-23% of uncultivated microorganisms likely engage in symbiotic relationships with other organisms, categorized as host-associated or obligate intracellular lifestyles, and are present in half of all known bacterial and archaeal phyla. We also identify genomic signatures of symbiotic lifestyles, including the loss of certain metabolic functions and the differential presence of metabolic modules that may enable host-dependent living. The symclatron software and the SymGs catalog represent valuable resources for studying symbioses, potentially facilitating future mechanistic investigations and engineering of host-microorganism associations.

Journal Article

A Graph Contrastive Learning Method for Enhancing Genome Recovery in Complex Microbial Communities.

Accurate genome binning is essential for resolving microbial community structure and functional potential from metagenomic data. However, existing approaches-primarily reliant on tetranucleotide frequency (TNF) and abundance profiles-often perform sub-optimally in the face of complex community compositions, low-abundance taxa, and long-read sequencing datasets. To address these limitations, we present MBGCCA, a novel metagenomic binning framework that synergistically integrates graph neural networks (GNNs), contrastive learning, and information-theoretic regularization to enhance binning accuracy, robustness, and biological coherence. MBGCCA operates in two stages: (1) multimodal information integration, where TNF and abundance profiles are fused via a deep neural network trained using a multi-view contrastive loss, and (2) self-supervised graph representation learning, which leverages assembly graph topology to refine contig embeddings. The contrastive learning objective follows the InfoMax principle by maximizing mutual information across augmented views and modalities, encouraging the model to extract globally consistent and high-information representations. By aligning perturbed graph views while preserving topological structure, MBGCCA effectively captures both global genomic characteristics and local contig relationships. Comprehensive evaluations using both synthetic and real-world datasets-including wastewater and soil microbiomes-demonstrate that MBGCCA consistently outperforms state-of-the-art binning methods, particularly in challenging scenarios marked by sparse data and high community complexity. These results highlight the value of entropy-aware, topology-preserving learning for advancing metagenomic genome reconstruction.

canonical correlation analysis

Promises and pitfalls of long-read sequencing for resolving microbial complexity.

Long-read sequencing (LRS) has driven a transition in microbial genomics, overcoming the assembly fragmentation inherent to short-read sequencing. This review elucidates the impact of LRS across isolate genomics, metagenomics, and multi-omics domains. By spanning extensive repetitive regions, LRS facilitates the reconstruction of circular chromosomes and precisely resolves mobile genetic elements (MGEs). In metagenomics, LRS enables strain-level resolution, the recovery of circular metagenome-assembled genomes, and the precise localization of MGEs within host replicons. Furthermore, the single-molecule, amplification-free properties of LRS provide enhanced resolution of native epigenetic modifications and full-length transcriptomes. Despite these advancements, widespread implementation remains constrained by multidimensional challenges, including stringent high-molecular-weight DNA requirements, depth deficits, and computational overhead. Nevertheless, LRS is increasingly becoming the method of choice for isolate genomics and metagenomics. As detection technologies and algorithms progress, LRS will further improve our ability to decipher the structural and functional diversity of microbial ecosystems.

Metagenomics

Comparative metagenomic assessment of Illumina-compatible library preparation methods, short-read lengths, and PacBio HiFi sequencing reveals differences in microbial and functional diversity recovery from a complex environmental sample.

UNLABELLED: Metagenomics enables comprehensive exploration of microbial communities but is influenced by library preparation and sequencing technologies, affecting recovery of microbial genomes and proteins. Here, we benchmarked six Illumina-compatible short-read library preparation conditions in triplicate at 2 × 150 bp and 2 × 250 bp read lengths alongside PacBio HiFi long-read sequencing using a composite environmental sample of marine mangrove sediment and terrestrial palm tree soil. Longer short reads (2 × 250 bp) combined with optimal library preparation approaches improved assembly quality, protein detection, and metagenome-assembled genome (MAG) recovery, achieving results approaching those of long-read sequencing. TruSeq libraries at 2 × 250 bp recovered more than sevenfold more unique proteins than the same kit at 2 × 150 bp (811,701 vs 110,108) using the same number of sequencing reads, while recovering a comparable number of high-quality MAGs to PacBio HiFi long-read sequencing (11 vs 18) and surpassing it in protein discovery by almost 10-fold (811,701 vs 87,745) at less than half of the sequencing cost. Furthermore, biosynthetic gene cluster analysis identified 46 biosynthetic gene clusters in TruSeq-250PE assemblies compared to 38 in PacBio HiFi, with several showing no close match in the MIBiG database. Although long reads yield more contiguity and complete genomes, longer short reads offer a cost-effective, scalable alternative for uncovering microbial and functional diversity. These findings provide critical guidance for metagenomic experimental design, demonstrating that strategic selection of library preparation chemistry and sequencing parameters can reveal more unknown microbial information in complex biomes without requiring additional sequencing depth. IMPORTANCE: Metagenomic outcomes are strongly influenced by library preparation and sequencing strategies, yet their combined effects in complex environmental samples remain poorly defined. Here, we provide the first direct comparison of Illumina NovaSeq short-read metagenomic sequencing at 2 × 150 bp and 2 × 250 bp across multiple library preparation kits, alongside PacBio HiFi long-read sequencing. We show that sequencing read length and library preparation critically shape assembly quality, protein recovery, and metagenome-assembled genome (MAG) reconstruction. These findings demonstrate that short-read sequencing at 2 × 250 bp, with appropriate library preparation, can match long-read technologies in MAG recovery while substantially surpassing them in protein discovery. With less than half of the sequencing price and a 3.5-fold reduction in cost per gigabase of usable data, this method facilitates more accessible large-scale metagenomic analysis within complex environmental systems.

Metagenomics

Predictions of rhizosphere microbiome dynamics with a genome-informed and trait-based energy budget model.

Soil microbiomes are highly diverse, and to improve their representation in biogeochemical models, microbial genome data can be leveraged to infer key functional traits. By integrating genome-inferred traits into a theory-based hierarchical framework, emergent behaviour arising from interactions of individual traits can be predicted. Here we combine theory-driven predictions of substrate uptake kinetics with a genome-informed trait-based dynamic energy budget model to predict emergent life-history traits and trade-offs in soil bacteria. When applied to a plant microbiome system, the model accurately predicted distinct substrate-acquisition strategies that aligned with observations, uncovering resource-dependent trade-offs between microbial growth rate and efficiency. For instance, inherently slower-growing microorganisms, favoured by organic acid exudation at later plant growth stages, exhibited enhanced carbon use efficiency (yield) without sacrificing growth rate (power). This insight has implications for retaining plant root-derived carbon in soils and highlights the power of data-driven, trait-based approaches for improving microbial representation in biogeochemical models.

Rhizosphere

Widespread marine and freshwater distributions of active sulfoquinovose-degrading bacteria.

Sulfoquinovose (SQ), a sulfonated sugar produced on a gigaton scale each year, contributes to global sulfur cycling, yet the microbes and pathways mediating its turnover in the environment have been inferred largely from genomic potential rather than direct activity. Here, we coupled incubations of environmental samples with 13C-labeled SQ to DNA-stable isotope probing to identify active SQ carbon assimilators across estuary, mangrove, and lake ecosystems. In estuarine communities, Vibrio and Cognatishimia incorporated SQ-derived 13C; Novosphingobium dominated in the mangrove, and Agrobacterium in the lake. Pure-culture experiments, coupled with comparative proteomics and gene knockout validation, demonstrated that Vibrio strains degrade SQ via modified sulfoglycolytic Embden-Meyerhof-Parnas and Entner-Doudoroff pathways to produce the environmentally significant organosulfur 2,3-dihydroxypropanesulfonate. Comparative genomic analyses suggested that closely related genome representatives of Novosphingobium, Cognatishimia, and Agrobacterium encode the sulfolytic SQ monooxygenase pathway. A global survey of aquatic microbial genomes indicated that over 9% harbor SQ degradation clusters, supporting a widespread distribution of bacterial SQ catabolic potential in aquatic environments.

Fresh Water

Genome mining for new enediyne antibiotics.

Enediyne antibiotics epitomize nature's chemical creativity. They contain intricate molecular architectures that are coupled with potent biological activities involving double-stranded DNA scission. The recent explosion in microbial genome sequences has revealed a large reservoir of novel enediynes. However, while hundreds of enediyne biosynthetic gene clusters (BGCs) can be detected, less than two dozen natural products have been characterized to date as many clusters remain silent or sparingly expressed under standard laboratory growth conditions. This review focuses on four distinct strategies, which have recently enabled discoveries of novel enediynes: phenotypic screening from rare sources, biosynthetic manipulation, genomic signature-based PCR screening, and DNA-cleavage assays coupled with activation of silent BGCs via high-throughput elicitor screening. With an abundance of enediyne BGCs and emerging approaches for accessing them, new enediyne natural products and further insights into their biogenesis are imminent.

Enediynes