Search PubMedSearch

Biomedical subjects

Martin Steinegger

Publications and source records attributed to Martin Steinegger.

2 recordsLinked to original sources

Easy and interactive taxonomic profiling with Metabuli App.

SUMMARY: Accurate metagenomic taxonomic profiling is critical for understanding microbial communities. However, computational analysis often requires command-line proficiency and high-performance computing resources. To lower these barriers, we developed Metabuli App, an all-in-one desktop application that efficiently runs taxonomic profiling locally on a consumer-grade computer. It features user-friendly graphical interfaces for custom database curation, raw read quality control (QC), taxonomic profiling, and interactive result visualization. AVAILABILITY AND IMPLEMENTATION: GPLv3-licensed source code and prebuilt apps for Windows, macOS, and Linux are available at https://github.com/steineggerlab/Metabuli-App and are archived at https://doi.org/10.5281/zenodo.15876171. Analysis scripts are available at https://github.com/jaebeom-kim/metabuli-app-analysis. The Sankey-based taxonomy visualization component is available at https://github.com/steineggerlab/taxoview for easy integration into other web projects.

Software

Logan: Planetary-Scale Genome Assembly Surveys Life's Diversity.

The breadth of life's diversity is unfathomable, but public nucleic acid sequencing data offers a window into the dispersion and evolution of genetic diversity across Earth. However the rapid growth and accumulation of sequence data have outpaced efficient analysis capabilities. The largest collection of freely available sequencing data is the Sequence Read Archive (SRA), comprising 27.3 million datasets or 5 × 1016 basepairs. To realize the potential of the SRA, we constructed Logan, a massive sequence assembly transforming short reads into long contigs and compressing the data over 100-fold, enabling highly efficient petabase-scale analysis. We created Logan-Search, a k-mer index of Logan for free planetary-scale sequence search, returning matches in minutes. We used Logan contigs to identify >200 million plastic-degrading enzyme homologs, and validate novel enzymes with catalytic activities exceeding current reference standards. Further, we vastly expand the known diversity of proteins (30-fold over UniRef50), plasmids (22-fold over PLSDB), P4 satellites (4.5-fold), and the recently described Obelisk RNA elements (3.7-fold). Logan also enables ecological and biomedical data mining, such as global tracking of antimicrobial resistance genes and the characterization of viral reactivation across millions of human BioSamples. By transforming the SRA, Logan democratizes access to the world's public genetic data and opens frontiers in biotechnology, molecular ecology, and global health.

Journal Article