Search PubMed⌕ Search

Biomedical subjects

Antonio Pedro Camargo

Publications and source records attributed to Antonio Pedro Camargo.

3 recordsLinked to original sources

Computational tool choice impacts CRISPR spacer-protospacer detection.

MOTIVATION: CRISPR spacer-protospacer matching is widely used to infer host-virus interactions in microbial and viromics studies, but the choice of sequence search or alignment tool and its reporting behavior is often under-evaluated for this specific task. RESULTS: Using synthetic, semi-synthetic, and real datasets, we benchmarked commonly used tools and observed substantial differences in recall, runtime, and resource usage across distance metrics and thresholds. Our analyses support practical defaults for large-scale spacer-target matching and clarify trade-offs between exhaustive and heuristic approaches. AVAILABILITY: Source code and benchmark workflows are available at https://github.com/UriNeri/spacer_matching_bench. Data and run artifacts are archived on Zenodo (https://doi.org/10.5281/zenodo.15171878).

Software↗

Logan: Planetary-Scale Genome Assembly Surveys Life's Diversity.

The breadth of life's diversity is unfathomable, but public nucleic acid sequencing data offers a window into the dispersion and evolution of genetic diversity across Earth. However the rapid growth and accumulation of sequence data have outpaced efficient analysis capabilities. The largest collection of freely available sequencing data is the Sequence Read Archive (SRA), comprising 27.3 million datasets or 5 × 1016 basepairs. To realize the potential of the SRA, we constructed Logan, a massive sequence assembly transforming short reads into long contigs and compressing the data over 100-fold, enabling highly efficient petabase-scale analysis. We created Logan-Search, a k-mer index of Logan for free planetary-scale sequence search, returning matches in minutes. We used Logan contigs to identify >200 million plastic-degrading enzyme homologs, and validate novel enzymes with catalytic activities exceeding current reference standards. Further, we vastly expand the known diversity of proteins (30-fold over UniRef50), plasmids (22-fold over PLSDB), P4 satellites (4.5-fold), and the recently described Obelisk RNA elements (3.7-fold). Logan also enables ecological and biomedical data mining, such as global tracking of antimicrobial resistance genes and the characterization of viral reactivation across millions of human BioSamples. By transforming the SRA, Logan democratizes access to the world's public genetic data and opens frontiers in biotechnology, molecular ecology, and global health.

Journal Article↗

CoverM: read alignment statistics for metagenomics.

SUMMARY: Genome-centric analysis of metagenomic samples is a powerful method for understanding the function of microbial communities. Calculating read coverage is a central part of analysis, enabling differential coverage binning for recovery of genomes and estimation of microbial community composition. Coverage is determined by processing read alignments to reference sequences of either contigs or genomes. Per-reference coverage is typically calculated in an ad-hoc manner, with each software package providing its own implementation and specific definition of coverage. Here we present a unified software package CoverM which calculates several coverage statistics for contigs and genomes in an ergonomic and flexible manner. It uses "Mosdepth arrays" for computational efficiency and avoids unnecessary I/O overhead by calculating coverage statistics from streamed read alignment results. AVAILABILITY AND IMPLEMENTATION: CoverM is free software available at https://github.com/wwood/coverm. CoverM is implemented in Rust, with Python (https://github.com/apcamargo/pycoverm) and Julia (https://github.com/JuliaBinaryWrappers/CoverM_jll.jl) interfaces.

Metabolomics↗