Search PubMedSearch

Biomedical subjects

Steven L Salzberg

Publications and source records attributed to Steven L Salzberg.

7 recordsLinked to original sources

Challenges and Opportunities in Analyzing Cancer-Associated Microbiomes.

The study of cancer-associated microbiomes has gained significant attention in recent years, spurred by advances in high-throughput sequencing and metagenomic analysis. Microbiome research holds promise for identifying noninvasive biomarkers and possibly new paradigms for cancer treatment. In this review, we explore the key computational challenges and opportunities in analyzing cancer-associated microbiomes (in tumor/normal tissues and other body sites, e.g., gut, oral, and skin), focusing on sequencing-driven strategies and associated considerations for taxonomic and functional characterization. The discussion covers the strengths and limitations of current analysis tools for identifying contamination, determining compositional bias, and resolving species and strains, as well as the statistical, metabolic, and network inferences that are essential to uncover host-microbiome interactions. Several key considerations are required to guide the choice of databases used for metagenomic analysis in such studies. Recent advances in spatial and single-cell technologies have provided insights into cancer-associated microbiomes, and Artificial Intelligence-driven protein function prediction might enable rapid advances in this field. Finally, we provide a perspective on how the field can evolve to manage the ever-growing size of datasets and generate robust and testable hypotheses. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

Comprehensive Transcriptome Annotation of Thousands of HIV-1 Genomes.

Alternative splicing in HIV-1 has been a central focus of decades of research, uncovering key mechanisms of viral gene regulation, immune evasion, and therapeutic response - yet, no reference resource has existed to support transcriptome-wide analysis, limiting adoption of modern computational methods. We present HIV Atlas (https://ccb.jhu.edu/HIV_Atlas), the first reference-quality annotation of HIV-1 and SIV transcriptional diversity. We manually curated transcriptomes for HIV-1HXB2 and SIVmac239 and developed Vira, an automated annotation-transfer method specifically designed to address unique challenges of viral genome biology, to generate high-quality annotations for 2,077 complete HIV-1 genomes. Using the resources presented in our work, we evaluated conservation of splice sites, revealing near-perfect preservation of major donors and acceptors. Furthermore, using several public datasets, we demonstrate how HIV Atlas enhances methodology, improves the quality and novelty of results, and opens novel avenues for research, supporting more accurate and comprehensive analyses of bulk, single-cell, and spatial RNA-seq in HIV-1 studies.

Journal Article

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article

Efficient evidence-based genome annotation with EviAnn.

For many years, machine learning-based ab initio gene finding approaches have been central components of eukaryotic genome annotation pipelines, and they remain so today. The reliance on these approaches was originally sustained by the high cost and low availability of gene expression data, a primary source of evidence for gene annotation along with protein homology. However, innovations in modern sequencing technologies have revolutionized the acquisition of gene expression data, allowing scientists to rely more heavily on this class of evidence. In addition, proteins found in a multitude of well-annotated genomes represent another invaluable resource for gene annotation. Existing annotation packages often underutilize these data sources, which prompted us to develop EviAnn (Evidence-based Annotator), a novel evidence-based eukaryotic gene annotation system. EviAnn takes a strongly data-driven approach, building the exon-intron structure of genes from transcript alignments or protein-sequence homology rather than from purely ab initio gene finding techniques. We show that when provided with the same input data, EviAnn consistently outperforms current state-of-the-art packages including BRAKER3, MAKER2, and FINDER, while utilizing considerably less computer time. Annotation of a mammalian genome can be completed in less than an hour on a single multi-core server. EviAnn is freely available under an open-source license from https://github.com/alekseyzimin/EviAnn_release and from Bioconda as "eviann".

Journal Article

OpenSpliceAI: An efficient, modular implementation of SpliceAI enabling easy retraining on non-human species.

The SpliceAI deep learning system is currently one of the most accurate methods for identifying splicing signals directly from DNA sequences. However, its utility is limited by its reliance on older software frameworks and human-centric training data. Here we introduce OpenSpliceAI, a trainable, open-source version of SpliceAI implemented in PyTorch to address these challenges. OpenSpliceAI supports both training from scratch and transfer learning, enabling seamless retraining on species-specific datasets and mitigating human-centric biases. Our experiments show that it achieves faster processing speeds and lower memory usage than the original SpliceAI code, allowing large-scale analyses of extensive genomic regions on a single GPU. Additionally, OpenSpliceAI's flexible architecture makes for easier integration with established machine learning ecosystems, simplifying the development of custom splicing models for different species and applications. We demonstrate that OpenSpliceAI's output is highly concordant with SpliceAI. In silico mutagenesis (ISM) analyses confirm that both models rely on similar sequence features, and calibration experiments demonstrate similar score probability estimates.

Journal Article

Genomic variability in Zika virus in GBS cases in Colombia.

Major clusters of Guillain-Barré Syndrome (GBS) emerged during the Zika virus (ZIKV) outbreaks in the South Pacific and the Americas from 2014 to 2016. The factors contributing to GBS susceptibility in ZIKV infection remain unclear, although considerations of viral variation, patient susceptibility, environmental influences, and other potential factors have been hypothesized. Studying the role of viral genetic factors has been challenging due to the low viral load and rapid viral clearance from the blood after the onset of Zika symptoms. The prolonged excretion of ZIKV in urine by the time of GBS onset, when the virus is no longer present in the blood, provides an opportunity to unravel whether specific ZIKV mutations are related to the development of GBS in certain individuals. This study aimed to investigate the association between specific ZIKV genotypes and the development of GBS, taking advantage of a unique collection of ZIKV-positive urine samples obtained from GBS cases and controls during the 2016 ZIKV outbreak in Colombia. Utilizing Oxford-Nanopore technology, we conducted complete genome sequencing of ZIKV in biological samples from 15 patients with GBS associated with ZIKV and 17 with ZIKV infection without neurological complications. ZIKV genotypes in Colombia exhibited distribution across three clades (average bootstrap of 90.9±14.9%), with two clades dominating the landscape. A comparative analysis of ZIKV genomes from GBS and non-neurological complications, alongside 1368 previously reported genomes, revealed no significant distinctions between the two groups. Both genotypes were similarly distributed among observed clades in Colombia. Furthermore, no variations were identified in the amino acid composition of the viral genome between the two groups. Our findings suggest that GBS in ZIKV infection is perhaps associated with patient susceptibility and/or other para- or post-infectious immune-mediated mechanisms rather than with specific ZIKV genome variations.

Zika Virus