Search PubMedSearch

Biomedical subjects

Noam Auslander

Publications and source records attributed to Noam Auslander.

2 recordsLinked to original sources

Uncovering viral protein acquisition events and human-specific folds with pairwise comparisons of predicted protein structures.

Pairwise sequence comparisons are at the center of molecular evolutionary analyses. However, viral pairwise comparisons are challenging because extreme mutation rates and evolutionary pressure cause genomes to diverge rapidly, limiting detectable sequence similarity to fewer than 3% of virus pairs. To overcome these limitations, we compared viruses based on structural similarity, using predicted protein structures from ColabFold and Foldseek to define protein fold clusters. We represented each virus genome by its protein structural content. Pairwise similarities between viruses were then quantified using the Jaccard index based on the presence or absence of protein fold clusters. Using a recently established viral protein fold database, we compared all pairs of eukaryotic viruses in RefSeq. This approach increased the proportion of comparable viral genome pairs from 2.4% to 16.5%. Using this protein-fold representation of viruses, we were able to accurately predict viral families with an average sensitivity of 85.9%. Investigation of viral families showing limited sensitivity with this approach uncovered a laterally transferred structural cluster (Rep/NS1) broadly shared across diverse viral families and found in the avian lineage of adenoviruses. Sequence homology suggests that this Rep was acquired from Parvoviridae, but the protein is mutant in the ATPase active site, indicating possible exaptation toward a purely DNA-binding function. In Gammapapillomaviruses, several E4 clusters were associated with human tropism. In summary, by representing viruses with structural protein clusters, we can classify highly divergent viruses, trace lateral gene transfer, and uncover features associated with viral host range.

Humans

kMermaid: Ultrafast metagenomic read assignment to protein clusters by hashing of amino acid k-mer frequencies.

Shotgun metagenomic sequencing can determine both the taxonomic and functional content of microbiomes. However, functional classification for metagenomic reads remains highly challenging as protein mapping tools require substantial computational resources and yield ambiguous classifications when short reads map to homologous proteins originating from different bacteria. Here we introduce kMermaid for the purpose of uniquely mapping bacterial short reads to taxa-agnostic clusters of homologous proteins, which can then be used for downstream analysis tasks such as read quantification and pathway or global functional analysis. Using a nested hash map containing amino acid k-mer profiles as a model for protein assignment, kMermaid achieves the sensitivity of popular existing protein mapping tools while remaining highly resource efficient. We evaluate kMermaid on simulated data and data from human fecal samples as well as demonstrate the utility of kMermaid for classifying reads originating from new, unseen proteins. kMermaid allows for highly accurate, unambiguous and ultrafast metagenomic read assignment into protein clusters, with a fixed memory usage, and can easily be employed on a typical computer.

Metagenomics