Search PubMedSearch

Biomedical subjects

Barak A Cohen

Publications and source records attributed to Barak A Cohen.

3 recordsLinked to original sources

Systematic evaluation of metatranscriptomic differential gene expression in silico, in vitro, and in vivo enables elucidation of inter-species cross-feeding.

Metatranscriptomic (MTX) sequencing quantifies gene expression from the collective genomes of microbial communities (microbiomes), enabling assessment of functional activity rather than functional potential. While differential expression testing is instrumental to RNA-sequencing analysis, current metatranscriptomic approaches have been benchmarked only on simulated data and not under real operating conditions, resulting in a lack of standard practices. Here, we evaluate the performance of statistical differential expression methods on both simulated datasets and data collected from real bacterial 'mock communities' designed for this purpose. We assess the robustness of individual methods to organisms' low relative abundance, differential abundance, low prevalence, and transcription rate changes, showing that no existing methods perform adequately across all confounding conditions. We then apply the same approaches to metatranscriptomic datasets generated from gnotobiotic mice colonized with defined consortia of human bacterial strains and show that the method nominated by our mock community comparisons successfully inferred cross-feeding dynamics which were validated in vitro. We conclude that MTX method benchmarking on real, not simulated, datasets can and should optimize model implementation, enabling inference and validation of cross-feeding and other inter-species and host-microbe dynamics from in vivo studies.

Journal Article

Active learning of enhancers and silencers in the developing neural retina.

Deep learning is a promising strategy for modeling cis-regulatory elements. However, models trained on genomic sequences often fail to explain why the same transcription factor can activate or repress transcription in different contexts. To address this limitation, we developed an active learning approach to train models that distinguish between enhancers and silencers composed of binding sites for the photoreceptor transcription factor cone-rod homeobox (CRX). After training the model on nearly all bound CRX sites from the genome, we coupled synthetic biology with uncertainty sampling to generate additional rounds of informative training data. This allowed us to iteratively train models on data from multiple rounds of massively parallel reporter assays. The ability of the resulting models to discriminate between CRX sites with identical sequence but opposite functions establishes active learning as an effective strategy to train models of regulatory DNA. A record of this paper's transparent peer review process is included in the supplemental information.

Retina

Active learning of enhancer and silencer regulatory grammar in photoreceptors.

Cis-regulatory elements (CREs) direct gene expression in health and disease, and models that can accurately predict their activities from DNA sequences are crucial for biomedicine. Deep learning represents one emerging strategy to model the regulatory grammar that relates CRE sequence to function. However, these models require training data on a scale that exceeds the number of CREs in the genome. We address this problem using active machine learning to iteratively train models on multiple rounds of synthetic DNA sequences assayed in live mammalian retinas. During each round of training the model actively selects sequence perturbations to assay, thereby efficiently generating informative training data. We iteratively trained a model that predicts the activities of sequences containing binding motifs for the photoreceptor transcription factor Cone-rod homeobox (CRX) using an order of magnitude less training data than current approaches. The model's internal confidence estimates of its predictions are reliable guides for designing sequences with high activity. The model correctly identified critical sequence differences between active and inactive sequences with nearly identical transcription factor binding sites, and revealed order and spacing preferences for combinations of motifs. Our results establish active learning as an effective method to train accurate deep learning models of cis-regulatory function after exhausting naturally occurring training examples in the genome.

Journal Article