Search PubMedSearch

Biomedical subjects

Anshul Kundaje

Publications and source records attributed to Anshul Kundaje.

9 recordsLinked to original sources

Flexible use of conserved motifs constrains genome access in cell type evolution.

Cell types can be organized into related families, but the regulatory mechanisms that define and maintain these families across deep evolutionary time remain unknown. Here, combining single-nucleus multi-omic sequencing with deep learning to analyse the accessible genomes of two groups of vastly divergent animals including flatworms and vertebrates, we find that hundreds of accessibility-dictating sequence motifs partition into distinct yet conserved sets, or 'vocabularies', each associated with a specific cell type family. However, combinatorial relationships among these motifs preferred by individual cell types are largely species specific. Deep-learning models trained on one species accurately predict family-level chromatin accessibility in distantly related species, albeit frequently rely on different motifs from shared vocabularies to reach convergent predictions. By contrast, models trained on individual cell types within a family lose cross-species predictive power, indicating that the regulatory syntax governing cell type-level identity evolves rapidly. We propose a 'collective maintenance' model in which motif vocabularies defining cell type families are evolutionarily stable, while recombination of these motifs generates cell type-specific regulatory programmes. This suggests that family identity is maintained collectively by large, conserved pools of regulatory factors, analogous to the logic of developmental homology, where character identity persists through network-level conservation despite extensive rewiring.

Journal Article

Massively parallel characterization and predictive modelling of neuronal regulatory variation.

Disease-associated variants reside frequently in noncoding cis-regulatory elements (CREs), yet their functional consequences remain poorly understood. We performed a large-scale lentiMPRA in human excitatory neurons, quantifying the impact of >46,000 naturally occurring variants across >27,000 candidate CREs near 524 disease-associated genes. These data improved regulatory variant effect predictions beyond state-of-the-art models. Significant allelic effects occurred at comparable rates across common, rare, and singleton variants, demonstrating that, within MPRA-measurable effects, population frequency carries limited information about per-variant regulatory impact. Variant effect detectability and magnitude were governed primarily by baseline activity of the enclosing regulatory element and local sequence context. Regulatory effects were distributed across numerous transcription factors rather than concentrated in master regulators, consistent with a combinatorial enhancer architecture. We establish a large-scale functional variant catalog and provide a complementary benchmark and resource for developing and evaluating models of noncoding regulatory variation.

Journal Article

An encyclopedia of human enhancer-gene regulatory interactions.

Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1-6. Here we create and evaluate a resource of more than 92 million enhancer-gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element-gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study (GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer-promoter contacts, additional features that guide enhancer-promoter communication include promoter class and enhancer-enhancer synergy. These genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.

Humans

Sensitive, direct detection of non-coding off-target base editor unwinding and editing in primary cells.

Base editors create precise nucleotide changes in DNA, but their off-target activity remains challenging to quantify. Here, we develop and deploy a direct, in cellulo sequencing assay that simultaneously measures both Cas9-mediated unwinding and deaminase editing of genomic DNA (beCasKAS). Our strategy nominates >460-fold more potential off-target sites than other methods by enriching for Cas9-dependent R-loops immediately preceding editing. Using beCasKAS in primary human T-cells, we observe that mRNA-encoded ABE8e and PAMless ABE8e-SpRY base editors have distinct off-target profiles that can be mitigated by optimizing mRNA dose. Finally, we combine beCasKAS with base-resolution deep learning models to risk-stratify off-target edits by their likelihood of epigenetic dysregulation. Collectively, beCasKAS offers a sensitive and facile tool to optimize the balance between base editor on- and off-target activity.

Journal Article

Enhancer-targeting CRISPR screens at coronary artery disease loci suggest shared mechanisms of disease risk.

To systematically identify causal genetic mechanisms that confer risk for coronary artery disease (CAD) in GWAS loci, we mapped genome-wide variant-to-enhancer-to-gene (V2E2G) links in vascular smooth muscle cells (SMC). Enhancers identified by active chromatin features, and further prioritized by base-resolution deep learning models of chromatin accessibility in 108 CAD loci, were studied with CRISPRi targeting and Direct-Capture Targeted Perturb-seq (DC-TAP-seq) evaluation of 470 genes. Seventy-six V2E2G links were identified for 59 candidate CAD genes representing gene programs including epithelial-mesenchymal transformation, ubiquitination, and protein folding as well as BMP and TGFB signaling. Similar methods employed with an independent focused screen targeting one candidate locus at 9p21.3 identified 10 enhancers regulating expression of multiple genes at this location. Detailed molecular studies revealed that two enhancers mediating transcription factor binding and transcriptional regulation contribute to ancestry-specific and sex-specific risk for CAD and the surrogate biomarker vascular calcification. Together, these studies advance our identification of GWAS CAD V2E2G links across the genome, and specific mechanisms of risk at the complex 9p21.3 locus.

Journal Article

Genetic risk factors modulate the association between physical activity and colorectal cancer.

BACKGROUND: Physical activity (PA) is an established protective factor for colorectal cancer (CRC), but it is unclear if genetic variants modify this effect. To investigate this possibility, we conducted a genome-wide gene-PA interaction analysis. METHODS: Using logistic regression and two-step and joint tests, we analyzed interactions between common genetic variants across the genome and PA in relation to CRC risk. Self-reported PA levels were categorized as active (&#x2265; 8.75 MET-h/wk) vs. inactive (< 8.75 MET-h/wk) and as study- and sex-specific quartiles of activity. RESULTS: PA had an overall protective effect on CRC (OR [active vs. inactive] = 0.85; 95%CI = 0.81-0.90). The two-step GxE method identified an interaction between rs4779584, an intergenic variant near the GREM1 and SCG5 genes, and PA for CRC risk (p-interaction = 2.6&#xd7;10- 8). Stratification by genotype at this locus showed a significant reduction in CRC risk by 20% in active vs. inactive participants with the CC genotype (OR = 0.80; 95%CI = 0.75-0.85), but no significant PA-CRC association among CT or TT carriers. When PA was modeled as quartiles, the 1-d.f. GxE test identified that rs56906466, an intergenic variant near the KCNG1 gene, modified the association between PA and CRC (p-interaction = 3.5&#xd7;10- 8). Stratification at this locus showed that increase in PA (highest vs. lowest quartile) was associated with a lower CRC risk solely among TT carriers (OR = 0.77; 95%CI = 0.72-0.82). CONCLUSIONS: In summary, we identified two genetic variants that modified the association between PA and CRC risk. One of them, related to GREM1 and SCG5, suggests that the bone morphogenetic protein (BMP)-related, inflammatory, and/or insulin signaling pathways may be associated with the protective influence of PA on colorectal carcinogenesis.

GWAS

Prediction and functional interpretation of inter-chromosomal genome architecture from DNA sequence with TwinC.

Three-dimensional nuclear DNA architecture comprises well-studied intra-chromosomal (cis) folding and less characterized inter-chromosomal (trans) interfaces. Current predictive models of 3D genome folding can effectively infer pairwise cis-chromatin interactions from the primary DNA sequence but generally ignore trans contacts. There is an unmet need for robust models of trans-genome organization that provide insights into their underlying principles and functional relevance. We present TwinC, an interpretable convolutional neural network model that reliably predicts trans contacts measurable through proximity ligation-dependent (in situ and intact Hi-C) and independent (DNA SPRITE) genome-wide chromatin conformation assays. . TwinC uses a paired sequence design from replicate Hi-C experiments to learn single base pair relevance in trans interactions across two stretches of DNA. The method achieves high predictive accuracy (AUROC=0.80) on a cross-chromosomal test set from in situ and intact Hi-C experiments in heart tissue. Furthermore, we train TwinC using in situ Hi-C data from the widely used GM12878 cell line and validate its performance with orthogonal DNA SPRITE assay in the same cell type. Mechanistically, the neural network learns the importance of compartments, chromatin accessibility, clustered transcription factor binding and G-quadruplexes in forming trans contacts. In summary, TwinC models and interprets trans genome architecture, shedding light on this poorly understood aspect of gene regulation.

Journal Article

An updated compendium and reevaluation of the evidence for nuclear transcription factor occupancy over the mitochondrial genome.

In most eukaryotes, mitochondrial organelles contain their own genome, usually circular, which is the remnant of the genome of the ancestral bacterial endosymbiont that gave rise to modern mitochondria. Mitochondrial genomes are dramatically reduced in their gene content due to the process of endosymbiotic gene transfer to the nucleus; as a result most mitochondrial proteins are encoded in the nucleus and imported into mitochondria. This includes the components of the dedicated mitochondrial transcription and replication systems and regulatory factors, which are entirely distinct from the information processing systems in the nucleus. However, since the 1990s several nuclear transcription factors have been reported to act in mitochondria, and previously we identified 8 human and 3 mouse transcription factors (TFs) with strong localized enrichment over the mitochondrial genome using ChIP-seq (Chromatin Immunoprecipitation) datasets from the second phase of the ENCODE (Encyclopedia of DNA Elements) Project Consortium. Here, we analyze the greatly expanded in the intervening decade ENCODE compendium of TF ChIP-seq datasets (a total of 6,153 ChIP experiments for 942 proteins, of which 763 are sequence-specific TFs) combined with interpretative deep learning models of TF occupancy to create a comprehensive compendium of nuclear TFs that show evidence of association with the mitochondrial genome. We find some evidence for chrM occupancy for 50 nuclear TFs and two other proteins, with bZIP TFs emerging as most likely to be playing a role in mitochondria. However, we also observe that in cases where the same TF has been assayed with multiple antibodies and ChIP protocols, evidence for its chrM occupancy is not always reproducible. In the light of these findings, we discuss the evidential criteria for establishing chrM occupancy and reevaluate the overall compendium of putative mitochondrial-acting nuclear TFs.

Genome, Mitochondrial

CasKAS: direct profiling of genome-wide dCas9 and Cas9 specificity using ssDNA mapping.

Detecting and mitigating off-target activity is critical to the practical application of CRISPR-mediated genome and epigenome editing. While numerous methods have been developed to map Cas9 binding specificity genome-wide, they are generally time-consuming and/or expensive, and not applicable to catalytically dead CRISPR enzymes. We have developed CasKAS, a rapid, inexpensive, and facile assay for identifying off-target CRISPR enzyme binding and cleavage by chemically mapping the unwound single-stranded DNA structures formed upon binding of a sgRNA-loaded Cas9 protein. We demonstrate this method in both in vitro and in vivo contexts.

CRISPR-Cas Systems