Search PubMed⌕ Search

Biomedical subjects

Ramana V Davuluri

Publications and source records attributed to Ramana V Davuluri.

At least 19 recordsLinked to original sources

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings.

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Sequence Analysis, DNA↗

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article↗

DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e., DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 23 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. Model, codes, and data is publicly available at https://github.com/MAGlCS-LAB/DNABERT_S.

Journal Article↗

A mixture model-based discriminate analysis for identifying ordered transcription factor binding site pairs in gene promoters directly regulated by estrogen receptor-alpha.

MOTIVATION: To detect and select patterns of transcription factor binding sites (TFBSs) which distinguish genes directly regulated by estrogen receptor-alpha (ERalpha), we developed an innovative mixture model-based discriminate analysis for identifying ordered TFBS pairs. RESULTS: Biologically, our proposed new algorithm clearly suggests that TFBSs are not randomly distributed within ERalpha target promoters (P-value < 0.001). The up-regulated targets significantly (P-value < 0.01) possess TFBS pairs, (DBP, MYC), (DBP, MYC/MAX heterodimer), (DBP, USF2) and (DBP, MYOGENIN); and down-regulated ERalpha target genes significantly (P-value < 0.01) possess TFBS pairs, such as (DBP, c-ETS1-68), (DBP, USF2) and (DBP, MYOGENIN). Statistically, our proposed mixture model-based discriminate analysis can simultaneously perform TFBS pattern recognition, TFBS pattern selection, and target class prediction; such integrative power cannot be achieved by current methods. AVAILABILITY: The software is available on request from the authors. CONTACT: lali@iupui.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Prognostic DNA methylation biomarkers in ovarian cancer.

PURPOSE: Aberrant DNA methylation, now recognized as a contributing factor to neoplasia, often shows definitive gene/sequence preferences unique to specific cancer types. Correspondingly, distinct combinations of methylated loci can function as biomarkers for numerous clinical correlates of ovarian and other cancers. EXPERIMENTAL DESIGN: We used a microarray approach to identify methylated loci prognostic for reduced progression-free survival (PFS) in advanced ovarian cancer patients. Two data set classification algorithms, Significance Analysis of Microarray and Prediction Analysis of Microarray, successfully identified 220 candidate PFS-discriminatory methylated loci. Of those, 112 were found capable of predicting PFS with 95% accuracy, by Prediction Analysis of Microarray, using an independent set of 40 advanced ovarian tumors (from 20 short-PFS and 20 long-PFS patients, respectively). Additionally, we showed the use of these predictive loci using two bioinformatics machine-learning algorithms, Support Vector Machine and Multilayer Perceptron. CONCLUSION: In this report, we show that highly prognostic DNA methylation biomarkers can be successfully identified and characterized, using previously unused, rigorous classifying algorithms. Such ovarian cancer biomarkers represent a promising approach for the assessment and management of this devastating disease.

Adenocarcinoma↗

Genome-wide analysis of core promoter elements from conserved human and mouse orthologous pairs.

BACKGROUND: The canonical core promoter elements consist of the TATA box, initiator (Inr), downstream core promoter element (DPE), TFIIB recognition element (BRE) and the newly-discovered motif 10 element (MTE). The motifs for these core promoter elements are highly degenerate, which tends to lead to a high false discovery rate when attempting to detect them in promoter sequences. RESULTS: In this study, we have performed the first analysis of these core promoter elements in orthologous mouse and human promoters with experimentally-supported transcription start sites. We have identified these various elements using a combination of positional weight matrices (PWMs) and the degree of conservation of orthologous mouse and human sequences--a procedure that significantly reduces the false positive rate of motif discovery. Our analysis of 9,010 orthologous mouse-human promoter pairs revealed two combinations of three-way synergistic effects, TATA-Inr-MTE and BRE-Inr-MTE. The former has previously been putatively identified in human, but the latter represents a novel synergistic relationship. CONCLUSION: Our results demonstrate that DNA sequence conservation can greatly improve the identification of functional core promoter elements in the human genome. The data also underscores the importance of synergistic occurrence of two or more core promoter elements. Furthermore, the sequence data and results presented here can help build better computational models for predicting the transcription start sites in the promoter regions, which remains one of the most challenging problems.

Animals↗

Combinatorial analysis of transcription factor partners reveals recruitment of c-MYC to estrogen receptor-alpha responsive promoters.

In breast cancer and normal estrogen target tissues, estrogen receptor-alpha (ERalpha) signaling results in the establishment of spatiotemporal patterns of gene expression. Whereas primary target gene regulation by ERalpha involves recruitment of coregulatory proteins, coactivators, or corepressors, activation of these downstream promoters by receptor signaling may also involve partnership of ERalpha with other transcription factors. By using an integrated, genome-wide approach that involves ChIP-chip and computational modeling, we uncovered 13 ERalpha-responsive promoters containing both ERalpha and c-MYC binding elements located within close proximity (13-214 bp) to each other. Estrogen stimulation enhanced the c-MYC-ERalpha interaction and facilitated the association of ERalpha, c-MYC, and the coactivator TRRAP with these estrogen-responsive promoters, resulting in chromatin remodeling and increased transcription. These results suggest that ERalpha and c-MYC physically interact to stabilize the ERalpha-coactivator complex, thereby permitting other signal transduction pathways to fine-tune estrogen-mediated signaling networks.

Animals↗

MPromDb: an integrated resource for annotation and visualization of mammalian gene promoters and ChIP-chip experimental data.

We have developed Mammalian Promoter Database (MPromDb), a novel database that integrates gene promoters with experimentally supported annotation of transcription start sites, cis-regulatory elements, CpG islands and chromatin immunoprecipitation microarray (ChIP-chip) experimental results with intuitively designed presentation. Release 1.0 of MPromDb currently contains 36,407 promoters and first exons (19,170 from human, 15,953 from mouse and 1284 from rat), 3739 transcription factor (TF)-binding sites (2027 from human, 1181 mouse and 531 rat) and 224 TFs with links to PubMed and GenBank references. Target promoters of TFs that have been identified by ChIP-chip assay are integrated into the database. MPromDb serves as a portal for genome-wide promoter analysis of data generated by ChIP-chip experimental studies. MPromDb can be accessed from http://bioinformatics.med.ohio-state.edu/MPromDb.

Animals↗

AGRIS and AtRegNet. a platform to link cis-regulatory elements and transcription factors into regulatory networks.

Gene regulatory pathways converge at the level of transcription, where interactions among regulatory genes and between regulators and target genes result in the establishment of spatiotemporal patterns of gene expression. The growing identification of direct target genes for key transcription factors (TFs) through traditional and high-throughput experimental approaches has facilitated the elucidation of regulatory networks at the genome level. To integrate this information into a Web-based knowledgebase, we have developed the Arabidopsis Gene Regulatory Information Server (AGRIS). AGRIS, which contains all Arabidopsis (Arabidopsis thaliana) promoter sequences, TFs, and their target genes and functions, provides the scientific community with a platform to establish regulatory networks. AGRIS currently houses three linked databases: AtcisDB (Arabidopsis thaliana cis-regulatory database), AtTFDB (Arabidopsis thaliana transcription factor database), and AtRegNet (Arabidopsis thaliana regulatory network). AtTFDB contains 1,690 Arabidopsis TFs and their sequences (protein and DNA) grouped into 50 (October 2005) families with information on available mutants in the corresponding genes. AtcisDB consists of 25,806 (September 2005) promoter sequences of annotated Arabidopsis genes with a description of putative cis-regulatory elements. AtRegNet links, in direct interactions, several hundred genes with the TFs that control their expression. The current release of AtRegNet contains a total of 187 (September 2005) direct targets for 66 TFs. AGRIS can be accessed at http://Arabidopsis.med.ohio-state.edu.

Arabidopsis↗

Changes in gene expression associated with loss of function of the NSDHL sterol dehydrogenase in mouse embryonic fibroblasts.

Seven human disorders of postsqualene cholesterol biosynthesis have been described. One of these, congenital hemidysplasia with ichthyosiform nevus and limb defects (CHILD) syndrome, results from mutations in the X-linked gene NADH sterol dehydrogenase-like (NSDHL) encoding a sterol dehydrogenase. A series of mutant alleles of the murine Nsdhl gene are carried by bare patches (Bpa) mice, with Bpa(1H) representing a null allele. Heterozygous Bpa(1H) females display skin and skeletal abnormalities in a distribution reflecting random X inactivation, whereas hemizygous male embryos die before embryonic day 10.5. To investigate the molecular basis of defects associated with perturbations in cholesterol biosynthesis, microarray analysis was performed comparing gene expression in embryonic fibroblasts expressing the Bpa(1H) allele versus wild-type (wt) cells. Labeled cDNAs from cells grown in normal serum or lipid-depleted serum (LDS) were hybridized to microarrays containing 22,000 mouse genes. Among 44 genes that showed higher expression in the Bpa(1H) versus wt cells grown in LDS, 11 function in cholesterol biosynthesis, 7 are involved in fatty acid synthesis, 3 (Srebp2, Insig1, and Orf11) encode sterol-regulatory proteins, and 2 (Ldlr and StarD4) are lipid transporters. Of the 21 remaining genes, 16 are known genes, some of which have been implicated previously in cholesterol homeostasis or lipid-mediated signaling, and 5 are uncharacterized cDNA clones.

3-Hydroxysteroid Dehydrogenases↗

Identifying estrogen receptor alpha target genes using integrated computational genomics and chromatin immunoprecipitation microarray.

The estrogen receptor alpha (ERalpha) regulates gene expression by either direct binding to estrogen response elements or indirect tethering to other transcription factors on promoter targets. To identify these promoter sequences, we conducted a genome-wide screening with a novel microarray technique called ChIP-on-chip. A set of 70 candidate ERalpha loci were identified and the corresponding promoter sequences were analyzed by statistical pattern recognition and comparative genomics approaches. We found mouse counterparts for 63 of these loci and classified 42 (67%) as direct ERalpha targets using classification and regression tree (CART) statistical model, which involves position weight matrix and human-mouse sequence similarity scores as model parameters. The remaining genes were considered to be indirect targets. To validate this computational prediction, we conducted an additional ChIP-on-chip assay that identified acetylated chromatin components in active ERalpha promoters. Of the 27 loci upregulated in an ERalpha-positive breast cancer cell line, 20 having mouse counterparts were correctly predicted by CART. This integrated approach, therefore, sets a paradigm in which the iterative process of model refinement and experimental verification will continue until an accurate prediction of promoter target sequences is derived.

Animals↗

Loss of estrogen receptor signaling triggers epigenetic silencing of downstream targets in breast cancer.

Alterations in histones, chromatin-related proteins, and DNA methylation contribute to transcriptional silencing in cancer, but the sequence of these molecular events is not well understood. Here we demonstrate that on disruption of estrogen receptor (ER) alpha signaling by small interfering RNA, polycomb repressors and histone deacetylases are recruited to initiate stable repression of the progesterone receptor (PR) gene, a known ERalpha target, in breast cancer cells. The event is accompanied by acquired DNA methylation of the PR promoter, leaving a stable mark that can be inherited by cancer cell progeny. Reestablishing ERalpha signaling alone was not sufficient to reactivate the PR gene; reactivation of the PR gene also requires DNA demethylation. Methylation microarray analysis further showed that progressive DNA methylation occurs in multiple ERalpha targets in breast cancer genomes. The results imply, for the first time, the significance of epigenetic regulation on ERalpha target genes, providing new direction for research in this classical signaling pathway.

Base Sequence↗

OMGProm: a database of orthologous mammalian gene promoters.

SUMMARY: Sequence comparisons between human and rodents are increasingly being used for the identification of gene regulatory regions. The effectiveness of such an approach largely depends on the quality and availability of promoter sequences. We developed OMGProm by integrating three data sources: (1) experimentally supported full-length cDNA, promoter and first exon sequences; (2) homology information from HomoloGene and (3) the human and mouse genomic sequences. The current version of OMGProm contains 8550 promoter pairs of 6373 orthologous human and mouse genes, where supporting experimental evidence for transcription start site annotation exists in at least one species.

Animals↗

Role of cancer-associated stromal fibroblasts in metastatic colon cancer to the liver and their expression profiles.

The cancer microenvironment and interaction between cancer and stromal cells play critical roles in tumor development and progression. The molecular features of cancer stroma are less well understood than those of cancer cells. Cancer-associated stromal fibroblasts are the predominant component of stroma associated with colon cancer and its functions remain unclear. Fibroblast cell cultures were established from metastatic colon cancer in liver, liver away from the metastatic lesions, and skin from three patients with metastatic colorectal cancer. We generated expression profiles of cancer-associated fibroblasts using oligochip arrays and compared them to those of uninvolved fibroblasts. The conditioned media from the cancer-associated fibroblast cultures enhanced proliferation of colon cancer cell line HCT116 to a greater extent than cultures from uninvolved fibroblasts. In microarray expression analysis, cancer-associated fibroblasts clustered tightly into one group and skin fibroblasts into another. Approximately 170 of 22,000 genes were up-regulated in cancer-associated fibroblasts (fold change > 2, P < 0.05) as compared to skin fibroblasts, including many genes encoding cell adhesion molecules, growth factors, and COX2. By immunohistochemistry in-vivo, we confirmed COX2 and TGFB2 expression in cancer-associated fibroblasts in metastatic colon cancer. The distinct molecular expression profiles of cancer-associated fibroblasts in colon cancer metastasis support the notion that these fibroblasts form a favorable microenvironment for cancer cells.

Cell Line, Tumor↗

Papillary and follicular thyroid carcinomas show distinctly different microarray expression profiles and can be distinguished by a minimum of five genes.

PURPOSE: We have previously conducted independent microarray expression analyses of the two most common types of nonmedullary thyroid carcinoma, namely papillary thyroid carcinoma (PTC) and follicular thyroid carcinoma (FTC). In this study, we sought to combine our data sets to shed light on the similarities and differences between these tumor types. MATERIALS AND METHODS: Microarray data from six PTCs, nine FTCs, and 13 normal thyroid samples were normalized to remove interlaboratory variability and then analyzed by unsupervised clustering, t test, and by comparison of absolute and change calls. Expression changes in four genes not previously implicated in thyroid carcinogenesis were verified by reverse transcriptase polymerase chain reaction on these same samples, together with eight additional FTC tumors. RESULTS: PTCs showed two distinct groups of genes that were either over- or underexpressed compared with normal thyroid, whereas the predominant changes in FTCs were of decreased expression. Five genes could collectively distinguish the two tumor types. PTCs showed overexpression of CITED1, claudin-10 (CLDN10), and insulin-like growth factor binding protein 6 (IGFBP6) but showed no change in expression of caveolin-1 (CAV1) or -2 (CAV2); conversely, FTCs did not express CLDN10 and had decreased expression of IGFBP6 and/or CAV1 and CAV2. CONCLUSION: PTC and FTC show distinctive microarray expression profiles, suggesting that either they have different molecular origins or they diverge distinctly from a common origin. Furthermore, if verified in a larger series of tumors, these genes could, in combination with known tumor-specific chromosome translocations, form the basis of a valuable diagnostic tool.

Adenocarcinoma, Follicular↗

Acute myeloid leukemia with complex karyotypes and abnormal chromosome 21: Amplification discloses overexpression of APP, ETS2, and ERG genes.

Molecular mechanisms of leukemogenesis have been successfully unraveled by studying genes involved in simple rearrangements including balanced translocations and inversions. In contrast, little is known about genes altered in complex karyotypic abnormalities. We studied acute myeloid leukemia (AML) patients with complex karyotypes and abnormal chromosome 21. High-resolution bacterial artificial chromosome (BAC) array-based comparative genomic hybridization disclosed amplification predominantly in the 25- to 30-megabase (MB) region that harbors the APP gene (26.3 MB) and at position 38.7-39.1 MB that harbors the transcription factors ERG and ETS2. Using oligonucleotide arrays, APP was by far the most overexpressed gene (mean fold change 19.74, P = 0.0003) compared to a control group of AML with normal cytogenetics; ERG and ETS2 also ranked among the most highly expressed chromosome 21 genes. Overexpression of APP and ETS2 correlated with genomic amplification, but high APP expression occurred even in a subset of AML patients with normal cytogenetics (10 of 64, 16%). APP encodes a glycoprotein of unknown function previously implicated in Alzheimer's disease, but not in AML. We hypothesize that APP and the transcription factors ERG and ETS2 are altered by yet unknown molecular mechanisms involved in leukemogenesis. Our results highlight the value of molecularly dissecting leukemic cells with complex karyotypes.

Acute Disease↗

PML is a direct p53 target that modulates p53 effector functions.

The p53 tumor suppressor promotes cell cycle arrest or apoptosis in response to stress. Previous work suggests that the promyelocytic leukemia gene (PML) can act upstream of p53 to enhance transcription of p53 targets by recruiting p53 to nuclear bodies (NBs). We show that PML is itself a p53 target gene that also acts downstream of p53 to potentiate its antiproliferative effects. Hence, p53 is required for PML induction in response to oncogenes and DNA damaging chemotherapeutics. Furthermore, the PML gene contains p53 binding sites that confer p53 responsiveness to a heterologous reporter and can bind p53 in vitro and in vivo. Finally, cells lacking PML show a reduced propensity to undergo senescence or apoptosis in response to p53 activation, despite the induction of several p53 target genes. These results identify an additional element of PML regulation and establish PML as a mediator of p53 tumor suppressor functions.

Animals↗

Genomic structure and alternative splicing of murine R2B receptor protein tyrosine phosphatases (PTPkappa, mu, rho and PCP-2).

BACKGROUND: Four genes designated as PTPRK (PTPkappa), PTPRL/U (PCP-2), PTPRM (PTPmu) and PTPRT (PTPrho) code for a subfamily (type R2B) of receptor protein tyrosine phosphatases (RPTPs) uniquely characterized by the presence of an N-terminal MAM domain. These transmembrane molecules have been implicated in homophilic cell adhesion. In the human, the PTPRK gene is located on chromosome 6, PTPRL/U on 1, PTPRM on 18 and PTPRT on 20. In the mouse, the four genes ptprk, ptprl, ptprm and ptprt are located in syntenic regions of chromosomes 10, 4, 17 and 2, respectively. RESULTS: The genomic organization of murine R2B RPTP genes is described. The four genes varied greatly in size ranging from approximately 64 kb to approximately 1 Mb, primarily due to proportional differences in intron lengths. Although there were also minor variations in exon length, the number of exons and the phases of exon/intron junctions were highly conserved. In situ hybridization with digoxigenin-labeled cRNA probes was used to localize each of the four R2B transcripts to specific cell types within the murine central nervous system. Phylogenetic analysis of complete sequences indicated that PTPrho and PTPmu were most closely related, followed by PTPkappa. The most distant family member was PCP-2. Alignment of RPTP polypeptide sequences predicted putative alternatively spliced exons. PCR experiments revealed that five of these exons were alternatively spliced, and that each of the four phosphatases incorporated them differently. The greatest variability in genomic organization and the majority of alternatively spliced exons were observed in the juxtamembrane domain, a region critical for the regulation of signal transduction. CONCLUSIONS: Comparison of the four R2B RPTP genes revealed virtually identical principles of genomic organization, despite great disparities in gene size due to variations in intron length. Although subtle differences in exon length were also observed, it is likely that functional differences among these genes arise from the specific combinations of exons generated by alternative splicing.

Alternative Splicing↗