Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptomic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

A comparison of transcriptomic and metabonomic technologies for identifying biomarkers predictive of two-year rodent cancer bioassays.

Two-year rodent bioassays play a central role in evaluating the carcinogenic potential of both commercial products and environmental contaminants. The bioassays are expensive and time consuming, requiring years to complete and costing $2-4 million. In this study, we compare transcriptomic and metabonomic technologies for discovering biomarkers that can efficiently and economically identify chemical carcinogens without performing a standard two-year rodent bioassay. Animals were exposed subchronically to two chemicals (one genotoxic and one nongenotoxic) that were positive for lung and liver tumors in a standard two-year bioassay, two chemicals that were negative, and two control groups. Microarray analysis performed on liver and lung tissues identified multiple biomarkers in each tissue that could discriminate between carcinogenic and noncarcinogenic treatments. The discriminating biomarkers shared a common expression profile among carcinogenic treatments despite different genotoxicity categories and potential modes of action, suggesting that they reflect underlying cellular changes in the transition toward neoplasia. Statistical classification analysis exhibited 100% accuracy in both tissues when the number of genes was less than 5000. Additional genes reduced the predictive accuracy of the model. Serum samples were analyzed by 1H nuclear magnetic resonance (NMR) spectroscopy, and chemical-specific metabolites were removed from the spectra. The statistical classification analysis of the endogenous serum metabolites showed relatively low predictive accuracy with few metabolites in the model, but the accuracy increased to a maximum of 94% when all metabolites were added. These results suggest that individual endogenous metabolites are relatively poor biomarkers, but the metabolite profile as a whole is altered following carcinogen treatment.

Animals↗

Sinorhizobium meliloti differentiation during symbiosis with alfalfa: a transcriptomic dissection.

Sinorhizobium meliloti is a soil bacterium able to induce the formation of nodules on the root of specific legumes, including alfalfa (Medicago sativa). Bacteria colonize nodules through infection threads, invade the plant intracellularly, and ultimately differentiate into bacteroids capable of reducing atmospheric nitrogen to ammonia, which is directly assimilated by the plant. As a first step to describe global changes in gene expression of S. meliloti during the symbiotic process, we used whole genome microarrays to establish the transcriptome profile of bacteria from nodules induced by a bacterial mutant blocked at the infection stage and from wild-type nodules harvested at various timepoints after inoculation. Comparison of these profiles to those of cultured bacteria grown either to log or stationary phase as well as examination of a number of genes with known symbiotic transcription patterns allowed us to correlate global gene-expression patterns to three known steps of symbiotic bacteria bacteroid differentiation, i.e., invading bacteria inside infection threads, young differentiating bacteroids, and fully differentiated, nitrogen-fixing bacteroids. Finally, analysis of individual gene transcription profiles revealed a number of new potential symbiotic genes.

Bacterial Proteins↗

Transcriptome profiling in root nodules and arbuscular mycorrhiza identifies a collection of novel genes induced during Medicago truncatula root endosymbioses.

Transcriptome profiling based on cDNA array hybridizations and in silico screening was used to identify Medicago truncatula genes induced in both root nodules and arbuscular mycorrhiza (AM). By array hybridizations, we detected several hundred genes that were upregulated in the root nodule and the AM symbiosis, respectively, with a total of 75 genes being induced during both interactions. The second approach based on in silico data mining yielded several hundred additional candidate genes with a predicted symbiosis-enhanced expression. A subset of the genes identified by either expression profiling tool was subjected to quantitative real-time reverse-transcription polymerase chain reaction for a verification of their symbiosis-induced expression. That way, induction in root nodules and AM was confirmed for 26 genes, most of them being reported as symbiosis-induced for the first time. In addition to delivering a number of novel symbiosis-induced genes, our approach identified several genes that were induced in only one of the two root endosymbioses. The spatial expression patterns of two symbiosis-induced genes encoding an annexin and a beta-tubulin were characterized in transgenic roots using promoter-reporter gene fusions.

Annexins↗

Uncomplicated human obesity is associated with a specific cardiac transcriptome: involvement of the Wnt pathway.

A dramatic increase in obesity prevalence and cardiovascular morbidity is expected for the coming years. However, with relevance to the heart, little is known about the specific contribution of obesity on associated morbidity. Consequently, global analysis of gene regulations in human heart was undertaken to monitor molecular regulations related to obesity or to obesity-related hypertension. Transcriptome analysis using cDNA arrays was performed in right appendage biopsies from obese patients (n=5), from patients with arterial hypertension with (n=5) or without obesity (n=5), and from 5 leans. All biopsies came from patients that had cardiac surgery and coronary bypass. Statistical analysis of the data revealed 2686 differentially expressed genes out of 11,500 when compared with lean tissues. Differential expression was verified by real-time PCR in 84% of 50 randomly chosen genes. Among genes encountered, 397 were specifically regulated in obese, 1,299 in non-obese hypertensive, and 355 in obese hypertensive patients, respectively, whereas an additional set of 153 genes was differentially expressed in all these situations. Ontology analysis, hierarchical clustering, and molecular pathway analysis indicated that the heart molecular picture of obesity differs clearly from that observed for obesity-related hypertension or arterial hypertension. Clearly, the Wnt pathway known to be involved in cardiac hypertrophy mechanisms, showed opposite regulation in obese heart compared with hypertensive heart and potentially prevented the development of cardiac remodeling in obese patients. All over, this work shows that uncomplicated obesity has a strong impact on cardiac gene expression, which could be considered as precursor signs for future cardiac disease and also demonstrates that obesity-related hypertension generates a heart-molecular-distinct phenotype that cannot be predicted by a simple sum of the impact of obesity and arterial hypertension on gene expression.

Cardiomegaly↗

Transcriptome profiling and the pathogenesis of diabetic complications.

Diabetes is an escalating problem worldwide and a major cause of vascular disease, renal failure, and blindness, among other complications. The cellular mediators of high glucose-induced injury include activation of protein kinase C, accumulation of cell sorbitol from increased flux through the aldose reductase pathway, and generation of advanced glycosylation end products and reactive oxygen species, among others. Current strategies for preventing and slowing the progression of the macrovascular and microvascular complications of diabetes include optimization of glycemic control and BP, angiotensin-converting enzyme inhibitors and angiotensin II blockers, and HMG CoA reductase inhibitors. However, there is an urgent need to develop new therapeutic strategies, as these interventions, although they may slow, rarely halt the progression of diabetic complications. Central to this process is the elucidation of the molecular events that drive this complex disease and that are potential therapeutic targets. This review discusses the promise offered in this regard by global monitoring of cellular or tissue mRNA expression (so-called transcriptomics) and illustrates the potential of this approach by focusing on recent studies on the pathogenesis of diabetic nephropathy.

Actins↗

The role of initial trauma in the host's response to injury and hemorrhage: insights from a correlation of mathematical simulations and hepatic transcriptomic analysis.

Trauma and hemorrhagic shock (HS) elicit severe physiological disturbances that predispose the victims to subsequent organ dysfunction and death. The general lack of effective therapeutic options for these patients is mainly due to the complex interplay of interacting inflammatory and physiological elements working at multiple levels. Systems biology has emerged as a new paradigm that allows the study of large portions of physiological networks simultaneously. Seeking a better understanding of the interplay among known inflammatory pathways, we constructed a mathematical model encompassing the dynamics of the acute inflammatory response that incorporates the intertwined effects of inflammation and global tissue damage. The model was calibrated using data from C57Bl/6 mice subjected to endotoxemia, sham operation (i.e., surgical trauma induced by cannulation [ST]) or ST + HS+ resuscitation (ST-HS-R). An in silico simulation, made at whole-organism level, suggested that similar pathways of different magnitudes were operant as the degree of total body damage increased. We sought to validate this hypothesis by subjecting mice to HS and comparing the models predictions to circulating markers of inflammation and tissue injury as well as the global transcriptomic response of the liver. C57Bl/6 mice were subjected to ST or ST-HS (without resuscitation). Liver gene expression was assessed using an Affymetrix DNA microarray (GeneChip Mouse Expression Set 430A, Affymetrix, Santa Clara, CA), which contains 22,621 probe sets and effectively interrogates 12,341 mouse genes. The microarray data sets were subjected to hierarchical clustering and pathway analysis. In agreement with model predictions, circulating levels of inflammation/tissue injury markers and the microarray analysis both demonstrated that ST alone accounts for a substantial proportion of the observed phenotypic and genetic/molecular changes versus untreated animals. The addition of HS further increased the magnitude of gene expression, but relatively few additional genes were recruited. Mathematical simulations and DNA microarrays, both systems biology tools, may provide valuable insight into the complex global physiological interactions that occur in response to trauma and hemorrhagic shock.

Animals↗

Tertiary lymphoid structure transcriptomic signatures show limited and cohort-dependent value for predicting axillary nodal involvement in oestrogen receptor-positive luminal breast cancer.

Tertiary lymphoid structures (TLS) are associated with prognosis in solid tumours. Their value for predicting axillary nodal involvement in oestrogen receptor-positive luminal breast cancer remains uncertain. Three published TLS signatures were scored by single-sample gene set enrichment analysis in oestrogen receptor-positive luminal tumours. The Cancer Genome Atlas Breast Invasive Carcinoma cohort (TCGA-BRCA) included 632 cases, of which 379 met strict consensus. METABRIC included 1086 cases, of which 663 met strict consensus. Logistic models adjusted for age and pathological tumour stage. Strict consensus, majority vote, and continuous scores were compared. Performance assessment included bootstrapped changes in area under the receiver-operating-characteristic curve, Brier scores, calibration, and decision-curve analysis. Survival was evaluated in METABRIC and explored in TCGA-BRCA. Strict-consensus TLS status was not associated with nodal positivity in TCGA-BRCA (adjusted odds ratio: 0.95, 95% confidence interval: 0.62-1.45, P = 0.822). METABRIC was similar (odds ratio: 0.76, 95% confidence interval: 0.55-1.06, P = 0.105). Full-cohort METABRIC analyses detected small majority-vote and continuous-score associations, absent in TCGA-BRCA. Across specifications, bootstrapped changes in area under the receiver-operating-characteristic curve ranged from 0.0002 to 0.0089, with minimal Brier-score improvement and no stable decision-curve benefit. In METABRIC, the univariable overall survival association attenuated after age adjustment (hazard ratio: 1.33-1.10). TCGA-BRCA survival analyses were nonsignificant. TLS transcriptomic signals showed small, cohort-dependent associations with nodal status but no reproducible or clinically meaningful incremental predictive value. These data do not support replacing sentinel lymph node biopsy with a TLS signature in oestrogen receptor-positive luminal breast cancer.

breast cancer↗

Transcriptome and proteome analysis of Bacillus subtilis gene expression in response to superoxide and peroxide stress.

The Gram-positive soil bacterium Bacillus subtilis responds to oxidative stress by the activation of different cellular defence mechanisms. These are composed of scavenging enzymes as well as protection and repair systems organized in highly sophisticated networks. In this study, the peroxide and the superoxide stress stimulons of B. subtilis were characterized by means of transcriptomics and proteomics. The results demonstrate that oxidative-stress-responsive genes can be classified into two groups. One group encompasses genes which show similar expression patterns in the presence of both reactive oxygen species. Examples are members of the PerR and the Fur regulon which were induced by peroxide and superoxide stress. Similarly, both kinds of stress stimulated the activation of the stringent response. The second group is composed of genes primarily responding to one stimulus, like the members of the SOS regulon which were particularly upregulated in the presence of peroxide, and many genes involved in sulfate assimilation and methionine biosynthesis which were only induced by superoxide. Several genes encoding proteins of unknown function could be assigned to one of these groups.

Bacillus subtilis↗

Diverse roles for HspR in Campylobacter jejuni revealed by the proteome, transcriptome and phenotypic characterization of an hspR mutant.

Campylobacter jejuni is a leading cause of bacterial gastroenteritis in the developed world. The role of a homologue of the negative transcriptional regulatory protein HspR, which in other organisms participates in the control of the heat-shock response, was investigated. Following inactivation of hspR in C. jejuni, members of the HspR regulon were identified by DNA microarray transcript profiling. In agreement with the predicted role of HspR as a negative regulator of genes involved in the heat-shock response, it was observed that the transcript amounts of 13 genes were increased in the hspR mutant, including the chaperone genes dnaK, grpE and clpB, and a gene encoding the heat-shock regulator HrcA. Proteomic analysis also revealed increased synthesis of the heat-shock proteins DnaK, GrpE, GroEL and GroES in the absence of HspR. The altered expression of chaperones was accompanied by heat sensitivity, as the hspR mutant was unable to form colonies at 44 degrees C. Surprisingly, transcriptome analysis also revealed a group of 17 genes with lower transcript levels in the hspR mutant. Of these, eight were predicted to be involved in the formation of the flagella apparatus, and the decreased expression is likely to be responsible for the reduced motility and ability to autoagglutinate that was observed for hspR mutant cells. Electron micrographs showed that mutant cells were spiral-shaped and carried intact flagella, but were elongated compared to wild-type cells. The inactivation of hspR also reduced the ability of Campylobacter to adhere to and invade human epithelial INT-407 cells in vitro, possibly as a consequence of the reduced motility or lower expression of the flagellar export apparatus in hspR mutant cells. It was concluded that, in C. jejuni, HspR influences the expression of several genes that are likely to have an impact on the ability of the bacterium to successfully survive in food products and subsequently infect the consumer.

Bacterial Proteins↗

Influence of the regulatory protein RsmA on cellular functions in Pseudomonas aeruginosa PAO1, as revealed by transcriptome analysis.

RsmA is a posttranscriptional regulatory protein in Pseudomonas aeruginosa that works in tandem with a small non-coding regulatory RNA molecule, RsmB (RsmZ), to regulate the expression of several virulence-related genes, including the N-acyl-homoserine lactone synthase genes lasI and rhlI, and the hydrogen cyanide and rhamnolipid biosynthetic operons. Although these targets of direct RsmA regulation have been identified, the full impact of RsmA on cellular activities is not as yet understood. To address this issue the transcriptome profiles of P. aeruginosa PAO1 and an isogenic rsmA mutant were compared. Loss of RsmA altered the expression of genes involved in a variety of pathways and systems important for virulence, including iron acquisition, biosynthesis of the Pseudomonas quinolone signal (PQS), the formation of multidrug efflux pumps, and motility. Not all of these effects can be explained through the established regulatory roles of RsmA. This study thus provides both a first step towards the identification of further genes under RsmA posttranscriptional control in P. aeruginosa and a fuller understanding of the broader impact of RsmA on cellular functions.

4-Butyrolactone↗

Adaptation of Bacillus subtilis to growth at low temperature: a combined transcriptomic and proteomic appraisal.

The soil bacterium Bacillus subtilis frequently encounters a reduction in temperature in its natural habitats. Here, a combined transcriptomic and proteomic approach has been used to analyse the adaptational responses of B. subtilis to low temperature. Propagation of B. subtilis in minimal medium at 15 degrees C triggered the induction of 279 genes and the repression of 301 genes in comparison to cells grown at 37 degrees C. The analysis thus revealed profound adjustments in the overall gene expression profile in chill-adapted cells. Important transcriptional changes in low-temperature-grown cells comprise the induction of the SigB-controlled general stress regulon, the induction of parts of the early sporulation regulons (SigF, SigE and SigG) and the induction of a regulatory circuit (RapA/PhrA and Opp) that is involved in the fine-tuning of the phosphorylation status of the Spo0A response regulator. The analysis of chill-stress-repressed genes revealed reductions in major catabolic (glycolysis, oxidative phosphorylation, ATP synthesis) and anabolic routes (biosynthesis of purines, pyrimidines, haem and fatty acids) that likely reflect the slower growth rates at low temperature. Low-temperature repression of part of the SigW regulon and of many genes with predicted functions in chemotaxis and motility was also noted. The proteome analysis of chill-adapted cells indicates a major contribution of post-transcriptional regulation phenomena in adaptation to low temperature. Comparative analysis of the previously reported transcriptional responses of cold-shocked B. subtilis cells with this data revealed that cold shock and growth in the cold constitute physiologically distinct phases of the adaptation of B. subtilis to low temperature.

Adaptation, Physiological↗

Transcriptomic and proteomic analyses of the pMOL30-encoded copper resistance in Cupriavidus metallidurans strain CH34.

The four replicons of Cupriavidus metallidurans CH34 (the genome sequence was provided by the US Department of Energy-University of California Joint Genome Institute) contain two gene clusters putatively encoding periplasmic resistance to copper, with an arrangement of genes resembling that of the copSRABCD locus on the 2.1 Mb megaplasmid (MPL) of Ralstonia solanacearum, a closely related plant pathogen. One of the copSRABCD clusters was located on the 2.6 Mb MPL, while the second was found on the pMOL30 (234 kb) plasmid as part of a larger group of genes involved in copper resistance, spanning 17 857 bp in total. In this region, 19 ORFs (copVTMKNSRABCDIJGFLQHE) were identified based on the sequencing of a fragment cloned in an IncW vector, on the preliminary annotation by the Joint Genome Institute, and by using transcriptomic and proteomic data. When introduced into plasmid-cured derivatives of C. metallidurans CH34, the cop locus was able to restore the wild-type MIC, albeit with a biphasic survival curve, with respect to applied Cu(II) concentration. Quantitative-PCR data showed that the 19 ORFs were induced from 2- to 1159-fold when cells were challenged with elevated Cu(II) concentrations. Microarray data showed that the genes that were most induced after a Cu(II) challenge of 0.1 mM belonged to the pMOL30 cop cluster. Megaplasmidic cop genes were also induced, but at a much lower level, with the exception of the highly expressed MPL copD. Proteomic data allowed direct observation on two-dimensional gel electrophoresis, and via mass spectrometry, of pMOL30 CopK, CopR, CopS, CopA, CopB and CopC proteins. Individual cop gene expression depended on both the Cu(II) concentration and the exposure time, suggesting a sequential scheme in the resistance process, involving genes such as copK and copT in an initial phase, while other genes, such as copH, seem to be involved in a late response phase. A concentration of 0.4 mM Cu(II) was the highest to induce maximal expression of most cop genes.

Amino Acid Sequence↗

The arginine regulon of Escherichia coli: whole-system transcriptome analysis discovers new genes and provides an integrated view of arginine regulation.

Analysis of the response to arginine of the Escherichia coli K-12 transcriptome by microarray hybridization and real-time quantitative PCR provides the first coherent quantitative picture of the ArgR-mediated repression of arginine biosynthesis and uptake genes. Transcriptional repression was shown to be the major control mechanism of the biosynthetic genes, leaving only limited room for additional transcriptional or post-transcriptional regulation. The art genes, encoding the specific arginine uptake system, are subject to ArgR-mediated repression, with strong repression of artJ, encoding the periplasmic binding protein of the system. The hisJQMP genes of the histidine transporter (part of the lysine-arginine-ornithine uptake system) were discovered to be a part of the arginine regulon. Analysis of their control region with reporter gene fusions and electrophoretic mobility shift in the presence of pure ArgR repressor showed the involvement in repression of the ArgR protein and an ARG box 120 bp upstream of hisJ. No repression of the genes of the third uptake system, arginine-ornithine, was observed. Finally, comparison of the time course of arginine repression of gene transcription with the evolution of the specific activities of the cognate enzymes showed that while full genetic repression was achieved 2 min after arginine addition, enzyme concentrations were diluted at the rate of cell division. This emphasizes the importance of feedback inhibition of the first enzymic step in the pathway in controlling the metabolic flow through biosynthesis in the period following the onset of repression.

Amino Acid Transport Systems, Basic↗

Transcriptome profile of murine gammaherpesvirus-68 lytic infection.

The murine gammaherpesvirus-68 genome encodes 73 protein-coding open reading frames with extensive similarities to human gamma(2) herpesviruses, as well as unique genes and cellular homologues. We performed transcriptome analysis of stage-specific viral RNA during permissive infection using an oligonucleotide-based microarray. Using this approach, M4, K3, ORF38, ORF50, ORF57 and ORF73 were designated as immediate-early genes based on cycloheximide treatment. The microarray analysis also identified 10 transcripts with early expression kinetics, 32 transcripts with early-late expression kinetics and 29 transcripts with late expression kinetics. The latter group consisted mainly of structural proteins, and showed high expression levels relative to other viral transcripts. Moreover, we detected all eight tRNA-like transcripts in the presence of cycloheximide and phosphonoacetic acid. Lytic infection with MHV-68 also resulted in a significant reduction in the expression of cellular transcripts included in the DNA chip. This global approach to viral transcript analysis offers a powerful system for examining molecular transitions between lytic and latent virus infections associated with disease pathogenesis.

Animals↗

Transcriptomal analysis of varicella-zoster virus infection using long oligonucleotide-based microarrays.

Varicella-zoster virus (VZV) is a human herpes virus that causes varicella as a primary infection and herpes zoster following reactivation of the virus from a latent state in trigeminal and spinal ganglia. In order to study the global pattern of VZV gene transcription, VZV microarrays using 75-base oligomers to 71 VZV open reading frames (ORFs) were designed and validated. The long-oligonucleotide approach maximizes the stringency of detection and polarity of gene expression. To optimize sensitivity, microarrays were hybridized to target RNA and the extent of hybridization measured using resonance light scattering. Microarray data were normalized to a subset of invariant ranked host-encoded positive-control genes and the data subjected to robust formal statistical analysis. The programme of viral gene expression was determined for VZV (Dumas strain)-infected MeWo cells and SVG cells (an immortalized human astrocyte cell line) 72 h post-infection. Marked quantitative and qualitative differences in the viral transcriptome were observed between the two different cell types using the Dumas laboratory-adapted strain. Oligonucleotide-based VZV arrays have considerable promise as a valuable tool in the analysis of viral gene transcription during both lytic and latent infections, and the observed heterogeneity in the global pattern of viral gene transcription may also have diagnostic potential.

Animals↗

Single-cell transcriptomic atlas of Alzheimer's disease middle temporal gyrus reveals region, cell type and sex specificity of gene expression with novel genetic risk for MERTK in female.

Alzheimer's disease, the most common age-related neurodegenerative disease, is closely associated with both amyloid-ß plaque and neuroinflammation. Two thirds of Alzheimer's disease patients are females and they have a higher disease risk. Moreover, women with Alzheimer's disease have more extensive brain histological changes than men along with more severe cognitive symptoms and neurodegeneration. To identify how sex difference induces structural brain changes, we performed unbiased massively parallel single nucleus RNA sequencing on Alzheimer's disease and control brains focusing on the middle temporal gyrus, a brain region strongly affected by the disease but not previously studied with these methods. We identified a subpopulation of selectively vulnerable layer 2/3 excitatory neurons that that were RORB-negative and CDH9-expressing. This vulnerability differs from that reported for other brain regions, but there was no detectable difference between male and female patterns in middle temporal gyrus samples. Disease-associated, but sex-independent, reactive astrocyte signatures were also present. In clear contrast, the microglia signatures of diseased brains differed between males and females. Combining single cell transcriptomic data with results from genome-wide association studies (GWAS), we identified MERTK genetic variation as a risk factor for Alzheimer's disease selectively in females. Taken together, our single cell dataset revealed a unique cellular-level view of sex-specific transcriptional changes in Alzheimer's disease, illuminating GWAS identification of sex-specific Alzheimer's risk genes. These data serve as a rich resource for interrogation of the molecular and cellular basis of Alzheimer's disease.

Journal Article↗

Inferring Metabolic States from Single Cell Transcriptomic Data via Geometric Deep Learning.

The ability to measure gene expression at single-cell resolution has elevated our understanding of how biological features emerge from complex and interdependent networks at molecular, cellular, and tissue scales. As technologies have evolved that complement scRNAseq measurements with things like single-cell proteomic, epigenomic, and genomic information, it becomes increasingly apparent how much biology exists as a product of multimodal regulation. Biological processes such as transcription, translation, and post-translational or epigenetic modification impose both energetic and specific molecular demands on a cell and are therefore implicitly constrained by the metabolic state of the cell. While metabolomics is crucial for defining a holistic model of any biological process, the chemical heterogeneity of the metabolome makes it particularly difficult to measure, and technologies capable of doing this at single-cell resolution are far behind other multiomics modalities. To address these challenges, we present GEFMAP (Gene Expression-based Flux Mapping and Metabolic Pathway Prediction), a method based on geometric deep learning for predicting flux through reactions in a global metabolic network using transcriptomics data, which we ultimately apply to scRNAseq. GEFMAP leverages the natural graph structure of metabolic networks to learn both a biological objective for each cell and estimate a mass-balanced relative flux rate for each reaction in each cell using novel deep learning models.

Preprint↗

Transcriptome-Wide Root Causal Inference.

Root causal genes correspond to the first gene expression levels perturbed during pathogenesis by genetic or non-genetic factors. Targeting root causal genes has the potential to alleviate disease entirely by eliminating pathology near its onset. No existing algorithm discovers root causal genes from observational data alone. We therefore propose the Transcriptome-Wide Root Causal Inference (TWRCI) algorithm that identifies root causal genes and their causal graph using a combination of genetic variant and unperturbed bulk RNA sequencing data. TWRCI uses a novel competitive regression procedure to annotate cis and trans-genetic variants to the gene expression levels they directly cause. The algorithm simultaneously recovers a causal ordering of the expression levels to pinpoint the underlying causal graph and estimate root causal effects. TWRCI outperforms alternative approaches across a diverse group of metrics by directly targeting root causal genes while accounting for distal relations, linkage disequilibrium, patient heterogeneity and widespread pleiotropy. We demonstrate the algorithm by uncovering the root causal mechanisms of two complex diseases, which we confirm by replication using independent genome-wide summary statistics.

Journal Article↗