Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptomic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

GSK3B inhibition partially reverses brain ethanol-induced transcriptomic changes in C57BL/6J mice: Expression network co-analysis with human genome-wide association studies.

Alcohol use disorder (AUD) is a chronic behavioral disease with greater than 50% of its risk due to complex genetic contributions. Existing pharmacological and behavioral treatments for AUD are minimally effective and underutilized. Animal model behavioral genetics and human genome-wide association studies have begun to identify individual genes contributing to the progressive compulsive consumption of ethanol that occurs with AUD, promising possible new therapeutic targets. Our laboratory has previously identified Gsk3b as a central member in a network of ethanol-responsive genes in mouse prefrontal cortex, which altered ethanol consumption with genetic manipulation and was also significantly associated with risk for alcohol dependence in human genome-wide association studies. Here we perform detailed brain RNA sequencing transcriptomic studies to characterize a highly specific and clinically available GSK3B pharmacological inhibitor, tideglusib, as a possible therapeutic for clinical trials on treatment of AUD. A model of chronic intermittent ethanol consumption was used to study gene expression changes in prefrontal cortex and nucleus accumbens in the presence or absence of tideglusib treatment. Multivariate analysis of differentially expressed genes showed that tideglusib largely reversed ethanol- induced expression changes for two prominent clusters of genes in both prefrontal cortex and nucleus accumbens. Bioinformatic analysis showed these genes to have prominent roles in neuronal functioning and synaptic activity. Additionally, mouse brain differential gene expression data was analyzed together with human protein-protein interaction and genome-wide association studies on AUD to derive networks responding to tideglusib and relevant to human genetic risk for alcohol dependence. These studies identified discrete networks significantly enriched with genes provisionally associated with AUD, and provide key information on central hubs of such networks. Together these studies document tideglusib as a major modulator of chronic ethanol consumption-evoked brain gene expression signatures, and identify possible new targets for therapeutic modulation of AUD.

Journal Article↗

Multi-ancestry genome-wide and transcriptome-wide association analyses identified new risk loci and genes for inflammatory bowel disease.

To advance genetic understanding of inflammatory bowel disease (IBD), we conducted genome-wide association meta-analyses of 63,415 IBD cases of European and East Asian descendants and identified 90 previously unknown risk loci. Integrating multi-ancestry transcriptome-wide association studies (TWAS), cell type-specific TWAS, alternative splicing (AS-WAS), and alternative polyadenylation (APA-WAS) analyses using RNA-seq data from normal colon tissues of 707 European and 364 East Asian individuals, we uncovered 506 high-confidence IBD risk genes, including 384 not previously reported. These genes converge on immune regulation, microbial interaction, and other pathways central to IBD pathogenesis, with over half showing transcriptional dysregulation supported by single-cell and spatial omics analyses. Notably, 46 risk genes are targeted by 225 drugs that have been approved or in Phase II/III trials, including sulfasalazine already used in IBD therapy. Our study findings deepen the understanding of IBD genetics and support the development of precision medicine for its prevention and treatment.

GWAS↗

Identification and analysis of chromodomain-containing proteins encoded in the mouse transcriptome.

The chromodomain is 40-50 amino acids in length and is conserved in a wide range of chromatic and regulatory proteins involved in chromatin remodeling. Chromodomain-containing proteins can be classified into families based on their broader characteristics, in particular the presence of other types of domains, and which correlate with different subclasses of the chromodomains themselves. Hidden Markov model (HMM)-generated profiles of different subclasses of chromodomains were used here to identify sequences encoding chromodomain-containing proteins in the mouse transcriptome and genome. A total of 36 different loci encoding proteins containing chromodomains, including 17 novel loci, were identified. Six of these loci (including three apparent pseudogenes, a novel HP1 ortholog, and two novel Msl-3 transcription factor-like proteins) are not present in the human genome, whereas the human genome contains four loci (two CDY orthologs and two apparent CDY pseudogenes) that are not present in mouse. A number of these loci exhibit alternative splicing to produce different isoforms, including 43 novel variants, some of which lack the chromodomain. The likely functions of these proteins are discussed in relation to the known functions of other chromodomain-containing proteins within the same family.

Acetyltransferases↗

Systematic expression profiling of the mouse transcriptome using RIKEN cDNA microarrays.

The number of known mRNA transcripts in the mouse has been greatly expanded by the RIKEN Mouse Gene Encyclopedia project. Validation of their reproducible expression in a tissue is an important contribution to the study of functional genomics. In this report, we determine the expression profile of 57,931 clones on 20 mouse tissues using cDNA microarrays. Of these 57,931 clones, 22,928 clones correspond to the FANTOM2 clone set. The set represents 20,234 transcriptional units (TUs) out of 33,409 TUs in the FANTOM2 set. We identified 7206 separate clones that satisfied stringent criteria for tissue-specific expression. Gene Ontology terms were assigned for these 7206 clones, and the proportion of 'molecular function' ontology for each tissue-specific clone was examined. These data will provide insights into the function of each tissue. Tissue-specific gene expression profiles obtained using our cDNA microarrays were also compared with the data extracted from the GNF Expression Atlas based on Affymetrix microarrays. One major outcome of the RIKEN transcriptome analysis is the identification of numerous nonprotein-coding mRNAs. The expression profile was also used to obtain evidence of expression for putative noncoding RNAs. In addition, 1926 clones (70%) of 2768 clones that were categorized as "unknown EST," and 1969 (58%) clones of 3388 clones that were categorized as "unclassifiable" were also shown to be reproducibly expressed.

Animals↗

Quantitative assessment of transcriptome differences between brain territories.

Transcriptome analysis of mammalian brain structures is a potentially powerful approach in addressing the diversity of cerebral functions. Here, we used a microassay for serial analysis of gene expression (SAGE) to generate quantitative mRNA expression profiles of normal adult mouse striatum, nucleus accumbens, and somatosensory cortex. Comparison of these profiles revealed 135 transcripts heterogeneously distributed in the brain. Among them, a majority (78), although matching a registered sequence, are novel regional markers. To improve the anatomical resolution of our analysis, we performed in situ hybridization and observed unique expression patterns in discrete brain regions for a number of candidates. We assessed the distribution of the new markers in peripheral tissues using quantitative RT-PCR, Northern hybridization, and published SAGE data. In most cases, expression was higher in the brain than in peripheral tissues. Because the markers were selected according to their expression level, without reference to prior knowledge, our studies provide an unbiased, comprehensive molecular signature for various mammalian brain structures that can be used to investigate their plasticity under a variety of circumstances.

Animals↗

The human transcriptome map reveals extremes in gene density, intron length, GC content, and repeat pattern for domains of highly and weakly expressed genes.

The chromosomal gene expression profiles established by the Human Transcriptome Map (HTM) revealed a clustering of highly expressed genes in about 30 domains, called ridges. To physically characterize ridges, we constructed a new HTM based on the draft human genome sequence (HTMseq). Expression of 25,003 genes can be analyzed online in a multitude of tissues (http://bioinfo.amc.uva.nl/HTMseq). Ridges are found to be very gene-dense domains with a high GC content, a high SINE repeat density, and a low LINE repeat density. Genes in ridges have significantly shorter introns than genes outside of ridges. The HTMseq also identifies a significant clustering of weakly expressed genes in domains with fully opposite characteristics (antiridges). Both types of domains are open to tissue-specific expression regulation, but the maximal expression levels in ridges are considerably higher than in antiridges. Ridges are therefore an integral part of a higher order structure in the genome related to transcriptional regulation.

Base Composition↗

Novel RNAs identified from an in-depth analysis of the transcriptome of human chromosomes 21 and 22.

In this report, we have achieved a richer view of the transcriptome for Chromosomes 21 and 22 by using high-density oligonucleotide arrays on cytosolic poly(A)(+) RNA. Conservatively, only 31.4% of the observed transcribed nucleotides correspond to well-annotated genes, whereas an additional 4.8% and 14.7% correspond to mRNAs and ESTs, respectively. Approximately 85% of the known exons were detected, and up to 21% of known genes have only a single isoform based on exon-skipping alternative expression. Overall, the expression of the well-characterized exons falls predominately into two categories, uniquely or ubiquitously expressed with an identifiable proportion of antisense transcripts. The remaining observed transcription (49.0%) was outside of any known annotation. These novel transcripts appear to be more cell-line-specific and have lower and less variation in expression than the well-characterized genes. Novel transcripts were further characterized based on their distance to annotations, transcript size, coding capacity, and identification as antisense to intronic sequences. By RT-PCR, 126 novel transcripts were independently verified, resulting in a 65% verification rate. These observations strongly support the argument for a re-evaluation of the total number of human genes and an alternative term for "gene" to encompass these growing, novel classes of RNA transcripts in the human genome.

Cell Line↗

Dual-contrastive learning for spatial domain identification in spatial transcriptomics with STAMGC.

Spatial transcriptomics (STs) have become a valuable approach for understanding the growth and development of organisms. Despite the recent emergence of numerous ST models, accurately identifying spatial domains remains challenging owing to the trade-off between preserving local details and reducing noise. Here, we introduce STAMGC, which is a dual-contrastive learning framework built upon graph convolutional networks. This model leverages regional and topological contrastive learning to jointly optimize the model, effectively reducing the noise in spatial domain identification and enhancing the extraction of detailed features. In this study, Gaussian smoothing, originally developed in the image processing field, is introduced to process ST data, providing a foundation for region-level contrastive learning by mitigating spatial discontinuities of gene expression signals. Experimental results indicate that STAMGC outperforms existing methods across multiple data sets according to comprehensive evaluations. Furthermore, STAMGC not only identifies finer structures in the mouse brain but also brings new discoveries for human breast cancer research.

Journal Article↗

Widespread RNA editing of embedded alu elements in the human transcriptome.

More than one million copies of the approximately 300-bp Alu element are interspersed throughout the human genome, with up to 75% of all known genes having Alu insertions within their introns and/or UTRs. Transcribed Alu sequences can alter splicing patterns by generating new exons, but other impacts of intragenic Alu elements on their host RNA are largely unexplored. Recently, repeat elements present in the introns or 3'-UTRs of 15 human brain RNAs have been shown to be targets for multiple adenosine to inosine (A-to-I) editing. Using a statistical approach, we find that editing of transcripts with embedded Alu sequences is a global phenomenon in the human transcriptome, observed in 2674 ( approximately 2%) of all publicly available full-length human cDNAs (n = 128,406), from >250 libraries and >30 tissue sources. In the vast majority of edited RNAs, A-to-I substitutions are clustered within transcribed sense or antisense Alu sequences. Edited bases are primarily associated with retained introns, extended UTRs, or with transcripts that have no corresponding known gene. Therefore, Alu-associated RNA editing may be a mechanism for marking nonstandard transcripts, not destined for translation.

Alternative Splicing↗

Organization of the Caenorhabditis elegans small non-coding transcriptome: genomic features, biogenesis, and expression.

Recent evidence points to considerable transcription occurring in non-protein-coding regions of eukaryote genomes. However, their lack of conservation and demonstrated function have created controversy over whether these transcripts are functional. Applying a novel cloning strategy, we have cloned 100 novel and 61 known or predicted Caenorhabditis elegans full-length ncRNAs. Studying the genomic environment and transcriptional characteristics have shown that two-thirds of all ncRNAs, including many intronic snoRNAs, are independently transcribed under the control of ncRNA-specific upstream promoter elements. Furthermore, the transcription levels of at least 60% of the ncRNAs vary with developmental stages. We identified two new classes of ncRNAs, stem-bulge RNAs (sbRNAs) and snRNA-like RNAs (snlRNAs), both featuring distinct internal motifs, secondary structures, upstream elements, and high and developmentally variable expression. Most of the novel ncRNAs are conserved in Caenorhabditis briggsae, but only one homolog was found outside the nematodes. Preliminary estimates indicate that the C. elegans transcriptome contains approximately 2700 small non-coding RNAs, potentially acting as regulatory elements in nematode development.

Animals↗

Decoding the fine-scale structure of a breast cancer genome and transcriptome.

A comprehensive understanding of cancer is predicated upon knowledge of the structure of malignant genomes underlying its many variant forms and the molecular mechanisms giving rise to them. It is well established that solid tumor genomes accumulate a large number of genome rearrangements during tumorigenesis. End Sequence Profiling (ESP) maps and clones genome breakpoints associated with all types of genome rearrangements elucidating the structural organization of tumor genomes. Here we extend the ESP methodology in several directions using the breast cancer cell line MCF-7. First, targeted ESP is applied to multiple amplified loci, revealing a complex process of rearrangement and co-amplification in these regions reminiscent of breakage/fusion/bridge cycles. Second, genome breakpoints identified by ESP are confirmed using a combination of DNA sequencing and PCR. Third, in vitro functional studies assign biological function to a rearranged tumor BAC clone, demonstrating that it encodes anti-apoptotic activity. Finally, ESP is extended to the transcriptome identifying four novel fusion transcripts and providing evidence that expression of fusion genes may be common in tumors. These results demonstrate the distinct advantages of ESP including: (1) the ability to detect all types of rearrangements and copy number changes; (2) straightforward integration of ESP data with the annotated genome sequence; (3) immortalization of the genome; (4) ability to generate tumor-specific reagents for in vitro and in vivo functional studies. Given these properties, ESP could play an important role in a tumor genome project.

Breast Neoplasms↗

Systematic characterization of the zinc-finger-containing proteins in the mouse transcriptome.

Zinc-finger-containing proteins can be classified into evolutionary and functionally divergent protein families that share one or more domains in which a zinc ion is tetrahedrally coordinated by cysteines and histidines. The zinc finger domain defines one of the largest protein superfamilies in mammalian genomes;46 different conserved zinc finger domains are listed in InterPro (http://www.ebi.ac.uk/InterPro). Zinc finger proteins can bind to DNA, RNA, other proteins, or lipids as a modular domain in combination with other conserved structures. Owing to this combinatorial diversity, different members of zinc finger superfamilies contribute to many distinct cellular processes, including transcriptional regulation, mRNA stability and processing, and protein turnover. Accordingly, mutations of zinc finger genes lead to aberrations in a broad spectrum of biological processes such as development, differentiation, apoptosis, and immunological responses. This study provides the first comprehensive classification of zinc finger proteins in a mammalian transcriptome. Specific detailed analysis of the SP/Krüppel-like factors and the E3 ubiquitin-ligase RING-H2 families illustrates the importance of such an analysis for a more comprehensive functional classification of large protein families. We describe the characterization of a new family of C2H2 zinc-finger-containing proteins and a new conserved domain characteristic of this family, the identification and characterization of Sp8, a new member of the Sp family of transcriptional regulators, and the identification of five new RING-H2 proteins.

Alternative Splicing↗

Comprehensive analysis of the mouse metabolome based on the transcriptome.

The complete set of cDNAs encoding the enzymes of known metabolic pathways has not previously been available for any mammal. Here, transcripts encoding the metabolic pathways of the mouse (mouse metabolome) were reconstructed by making use of the KEGG metabolic pathway database and gene ontology (GO) assignment to the mouse representative transcript and protein set (RTPS), which contains all available mouse transcript sequences including the FANTOM set of RIKEN mouse cDNA clones. By assigning EC numbers extracted from the molecular function ontology in GO, the known mouse transcriptome was predicted to encode enzymes with 726 unique EC numbers. Of these, 648 EC numbers were newly assigned based on the FANTOM set. The mouse metabolome confirmed by cDNA analysis includes almost all of the enzymes of well known pathways such as the tricarboxylic acid cycle and urea cycle. On the other hand, analysis of enzymes required for the tryptophan metabolism pathway revealed a lack of connectivity, indicating that cDNAs/genes encoding several key enzymes remain to be identified. The information derived from coexpression from the cDNA microarray analysis of enzymes of known function may lead to identification of the missing components of the metabolome, and will add new insights into the connectivity of the mammalian metabolic pathways.

Amino Acids↗

Kinesin superfamily proteins (KIFs) in the mouse transcriptome.

In the post genomic era where virtually all the genes and the proteins are known, an important task is to provide a comprehensive analysis of the expression of important classes of genes, such as those that are required for intracellular transport. We report the comprehensive analysis of the Kinesin Superfamily, which is the first and only large protein family whose constituents have been completely identified and confirmed in silico and at the cDNA, mRNA level. In FANTOM2, we have found 90 clones from 33 Kinesin Superfamily Protein (KIF) gene loci. The clones were analyzed in reference to sequence state, library of origin, detection methods, and alternative splicing. More than half of the representative transcriptional units (TU) were full length. The FANTOM2 library also contains novel splice variants previously unreported. We have compared and evaluated various protein classification tools and protein search methods using this data set. This report provides a foundation for future research of the intracellular transport along microtubules and proves the significance of intracellular transport protein transcripts as part of the transcriptome.

Alternative Splicing↗

Transcriptome profiling of sulfur-responsive genes in Arabidopsis reveals global effects of sulfur nutrition on multiple metabolic pathways.

Sulfate is a macronutrient required for cell growth and development. Arabidopsis has two high-affinity sulfate transporters (SULTR1;1 and SULTR1;2) that represent the sulfate uptake activities at the root surface. Sulfur limitation (-S) response relevant to the function of SULTR1;2 was elucidated in this study. We have isolated a novel T-DNA insertion allele defective in the SULTR1;2 sulfate transporter. This mutant, sel1-10, is allelic with the sel1 mutants identified previously in a screen for increased tolerance to selenate, a toxic analog of sulfate (Shibagaki et al., 2002). The abundance of SULTR1;1 mRNA was significantly increased in the sel1-10 mutant; however, this compensatory up-regulation of SULTR1;1 was not sufficient to restore the growth. The sulfate content of the mutant was 10% to 20% of the wild type, suggesting that induction of SULTR1;1 is not fully complementing the function of SULTR1;2 and that SULTR1;2 serves as the major facilitator for the acquisition of sulfate in Arabidopsis roots. Transcriptome analysis of approximately 8,000 Arabidopsis genes in the sel1-10 mutant suggested that dysfunction of the SULTR1;2 transporter can mimic general -S symptoms. Hierarchal clustering of sulfur responsive genes in the wild type and mutant indicated that sulfate uptake, reductive sulfur assimilation, and turnover of secondary sulfur metabolites are activated under -S. The profiles of -S-responsive genes further suggested induction of genes that may alleviate oxidative damage and generation of reactive oxygen species caused by shortage of glutathione.

Arabidopsis↗

SAGE analysis of transcriptome responses in Arabidopsis roots exposed to 2,4,6-trinitrotoluene.

Serial analysis of gene expression was used to profile transcript levels in Arabidopsis roots and assess their responses to 2,4,6-trinitrotoluene (TNT) exposure. SAGE libraries representing control and TNT-exposed seedling root transcripts were constructed, and each was sequenced to a depth of roughly 32,000 tags. More than 19,000 unique tags were identified overall. The second most highly induced tag (27-fold increase) represented a glutathione S-transferase. Cytochrome P450 enzymes, as well as an ABC transporter and a probable nitroreductase, were highly induced by TNT exposure. Analyses also revealed an oxidative stress response upon TNT exposure. Although some increases were anticipated in light of current models for xenobiotic metabolism in plants, evidence for unsuspected conjugation pathways was also noted. Identifying transcriptome-level responses to TNT exposure will better define the metabolic pathways plants use to detoxify this xenobiotic compound, which should help improve phytoremediation strategies directed at TNT and other nitroaromatic compounds.

ATP-Binding Cassette Transporters↗

The Arabidopsis root transcriptome by serial analysis of gene expression. Gene identification using the genome sequence.

Large-scale identification of genes expressed in roots of the model plant Arabidopsis was performed by serial analysis of gene expression (SAGE), on a total of 144,083 sequenced tags, representing at least 15,964 different mRNAs. For tag to gene assignment, we developed a computational approach based on 26,620 genes annotated from the complete sequence of the genome. The procedure selected warrants the identification of the genes corresponding to the majority of the tags found experimentally, with a high level of reliability, and provides a reference database for SAGE studies in Arabidopsis. This new resource allowed us to characterize the expression of more than 3,000 genes, for which there is no expressed sequence tag (EST) or cDNA in the databases. Moreover, 85% of the tags were specific for one gene. To illustrate this advantage of SAGE for functional genomics, we show that our data allow an unambiguous analysis of most of the individual genes belonging to 12 different ion transporter multigene families. These results indicate that, compared with EST-based tag to gene assignment, the use of the annotated genome sequence greatly improves gene identification in SAGE studies. However, more than 6,000 different tags remained with no gene match, suggesting that a significant proportion of transcripts present in the roots originate from yet unknown or wrongly annotated genes. The root transcriptome characterized in this study markedly differs from those obtained in other organs, and provides a unique resource for investigating the functional specificities of the root system. As an example of the use of SAGE for transcript profiling in Arabidopsis, we report here the identification of 270 genes differentially expressed between roots of plants grown either with NO3- or NH4NO3 as N source.

Arabidopsis↗

Evaluation of monocot and eudicot divergence using the sugarcane transcriptome.

Over 40,000 sugarcane (Saccharum officinarum) consensus sequences assembled from 237,954 expressed sequence tags were compared with the protein and DNA sequences from other angiosperms, including the genomes of Arabidopsis and rice (Oryza sativa). Approximately two-thirds of the sugarcane transcriptome have similar sequences in Arabidopsis. These sequences may represent a core set of proteins or protein domains that are conserved among monocots and eudicots and probably encode for essential angiosperm functions. The remaining sequences represent putative monocot-specific genetic material, one-half of which were found only in sugarcane. These monocot-specific cDNAs represent either novelties or, in many cases, fast-evolving sequences that diverged substantially from their eudicot homologs. The wide comparative genome analysis presented here provides information on the evolutionary changes that underlie the divergence of monocots and eudicots. Our comparative analysis also led to the identification of several not yet annotated putative genes and possible gene loss events in Arabidopsis.

Arabidopsis↗