Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

A high-resolution map of active promoters in the human genome.

In eukaryotic cells, transcription of every protein-coding gene begins with the assembly of an RNA polymerase II preinitiation complex (PIC) on the promoter. The promoters, in conjunction with enhancers, silencers and insulators, define the combinatorial codes that specify gene expression patterns. Our ability to analyse the control logic encoded in the human genome is currently limited by a lack of accurate information regarding the promoters for most genes. Here we describe a genome-wide map of active promoters in human fibroblast cells, determined by experimentally locating the sites of PIC binding throughout the human genome. This map defines 10,567 active promoters corresponding to 6,763 known genes and at least 1,196 un-annotated transcriptional units. Features of the map suggest extensive use of multiple promoters by the human genes and widespread clustering of active promoters in the genome. In addition, examination of the genome-wide expression profile reveals four general classes of promoters that define the transcriptome of the cell. These results provide a global view of the functional relationships among transcriptional machinery, chromatin structure and gene expression in human cells.

Chromatin↗

Genome-wide epigenomic atlas and multi-omics responses of Eriocheir sinensis to natural extreme heat.

BACKGROUND: Global climate warming has led to increasingly frequent and prolonged extreme summer heat events, posing severe environmental challenges to aquaculture systems. Extreme summer heat can disrupt the performance of pond-cultured ectotherms. The Chinese mitten crab (Eriocheir sinensis) is an economically important freshwater crustacean, but coordinated molecular differences following contrasting natural summers remain incompletely characterized. RESULTS: We performed a comprehensive multi-omics analysis integrating meteorological monitoring, mRNA/lncRNA transcriptomics, small-RNA profiling of miRNAs, DNA methylomics, and LC-MS metabolomics in E. sinensis populations collected from Yancheng, China, between 2020 and 2024. Across the ten farms, survival was significantly lower in 2024, whereas yield and the proportion of large individuals showed nonsignificant downward trends. Gene-set analyses showed negative enrichment of cellular heat-response, protein-folding, oxidative-phosphorylation, and mitochondrial ATP-production terms in the 2024 cohort at the time of sampling. The integrated transcript annotation contained 72,240 lncRNAs and 63,833 mRNAs, and CpG was the predominant methylation context. Differential methylation analysis identified 73 regions and 185 cytosines, with hypomethylated events predominating within the significant subset. Metabolomic profiles differed between annual cohorts and mapped to carbohydrate, lipid, and amino-acid pathways. Cross-omics integration prioritized eight candidate genes-ADCY9, UNC79, UBN1, IFT52, ACO2, LOC126986070, LOC127001126, and LOC126997895-and qPCR reproduced the reported directions of expression for selected RNAs. CONCLUSION: This study provides the first integrative multi-omics framework for understanding chronic heat adaptation in E. sinensis. By linking transcriptomic, epigenomic, and metabolic remodeling, we elucidate the molecular mechanisms underlying energy imbalance, epigenetic reprogramming, and immune dysregulation during prolonged thermal stress. These findings offer valuable insights and genomic resources for breeding heat-tolerant crab strains and improving aquaculture resilience under ongoing climate change.

DNA methylation↗

Distinct molecular responses to acute cold exposure revealed by comparative transcriptomic and metabolomic profiling in the bay scallop Argopecten irradians.

Acute cold stress can elicit distinct molecular responses even when bay scallop populations show similar phenotypic outcomes. We compared a seventh-generation fast-growing bay scallop line (BS) with a commercial control population (CC) during a 72-h acute cold exposure at -1 ± 0.3 °C. RNA-seq was used as the discovery layer, representative BS cold-responsive genes were evaluated by qRT-PCR, and paired LC-MS profiles provided a comparative metabolic layer. At baseline, 138 genes differed between BS and CC; after cold exposure, 134 of these baseline differences disappeared and 61 of 65 cold-state differences newly emerged. BS showed a larger transcriptomic response magnitude than CC, with 1129 cold-responsive genes compared with 28 genes in CC, and this ordering remained robust across multiple sensitivity analyses. Survival after 72 h was identical in BS and CC (83/90, 92.2% in each population). Biochemical responses were time-dependent and marker-specific: CAT, LZM, T-SOD and T-AOC showed population-by-time interactions, whereas GSH-Px and MDA did not, and the 72-h differences were not consistently favourable to BS. Metabolomic cold effects were strongly concordant between populations, and no feature showed a significant population-by-cold interaction. Features putatively assigned to arachidonic acid metabolism were enriched, but this provider-annotated pathway signal remains exploratory because authentic-standard confirmation was not performed. These findings indicate population-specific differences in molecular responsiveness but do not establish superior cold tolerance in BS.

Animals↗

Sepsis gene expression profiling: murine splenic compared with hepatic responses determined by using complementary DNA microarrays.

OBJECTIVE: DNA microarrays allow genome-wide assessment of changes in relative messenger RNA abundance and thus can be used to monitor changes in gene expression. The aim of this series of experiments was to gain experience in sepsis gene expression profiling in a well-accepted model of murine polymicrobial abdominal sepsis and begin characterizing (in the parlance of genomics) the sepsis "transcriptome." DESIGN: Prospective animal study. SETTING: University-based animal research facility.SUBJECTS C57BL/6 mice. INTERVENTIONS: After induction of general anesthesia, cecal ligation and puncture were performed to induce peritonitis and polymicrobial sepsis. The control group had sham laparotomy only. Three samples of spleen and liver were collected from septic and sham animals at 24 hrs after laparotomy. Changes in expression were measured for 588 annotated mouse genes by using a commercially available complementary DNA microarray kit. MEASUREMENTS AND MAIN RESULTS: Broad-scale gene expression profiles were characterized for septic liver and spleen and compared with sham controls. The analytical tools used included commercially available software packages and a novel analysis program. Very little overlap was observed in the septic gene expression profiles of these two organs. Most of the genes identified have previously been linked to regulation of the inflammatory response; importantly, however, some have not. In addition, hierarchical cluster analysis showed that cecal ligation and puncture at 24 hrs induced coordinate expression of genes that alter cell signaling and survival pathways in spleen, consistent with previously published reports of sepsis-induced splenocyte apoptosis. The current limitations of microarray analysis as reflected in these studies are also discussed. CONCLUSIONS: Microarray technology provides a powerful new tool for rapidly analyzing tissue-specific changes in gene expression induced by sepsis in animal models. To our knowledge, these data constitute the first report on the use of microarrays to determine the sepsis transcriptome.

Animals↗

Genomic exploration of the hemiascomycetous yeasts: 4. The genome of Saccharomyces cerevisiae revisited.

Since its completion more than 4 years ago, the sequence of Saccharomyces cerevisiae has been extensively used and studied. The original sequence has received a few corrections, and the identification of genes has been completed, thanks in particular to transcriptome analyses and to specialized studies on introns, tRNA genes, transposons or multigene families. In order to undertake the extensive comparative sequence analysis of this program, we have entirely revisited the S. cerevisiae sequence using the same criteria for all 16 chromosomes and taking into account publicly available annotations for genes and elements that cannot be predicted. Comparison with the other yeast species of this program indicates the existence of 50 novel genes in segments previously considered as 'intergenic' and suggests extensions for 26 of the previously annotated genes.

Ascomycota↗

LNCIB human full-length cDNAs collection: towards a better comprehension of the human transcriptome.

LNCIB has been producing a variety of human full-length-enriched, normalized and subtracted cDNA libraries from various cell lines and tissues in different developmental stages by using the CAP-Trapper method. By sequencing 23000 clones of these libraries we identified a pool of about 5800 good quality unique cDNAs. After BLAST analysis on Human RefSeq/Unigene databases, 1717 of these sequences remained with no or poor annotation. We show that cross-species comparative BLAST resulted as a valid tool for the annotation of orthologous genes.

Databases, Factual↗

Microbiology Galaxy Lab: The first community-driven gateway for reproducible and FAIR analysis of microbial data.

The explosion of microbial omics data has outpaced the ability of many researchers to analyze it, with complex tools and limited computational resources creating barriers to discovery. To address this gap, we present the Microbiology Galaxy Lab: a free, globally accessible, community-supported platform that combines state-of-the-art analytical power with user-friendly accessibility. Supported by the Galaxy and global microbiology communities, this platform integrates over 315 tool suites and 115 curated workflows, enabling comprehensive metabarcoding, (meta)genomic, (meta)transcriptomic, and (meta)proteomic data analysis within a FAIR-aligned environment. It also supports research in the health and infectious disease sectors, as well as in environmental microbiology. The platform's utility is exemplified through various use cases, including antimicrobial resistance tracking, biomarker prediction, microbiome classification, and functional annotation of key microbes. Built on reproducibility and community engagement, it supports creation, sharing, and updating of best-practice workflows. Over 35 tutorials and learning paths empower scientists, fostering an ecosystem that keeps resources at the forefront of microbial science. The Microbiology Galaxy Lab enables collective analysis, democratising research, thereby accelerating discovery across the global microbiology community (microbiology.usegalaxy.org, .eu, .org.au, .fr).

Journal Article↗

Decoding the fine-scale structure of a breast cancer genome and transcriptome.

A comprehensive understanding of cancer is predicated upon knowledge of the structure of malignant genomes underlying its many variant forms and the molecular mechanisms giving rise to them. It is well established that solid tumor genomes accumulate a large number of genome rearrangements during tumorigenesis. End Sequence Profiling (ESP) maps and clones genome breakpoints associated with all types of genome rearrangements elucidating the structural organization of tumor genomes. Here we extend the ESP methodology in several directions using the breast cancer cell line MCF-7. First, targeted ESP is applied to multiple amplified loci, revealing a complex process of rearrangement and co-amplification in these regions reminiscent of breakage/fusion/bridge cycles. Second, genome breakpoints identified by ESP are confirmed using a combination of DNA sequencing and PCR. Third, in vitro functional studies assign biological function to a rearranged tumor BAC clone, demonstrating that it encodes anti-apoptotic activity. Finally, ESP is extended to the transcriptome identifying four novel fusion transcripts and providing evidence that expression of fusion genes may be common in tumors. These results demonstrate the distinct advantages of ESP including: (1) the ability to detect all types of rearrangements and copy number changes; (2) straightforward integration of ESP data with the annotated genome sequence; (3) immortalization of the genome; (4) ability to generate tumor-specific reagents for in vitro and in vivo functional studies. Given these properties, ESP could play an important role in a tumor genome project.

Breast Neoplasms↗

Transcriptome analysis of human gastric cancer.

To elucidate the genetic events associated with gastric cancer, 124,704 cDNA clones were collected from 37 human gastric cDNA libraries, including 20 full-length enriched cDNA libraries of gastric cancer cell lines and tissues from Korean patients. An analysis of the collected ESTs revealed that 97,930 high-quality ESTs coalesced into 13,001 clusters, of which 11,135 clusters (85.6%) were annotated to known ESTs. The analysis of the full-length cDNAs also revealed that 4862 clusters (51.7%) contained at least one putative full-length cDNA clone with an initiation codon, with the average length of the 5' UTR of 140 bp. A large number appear to have a diverse transcription start site (TSS). An examination of the TSS of some genes, such as TEGT and GAPD, using 5' RACE revealed that the predicted TSSs are actually found in human gastric cancer cells and that several TSSs differ depending on the specific gastric cell line. Furthermore, of the human gastric ESTs, 766 genes (9.5%) were present as putative alternatively spliced variants. Confirmation of the predicted spliced isoforms using RT-PCR showed that the predicted isoforms exist in gastric cancer cells and some isoforms coexist in gastric cell lines. These results provide potentially useful information for elucidating the molecular mechanisms associated with gastric oncogenesis.

5' Untranslated Regions↗

Cancer target discovery using SAGE.

Cancer is a genetic disease. Genetic events including mutations, chromosomal gains, losses and rearrangements, along with epigenetic alterations, lead to significant transcriptional changes in cancer cells. Changes in the expression of many genes associated with the onset and progression of cancer likely contribute to the cancerous phenotype. SAGE (Serial Analysis of Gene Expression) is an expression profiling method that allows for global, unbiased and quantitative characterisation of transcriptomes. The expression of thousands of genes can be analysed simultaneously without prior knowledge of their sequence, thus leading to the discovery of novel transcripts. In addition to characterising normal and malignant gene expression patterns, SAGE can be used to identify downstream targets of tumour suppressors and oncogenes and further annotate genomes. Comprehensive analyses of expression profiles using SAGE will yield many new diagnostic and prognostic markers as well as therapeutic targets in cancer.

Antineoplastic Agents↗

Morphological, Physiological and Transcriptomic Changes in Response to Water Deficit Stress in Brassica napus L.

Yield losses due to water-deficit (WD) conditions, especially during the reproductive stages of plant development, pose a significant threat to global canola (Brassica napus L.) production. Therefore, it is critical to investigate traits contributing to improved productivity under increased WD conditions. Here we present phenotypic, physiological and transcriptomic changes in response to WD across contrasting canola accessions exhibiting variation in drought resistance-related traits. WD significantly reduced shoot biomass, plant height, harvest index, leaf water content, photosynthetic CO2 assimilation rate, intrinsic water-use efficiency and carbon isotope discrimination. WD caused 49 to 100% of the seed yield reduction: the minimum seed yield reduction (49.66%) was observed in a doubled-haploid (DH) line, 06-5101.137, while the maximum yield reduction (94.1 to 100%) occurred in the late-flowering DH lines (06.5101.088 and 06-5101.306). Seed yield showed a positive correlation (r = 0.29 to 0.95) with shoot biomass and harvest index, leaf water content, photosynthetic CO2 assimilation rate, intrinsic water use efficiency and carbon isotope discrimination. However, it showed negative correlations with days to flower, leaf specific weight, root length, root biomass (r = -0.04 to -0.79) across water treatments. The specific leaf transcriptome analysis of the two parental lines of DH population that exhibit variation for effective water use under well-watered and water-deficient conditions revealed different categories of differentially expressed genes (DEGs): WD-responsive DEGs in BC1329 parental line (1116) and BC9102 (1205) with 754 and 853 DEGs unique to BC1329 and BC9102, respectively, WD-responsive DEGs (906), genotype-dependent DEGs (8465) and genotype × treatment interaction DEGs (353). DEG annotations revealed that the WD-treatment-affected genes were involved in stress responses and growth and development. We further located 235 DEGs within the QTL regions underlying agronomic and physiological performance. Our study provides a conceptual framework for the morphological, physiological and molecular determinants involved in water-use efficiency. Seedlings' traits with high heritability values, such as shoot biomass, leaf weight, leaf water content and Δ13C, serve as proxies for trait-based selection for improved seed yield under both water-limited and non-water-limited conditions.

Brassica napus↗

Arabidopsis thaliana full genome longmer microarrays: a powerful gene discovery tool for agriculture and forestry.

Sequenced plant genomes provide a large reservoir of known genes with potential for use in crop and tree improvement, but assignment of specific functions to annotated genes in sequenced plant genomes remains a challenge. Furthermore, most plant genes belong to families encoding proteins with related but distinct functions. In this commentary, we discuss our development of Arabidopsis spotted whole genome longmer oligonucleotide microarrays, and their use in global transcription profiling. We show that longmer array based transcriptome analysis in Arabidopsis can be used as an efficient and effective gene discovery and functional genomics tool, particularly for functional analyses of members of large gene families. We discuss experiments that focus on gene families involved in phenylpropanoid natural product biosynthesis and fiber differentiation. These analyses have helped to elucidate functions of individual gene family members, and have identified new candidate genes involved in fiber development and differentiation. Results obtained by these studies in Arabidopsis can be used as the basis for gene discovery in commercially important plants, and we have focused our attention on Populus trichocarpa (poplar), a species important in forestry and agroforestry for which complete genome sequence information is available.

Agriculture↗

Sputnik: a database platform for comparative plant genomics.

Two million plant ESTs, from 20 different plant species, and totalling more than one 1000 Mbp of DNA sequence, represents a formidable transcriptomic resource. Sputnik uses the potential of this sequence resource to fill some of the information gap in the un-sequenced plant genomes and to serve as the foundation for in silicio comparative plant genomics. The complexity of the individual EST collections has been reduced using optimised EST clustering techniques. Annotation of cluster sequences is performed by exploiting and transferring information from the comprehensive knowledgebase already produced for the completed model plant genome (Arabidopsis thaliana) and by performing additional state of-the-art sequence analyses relevant to today's plant biologist. Functional predictions, comparative analyses and associative annotations for 500 000 plant EST derived peptides make Sputnik (http://mips.gsf.de/proj/sputnik/) a valid platform for contemporary plant genomics.

Databases, Nucleic Acid↗

The Gene Resource Locator: gene locus maps for transcriptome analysis.

Since the advent of the draft human genome sequence there has been growing interest in transcriptome analysis based on genomic data. The Gene Resource Locator (GRL) assembles gene maps that include information on gene-expression patterns, cis-elements in regulatory regions and alternatively spliced transcripts. The database was constructed using customized software, and currently contains 2.2 million alignments (exon-intron structures). The alignments have been annotated and integrated into a system that encompasses approximately 90 000 EST loci sharing common exons, 8091 alternatively spliced transcript groups, 10 801 expression-profile groups, 8066 candidate regulatory regions in full-length cDNAs, and 1 million SNP loci. We have used Flash technology to build a dynamic web viewer that facilitates browsing through the millions of alignments. All of the information is available through the World Wide Web at the Gene Resource Locator web site (http://grl.gi.k.u-tokyo.ac.jp).

Alternative Splicing↗

Gene-dosage effect on chromosome 21 transcriptome in trisomy 21: implication in Down syndrome cognitive disorders.

In the era of human functional genomics, the chromosome 21 has represented a prototype for pioneering global biotechnologies. Its relatively low gene content enabled studying Down syndrome at the chromosomal scale, for which the last years have seen intense research activity aiming at genotype-phenotype correlations. The global gene-dose dependent upregulation of gene expression seen in the context of trisomy and preliminary functional annotation of chromosome 21 genes points towards candidate genes and molecular pathways potentially associated with the cognitive defects observed in Down syndrome.

Chromosome Mapping↗

Functional Annotation of the Major Histocompatibility Complex Locus.

The human major histocompatibility complex (MHC) locus has the greatest density of disease-associations in the human genome, including links to over 100 polygenic disorders. Its complex haplotype structure, rich gene density, and high degree of linkage disequilibrium combine to make deciphering the gene regulatory logic of the MHC locus extremely challenging. Employing complementary high-throughput CRISPR interference (CRISPRi) and activation (CRISPRa) epigenetic screens coupled with single-cell transcriptome profiling across three distinct human cell types, we identified hundreds of new connections between cis -regulatory elements (CREs) and their target genes in this locus. These CRE-gene links are largely cell type-specific and act as enhancers. Additionally, some CREs have complex features, including harboring both active and repressive histone marks, lacking chromatin accessibility, targeting multiple genes, or acting as silencers. Computational methods fail to predict a majority of these CRE-gene connections. These findings emphasize the potential for functional perturbation experiments to dissect complex loci and reveal shared and cell type-specific regulatory mechanisms relevant to genomics of complex diseases. Collectively, this study provides a unique resource for understanding the complex regulatory landscape within the MHC locus and supports the need for creating new models that encompass CRE-gene interactions, cell type-specific gene expression, and disease genetics in the noncoding genome.

Journal Article↗

Benchmarking the CATMA microarray. A novel tool for Arabidopsis transcriptome analysis.

Transcript profiling is crucial to study biological systems, and various platforms have been implemented to survey mRNAs at the genome scale. We have assessed the performance of the CATMA microarray designed for Arabidopsis (Arabidopsis thaliana) transcriptome analysis and compared it with the Agilent and Affymetrix commercial platforms. The CATMA array consists of gene-specific sequence tags of 150 to 500 bp, the Agilent (Arabidopsis 2) array of 60mer oligonucleotides, and the Affymetrix gene chip (ATH1) of 25mer oligonucleotide sets. We have matched each probe repertoire with the Arabidopsis genome annotation (The Institute for Genomic Research release 5.0) and determined the correspondence between them. Array performance was analyzed by hybridization with labeled targets derived from eight RNA samples made of shoot total RNA spiked with a calibrated series of 14 control transcripts. CATMA arrays showed the largest dynamic range extending over three to four logs. Agilent and Affymetrix arrays displayed a narrower range, presumably because signal saturation occurred for transcripts at concentrations beyond 1,000 copies per cell. Sensitivity was comparable for all three platforms. For Affymetrix GeneChip data, the RMA software package outperformed Microarray Suite 5.0 for all investigated criteria, confirming that the information provided by the mismatch oligonucleotides has no added value. In addition, taking advantage of replicates in our dataset, we conducted a robust statistical analysis of the platform propensity to yield false positive and false negative differentially expressed genes, and all gave satisfactory results. The results establish the CATMA array as a mature alternative to the Affymetrix and Agilent platforms.

Arabidopsis↗

Modification of the transcriptomic response to renal ischemia/reperfusion injury by lipoxin analog.

BACKGROUND: Lipoxins are lipoxygenase-derived eicosanoids with anti-inflammatory and proresolution bioactivities in vitro and in vivo. We have previously demonstrated that the stable synthetic LXA4 analog 15-epi-16-(FPhO)-LXA4-Me is renoprotective in murine renal ischemia/reperfusion injury, as gauged by lower serum creatinine, attenuated leukocyte infiltration, and reduced morphologic tubule injury. METHODS: We employed complementary oligonucleotide microarray and bioinformatic analyses to probe the transcriptomic events that underpin lipoxin renoprotection in this setting. RESULTS: Microarray-based analysis identified three broad categories of genes whose mRNA levels are altered in response to ischemia/reperfusion injury, including known genes previously implicated in the pathogenesis of ischemia/reperfusion injury [e.g., intercellular adhesion molecule-1 (ICAM-1), p21, KIM-1], known genes not previously associated with ischemia/reperfusion injury, and cDNAs representing yet uncharacterized genes. Characterization of expressed sequence tags (ESTs) displayed on microarrays represents a major challenge in studies of global gene expression. A bioinformatic annotation pipeline successfully annotated a large proportion of ESTs modulated during ischemia/reperfusion injury. The differential expression of a representative group of these ischemia/reperfusion injury-modulated genes was confirmed by real-time polymerase chain reaction. Prominent among the up-regulated genes were claudin-1, -3, and -7, and ADAM8. Interestingly, the former response was claudin-specific and was not observed with other claudins expressed by the kidney (e.g., claudin-8 and -6) or indeed with other components of the renal tight junctions (e.g., occludin and junctional adhesion molecule). Noteworthy among the down-regulated genes was a cluster of transport proteins (e.g., aquaporin-1) and the zinc metalloendopeptidase meprin-1 beta implicated in renal remodeling. CONCLUSION: Treatment with the lipoxin analog 15-epi-16-(FPhO)-LXA4-Me prior to injury modified the expression of many differentially expressed pathogenic mediators, including cytokines, growth factors, adhesion molecules, and proteases, suggesting a renoprotective action at the core of the pathophysiology of acute renal failure (ARF). Importantly, this lipoxin-modulated transcriptomic response included many genes expressed by renal parenchymal cells and was not merely a reflection of a reduced renal mRNA load resulting from attenuated leukocyte recruitment. The data presented herein suggest a framework for understanding drivers of kidney injury in ischemia/reperfusion and the molecular basis for renoprotection by lipoxins in this setting.

ADAM Proteins↗