Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Genomic exploration of the hemiascomycetous yeasts: 4. The genome of Saccharomyces cerevisiae revisited.

Since its completion more than 4 years ago, the sequence of Saccharomyces cerevisiae has been extensively used and studied. The original sequence has received a few corrections, and the identification of genes has been completed, thanks in particular to transcriptome analyses and to specialized studies on introns, tRNA genes, transposons or multigene families. In order to undertake the extensive comparative sequence analysis of this program, we have entirely revisited the S. cerevisiae sequence using the same criteria for all 16 chromosomes and taking into account publicly available annotations for genes and elements that cannot be predicted. Comparison with the other yeast species of this program indicates the existence of 50 novel genes in segments previously considered as 'intergenic' and suggests extensions for 26 of the previously annotated genes.

Ascomycota↗

LNCIB human full-length cDNAs collection: towards a better comprehension of the human transcriptome.

LNCIB has been producing a variety of human full-length-enriched, normalized and subtracted cDNA libraries from various cell lines and tissues in different developmental stages by using the CAP-Trapper method. By sequencing 23000 clones of these libraries we identified a pool of about 5800 good quality unique cDNAs. After BLAST analysis on Human RefSeq/Unigene databases, 1717 of these sequences remained with no or poor annotation. We show that cross-species comparative BLAST resulted as a valid tool for the annotation of orthologous genes.

Databases, Factual↗

Microbiology Galaxy Lab: The first community-driven gateway for reproducible and FAIR analysis of microbial data.

The explosion of microbial omics data has outpaced the ability of many researchers to analyze it, with complex tools and limited computational resources creating barriers to discovery. To address this gap, we present the Microbiology Galaxy Lab: a free, globally accessible, community-supported platform that combines state-of-the-art analytical power with user-friendly accessibility. Supported by the Galaxy and global microbiology communities, this platform integrates over 315 tool suites and 115 curated workflows, enabling comprehensive metabarcoding, (meta)genomic, (meta)transcriptomic, and (meta)proteomic data analysis within a FAIR-aligned environment. It also supports research in the health and infectious disease sectors, as well as in environmental microbiology. The platform's utility is exemplified through various use cases, including antimicrobial resistance tracking, biomarker prediction, microbiome classification, and functional annotation of key microbes. Built on reproducibility and community engagement, it supports creation, sharing, and updating of best-practice workflows. Over 35 tutorials and learning paths empower scientists, fostering an ecosystem that keeps resources at the forefront of microbial science. The Microbiology Galaxy Lab enables collective analysis, democratising research, thereby accelerating discovery across the global microbiology community (microbiology.usegalaxy.org, .eu, .org.au, .fr).

Journal Article↗

Cancer target discovery using SAGE.

Cancer is a genetic disease. Genetic events including mutations, chromosomal gains, losses and rearrangements, along with epigenetic alterations, lead to significant transcriptional changes in cancer cells. Changes in the expression of many genes associated with the onset and progression of cancer likely contribute to the cancerous phenotype. SAGE (Serial Analysis of Gene Expression) is an expression profiling method that allows for global, unbiased and quantitative characterisation of transcriptomes. The expression of thousands of genes can be analysed simultaneously without prior knowledge of their sequence, thus leading to the discovery of novel transcripts. In addition to characterising normal and malignant gene expression patterns, SAGE can be used to identify downstream targets of tumour suppressors and oncogenes and further annotate genomes. Comprehensive analyses of expression profiles using SAGE will yield many new diagnostic and prognostic markers as well as therapeutic targets in cancer.

Antineoplastic Agents↗

Morphological, Physiological and Transcriptomic Changes in Response to Water Deficit Stress in Brassica napus L.

Yield losses due to water-deficit (WD) conditions, especially during the reproductive stages of plant development, pose a significant threat to global canola (Brassica napus L.) production. Therefore, it is critical to investigate traits contributing to improved productivity under increased WD conditions. Here we present phenotypic, physiological and transcriptomic changes in response to WD across contrasting canola accessions exhibiting variation in drought resistance-related traits. WD significantly reduced shoot biomass, plant height, harvest index, leaf water content, photosynthetic CO2 assimilation rate, intrinsic water-use efficiency and carbon isotope discrimination. WD caused 49 to 100% of the seed yield reduction: the minimum seed yield reduction (49.66%) was observed in a doubled-haploid (DH) line, 06-5101.137, while the maximum yield reduction (94.1 to 100%) occurred in the late-flowering DH lines (06.5101.088 and 06-5101.306). Seed yield showed a positive correlation (r = 0.29 to 0.95) with shoot biomass and harvest index, leaf water content, photosynthetic CO2 assimilation rate, intrinsic water use efficiency and carbon isotope discrimination. However, it showed negative correlations with days to flower, leaf specific weight, root length, root biomass (r = -0.04 to -0.79) across water treatments. The specific leaf transcriptome analysis of the two parental lines of DH population that exhibit variation for effective water use under well-watered and water-deficient conditions revealed different categories of differentially expressed genes (DEGs): WD-responsive DEGs in BC1329 parental line (1116) and BC9102 (1205) with 754 and 853 DEGs unique to BC1329 and BC9102, respectively, WD-responsive DEGs (906), genotype-dependent DEGs (8465) and genotype × treatment interaction DEGs (353). DEG annotations revealed that the WD-treatment-affected genes were involved in stress responses and growth and development. We further located 235 DEGs within the QTL regions underlying agronomic and physiological performance. Our study provides a conceptual framework for the morphological, physiological and molecular determinants involved in water-use efficiency. Seedlings' traits with high heritability values, such as shoot biomass, leaf weight, leaf water content and Δ13C, serve as proxies for trait-based selection for improved seed yield under both water-limited and non-water-limited conditions.

Brassica napus↗

Sputnik: a database platform for comparative plant genomics.

Two million plant ESTs, from 20 different plant species, and totalling more than one 1000 Mbp of DNA sequence, represents a formidable transcriptomic resource. Sputnik uses the potential of this sequence resource to fill some of the information gap in the un-sequenced plant genomes and to serve as the foundation for in silicio comparative plant genomics. The complexity of the individual EST collections has been reduced using optimised EST clustering techniques. Annotation of cluster sequences is performed by exploiting and transferring information from the comprehensive knowledgebase already produced for the completed model plant genome (Arabidopsis thaliana) and by performing additional state of-the-art sequence analyses relevant to today's plant biologist. Functional predictions, comparative analyses and associative annotations for 500 000 plant EST derived peptides make Sputnik (http://mips.gsf.de/proj/sputnik/) a valid platform for contemporary plant genomics.

Databases, Nucleic Acid↗

The Gene Resource Locator: gene locus maps for transcriptome analysis.

Since the advent of the draft human genome sequence there has been growing interest in transcriptome analysis based on genomic data. The Gene Resource Locator (GRL) assembles gene maps that include information on gene-expression patterns, cis-elements in regulatory regions and alternatively spliced transcripts. The database was constructed using customized software, and currently contains 2.2 million alignments (exon-intron structures). The alignments have been annotated and integrated into a system that encompasses approximately 90 000 EST loci sharing common exons, 8091 alternatively spliced transcript groups, 10 801 expression-profile groups, 8066 candidate regulatory regions in full-length cDNAs, and 1 million SNP loci. We have used Flash technology to build a dynamic web viewer that facilitates browsing through the millions of alignments. All of the information is available through the World Wide Web at the Gene Resource Locator web site (http://grl.gi.k.u-tokyo.ac.jp).

Alternative Splicing↗

Functional Annotation of the Major Histocompatibility Complex Locus.

The human major histocompatibility complex (MHC) locus has the greatest density of disease-associations in the human genome, including links to over 100 polygenic disorders. Its complex haplotype structure, rich gene density, and high degree of linkage disequilibrium combine to make deciphering the gene regulatory logic of the MHC locus extremely challenging. Employing complementary high-throughput CRISPR interference (CRISPRi) and activation (CRISPRa) epigenetic screens coupled with single-cell transcriptome profiling across three distinct human cell types, we identified hundreds of new connections between cis -regulatory elements (CREs) and their target genes in this locus. These CRE-gene links are largely cell type-specific and act as enhancers. Additionally, some CREs have complex features, including harboring both active and repressive histone marks, lacking chromatin accessibility, targeting multiple genes, or acting as silencers. Computational methods fail to predict a majority of these CRE-gene connections. These findings emphasize the potential for functional perturbation experiments to dissect complex loci and reveal shared and cell type-specific regulatory mechanisms relevant to genomics of complex diseases. Collectively, this study provides a unique resource for understanding the complex regulatory landscape within the MHC locus and supports the need for creating new models that encompass CRE-gene interactions, cell type-specific gene expression, and disease genetics in the noncoding genome.

Journal Article↗

Benchmarking the CATMA microarray. A novel tool for Arabidopsis transcriptome analysis.

Transcript profiling is crucial to study biological systems, and various platforms have been implemented to survey mRNAs at the genome scale. We have assessed the performance of the CATMA microarray designed for Arabidopsis (Arabidopsis thaliana) transcriptome analysis and compared it with the Agilent and Affymetrix commercial platforms. The CATMA array consists of gene-specific sequence tags of 150 to 500 bp, the Agilent (Arabidopsis 2) array of 60mer oligonucleotides, and the Affymetrix gene chip (ATH1) of 25mer oligonucleotide sets. We have matched each probe repertoire with the Arabidopsis genome annotation (The Institute for Genomic Research release 5.0) and determined the correspondence between them. Array performance was analyzed by hybridization with labeled targets derived from eight RNA samples made of shoot total RNA spiked with a calibrated series of 14 control transcripts. CATMA arrays showed the largest dynamic range extending over three to four logs. Agilent and Affymetrix arrays displayed a narrower range, presumably because signal saturation occurred for transcripts at concentrations beyond 1,000 copies per cell. Sensitivity was comparable for all three platforms. For Affymetrix GeneChip data, the RMA software package outperformed Microarray Suite 5.0 for all investigated criteria, confirming that the information provided by the mismatch oligonucleotides has no added value. In addition, taking advantage of replicates in our dataset, we conducted a robust statistical analysis of the platform propensity to yield false positive and false negative differentially expressed genes, and all gave satisfactory results. The results establish the CATMA array as a mature alternative to the Affymetrix and Agilent platforms.

Arabidopsis↗

Modification of the transcriptomic response to renal ischemia/reperfusion injury by lipoxin analog.

BACKGROUND: Lipoxins are lipoxygenase-derived eicosanoids with anti-inflammatory and proresolution bioactivities in vitro and in vivo. We have previously demonstrated that the stable synthetic LXA4 analog 15-epi-16-(FPhO)-LXA4-Me is renoprotective in murine renal ischemia/reperfusion injury, as gauged by lower serum creatinine, attenuated leukocyte infiltration, and reduced morphologic tubule injury. METHODS: We employed complementary oligonucleotide microarray and bioinformatic analyses to probe the transcriptomic events that underpin lipoxin renoprotection in this setting. RESULTS: Microarray-based analysis identified three broad categories of genes whose mRNA levels are altered in response to ischemia/reperfusion injury, including known genes previously implicated in the pathogenesis of ischemia/reperfusion injury [e.g., intercellular adhesion molecule-1 (ICAM-1), p21, KIM-1], known genes not previously associated with ischemia/reperfusion injury, and cDNAs representing yet uncharacterized genes. Characterization of expressed sequence tags (ESTs) displayed on microarrays represents a major challenge in studies of global gene expression. A bioinformatic annotation pipeline successfully annotated a large proportion of ESTs modulated during ischemia/reperfusion injury. The differential expression of a representative group of these ischemia/reperfusion injury-modulated genes was confirmed by real-time polymerase chain reaction. Prominent among the up-regulated genes were claudin-1, -3, and -7, and ADAM8. Interestingly, the former response was claudin-specific and was not observed with other claudins expressed by the kidney (e.g., claudin-8 and -6) or indeed with other components of the renal tight junctions (e.g., occludin and junctional adhesion molecule). Noteworthy among the down-regulated genes was a cluster of transport proteins (e.g., aquaporin-1) and the zinc metalloendopeptidase meprin-1 beta implicated in renal remodeling. CONCLUSION: Treatment with the lipoxin analog 15-epi-16-(FPhO)-LXA4-Me prior to injury modified the expression of many differentially expressed pathogenic mediators, including cytokines, growth factors, adhesion molecules, and proteases, suggesting a renoprotective action at the core of the pathophysiology of acute renal failure (ARF). Importantly, this lipoxin-modulated transcriptomic response included many genes expressed by renal parenchymal cells and was not merely a reflection of a reduced renal mRNA load resulting from attenuated leukocyte recruitment. The data presented herein suggest a framework for understanding drivers of kidney injury in ischemia/reperfusion and the molecular basis for renoprotection by lipoxins in this setting.

ADAM Proteins↗

Genome-Wide Identification and Characterization of the TBL Gene Family and Temporal Expression Dynamics During Powdery Mildew Infection in Cucumber (Cucumis sativus).

Cell-wall polysaccharide O-acetylation contributes to cell-wall assembly, organ development, and plant-pathogen interactions, but the cucumber TBL gene family remains poorly characterized. Here, 37 CsTBL genes were identified genome-wide and analyzed using phylogenetic, syntenic, conserved-motif, gene-structure, promoter, protein-structure, Gene Ontology, and transcriptome approaches, followed by RT-qPCR analysis after powdery mildew inoculation. All CsTBL proteins contained the conserved GDS and DxxH motifs, whereas accessory motifs and predicted structural features varied among clades. Intraspecific analysis identified dispersed, WGD/segmental, and tandem duplication categories, and cross-species synteny was more extensive with melon than with Arabidopsis. Homology-derived annotations associated CsTBL genes with cell-wall polysaccharide metabolism, Golgi/endomembrane compartments, and O-acetyltransferase activity, including six genes assigned to xylan O-acetyltransferase-related annotations. Expression profiling revealed tissue- and developmental-stage-dependent patterns, whereas the publicly available powdery mildew RNA-seq dataset provided descriptive temporal expression profiles in Podosphaera xanthii-inoculated samples. Independent RT-qPCR analysis using time-matched mock controls revealed distinct post-inoculation responses among six selected genes. Relative to the corresponding mock controls, CsTBL2 was consistently repressed; CsTBL15 showed transient induction at 1 dpi followed by repression; CsTBL24 exhibited a biphasic response; CsTBL25 was induced at all sampled post-inoculation time points; CsTBL26 showed progressive induction; and CsTBL30 reached its highest observed expression level at 3 dpi. Integrated functional annotation and expression evidence highlighted CsTBL26 as a priority candidate for further functional characterization, while CsTBL24 and CsTBL25 represented fruit-associated candidates with distinct powdery mildew responses; CsTBL30 remained an additional strongly infection-responsive candidate. These findings provide an evolutionary and expression-based framework for the functional characterization of the cucumber TBL gene family.

O-acetylation↗

Transcription profiling of renal cell carcinoma.

AIMS: Our aim was to prepare a comprehensive catalogue of the changes in gene expression accompanying the development and progression of renal cell carcinoma, and to correlate these with histo-pathological, cytogenetic and clinical findings. METHODS: mRNA samples from paired neoplastic and non-cancerous human kidney tissue were labeled and hybridized in duplicate against high-density cDNA arrays. Two array technologies were used: 31,500-element transcriptome-wide nylon arrays for hybridization with 37 radioactively labelled sample pairs, and 4200-element kidney- and cancer-specific glass microarrays for hybridization with 19 fluorescently labelled sample pairs. RESULTS: We identified more than 1700 cDNA clones that show differential transcription levels in kidney tumor tissue compared to normal kidney tissue. The functional classification of 389 annotated genes provided views of the changes in the activities of specific biological processes in renal cancer. Among the biological processes with a large proportion of up-regulated genes we found cell adhesion, signal transduction, and nucleotide metabolism. Down-regulated processes included small molecule transport, ion homeostasis, and oxygen and radical metabolism. Furthermore, we explored the feasibility of molecular diagnosis for renal cell tumors using cDNA microarrays on glass slides, investigating the association of transcription levels with tumor type, progression, and a putative prognostic variable. The experimental data is available from the GEO gene expression database (http://www.ncbi.nlm.nih.gov/geo; accession no. GSE3), and a comprehensive presentation of the results is available in the web supplement (http://www.dkfz-heidelberg.de/abt0840/whuber/rcc). CONCLUSION: Transcription profiling using high-density cDNA arrays is a powerful method with the potential to improve cancer diagnosis and prognosis. The identification and classification of differentially transcribed genes, as described in our study, is the beginning of a more complete understanding of kidney cancer.

Carcinoma, Renal Cell↗

Development of a citrus genome-wide EST collection and cDNA microarray as resources for genomic studies.

A functional genomics project has been initiated to approach the molecular characterization of the main biological and agronomical traits of citrus. As a key part of this project, a citrus EST collection has been generated from 25 cDNA libraries covering different tissues, developmental stages and stress conditions. The collection includes a total of 22,635 high-quality ESTs, grouped in 11,836 putative unigenes, which represent at least one third of the estimated number of genes in the citrus genome. Functional annotation of unigenes which have Arabidopsis orthologues (68% of all unigenes) revealed gene representation in every major functional category, suggesting that a genome-wide EST collection was obtained. A Citrus clementina Hort. ex Tan. cv. Clemenules genomic library, that will contribute to further characterization of relevant genes, has also been constructed. To initiate the analysis of citrus transcriptome, we have developed a cDNA microarray containing 12,672 probes corresponding to 6875 putative unigenes of the collection. Technical characterization of the microarray showed high intra- and inter-array reproducibility, as well as a good range of sensitivity. We have also validated gene expression data achieved with this microarray through an independent technique such as RNA gel blot analysis.

Citrus↗

Large-scale analysis of the human and mouse transcriptomes.

High-throughput gene expression profiling has become an important tool for investigating transcriptional activity in a variety of biological samples. To date, the vast majority of these experiments have focused on specific biological processes and perturbations. Here, we have generated and analyzed gene expression from a set of samples spanning a broad range of biological conditions. Specifically, we profiled gene expression from 91 human and mouse samples across a diverse array of tissues, organs, and cell lines. Because these samples predominantly come from the normal physiological state in the human and mouse, this dataset represents a preliminary, but substantial, description of the normal mammalian transcriptome. We have used this dataset to illustrate methods of mining these data, and to reveal insights into molecular and physiological gene function, mechanisms of transcriptional regulation, disease etiology, and comparative genomics. Finally, to allow the scientific community to use this resource, we have built a free and publicly accessible website (http://expression.gnf.org) that integrates data visualization and curation of current gene annotations.

Animals↗

Odon: an ultra-fast viewer for spatial proteomics.

MOTIVATION: Multiplexed spatial proteomics and spatial transcriptomics generate large, high-dimensional imaging datasets that are challenging to visualize efficiently, particularly at whole-slide and cohort scale. Visualization is an essential step for rapid detection of staining artefacts, such as protein aggregates or non-specific staining. RESULTS: Here, we present Odon, a native Rust desktop viewer designed for rapid, interactive exploration of multiplex imaging data on a standard laptop. Odon is primarily built around the OME-Zarr imaging format, and supports annotations via GeoJSON and GeoParquet, with secondary support for SpatialData, Xenium containers, and TIFF. Data can be stored locally or streamed directly from HTTP or S3-compatible object storage using viewport-driven tile loading. Odon incorporates a highly optimized rendering engine designed for viewport-driven tile loading and GPU-based compositing. In scripted benchmarks using synthetic multiplex OME-Zarr datasets, Odon showed lower peak memory use, lower affine-derived zoom-step error, and faster warm-start image loading than napari and QuPath under the tested conditions. Its GPU-based compositing pipeline also enables smooth rendering and interaction with >1 000 000 segmented cells. Odon further supports integrated visual analytics, including live thresholding and cell selection, and a mosaic mode for simultaneous viewing of hundreds of regions of interest in cohort and tissue microarray studies. Together, these features establish Odon as a high-performance platform for scalable visualization of spatial proteomics data. AVAILABILITY AND IMPLEMENTATION: Source code and compiled installers are available at https://github.com/alexcoulton/odon.

Proteomics↗

Transcriptome profiling of adult zebrafish at the late stage of chronic tuberculosis due to Mycobacterium marinum infection.

The Mycobacterium marinum-zebrafish infection model was used in this study for analysis of a host transcriptome response to mycobacterium infection at the organismal level. RNA isolated from adult zebrafish that showed typical signs of fish tuberculosis due to a chronic progressive infection with M. marinum was compared with RNA from healthy fish in microarray analyses. Spotted oligonucleotide sets (designed by Sigma-Compugen and MWG) and Affymetrix GeneChips were used, in total comprising 45,465 zebrafish transcript annotations. Based on a detailed comparative analysis and quantitative reverse transcriptase-PCR analysis, we present a validated reference set of 159 genes whose regulation is strongly affected by mycobacterial infection in the three types of microarrays analyzed. Furthermore, we analyzed the separate datasets of the microarrays with special emphasis on the expression profiles of immune-related genes. Upregulated genes include many known components of the inflammatory response and several genes that have previously been implicated in the response to mycobacterial infections in cell cultures of other organisms. Different marker genes of the myeloid lineage that have been characterized in zebrafish also showed increased expression. Furthermore, the zebrafish homologs of many signal transduction genes with relationship to the immune response were induced by M. marinum infection. Future functional analysis of these genes may contribute to understanding the mechanisms of mycobacterial pathogenesis. Since a large group of genes linked to immune responses did not show altered expression in the infected animals, these results suggest specific responses in mycobacterium-induced disease.

Animals↗

A method for automated detection of gene expression required for the establishment of a digital transcriptome-wide gene expression atlas.

Acquiring information about the expression of a gene in different cell populations and tissues can provide key insight into the function of the gene. A high-throughput in situ hybridization (ISH) method was recently developed for rapid and reproducible acquisition of gene expression patterns in serial tissue sections at cellular resolution. Characterizing and analysing expression patterns on thousands of sections requires efficient methods for locating cells and estimating the level of expression in each cell. Such cellular quantification is an essential step in both annotating and quantitatively comparing high-throughput ISH results. Here we describe a novel automated and efficient methodology for performing this quantification on postnatal mouse brain.

Animals↗

Colorectal Liver Metastasis Pathomics Model: Integrating Single-Cell and Spatial Transcriptome Analysis With Pathomics for Predicting Liver Metastasis in Colorectal Cancer.

The liver is the primary target organ for hematologic metastasis of colorectal cancer (CRC), and CRC liver metastasis (CRLM) often precludes radical resection, making it the leading cause of death in patients with CRC. To improve the identification and prediction of liver metastasis risk, we identified a cell type of liver metastasis--triggering malignant cells (LMTMCs) through integrating single-cell RNA sequencing and spatial transcriptome analysis. Multiomics cell communication analysis indicated that the interaction between fibroblasts and LMTMCs through the COL1A1-CD44/SDC4 and LAMA4-CD44 signaling axes could promote CRLM. By applying the one-class logistic regression algorithm, we developed a CRLM scoring system in the bulk RNA-sequencing data according to the abundance of LMTMCs in each individual. Using the grouping labels derived from the CRLM scoring system in the bulk data and the corresponding whole-slide images without any manual annotations at the region or pixel level, processed via slide-level weakly supervised learning, a deep-learning model based on the ResNet18 architecture, called Colorectal Liver Metastasis Pathomics Model, was developed to predict the risk of liver metastasis in patients with CRC. The Colorectal Liver Metastasis Pathomics Model achieved an area under the curve of 0.84 at the internal test set of The Cancer Genome Atlas-CRC histology images. In the external independent validation sets, namely the Affiliated Hospital of Southwest Medical University and the Affiliated Traditional Chinese Medicine Hospital of Southwest Medical University cohorts, the areas under the curve were 0.89 and 0.72, respectively, indicating effective classification performances. This study provided new insights and tools for the early identification of CRLM and demonstrated the potential of combining multiomics with deep learning-based pathomics in cancer research.

Humans↗