Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Decoding the fine-scale structure of a breast cancer genome and transcriptome.

A comprehensive understanding of cancer is predicated upon knowledge of the structure of malignant genomes underlying its many variant forms and the molecular mechanisms giving rise to them. It is well established that solid tumor genomes accumulate a large number of genome rearrangements during tumorigenesis. End Sequence Profiling (ESP) maps and clones genome breakpoints associated with all types of genome rearrangements elucidating the structural organization of tumor genomes. Here we extend the ESP methodology in several directions using the breast cancer cell line MCF-7. First, targeted ESP is applied to multiple amplified loci, revealing a complex process of rearrangement and co-amplification in these regions reminiscent of breakage/fusion/bridge cycles. Second, genome breakpoints identified by ESP are confirmed using a combination of DNA sequencing and PCR. Third, in vitro functional studies assign biological function to a rearranged tumor BAC clone, demonstrating that it encodes anti-apoptotic activity. Finally, ESP is extended to the transcriptome identifying four novel fusion transcripts and providing evidence that expression of fusion genes may be common in tumors. These results demonstrate the distinct advantages of ESP including: (1) the ability to detect all types of rearrangements and copy number changes; (2) straightforward integration of ESP data with the annotated genome sequence; (3) immortalization of the genome; (4) ability to generate tumor-specific reagents for in vitro and in vivo functional studies. Given these properties, ESP could play an important role in a tumor genome project.

Breast Neoplasms↗

Transcriptome analysis of human gastric cancer.

To elucidate the genetic events associated with gastric cancer, 124,704 cDNA clones were collected from 37 human gastric cDNA libraries, including 20 full-length enriched cDNA libraries of gastric cancer cell lines and tissues from Korean patients. An analysis of the collected ESTs revealed that 97,930 high-quality ESTs coalesced into 13,001 clusters, of which 11,135 clusters (85.6%) were annotated to known ESTs. The analysis of the full-length cDNAs also revealed that 4862 clusters (51.7%) contained at least one putative full-length cDNA clone with an initiation codon, with the average length of the 5' UTR of 140 bp. A large number appear to have a diverse transcription start site (TSS). An examination of the TSS of some genes, such as TEGT and GAPD, using 5' RACE revealed that the predicted TSSs are actually found in human gastric cancer cells and that several TSSs differ depending on the specific gastric cell line. Furthermore, of the human gastric ESTs, 766 genes (9.5%) were present as putative alternatively spliced variants. Confirmation of the predicted spliced isoforms using RT-PCR showed that the predicted isoforms exist in gastric cancer cells and some isoforms coexist in gastric cell lines. These results provide potentially useful information for elucidating the molecular mechanisms associated with gastric oncogenesis.

5' Untranslated Regions↗

Cancer target discovery using SAGE.

Cancer is a genetic disease. Genetic events including mutations, chromosomal gains, losses and rearrangements, along with epigenetic alterations, lead to significant transcriptional changes in cancer cells. Changes in the expression of many genes associated with the onset and progression of cancer likely contribute to the cancerous phenotype. SAGE (Serial Analysis of Gene Expression) is an expression profiling method that allows for global, unbiased and quantitative characterisation of transcriptomes. The expression of thousands of genes can be analysed simultaneously without prior knowledge of their sequence, thus leading to the discovery of novel transcripts. In addition to characterising normal and malignant gene expression patterns, SAGE can be used to identify downstream targets of tumour suppressors and oncogenes and further annotate genomes. Comprehensive analyses of expression profiles using SAGE will yield many new diagnostic and prognostic markers as well as therapeutic targets in cancer.

Antineoplastic Agents↗

Morphological, Physiological and Transcriptomic Changes in Response to Water Deficit Stress in Brassica napus L.

Yield losses due to water-deficit (WD) conditions, especially during the reproductive stages of plant development, pose a significant threat to global canola (Brassica napus L.) production. Therefore, it is critical to investigate traits contributing to improved productivity under increased WD conditions. Here we present phenotypic, physiological and transcriptomic changes in response to WD across contrasting canola accessions exhibiting variation in drought resistance-related traits. WD significantly reduced shoot biomass, plant height, harvest index, leaf water content, photosynthetic CO2 assimilation rate, intrinsic water-use efficiency and carbon isotope discrimination. WD caused 49 to 100% of the seed yield reduction: the minimum seed yield reduction (49.66%) was observed in a doubled-haploid (DH) line, 06-5101.137, while the maximum yield reduction (94.1 to 100%) occurred in the late-flowering DH lines (06.5101.088 and 06-5101.306). Seed yield showed a positive correlation (r = 0.29 to 0.95) with shoot biomass and harvest index, leaf water content, photosynthetic CO2 assimilation rate, intrinsic water use efficiency and carbon isotope discrimination. However, it showed negative correlations with days to flower, leaf specific weight, root length, root biomass (r = -0.04 to -0.79) across water treatments. The specific leaf transcriptome analysis of the two parental lines of DH population that exhibit variation for effective water use under well-watered and water-deficient conditions revealed different categories of differentially expressed genes (DEGs): WD-responsive DEGs in BC1329 parental line (1116) and BC9102 (1205) with 754 and 853 DEGs unique to BC1329 and BC9102, respectively, WD-responsive DEGs (906), genotype-dependent DEGs (8465) and genotype × treatment interaction DEGs (353). DEG annotations revealed that the WD-treatment-affected genes were involved in stress responses and growth and development. We further located 235 DEGs within the QTL regions underlying agronomic and physiological performance. Our study provides a conceptual framework for the morphological, physiological and molecular determinants involved in water-use efficiency. Seedlings' traits with high heritability values, such as shoot biomass, leaf weight, leaf water content and Δ13C, serve as proxies for trait-based selection for improved seed yield under both water-limited and non-water-limited conditions.

Brassica napus↗

Arabidopsis thaliana full genome longmer microarrays: a powerful gene discovery tool for agriculture and forestry.

Sequenced plant genomes provide a large reservoir of known genes with potential for use in crop and tree improvement, but assignment of specific functions to annotated genes in sequenced plant genomes remains a challenge. Furthermore, most plant genes belong to families encoding proteins with related but distinct functions. In this commentary, we discuss our development of Arabidopsis spotted whole genome longmer oligonucleotide microarrays, and their use in global transcription profiling. We show that longmer array based transcriptome analysis in Arabidopsis can be used as an efficient and effective gene discovery and functional genomics tool, particularly for functional analyses of members of large gene families. We discuss experiments that focus on gene families involved in phenylpropanoid natural product biosynthesis and fiber differentiation. These analyses have helped to elucidate functions of individual gene family members, and have identified new candidate genes involved in fiber development and differentiation. Results obtained by these studies in Arabidopsis can be used as the basis for gene discovery in commercially important plants, and we have focused our attention on Populus trichocarpa (poplar), a species important in forestry and agroforestry for which complete genome sequence information is available.

Agriculture↗

Sputnik: a database platform for comparative plant genomics.

Two million plant ESTs, from 20 different plant species, and totalling more than one 1000 Mbp of DNA sequence, represents a formidable transcriptomic resource. Sputnik uses the potential of this sequence resource to fill some of the information gap in the un-sequenced plant genomes and to serve as the foundation for in silicio comparative plant genomics. The complexity of the individual EST collections has been reduced using optimised EST clustering techniques. Annotation of cluster sequences is performed by exploiting and transferring information from the comprehensive knowledgebase already produced for the completed model plant genome (Arabidopsis thaliana) and by performing additional state of-the-art sequence analyses relevant to today's plant biologist. Functional predictions, comparative analyses and associative annotations for 500 000 plant EST derived peptides make Sputnik (http://mips.gsf.de/proj/sputnik/) a valid platform for contemporary plant genomics.

Databases, Nucleic Acid↗

The Gene Resource Locator: gene locus maps for transcriptome analysis.

Since the advent of the draft human genome sequence there has been growing interest in transcriptome analysis based on genomic data. The Gene Resource Locator (GRL) assembles gene maps that include information on gene-expression patterns, cis-elements in regulatory regions and alternatively spliced transcripts. The database was constructed using customized software, and currently contains 2.2 million alignments (exon-intron structures). The alignments have been annotated and integrated into a system that encompasses approximately 90 000 EST loci sharing common exons, 8091 alternatively spliced transcript groups, 10 801 expression-profile groups, 8066 candidate regulatory regions in full-length cDNAs, and 1 million SNP loci. We have used Flash technology to build a dynamic web viewer that facilitates browsing through the millions of alignments. All of the information is available through the World Wide Web at the Gene Resource Locator web site (http://grl.gi.k.u-tokyo.ac.jp).

Alternative Splicing↗

Gene-dosage effect on chromosome 21 transcriptome in trisomy 21: implication in Down syndrome cognitive disorders.

In the era of human functional genomics, the chromosome 21 has represented a prototype for pioneering global biotechnologies. Its relatively low gene content enabled studying Down syndrome at the chromosomal scale, for which the last years have seen intense research activity aiming at genotype-phenotype correlations. The global gene-dose dependent upregulation of gene expression seen in the context of trisomy and preliminary functional annotation of chromosome 21 genes points towards candidate genes and molecular pathways potentially associated with the cognitive defects observed in Down syndrome.

Chromosome Mapping↗

Annotating the human proteome: beyond establishing a parts list.

The completion of the human genome has shifted the attention from deciphering the sequence to the identification and characterisation of the functional components, including genes. Improved gene prediction algorithms, together with the existing transcript and protein information, have enabled the identification of most exons in a genome. Availability of the 'parts list' has fostered the development of experimental approaches to systematically interrogate gene function on the genome, transcriptome and proteome level. Studying gene function at the protein level is vital to the understanding of how cells perform their functions as variations in protein isoforms and protein quantity which may underlie a change in phenotype can often not be deduced from sequence or transcript level genomics experiments alone. Recent advancements in proteomics have afforded technologies capable of measuring protein expression, post-translational modifications of these proteins, their subcellular localisation and assembly into complexes and pathways. Although an enormous amount of data already exists on the function of many human proteins, much of it is scattered over multiple resources. Public domain databases are therefore required to manage and collate this information and present it to the user community in both a human and machine readable manner. Of special importance here is the integration of heterogeneous data to facilitate the creation of resources that go beyond a mere parts list.

Humans↗

Transcriptome analysis of senescence in the flag leaf of wheat (Triticum aestivum L.).

The senescence process in wheat flag leaves was investigated over a time course from ear emergence until 50% yellowing of harvested leaf samples using an in-house fabricated cDNA microarray based on a 9K wheat unigene set. The top 1000 ranked differentially expressed probes were subjected to a cluster analysis and, from these, we selected 140 up-regulated genes with informative annotations. There was a considerable overlap between this list of genes and genes previously observed to be associated with senescence in other species, covering several functional categories involved in the degradation of macromolecules and nutrient remobilization, notably of nitrogen via the metabolism of carboxylic and amino acids. The up-regulation of a number of genes in this metabolism was confirmed by real-time polymerase chain reaction experiments. The data suggest a role for cytosolic/peroxisomal routes in the integration of the degradation of carbohydrates, fatty acids and proteins, leading to the remobilization of nitrogen. Illustrative examples of up-regulated genes comprise cytoplasmic aconitate hydratase and peroxisomal citrate synthase. The data support a protective role of the mitochondria towards oxidative cell damage via the up-regulation of the alternative oxidase, and possibly also involving the up-regulated succinate dehydrogenase. A number of up-regulated regulatory genes were also identified, notably NAC-domain and WRKY transcription factors. These factors have previously been identified as being associated with senescence in other species. The data support the notion that a generic senescence programme exists across monocot and dicot plant species. However, notable differences can also be recognized. We thus found transcriptional up-regulation of the biosynthetic pathway for benzoxazinoids, a group of graminaceous-specific secondary metabolites.

Aging↗

Functional Annotation of the Major Histocompatibility Complex Locus.

The human major histocompatibility complex (MHC) locus has the greatest density of disease-associations in the human genome, including links to over 100 polygenic disorders. Its complex haplotype structure, rich gene density, and high degree of linkage disequilibrium combine to make deciphering the gene regulatory logic of the MHC locus extremely challenging. Employing complementary high-throughput CRISPR interference (CRISPRi) and activation (CRISPRa) epigenetic screens coupled with single-cell transcriptome profiling across three distinct human cell types, we identified hundreds of new connections between cis -regulatory elements (CREs) and their target genes in this locus. These CRE-gene links are largely cell type-specific and act as enhancers. Additionally, some CREs have complex features, including harboring both active and repressive histone marks, lacking chromatin accessibility, targeting multiple genes, or acting as silencers. Computational methods fail to predict a majority of these CRE-gene connections. These findings emphasize the potential for functional perturbation experiments to dissect complex loci and reveal shared and cell type-specific regulatory mechanisms relevant to genomics of complex diseases. Collectively, this study provides a unique resource for understanding the complex regulatory landscape within the MHC locus and supports the need for creating new models that encompass CRE-gene interactions, cell type-specific gene expression, and disease genetics in the noncoding genome.

Journal Article↗

Benchmarking the CATMA microarray. A novel tool for Arabidopsis transcriptome analysis.

Transcript profiling is crucial to study biological systems, and various platforms have been implemented to survey mRNAs at the genome scale. We have assessed the performance of the CATMA microarray designed for Arabidopsis (Arabidopsis thaliana) transcriptome analysis and compared it with the Agilent and Affymetrix commercial platforms. The CATMA array consists of gene-specific sequence tags of 150 to 500 bp, the Agilent (Arabidopsis 2) array of 60mer oligonucleotides, and the Affymetrix gene chip (ATH1) of 25mer oligonucleotide sets. We have matched each probe repertoire with the Arabidopsis genome annotation (The Institute for Genomic Research release 5.0) and determined the correspondence between them. Array performance was analyzed by hybridization with labeled targets derived from eight RNA samples made of shoot total RNA spiked with a calibrated series of 14 control transcripts. CATMA arrays showed the largest dynamic range extending over three to four logs. Agilent and Affymetrix arrays displayed a narrower range, presumably because signal saturation occurred for transcripts at concentrations beyond 1,000 copies per cell. Sensitivity was comparable for all three platforms. For Affymetrix GeneChip data, the RMA software package outperformed Microarray Suite 5.0 for all investigated criteria, confirming that the information provided by the mismatch oligonucleotides has no added value. In addition, taking advantage of replicates in our dataset, we conducted a robust statistical analysis of the platform propensity to yield false positive and false negative differentially expressed genes, and all gave satisfactory results. The results establish the CATMA array as a mature alternative to the Affymetrix and Agilent platforms.

Arabidopsis↗

Modification of the transcriptomic response to renal ischemia/reperfusion injury by lipoxin analog.

BACKGROUND: Lipoxins are lipoxygenase-derived eicosanoids with anti-inflammatory and proresolution bioactivities in vitro and in vivo. We have previously demonstrated that the stable synthetic LXA4 analog 15-epi-16-(FPhO)-LXA4-Me is renoprotective in murine renal ischemia/reperfusion injury, as gauged by lower serum creatinine, attenuated leukocyte infiltration, and reduced morphologic tubule injury. METHODS: We employed complementary oligonucleotide microarray and bioinformatic analyses to probe the transcriptomic events that underpin lipoxin renoprotection in this setting. RESULTS: Microarray-based analysis identified three broad categories of genes whose mRNA levels are altered in response to ischemia/reperfusion injury, including known genes previously implicated in the pathogenesis of ischemia/reperfusion injury [e.g., intercellular adhesion molecule-1 (ICAM-1), p21, KIM-1], known genes not previously associated with ischemia/reperfusion injury, and cDNAs representing yet uncharacterized genes. Characterization of expressed sequence tags (ESTs) displayed on microarrays represents a major challenge in studies of global gene expression. A bioinformatic annotation pipeline successfully annotated a large proportion of ESTs modulated during ischemia/reperfusion injury. The differential expression of a representative group of these ischemia/reperfusion injury-modulated genes was confirmed by real-time polymerase chain reaction. Prominent among the up-regulated genes were claudin-1, -3, and -7, and ADAM8. Interestingly, the former response was claudin-specific and was not observed with other claudins expressed by the kidney (e.g., claudin-8 and -6) or indeed with other components of the renal tight junctions (e.g., occludin and junctional adhesion molecule). Noteworthy among the down-regulated genes was a cluster of transport proteins (e.g., aquaporin-1) and the zinc metalloendopeptidase meprin-1 beta implicated in renal remodeling. CONCLUSION: Treatment with the lipoxin analog 15-epi-16-(FPhO)-LXA4-Me prior to injury modified the expression of many differentially expressed pathogenic mediators, including cytokines, growth factors, adhesion molecules, and proteases, suggesting a renoprotective action at the core of the pathophysiology of acute renal failure (ARF). Importantly, this lipoxin-modulated transcriptomic response included many genes expressed by renal parenchymal cells and was not merely a reflection of a reduced renal mRNA load resulting from attenuated leukocyte recruitment. The data presented herein suggest a framework for understanding drivers of kidney injury in ischemia/reperfusion and the molecular basis for renoprotection by lipoxins in this setting.

ADAM Proteins↗

Genomic fossils as a snapshot of the human transcriptome.

Processed pseudogenes (PPGs) are cDNA sequences that were generated through reverse transcription of mature, spliced mRNAs and have subsequently been reinserted at a new genomic location. These cDNA sequences are usually no longer transcribed and are considered "dead on arrival." Here we show that PPGs can be used to generate a map of the transcriptome. By analyzing thousands of human PPGs, we were able to discover hundreds of transcript variants so far unidentified. An experimental verification of a subset of these variants by RT-PCR indicates that most of them are still active in the human transcriptome. Furthermore, we demonstrate that PPGs can enable the identification of ancient splice variants that were expressed ancestrally but are now extinct. Our results show that the genome itself carries a "virtual cDNA library" that can readily be used to analyze both present and ancestral transcripts. Our approach can be applied to sequenced metazoan genomes to computationally annotate splicing variation even when expressed sequences are unavailable.

Alternative Splicing↗

Human protein reference database--2006 update.

Human Protein Reference Database (HPRD) (http://www.hprd.org) was developed to serve as a comprehensive collection of protein features, post-translational modifications (PTMs) and protein-protein interactions. Since the original report, this database has increased to >20 000 proteins entries and has become the largest database for literature-derived protein-protein interactions (>30 000) and PTMs (>8000) for human proteins. We have also introduced several new features in HPRD including: (i) protein isoforms, (ii) enhanced search options, (iii) linking of pathway annotations and (iv) integration of a novel browser, GenProt Viewer (http://www.genprot.org), developed by us that allows integration of genomic and proteomic information. With the continued support and active participation by the biomedical community, we expect HPRD to become a unique source of curated information for the human proteome and spur biomedical discoveries based on integration of genomic, transcriptomic and proteomic data.

Databases, Protein↗

Genome-Wide Identification and Characterization of the TBL Gene Family and Temporal Expression Dynamics During Powdery Mildew Infection in Cucumber (Cucumis sativus).

Cell-wall polysaccharide O-acetylation contributes to cell-wall assembly, organ development, and plant-pathogen interactions, but the cucumber TBL gene family remains poorly characterized. Here, 37 CsTBL genes were identified genome-wide and analyzed using phylogenetic, syntenic, conserved-motif, gene-structure, promoter, protein-structure, Gene Ontology, and transcriptome approaches, followed by RT-qPCR analysis after powdery mildew inoculation. All CsTBL proteins contained the conserved GDS and DxxH motifs, whereas accessory motifs and predicted structural features varied among clades. Intraspecific analysis identified dispersed, WGD/segmental, and tandem duplication categories, and cross-species synteny was more extensive with melon than with Arabidopsis. Homology-derived annotations associated CsTBL genes with cell-wall polysaccharide metabolism, Golgi/endomembrane compartments, and O-acetyltransferase activity, including six genes assigned to xylan O-acetyltransferase-related annotations. Expression profiling revealed tissue- and developmental-stage-dependent patterns, whereas the publicly available powdery mildew RNA-seq dataset provided descriptive temporal expression profiles in Podosphaera xanthii-inoculated samples. Independent RT-qPCR analysis using time-matched mock controls revealed distinct post-inoculation responses among six selected genes. Relative to the corresponding mock controls, CsTBL2 was consistently repressed; CsTBL15 showed transient induction at 1 dpi followed by repression; CsTBL24 exhibited a biphasic response; CsTBL25 was induced at all sampled post-inoculation time points; CsTBL26 showed progressive induction; and CsTBL30 reached its highest observed expression level at 3 dpi. Integrated functional annotation and expression evidence highlighted CsTBL26 as a priority candidate for further functional characterization, while CsTBL24 and CsTBL25 represented fruit-associated candidates with distinct powdery mildew responses; CsTBL30 remained an additional strongly infection-responsive candidate. These findings provide an evolutionary and expression-based framework for the functional characterization of the cucumber TBL gene family.

O-acetylation↗

Transcription profiling of renal cell carcinoma.

AIMS: Our aim was to prepare a comprehensive catalogue of the changes in gene expression accompanying the development and progression of renal cell carcinoma, and to correlate these with histo-pathological, cytogenetic and clinical findings. METHODS: mRNA samples from paired neoplastic and non-cancerous human kidney tissue were labeled and hybridized in duplicate against high-density cDNA arrays. Two array technologies were used: 31,500-element transcriptome-wide nylon arrays for hybridization with 37 radioactively labelled sample pairs, and 4200-element kidney- and cancer-specific glass microarrays for hybridization with 19 fluorescently labelled sample pairs. RESULTS: We identified more than 1700 cDNA clones that show differential transcription levels in kidney tumor tissue compared to normal kidney tissue. The functional classification of 389 annotated genes provided views of the changes in the activities of specific biological processes in renal cancer. Among the biological processes with a large proportion of up-regulated genes we found cell adhesion, signal transduction, and nucleotide metabolism. Down-regulated processes included small molecule transport, ion homeostasis, and oxygen and radical metabolism. Furthermore, we explored the feasibility of molecular diagnosis for renal cell tumors using cDNA microarrays on glass slides, investigating the association of transcription levels with tumor type, progression, and a putative prognostic variable. The experimental data is available from the GEO gene expression database (http://www.ncbi.nlm.nih.gov/geo; accession no. GSE3), and a comprehensive presentation of the results is available in the web supplement (http://www.dkfz-heidelberg.de/abt0840/whuber/rcc). CONCLUSION: Transcription profiling using high-density cDNA arrays is a powerful method with the potential to improve cancer diagnosis and prognosis. The identification and classification of differentially transcribed genes, as described in our study, is the beginning of a more complete understanding of kidney cancer.

Carcinoma, Renal Cell↗

Development of a citrus genome-wide EST collection and cDNA microarray as resources for genomic studies.

A functional genomics project has been initiated to approach the molecular characterization of the main biological and agronomical traits of citrus. As a key part of this project, a citrus EST collection has been generated from 25 cDNA libraries covering different tissues, developmental stages and stress conditions. The collection includes a total of 22,635 high-quality ESTs, grouped in 11,836 putative unigenes, which represent at least one third of the estimated number of genes in the citrus genome. Functional annotation of unigenes which have Arabidopsis orthologues (68% of all unigenes) revealed gene representation in every major functional category, suggesting that a genome-wide EST collection was obtained. A Citrus clementina Hort. ex Tan. cv. Clemenules genomic library, that will contribute to further characterization of relevant genes, has also been constructed. To initiate the analysis of citrus transcriptome, we have developed a cDNA microarray containing 12,672 probes corresponding to 6875 putative unigenes of the collection. Technical characterization of the microarray showed high intra- and inter-array reproducibility, as well as a good range of sensitivity. We have also validated gene expression data achieved with this microarray through an independent technique such as RNA gel blot analysis.

Citrus↗