Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

388 records · Page 22Linked to original sources

[Transcriptomes for serial analysis of gene expression].

The availability of the sequences for whole genomes is changing our understanding of cell biology. Functional genomics refers to the comprehensive analysis, at the protein level (proteome) and at the mRNA level (transcriptome) of all events associated with the expression of whole sets of genes. New methods have been developed for transcriptome analysis. Serial Analysis of Gene Expression (SAGE) is based on the massive sequential analysis of short cDNA sequence tags. Each tag is derived from a defined position within a transcript. Its size (14 bp) is sufficient to identify the corresponding gene and the number of times each tag is observed provides an accurate measurement of its expression level. Since tag populations can be widely amplified without altering their relative proportions, SAGE may be performed with minute amounts of biological extract. Dealing with the mass of data generated by SAGE necessitates computer analysis. A software is required to automatically detect and count tags from sequence files. Criterias allowing to assess the quality of experimental data can be included at this stage. To identify the corresponding genes, a database is created registering all virtual tags susceptible to be observed, based on the present status of the genome knowledge. By using currently available database functions, it is easy to match experimental and virtual tags, thus generating a new database registering identified tags, together with their expression levels. As an open system, SAGE is able to reveal new, yet unknown, transcripts. Their identification will become increasingly easier with the progress of genome annotation. However, their direct characterization can be attempted, since tag information may be sufficient to design primers allowing to extend unknown sequences. A major advantage of SAGE is that, by measuring expression levels without reference to an arbitrary standard, data are definitively acquired and cumulative. All publicly available data can thus be stored in a unique database, facilitating whole-genome analysis of differential expression between cell types, normal and diseased samples, or samples with and without drug treatment. SAGE data are readily amenable to statistical comparisons, allowing to determine the level of confidence of the observed variations. A major limitation of SAGE is that, because each analysis is obligatory performed on the whole set of expressed genes, it can hardly be performed on multiple samples, for example in kinetics studies or to compare the effects of large numbers of drugs. To overcome this limitation, high-throughput detection of a subset of mRNAs is more rapidly performed by parallel hybridization of mRNAs on arrays of nucleic acids immobilized on solid supports. From this point of view, a SAGE platform is a powerful instrument for selecting the most informative subset of genes, assembling them to design microarrays dedicated to a specific problem and calibrating measurement by comparison with a standard cell model for which SAGE data are available. This approach is an attractive alternative to strategies based exclusively on pangenomic arrays. A very large amount of SAGE data are already available and the problem is now to extract their biological meaning. Knowledge on metabolic pathways is already organized so that its successful integration in a SAGE platform can be undertaken. For other cell components and pathways, the problem lies on the lack of controlled vocabulary to describe gene activities, starting form a clear definition of the concept of biological function itself. Progress in gene and cell ontology is expected to facilitate computer-based extraction of biological knowledge from existing and forthcoming SAGE data.

Animals↗

Poplar carbohydrate-active enzymes. Gene identification and expression analyses.

Over 1,600 genes encoding carbohydrate-active enzymes (CAZymes) in the Populus trichocarpa (Torr. & Gray) genome were identified based on sequence homology, annotated, and grouped into families of glycosyltransferases, glycoside hydrolases, carbohydrate esterases, polysaccharide lyases, and expansins. Poplar (Populus spp.) had approximately 1.6 times more CAZyme genes than Arabidopsis (Arabidopsis thaliana). Whereas most families were proportionally increased, xylan and pectin-related families were underrepresented and the GT1 family of secondary metabolite-glycosylating enzymes was overrepresented in poplar. CAZyme gene expression in poplar was analyzed using a collection of 100,000 expressed sequence tags from 17 different tissues and compared to microarray data for poplar and Arabidopsis. Expression of genes involved in pectin and hemicellulose metabolism was detected in all tissues, indicating a constant maintenance of transcripts encoding enzymes remodeling the cell wall matrix. The most abundant transcripts encoded sucrose synthases that were specifically expressed in wood-forming tissues along with cellulose synthase and homologs of KORRIGAN and ELP1. Woody tissues were the richest source of various other CAZyme transcripts, demonstrating the importance of this group of enzymes for xylogenesis. In contrast, there was little expression of genes related to starch metabolism during wood formation, consistent with the preferential flux of carbon to cell wall biosynthesis. Seasonally dormant meristems of poplar showed a high prevalence of transcripts related to starch metabolism and surprisingly retained transcripts of some cell wall synthesis enzymes. The data showed profound changes in CAZyme transcriptomes in different poplar tissues and pointed to some key differences in CAZyme genes and their regulation between herbaceous and woody plants.

Arabidopsis↗

Characterization of a normalized cDNA library from bovine intestinal muscle and epithelial tissues.

Tissue-specific cDNA library sequences (expressed sequence tags, or EST) yield a detailed snapshot of gene expression and are useful in developing second-generation molecular resources (i.e., microarrays) for gene expression profiling. The objective of this study was to develop and characterize an intestine-specific cDNA library to examine the transcriptome of the bovine gut and identify expressed genes that influence ruminant nutrition and health. We describe BARC-8BOV, a normalized cDNA library developed from mRNA isolated from four distinct intestinal locations (duodenal, jejunal and ileal small intestine, colon) of Holstein dairy cattle resulting in 19,110 5'-EST deposited into the NCBI GenBank EST database. Assembly and clustering of these 19,110 clone sequences yielded 11,208 unique elements (3,419 contigs and 7,789 singletons) with an average length of 695 base pairs. Analysis strongly suggests normalization and tissue pooling were effective at increasing the discovery rate of new bovine sequence. A total of 1,123 sequence elements not previously identified in cattle, but with similarity to known genes in other animal species, were identified and shown to be involved in numerous critical biological processes. An additional 745 transcripts were not previously represented as EST in nucleotide or protein databases, and further analysis of these could lead to the identification of gut-specific transcript variants of known genes or potentially the discovery of novel bovine genes. Of the 11,208 assembled sequences, 11,034, or 98.4%, match sequences present in the bovine DNA trace archive at NCBI, and add to a bovine EST database previously lacking significant gut tissue representation. Ultimately, these data will also contribute in efforts to annotate the bovine genome.

Animals↗

Shared genetic architecture of smoking dependence and Crohn's disease: A cross-trait analysis of GWAS summary statistics.

INTRODUCTION: Smoking dependence (SD) and Crohn's disease (CD) are epidemiologically associated, but whether this relationship reflects shared genetic susceptibility remains unclear. METHODS: We conducted a cross-trait genetic analysis of SD and CD using publicly available genome-wide association study (GWAS) summary statistics from European-ancestry populations. Genome-wide genetic correlation was estimated using linkage disequilibrium score regression (LDSC) and high-definition likelihood (HDL). Pleiotropic variants were identified using PLACO and mapped to genomic loci using FUMA. Regional signal sharing was assessed by Bayesian colocalization. Functional analyses included stratified LDSC, Multi-marker Analysis of GenoMic Annotation (MAGMA), GTEx tissue analysis, and Metascape. Expression-linked candidate genes were prioritized using expression quantitative trait locus (eQTL)-based summary-data-based Mendelian randomization (SMR) with heterogeneity in dependent instruments (HEIDI) testing. Genetically informed spatial mapping of cells for complex traits (gsMap) was used for spatial mapping. RESULTS: SD and CD showed positive genetic correlation by LDSC (rg=0.2090, p=0.0008) and HDL (rg=0.3817, p=0.00106). PLACO identified 81 genome-wide significant pleiotropic SNPs, which were mapped by FUMA to three loci at 1p31.3, 5p13.1, and 12q12, represented by rs11209031, rs1395152, and rs17467116, respectively. MAGMA identified 22 FDR-significant genes, four of which remained Bonferroni significant: LRRK2, TNFRSF6B, ZGPAT, and RP4-583P15.15. Cross-trait tissue analysis showed significant enrichment of the shared genetic signal in whole blood and small intestine, while gene-set analysis highlighted inflammatory response (pbon=1.86×10-5) and T-helper 17 cell differentiation (pbon=7.37×10-4). SMR/HEIDI analysis further prioritized RPS6KB1 as a shared expression-linked candidate. Spatial mapping revealed a prominent signal in the embryonic gastrointestinal tract and gene-specific regional patterns involving LRRK2 and SLC2A13 in the adult mouse brain. CONCLUSIONS: SD and CD showed measurable shared genetic susceptibility, with convergent evidence from pleiotropic loci, immune-inflammatory pathway enrichment, tissue-level associations, and spatial transcriptomic mapping.

Crohn's disease↗

In silico analysis of SH3BP2 genomic alterations and expression profiles in CRC.

AIM: Colorectal cancer (CRC) is a widespread health issue that attains high mortality. The adaptor protein SH3BP2 amplification results in metabolic changes, oxidative stress, NK cell activity, and inflammation. The NK cells are capable of destroying tumor cells without prior activation, help prevent metastasis, and have prognostic value. Targeting SH3BP2 to regulate NK cell activity in the TME could enhance CRC-based immunotherapy. MATERIALS AND METHODS: The cancer hallmark tool helps in understanding SH3BP2 hallmark annotation. Utilizing the STRING tool and the KEGG pathway, protein functional enrichment and PPI networking were analyzed. TIMER 2.0 was used for immune cell infiltration correlation analysis, and UALCAN was used for CPTAC-based protein expression profiling. RESULTS AND CONCLUSIONS: The GEO (GSE9348) dataset showed SH3BP2 is upregulated in CRC (log2 fold change = 1.18). GEO, TCGA, and cBioPortal revealed SH3BP2 alterations in CRC cases, potentially aiding immune evasion. Mutations in SH3BP2 influence cancer growth, suppressing tumors or promoting them by activating NF-κB and affecting immune responses through WNT/β-catenin, PI3K, MAPK, and JAK-STAT pathways. Overall, SH3BP2 plays a key role in cancer growth and immune regulation, making it a promising target for CRC therapy. Further experimental validation is needed to demonstrate its diagnostic and therapeutic potency.

Humans↗

Hierarchical modeling of tumor subtypes in cell lines using large-scale genomic datasets.

Cancer cell lines (CLs) are widely used to study tumor biology and drug response, yet their translational relevance is often limited by inaccurate subtype annotations. Existing CL-tumor matching approaches are frequently constrained by flat classification schemes, weak subtype definitions, and the exclusion of normal tissue references, leading to potential confounding of tumor-specific and tissue-of-origin signals. To address these limitations, a hierarchical classification (HC) framework is presented in which CLs are aligned with patient tumors across biological resolutions, from organ to molecular subtype. Gene expression profiles from 802 CLs, 5,612 tumors from The Cancer Genome Atlas (TCGA) , and 8,939 non-cancerous tissues were integrated to separate oncogenic signals from tissue-specific signals. Node-specific features were selected using maximum relevance minimum redundancy, and balanced accuracies of 89% in cross-validation and 75%, and 80% on external datasets were achieved. Through the framework, 43 CLs were reassigned, and clinically relevant underrepresented subtypes were identified.

cancer cell lines↗

Real-time RT-PCR profiling of over 1400 Arabidopsis transcription factors: unprecedented sensitivity reveals novel root- and shoot-specific genes.

Summary To overcome the detection limits inherent to DNA array-based methods of transcriptome analysis, we developed a real-time reverse transcription (RT)-PCR-based resource for quantitative measurement of transcripts for 1465 Arabidopsis transcription factors (TFs). Using closely spaced gene-specific primer pairs and SYBR Green to monitor amplification of double-stranded DNA (dsDNA), transcript levels of 83% of all target genes could be measured in roots or shoots of young Arabidopsis wild-type plants. Only 4% of reactions produced non-specific PCR products. The amplification efficiency of each PCR was determined from the log slope of SYBR Green fluorescence versus cycle number in the exponential phase, and was used to correct the readout for each primer pair and run. Measurements of transcript abundance were quantitative over six orders of magnitude, with a detection limit equivalent to one transcript molecule in 1000 cells. Transcript levels for different TF genes ranged between 0.001 and 100 copies per cell. Only 13% of TF transcripts were undetectable in these organs. For comparison, 22K Arabidopsis Affymetrix chips detected less than 55% of TF transcripts in the same samples, the range of transcript levels was compressed by a factor more than 100, and the data were less accurate especially in the lower part of the response range. Real-time RT-PCR revealed 35 root-specific and 52 shoot-specific TF genes, most of which have not been identified as organ-specific previously. Finally, many of the TF transcripts detected by RT-PCR are not represented in Arabidopsis EST (expressed sequence tag) or Massively Parallel Signature Sequencing (MPSS) databases. These genes can now be annotated as expressed.

Arabidopsis↗

Multimodal profiling reveals tissue-directed signatures of human immune cells altered with age.

The immune system comprises multiple cell lineages and subsets maintained in tissues throughout the lifespan, with unknown effects of tissue and age on immune cell function. Here we comprehensively profiled RNA and surface protein expression of over 1.25 million immune cells from blood and lymphoid and mucosal tissues from 24 organ donors aged 20-75 years. We annotated major lineages (T cells, B cells, innate lymphoid cells and myeloid cells) and corresponding subsets using a multimodal classifier and probabilistic modeling for comparison across tissue sites and age. We identified dominant site-specific effects on immune cell composition and function across lineages; age-associated effects were manifested by site and lineage for macrophages in mucosal sites, B cells in lymphoid organs, and circulating T cells and natural killer cells across blood and tissues. Our results reveal tissue-specific signatures of immune homeostasis throughout the body, from which to define immune pathologies across the human lifespan.

Humans↗

Large-scale transcriptome analyses reveal new genetic marker candidates of head, neck, and thyroid cancer.

A detailed genome mapping analysis of 213,636 expressed sequence tags (EST) derived from nontumor and tumor tissues of the oral cavity, larynx, pharynx, and thyroid was done. Transcripts matching known human genes were identified; potential new splice variants were flagged and subjected to manual curation, pointing to 788 putatively new alternative splicing isoforms, the majority (75%) being insertion events. A subset of 34 new splicing isoforms (5% of 788 events) was selected and 23 (68%) were confirmed by reverse transcription-PCR and DNA sequencing. Putative new genes were revealed, including six transcripts mapped to well-studied chromosomes such as 22, as well as transcripts that mapped to 253 intergenic regions. In addition, 2,251 noncoding intronic RNAs, eventually involved in transcriptional regulation, were found. A set of 250 candidate markers for loss of heterozygosis or gene amplification was selected by identifying transcripts that mapped to genomic regions previously known to be frequently amplified or deleted in head, neck, and thyroid tumors. Three of these markers were evaluated by quantitative reverse transcription-PCR in an independent set of individual samples. Along with detailed clinical data about tumor origin, the information reported here is now publicly available on a dedicated Web site as a resource for further biological investigation. This first in silico reconstruction of the head, neck, and thyroid transcriptomes points to a wealth of new candidate markers that can be used for future studies on the molecular basis of these tumors. Similar analysis is warranted for a number of other tumors for which large EST data sets are available.

Alternative Splicing↗

Sodium Overload-Related Molecular Subtypes and a Four-Gene Prognostic Signature Predict Survival, Immune Landscape, and Therapeutic Response in Acute Myeloid Leukemia.

Sodium overload has recently emerged as a critical metabolic stressor involved in cancer progression; however, its molecular characteristics and clinical relevance in acute myeloid leukemia (AML) remain unexplored. RNA-seq data sets, clinical annotations, and mutational profiles of AML patients were annotations from The Cancer Genome Atlas and integrated with Genotype-Tissue Expression normal samples. Sodium overload-related genes (SORGs) were obtained from GeneCards. Differentially expressed SORGs (DESORGs) screened by applying the limma statistical model, followed by univariate Cox proportional hazards regression, consensus clustering, functional enrichment, immune infiltration analysis, and pathway evaluation. A prognostic signature was developed through least absolute shrinkage and selection operator regression followed by multivariate Cox modeling. The model's performance was further verified in two external GEO data sets (GSE71014 and GSE37642). Nomogram construction, subgroup analysis, tumor mutational burden (TMB) assessment, drug sensitivity prediction, transcription factor (TF) analysis, and competing endogenous RNA (ceRNA) network analyses were also performed. A total of 57 DESORGs were identified, and 2 sodium overload-related molecular subtypes exhibited distinct survival, immune infiltration, and inflammatory pathway activation. A robust four-gene signature (DOCK1, GABRE, HTR7, ACSM1) stratified patients into high- and low-risk categories with significantly different survival across training and validation cohorts. High-risk patients displayed increased immune infiltration, higher TMB, reduced sensitivity to multiple chemotherapeutic drugs, and inferior predicted response to PD-L1 blockade. TF and ceRNA networks revealed multilayered transcriptional and post-transcriptional regulation of the signature genes. This study identifies sodium overload-related molecular heterogeneity in AML and establishes a validated four-gene prognostic signature that integrates genomic, immunologic, and therapeutic features, offering potential utility for personalized risk assessment and treatment optimization.

Humans↗