Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Optimized protocol for linear RNA amplification and application to gene expression profiling of human renal biopsies.

Gene expression analysis using high-density cDNA or oligonucleotide arrays is a rapidly emerging tool for transcriptomics, the analysis of the transcriptional state of a cell or organ. One of the limitations of current methodologies is the requirement of a relatively large amount of total or polyadenylated RNA as starting material. Standard array hybridization protocols require 5-15 micrograms labeled RNA. To obtain these quantities from small amounts of starting RNA material, RNA can be amplified in a linear fashion. Here we introduce an optimized protocol for rapid and easy-to-use amplification of as little as 1 ng total RNA. Our analysis shows that this method is linear and highly reproducible and that it preserves similarities as well as dissimilarities between normal and disease-related samples. We applied this technique to the RNA expression profiling of human renal allograft biopsies with normal histology and compared them to the profiles of renal biopsies with histological evidence of chronic transplant nephropathy or chronic rejection. Among others, complement component C1r was found to be significantly up-regulated in chronic rejection and chronic transplant nephropathy biopsies compared to normal samples, while fructose-1,6-biphosphatase showed lower-than-normal expression.

Gene Expression Profiling↗

Multi-omics integration uncovers epigenetic control of metabolic reprogramming in triple-negative breast cancer.

Triple-negative breast cancer (TNBC) is an aggressive subtype characterized by the absence of estrogen, progesterone, and HER2 receptors, limiting effective targeted therapies. Increasing evidence suggests that metabolic reprogramming, a hallmark of TNBC progression, is driven by underlying epigenetic mechanisms such as DNA methylation. The represented study performed an integrative analysis of transcriptomic (RNA-seq) and methylome data to uncover the metabolic-epigenetic interplay in TNBC. Differential gene expression analysis using DESeq2 revealed significant dysregulation of key metabolic genes, including upregulation of genes encoding glycolytic and serine biosynthesis enzymes and downregulation of metabolic tumor suppressors. Genome-wide methylation profiling identified extensive cytosine-phosphate-guanine (CpG) hypermethylation events associated with transcriptional repression, particularly in promoter regions. Integrative analysis pinpointed a subset of metabolism-related genes exhibiting both differential expression and methylation, such as FBP1, RASSF1A, and PHGDH. Pathway enrichment analysis highlighted aberrations in glycolysis/gluconeogenesis, fatty acid metabolism, and one-carbon pathways (adjusted p&#x2009;<&#x2009;0.01). Importantly, TNBC patients with hypermethylated metabolic gene signatures displayed significantly shorter overall survival (log-rank p&#x2009;<&#x2009;0.05). These findings reveal that DNA methylation-driven metabolic dysregulation contributes to TNBC aggressiveness and may provide novel biomarkers and therapeutic targets at the metabolic-epigenetic interface.

Humans↗

Bacillus subtilis functional genomics: genome-wide analysis of the DegS-DegU regulon by transcriptomics and proteomics.

The DegS-DegU two-component regulatory system of Bacillus subtilis controls various processes that characterize the transition from the exponential to the stationary growth phase, including the induction of extracellular degradative enzymes, expression of late competence genes and down-regulation of the sigma(D) regulon. The degU32(Hy) mutation stabilizes the phosphorylated form of DegU (DegU-P), resulting in overproduction of several extracellular degradative enzymes. In this study, the pleiotropic DegS-DegU regulon was characterized by combining proteomic and transcriptomic approaches. A comparative analysis of wild-type B. subtilis and the degU32(Hy) mutant grown in complex medium was performed during the exponential and in the stationary growth phase. Besides genes already known to be under the control of DegU-P, novel putative members of this regulon were identified. Although the degU32(Hy) mutant is assumed to contain high levels of phosphorylated DegU in the exponential as well as in the stationary growth phase, many genes known to be positively regulated by DegU-P did not show enhanced expression in the mutant strain during exponential growth. This is consistent with the fact that most genes belonging to the DegS-DegU regulon are subject to multiple regulation; this is also reflected in the strong stationary-phase induction of these genes in the mutant strain. As expected, during the exponential growth phase, the sigma(D) regulon was expressed at significantly lower levels in the degU32(Hy) mutant than in the wild type.

Bacillus subtilis↗

Gene family analysis of the Arabidopsis pollen transcriptome reveals biological implications for cell growth, division control, and gene expression regulation.

Upon germination, pollen forms a tube that elongates dramatically through female tissues to reach and fertilize ovules. While essential for the life cycle of higher plants, the genetic basis underlying most of the process is not well understood. We previously used a combination of flow cytometry sorting of viable hydrated pollen grains and GeneChip array analysis of one-third of the Arabidopsis (Arabidopsis thaliana) genome to define a first overview of the pollen transcriptome. We now extend that study to approximately 80% of the genome of Arabidopsis by using Affymetrix Arabidopsis ATH1 arrays and perform comparative analysis of gene family and gene ontology representation in the transcriptome of pollen and vegetative tissues. Pollen grains have a smaller and overall unique transcriptome (6,587 genes expressed) with greater proportions of selectively expressed (11%) and enriched (26%) genes than any vegetative tissue. Relative gene ontology category representations in pollen and vegetative tissues reveal a functional skew of the pollen transcriptome toward signaling, vesicle transport, and the cytoskeleton, suggestive of a commitment to germination and tube growth. Cell cycle analysis reveals an accumulation of G2/M-associated factors that may play a role in the first mitotic division of the zygote. Despite the relative underrepresentation of transcription-associated transcripts, nonclassical MADS box genes emerge as a class with putative unique roles in pollen. The singularity of gene expression control in mature pollen grains is further highlighted by the apparent absence of small RNA pathway components.

Arabidopsis↗

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis↗

Genome-wide identification of potassium transporters and channels in Malus domestica genome.

Potassium (K+) is an essential nutrient for plants. It contributes to most physiological and biochemical pathways for plant metabolism, growth, and development. It is the most available plant nutrient, comprising 10&#x2013;15% of plant weight. Plants have a sophisticated system of K+ transporters and channels for distribution in plant body. Apple is one of the most consumed fruits in the world. Its fruit quality and yield are positively affected by K+. However, limited information is available about K+ transport systems in Apple. In this study, 47 candidate genes (26&#xa0;K+ transporters and 21&#xa0;K+ channels) have been identified in Apple (Malus domestica) genome. The phylogenetic comparisons with other plants (Glycine max, Arabidopsis thaliana, and Oryza sativa) indicated that the K+ transport system is much conserved among different plants. The analysis of Gene structure showed the presence of specific introns and exon patterns for these gene families. Transcriptomic data analysis and RT-qPCR demonstrated significant variations in the transcript abundance of these genes in response to abiotic stresses. The current project represents the first report about the K+ transport system in Apple. Therefore, it may act as a starting point for further functional characterizations.

Malus↗

Integrative cross-tissue transcriptome-wide association and metabolomic analysis reveals novel genetic risk loci for aortic aneurysm.

BACKGROUND: Aortic aneurysm (AA) is a life-threatening cardiovascular condition with a strong genetic component, however, its molecular mechanisms remain poorly understood. Although genome-wide association studies (GWAS) have identified numerous risk loci, most prior studies have investigated genetic and metabolic factors separately, leaving the causal pathways from genetic variants to disease largely unexplored. METHODS: We established an integrative framework combining cross-tissue transcriptome-wide association studies (TWAS) with metabolomic mediation analysis. First, we integrated GWAS data from FinnGen R12 with multi-tissue expression quantitative trait loci (eQTL) data from Genotype-Tissue Expression Project (GTEx) V8, then performed cross-tissue TWAS using the Unified Test for MOlecular SignaTures (UTMOST) and single-tissue validation with the Functional Summary-based Imputation (FUSION) to prioritize susceptibility genes. Second, we applied Mendelian randomization (MR), colocalization, and Fine-mapping Of CaUsal gene Sets (FOCUS) to assess causality and identify high-confidence genes. Third, we performed metabolite mediation analysis to uncover metabolic pathways linking genetic variants to disease risk. Finally, we validated key findings in mouse models of thoracic aortic aneurysm (TAA) and abdominal aortic aneurysm (AAA) using Quantitative Real-Time Reverse Transcription Polymerase Chain Reaction (RT-qPCR) and Western blotting. RESULTS: We identified multiple novel susceptibility genes for AA and its subtypes. Key genes included ADH family members (ADH1A, ADH1B, ADH4, ADH6) and ZNF827, which showed cross-subtype associations with strong colocalization evidence in vascular tissues. Metabolite mediation analysis revealed significant pathways involving N-acetylphenylalanine and methionine sulfoxide. Functional enrichment revealed distinct biological mechanisms: AA and AAA were primarily associated with metabolic pathways, whereas TAA-related genes were enriched in developmental and contractile processes. PheWAS indicated no significant off-target associations. Critically, experimental validation in mouse models confirmed significant upregulation of ZNF827 in TAA and ADH6 in AAA at both mRNA and protein levels, corroborating the genetic predictions. CONCLUSION: This integrated cross-omics analysis identifies novel genetic loci and, crucially, uncovers specific nutrient-related metabolic pathways that mediate genetic risk. These findings provide a mechanistic basis for future nutritional and metabolic intervention studies in AA and its subtypes.

MAGMA↗

Visualizing chromosomes as transcriptome correlation maps: evidence of chromosomal domains containing co-expressed genes--a study of 130 invasive ductal breast carcinomas.

Completion of the working draft of the human genome has made it possible to analyze the expression of genes according to their position on the chromosomes. Here, we used a transcriptome data analysis approach involving for each gene the calculation of the correlation between its expression profile and those of its neighbors. We used the U133 Affymetrix transcriptome data set for a series of 130 invasive ductal breast carcinomas to construct chromosomal maps of gene expression correlation (transcriptome correlation map). This highlighted nonrandom clusters of genes along the genome with correlated expression in tumors. Some of the gene clusters identified by this method probably arose because of genetic alterations, as most of the chromosomes with the highest percentage of correlated genes (1q, 8p, 8q, 16p, 16q, 17q, and 20q) were also the most frequent sites of genomic alterations in breast cancer. Our analysis showed that several known breast tumor amplicons (at 8p11-p12, 11q13, and 17q12) are located within clusters of genes with correlated expression. Using hierarchical clustering on samples and a Treeview representation of whole chromosome arms, we observed a higher-order organization of correlated genes, sometimes involving very large chromosomal domains that could extend to a whole chromosome arm. Transcription correlation maps are a new way of visualizing transcriptome data. They will help to identify new genes involved in tumor progression and new mechanisms of gene regulation in tumors.

Breast Neoplasms↗

Single-Cell Transcriptome-Wide Mendelian Randomization and Colocalization Uncover Potential Immunocytes-Related Therapeutic Targets for Obesity.

Weight-loss treatment is crucial for individuals with obesity to prevent various complications. The role of Immune cells in obesity has been recently recognized, whereas its translation into therapy requires identifying key target genes. We performed Mendelian randomization (MR) analysis to assess causal relationships between expression quantitative trait loci (eQTL) of 14 immune cells and obesity-related traits (obesity, body mass index and body fat percentage), and validated the results in colocalization analysis. For the putative causal genes identified by the MR and colocalization analyses, we conducted pathway enrichment, differential expressed gene (DEG) analysis and search of druggable evidence, and utilized a Tier system to prioritize drug targets for obesity. MR and colocalization evidence was observed for 1630 genes associated with one or more obesity-related traits, mainly expressed in CD4+ naive/central memory T cells and enriched in antigen processing and presentation pathways. Forty-one genes showed causal relationship with all three outcomes, among which 19 genes have not been reported for obesity previously. DEG analysis using single-cell RNA sequencing data of blood or adipose tissue indicated that the differential expression of UBE2Z in monocytes, ZCCHC7 in T cells, and FNBP4 in B cells between lean and obese individuals were consistent with the MR results. By searching drug-gene interaction databases, we found targeted drugs for PYGB and PRUNE1, and PYGB was the top gene ranked in the Tier system. This study provides evidence for the involvement of immune cells in obesity, and the potential cell-specific, immune-related targets for obesity treatment.

Obesity↗

An immune exhaustion signature predicts prognosis and identifies patients with diffuse large B-cell lymphoma (DLBCL) who derive preferential benefit from chimeric antigen receptor (CAR)-T cell therapy.

BACKGROUND: The tumor microenvironment (TME) is a key determinant of prognosis in diffuse large B-cell lymphoma (DLBCL). While T-cell exhaustion is implicated in therapeutic failure, its precise molecular hallmarks and utility for predicting response to modern immunotherapies, such as chimeric antigen receptor (CAR)-T cell therapy, remain unclear. METHODS: We performed an integrative analysis of transcriptomic and clinical data from multiple DLBCL cohorts (The Cancer Genome Atlas [TCGA], GSE181063, GSE10846, GSE248835, GSE182434). We used unsupervised clustering, exploratory analysis of single-cell RNA sequencing data, and the least absolute shrinkage and selection operator for variable selection (LASSO-Cox) regression to characterize the exhausted TME, construct a prognostic model, and evaluate its predictive value for CAR-T cell therapy. The model's dynamic behavior was assessed in a proof-of-concept longitudinal cohort of patients treated with the T-cell-engaging bispecific antibody glofitamab. RESULTS: We identified a "high-exhaustion" subtype associated with significantly poorer overall survival (OS; log-rank P = 0.016). Based on this, we developed a five-gene immune exhaustion-Related Prognostic Score (IERPS) that served as a robust independent predictor of poor OS across multiple cohorts. Critically, in a cohort of 256 relapsed/refractory patients, the IERPS was strongly prognostic for event-free survival (EFS) in the standard-of-care (SOC) arm (HR = 2.02, 95% confidence interval [95% CI]: 1.07-3.81, P = 0.029) but lost prognostic significance in the CAR-T arm (HR = 0.70, 95 % CI: 0.35-1.40, P = 0.314). This significant interaction suggests that CAR-T cell therapy may abrogate the poor prognosis associated with a high IERPS. Biologically, exploratory single-cell analysis (n = 4 samples) defined the high-IERPS state by hallmarks of classical T-cell exhaustion, and a descriptive case study showed the score dynamically tracked clinical response to glofitamab. CONCLUSIONS: A state of active T-cell exhaustion and a suppressive TME drive the adverse immune phenotype in DLBCL. Our IERPS model captures this dysfunctional state, acting as a powerful prognostic tool and, more importantly, as a potential predictive biomarker to identify high-risk patients who appear to overcome their inherently poor prognosis through CAR-T cell therapy.

Biomarkers↗

Enhancing and accelerating cell type deconvolution of large-scale spatial transcriptomics slices with dual network model.

MOTIVATION: Cell type deconvolution deciphers spatial distribution of mRNA transcripts at single cell level by integrating single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics data to infer mixture of cell types of spots in slices. Current algorithms are criticized for neglecting connection between scRNA-seq and spatial transcriptomics data, as well as time-consuming, hampering their application to large-scale datasets. RESULTS: In this study, we propose a joint learning nonnegative matrix factorization algorithm for fast cell type deconvolution (aka jMF2D), which integrates scRNA-seq and spatial transcriptomics data with network models. To bridge scRNA-seq and spatial transcriptomics data, jMF2D jointly learns cell type similarity network to enhance quality of signatures of cell types, thereby promoting accuracy and efficiency of deconvolution. Experiments demonstrate that jMF2D outperforms state-of-the-art baselines in terms of accuracy by saving about 90% running time on various datasets generated by different platforms. Furthermore, it can also facilitates the identification of spatial domains and bio-marker genes, providing an efficient and effective model for analyzing spatial transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The software is coded using python, and is free available for academic https://github.com/xkmaxidian/jMF2D.

Algorithms↗

Applications for microarrays in renal biology and medicine.

Groundbreaking recent developments, such as the near completion of human and mouse genome sequencing efforts and the emergence of robust microarray (gene chip) technologies, enabling comprehensive analysis of transcriptomes, provide new opportunities of unprecedented scale for researchers of kidney biology and disease. Combined with advanced computational and mathematical approaches for microarray data analysis, microarray applications promise to revolutionize our understanding of molecular mechanisms of kidney development and renal pathogenesis. New knowledge in this field will facilitate new approaches for molecular diagnostics, drug discovery, and eventually "personalized" renal medicine. In this review, we outline current and future research applications of microarray and computational approaches in renal biology and disease. We describe basic steps in microarray data analysis and introduce advanced computational approaches to optimize data mining of vast microarray datasets.

Genomics↗

Pitavastatin, Procollagen Pathways, and Plaque Stabilization in Patients With HIV: A Secondary Analysis of the REPRIEVE Randomized Clinical Trial.

IMPORTANCE: In a mechanistic substudy of the Randomized Trial to Prevent Vascular Events in HIV (REPRIEVE) randomized clinical trial, pitavastatin reduced noncalcified plaque (NCP) volume, but specific protein and gene pathways contributing to changes in coronary plaque remain unknown. OBJECTIVE: To use targeted discovery proteomics and transcriptomics approaches to interrogate biological pathways beyond low-density lipoprotein cholesterol (LDL-C), relating statin outcomes to reduce NCP volume and promote plaque stabilization among people with HIV (PWH). DESIGN, SETTING, AND PARTICIPANTS: This was a post hoc analysis of the double-blind, placebo-controlled, REPRIEVE randomized clinical trial. Participants underwent coronary computed tomography angiography (CTA), plasma protein analysis, and transcriptomic analysis at baseline and 2-year follow-up. The trial enrolled PWH from April 2015 to February 2018 at 31 US research sites. PWH without known cardiovascular diseases taking antiretroviral therapy and with low to moderate 10-year cardiovascular risk were eligible. Data analyses were conducted from October 2023 to February 2024. INTERVENTION: Oral pitavastatin calcium, 4 mg per day. MAIN OUTCOMES AND MEASURES: Relative change in plasma proteomics, transcriptomics, and noncalcified plaque volume among those receiving treatment vs placebo. RESULTS: Among 558 individuals (mean [SD] age, 51 [6] years; 455 male [82%]) included in the proteomics assessment, 272 (48.7%) received pitavastatin and 286 (51.3%) received placebo. After adjusting for false discovery rates, pitavastatin increased abundance of procollagen C-endopeptidase enhancer 1 (PCOLCE), neuropilin 1 (NRP-1), major histocompatibility complex class I polypeptide-related sequence A (MIC-A) and B (MIC-B), and decreased abundance of tissue factor pathway inhibitor (TFPI), tumor necrosis factor ligand superfamily member 10 (TRAIL), angiopoietin-related protein 3 (ANGPTL3), and mannose-binding protein C (MBL2). Among these proteins, the association of pitavastatin with PCOLCE (a rate-limiting enzyme of collagen deposition) was greatest, with an effect size of 24.3% (95% CI, 18.0%-30.8%; P&#x2009;<&#x2009;.001). In a transcriptomic analysis, individual collagen genes and collagen gene sets showed increased expression. Among the 195 individuals with plaque at baseline (88 [45.1%] taking pitavastatin, 107 [54.9%] taking placebo), changes in NCP volume were most strongly associated with changes in PCOLCE (%change NCP volume/log2-fold change&#x2009;=&#x2009;-31.9%; 95% CI, -42.9% to -18.7%; P&#x2009;<&#x2009;.001), independent of changes in LDL-C level. Increases in PCOLCE related most strongly to change in the fibro-fatty (<130 Hounsfield units) component of NCP (%change fibro-fatty volume/log2-fold change&#x2009;=&#x2009;-38.5%; 95% CI, -58.1% to -9.7%; P&#x2009;=&#x2009;.01) with a directionally opposite, although nonsignificant, increase in calcified plaque (%change calcified volume/log2-fold change&#x2009;=&#x2009;34.4%; 95% CI, -7.9% to 96.2%; P&#x2009;=&#x2009;.12). CONCLUSIONS AND RELEVANCE: Results of this secondary analysis of the REPRIEVE randomized clinical trial suggest that PCOLCE may be associated with the atherosclerotic plaque stabilization effects of statins by promoting collagen deposition in the extracellular matrix transforming vulnerable plaque phenotypes to more stable coronary lesions. TRIAL REGISTRATION: ClinicalTrials.gov Identifier: NCT02344290.

Humans↗

Gene discovery using the serial analysis of gene expression technique: implications for cancer research.

Cancer is a genetic disease. As such, our understanding of the pathobiology of tumors derives from analyses of the genes whose mutations are responsible for those tumors. The cancer phenotype, however, likely reflects the changes in the expression patterns of hundreds or even thousands of genes that occur as a consequence of the primary mutation of an oncogene or a tumor suppressor gene. Recently developed functional genomic approaches, such as DNA microarrays and serial analysis of gene expression (SAGE), have enabled researchers to determine the expression level of every gene in a given cell population, which represents that cell population's entire transcriptome. The most attractive feature of SAGE is its ability to evaluate the expression pattern of thousands of genes in a quantitative manner without prior sequence information. This feature has been exploited in three extremely powerful applications of the technology: the definition of transcriptomes, the analysis of differences between the gene expression patterns of cancer cells and their normal counterparts, and the identification of downstream targets of oncogenes and tumor suppressor genes. Comprehensive analyses of gene expression not only will further understanding of growth regulatory pathways and the processes of tumorigenesis but also may identify new diagnostic and prognostic markers as well as potential targets for therapeutic intervention.

Gene Expression Profiling↗

An isoform-resolution transcriptomic atlas of colorectal cancer from long-read single-cell sequencing.

Colorectal cancer (CRC) ranks as the second leading cause of cancer deaths globally. In recent years, short-read single-cell RNA sequencing (scRNA-seq) has been instrumental in deciphering tumor heterogeneities. However, these studies only enable gene-level quantification but neglect alterations in transcript structures arising from alternative end processing or splicing. In this study, we integrated short- and long-read scRNA-seq of CRC samples to build an isoform-resolution CRC transcriptomic atlas. We identified 394 dysregulated transcript structures in tumor epithelial cells, including 299 resulting from various combinations of splicing events. Second, we characterized genes and isoforms associated with epithelial lineages and subpopulations exhibiting distinct prognoses. Among 31,935 isoforms with novel junctions, 330 were supported by The Cancer Genome Atlas RNA-seq and mass spectrometry data. Finally, we built an algorithm that integrated novel peptides derived from open reading frames of recurrent tumor-specific transcripts with mass spectrometry data and identified recurring neoepitopes that may aid the development of cancer vaccines.

Colorectal Neoplasms↗

Integrated analysis of the genome and the transcriptome by FANTOM.

The key to reliable annotation of a mammalian genome is broad characterisation of the transcriptional output, the transcriptome. FANTOM, the functional annotation of mouse cDNA, is a large-scale analysis of both the genome and the transcriptome of the mouse. In the early days of this work, the transcripts were characterised using our sophisticated methods. After the timely release of the first draft of mouse genome sequences, interesting information was obtained by its integration with these one-by-one annotations. Moreover, each transcript included its expression profile. Here, the two integrated annotation methods used by FANTOM are reviewed: one-by-one and categorised. One-by-one annotation refers to naming carried out based on well-known transcripts or its fragments using the top-down-style pipeline developed mostly by the FANTOM project. Categorised annotation, which refers to transcript grouping, not only helps naming of unknown transcripts, but will be the most utilised method for integration of the genome and the transcriptome from now on.

Abstracting and Indexing↗

Comprehensive Transcriptome Annotation of Thousands of HIV-1 Genomes.

Alternative splicing in HIV-1 has been a central focus of decades of research, uncovering key mechanisms of viral gene regulation, immune evasion, and therapeutic response - yet, no reference resource has existed to support transcriptome-wide analysis, limiting adoption of modern computational methods. We present HIV Atlas (https://ccb.jhu.edu/HIV_Atlas), the first reference-quality annotation of HIV-1 and SIV transcriptional diversity. We manually curated transcriptomes for HIV-1HXB2 and SIVmac239 and developed Vira, an automated annotation-transfer method specifically designed to address unique challenges of viral genome biology, to generate high-quality annotations for 2,077 complete HIV-1 genomes. Using the resources presented in our work, we evaluated conservation of splice sites, revealing near-perfect preservation of major donors and acceptors. Furthermore, using several public datasets, we demonstrate how HIV Atlas enhances methodology, improves the quality and novelty of results, and opens novel avenues for research, supporting more accurate and comprehensive analyses of bulk, single-cell, and spatial RNA-seq in HIV-1 studies.

Journal Article↗

The use of Open Reading frame ESTs (ORESTES) for analysis of the honey bee transcriptome.

BACKGROUND: The ongoing efforts to sequence the honey bee genome require additional initiatives to define its transcriptome. Towards this end, we employed the Open Reading frame ESTs (ORESTES) strategy to generate profiles for the life cycle of Apis mellifera workers. RESULTS: Of the 5,021 ORESTES, 35.2% matched with previously deposited Apis ESTs. The analysis of the remaining sequences defined a set of putative orthologs whose majority had their best-match hits with Anopheles and Drosophila genes. CAP3 assembly of the Apis ORESTES with the already existing 15,500 Apis ESTs generated 3,408 contigs. BLASTX comparison of these contigs with protein sets of organisms representing distinct phylogenetic clades revealed a total of 1,629 contigs that Apis mellifera shares with different taxa. Most (41%) represent genes that are in common to all taxa, another 21% are shared between metazoans (Bilateria), and 16% are shared only within the Insecta clade. A set of 23 putative genes presented a best match with human genes, many of which encode factors related to cell signaling/signal transduction. 1,779 contigs (52%) did not match any known sequence. Applying a correction factor deduced from a parallel analysis performed with Drosophila melanogaster ORESTES, we estimate that approximately half of these no-match ESTs contigs (22%) should represent Apis-specific genes. CONCLUSIONS: The versatile and cost-efficient ORESTES approach produced minilibraries for honey bee life cycle stages. Such information on central gene regions contributes to genome annotation and also lends itself to cross-transcriptome comparisons to reveal evolutionary trends in insect genomes.

Animals↗