Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

The UTRs of Leishmania donovani vary in length and are enriched in potential regulatory structures.

Leishmania spp. regulate gene expression largely post-transcriptionally, yet untranslated regions (UTRs) remain poorly delineated. We generated high-quality genome and transcriptome datasets for Leishmania donovani strain 1S2D (Ld1S) by combining PacBio HiFi de novo assembly with Oxford Nanopore direct RNA sequencing of promastigotes and axenic amastigotes. The genome assembly consists of 65 scaffolds totaling ~33.3 Mb. Structural comparisons to LdBPK282A1 revealed numerous rearrangements, including some reshuffling genes among polycistronic transcription units and validated by polycistronic reads from RNA sequencing. Promastigote and amastigote RNA sequencing produced 469,010 and 46,729 monocistronic reads containing a spliced-leader and a polyA tail sequences, defining 8,479 transcripts and supporting 7,415 of the 7,969 annotated protein coding genes, as well as 604 putative long non-coding RNAs. We annotated UTRs for 4,921 genes and observed that putative RNA G-quadruplexes were markedly enriched in UTRs. We also noted that 31.9% and 11.5% were expressed into multiple isoforms in promastigotes and amastigotes, respectively. Collectively, these data provide a comprehensive annotation of L. donovani genes and their UTRs and reveal widespread and stage-specific UTR length polymorphisms, and, overall, points to an important role of 3' UTR in post-transcriptional regulation in L. donovani.

Journal Article↗

High-Content CRISPR Screening: Methods and Applications.

Clustered regularly interspaced short palindromic repeats (CRISPR)-Cas9 screening has become a central technology in functional genomics, enabling genome-scale interrogation via pooled perturbations. Early CRISPR screens employed survival or simple phenotypic readouts to identify essential genes and drug resistance mechanisms. However, as biological questions have shifted toward understanding regulatory networks, cellular heterogeneity, and context-dependent gene functions, there has been increasing demand for screening strategies capable of capturing complex cellular phenotypes beyond cell fitness. Recent advances in single-cell sequencing, high-content imaging, and spatial transcriptomics have expanded the resolution of CRISPR screening by enabling multidimensional phenotypic characterization following genetic perturbation. By integrating pooled perturbations with diverse readouts, these approaches systematically map targeted gene edits to transcriptional states, cellular phenotypes, and microenvironmental contexts. Meanwhile, innovations in library design, delivery, and computational pipelines have further improved the robustness and interpretability of high-content screening platforms. This review synthesizes the methodological evolution of CRISPR screening, emphasizing advances in perturbation strategies, delivery systems, and multimodal readouts. Representative applications spanning oncology, immunotherapy, developmental biology, neurobiology, and infectious diseases are delineated to demonstrate refined gene network annotations. Additionally, existing technical bottlenecks, such as scalability, cost constraints, and in vivo limitations, are critically assessed. Finally, future directions are proposed to facilitate the development of precise medicine.

CRISPR screening↗

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article↗

An Annotated Biobank of Triple-Negative Breast Cancer Patient-Derived Xenografts Features Treatment-Naïve and Longitudinal Samples during Neoadjuvant Chemotherapy.

UNLABELLED: Triple-negative breast cancer (TNBC) that fails to respond to neoadjuvant chemotherapy (NACT) can be lethal. Developing effective strategies to eradicate chemoresistant disease requires experimental models that recapitulate the heterogeneity characteristic of TNBC. To that end, we established a biobank of 92 orthotopic patient-derived xenograft (PDX) models of TNBC from the tumors of 75 patients enrolled in A Robust TNBC Evaluation fraMework to Improve Survival clinical trial (ARTEMIS, NCT02276443), including 12 longitudinal sets generated from serial patient biopsies collected throughout NACT treatment and from metastatic disease. Models were established from both chemosensitive and chemoresistant tumors, and nearly 30% of the PDX models were capable of metastasizing to the lungs. Comprehensive molecular profiling demonstrated conservation of genomes and transcriptomes between patient and corresponding PDX tumors, with representation of all major transcriptional subtypes. Transcriptional changes observed in the longitudinal PDX models highlighted dysregulation in pathways associated with DNA integrity, extracellular matrix interactions, the ubiquitin-proteasome system, epigenetics, and inflammatory signaling. These alterations revealed a complex network of adaptations associated with chemoresistance. Overall, this PDX biobank provides a valuable tool for tackling the most pressing issues facing the clinical management of TNBC. SIGNIFICANCE: The development of a patient-derived xenograft biobank that comprehensively captures the genomic and transcriptional diversity of triple-negative breast cancer promises to be a robust resource to investigate and overcome chemoresistance and metastasis.

Animals↗

Moderate expression and activity of flocculins underlie the characteristic flocculation phenotype of Saccharomyces pastorianus.

Flocculation is a key technological trait in lager brewing, governing fermentation performance, yeast recovery, and beer quality. In the allo-aneuploid hybrid yeast Saccharomyces pastorianus, the genetic basis of flocculation remains poorly resolved due to its complex dual sub-genome architecture. Here, we systematically re-annotated and functionally characterized the complete FLO gene repertoire of the Group II strain CBS 1483. Thirteen FLO genes were identified, including allelic variants and a previously uncharacterized adhesin, Flo12, containing a Hyphal_reg_CWP domain instead of the canonical PA14 lectin-binding domain. Structural modeling revealed strong conservation of Ca²+-binding residues in PA14 domains, alongside repeat-region diversification likely contributing to functional variability. Using optogenetic expression in a FLO-null background, we demonstrated that SpcI-FLO9-1 and SpcI-FLO9-2_1 are the strongest drivers of flocculation, exhibiting NewFlo-like sugar sensitivity. Transcriptomic analysis during 17°P wort fermentation showed dynamic induction of these genes coinciding with flocculation onset. Surprisingly, deletion of both loci in CBS 1483 did not abolish but only delayed sedimentation in wort, accompanied by improved maltose utilization and attenuation. These findings reveal functional redundancy and compensatory mechanisms within the FLO network of lager yeast, highlighting the genetic complexity underlying flocculation, and providing a molecular framework to inform yeast selection, strain development, and optimization of the lager fermentation processes.IMPORTANCEFlocculation, the process by which yeast cells aggregate and settle, is essential for producing clear, high-quality lager beer, and for efficient yeast recovery during brewing. However, the genetic basis of this trait in lager yeast has remained poorly understood because these strains possess unusually complex hybrid genomes. In this study, we systematically identified and characterized the complete set of flocculation genes in the industrial lager yeast Saccharomyces pastorianus CBS 1483. We demonstrated that lager yeast flocculation is not controlled by a single dominant gene, but instead emerges from the combined action of several moderately active adhesion proteins that are expressed at low levels during fermentation. Surprisingly, deleting the two strongest candidate genes only delayed, rather than eliminated, sedimentation, revealing a robust compensatory network that preserves brewing performance. These findings refine the current understanding of yeast flocculation and provide a molecular framework for developing brewing strains with improved fermentation efficiency, product consistency, and flavor quality.

Saccharomyces pastorianus↗

Enhanced chromatin compaction is associated with de novo expression of a nuclear microprotein, global loss of H3 acetylation and local transcriptional changes in retinal rod photoreceptors.

We have limited understanding of how aging alters gene expression and remodels cellular architecture in post-mitotic neurons. The inverted nuclear organization of mouse rod photoreceptors provides a unique model to gain mechanistic insights into age-associated decline in neuronal function. We have generated and integrated multi-omic datasets including 3D-genome topology, histone modifications, chromatin accessibility, DNA methylation and transcriptome of rod photoreceptors from young- and aged-mice. We show that aging drives global chromatin compaction, with regional alterations enriched at active chromatin. Epigenomic and transcriptional changes broadly correlate with chromatin dynamics as validated by high resolution microscopy. We uncover a megabase-sized genomic region with multi-level alterations, including de novo transcription of Gm7239, which encodes a functional microprotein carrying histone acetyltransferase-inhibitor domain. Overexpression of Gm7239 is associated with global loss of histone H3 acetylation, highlighting a potential new axis of genomic regulation in aging. Finally, we identify multiple significant local transcriptional alterations in non-annotated regions and genes associated with age-related macular degeneration. Our studies link age-related chromatin landscape changes with gene expression that may influence rod function and vulnerability to diseases.

Journal Article↗

Impact of alternative initiation, splicing, and termination on the diversity of the mRNA transcripts encoded by the mouse transcriptome.

We analyzed the FANTOM2 clone set of 60,770 RIKEN full-length mouse cDNA sequences and 44,122 public mRNA sequences. We developed a new computational procedure to identify and classify the forms of splice variation evident in this data set and organized the results into a publicly accessible database that can be used for future expression array construction, structural genomics, and analyses of the mechanism and regulation of alternative splicing. Statistical analysis shows that at least 41% and possibly as much as 60% of multiexon genes in mouse have multiple splice forms. Of the transcription units with multiple splice forms, 49% contain transcripts in which the apparent use of an alternative transcription start (stop) is accompanied by alternative splicing of the initial (terminal) exon. This implies that alternative transcription may frequently induce alternative splicing. The fact that 73% of all exons with splice variation fall within the annotated coding region indicates that most splice variation is likely to affect the protein form. Finally, we compared the set of constitutive (present in all transcripts) exons with the set of cryptic (present only in some transcripts) exons and found statistically significant differences in their length distributions, the nucleotide distributions around their splice junctions, and the frequencies of occurrence of several short sequence motifs.

Alternative Splicing↗

A single-cell transcriptomic atlas of the pigtail macaque placenta in late gestation.

The placenta is a complex organ with multiple immune and non-immune cell types that promote fetal tolerance and facilitate the transfer of nutrients and oxygen. The nonhuman primate (NHP) is a key experimental model for studying human pregnancy complications, in part due to similarities in placental structure, which makes it essential to understand how single-cell populations compare across the human and NHP maternal-fetal interface. We constructed a single-cell RNA-Seq (scRNA-Seq) atlas of the placenta from the pigtail macaque ( Macaca nemestrina ) in the third trimester, comprising three different tissues at the maternal-fetal interface: the chorionic villi (placental disc), chorioamniotic membranes, and the maternal decidua. Each tissue was separately dissociated into single cells and processed through the 10X Genomics and Seurat pipeline, followed by aggregation, unsupervised clustering, and cluster annotation. Next, we determined the maternal-fetal origins of cell populations and analyzed single-cell RNA trajectory, Gene Ontology enrichment, and cell-cell communication. Single-cell populations in the pigtail macaque were strikingly similar in their identity and frequency to those found in the human placenta, including cells from trophoblast, stromal cell, immune, and macrophage lineages. An advantage of our approach was the deep sequencing of three tissues at the maternal-fetal interface, which yielded a rich diversity of common and rare single-cell populations. The third-trimester pigtail macaque single-cell atlas enables the identification of cellular subclusters analogous to those in humans and provides a powerful resource for understanding experimental perturbations on the NHP placenta.

Journal Article↗

The Schistosoma mansoni gene index: gene discovery and biology by reconstruction and analysis of expressed gene sequences.

Expressed sequence tag (EST) sequencing and analysis is a primary research tool to identify and characterize the Schistosoma mansoni transcriptome. As part of our gene discovery effort, a total of 5,793 ESTs have been generated from clones selected randomly from complementary DNA (cDNA) libraries constructed from male and female adult worms. Assembly analysis of all the 16,813 public S. mansoni ESTs has identified 1,920 distinct tentative consensus sequences (TCs) and 5,571 nonoverlapping ESTs (singletons). Of these, 376 TCs (20%) and 1,449 singletons (26%) are unique to the SUNY/TIGR sequencing effort. Tentative consensus sequences and singletons were distributed into various categories of biological roles associated with cell structure, metabolism, protein fate, signal transduction, transcription, protein synthesis, transporters, and cell growth. The TCs and singletons represent transcripts that can be used as a resource for functional annotation of genomic sequence data, comparative sequence analysis, and cDNA clone selection for microarray projects. The utility of EST analysis is demonstrated by identifying new protease genes, which may be involved in hemoglobin degradation.

Amino Acid Sequence↗

Receptor-defined targeting of a genomically unique melanoma-enriched noncanonical antigen.

Effective T cell-based immunotherapies require functional receptors that can be engineered and redeployed to recognize tumor-restricted antigens. Noncanonical peptides arising from transcription outside annotated protein-coding regions expand the antigenic landscape of cancer; however, systematic strategies to biologically prioritize and functionally validate such targets remain underdeveloped. Here, we integrated de novo transcript analysis, exon-resolved quantification, RNA in situ hybridization, and immunopeptidomics to identify melanoma-associated noncanonical transcripts and advance candidates through receptor-level validation. Among three recurrent melanoma-associated transcripts, EVA003 emerged as a lead target based on its distinct repeat-enriched genomic architecture, consistent tumor-enriched exon-level expression across independent datasets, and a genomically unique immunogenic core sequence. We demonstrate endogenous presentation of EVA003-derived peptides on HLA-A*03:01 and detect specific reactivity in patient-derived tumor-infiltrating lymphocytes. Single-cell transcriptomic profiling identified a dominant peptide-reactive clonotype, enabling isolation of a naturally occurring T cell receptor. Transfer of this receptor into healthy donor T cells conferred antigen-dependent activation and cytotoxicity against both peptide-pulsed targets and melanoma cells expressing EVA003 endogenously. Together, these findings establish a biologically informed strategy for prioritizing noncanonical tumor antigens and demonstrate that genomically unique, tumor-enriched noncanonical peptides can be presented to molecularly defined receptors capable of mediating cancer cell killing. These findings support the integration of prioritized noncanonical antigens into engineered T cell therapeutic strategies.

Humans↗

Structural characterization and predicted biosynthetic pathway of the polysaccharide component of bioflocculant from starch-degrading Bacillus subtilis ZHX3.

Polysaccharides-based bioflocculant is a promising eco-friendly alternative to conventional flocculants, yet their application is limited by high production cost. Understanding the biosynthetic pathway is essential for targeted strain improvement. In this study, we characterized polysaccharides structure of bioflocculant MBF-ZHX3 from Bacillus subtilis ZHX3 and predicted its biosynthetic pathway via genomic analysis combined with quantitative real-time PCR (qPCR). Two purified polysaccharide fractions, PS1-1 (5982 Da) and PS2-1 (17,577 Da), were obtained. Both were mainly composed of glucose, with a backbone of →4)-α-D-Glcp-(1 → and α-D-Glcp-(1 → branches attached at O-6. Whole-genome sequencing revealed a circular chromosome of 4,122,369 bp and two plasmids. Functional annotation showed high carbohydrate metabolism activity, with 284 genes (9.52%) and 264 genes (11.28%) assigned to carbohydrate metabolism in the COG and KEGG database, respectively. A complete eps gene cluster consisting of 15 open reading frames was identified. qPCR showed that key genes involved in substrate uptake (ptsG, malP, mdxEFG-msmX) and nucleotide sugar synthesis (pgcA, gtaB) were significantly upregulated. The priming glycosyltransferase (GT) epsL and the primary GT epsF were upregulated, along with the flippase epsK, polymerase epsG, and chain-length regulators epsA and epsB. Based on these findings, we propose a putative biosynthetic pathway for the polysaccharide component of MBF-ZHX3, and identify epsL, epsF, and epsG as prioritized targets for future genetic engineering. This work provides an integrated structural-genomic-transcriptomic framework that can guide rational strain improvement to enhance bioflocculant production.

Polysaccharides structure↗

Genomic landscape of hepatocellular carcinoma in Egyptian patients by whole exome sequencing.

BACKGROUND: Hepatocellular carcinoma (HCC) is the most common primary liver cancer. Chronic hepatitis and liver cirrhosis lead to accumulation of genetic alterations driving HCC pathogenesis. This study is designed to explore genomic landscape of HCC in Egyptian patients by whole exome sequencing. METHODS: Whole exome sequencing using Ion Torrent was done on 13 HCC patients, who underwent surgical intervention (7 patients underwent living donor liver transplantation (LDLT) and 6 patients had surgical resection}. RESULTS: Mutational signature was mostly S1, S5, S6, and S12 in HCC. Analysis of highly mutated genes in both HCC and Non-HCC revealed the presence of highly mutated genes in HCC (AHNAK2, MUC6, MUC16, TTN, ZNF17, FLG, MUC12, OBSCN, PDE4DIP, MUC5b, and HYDIN). Among the 26 significantly mutated HCC genes-identified across 10 genome sequencing studies-in addition to TCGA, APOB and RP1L1 showed the highest number of mutations in both HCC and Non-HCC tissues. Tier 1, Tier 2 variants in TCGA SMGs in HCC and Non-HCC (TP53, PIK3CA, CDKN2A, and BAP1). Cancer Genome Landscape analysis revealed Tier 1 and Tier 2 variants in HCC (MSH2) and in Non-HCC (KMT2D and ATM). For KEGG analysis, the significantly annotated clusters in HCC were Notch signaling, Wnt signaling, PI3K-AKT pathway, Hippo signaling, Apelin signaling, Hedgehog (Hh) signaling, and MAPK signaling, in addition to ECM-receptor interaction, focal adhesion, and calcium signaling. Tier 1 and Tier 2 variants KIT, KMT2D, NOTCH1, KMT2C, PIK3CA, KIT, SMARCA4, ATM, PTEN, MSH2, and PTCH1 were low frequency variants in both HCC and Non-HCC. CONCLUSION: Our results are in accordance with previous studies in HCC regarding highly mutated genes, TCGA and specifically enriched pathways in HCC. Analysis for clinical interpretation of variants revealed the presence of Tier 1 and Tier 2 variants that represent potential clinically actionable targets. The use of sequencing techniques to detect structural variants and novel techniques as single cell sequencing together with multiomics transcriptomics, metagenomics will integrate the molecular pathogenesis of HCC in Egyptian patients.

Humans↗

The genexpress IMAGE knowledge base of the human muscle transcriptome: a resource of structural, functional, and positional candidate genes for muscle physiology and pathologies.

Sequence, gene mapping, and expression data corresponding to 910 genes transcribed in human skeletal muscle have been integrated to form the muscle module of the Genexpress IMAGE Knowledge Base. Based on cDNA array hybridization, a set of 14 transcripts preferentially or specifically expressed in muscle have been selected and characterized in more detail: Their pattern of expression was confirmed by Northern blot analysis; their structure was further characterized by full-insert cDNA sequencing and cDNA extension; the map location of the corresponding genes was refined by radiation hybrid mapping. Five of the 14 selected genes appear as interesting positional and functional candidate genes to study in relation with muscle physiology and/or specific orphan muscular pathologies. One example is discussed in more detail. The expression profiling data and the associated Genexpress Index2 entries for the 910 genes and the detailed characterization of the 14 selected transcripts are available from a dedicated Web server at. The database has been organized to provide the users with a working space where they can find curated, annotated, integrated data for their genes of interest. Different navigation routes to exploit the resource are discussed.

Base Sequence↗

Temperature-regulated transcription in the pathogenic fungus Cryptococcus neoformans.

The basidiomycete fungus Cryptococcus neoformans is an opportunistic pathogen of worldwide importance that causes meningitis, leading to death in immunocompromised individuals. Unlike many basidiomycete fungi, C. neoformans is thermotolerant, and its ability to grow at 37 degrees C is considered to be a virulence factor. We used serial analysis of gene expression (SAGE) to characterize the transcriptomes of C. neoformans strains that represent two varieties with different polysaccharide capsule serotypes. These include a serotype D strain of the C. neoformans variety neoformans and a serotype A strain of variety grubii. In this report, we describe the construction and characterization of SAGE libraries from each strain grown at 25 degrees C and 37 degrees C. The SAGE data reveal transcriptome differences between the two strains, even at this early stage of analysis, and identify sets of genes with higher transcript levels at 25 degrees C or 37 degrees C. Notably, growth at the lower temperature increased transcript levels for histone genes, indicating a general influence of temperature on chromatin structure. At 37 degrees C, we noted elevated transcript levels for several genes encoding heat shock proteins and translation machinery. Some of these genes may play a role in temperature-regulated phenotypes in C. neoformans, such as the adaptation of the fungus to growth in the host and the dimorphic transition between budding and filamentous growth. Overall, this work provides the most comprehensive gene expression data available for C. neoformans; this information will be a critical resource both for gene discovery and genome annotation in this pathogen.

Blotting, Northern↗

eVOC: a controlled vocabulary for unifying gene expression data.

Expression data contribute significantly to the biological value of the sequenced human genome, providing extensive information about gene structure and the pattern of gene expression. ESTs, together with SAGE libraries and microarray experiment information, provide a broad and rich view of the transcriptome. However, it is difficult to perform large-scale expression mining of the data generated by these diverse experimental approaches. Not only is the data stored in disparate locations, but there is frequent ambiguity in the meaning of terms used to describe the source of the material used in the experiment. Untangling semantic differences between the data provided by different resources is therefore largely reliant on the domain knowledge of a human expert. We present here eVOC, a system which associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. We have curated and annotated 7016 cDNA libraries represented in dbEST, as well as 104 SAGE libraries,with expression information,and provide this as an integrated, public resource that allows the linking of transcripts and libraries with expression terms. Both the vocabularies and the vocabulary-annotated libraries can be retrieved from http://www.sanbi.ac.za/evoc/. Several groups are involved in developing this resource with the aim of unifying transcript expression information.

Animals↗

Dietary effects of arachidonate-rich fungal oil and fish oil on murine hepatic and hippocampal gene expression.

BACKGROUND: The functions, actions, and regulation of tissue metabolism affected by the consumption of long chain polyunsaturated fatty acids (LC-PUFA) from fish oil and other sources remain poorly understood; particularly how LC-PUFAs affect transcription of genes involved in regulating metabolism. In the present work, mice were fed diets containing fish oil rich in eicosapentaenoic acid and docosahexaenoic acid, fungal oil rich in arachidonic acid, or the combination of both. Liver and hippocampus tissue were then analyzed through a combined gene expression- and lipid- profiling strategy in order to annotate the molecular functions and targets of dietary LC-PUFA. RESULTS: Using microarray technology, 329 and 356 dietary regulated transcripts were identified in the liver and hippocampus, respectively. All genes selected as differentially expressed were grouped by expression patterns through a combined k-means/hierarchical clustering approach, and annotated using gene ontology classifications. In the liver, groups of genes were linked to the transcription factors PPARalpha, HNFalpha, and SREBP-1; transcription factors known to control lipid metabolism. The pattern of differentially regulated genes, further supported with quantitative lipid profiling, suggested that the experimental diets increased hepatic beta-oxidation and gluconeogenesis while decreasing fatty acid synthesis. Lastly, novel hippocampal gene changes were identified. CONCLUSIONS: Examining the broad transcriptional effects of LC-PUFAs confirmed previously identified PUFA-mediated gene expression changes and identified novel gene targets. Gene expression profiling displayed a complex and diverse gene pattern underlying the biological response to dietary LC-PUFAs. The results of the studied dietary changes highlighted broad-spectrum effects on the major eukaryotic lipid metabolism transcription factors. Further focused studies, stemming from such transcriptomic data, will need to dissect the transcription factor signaling pathways to fully explain how fish oils and arachidonic acid achieve their specific effects on health.

Journal Article↗

Contrasting regulation of protein-coding genes and lncRNA homeologs in allotetraploid Coffea arabica.

A chromosome-level Bourbon assembly revealed that protein-coding homeologs are predominantly co-regulated between subgenomes. In contrast, intergenic lncRNAs display a modest, but statistically consistent bias toward subgenome E across diverse developmental and stress contexts. Coffea arabica is an allotetraploid species derived from natural hybridization between C. canephora and C. eugenioides, which contributed the C and E subgenomes, respectively. This genomic origin poses major challenges for genome assembly, annotation, and the interpretation of gene regulation. In this study, a high-quality genome assembly of C. arabica was generated and annotated, with particular emphasis on identifying protein-coding genes and intergenic long non-coding RNAs (lincRNAs). Homeologous relationships between genes from the C and E subgenomes were established, providing a robust framework to investigate subgenomic conservation and regulatory divergence. Using an extensive collection of publicly available RNA-seq libraries spanning multiple developmental stages, tissues, and environmental conditions, the relative transcriptional contribution of each subgenome was evaluated. On a global scale, gene expression was largely balanced between subgenomes, with no consistent evidence of subgenome dominance. While protein-coding genes showed comparable regulatory behavior across subgenomes, lincRNAs exhibited a more asymmetric expression pattern, suggesting higher subgenome-specific expression that is interpreted here as a consistent directional tendency rather than as evidence of subgenome dominance. Together, these results provide new insights into the regulatory architecture of the C. arabica genome and establish a foundational genomic and transcriptomic resource for future functional studies and crop improvement efforts.

Coffea↗

The machine-learning classifier ALLCatchR2 identifies 20 T-ALL subtypes across cohorts and age groups.

T-cell acute lymphoblastic leukemia (T-ALL) comprises molecularly diverse subtypes, but robust cross-cohort validations and operational gene-expression definitions are lacking. To establish a gene-expression-anchored framework for T-ALL subtyping, we aggregated 2314 transcriptomes (15 cohorts, age: 0.8-90.8 years). An extended unsupervised approach defined 17 main clusters and 3 subclusters in samples with high blast fractions. Supervised analyses added an overarching immature T-ALL (early T cell precursor [ETP]-like) definition and resolved the LMO2 &#x3b3;&#x3b4;-like subtype. All clusters contained samples from at least two cohorts. Characteristic genomic driver enrichments were consistent across cohorts, while gene-expression clusters did not correspond exclusively to single driver events but also reflected developmental origins. A machine-learning classifier based on ALLCatchR, our B-cell acute lymphoblastic leukemia (B-ALL) classifier, identified these 20 transcriptomic subtypes and the immature T-ALL (ETP-like) signature with 0.995-1.0 accuracy in a validation set (n&#x2009;=&#x2009;203). Testing the classifier on a second hold-out data set (n&#x2009;=&#x2009;265 samples) showed that 92.7% of predictions matched with corresponding driver alterations. Across all samples, 83.2% of cases received high-confidence predictions, 7.3% candidate predictions, and 9.5% remained unclassified, largely because of low blast fractions. We identified a novel gene-expression cluster markedly enriched (P&#x2009;<&#x2009;0.001) for clonal hematopoiesis mutations (IDH2 R140Q, DNMT3A) and a stem-/progenitor cell-like gene expression. This novel clonal hematopoiesis-related T-ALL subtype was observed in six cohorts and accounted for 8.9% of adults and 39.5% of patients aged >50 years. We extended&#xa0;ALLCatchR into ALLCatchR2, a free R package that now enables B-/T-lineage separation, gene-expression subtyping, blast estimation, and developmental annotation to harmonize T-ALL classification across studies and clinical contexts.

Journal Article↗