Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding↗

NextLongIso: a comprehensive Nextflow pipeline for multi-dimensional long-read RNA-seq analysis.

SUMMARY: Long-read RNA sequencing technologies, including Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), enable direct characterization of full-length transcripts and transcriptome complexity. However, analysis of long-read RNA-seq data remains fragmented across multiple tools, limiting the ability to obtain a unified view of transcript structure, expression, and regulatory variation in long-read transcriptomes. We present NextLongIso, a scalable and reproducible Nextflow pipeline that enables coordinated analysis of multiple layers of transcript regulation. Rather than focusing solely on transcript reconstruction, NextLongIso integrates transcript discovery with downstream regulatory analyses to jointly characterize alternative splicing, isoform switching, transcript boundary dynamics (including alternative promoters and polyadenylation), and transposable element-associated transcription from both PacBio and ONT datasets. By eliminating complex cross-tool data harmonization, this unified framework facilitates the transition from transcript identification to functional interpretation of transcriptomic variation. AVAILABILITY AND IMPLEMENTATION: NextLongIso is implemented in Nextflow and is freely available at github: https://github.com/YidanSunResearchLab/nf-LongIso.git and Zenodo: https://doi.org/10.5281/zenodo.21049837.

Software↗

Root growth promotion by Penicillium melinii : mechanistic insights and agricultural applications.

This study characterizes Penicillium melinii , an endophytic fungus isolated from Arabidopsis thaliana roots, as a plant growth-promoting fungus with potential use as a model to study root development and as a biostimulant for sustainable agriculture. Although endophytes are known to promote plant growth, the underlying molecular mechanisms often remain poorly understood. Here, we aimed to elucidate how P. melinii enhances root system development and to assess its applicability across different crops. Phenotypic assays were conducted in Arabidopsis, quinoa and tomato under in vitro , greenhouse and field conditions. Root architecture and biomass were quantified using image-based phenotyping. Transcriptomic and phytohormone profiling assessed plant responses, and fungal genome sequencing coupled with secretome analysis was used to identify candidate effectors and metabolic traits. P. melinii consistently promoted root growth and increased plant biomass across species and environments, both in vitro and in the greenhouse. In tomato field trials, this translated into a significant increase in yield. The fungus colonized root surfaces without vascular penetration and triggered a mild transcriptomic response: early activation of stress-response genes followed by their attenuation and sustained upregulation of auxin-related pathways. Notably, the interaction modulates the SLR-ARF-LBD pathway and the number of pre-branch sites probably through increased auxin signalling in the oscillation zone. Additional hormonal changes were limited and mainly associated with the attenuation of the plant response to microorganisms. P. melinii enhances lateral root formation through a subtle molecular and metabolic dialogue with the host plant, underscoring its relevance as a model for studying root developmental plasticity. Its strong and reproducible growth-promoting effect, demonstrated with different fungal strains and under controlled and field conditions, supports its potential as a biostimulant for sustainable crop production.

Journal Article↗

Analysis of bovine mammary gland EST and functional annotation of the Bos taurus gene index.

Functional genomic studies of the mammary gland require an appropriate collection of cDNA sequences to assess gene expression patterns from the different developmental and operational states of underlying cell types. To better capture the range of gene expression, a normalized cDNA library was constructed from pooled bovine mammary tissues, and 23,202 expressed sequence tags (EST) were produced and deposited into GenBank. Assembly of these EST with sequences in the Bos taurus Gene Index (BtGI) helped to form 5751 of the current 23,883 tentative consensus (TC) sequences. The majority (87%) of these 5751 assemblies contained only one to three mammary-derived EST. In contrast, 18% of the mammary EST assembled with TC sequences corresponding to 12 genes. These results suggest library normalization was only partially effective, because the reduction in EST for genes abundantly transcribed during lactation could be attributed to pooling. For better assessment of novel content in the mammary library and to add to existing annotation of all bovine sequence elements, gene ontology assignments, and comparative sequence analyses against human genome sequence, human and rodent gene indices, and an index of orthologous alignments of genes across eukaryotes (TOGA) were performed, and results were added to existing BtGI annotation. Over 35,000 of the bovine elements significantly matched human genome sequence, and the positions of some alignments (3%) were unique relative to those using human expressed sequences. Because 3445 TC sequences had no significant match with any data set, mammary-derived cDNA clones representing 23 of these elements were analyzed further for expression and novelty. Only one clone met criteria suggesting the corresponding gene was a divergent ortholog or expressed sequence unique to cattle. These results demonstrate that bovine sequence expression data serve as a resource for characterizing mammalian transcriptomes and identifying those genes potentially unique to ruminants.

Animals↗

Long-range heterogeneity at the 3' ends of human mRNAs.

The publication of a draft of the human genome and of large collections of transcribed sequences has made it possible to study the complex relationship between the transcriptome and the genome. In the work presented here, we have focused on mapping mRNA 3' ends onto the genome by use of the raw data generated by the expressed sequence tag (EST) sequencing projects. We find that at least half of the human genes encode multiple transcripts whose polyadenylation is driven by multiple signals. The corresponding transcript 3' ends are spread over distances in the kilobase range. This finding has profound implications for our understanding of gene expression regulation and of the diversity of human transcripts, for the design of cDNA microarray probes, and for the interpretation of gene expression profiling experiments.

3' Flanking Region↗

A High-Resolution Stereo-Seq Spatial Transcriptomic Resource for Adult Holstein Cattle Liver.

The bovine liver is a highly compartmentalized organ that plays essential roles in continuous gluconeogenesis and nitrogen recycling; however, its spatial molecular architecture has remained largely uncharacterized due to the limitations of traditional bulk and single-cell approaches. To address this gap, Spatial Enhanced Resolution Omics-sequencing (Stereo-seq) was utilized to generate a subcellular-resolution (500 nm) transcriptomic map of an adult Holstein cattle liver, and a refined reference-guided workflow was implemented to overcome standard annotation limitations in livestock. Raw sequencing data were processed using the Stereo-seq Analysis Workflow and analyzed with Stereopy, Seurat, SingleR, and reference-guided workflows. Spatial aggregation was evaluated at Bin20, Bin50, Bin100, Bin150, and Bin200. Increasing bin size increased molecular identifier counts and detected-gene complexity while progressively reducing spatial granularity. Bin50, corresponding to 50 × 50 DNA nanoballs and an approximate nominal footprint of 25 × 25 µm, was therefore selected as a practical intermediate aggregation level for the primary analyses. Quality-control assessment, Leiden clustering, UMAP visualization, reference-based cell-type annotation, cluster-marker analysis, and spatial mapping of canonical hepatic genes demonstrated preservation of biologically interpretable liver transcriptional organization. Raw sequencing data processed spatial matrices, annotated objects, and analysis code are publicly available to support reanalysis and computational benchmarking. In summary, we present a Stereo-seq spatial transcriptomic resource generated from liver tissue of an adult Holstein cow. This initial resource provides a valuable foundation for future studies of bovine liver biology, comparative genomics, and the spatial basis of livestock health and production traits.

Animals↗

Serial analysis of gene expression in murine fetal thymocyte cell lines.

FTL-1, -3 and -10 are three murine day 14 fetal thymocyte cell lines produced in order to model developmental stages within early (CD3-CD4-CD8-) thymocyte differentiation. In this study, we used the serial analysis of gene expression (SAGE) method to perform a systematic analysis of transcripts present in these three cell lines. A total of 77,313 SAGE tags were sequence identified from the three cell lines, representing 24,645 unique transcripts. Differentially expressed mRNA transcripts representing different gene classes were identified, including T cell functional genes, cytokine receptors, adhesion molecules and transcription factors. These results may serve as a model of the transcriptome of early thymocyte differentiation. A large number of unknown expressed sequence tags were also found to be differentially expressed. In order to validate the SAGE data, selected differentially expressed transcripts identified by SAGE were analyzed by quantitative RT-PCR in normal murine double-negative stage DN1-4 thymocytes. Expression of the transcription factors RUNX2 and PHD finger protein 2 and of the IGF type 1 receptor was shown to have differentially regulated expression patterns in sorted DN1-4 cells. These genes, and others identified by this analysis, are likely to play important roles in the development of T cells.

Animals↗

Genomic subtypes of non-muscle-invasive bladder cancer: guiding immunotherapy decision-making for patients exposed to aristolochic acid.

BACKGROUND: The limited genomic data on non-muscle-invasive bladder cancer (NMIBC) hampers our understanding of its carcinogenesis and development. Specifically, Aristolochic acid (AA), a potent human carcinogenic compound from aristolochia plants and commonly found in Chinese herbal medicine, has been extensively documented as being closely associated with the onset and progression of bladder cancer. However, the field of AA-induced NMIBC remains largely unexplored in terms of its genomic and molecular characteristics, as well as clinical therapeutic strategies. METHODS: To bridge this knowledge gap, we conducted a comprehensive study using a cohort of 81 NMIBC samples. We performed whole-exome sequencing (WES) and RNA sequencing (RNA-seq) to obtain detailed genomic and transcriptomic data. We subjected these datasets to genomic analysis and subtype analysis to gain valuable insights into NMIBC. RESULTS: By temporally dissecting mutations in NMIBC specimens, we identified a comprehensive mutational landscape of NMIBC and the associations of these mutations with recurrence-free survival. Additionally, we discerned four genomic subtypes of NMIBC: AA-like, FGFR3/HRAS, FGFR3 & chr9Del, and genome instability (GI). The AA-like subtype presented a high frequency of gene mutations along with a pronounced AA mutagenesis signature of SBS22 (Fisher test: P-value 3.5e-4, OR 25.25) even after temporal dissection. The FGFR3/HRAS subtype exhibited FGFR3 or HRAS mutations with few copy number alterations (CNAs). The FGFR3 & chr9Del subtype was characterized by the co-occurrence of chr9p and chr9q deletions as well as FGFR3 mutations, while the GI subtype showed a high frequency of CNAs. Notably, the AA-like and GI subtypes demonstrated better outcomes after immunotherapy, whereas the FGFR3/HRAS subtype showed poorer outcomes. CONCLUSIONS: Our findings provide novel perspectives on the genomics of NMIBC, unveiling four prominent genomic subtypes, each showing different outcomes following immunotherapy. TRIAL REGISTRATION: No. 2019PHB268-01 (retrospectively registered on February 14, 2020).

Humans↗

Transcriptome of channel catfish (Ictalurus punctatus): initial analysis of genes and expression profiles of the head kidney.

Analysis of expressed sequence tags (ESTs) is an efficient approach for gene discovery, expression profiling, and development of resources useful for functional genomics studies. As part of the transcriptome analysis in channel catfish (Ictalurus punctatus), we have conducted EST analysis using a cDNA library made from the head kidney. We analysed 2228 EST clones. Orthologues were established for 1495 (67.1%) clones representing 748 genes, of which 545 (36.5%) clones were singletons. The remaining 733 (32.9%) clones represent unknown gene clones, for which the number of genes has not yet been determined.

Animals↗

Fasudil induces anti-inflammatory transcriptomic changes and increased proliferation in human trisomy 21 neural progenitor cells.

Down syndrome (DS) results from trisomy for human chromosome 21 and is the most frequent genetic cause of intellectual disability. No effective treatments currently exist that improve neurodevelopment and cognition. Atypical brain development in individuals with DS is apparent before birth, which suggests that the optimal time to begin administration of therapies is prenatally. Human neural progenitor cell (NPC) cultures provide a tractable in vitro model system to examine the effects of trisomy 21 (T21) on neurodevelopment and to measure the effects of pharmacological interventions. Here, we report the results of preclinical studies evaluating 24 candidate therapies. RNA sequencing analyses found that euploid and T21 NPCs showed different transcriptomic responses to five candidate pharmacotherapies. The Rho-associated coiled-coil kinase inhibitor fasudil increased proliferation of T21 NPCs, reduced expression of inflammatory pathway genes in T21 NPCs, and reduced markers of inflammation in LPS-stimulated microglial model systems. These results demonstrate that fasudil can alter multiple T21-associated abnormalities in a beneficial manner, suggesting that fasudil warrants further study as a candidate prenatal pharmacotherapy for DS.

Down Syndrome↗

Spatial mapping of RNA turnover kinetics in the mouse brain.

Gene regulation requires coordinated control of RNA synthesis and degradation, yet measuring RNA turnover across intact tissues remains challenging. Here we present spatial NT-seq, a method that combines transgenesis-free metabolic RNA labeling with in situ chemical recoding on spatial transcriptomics platforms to co-map newly synthesized and pre-existing RNAs. Applying spatial NT-seq to the mouse brain reveals pronounced regional heterogeneity in RNA turnover and identifies the dentate gyrus as a spatial hotspot marked by coordinated upregulation of basal RNA synthesis and decay. Moreover, spatial NT-seq uncovers rapid, brain region-specific transcriptional and post-transcriptional responses to electroconvulsive stimulation, a clinically relevant treatment for refractory depression. Finally, we leverage computational modeling to identify sequence features and post-transcriptional regulators that shape transcriptome-wide mRNA stability across spatial and cellular contexts in the mouse brain. Together, this integrated 'in vivo timescope' framework provides a spatially resolved view of RNA turnover kinetics and reveals the regulatory architecture of RNA stability in vivo.

Journal Article↗

Prominent Movement Disorders in RNU2-2-Related Spliceosomopathy.

Pediatric movement disorders often overlap with neurodevelopmental diseases, suggesting shared molecular mechanisms. Variants in small nuclear RNA (snRNA) genes encoding spliceosome components have recently been associated with neurodevelopmental disorders, termed "RNUopathies." We analyzed genome sequencing data from 14 patients with undiagnosed pediatric movement disorders for pathogenic variants in snRNA genes. We identified recurrent de novo RNU2-2 variants (n.35A > G and n.4G > A) in two patients with intellectual disability, epilepsy, and hyperkinetic movement disorders. RNA sequencing of fibroblasts in one patient showed no characteristic transcriptomic signature. Spliceosomopathies should be considered in neurodevelopmental disorders and developmental and epileptic encephalopathies with hyperkinetic features.

Humans↗

Integrative Transcriptomic and Proteomic Profiling Identifies S100P as a Potential Functional Biomarker for Sessile Serrated Lesions.

BACKGROUND: Sessile serrated lesions (SSLs) account for 15% of colorectal cancers (CRCs) but detection remains difficult due to flat morphology, mucinous features, and subtle histology. AIMS: This study aimed to identify novel and functionally relevant biomarkers of SSLs using transcriptomic screening and multi-omics validation. METHODS: Paired SSL and normal mucosa specimens (n = 6) underwent RNA sequencing. Differentially expressed genes (DEGs) were filtered for membrane or secretory proteins and validated across TCGA and adenoma transcriptomes. Functional significance was assessed using CRISPR dependency profiling, proteotranscriptomic concordance, pharmacogenomic sensitivity, and connectivity map analysis. RESULTS: We identified 216 upregulated genes in SSLs, including 68 encoding secretory/membrane proteins that better discriminated SSLs from controls and were enriched for adhesion and neuronal signaling while suppressing TNFα-NFκB inflammatory pathways. Cross-cohort comparison revealed five overlapping candidates between SSLs and TCGA CMS1 tumors. Among them, S100P emerged as the primary biomarker candidate, showing consistent upregulation in SSLs and CMS1 tumors while remaining low in normal mucosa and conventional adenomas. TFF1 also showed RNA-level upregulation but appeared more context-dependent. S100P demonstrated strong RNA-protein concordance in CRC cell-line profiling, supporting its detectability as a biomarker candidate. Pharmacogenomic profiling of LS411N cells revealed marked sensitivity to SN-38 and fluoropyrimidines, consistent with serrated CRC vulnerabilities. Connectivity map analysis identified perturbations, including MAPK1 and histone acetyltransferase suppression, that may reverse parts of the SSL transcriptional program. CONCLUSION: These findings prioritize S100P as a promising biomarker candidate for SSLs that warrants further validation in larger cohorts and clinically applicable platforms.

Humans↗

Omics in optic neuropathies: From molecular landscapes to personalized therapeutics.

Optic neuropathies comprise a heterogeneous group of disorders involving transient or permanent injury to retinal ganglion cells (RGCs) and their axons. Clinically, these neurodegenerative conditions manifest as dyschromatopsia, decreased visual acuity, and visual field defects, and in severe cases may ultimately lead to blindness and disability. The marked heterogeneity across disease subtypes, incompletely understood etiologies, and complex pathogenic mechanisms pose substantial challenges to precise diagnosis and effective treatment. Recent advances in omics technologies - including genomics, transcriptomics, proteomics, metabolomics, lipidomics, single-cell and spatial sequencing, and integrative multi-omics approaches - have ushered optic nerve degenerative disease research into an era of high-resolution comprehensive investigation. In this review, we summarize representative applications of omics approaches to elucidate genetic alterations, signaling dysregulation, metabolic reprogramming, and immune responses in optic neuropathies. We further discuss the emerging potential of multi-omics in identifying early diagnostic biomarkers and informing individualized therapeutic strategies. Finally, we provide a forward-looking perspective on the future trajectory of omics technologies and their prospects in both fundamental research and clinical translation, with the overarching aim of accelerating the bench-to-bedside transition in this critical eye disease field.

biomarkers↗

Transcriptome of a mouse kidney cortical collecting duct cell line: effects of aldosterone and vasopressin.

Aldosterone and vasopressin are responsible for the final adjustment of sodium and water reabsorption in the kidney. In principal cells of the kidney cortical collecting duct (CCD), the integral response to aldosterone and the long-term functional effects of vasopressin depend on transcription. In this study, we analyzed the transcriptome of a highly differentiated mouse clonal CCD principal cell line (mpkCCD(cl4)) and the changes in the transcriptome induced by aldosterone and vasopressin. Serial analysis of gene expression (SAGE) was performed on untreated cells and on cells treated with either aldosterone or vasopressin for 4 h. The transcriptomes in these three experimental conditions were determined by sequencing 169,721 transcript tags from the corresponding SAGE libraries. Limiting the analysis to tags that occurred twice or more in the data set, 14,654 different transcripts were identified, 3,642 of which do not match known mouse sequences. Statistical comparison (at P < 0.05 level) of the three SAGE libraries revealed 34 AITs (aldosterone-induced transcripts), 29 ARTs (aldosterone-repressed transcripts), 48 VITs (vasopressin-induced transcripts) and 11 VRTs (vasopressin-repressed transcripts). A selection of the differentially-expressed, hormone-specific transcripts (5 VITs, 2 AITs and 1 ART) has been validated in the mpkCCD(cl4) cell line either by Northern blot hybridization or reverse transcription-PCR. The hepatocyte nuclear transcription factor HNF-3-alpha (VIT39), the receptor activity modifying protein RAMP3 (VIT48), and the glucocorticoid-induced leucine zipper protein (GILZ) (AIT28) are candidate proteins playing a role in physiological responses of this cell line to vasopressin and aldosterone.

Aldosterone↗

[Cancer genome or the development of molecular portraits of tumors].

The rapid development of cancer genomics is due to important progresses in oncogenesis, human genome sequencing and emergence of new technologies in genome and transcriptome analysis. In this context, the aim of the French program 'Cartes d'Identites des Tumeurs--Molecular Portraits of Tumors' is to build a public data base containing a pan genome assessment of genome and transcriptome alterations in the major types of tumors as well as in relevant normal cells and experimental models. Data mining is done in the context of genome annotations and clinical and biological informations attached to the enrolled samples. The goal of the program is to define new tests useful for diagnostic procedures in clinical laboratories and new targets for biological treatments of tumors.

France↗

The chromosome-level genome of Stylosanthes guianensis provides insights into genome evolution and environmental adaptation.

Stylosanthes guianensis is a leguminous forage crop of significant economic importance, primarily distributed in tropical and subtropical regions. It exhibits strong adaptability to various stresses, yet the genetic basis underlying this trait remains unclear. In this study, we constructed the first chromosome-scale reference genome of S. guianensis using a combination of Nanopore and Hi-C sequencing technologies. The assembled genome size is 1254&#x2009;Mb, with 10 pseudochromosomes. Using Nanopore full-length transcriptome data, we generated high-quality transcript-level gene annotations, identifying 36&#x2009;585 gene models and 110&#x2009;601 transcripts. The repetitive sequences in S. guianensis account for 79.16% of the genome, with the extensive expansion of Gypsy elements in long terminal repeats contributing to its genome size enlargement. Comparative genomic and transcriptomic analyses revealed that flavonoid metabolism plays a pivotal role in stress adaptation, providing new insights into the genetic basis of stress tolerance. Additionally, we generated whole-genome methylation profiles under cold treatment and control conditions, offering valuable data for future epigenomic research. These findings provide essential molecular resources for understanding stress resilience in S. guianensis and advancing its molecular breeding.

Genome, Plant↗

Transcriptome atlases of rat brain regions and their adaptation to diabetes resolution following gastrectomy in the Goto-Kakizaki rat.

Brain regions drive multiple physiological functions through specific gene expression patterns that adapt to environmental influences, drug treatments and disease conditions. To generate a detailed atlas of the brain transcriptome in the context of diabetes, we carried out RNA sequencing in hypothalamus, hippocampus, brainstem and striatum of the Goto-Kakizaki (GK) rat model of spontaneous type 2 diabetes, which was applied to identify gene transcription adaptation to improved glycemic control following vertical sleeve gastrectomy (VSG) in the GK. Over 19,000 distinct transcripts were detected in the rat brain, including 2794 which were consistently expressed in the four brain regions. Region-specific gene expression was identified in hypothalamus (n&#x2009;=&#x2009;477), hippocampus (n&#x2009;=&#x2009;468), brainstem (n&#x2009;=&#x2009;1173) and striatum (n&#x2009;=&#x2009;791), resulting in differential regulation of biological processes between regions. Differentially expressed genes between VSG and sham operated rats were only found in the hypothalamus and were predominantly involved in the regulation of endothelium and extracellular matrix. These results provide a detailed atlas of regional gene expression in the diabetic rat brain and suggest that the long term effects of gastrectomy-promoted diabetes remission involve functional changes in the hypothalamus endothelium.

Animals↗