Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Morphological, Physiological and Transcriptomic Changes in Response to Water Deficit Stress in Brassica napus L.

Yield losses due to water-deficit (WD) conditions, especially during the reproductive stages of plant development, pose a significant threat to global canola (Brassica napus L.) production. Therefore, it is critical to investigate traits contributing to improved productivity under increased WD conditions. Here we present phenotypic, physiological and transcriptomic changes in response to WD across contrasting canola accessions exhibiting variation in drought resistance-related traits. WD significantly reduced shoot biomass, plant height, harvest index, leaf water content, photosynthetic CO2 assimilation rate, intrinsic water-use efficiency and carbon isotope discrimination. WD caused 49 to 100% of the seed yield reduction: the minimum seed yield reduction (49.66%) was observed in a doubled-haploid (DH) line, 06-5101.137, while the maximum yield reduction (94.1 to 100%) occurred in the late-flowering DH lines (06.5101.088 and 06-5101.306). Seed yield showed a positive correlation (r = 0.29 to 0.95) with shoot biomass and harvest index, leaf water content, photosynthetic CO2 assimilation rate, intrinsic water use efficiency and carbon isotope discrimination. However, it showed negative correlations with days to flower, leaf specific weight, root length, root biomass (r = -0.04 to -0.79) across water treatments. The specific leaf transcriptome analysis of the two parental lines of DH population that exhibit variation for effective water use under well-watered and water-deficient conditions revealed different categories of differentially expressed genes (DEGs): WD-responsive DEGs in BC1329 parental line (1116) and BC9102 (1205) with 754 and 853 DEGs unique to BC1329 and BC9102, respectively, WD-responsive DEGs (906), genotype-dependent DEGs (8465) and genotype × treatment interaction DEGs (353). DEG annotations revealed that the WD-treatment-affected genes were involved in stress responses and growth and development. We further located 235 DEGs within the QTL regions underlying agronomic and physiological performance. Our study provides a conceptual framework for the morphological, physiological and molecular determinants involved in water-use efficiency. Seedlings' traits with high heritability values, such as shoot biomass, leaf weight, leaf water content and Δ13C, serve as proxies for trait-based selection for improved seed yield under both water-limited and non-water-limited conditions.

Brassica napus↗

The Gene Resource Locator: gene locus maps for transcriptome analysis.

Since the advent of the draft human genome sequence there has been growing interest in transcriptome analysis based on genomic data. The Gene Resource Locator (GRL) assembles gene maps that include information on gene-expression patterns, cis-elements in regulatory regions and alternatively spliced transcripts. The database was constructed using customized software, and currently contains 2.2 million alignments (exon-intron structures). The alignments have been annotated and integrated into a system that encompasses approximately 90 000 EST loci sharing common exons, 8091 alternatively spliced transcript groups, 10 801 expression-profile groups, 8066 candidate regulatory regions in full-length cDNAs, and 1 million SNP loci. We have used Flash technology to build a dynamic web viewer that facilitates browsing through the millions of alignments. All of the information is available through the World Wide Web at the Gene Resource Locator web site (http://grl.gi.k.u-tokyo.ac.jp).

Alternative Splicing↗

Functional Annotation of the Major Histocompatibility Complex Locus.

The human major histocompatibility complex (MHC) locus has the greatest density of disease-associations in the human genome, including links to over 100 polygenic disorders. Its complex haplotype structure, rich gene density, and high degree of linkage disequilibrium combine to make deciphering the gene regulatory logic of the MHC locus extremely challenging. Employing complementary high-throughput CRISPR interference (CRISPRi) and activation (CRISPRa) epigenetic screens coupled with single-cell transcriptome profiling across three distinct human cell types, we identified hundreds of new connections between cis -regulatory elements (CREs) and their target genes in this locus. These CRE-gene links are largely cell type-specific and act as enhancers. Additionally, some CREs have complex features, including harboring both active and repressive histone marks, lacking chromatin accessibility, targeting multiple genes, or acting as silencers. Computational methods fail to predict a majority of these CRE-gene connections. These findings emphasize the potential for functional perturbation experiments to dissect complex loci and reveal shared and cell type-specific regulatory mechanisms relevant to genomics of complex diseases. Collectively, this study provides a unique resource for understanding the complex regulatory landscape within the MHC locus and supports the need for creating new models that encompass CRE-gene interactions, cell type-specific gene expression, and disease genetics in the noncoding genome.

Journal Article↗

Genome-Wide Identification and Characterization of the TBL Gene Family and Temporal Expression Dynamics During Powdery Mildew Infection in Cucumber (Cucumis sativus).

Cell-wall polysaccharide O-acetylation contributes to cell-wall assembly, organ development, and plant-pathogen interactions, but the cucumber TBL gene family remains poorly characterized. Here, 37 CsTBL genes were identified genome-wide and analyzed using phylogenetic, syntenic, conserved-motif, gene-structure, promoter, protein-structure, Gene Ontology, and transcriptome approaches, followed by RT-qPCR analysis after powdery mildew inoculation. All CsTBL proteins contained the conserved GDS and DxxH motifs, whereas accessory motifs and predicted structural features varied among clades. Intraspecific analysis identified dispersed, WGD/segmental, and tandem duplication categories, and cross-species synteny was more extensive with melon than with Arabidopsis. Homology-derived annotations associated CsTBL genes with cell-wall polysaccharide metabolism, Golgi/endomembrane compartments, and O-acetyltransferase activity, including six genes assigned to xylan O-acetyltransferase-related annotations. Expression profiling revealed tissue- and developmental-stage-dependent patterns, whereas the publicly available powdery mildew RNA-seq dataset provided descriptive temporal expression profiles in Podosphaera xanthii-inoculated samples. Independent RT-qPCR analysis using time-matched mock controls revealed distinct post-inoculation responses among six selected genes. Relative to the corresponding mock controls, CsTBL2 was consistently repressed; CsTBL15 showed transient induction at 1 dpi followed by repression; CsTBL24 exhibited a biphasic response; CsTBL25 was induced at all sampled post-inoculation time points; CsTBL26 showed progressive induction; and CsTBL30 reached its highest observed expression level at 3 dpi. Integrated functional annotation and expression evidence highlighted CsTBL26 as a priority candidate for further functional characterization, while CsTBL24 and CsTBL25 represented fruit-associated candidates with distinct powdery mildew responses; CsTBL30 remained an additional strongly infection-responsive candidate. These findings provide an evolutionary and expression-based framework for the functional characterization of the cucumber TBL gene family.

O-acetylation↗

Large-scale analysis of the human and mouse transcriptomes.

High-throughput gene expression profiling has become an important tool for investigating transcriptional activity in a variety of biological samples. To date, the vast majority of these experiments have focused on specific biological processes and perturbations. Here, we have generated and analyzed gene expression from a set of samples spanning a broad range of biological conditions. Specifically, we profiled gene expression from 91 human and mouse samples across a diverse array of tissues, organs, and cell lines. Because these samples predominantly come from the normal physiological state in the human and mouse, this dataset represents a preliminary, but substantial, description of the normal mammalian transcriptome. We have used this dataset to illustrate methods of mining these data, and to reveal insights into molecular and physiological gene function, mechanisms of transcriptional regulation, disease etiology, and comparative genomics. Finally, to allow the scientific community to use this resource, we have built a free and publicly accessible website (http://expression.gnf.org) that integrates data visualization and curation of current gene annotations.

Animals↗

Odon: an ultra-fast viewer for spatial proteomics.

MOTIVATION: Multiplexed spatial proteomics and spatial transcriptomics generate large, high-dimensional imaging datasets that are challenging to visualize efficiently, particularly at whole-slide and cohort scale. Visualization is an essential step for rapid detection of staining artefacts, such as protein aggregates or non-specific staining. RESULTS: Here, we present Odon, a native Rust desktop viewer designed for rapid, interactive exploration of multiplex imaging data on a standard laptop. Odon is primarily built around the OME-Zarr imaging format, and supports annotations via GeoJSON and GeoParquet, with secondary support for SpatialData, Xenium containers, and TIFF. Data can be stored locally or streamed directly from HTTP or S3-compatible object storage using viewport-driven tile loading. Odon incorporates a highly optimized rendering engine designed for viewport-driven tile loading and GPU-based compositing. In scripted benchmarks using synthetic multiplex OME-Zarr datasets, Odon showed lower peak memory use, lower affine-derived zoom-step error, and faster warm-start image loading than napari and QuPath under the tested conditions. Its GPU-based compositing pipeline also enables smooth rendering and interaction with >1 000 000 segmented cells. Odon further supports integrated visual analytics, including live thresholding and cell selection, and a mosaic mode for simultaneous viewing of hundreds of regions of interest in cohort and tissue microarray studies. Together, these features establish Odon as a high-performance platform for scalable visualization of spatial proteomics data. AVAILABILITY AND IMPLEMENTATION: Source code and compiled installers are available at https://github.com/alexcoulton/odon.

Proteomics↗

Colorectal Liver Metastasis Pathomics Model: Integrating Single-Cell and Spatial Transcriptome Analysis With Pathomics for Predicting Liver Metastasis in Colorectal Cancer.

The liver is the primary target organ for hematologic metastasis of colorectal cancer (CRC), and CRC liver metastasis (CRLM) often precludes radical resection, making it the leading cause of death in patients with CRC. To improve the identification and prediction of liver metastasis risk, we identified a cell type of liver metastasis--triggering malignant cells (LMTMCs) through integrating single-cell RNA sequencing and spatial transcriptome analysis. Multiomics cell communication analysis indicated that the interaction between fibroblasts and LMTMCs through the COL1A1-CD44/SDC4 and LAMA4-CD44 signaling axes could promote CRLM. By applying the one-class logistic regression algorithm, we developed a CRLM scoring system in the bulk RNA-sequencing data according to the abundance of LMTMCs in each individual. Using the grouping labels derived from the CRLM scoring system in the bulk data and the corresponding whole-slide images without any manual annotations at the region or pixel level, processed via slide-level weakly supervised learning, a deep-learning model based on the ResNet18 architecture, called Colorectal Liver Metastasis Pathomics Model, was developed to predict the risk of liver metastasis in patients with CRC. The Colorectal Liver Metastasis Pathomics Model achieved an area under the curve of 0.84 at the internal test set of The Cancer Genome Atlas-CRC histology images. In the external independent validation sets, namely the Affiliated Hospital of Southwest Medical University and the Affiliated Traditional Chinese Medicine Hospital of Southwest Medical University cohorts, the areas under the curve were 0.89 and 0.72, respectively, indicating effective classification performances. This study provided new insights and tools for the early identification of CRLM and demonstrated the potential of combining multiomics with deep learning-based pathomics in cancer research.

Humans↗

Identifying Co-Expressed lncRNAs Correlated With Traits of Interest in an Animal Model for Metabolic Diseases in Humans.

Nutrigenomics investigates how nutrients modulate gene expression. Among them, fatty acids (FA) play important roles in regulating gene transcription, while long non-coding RNAs (lncRNAs) may be associated with gene regulation and metabolic diseases. This study aimed to analyze the hepatic transcriptome of pigs, a species frequently used as a model for nutrigenomic studies, to identify novel lncRNAs and their potential target genes in response to diets containing different sources of FA. Seventy-two pigs were fed four diets supplemented with 1.5% soybean oil (control), 3% canola oil, 3% fish oil, and 3% soybean oil. RNA sequencing of liver samples was performed to identify novel lncRNAs. Weighted Gene Co-expression Network Analysis (WGCNA) was used to identify modules associated with phenotypic traits related to lipid metabolism and inflammation. Functional enrichment analyses were then conducted to annotate genes within these modules using Gene Ontology (GO) terms and to assess overlap with Quantitative Trait Loci (QTL). The results revealed 106 novel lncRNAs potentially regulating genes associated with lipid metabolism and immune responses in pigs fed diets with different FA sources. These findings enhance understanding of the regulatory role of lncRNAs in pigs and reinforce their relevance as models for human metabolic diseases.

Animals↗

Multi-omics analysis reveals distinct spatial compartmentalization of lung repair niches in pediatric ARDS.

BACKGROUND: Pediatric acute respiratory distress syndrome (PARDS), often triggered by viral infections, is a life-threatening condition. Despite its severity, children demonstrate significantly better survival rates and superior lung repair compared to adults. However, the mechanisms underlying this age-specific advantage remain incompletely understood. PATIENTS AND METHODS: We conducted a pilot multi-omics study of influenza-associated PARDS integrating single-cell RNA sequencing (scRNA-seq) of pediatric lung tissue and bronchoalveolar lavage fluid (BALF), spatial transcriptomics, and plasma proteomics. Analyses were harmonized with the Human Lung Cell Atlas (HLCA) reference, reanalysis of public pediatric PARDS airway scRNA-seq, and contextual comparisons to adult lethal COVID-19 lung. RESULTS: Tissue scRNA-seq and spatial data indicated outcome-linked divergence in PARDS. Survivor showed spatially restricted repair with preserved alveolar type II (AT2) cells, AT2-to-alveolar type I (AT1) differentiation signatures, and higher KRT17, whereas fatal case and adults exhibited diffuse immune activation with pro-fibrotic and pro-apoptotic signaling. In BALF, KRT17-positive airway stress–repair epithelial cells (hillock-like) increased from the acute to recovery phase, and plasma proteomics showed higher circulating KRT17 in survivors. HLCA-based label transfer strengthened cell-type definitions and enabled pediatric–adult comparisons suggesting biological and developmental differences; the adult lethal COVID-19 atlas provided a benchmark with attenuated epithelial repair and prominent collagen CTHRC1-pathologic fibroblasts. Fibroblast programs were regionally compartmentalized, with injury-enriched CTHRC1+ states versus alveolar fibroblasts in preserved areas, and showed stronger injury–homeostasis anti-correlation in fatalities. Myeloid remodeling included BALF transitions from FCN1-high inflammatory states toward FABP4-positive resident-like states, consistent with public pediatric datasets showing reduced inflammatory and interferon-stimulated gene (ISG) modules and severity-linked increases in aged neutrophils. CONCLUSIONS: This pilot multi-omics case series outlines putative pediatric lung repair niches in influenza-associated PARDS. KRT17-positive transitional epithelium, preserved AT2 differentiation, and restoration of resident-like macrophages may align with recovery, whereas diffuse immune activation and CTHRC1-enriched fibroblast programs may accompany worse outcomes. HLCA-guided annotations and adult benchmarks indicate possible age-related differences, warranting validation in larger multi-center cohorts.

Humans↗

ARISE: RNA-anchored shared-edge topology and hierarchical fusion for spatial multi-omics integration.

MOTIVATION: Spatial multi-omics technologies jointly profile transcriptomes, proteins and chromatin accessibility in situ, enabling integrative analysis of tissue organization across molecular layers. However, most existing graph-based integration methods rely on independently constructed modality-specific k-nearest-neighbor graphs. When auxiliary modalities are sparse or noisy, these graphs can become topologically discordant, propagate spurious edges, weaken cross-modal alignment, and reduce spatial domain resolution. RESULTS: We present Anchored RNA for Integrated Spatial Embedding (ARISE), an RNA expression anchored framework for spatial multi-omics integration. ARISE defines a shared-edge topology by intersecting RNA feature-similarity and spatial-proximity graphs, encodes auxiliary modalities on this common scaffold, and integrates them through inside-out hierarchical fusion. We further show theoretically that graph intersection minimizes false-positive edges within a broad class of k-of-r graph fusion rules, providing a principled basis for topology anchoring. Across various spatial multi-omics benchmarks spanning simulated and real datasets in bi-modal and tri-modal settings, ARISE improves spatial domain identification, cross-modal consistency, and preservation of tissue structure relative to existing methods. Furthermore, the learned representation supports biologically meaningful downstream analyses, including marker-based domain annotation, pathway enrichment, and cis-regulatory inference, indicating that ARISE yields a robust and interpretable framework for spatial multi-omics integration. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/XiangxiangWang-code/ARISE. The archived version used in this study is available at https://doi.org/10.6084/m9.figshare.32686137.v2.

Multiomics↗

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis↗

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32 Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59 Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant↗

Identifying JAK2 and ANXA5 as Key Genes Linking Obstructive Sleep Apnea and Oxidative Stress via Machine Learning and Multilayer Transcriptomic Integration With Functional Validation.

Obstructive sleep apnea (OSA) is a common and severe sleep disorder closely associated with oxidative stress (OS). This study aims to identify and validate potential OS-related genes associated with OSA through bioinformatics methods. We successfully identified OS-related differentially expressed genes (OS-DEGs) by combining the limma test, weighted correlation network analysis (WGCNA), and OS-related genes from the GeneCards database. Key genes and potential biological roles were further identified using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), enrichment analysis, protein-protein interaction (PPI) network analysis, Lasso regression analysis, random forest algorithm, and support vector machine recursive feature elimination (SVM-RFE) method. Evaluate and validate the accuracy of key genes through receiver operating characteristic (ROC) curve analysis. The human single-cell RNA sequencing (scRNA-seq) dataset is used for cell classification annotation, analysis of key gene single-cell expression profiles, and virtual gene knockout experiments based on the scTenifoldKnk algorithm. Integrating scRNA-seq sequencing, pseudotime trajectory inference, cell-cell communication analysis, and bulk immune infiltration deconvolution reveals monocyte subtype remodeling in OSA. Finally, the expression levels of key genes in clinical samples were validated using real-time quantitative PCR (RT-qPCR) and Western blotting. A total of 57 common DEGs, indicating significant enrichment in OS, inflammation, and tumor pathways, particularly prominent in the immunometabolism pathway. By integrating DEGs, WGCNA, PPI results, and machine learning methods, key genes Janus kinase 2 (JAK2) and ANXA5 were screened out. JAK2 was significantly upregulated under disease conditions, while ANXA5 was significantly downregulated. ROC curve exhibited high accuracy (area under the curve [AUC] > 0.85). Human scRNA-seq analysis revealed that key genes were predominantly highly expressed in monocytes. Virtual knockout experiments demonstrated that these key genes play a crucial role in regulating immune responses and inflammatory reactions. PPI networks and enrichment analysis verified that downstream genes S100P, ALOX5AP, PROK2, and PADI4 may collaboratively participate in immune response and inflammation regulation. Finally, clinical sample experiment further validated the results of bioinformatics analysis. This study provides new research insights for the diagnosis, mechanism research, and treatment development of OSA in the future by integrating multilayer transcriptomic and machine learning techniques.

Humans↗

Transcriptome analysis of the diseased intervertebral disc tissue in patients with spinal tuberculosis.

OBJECTIVE: To investigate the differential expression genes (DEGs) in spinal tuberculosis using transcriptomics, with the aim of identifying novel therapeutic targets and prognostic indicators for the clinical management of spinal tuberculosis. METHODS: Patients who visited the Department of Orthopedics at the Second Hospital, Lanzhou University from January 2021 to May 2023 were enrolled. Based on the inclusion and exclusion criteria, there were 5 patients in the test group and 5 patients in the control group. Total RNA was extracted and paired-end sequencing was conducted on the sequencing platform. After processing the sequencing data with clean reads and annotating the reference genome, FPKM normalization and differential expression analysis were performed. The DEGs and long non-coding RNAs (LncRNAs) were analyzed for Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Ontology (GO) enrichment. The cis-regulation of differentially expressed mRNAs (DE mRNAs) by LncRNAs was predicted and analyzed to establish a co-expression network. RESULTS: This study identified 2366 DEGs, with 974 genes significantly upregulated and 1392 genes significantly downregulated. The upregulated genes are associated with cytokine-cytokine receptor interactions, tuberculosis, and TNF-α signaling pathways, primarily enriched in biological processes such as immunity and inflammation. The downregulated genes are related to muscle development, contraction, fungal defense response, and collagen metabolism processes. Analysis of LncRNAs from bone tuberculosis RNA-seq data detected a total of 3652 LncRNAs, with 356 significantly upregulated and 184 significantly downregulated. Further analysis identified 311 significantly different LncRNAs that could cis-regulate 777 target genes, enriched in pathways such as muscle contraction, inflammatory response, and immune response, closely related to bone tuberculosis. There are 51 genes enriched in the immune response pathway regulated by cis-acting LncRNAs. LncRNAs that regulate immune response-related genes, such as upregulated RP11-451G4.2, RP11-701P16.5, AC079767.4, AC017002.1, LINC01094, CTA-384D8.35, and AC092484.1, as well as downregulated RP11-2C24.7, may serve as potential prognostic and therapeutic targets. CONCLUSION: The DE mRNAs and LncRNAs in spinal tuberculosis are both associated with immune regulatory pathways. These pathways promote or inhibit the tuberculosis infection and development at the mechanistic level and play an important role in the process of tuberculosis transferring to bone tissue.

Humans↗

Insights into dill (Anethum graveolens) flavor formation via integrative analysis of chromosomal-scale genome, metabolome and transcriptome.

INTRODUCTION: Dill (Anethum graveolens) is a significant medicinal herb belonging to the Apiaceae family. Owing to its high levels of volatile organic compounds (VOCs), dill is commonly utilized for essential oil extraction and medicine purpose. However, the biosynthesis of the crucial VOC in dill remains obscure. OBJECTIVES: Identify the key VOCs related to the flavor formation in dill and dissect the regulatory mechanism of their synthesis. METHODS: The dill chromosomal-level genome was constructed by PacBio HiFi, Hi-C, and BGISEQ second generation sequencing and assembly. The VOCs in dill leaves were identified through GC-MS. The potential mechanism involved in regulating the VOC accumulation in dill flavor formation was analyzed by multi-omics analysis. RESULTS: A 1.17 Gb chromosome-scale genome of dill with a contig N50 of 10.78 Mb was constructed. A total of 46,538 genes were annotated across 11 assembled chromosomes. Comparative genomics analysis suggested that transposable element insertions, especially LTR-Gypsy, have contributed to the evolution and expansion of the dill genome. The flavor formation of dill was mainly attributed to terpenoids, especially α-phellandrene, β-ocimene, and o-cymene. The contribution of expansion and replication of terpenoid synthesis pathway genes, especially terpene synthase (TPS), to the abundant terpenoid production of dill was identified. Differential gene expression patterns observed at various developmental stages and tissues provided key candidate genes for the regulation of terpenoid synthesis, as well as transcription factors. The different accumulation of esters and aromatics also affected the flavor formation of dill. The key genes implicated in the synthesis of anethole, namely AIS and AMT were further identified. CONCLUSION: This study constructed the chromosome level genome and identified the main VOCs and related key genes in flavor formation of dill, shedding lights on our understanding of terpenoid biosynthesis but also offered guidance for future genetic research on molecular breeding in Anethum graveolens.

Transcriptome↗

ceRNA network of lncRNAs and mRNAs in OSF-to-OSCC progression: Diagnostic biomarkers and functional pathways.

BACKGROUND: Oral submucous fibrosis (OSF) is a chronic potentially malignant disorder that can progress to oral squamous cell carcinoma (OSCC). Although dysregulated non-coding RNAs have been implicated in oral carcinogenesis, the competing endogenous RNA (ceRNA)-mediated regulatory mechanisms underlying OSF-to-OSCC progression remain poorly understood. This study aimed to identify candidate regulatory molecules and construct a putative lncRNA-miRNA-mRNA network associated with malignant transformation. METHODS: Publicly available microarray datasets (GSE117973 and GSE125866) were analyzed to identify differentially expressed genes between OSF and OSCC. Differentially expressed transcripts were classified into mRNAs and lncRNAs based on public transcript annotations. Highly correlated lncRNA-mRNA pairs were identified using Pearson correlation analysis and integrated with multiMiR-supported miRNA-mRNA interactions obtained from public databases to construct a putative ceRNA regulatory network. Functional characterization focused on apoptosis, epithelial-mesenchymal transition (EMT), and immune checkpoint-related pathways. Receiver operating characteristic (ROC) analysis was performed to evaluate diagnostic performance, and selected biomarkers were externally validated using The Cancer Genome Atlas (TCGA) OSCC cohort. RESULTS: Integrated transcriptomic analysis identified several dysregulated mRNAs and lncRNAs associated with OSF-to-OSCC progression. Network analysis highlighted TBC1D3B, RREB1, TEAD3, SREBF1, TMEM41B, FOXK2, and KIAA1958 as prominent hub genes within the putative regulatory network. Functional analyses demonstrated significant associations with apoptosis-, EMT-, and immune checkpoint-related genes, suggesting potential involvement in multiple biological processes contributing to malignant transformation. Several hub genes exhibited strong diagnostic performance, with ROC analysis yielding AUC values ranging from 0.891 to 1.000, indicating excellent discrimination between OSF and OSCC samples. External validation using TCGA further supported the relevance of the identified biomarkers in OSCC. CONCLUSIONS: This study provides a comprehensive transcriptomic framework describing putative lncRNA-miRNA-mRNA regulatory interactions associated with OSF progression to OSCC. The identified hub genes and regulatory networks represent candidate biomarkers for early detection and provide a foundation for future mechanistic and experimental validation. As the proposed ceRNA interactions are computationally inferred, further biological validation is required before clinical application.

RNA, Long Noncoding↗

[DNA arrays: technological aspects and applications].

The Human Genome Project has allowed considerable progress in the construction of physical and genetic maps and the identification of genes involved in human sicknesses. The accelerated accumulation of biological information and knowledge is due in large part to the sequencing projects of other organisms, which in fact paved the way for the Human Genome Project. In parallel, recently developed techniques which take advantage of genomic sequences allow large scale molecular analyses resulting in the functional annotation of many of the proteins represented by these genes. This is the goal of functional genomics. These progresses are at the origin of the present revolution in biomedical research. DNA microarrays are playing a dominant role compared to the other developing technologies since they are relatively easy to make and use and are applicable to numerous scientific inquiries. They allow the simultaneous analysis of several thousands of genes in biological samples from sick or healthy tissues, at the genome or transcriptome level. The data obtained is expected to result in major advances in the health sciences. In addition to an improved understanding of the complex molecular interaction networks of healthy cells and tissues, a more precise genetic characterization of the molecular mechanisms involved in pathology should result in the identification of new therapeutic targets and the development of new medicines. The genetic profiles thus obtained should also permit the definition of new pathologic subclasses not recognizable by traditional clinical factors, as well as new markers for susceptibility to certain illnesses, and new prognostic markers or methods of predicting responses to treatment. In this article, we present the different approaches and potential applications of DNA microarray technology, in particular as applied to cancer research.

Chromosome Mapping↗

Identification of a putative RocS homolog through phenotypic profiling of uncharacterized essential genes in Streptococcus mutans.

Genome-wide viability catalogs produced by transposon sequencing (Tn-seq) and CRISPR interference (CRISPRi) have successfully mapped the essential genome of Streptococcus mutans . In this study, we combined predictive bioinformatics, conditional CRISPRi transcriptional silencing, transmission electron microscopy, transcriptomics, and genetic suppressor screens to investigate nine poorly characterized essential genes in S. mutans . From this screen, phenotypic and genetic analyses identified SMU_393 as a functional homolog of the pneumococcal chromosome segregation factor, RocS. Depletion of SMU_393 resulted in abnormal cell widening, hypersensitivity to DNA damage, and a significant subpopulation of anucleate cells. These phenotypes were bypassed by a spontaneous surface-exposed missense mutation ( dnaA Q197E ) within the AAA+ ATPase domain of the replication initiator. Together, this study refines annotations within the S. mutans essential genome and provides genetic insights into streptococcal chromosome segregation and cell cycle control.

Journal Article↗