Search PubMedSearch

SEARCH · Search PubMed

Results for “Transcriptome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Transcriptome mining and comparative genomics reveal 36 putative novel marafivirus species and conserved evolution of the marafibox regulatory element.

BACKGROUND: Marafiviruses are plant-infecting RNA viruses associated with several economically important crops, but their genomic diversity remains incompletely characterized. OBJECTIVE: This study aimed to identify previously unrecognized marafivirus genomes and investigate their genomic features and evolutionary relationships. METHODS: Publicly available plant transcriptome datasets were systematically mined to detect marafivirus-like sequences. Recovered genomes were analyzed using comparative sequence analysis, phylogenetic reconstruction, and genome organization characterization. RESULTS: A total of 62 marafivirus-like genomes were recovered from 33 independent sources representing diverse plant hosts. Polyprotein-based comparative and phylogenetic analyses grouped these genomes into 36 lineages likely representing novel species. All newly identified viruses clustered within the Marafivirus clade. Genome organization analysis revealed conserved polyprotein architecture and widespread presence of the marafibox promoter element. Conservation of additional open reading frames among closely related isolates aided identification of potentially functional genes. CONCLUSION: These findings substantially expand the known diversity of marafiviruses and demonstrate the effectiveness of transcriptome mining for discovering previously unrecognized plant viruses.

Phylogeny

Diverse RNA viruses discovered in multiple seagrass species.

Seagrasses are marine angiosperms that form highly productive and diverse ecosystems. These ecosystems, however, are declining worldwide. Plant-associated microbes affect critical functions like nutrient uptake and pathogen resistance, which has led to an interest in the seagrass microbiome. However, despite their significant role in plant ecology, viruses have only recently garnered attention in seagrass species. In this study, we produced original data and mined publicly available transcriptomes to advance our understanding of RNA viral diversity in Zostera marina, Zostera muelleri, Zostera japonica, and Cymodocea nodosa. In Z. marina, we present evidence for additional Zostera marina amalgavirus 1 and 2 genotypes, and a complete genome for an alphaendornavirus previously evidenced by an RNA-dependent RNA polymerase gene fragment. In Z. muelleri, we present evidence for a second complete alphaendornavirus and near complete furovirus. Both are novel, and, to the best of our knowledge, this marks the first report of a furovirus infection naturally occurring outside of cereal grasses. In Z. japonica, we discovered genome fragments that belong to a novel strain of cucumber mosaic virus, a prolific pathogen that depends largely on aphid vectoring for host-to-host transmission. Lastly, in C. nodosa, we discovered two contigs that belong to a novel virus in the family Betaflexiviridae. These findings expand our knowledge of viral diversity in seagrasses and provide insight into seagrass viral ecology.

RNA Viruses

MED12-STAT1-TAP2 axis regulates CD8 + T cell cytotoxicity and mediates immunotherapy outcome in non-small cell lung cancer.

Although immunotherapy for late-stage non-small cell lung carcinoma (NSCLC) has been clinically utilized, its prognosis remains highly heterogeneous, prompting us to investigate novel predictive immunotherapy biomarkers for NSCLC. We analyzed the correlations between MED12 nonsynonymous mutations and survival, clinical, genomic, transcriptomic information, and immune infiltration information through data mining across multiple datasets. We also investigated the mechanism of MED12 using luciferase assay, Western blot, ChIP-PCR, and siRNA. MED12 is significantly associated with survival in completely independent immunotherapy datasets, including MSKCC (N = 350), Naiyer2015 (N = 34), our own (N = 295) and the pan-cancer dataset, but not in the TCGA dataset, where patients received non-immunotherapy regimens. Mutations in MED12 showed no significant correlation with known metrics (TMB, IPS/CTLA4/PD1 status, PD-1/PD-L1 expression, and TCR/BCR status) or DNA Damage Repair (DDR) pathway mutations, yet they carried independent prognostic information according to the Cox multivariate regression. On the other hand, MED12 mutation is significantly associated with multiple immune-related pathways and immune infiltration of CD8 + T cells and activated NK cells. Lactate dehydrogenase assay revealed that knockdown of TAP2 restored the upregulation of CD8 + T cell cytotoxicity triggered by MED12 knockdown. ChIP-PCR, luciferase assay and siRNA knock down assay indicate that MED12 binds to the promoter region of STAT1 to suppress its transcription, while the transcription factor STAT1 promotes the transcription of TAP2, thus inhibiting the antigen processing and presentation. Collectively, MED12 mutation is an independent and valuable biomarker for predicting the response to immune checkpoint inhibitor (ICI)therapy in NSCLC by modulating CD8 + T cell cytotoxicity via the STAT1/TAP2 axis.

Humans

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus

eVOC: a controlled vocabulary for unifying gene expression data.

Expression data contribute significantly to the biological value of the sequenced human genome, providing extensive information about gene structure and the pattern of gene expression. ESTs, together with SAGE libraries and microarray experiment information, provide a broad and rich view of the transcriptome. However, it is difficult to perform large-scale expression mining of the data generated by these diverse experimental approaches. Not only is the data stored in disparate locations, but there is frequent ambiguity in the meaning of terms used to describe the source of the material used in the experiment. Untangling semantic differences between the data provided by different resources is therefore largely reliant on the domain knowledge of a human expert. We present here eVOC, a system which associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. We have curated and annotated 7016 cDNA libraries represented in dbEST, as well as 104 SAGE libraries,with expression information,and provide this as an integrated, public resource that allows the linking of transcripts and libraries with expression terms. Both the vocabularies and the vocabulary-annotated libraries can be retrieved from http://www.sanbi.ac.za/evoc/. Several groups are involved in developing this resource with the aim of unifying transcript expression information.

Animals

Metabolic Dysregulation of the Lysophospholipid/Autotaxin Axis in the Chromosome 9p21 Gene SNP rs10757274.

BACKGROUND: Common chromosome 9p21 single nucleotide polymorphisms (SNPs) increase coronary heart disease risk, independent of traditional lipid risk factors. However, lipids comprise large numbers of structurally related molecules not measured in traditional risk measurements, and many have inflammatory bioactivities. Here, we applied lipidomic and genomic approaches to 3 model systems to characterize lipid metabolic changes in common Chr9p21 SNPs, which confer ≈30% elevated coronary heart disease risk associated with altered expression of ANRIL, a long ncRNA. METHODS: Untargeted and targeted lipidomics was applied to plasma from NPHSII (Northwick Park Heart Study II) homozygotes for AA or GG in rs10757274, followed by correlation and network analysis. To identify candidate genes, transcriptomic data from shRNA downregulation of ANRIL in HEK-293 cells was mined. Transcriptional data from vascular smooth muscle cells differentiated from induced pluripotent stem cells of individuals with/without Chr9p21 risk, nonrisk alleles, and corresponding knockout isogenic lines were next examined. Last, an in-silico analysis of miRNAs was conducted to identify how ANRIL might control lysoPL (lysophosphospholipid)/lysoPA (lysophosphatidic acid) genes. RESULTS: Elevated risk GG correlated with reduced lysoPLs, lysoPA, and ATX (autotaxin). Five other risk SNPs did not show this phenotype. LysoPL-lysoPA interconversion was uncoupled from ATX in GG plasma, suggesting metabolic dysregulation. Significantly altered expression of several lysoPL/lysoPA metabolizing enzymes was found in HEK cells lacking ANRIL. In the vascular smooth muscle cells data set, the presence of risk alleles associated with altered expression of several lysoPL/lysoPA enzymes. Deletion of the risk locus reversed the expression of several lysoPL/lysoPA genes to nonrisk haplotype levels. Genes that were altered across both cell data sets were DGKA, MBOAT2, PLPP1, and LPL. The in-silico analysis identified 4 ANRIL-regulated miRNAs that control lysoPL genes as miR-186-3p, miR-34a-3p, miR-122-5p, and miR-34a-5p. CONCLUSIONS: A Chr9p21 risk SNP associates with complex alterations in immune-bioactive phospholipids and their metabolism. Lipid metabolites and genomic pathways associated with coronary heart disease pathogenesis in Chr9p21 and ANRIL-associated disease are demonstrated.

Chromosomes, Human, Pair 9

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease

Allelic variation and light-responsive regulation of FaMYB10-2 underlie tissue-specific anthocyanin accumulation in strawberry.

Anthocyanins critically determine fruit color, nutrition, and stress resilience in cultivated strawberry (Fragaria &#xd7; ananassa), directly influencing consumer preference. Despite complex genetic and environmental regulation of their biosynthesis, the basis for tissue-specific pigmentation, notably the widespread occurrence of red skin and pale flesh, remains poorly understood. We integrated genomic, transcriptomic, and functional analyses across 200 cultivars to dissect receptacle pigmentation regulation. Approaches included FaMYB10-2 allele mining, promoter structural variant (SV) identification, expression profiling, regulatory interaction assays, and characterization of upstream light-responsive factors. FaMYB10-2 was identified as the key R2R3-MYB regulator of fruit anthocyanin biosynthesis. Alleles FaMYB10-2.2 and FaMYB10-2.3 encode truncated proteins retaining bHLH-binding capacity but lacking activation domains, functioning as dominant-negative repressors. A promoter SV 986&#x2005;bp upstream of FaMYB10-2 was associated with reduced pale fruit due to cis-regulatory divergence. The SV (Alt) allele is prevalent in Asian cultivars, while the Ref allele is enriched in Western germplasm. Crucially, a light-responsive FaHYH-FaWRKY71 cascade activates FaMYB10-2 and structural genes haplotype-dependently, compensating for weak MYB activity in the skin. Our findings reveal a multilayered regulatory system integrating allelic variation, cis-regulatory divergence, and environmental signals, advancing anthocyanin understanding and providing engineering targets for polyploid crop color improvement.

Fragaria

Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health record (EHR) data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; nine tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations (SHAP) identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct "subtissues" (clusters of samples); and gene-gene co-expression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six FDA-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large U.S. de-identified insurance-claims database (n = 364733), exposure to promethazine, one of the candidate drugs, was associated with a 57-62 % lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both p < 0.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multi-omics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Computational Biology

Expression regulation network in papillae of sea cucumbers: Whole-transcriptome and DNA methylation datasets.

To elucidate the expression regulation network of papilla size of sea cucumbers (Apostichopus japonicus), the whole-transcriptome and DNA methylome datasets of different sizes of papillae in sea cucumbers were generated. Average clean bases of whole-transcriptome (16.35&#x2009;G) and DNA methylome (28.92&#x2009;G) were obtained using RNA sequencing and whole-genome bisulfite sequencing techniques. A total of 3,188 ceRNA networks were also identified including 3,081 long non-coding RNAs (lncRNA)/microRNAs (miRNA)/mRNA networks and 107 circular RNA (circRNA)/miRNA/mRNA networks. Methylome data indicate that there were 3,307 and 3,776 differentially methylated regions (DMRs) with high-level methylation as well as 3,125 and 3,016 DMRs with low-level methylation in big papillae compared to small papillae. The identified DMRs were mainly distributed in introns, promotors, or exons. The whole-transcriptome and DNA methylome datasets generated from this study not only established a robust theoretical foundation (especially from the epigenetic aspect) for elucidating expression regulation network determining papilla size in sea cucumbers but also can be a valuable resource of biomarker mining for papilla appearance-based selective breeding in sea cucumbers.

DNA Methylation

Multi-Omics and Integrative Analytics in Natural Products Discovery.

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Biological Products

Identification of Novel Wraparound Transcripts in JC Polyomavirus.

JC polyomavirus (JCPyV) is a ubiquitous pathogen that causes progressive multifocal leukoencephalopathy (PML). Although a recent study using next-generation sequencing (NGS) provided detailed transcriptome atlases for polyomaviruses (PyVs) such as BK polyomavirus and simian virus 40, the transcriptome of JCPyV remains poorly characterized. Here, we conducted a comprehensive analysis using both short-read and long-read NGS technologies to construct a transcriptome atlas of JCPyV. RNA extracted from IMR-32 and HEK293 cells transfected with the circular JCPyV genome was analyzed, leading to the identification of 39 previously uncharacterized viral transcripts in addition to 12 known ones. Among the novel transcripts, we identified wraparound transcripts, conserved across PyVs, which are generated through continuous, multicyclic transcription of the circular viral genome. These included both late transcripts containing leader-to-leader repeated sequences and SuperT transcripts with multiple LxCxE motifs. Notably, wraparound transcripts, including SuperT transcripts, were also detected in brain tissues from PML patients. Collectively, this study significantly expands our understanding of the JCPyV transcriptome, revealing the expression of wraparound transcripts in PML lesions. These findings provide valuable insights into the molecular basis of JCPyV gene expression and PML pathogenesis, potentially facilitating the development of effective countermeasures against PML.

JC Virus

Gut microbiota-derived metabolites target C5AR1/KDM2A/HCAR3 axis in inflammatory bowel disease: a multi-machine learning algorithms and molecular docking study.

BACKGROUND: Inflammatory bowel disease (IBD) is a chronic recurrent disorder. Gut microbiota-derived metabolites regulate intestinal homeostasis, but their molecular mechanisms in IBD remain unclear. Current studies lack systematic "microbiota-metabolite-target" network mining with multi-method validation. This study integrates network pharmacology, three machine learning algorithms, and molecular docking to construct this regulatory network in IBD. METHODS: Transcriptome data were obtained from the Gene Expression Omnibus (GEO) database. Differentially expressed genes (DEGs) were identified using limma (p < 0.05, |log2FC| > 0.5). Weighted gene co-expression network analysis (WGCNA) with an optimal soft threshold of &#x3b2; = 7 was performed to identify key module genes. Candidate genes were obtained by intersecting DEGs, gut microbiota-associated genes from the gutMGene database, and WGCNA module genes. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were conducted to explore the functional roles of candidate genes. Core genes were identified using three machine learning algorithms (LASSO, Boruta, and SVM-RFE), followed by protein-protein interaction (PPI) network analysis. Molecular docking was performed to assess the binding affinities between hub proteins and gut microbiota-derived metabolites. RESULTS: A total of 885 DEGs were identified between the IBD and control groups, including 463 upregulated and 422 downregulated genes. WGCNA identified 280 key module genes from the purple and yellow modules. The intersection of DEGs, gut microbiota-associated genes, and WGCNA module genes yielded 19 core candidate genes. PPI network analysis combined with three machine learning algorithms jointly identified C5AR1, KDM2A, and HCAR3 as core hub genes. ROC curve analysis demonstrated that all three hub genes achieved AUC values greater than 0.7 in both the training and validation sets, indicating excellent diagnostic performance for IBD. Enrichment analysis revealed significant associations with the TNF, NF-&#x3ba;B, and IL-17 signaling pathways. Molecular docking confirmed stable binding of C5AR1 with 1,3-Diphenylpropan-2-Ol (-7.87 &#xb1; 0.83 kcal&#xb7;mol-&#xb9;) and HCAR3 with 3-Indolepropionic Acid (-6.35 &#xb1; 0.70 kcal&#xb7;mol-&#xb9;), both below -5.0 kcal&#xb7;mol-&#xb9;. CONCLUSION: This study first constructs a "gut microbiota-metabolite-hub gene" axis in IBD, providing a computational framework for microbiota-targeted precision therapy, and identifying C5AR1/KDM2A/HCAR3 as computationally predicted diagnostic biomarkers and 1,3-Diphenylpropan-2-Ol/3-Indolepropionic Acid as candidate intervention molecules that warrant further experimental validation.

Molecular Docking Simulation

Genome-wide characterisation of the myosin light chain gene family in Chinese perch (Siniperca chuatsi) and its expression patterns in muscle fibre types and injury response.

The Class II myosin light chain (myl) genes in Chinese perch (Siniperca chuatsi) have not yet been systematically characterised, and relationships with muscle fibre specification, development, and injury-associated remodelling remain unclear. In this study, fast and slow muscle fibres were initially distinguished using myofibrillar ATPase histochemistry. Subsequently, genome-wide mining identified 16 Class II myl genes, comprising eight essential and eight regulatory light-chain subunits. Their conserved-domain features, chromosomal distribution, phylogenetic relationships and expression profiles were analysed. Transcriptomic profiling showed that summed myl transcript abundance was higher in fast muscle than in slow muscle, accounting for 67.2% of the pooled myl transcript pool across the two muscle types (paired t-test, raw P&#xa0;=&#xa0;0.036). mylpfa, myl1 and mylz3 were the major fast-muscle-associated genes, whereas myl10, myl2b and myl13 were preferentially expressed in slow muscle at the transcript level. These patterns support these genes as candidate fibre-type-associated expression markers. Developmental profiling identified stage-associated myl expression patterns, including a possible expression shift between mylpfb and mylpfa. In the descriptive injury-repair time course (d0-d7), FPKM profiles indicated that fast-muscle-associated genes (mylpfa, mylz3 and myl1) were lower at d1 and recovered by d3, whereas several slow-muscle-associated genes showed biphasic transcript-level increases. The slow-muscle-associated RLC gene mylpfb showed a delayed expression peak at d7. Notably, the embryonic isoform myl6l showed a modest increase from approximately 2 FPKM at d0 to 4-5 FPKM after injury, suggesting a possible injury-associated expression pattern that requires further validation. Together, these findings provide a genome-wide description of the Chinese perch myl gene family and identify candidate fibre-type-associated genes and descriptive injury-associated isoform expression patterns.

Animals

A transcriptome-wide approach for rapid pathotype discrimination of Puccinia striiformis f. sp. tritici in north-western India.

Stripe rust of wheat caused by Puccinia striiformis f. sp. tritici (Pst) remains a major constraint to wheat production in India due to the rapid evolution and frequent emergence of virulent pathotypes. Rapid and reliable discrimination of Pst pathotypes is essential for effective resistance deployment and surveillance. In the present study, transcriptome-wide simple sequence repeats (SSRs) and single nucleotide polymorphisms (SNPs) were exploited to develop and validate molecular markers for pathotype-specific detection of Pst pathotypes prevalent in North India (110S119, 238S119, 46S119, 110S84 and 78S84). Microsatellite mining from 6103 core orthologous clusters comprising 51,127 transcripts mined 14,634 SSR loci, from which 93 primer pairs were synthesized. However, only three SSR markers exhibited polymorphism indicating limited discrimination potential of expressed sequence-derived (EST) SSRs for pathotype differentiation. In contrast, SNP discovery through stringent variant calling and filtration yielded 186 pathotype-specific homokaryotic SNPs, of which 56 high-confidence loci were selected for Kompetitive Allele-Specific PCR (KASP) assay development. A total of 48 KASP markers were synthesized and 14 demonstrated clear pathotype- or cluster-specific polymorphism representing substantially higher resolution than SSR markers. The high SNP-to-KASP conversion efficiency (~&#x2009;95%) and reproducible fluorescence-based clustering emphasize the robustness of KASP assay. Comparative evaluation revealed that SNP-based KASP markers provide superior discriminatory capacity for closely related Pst pathotypes and represent a promising complementary molecular approach for rapid identification of predominant Indian Pst pathotypes. The validated marker panel developed in this study can complement conventional virulence phenotyping and field pathogenomics approaches for surveillance of currently known pathotypes, while continued refinement may accommodate future changes in pathogen populations.

India

Genome-Wide Mining of lncRNAs Reveals Their Potential Regulatory Role in the Evolution of Viviparity.

Reproduction in vertebrates usually involves egg-laying (oviparity) or live-bearing (viviparity). Oviparity is the ancestral trait from which viviparity has independently evolved more than 100 times in squamate reptiles. This transition involves a series of physiological and structural changes, including the degeneration of eggshell and the evolution of a placenta and differences in the temporal and spatial expression patterns of some functional genes that drive the structural transformation. Long non-coding RNAs (lncRNAs) play important roles in the regulation of gene expression, yet it remains unclear whether they participate in gene expression shifts during the transition from oviparity to viviparity, and if so how. Therefore, we employ deep mining to identify novel lncRNAs of a closely related oviparous-viviparous pair of lizards (Phrynocephalus przewalskii and P. vlangalii). We construct cis- and trans-regulatory networks between lncRNAs and target genes using the transcriptomic data of oviduct or uteri tissues across reproductive periods. Results show that lncRNAs that regulate eggshell gland developmental genes in the oviparous lizard are lost or less expressed in the viviparous lizard. A number of lncRNAs involved in the regulation of placental development and embryo attachment in viviparous species have no orthologs&#xa0;in oviparous species, and others show little or no expression. Accordingly, lncRNAs may play important regulatory roles in the physiological and structural changes in the transition from oviparity to viviparity. These results open doors to the further elucidation of genetic regulatory networks.

Animals

Comparative transcriptomics uncovers poplar and fungal genetic determinants of ectomycorrhizal compatibility.

Ectomycorrhizal symbiosis supports tree growth and is crucial for nutrient cycling and temperate and boreal ecosystems functioning. The establishment of functional ectomycorrhiza (ECM) first requires the association of compatible partners. However, host and fungal genetic determinants governing mycorrhizal compatibility are unknown. To identify such factors in poplar and its fungal associates, we mined existing and de novo tree and fungal transcriptional datasets. We identified co-expressed genes enabling ECM symbiosis at early and mature stages of the interaction. These sets of genes can be divided into general fungal-sensing and ECM-specific components. We highlight the importance of fungal modulation of plant JA-related defenses and the regulation of secretory pathways for ECM compatibility, including upregulation of key fungal small secreted proteins, the downregulation of plant secreted peroxidases, and the downregulation of plant cell wall remodeling proteins concomitantly with the upregulation of fungal glycosyl hydrolases acting on pectin. Not only gene regulation, but also its temporal scale and dynamics seem to play a crucial role for mycorrhizal compatibility. The expression profile of the host Common Symbiosis Pathway and nutrient transporters was also studied, revealing constitutive levels of expression and moderate upregulation in compatible ECM interactions. Overall, these results underscore the importance of novel biological functions during the establishment of ECM symbiosis, help us gain insights into the molecular events determining mycorrhiza compatibility, and serve as a data-rich transcriptomic resource to open new research questions in the field.

Mycorrhizae

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis