Search PubMedSearch

SEARCH · Search PubMed

Results for “Transcriptome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

27 records · Page 2Linked to original sources

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis

Livestock Multi-Omics Integration: A Systematic Framework From Statistical Association to Causal Interpretation.

Livestock multi-omics integration is key to unraveling complex trait regulation, yet systematic, livestock-specific strategies remain scarce. This review traces the progression from single-omics accumulation to multi-dimensional integration, highlighting how large-scale genomic, epigenomic, and transcriptomic projects lay the foundation for functional dissection. We identify core impediments: extreme species diversity, marked data heterogeneity, limited sample sizes, and a pervasive reduction of multi-omics data to simplistic differential screens, resulting in low translational efficiency. We critically appraise four common pitfalls-overinterpreting correlation as causation, relegating proteomics to corroborating transcriptomics, incomplete microbiome-host integration lacking environmental context, and systematic neglect of metabolic fluxomics-and show how exposomics and fluxomics add necessary causal and dynamic dimensions. To address these, we propose a livestock-adapted three-tier analytical framework: (1) statistical association of cross-omics covariation patterns; (2) machine learning-driven feature mining and integrative modeling; and (3) causal interpretation encompassing Mendelian randomization, prior-knowledge-guided network inference, and physical causal evidence via fluxomics and metabolic control analysis. We further discuss how multimodal sequencing (single-cell, spatial, temporal) and generative AI can fundamentally mitigate heterogeneity and strengthen causal evidence. Finally, we outline future priorities in database standardization, livestock-specific benchmarking, and translational pipelines, charting a path from correlation-centric reporting to mechanistic causality and precision breeding.

Animals

Genetic and epigenetic underpinnings of biological aging: a multi-omics study integrating Mendelian randomization, spatial transcriptomics, and drug target discovery.

Inflammaging represents a hallmark of biological aging, yet the causal inflammatory mediators driving multi-dimensional epigenetic aging and their effector genes remain poorly characterized at the genetic level. We developed a four-tier analytical framework integrating causal screening, multi-omics effector gene mapping, spatial transcriptomics, and drug target evaluation. Two-sample Mendelian randomization (MR) of 91 circulating inflammatory proteins against six aging phenotypes identified IL-12B, IFNG, and IL-2 as the most robust pro-aging mediators with consistent effects across independent outcomes. Using multi-omics summary-based MR (SMR) as the core analytical engine, we integrated four-layer whole-blood molecular QTL resources eQTL (eQTLGen, n = 31,684), sQTL (GTEx, n = 755), pQTL (INTERVAL + SCALLOP, n = 34,232), and mQTL (McRae et al., n = 1,980) - with GWAS summary statistics for four epigenetic age acceleration measures. At a stringent threshold (P_SMR < 1&#xd7;10&#x207b;&#xb9;&#xb2;), seven high-confidence effector genes were identified: NHLRC1, TPMT, SELP, and RIPPLY3 for IEAA; ZNF373A and PLDN for HannumAA; and EDARADD for PhenoAA. The chromosome 6p21 NHLRC1-TPMT locus, overwhelmingly driven by methylation QTL signals (-log&#x2081;&#x2080;P = 26.06), emerged as the dominant genetic node of epigenetic aging. Spatial projection via gsMap onto a mouse E16.5 embryo atlas (121,767 cells) revealed preferential enrichment in smooth muscle and lung, with EDARADD showing marked specificity in mucosal epithelium. Cross-database drug target mining classified TPMT and SELP as repurposable known targets and NHLRC1 as a high-priority novel druggable candidate. This study provides multi-omics convergent causal evidence for inflammation-driven epigenetic aging and delivers genetically anchored targets for precision anti-aging intervention.

Aging

SwinePan for pig graph-based pangenome and multiomics data mining.

Pigs are one of the most important livestock species worldwide. Although multiple high-quality reference genomes exist, reliance on a single linear reference limits the detection of structural variants (SVs) and the characterization of population-specific genetic diversity. To address this limitation, we developed SwinePan, a comprehensive and integrated multiomics database for pigs built on a graph-based pangenome framework. SwinePan incorporates a variome derived from the graph-based pangenome, covering 2,598 individuals across 35 breeds, including 185,759 SVs, 117 million SNPs, and 6.8 million indels. The database also integrates transcriptomic data from liver, loin muscle, abdominal fat, and backfat, along with over 150,000 phenotypic records. The online toolkit deployed in SwinePan enables genome-wide association studies (GWAS), expression quantitative trait locus (eQTL) mapping, and colocalization, while interactive modules visualize population structure and multiomics associations, streamlining candidate gene and variant exploration. Additionally, two proof-of-concept analyses demonstrate how SwinePan pinpoints trait-associated loci and deciphers their potential regulatory mechanisms.

Journal Article

Systematic mining and characterization of metal transporter families regulating zinc homeostasis provide insights into metal homeostasis in Camellia sinensis.

BACKGROUND AND AIMS: Zinc is essential for tea plant growth and quality formation, yet its homeostatic mechanisms remain poorly understood. This study identified metal transporter families regulating zinc homeostasis, analyzed their evolution, structure, and expression, and clarified zinc uptake, transport, detoxification networks, and their links to metabolism. METHODS: This study identified zinc homeostasis-related metal transporter families in the tea plant genome, characterized their structural features and expression profiles across tissues and developmental stages through integrative bioinformatics and transcriptomic analyses, and delineated the molecular mechanisms underlying zinc uptake, translocation, and detoxification by systematically integrating published evidence. RESULTS: This study identified 74 metal transporter genes from six families: 13 CsZIPs, 12 CsNRAMPs, 10 CsHMAs, 10 CsYSLs, 14 CsMTPs, and 15 CsCAXs in the 'Shuchazao2' genome, revealing closer affinity to woody species than to Arabidopsis. These proteins exhibit conserved domains, diverse subcellular localizations (cell membrane, vacuole, chloroplast, and Golgi apparatus), and tissue-specific expression with abundant stress/hormone-responsive cis-elements. At the plant-soil interface, tea plants mobilize rhizospheric zinc via proton and organic acid secretion; CsYSLs, CsNRAMPs, and CsZIPs mediate zinc uptake, aided by arbuscular mycorrhizal fungi (AMF) and plant growth-promoting rhizobacteria (PGPR) that expand root absorption zones. Xylem CsHMAs and phloem CsYSLs coordinate root-to-shoot zinc translocation, and vacuolar transporters (CsMTPs, CsCAXs), cell wall immobilization, and antioxidant systems alleviate high-zinc stress injury. CONCLUSIONS: These findings collectively delineate an integrated zinc "acquisition-distribution-buffering" network in tea plants, offering a repertoire of candidate genes with potential utility in zinc biofortification breeding and improving acid soil adaptation. Further experimental validation, including tea&#xa0;transgenesis, zinc-stress qRT-PCR, and heterologous functional complementation, is essential to substantiate their biological roles.

Camellia sinensis

Biocontrol effect of a solid-state fermentation-derived extract mixture of Trichoderma asperellum on sunflower Sclerotinia rot and associated host defense responses.

Sclerotinia disease is a destructive fungal disease of sunflowers, soybeans, and other economically important crops, causing substantial yield loss and quality deterioration. Long-term reliance on dose-dependent broad-spectrum fungicides is constrained by resistance risks and potential environmental burdens, creating tension with the sustainability goal of "reducing pesticide use while improving efficacy." Here, we explore a Trichoderma spp.-based microbial disease management strategy. Whole-genome sequencing of Trichoderma asperellum TCS007 isolated from Antarctic marine sediments, coupled with genome mining, predicted diverse biosynthetic gene clusters putatively associated with siderophores, polyketides, nonribosomal peptides, and terpenoids; the corresponding metabolites are not chemically confirmed and require further validation. Using a solid-state fermentation workflow, we prepared a fermentation-derived extract mixture (TCS007-SSF-Ex). In vitro assays showed dose-dependent inhibition of Sclerotinia sclerotiorum by TCS007-SSF-Ex (EC50 = 1.252 mg/L), and microscopy revealed cellular damage-consistent changes, including organelle disruption and plasmolysis. Pathogen transcriptomic and metabolism-related analyses indicated broad perturbations in organelle biogenesis and metabolic processes, with significant alterations in pathways associated with succinate, D-glucose, and phenylacetate; these results are consistent with growth inhibition and reduced pathogenicity, but specific molecular targets and causal links remain to be validated. In vivo, under certain application conditions, triple applications increased APX activity (+492.5%) and &#x3b2;-1,3-glucanase activity (+419.6%). Collectively, this work supports a "pathogen suppression-host defense induction" framework and facilitates subsequent identification of active components and mechanistic validation.IMPORTANCESclerotinia diseases cause recurrent and economically important losses in oilseed crops, while long-term fungicide use is constrained by resistance risks and environmental burdens. Trichoderma-based biocontrol is a promising complementary strategy, yet evidence supporting metabolite-containing Trichoderma-derived preparations as immune elicitors remains less consolidated than that for living inoculants, and scalable production routes are still needed. Here, we examine an Antarctic marine sediment-derived strain, Trichoderma asperellum TCS007, and a solid-state fermentation (SSF)-derived extract mixture (TCS007-SSF-Ex) produced via solid-state fermentation. We combine in vitro antifungal assays, pathogen ultrastructural observations, and correlative omics analyses with in vivo measurements of sunflower defense enzymes (APX and &#x3b2;-1,3-glucanase) to evaluate a "pathogen suppression-host defense induction" framework. Our findings support the potential of SSF-derived Trichoderma metabolite mixtures for greener management of Sclerotinia disease and provide a foundation for future chemical identification of active components and mechanistic validation.

Ascomycota

A chromosome-scale genome of Capsicum pubescens provides insights into candidate terpene-associated gene clusters and pan variation of terpene synthases.

A chromosome-scale genome of Capsicum pubescens and comparative pan-TPS analysis support structural characterization and gene-level prioritization of a chromosome-9 terpene-associated candidate locus in this accession. Capsicum pubescens is one of the five domesticated Capsicum species, mainly cultivated in mid- to high-elevation regions of the Americas. Despite its distinctive morphology and fruit traits, genomic resources for C. pubescens remain less developed than those for the widely cultivated C. annuum. Here, we assembled a chromosome-scale reference genome for accession HNUCP0001, spanning 3.70&#xa0;Gb with a scaffold N50 of 278.01&#xa0;Mb. Comparative genomics revealed 679 significantly expanded gene families enriched in sesquiterpenoid and triterpenoid biosynthesis. Genome-wide biosynthetic gene-cluster mining identified multiple terpene-associated candidate loci, which were subsequently prioritized using genome-derived structural criteria and Capsicum pubescens-specific expression evidence. Subsequently, we curated the terpene synthase (TPS) repertoire and, across 16 Capsicum genomes, resolved 36 TPS orthogroups with pronounced presence/absence variation, highlighting dynamic lineage-specific diversification. Together, these analyses establish HNUCP0001 as an accession-specific genomic resource and provide a comparative framework for prioritizing terpene-associated TPS genes and candidate BGCs in Capsicum. These candidate loci, together with accession-level transcriptomic and metabolomic evidence, offer testable hypotheses for future functional studies of specialized terpenoid metabolism in C. pubescens.

Alkyl and Aryl Transferases

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI)&#xa0;framework that uses large language models (LLMs)&#xa0;to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them&#xa0;show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models

Mining the sHSP20 (small heat-shock protein) gene family in finger millet (Eleusine coracana (L.) Gaertn.): structural, evolutionary and predicted abiotic-stress-responsive insights.

Small heat-shock proteins (sHSPs, the HSP20 family) are ATP-independent molecular chaperones that hold partially unfolded substrates and protect the proteome during heat and other abiotic stresses; every member is defined by a conserved &#x3b1;-crystallin domain (ACD). Finger millet (Eleusine coracana) is a climate-resilient, calcium-rich allotetraploid cereal of the semi-arid tropics whose HSP20 repertoire had not been catalogued. The present study is an entirely computational (in silico) analysis of the chromosome-scale reference genome of finger millet (NCBI GenBank assembly GCA_032690845.1, cultivar KNE 796-S). Mining the predicted proteome with the ACD profile (Pfam PF00011) and confirming every candidate by NCBI CD-search recovered 76 non-redundant ACD-bearing HSP20 genes (EcHSP20-1-EcHSP20-76). Based on phylogeny and TargetP-predicted localization, the members were classified into ten subfamilies: seven cytosolic/nuclear classes (C-I to C-VII, 60 members) together with chloroplastic (11), mitochondrial (3) and endoplasmic-reticulum (2) groups. The proteins ranged from 110 to 355 amino acids (12.1-39.2&#xa0;kDa) with theoretical pI of 4.85-9.69. The 76 loci were distributed over 14 of the 18 chromosomes and were conspicuously absent from chromosomes 8&#xa0;A, 8B, 9&#xa0;A and 9B, with pronounced clustering on chromosomes 1, 2, 3 and 6. Duplication analysis detected 149 paralogous pairs (49 homoeologous, 80 segmental/dispersed and 18 tandem); 147 of 148 pairs for which substitution rates could be calculated returned Ka/Ks&#x2009;<&#x2009;1 (mean 0.20), indicating strong purifying selection consistent with retention after whole-genome/allopolyploid duplication. Promoter analysis (PlantCARE) revealed enrichment of abscisic-acid-responsive (ABRE), MYB/MYC drought-related, STRE, DRE, low-temperature (LTR) and methyl-jasmonate/salicylic-acid elements, whereas canonical heat-shock elements (HSE) were not recovered. Expression profiling against a public drought transcriptome (SRP081350) showed that about half of the genes (39 of 76) are transcribed in leaf tissue, the expressed fraction being dominated by the cytosolic class C-I. This first finger-millet HSP20 catalogue provides a verified, reproducible framework and nominates computationally predicted candidate genes for future functional work on thermotolerance in cereals.

Allotetraploid