Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Telomere-to-telomere genome of Phoebe chekiangensis reveals that age-dependent CHG hypomethylation promotes floral transition via MADS-box gene activation.

Phoebe species are renowned for their highly valuable 'golden thread' timber; however, their protracted juvenile phase presents a significant obstacle to mechanistic investigations of floral induction. Phoebe chekiangensis, a rare early-flowering representative within this genus, provides a unique model system for dissecting the vegetative-to-reproductive phase transition. Nevertheless, the absence of a high-quality reference genome has severely hindered molecular insights into its developmental regulation. Here, we present the first telomere-to-telomere (T2T) genome assembly for P. chekiangensis, comprising two completely gap-free haplotypes with contig N50 values exceeding 65 Mb, base-level quality scores >36, and Long Terminal Repeat Assembly Index scores surpassing the gold standard threshold of 20. Approximately 29 000 genes were annotated per haplotype, supported by a BUSCO completeness score of >97%. Age-resolved transcriptomic landscapes identified two MADS-box transcription factors, PcMADS5 (AP1-like) and PcMADS19.1 (SOC1-like), as core activators of the floral transition. Both genes triggered precocious flowering when ectopically expressed in Arabidopsis thaliana. Whole-genome bisulfite sequencing revealed a progressive, age-dependent decline in CHG (where H is A, C, or T) DNA methylation, which was particularly pronounced at the PcMADS19.1 locus. Notably, DML1/2, which mediate active DNA demethylation, were coordinately upregulated during the onset of reproductive growth. Chemical demethylation using 5-azacytidine further diminished CHG methylation and selectively enhanced PcMADS19.1 expression, confirming a causal relationship between CHG hypomethylation and transcriptional activation. This work delivers the first chromosome-scale T2T genome within the genus Phoebe and uncovers CHG demethylation as a previously unrecognized epigenetic switch governing reproductive competence in woody perennials.

Journal Article↗

Changes of DNA methylation and gene expression profile in placental villi and chorioamniotic membranes under preeclampsia.

BACKGROUND: Preeclampsia (PE) is a serious pregnancy complication with elusive pathogenesis. Although epigenetic dysregulation is implicated, its layer-specific placental roles are poorly defined. This study aimed to identify shared and layer-specific epigenetic alterations in PE by profiling DNA methylation and gene expression in placental villi (PV) and chorioamniotic membranes (CAM). RESEARCH DESIGN AND METHODS: PV and CAM samples were collected from 7 normal and 8 PE pregnancies, and three public DNA methylation datasets (GSE98224, GSE44667, GSE75196) were integrated. Differentially methylated genes (DMGs) and differentially expressed genes (DEGs) were identified based on whole-genome methylation and transcriptome sequencing. Layer-specific and shared gene sets were identified by cross-analysis, with functional annotation using Gene Ontology (GO). RESULTS: EM-seq revealed a hypermethylation-dominant, tissue-specific methylation landscape in PE placentas. Cross-tissue comparison identified shared DMGs between the two layers, including nine key genes consistently altered in public datasets. Integrated analysis in PV further identified 22 co-dysregulated genes, enriched in thermoregulation, maternal-fetal immunity, signal transduction, and cell differentiation. CONCLUSIONS: This study elucidates the shared and layer-specific dysregulation of gene networks at methylomic and transcriptomic levels in PE placenta. Comparing PV and CAM highlights placental epigenetic heterogeneity and dysfunction, offering novel clues for mechanistic research and layer-targeted therapies.

Humans↗

Genomic and transcriptomic characterization of genes expressed at 20 MPa by the marine actinobacterium Kocuria flava.

A marine hydrocarbonoclastic actinobacterium Kocuria flava IOS11 was isolated from 3500 m deep-sea water of the Indian Ocean. The isolate efficiently degraded phenanthrene (250 mg/L) achieving 82 and 98% of degradation at 0.1 MPa and 20 MPa, respectively within a period of 5 days. Whole genome, transcriptomee and metabolomic analysis elucidated its phenanthrene biodegradation efficiency under in situ deep-sea conditions. The genome sequence comprises 3.47 Mb distributed across 88 scaffolds with a high GC content of 74.30%. The genome analysis encoded 3126 genes including 3052 protein coding sequences with functional annotation identifying a broad array of genes associated with PAHs degradation, environmental stress adaptation, biosurfactant and siderophore synthesis. Transcriptome profiling under 0.1 and 20 MPa conditions with phenanthrene as a sole carbon source revealed enhanced expression of hydrocarbon degrading genes, transporters, biosurfactant associated enzymes and stress responsive genes including integrases, DNA repair protein Rad, alanine ligase, heat and cold shock proteins under high pressure conditions underscoring the deep-sea adaptation capabilities of the strain. The degradation pathway of phenanthrene was proposed through integrated genome, transcriptome and metabolomic analysis. These studies provided K. flava IOS11 as a metabolically versatile and pressure adapted bacterium with promising potential for bioremediation application in extreme marine environment.

Transcriptome↗

Multi-omics integration and colocalization analyses prioritize candidate molecular loci associated with hypothermia.

BACKGROUND: Hypothermia is a life-threatening condition lacking specific pharmacological treatments. This study aimed to prioritize genetically supported molecular loci associated with hypothermia and to explore their pharmacological tractability using multi-omics data. METHODS: Initially, 2532 druggable genes were curated from the Drug-Gene Interaction Database and established literature. These were cross-referenced with cis-eQTL and cis-pQTL datasets, encompassing 870,655 and 114,281 SNPs for blood, respectively, alongside 2379 shared SNPs across adipose, skeletal muscle, and heart tissues. Matched instrumental variables were integrated with hypothermia GWAS summary statistics for two-sample Mendelian randomization (MR) and Bayesian colocalization. Transcriptomic differential expression analysis (DEA) was subsequently conducted as an exploratory analysis of cold-exposure-associated expression changes. Database-derived compound annotations were systematically re-evaluated according to target specificity, established pharmacological mechanism, and concordance with the direction of the MR estimates. RESULTS: Among 671 gene-level MR tests, 36 genes reached nominal significance, whereas only ABCC8 remained significant after FDR correction. Colocalization was evaluable for 8 of these 36 genes, and 4 loci (COL18A1, SLC1A7, ADIPOQ, and MERTK) met the prespecified PP.H4>0.90 threshold. The remaining 28 loci were not evaluable because sufficient overlapping regional variants were unavailable after harmonization. Transcriptomic analysis identified altered expression of SLC1A3 and SLCO4A1 under cold exposure, although these findings did not directly validate the colocalization-supported loci. Re-evaluation of database-derived compound annotations did not identify any direct, selective, and directionally concordant drug-repurposing candidate for hypothermia. CONCLUSIONS: COL18A1, SLC1A7, ADIPOQ, and MERTK showed colocalization support among the 8 evaluable nominal MR-associated loci. Because colocalization coverage was limited, these genes should be regarded as preliminary candidate loci rather than established therapeutic targets. The pharmacological annotations were indirect, non-selective, unsupported, or directionally inconsistent and should be interpreted solely as hypothesis-generating information.

Bayesian colocalization↗

hypeR-GEM: connecting metabolite signatures to enzyme-coding genes via genome-scale metabolic models.

MOTIVATION: Enrichment analysis is a cornerstone of "omics" data interpretation, enabling researchers to connect analysis results to biological processes and generate testable hypotheses. Enrichment analysis in metabolomics poses distinct challenges for interpretation and multi-omics integration due to the lack of well-defined and consistent connections to well-curated gene-centered biological knowledge repositories. To address these challenges, we developed hypeR-GEM, a methodology and associated R package that adapts gene set enrichment analysis to metabolomics. hypeR-GEM leverages genome-scale metabolic models (GEMs) to infer reaction-based links between metabolites and enzyme-coding genes, enabling the mapping of metabolite signatures to gene signatures and their subsequent annotation via gene set enrichment analysis. RESULTS: We validated hypeR-GEM using paired metabolomics-proteomics and metabolomics-transcriptomics datasets by assessing whether genes mapped from metabolites significantly overlapped with differentially expressed proteins or transcripts. We further evaluated whether pathways enriched via hypeR-GEM-mapped genes corresponded to those derived from paired proteomic or transcriptomic data. In most datasets analyzed, both the predicted enzyme-coding genes and the associated enriched pathways showed significant concordance with independently derived omics signatures, supporting the utility and robustness of hypeR-GEM. Finally, we applied hypeR-GEM to the analysis of age-associated metabolic signatures from the New England Centenarian Study. The results revealed consistent enrichment of lipid-related pathways, aligning with the well-established role of lipid metabolism in aging, and highlighted additional pathways not captured in the metabolites' annotation, demonstrating hypeR-GEM's practical utility in a real-world use case. AVAILABILITY AND IMPLEMENTATION: The hypeR-GEM R package, documentation, and workflow examples are freely available at https://github.com/montilab/hypeR-GEM and archived at https://doi.org/10.5281/zenodo.20586748.

Metabolomics↗

The Annotated Blueprint: Integrated Functional Genomic Resources for a model Tetraploid Wheat Triticum turgidum cv. Kronos.

Triticum turgidum cv. Kronos is a tetraploid wheat cultivar that underpins one of the richest community platforms for functional genomics. Over the past decade, about 3,000 exome- and promoter-capture datasets, linked to mutagenized seed stocks, and transcriptomic and phenotypic resources have accumulated, yet the absence of a reference genome has constrained their impact. Here, we present a chromosome-scale reference genome of Kronos with high-confidence annotations, including manual curation of over 1,000 disease resistance (NLR) genes. This reference revealed previously hidden NLR diversity and clarified their genomic organization at chromosomal ends. Re-analysis of exome- and promoter-capture datasets enabled high-resolution mutation discovery in genes and regulatory regions that were previously inaccessible, uncovering the full standing variation present in Kronos mutant lines. We further re-curated transcriptomic and small RNA datasets, generating improved, genome-wide maps of microRNAs and phasiRNAs important for wheat development. Collectively, these resources elevate Kronos to reference quality and establish it as a versatile platform for functional and translational wheat research.

Journal Article↗

Coordinated inflammatory macrophage and vascular smooth muscle cell remodeling signatures in human atherosclerosis: An integrative single-cell and bulk transcriptomic analysis.

Atherosclerotic plaque progression is shaped by coordinated inflammatory and remodeling programs involving immune cells and vascular wall cells. Inflammatory macrophage activation and vascular smooth muscle cell (VSMC) phenotypic remodeling are central features of human atherosclerosis, but their transcriptomic relationships during plaque progression remain incompletely characterized. This study integrated single-cell and bulk transcriptomic datasets to examine highly inflammatory macrophage states, VSMC remodeling-related transcriptional programs, and candidate ligand-receptor expression patterns in human atherosclerotic plaques. Human atherosclerotic plaque single-cell RNA sequencing data from GSE260657 and bulk transcriptomic data from GSE28829 were analyzed. After quality control, 7628 cells were retained for single-cell analysis. Major cell types were annotated using canonical markers, followed by reclustering of macrophages and VSMC-related cells. Functional module scoring, differential expression analysis, Gene Ontology biological process enrichment, and Kyoto Encyclopedia of Genes and Genomes pathway analyses were performed to characterize macrophage transcriptional states. Slingshot was applied to infer VSMC pseudotime ordering. CellChat and NicheNet were used to prioritize candidate ligand-receptor expression patterns and ligand-associated VSMC target gene programs. External bulk transcriptomic analysis was performed to examine whether single-cell-derived inflammatory and remodeling signatures were represented at the tissue-transcriptome level during plaque progression. Macrophage reclustering identified a highly inflammatory macrophage state characterized by prominent inflammatory activation, cytokine-response, and stress-response features. Genes upregulated in this population were enriched in pathways related to tumor necrosis factor (TNF) response, nuclear factor kappa B signaling, leukocyte activation, cytokine signaling, lipid and atherosclerosis, toll-like receptor signaling, and inflammasome-associated inflammation. VSMC reclustering revealed contractile VSMCs, PTHLH+ synthetic VSMCs, KRT7+ VSMC-like cells, interferon-responsive VSMCs, pericyte-like mural cells, and osteogenic/modulated VSMCs. Pseudotime analysis showed a broad contractile-to-osteogenic/modulated transcriptional continuum accompanied by increased expression of remodeling-associated genes and selected inflammatory or remodeling-associated receptor genes. CellChat and NicheNet analyses prioritized candidate ligand-receptor and ligand-associated target gene expression patterns involving SPP1-CD44, TNF-TNFRSF1A, IL1B-IL1R1/IL1RAP, MIF-ACKR3, PDGFB-PDGFRB, and FN1-SDC1/ITGB1. In GSE28829, inflammatory macrophage-, osteogenic/modulated VSMC-, candidate ligand-receptor expression-, SPP1-CD44 candidate axis-, and NicheNet-prioritized target program-related signatures were more prominent in advanced plaques and were positively correlated with each other. This integrative transcriptomic analysis identified a highly inflammatory macrophage state and a VSMC remodeling continuum in human atherosclerotic plaques. Candidate ligand-receptor and ligand-associated target gene expression patterns linked inflammatory macrophage activation with osteogenic/modulated VSMC remodeling at the computational level. External bulk data further showed coordinated enrichment of inflammatory and remodeling signatures in advanced plaques. These findings provide a descriptive and hypothesis-generating transcriptomic framework for understanding inflammatory macrophage activation and VSMC remodeling in human atherosclerosis.

atherosclerosis↗

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus↗

Comparative analysis of histological and transcriptomic characteristics in caudal muscles of nile crocodiles (Crocodylus niloticus), siamese crocodiles (Crocodylus siamensis), and their hybrids.

Crocodylus niloticus and Crocodylus siamensis are high-value aquaculture species. C. niloticus is large-bodied but less abundant, while C. siamensis grows fast but is small-sized. Their hybrids combine parental advantages, yet relevant research is scarce. This study compared the histological and transcriptomic characteristics of the caudal muscle across the three taxa. HE staining indicated that C. niloticus had significantly larger myofiber diameters (p&#xa0;<&#xa0;0.05); C. siamensis had the smallest, and the myofiber density of hybrids was much closer to that of C. siamensis. Masson's trichrome staining indicated that C. niloticus had the thickest collagen fibers (p&#xa0;<&#xa0;0.05), C. siamensis the thinnest, and hybrids exhibited highly similar histological traits to C. siamensis. C. niloticus had higher LDH and SDH activities in caudal muscles, whereas the hybrid crocodile indicated the highest CK activity. Transcriptomic analysis identified numerous differentially expressed genes (DEGs), which were enriched in growth, muscle metabolism, and energy allocation pathways via GO/KEGG annotations. PPI analysis screened 24 hub genes related to energy metabolism. This study systematically reveals caudal muscle differences, providing insights into growth-related molecular mechanisms and theoretical support for crocodile artificial breeding.

Animals↗

Exploring the transcriptional crosstalk between adipose tissue and locoregional recurrence in breast cancer using independent component analysis.

Locoregional recurrence (LRR) poses a persistent clinical challenge in breast cancer, with emerging evidence implicating the tumor-associated adipose tissue in modulating recurrence risk. This study investigates shared transcriptional programs between adipose tissue and breast tumors and examines their association with disease-free survival (DFS), particularly in the context of reconstructive surgery where adipose tissue from different body compartments are commonly used. We analyzed bulk gene expression data from 5,691 breast tumors and 978 human adipose tissue samples from different body compartments using consensus-independent component analysis (c-ICA) to identify transcriptional components (TCs). Gene set enrichment analysis (GSEA) and copy number alteration profiling were used for biological annotation. Associations between TCs and DFS were evaluated through univariate Cox regression. Key findings were validated using spatial transcriptomic and single-cell RNA sequencing datasets. Among the 411 TCs identified, 332 showed biological enrichment, and 35 were significantly associated with DFS. Four DFS-associated TCs (TC257, TC350, TC371, TC400) were enriched for adipogenesis-related genes and exhibited heightened activity in high-grade, triple-negative tumors and in patients with elevated BMI. Notably, TC350 was highly active in adipose tissue from common reconstructive donor sites (abdomen, omentum, subcutis) but not in native breast adipose tissue. Spatial transcriptomic and single-cell analyses confirmed the increased activity of these adipogenesis-related TCs in tumor regions and adipose cells. TC350 included&#xa0;FABP4, a gene previously linked to poor prognosis in breast cancer and considered as a potential new therapeutic target. Adipose tissue-derived transcriptional programs influence breast cancer prognosis and this seems to differ by tissue origin. These findings generate a hypothesis that donor site selection for adipose tissue in reconstructive surgery may impact LRR risk through adipogenesis-associated mechanisms. Further research is warranted to elucidate the biological and clinical implications of adipose-tumor transcriptional interactions.

Humans↗

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals↗

Beyond the gene: isoform diversity as a key contributor to human brain disorders.

The human brain exhibits exceptional transcriptomic complexity, with alternative splicing, promoter usage, and polyadenylation generating extensive transcript-isoform diversity. Isoform dysregulation is increasingly implicated in neurodevelopmental and psychiatric disorders (NPDs), yet the landscape, function, and genetic regulation of brain isoforms remain poorly understood due to limitations of short-read RNA sequencing. Advances in long-read sequencing (LR-seq) enable scalable full-length transcriptome profiling with single-cell and spatial resolution across developmental stages. Here, we review recent progress in isoform discovery, quantification, functional annotation, and genetic regulation, highlighting emerging links to human neurodevelopment and disease. LR-seq studies have uncovered tens of thousands of previously unannotated brain isoforms, with neuronal maturation characterized by increased exon inclusion and progressive 3' untranslated region (3' UTR) lengthening. Isoform-resolved genetic mapping outperforms gene-level analyses for NPD gene discovery and mechanistic interpretation. We argue that a shift from gene-centric to isoform-centric frameworks is essential to fully capture regulatory complexity in human neurogenetics. Together, these advances establish isoform diversity as a fundamental yet underappreciated axis of brain gene regulation and a key entry point for dissecting NPD biology.

Humans↗

Combining Annotation Software to Identify Orthologous Genes (CASIO) Provides a New Dataset of Orthologous Genes for Swallowtail Butterflies.

With the massive increase in genomic resources, it is becoming increasingly popular to analyse thousands of loci across many species. However, many of the available genomes are not annotated, which hinders an efficient search for orthologous protein-coding genes. Here, we aim to develop a semi-automated pipeline and compare four genomic annotation methods (BRAKER2, BUSCO, Miniprot and Scipio). Our results highlight the importance of integrating multiple annotation tools to optimise ortholog detection and improve genomic studies. Each annotation method showed different strengths. BRAKER2 annotated a substantial number of genes. BUSCO, despite limitations inherent to its reference database, identified a higher number of orthologs. Miniprot exhibited notable flexibility in accommodating diverse protein datasets, whereas Scipio successfully recovered a considerable set of genes that were not detected by the other tools. The combination of these tools allowed for more comprehensive ortholog detection. Taking advantage of this pipeline, we developed a comprehensive dataset of orthologous genes for swallowtail butterflies (Lepidoptera: Papilionidae), called Papilionidae_odb, which will facilitate future studies, especially for a non-model group with abundant genomic data and few transcriptomic resources. We tested Papilionidae_odb by inferring a robust phylogenetic framework for Leptocircini using 142 complete genomes, which improved branch support for some phylogenetic relationships, although challenges remained in resolving relationships within certain species groups, likely due to rapid radiations. Our results highlight the complementary nature of the annotation methods and suggest that combining these tools can yield more accurate results in genomic research. This approach was implemented in a Snakemake workflow called CASIO (Combining Annotation Software to Identify Orthologous genes) and can easily be applied to other non-model groups to improve genomic datasets in diverse taxa where transcriptomic resources are still limited.

Animals↗

Deciphering the Role of LNX2 as a Potential Contributor to Neurodevelopmental Disorders.

BACKGROUND/OBJECTIVES: Attention-deficit/hyperactivity disorder (ADHD) is a common neurodevelopmental condition characterized by a complex and multifactorial genetic architecture. In this study, we report a male patient, born to non-consanguineous healthy parents, presenting with ADHD and oppositional defiant disorder (ODD). METHODS: Trio-based whole-exome sequencing (WES) was performed in the proband and both parents. Variant classification was performed according to American College of Medical Genetics and Genomics (ACMG) guidelines, and the potential pathogenicity of the identified variant was further assessed through multiple in silico prediction algorithms and protein structural analyses. RESULTS: WES identified a homozygous variant in the LNX2 gene (NM_153371.4: c.1165G>A, p.Ala389Thr), classified as a variant of uncertain significance (VUS) and supported by multiple in silico predictions. LNX2 is expressed during brain development and encodes an E3 ubiquitin ligase involved in neuronal differentiation and synaptic function. The identified variant is located within the PDZ2 domain, a functionally relevant region involved in protein-protein interactions. Although the variant is reported in population databases (gnomAD ID: rs148429804), it has not been associated with any clinical phenotype, and its presence in the homozygous state has been reported only once, remaining extremely rare and lacking clinical annotation. Structural modelling predicted localized rearrangement of the hydrogen-bonding network within the PDZ2 domain without major conformational changes. Integrative transcriptomic, and single-cell analyses further supported the biological relevance of LNX2 in neurodevelopment, highlighting its preferential association with neuronal projection-cell networks, synaptic vesicle trafficking pathways, and neuron-specific regulatory programs. CONCLUSION: Although the identified LNX2 variant cannot be considered causative for the patient's phenotype and a definitive disease-gene relationship cannot be established based on a single individual, the complementary genetic, structural, and transcriptomic findings support the biological plausibility of LNX2 as a candidate gene for neurodevelopmental disorders. Additional independent patients and functional studies will be required to clarify its contribution to human disease.

Child↗

Comparative transcriptomics of Venus flytrap (Dionaea muscipula) across stages of prey capture and digestion.

The Venus flytrap, Dionaea muscipula, is perhaps the world's best-known botanical carnivore. The act of prey capture and digestion along with its rapidly closing, charismatic traps make this species a compelling model for studying the evolution and fundamental biology of carnivorous plants. There is a growing body of research on the genome, transcriptome, and digestome of Dionaea muscipula, but surprisingly limited information on changes in trap transcript abundance over time since feeding. Here we present the results of a comparative transcriptomics project exploring the transcriptomic changes across seven timepoints in a 72-hour time series of prey digestion and three timepoints directly comparing triggered traps with and without prey items. We document a dynamic response to prey capture including changes in abundance of transcripts with Gene Ontology (GO) annotations related to digestion and nutrient uptake. Comparisons of traps with and without prey documented 174 significantly differentially expressed genes at 1 hour after triggering and 151 genes with significantly different abundances at 24 hours. Approximately 50% of annotated protein-coding genes in Venus flytrap genome exhibit change (10041 of 21135) in transcript abundance following prey capture. Whereas peak abundance for most of these genes was observed within 3 hours, an expression cluster of 3009 genes exhibited continuously increasing abundance over the 72-hour sampling period, and transcript for these genes with GO annotation terms including both catabolism and nutrient transport may continue to accumulate beyond 72 hours.

Droseraceae↗

A Comprehensive Bioinformatics Approach to Analysis of Variants: Variant Calling, Annotation, and Prioritization.

Next-Generation Sequencing (NGS), also known as high-throughput sequencing technologies, has enabled rapid and efficient sequencing of large amounts of DNA and RNA. These technologies have revolutionized the field of genomics, transcriptomics, and proteomics and have been widely used in cancer research, leading to advances in clinical diagnosis and treatment. Improvements in the NGS technologies enabled millions of fragments to be sequenced simultaneously in a time- and cost-effective manner and resulted in large amount of genomic data which require efficient analysis methods. Analysis of the genomic data requires both efficient computer resources and bioinformatics approaches. This chapter details a comprehensive computational approach and analysis steps for genomic data analysis.

Computational Biology↗

Multi-season analysis reveals hundreds of drought-responsive genes in sorghum.

Persistent drought affects global crop production and is becoming more severe in many parts of the world in recent decades. Deciphering how plants respond to drought will facilitate the development of flexible mitigation strategies. Sorghum bicolor L. Moench (sorghum), a major cereal crop and an emerging bioenergy crop, exhibits remarkable resilience to drought. To better understand the molecular traits that underlie sorghum's remarkable drought tolerance, we undertook a large-scale sorghum gene expression profiling effort, totaling nearly 1500 transcriptome profiles, across a 3-year field study with replicated plots in California's Central Valley. This study included time-resolved gene expression data from roots and leaves of two sorghum genotypes, BTx642 and RTx430, with different pre-flowering and post-flowering drought-tolerance adaptations under control and drought conditions. Quantification of genotype-specific drought tolerance effects was enabled by de novo sequencing, assembly, and annotation of both BTx642 and RTx430 genomes. These reference-quality genomes were used to construct a pangene set for characterizing conserved and genotype-specific expression. By integrating time-resolved transcriptomic responses to drought in the field across three consecutive years, we identified a set of 726 drought-responsive genes that responded similarly in all 3&#x2009;years of our field study. Functional enrichment analysis identified abiotic stress, secondary cell wall-related processes and metabolism as particularly affected under both types of drought stress. We also found that some glyoxylate cycle pathway genes, including malate synthase and isocitrate lyase, are differentially regulated particularly during post-flowering drought stress, implicating this pathway as potentially important for drought responsiveness. This expansive dataset represents a unique resource for sorghum and drought research communities and provides a methodological framework for the integration of multi-faceted time-resolved transcriptomic datasets.

Sorghum↗

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI)&#xa0;framework that uses large language models (LLMs)&#xa0;to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them&#xa0;show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models↗