Search PubMedSearch

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Comprehensive Transcriptome Annotation of Thousands of HIV-1 Genomes.

Alternative splicing in HIV-1 has been a central focus of decades of research, uncovering key mechanisms of viral gene regulation, immune evasion, and therapeutic response - yet, no reference resource has existed to support transcriptome-wide analysis, limiting adoption of modern computational methods. We present HIV Atlas (https://ccb.jhu.edu/HIV_Atlas), the first reference-quality annotation of HIV-1 and SIV transcriptional diversity. We manually curated transcriptomes for HIV-1HXB2 and SIVmac239 and developed Vira, an automated annotation-transfer method specifically designed to address unique challenges of viral genome biology, to generate high-quality annotations for 2,077 complete HIV-1 genomes. Using the resources presented in our work, we evaluated conservation of splice sites, revealing near-perfect preservation of major donors and acceptors. Furthermore, using several public datasets, we demonstrate how HIV Atlas enhances methodology, improves the quality and novelty of results, and opens novel avenues for research, supporting more accurate and comprehensive analyses of bulk, single-cell, and spatial RNA-seq in HIV-1 studies.

Journal Article

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames

PotatoRTD and TomatoRTD: Comprehensive Reference Transcript Datasets for Accurate Transcriptome Analysis and Isoform Discovery.

Transcriptome annotations provide essential information on transcript locations, sequences and structures, including transcription start, end sites and splice junctions. They underpin key biological analyses such as gene and transcript quantification, and the study of transcriptional and post-transcriptional regulation, including alternative transcription initiation, polyadenylation and splicing. Accurate characterisation of transcript isoforms is critical for understanding how gene expression relates to functional protein products. However, for many species-including Solanaceae crops such as potato and tomato-current annotations suffer from limited isoform coverage, with tens or hundreds of thousands of splice junctions and transcript isoforms missing. This undermines the completeness and accuracy of transcript-level analyses. Here, by generating Iso-seq and RNA-seq on a range of tissues and samples, we have produced transcriptome annotations for both potato and tomato with improved coverage, diversity, accurate splice junctions, and transcript start and end sites. We have also made these high-quality resources accessible through genome browsers. These enhanced annotations will enable more accurate transcriptome analyses, supporting higher-resolution and novel biological discoveries.

Solanum tuberosum

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals

Genetic evidence prioritizes circulating proteins for heart failure beyond shared BMI-related genetic liability.

BACKGROUND: Heart failure (HF) and body mass index (BMI) share substantial genetic architecture, which may lead genetically informed target discovery to preferentially identify adiposity-related pathways. We sought to identify circulating proteins associated with HF beyond this shared genetic component. METHODS: We applied GWAS-by-subtraction to overall HF, nonischemic HF, and nonischemic HF with reduced or preserved ejection fraction to derive BMI-related and BMI-subtracted HF components. We then performed proteome-wide cis-pQTL Mendelian randomization and colocalization using four independent proteomic cohorts, followed by tissue-specific eQTL colocalization, cardiac transcriptomic annotation, and druggability assessment. RESULTS: Compared with the original HF phenotypes, the BMI-subtracted components showed attenuated genetic correlations with BMI (0.045-0.147) while retaining 28 independent loci for overall HF and nine for nonischemic HF. Across 19,930 protein-HF tests, 11 associations involving nine proteins were prioritized by the Mendelian randomization and colocalization analyses. For example, a 1-SD increase in genetically predicted CELSR2 abundance was associated with lower overall HF risk (odds ratio, 0.96 [95% CI, 0.94-0.98]; P=8.6×10-7), whereas a 1-SD increase in genetically predicted CSF3 abundance was associated with higher nonischemic HF risk (odds ratio, 1.32 [95% CI, 1.18-1.48]; P=2.0×10-6). CELSR2 and TMEM106B colocalized with cis-eQTLs in failing left ventricular myocardium, and DAG1 showed cardiomyocyte enrichment with concordant downregulation in failing hearts. CONCLUSIONS: We identified nine circulating proteins associated with HF beyond the genetic component shared with BMI. These findings extend the range of genetically supported pathways implicated in HF and nominate candidate proteins for further mechanistic and therapeutic investigation.

Genetics

Ulmus minor response to Dutch elm disease: de novo transcriptome assembly and annotation.

Dutch elm disease (DED), caused by Ophiostoma novo-ulmi (ONU), has devastated elm populations across Europe and North America since the 20th century. In this work, a de novo transcriptome assembly of Ulmus minor in response to ONU is presented. We used two DED-resistant genotypes, MDV2.3 and VAD2, and one DED-susceptible genotype, MDV1, to capture responses to ONU at four time points post-inoculation (6, 24, 72, and 144 hours). RNA from collected samples was isolated and sequenced producing 60.88 M 100 bp paired-end reads per sample. We performed a de novo transcriptome assembly combining data from the three genotypes. The assembly was functionally annotated and validated through differential gene expression analysis of the response. This dataset provides a valuable resource for studying molecular mechanisms of DED resistance in elms, contributing to broadening our understanding of tree immunity and facilitating potential applications in functional annotation of future genome assemblies.

Transcriptome

Identification and characterization of G protein-coupled receptors in the nocturnal halictid bee Megalopta genalis.

G protein-coupled receptors (GPCRs) are one of the largest families of membrane proteins in insects, regulating vision, neural signal transduction, and various physiological behaviors. Megalopta genalis exhibits a unique facultatively eusocial lifestyle and possesses adaptations for nocturnal activity; however, its GPCR family has not yet been systematically characterized. In this study, we performed genome-wide identification, phylogenetic analysis, and expression profiling of GPCRs in M. genalis by integrating genomic annotation and transcriptomic analysis. The results showed that a total of 99 GPCRs were identified in the genome of M. genalis, which were classified into four major families. Here, we show that M. genalis has undergone lineage-specific GPCR repertoire remodeling, marked by the expansion of novel orphan receptors and the systematic loss of multiple receptor subtypes, such as the neuropeptide receptors MIP-R and NPFR. Moreover, opsins have formed a diverse array of combinations and non-GPCR odorant receptors have undergone significant expansion via tandem duplication. Together, these features may represent part of the molecular repertoire associated with the adaptation of M. genalis to a nocturnal lifestyle. Furthermore, transcriptomic analysis revealed distinct spatiotemporal expression divergence within each of the Mth/Mthl and Fz GPCR families, suggesting functional specialization across development and adult tissues. This study provides the first systematic identification and initial functional characterization of GPCRs in M. genalis, revealing an evolutionary pattern characterized by the coexistence of contraction and expansion within the GPCR family. These findings lay a foundation for further studies aimed at elucidating the roles of these GPCRs in regulating M. genalis physiology and behavior.

Animals

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta

A complete and near-perfect rhesus macaque reference genome: lessons from subtelomeric repeats and sequencing bias.

A truly complete, telomere-to-telomere (T2T), and error-free reference genome remains a foundational resource-and long-standing goal-for unbiased comparative and functional genomics. While recent T2T assemblies of humans and other primates have made substantial progress, most still contain thousands of base-level errors, particularly within highly repetitive regions. Here, we present T2T-MMU8v2.0, a near-perfect T2T assembly of the rhesus macaque (Macaca mulatta), representing the highest base-level accuracy reported in a primate genome to date. By employing an optimized ONT-only assembly strategy, we identify subtelomeric satellite-rich regions as the principal bottleneck to improving assembly quality, owing to technological biases in long-read platforms and limitations in current hybrid assembly frameworks. We discover 268 previously unannotated repeat families and resolve ~8 Mbp of SATR satellite arrays, with over 99-fold enrichment in historically misassembled subtelomeric regions. These satellites form four distinct genomic architectures, each with unique SATR satellite composition, segmental duplication organization, and epigenetic signatures, distinct from the subtelomeric architectures observed in hominid genomes. Notably, in contrast to the largely gene-poor subtelomeric regions in African hominids, the SATR architectures in macaques harbor 58 actively transcribed genes, supported by open chromatin and expression data, suggesting gene innovation within these repetitive regions. Functionally, T2T-MMU8v2.0 improves read mappability and accuracy across sequencing platforms, and results in a 19% improvement of transcription start site enrichment scores and 5,821 additional chromatin accessibility peaks on average, thereby enhancing variant detection, regulatory annotation, and transcriptomic resolution in population genetics or single-nucleus studies. Together, this work establishes a new benchmark for genomics, offers a roadmap for resolving complex repetitive regions, and reveals previously unrecognized features of subtelomeric genome structure and evolution.

Journal Article

The extracellular vesicle transcriptome provides tissue-specific functional genomic annotation relevant to disease susceptibility in obesity.

We characterized circulating extracellular vesicles (EVs) in obese and lean humans, identifying transcriptional cargo differentially expressed in obesity (277 unique genes; false discovery rate < 10%). Since circulating EVs may have broad origin, we compared this obesity EV transcriptome with expression from human visceral-adipose-tissue-derived EVs from freshly collected and cultured biopsies from the same obese individuals, observing high concordance. Using a comprehensive set of adipose-specific epigenomic and chromatin conformation assays, we found that the differentially expressed transcripts from the EVs were those regulated in adipose by body mass index-associated SNPs (p < 5 &#xd7; 10-8) from a large-scale genome-wide association study (GWAS). Using a phenome-wide association study of the regulatory SNPs for the EV-derived transcripts, we identified a substantial enrichment for inflammatory phenotypes, including type 2 diabetes. Collectively, these findings represent the convergence of the GWAS (genetics), epigenomics (transcript regulation), and EV (liquid biopsy) fields, enabling powerful future genomic studies of complex diseases.

Humans

Integrated multi-omics identification of m6A-SNP-related diagnostic biomarkers in amyotrophic lateral sclerosis.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) lacks reliable and minimally invasive biomarkers for early diagnosis. m6A-associated single-nucleotide polymorphisms (m6A-SNPs) may influence RNA methylation and gene expression, offering opportunities to identify clinically relevant diagnostic markers. METHODS: We integrated eQTLGen cis-eQTL data, RMVar m6A-SNP annotations, and ALS transcriptomic datasets to identify m6A-SNP-related genes. Random Forest and LASSO regression were combined to screen robust diagnostic markers. A nomogram was constructed and validated using independent cohorts. Immune infiltration, predicted m6A modification sites, and potential RBP-SNP interactions were assessed. Peripheral blood samples from ALS patients were used for exploratory validation of gene expression and global m6A levels. RESULTS: We identified 109 ALS-associated m6A-SNP-related genes with cis-eQTL signals and narrowed these to seven candidate diagnostic markers (TMED5, OXR1, BRI3, FEM1C, SUZ12, EIF2AK4, and TJAP1). The seven-gene model outperformed the individual markers in the training cohort and retained moderate discrimination in the independent validation cohort. ALS samples showed differences in inferred immune-cell composition, including monocytes, neutrophils, and T-cell subsets. The selected SNP loci were located near predicted m6A sites and annotated RBP-binding regions. Exploratory clinical validation showed significant upregulation of FEM1C and SUZ12 at both mRNA and protein levels, accompanied by reduced global m6A modification. CONCLUSIONS: Through multi-omics integration and exploratory clinical validation, this study identifies m6A-SNP-related candidate markers associated with ALS. The findings support further evaluation of m6A-related signatures for ALS discrimination and molecular characterization, while larger independent cohorts and additional calibration are required before clinical application.

Humans

De novo assembly of transcriptomes of six Hua species (Semisulcospiridae, Cerithioidea, Gastropoda).

Species in Semisulcospiridae are important in freshwater ecology and have great research value, yet their genomic resources remain very limited. Here, we present de novo assembled transcriptomes from six species of Hua in Semisulcospiridae, including Hua textrix (Heude, 1888), H. yangi L.-N. Du, J.-X. Yang & Chen, 2023, H. wujiangensis L.-N. Du, J.-X. Yang & Chen, 2023, and three undescribed species. Assembly was performed using Trinity, resulting in average contig lengths ranging from 716.6 to 883.3&#x2009;bp and transcript numbers ranging from 147,147 to 268,741. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis was used to assess the transcriptome completeness. The functional annotation of transcripts for each species had over 18,000 BLAST hits, 17,000 GO terms, 15,000 KEGG pathways, 8,000 Pfam accessions, and 140 COG functional categories. This study provides valuable transcriptomic resources for the six Hua species, which can be used for various research of Semisulcospiridae, including biodiversity, phylogeny, and comparative genomics.

Transcriptome

Sex- and development-specific transcriptomic profiling of venom and silk genes in the wolf spider Pardosa astrigera provides insights into ecological adaptation and predatory strategies.

Spider venom and silk glands are two major secretory systems that contribute to prey capture, defense, and reproduction, but their sex- and development-specific molecular regulation in wandering wolf spiders remains poorly understood. Here, the transcriptome of Pardosa astrigera, an important agricultural natural enemy in China, revealed significant sex- and development-associated molecular differentiation among adult females, adult males, and spiderlings. A total of 100,025 unigenes were obtained, of which 23,852 were functionally annotated, providing a comprehensive transcriptomic resource for this species. Differential expression patterns showed marked variation among groups, with 531, 1792, and 832 DEGs detected in PAF vs PAS, PAM vs PAS, and PAF vs PAM, respectively. These genes were mainly associated with metabolic, oxidation-reduction, cuticle development, MAPK signaling, and lysosome pathways. Fifteen co-expression modules revealed distinct expression patterns. The turquoise, pink, yellow, and red modules were development-related, whereas the blue module was male-biased. Venom- and spidroin-related genes were distributed across multiple modules, suggesting coordinated regulation. Overall, 42 venom peptides, 21 venom proteins, and 11 spidroins were identified. Representative genes showed strongly biased expression, including spiderling-biased U3_Pp1a and U5_Pp1e, female-biased U4_Pp1a, and male-biased SMase D_108750 and PaTuSp_108466. These findings reveal sex- and development-biased expression patterns of venom- and silk-related candidate genes in P. astrigera and may provide molecular insights into ecological adaptation and predatory strategies in wandering wolf spiders.

Animals

Integrated multi-omics profiling identifies aging-related molecular signatures and convergent interferon signaling in systemic lupus erythematosus.

BACKGROUND: Systemic lupus erythematosus (SLE) is characterized by chronic immune activation and molecular alterations that overlap with aging-related biological processes. However, how these alterations are organized across molecular layers and whether they converge on shared regulatory networks remain incompletely understood. METHODS: We performed an integrative multi-omics analysis combining in-house proteomic and phosphoproteomic data from 130 patients with SLE and 90 healthy controls (HCs) and publicly available transcriptomic datasets comprising 1,461 SLE patients. Proteins and phosphorylation sites were annotated using established aging-related gene resources. Differential protein abundance and phosphorylation changes were analyzed across disease-status and disease-activity comparisons. Nominal P-value thresholds were used for exploratory feature selection, whereas FDR-adjusted P values were used to assess robustness after multiple-testing correction. Kinase-substrate enrichment, transcription factor annotation, and cell-type-resolved transcriptomic comparison were used to explore potential regulatory programs. RESULTS: We identified 128 nominally altered proteins annotated to aging-related biological processes, including genomic instability, mitochondrial dysfunction, and epigenetic alterations. Phosphoproteomic analysis revealed 36 nominally altered phosphorylation sites, including previously unreported sites in IFI16 (S153, S780) and PKC&#x3b4; (S507, S664). Clustering analysis demonstrated heterogeneous protein co-regulation patterns across disease states. Kinase activity inference suggested altered activity of TBK1 and IKK&#x3b2;. TF analysis further highlighted STAT1, RELA, and PML as potential central nodes within the inferred regulatory network. Notably, these multi-omic alterations were not randomly distributed but showed convergence toward shared signaling pathways, particularly those related to interferon responses. CONCLUSIONS: This integrative multi-omics study identifies inflammatory and interferon-dominated molecular alterations in SLE PBMCs that overlap with aging-related biological processes and converge on shared regulatory networks. These findings provide a hypothesis-generating framework for investigating the intersection between chronic immune activation and aging-related molecular remodeling in SLE.

Humans

Species-specific transcriptomic changes upon respiratory syncytial virus infection in cotton rats.

The cotton rat (Sigmodon) is the gold standard pre-clinical small animal model for respiratory viral pathogens, especially for respiratory syncytial virus (RSV). However, without a reference genome or a published transcriptome, studies requiring gene expression analysis in cotton rats are severely limited. The aims of this study were to generate a comprehensive transcriptome from multiple tissues of two species of cotton rats that are commonly used as animal models (Sigmodon fulviventer and Sigmodon hispidus), and to compare and contrast gene expression changes and immune responses to RSV infection between the two species. Transcriptomes were assembled from lung, spleen, kidney, heart, and intestines for each species with a contig N50&#x2009;>&#x2009;1600. Annotation of contigs generated nearly 120,000 gene annotations for each species. The transcriptomes of S. fulviventer and S. hispidus were then used to assess immune response to RSV infection. We identified 238 unique genes that are significantly differentially expressed, including several genes implicated in RSV infection (e.g., Mx2, I27L2, LY6E, Viperin, Keratin 6A, ISG15, CXCL10, CXCL11, IRF9) as well as novel genes that have not previously described in RSV research (LG3BP, SYWC, ABEC1, IIGP1, CREB1). This study presents two comprehensive transcriptome references as resources for future gene expression analysis studies in the cotton rat model, as well as provides gene sequences for mechanistic characterization of molecular pathways. Overall, our results provide generalizable insights into the effect of host genetics on host-virus interactions, as well as identify new host therapeutic targets for RSV treatment and prevention.

Animals

Deciphering the Genetic Underpinnings of Liver Cirrhosis-Heart Failure Comorbidity Through Multi-Omics: CRIM1 as a Key Endothelial Mediator.

The co-occurrence of liver cirrhosis (LC) and heart failure (HF) poses considerable clinical challenges, yet the cellular and molecular determinants of this comorbidity remain poorly characterized. To address this, we developed an integrative multi-omics pipeline encompassing GWAS meta-analysis, gsMap-based spatial transcriptomic projection, GeneEnrich functional annotation, single-cell atlas construction, seismicGWAS and ECLIPSER cell-type scoring, eCAVIAR and fastenloc colocalization, hdWGCNA network inference, scTenifoldKnk in silico gene perturbation, and GCTA-COJO fine-mapping. Quality-controlled meta-analysis yielded 12,347,758 and 9,256,862 variant-level associations for LC and HF, respectively. Spatial projection confirmed preferential enrichment of disease signals within embryonic hepatic and cardiac compartments. Pathway analyses disclosed that LC-linked loci were concentrated in lipid metabolic programs, whereas HF-linked loci implicated mitochondrial bioenergetics and lysosomal degradation. At the cellular level, endothelial cells emerged as the dominant HF-associated population. Convergent evidence from five orthogonal algorithms pinpointed CRIM1 as the sole robustly supported shared gene, selectively enriched in HF endothelial cells; virtual perturbation further identified LCP1 and PTPRC as downstream regulatory nodes. Fine-mapping of the chromosome 2 locus harboring rs12476437 revealed multiple statistically independent signals in the vicinity of CRIM1. Collectively, these findings computationally prioritize the endothelial-CRIM1 axis as a previously unappreciated candidate mechanistic bridge between LC and HF requiring experimental validation.

Humans