Search PubMedSearch

SEARCH · Search PubMed

Results for “transcriptome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

EBV reactivation priming of the peripheral immune system in multiple sclerosis relapse.

Despite decades of research, the cellular and molecular events preceding multiple sclerosis (MS) relapse remain incompletely understood. Here, in this observational study of longitudinal blood samples from patients with relapsing-remitting MS, we used single-cell RNA sequencing, bulk transcriptomics, multiparameter flow cytometry and targeted viral reverse transcription quantitative polymerase chain reaction (RT-qPCR) to construct a time-resolved atlas of immune perturbations surrounding relapse. A reproducible pre-relapse signature in monocytes and B cells, emerging up to 3 months before clinical onset, was enriched for host genes responsive to Epstein-Barr virus (EBV) lytic reactivation factors. RT-qPCR confirmed elevated EBV LMP-1 transcripts in pre-relapse B cells, and flow cytometry demonstrated expansion of CD11c+ atypical B cell populations displaying EBV surface protein gp350. Pre-relapse transcriptional modules overlapped with MS genome-wide association study (GWAS) risk loci and EBNA-2-bound enhancers, suggesting that inherited MS susceptibility and EBV-responsive programs operate through shared regulatory elements. How this peripheral activation relates to central nervous system lesion formation remains to be established. These findings nonetheless suggest that EBV reactivation, when occurring within a genetically predisposed peripheral immune environment, is a proximal precursor of MS relapse.

Journal Article

Deciphering the Impact of Temperature on Pleiotropic Consequences of RNA Polymerase Mutations.

Despite occurring in an essential molecule, mutations in RNA polymerase readily emerge and elicit complex pleiotropic effects across different levels of biological organization, which are all modulated by environment. We investigated the impact of temperature on the effects of six mutations on sequence, structure, transcriptome, and organismal traits. We found temperature altered the transcriptomic response and key organismal traits such as growth rate and biofilm formation in a genotype-specific manner. Critically, mechanistic insights into the possible drivers of mutational effects emerged only when examining the relationships between different levels of organization: location of mutations in the tertiary structure and distance to key interacting molecules partly explained the observed transcriptomic differences, which in turn drove the impact of mutations on organismal traits. While falling short of capturing the full complexity of the system, our findings underscore the benefits of integrating insights across multiple biological levels to understand the relationship between environment and mutational effects in molecules with extensive pleiotropic effects.

Mutation

PSEUDO-RESPONSE REGULATOR 3b and transcription factor ABF3 modulate abscisic acid-dependent drought stress response in soybean.

The circadian system plays a pivotal role in facilitating the ability of crop plants to respond and adapt to fluctuations in their immediate environment effectively. Despite the increasing comprehension of PSEUDO-RESPONSE REGULATORs and their involvement in the regulation of diverse biological processes, including circadian rhythms, photoperiodic control of flowering, and responses to abiotic stress, the transcriptional networks associated with these factors in soybean (Glycine max (L.) Merr.) remain incompletely characterized. In this study, we provide empirical evidence highlighting the significance of GmPRR3b as a crucial mediator in regulating the circadian clock, drought stress response, and abscisic acid (ABA) signaling pathway in soybeans. A comprehensive analysis of DNA affinity purification sequencing and transcriptome data identified 795 putative target genes directly regulated by GmPRR3b. Among them, a total of 570 exhibited a significant correlation with the response to drought, and eight genes were involved in both the biosynthesis and signaling pathways of ABA. Notably, GmPRR3b played a pivotal role in the negative regulation of the drought response in soybeans by suppressing the expression of abscisic acid-responsive element-binding factor 3 (GmABF3). Additionally, the overexpression of GmABF3 exhibited an increased ability to tolerate drought conditions, and it also restored the hypersensitive phenotype of the GmPRR3b overexpressor. Consistently, studies on the manipulation of GmPRR3b gene expression and genome editing in plants revealed contrasting reactions to drought stress. The findings of our study collectively provide compelling evidence that emphasizes the significant contribution of the GmPRR3b-GmABF3 module in enhancing drought tolerance in soybean plants. Moreover, the transcriptional network of GmPRR3b provides valuable insights into the intricate interactions between this gene and the fundamental biological processes associated with plant adaptation to diverse environmental conditions.

Glycine max

Adaptive traits for chitin utilization in the saprotrophic aquatic chytrid fungus Rhizoclosmatium globosum.

The Chytridiomycota (chytrids) are early diverging fungi, many of which function in ecosystems as saprotrophs; however, associated adaptive traits are poorly understood. We focused on chitin degradation, a common ecosystem function of aquatic chytrids, using the model chitinophilic Rhizoclosmatium globosum and comparison of other chytrid genomes. Zoospores are chemotactic to the chitin monomer N-acetylglucosamine and accelerate development when grown with chitin. The R. globosum secretome is dominated by different glycoside hydrolase (GH) family GH18 chitinases, with abundance matching reciprocal transcriptome mRNA sequences. Models of the secreted chitinases indicate a range of sizes and domain configurations. Along with R. globosum, the genomes of other chitinophilic chytrids also have expanded inventories of GH-encoding genes responsible for chitin processing. Several R. globosum GH18 chitinases have bacteria-like chitin-binding module domains, also present in the genomes of other chitinophilic chytrids yet absent in non-chitinophilic chytrids. Chemotaxis, increased abundance and diversity of secreted chitinases, complemented with the acquisition of novel chitin-binding capability, are probably adaptive traits that facilitate chitin saprotrophy. Our study reveals the underpinning mechanisms that have supported the niche expansion of some chytrids to utilize lucrative chitin-rich particles in aquatic ecosystems and is a demonstration of the adaptive ability of this successful fungal group.

Chitin

HIF1A+CSF3R+ neutrophils-dominated hypoxic niche induced metabolic reprogramming for neoadjuvant therapy resistance in NSCLC.

BACKGROUND: Non-small cell lung cancer (NSCLC) is one of the frequently occurring cancers characterized by molecular heterogeneity and multiple immune cell infiltration patterns, which are associated with treatment sensitivity and resistance. However, the specific microenvironmental cells and their mechanisms that lead to treatment resistance in patients need to be explored in greater depth. METHODS: On the basis of patients receiving neoadjuvant therapy in our center, a multicenter, multicohort NSCLC spatial transcriptome, single-cell transcriptome, T-cell receptor repertoire sequencing, bulk RNA transcriptome, phosphorylated proteome, genome mutation, and clinical data were included for a comprehensive assessment of the therapeutic and prognostic impact of HIF1A+ CSF3R+ neutrophils in NSCLC. In vitro experiments validated the functional phenotype of HIF1A+ CSF3R+ neutrophils and co-localization interactions with other cellular subpopulations. Gradient boosting machine (GBM) constructed region of interest (ROI) models for evaluation. Computer-aided drug design (CADD) was used to predict targeted small molecule drugs, and in vivo mouse models were constructed to assess the effectiveness of the combination treatment regimen. RESULTS: Centered on HIF1A+ CSF3R+ neutrophils, recruited exhausted T cells and stromal cells form a hypoxic niche within the tumor region, which was enriched in non-response patients. ROI composed of these specific cellular subpopulations, associated with senescence and glycolysis, accurately predicting NSCLC progression, prognosis, and microenvironment composition. CADD analysis identified that platycodin-D2 specifically targeted CSF3R, reducing HIF1A expression and inhibiting neutrophil activity. Combining navitoclax, platycodin-D2 with anti-programmed cell death protein 1 (PD-1) significantly suppressed tumor proliferation and improved the immunosuppressive microenvironment. CONCLUSION: Our study emphasized the role of HIF1A+ CSF3R+ neutrophils in immunotherapeutic resistance of NSCLC, constructed a microenvironmental immune dysregulation network in a hypoxic ecological niche with HIF1A+ CSF3R+ neutrophils as the center. Platycodin-D2 specifically targeted HIF1A+ CSF3R+ neutrophils, enhancing the efficacy of anti-PD-1 therapy in NSCLC.

Humans

Genetic variation influences food-sharing sociability in honey bees.

Individual variation in sociability is a central feature of every society. This includes honey bees, with some individuals well connected and sociable, and others at the periphery of their colony's social network. However, the genetic and molecular bases of sociability are poorly understood. Trophallaxis-a behavior involving sharing liquid with nutritional and signaling properties-comprises a social interaction and a proxy for sociability in honey bee colonies: more sociable bees engage in more trophallaxis. Here, we identify genetic and molecular mechanisms of trophallaxis-based sociability by combining genome sequencing, brain transcriptomics, and automated behavioral tracking. A genome-wide association study (GWAS) identified 18 single nucleotide polymorphisms (SNPs) associated with variation in sociability. Several SNPs were localized to genes previously associated with sociability in other species, including in the context of human autism, suggesting shared molecular mechanisms of sociability. Variation in sociability also was linked to differential brain gene expression, particularly genes associated with neural signaling and development. Using comparative genomic and transcriptomic approaches, we also detected evidence for divergent mechanisms underpinning sociability across species, including those related to reward sensitivity and encounter probability. These results highlight both potential evolutionary conservation of the molecular roots of sociability and points of divergence.

Animals

The maternal-to-zygotic transition is a critical window for PFOA-induced disruption of developmental programming.

Early embryogenesis is governed by precisely timed gene regulatory programs that coordinate cell fate specification, tissue patterning, and morphogenesis. The maternal-to-zygotic transition (MZT) represents a pivotal developmental milestone during which regulatory control shifts from maternally deposited transcripts to activation of the zygotic genome. Disruption of this transition has the potential to alter developmental trajectories with lasting consequences. Per- and polyfluoroalkyl substances (PFAS), environmentally persistent contaminants, have been linked to developmental abnormalities, yet their impact on core embryonic gene regulatory networks especially with exposure during MZT is not well understood. Using zebrafish (Danio rerio), a tractable vertebrate model and New Approach Methodology (NAM), we investigated how PFAS exposure during the MZT alters early developmental programming. Embryos were exposed starting at different times before and within the MZT time window and collected at 24 h post-fertilization (hpf) for transcriptomic analysis. Targeted qRT-PCR revealed dysregulation of genes controlling transcriptional activation, lineage specification, proliferation, and differentiation. Whole-transcriptome RNA sequencing (RNA-seq) further identified widespread perturbations in gene networks governing transcriptional regulation, cell signaling, and embryonic morphogenesis. Temporal analysis revealed that exposure beginning at 3.5 hpf, followed by 8 hpf, corresponding to early zygotic genome activation and near completion of zygotic activation, respectively, resulted in the greatest differential gene expression changes at 24 hpf. Consistent with these early gene regulatory perturbations, larvae exposed starting at 8 hpf also exhibited altered behavior at 5 days post-fertilization. Together, these findings demonstrate that PFAS exposure during MZT disrupts the establishment of embryonic gene regulatory networks, linking environmental toxicant exposure to altered developmental patterning and organismal outcomes. This work underscores the vulnerability of early developmental transitions to environmental perturbation and positions MZT as a critical window of susceptibility during development.

NAMs (new approach methodologies)

Prevalence and chronology of colibactin-associated mutational processes and their microbiome spectra in Japanese colorectal cancer.

The incidence of colorectal cancer (CRC) has risen in recent decades, with a disproportionate increase observed among younger individuals in Japan and other countries. The etiological contribution of the gut microbiota to CRC pathogenesis is recognized, yet the mechanisms involved remain to be fully clarified. Here we integrated whole-genome sequencing (WGS) and transcriptome profiling of CRC with whole-genome metagenomic sequencing of fecal samples to interrogate host-microbiome interactions at high resolution. Application of interpretable artificial intelligence enabled the stratification of CRC into four distinct microbiome-informed subtypes. WGS analysis identified mutational signatures SBS88 and ID18, linked to colibactin exposure, as early clonal events detected in 44.8% of non-hypermutated patients. Notably, these signatures were significantly more frequent among patients born after the 1960s. Microbiome-based subclassification revealed subtype-specific clinical and molecular features. Collectively, our findings indicate that colibactin exposure constitutes a prevalent and potentially modifiable risk factor for CRC in the Japanese population.

Humans

Molecular hallmarks of excitatory and inhibitory neuronal resilience to Alzheimer's disease.

BACKGROUND: A significant proportion of individuals maintain cognition despite extensive Alzheimer's disease (AD) pathology, known as cognitive resilience. Understanding the molecular mechanisms that protect these individuals could reveal therapeutic targets for AD. METHODS: This study defines molecular and cellular signatures of cognitive resilience by integrating bulk RNA and single-cell transcriptomic data with genetics across multiple brain regions. We analyzed data from the Religious Order Study and the Rush Memory and Aging Project (ROSMAP), including bulk RNA sequencing (n = 631 individuals) and multiregional single-nucleus RNA sequencing (n = 48 individuals). Subjects were categorized into AD, resilient, and control based on β-amyloid and tau pathology, and cognitive status. We identified and prioritized protected cell populations using whole-genome sequencing-derived genetic variants, transcriptomic profiling, and cellular composition. RESULTS: Transcriptomics and polygenic risk analysis position resilience as an intermediate AD state. Only GFAP and KLF4 expression distinguished resilience from controls at tissue level, whereas differential expression of genes involved in nucleic acid metabolism and signaling differentiated AD and resilient brains. At the cellular level, resilience was characterized by broad downregulation of LINGO1 expression and reorganization of chaperone pathways, specifically downregulation of Hsp90 and upregulation of Hsp40, Hsp70, and Hsp110 families in excitatory neurons. MEF2C, ATP8B1, and RELN emerged as key markers of resilient neurons. Excitatory neuronal subtypes in the entorhinal cortex (ATP8B+ and MEF2Chigh) exhibited unique resilience signaling through activation of neurotrophin (BDNF-NTRK2, modulated by LINGO1) and angiopoietin (ANGPT2-TEK) pathways. MEF2C+ inhibitory neurons were over-represented in resilient brains, and the expression of genes associated with rare genetic variants revealed vulnerable somatostatin (SST) cortical interneurons that survive in AD resilience. The maintenance of excitatory-inhibitory balance emerges as a key characteristic of resilience. CONCLUSIONS: We have defined molecular and cellular hallmarks of cognitive resilience, an intermediate state in the AD continuum. Resilience mechanisms include preserved neuronal function, balanced network activity, and activation of neurotrophic survival signaling. Specific excitatory neuronal populations appear to play a central role in mediating cognitive resilience, while a subset of vulnerable interneurons likely provides compensation against AD-associated hyperexcitability. This study offers a framework to leverage natural protective mechanisms to mitigate neurodegeneration and preserve cognition in AD.

Humans

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta

Nonadditive gene expression and reduced homoeolog expression bias in an intraspecific hexaploid wheat hybrid.

BACKGROUND: Intraspecific hybridization in allopolyploid plants can generate additive and nonadditive changes in gene expression through interactions between divergent parental genomes. However, how it simultaneously affects gene expression and the relative expression of homoeologs in higher-order polyploids is less well understood. To study this, we sequenced seedling leaf transcriptomes and profiled gene body methylation in two hexaploid wheat (Triticum aestivum L.) cultivars and their F₁ hybrids. RESULTS: Although only 4.3% of genes differed in expression between the parents, 22.3% deviated from mid-parent expression in the hybrids, with many showing transgressive expression. 32.1% of triads contained at least one homoeolog that deviated from mid-parent expression, and all three homoeologs deviated in 11% of triads, substantially more than expected by chance. Triads in which all three homoeologs were overexpressed also showed reduced differences in expression among homoeologs. Greater parental divergence in relative homoeolog expression was associated with nonadditive expression. Genes lacking gene body methylation were also more likely to show dominant or transgressive expression, whereas gene body methylation was associated with more balanced homoeolog expression and additive or conserved expression. CONCLUSIONS: Intraspecific hybridization in hexaploid wheat, even without a change in ploidy, was associated with widespread nonadditive gene expression and altered relative homoeolog expression within triads. These responses were associated with parental differences in homoeolog expression and the absence of gene body methylation. Although our findings are limited to seedling leaves from a single intraspecific cross, they provide a basis for testing the generality of these patterns across tissues, developmental stages, and genetic backgrounds.

Triticum

Colorectal Liver Metastasis Pathomics Model: Integrating Single-Cell and Spatial Transcriptome Analysis With Pathomics for Predicting Liver Metastasis in Colorectal Cancer.

The liver is the primary target organ for hematologic metastasis of colorectal cancer (CRC), and CRC liver metastasis (CRLM) often precludes radical resection, making it the leading cause of death in patients with CRC. To improve the identification and prediction of liver metastasis risk, we identified a cell type of liver metastasis--triggering malignant cells (LMTMCs) through integrating single-cell RNA sequencing and spatial transcriptome analysis. Multiomics cell communication analysis indicated that the interaction between fibroblasts and LMTMCs through the COL1A1-CD44/SDC4 and LAMA4-CD44 signaling axes could promote CRLM. By applying the one-class logistic regression algorithm, we developed a CRLM scoring system in the bulk RNA-sequencing data according to the abundance of LMTMCs in each individual. Using the grouping labels derived from the CRLM scoring system in the bulk data and the corresponding whole-slide images without any manual annotations at the region or pixel level, processed via slide-level weakly supervised learning, a deep-learning model based on the ResNet18 architecture, called Colorectal Liver Metastasis Pathomics Model, was developed to predict the risk of liver metastasis in patients with CRC. The Colorectal Liver Metastasis Pathomics Model achieved an area under the curve of 0.84 at the internal test set of The Cancer Genome Atlas-CRC histology images. In the external independent validation sets, namely the Affiliated Hospital of Southwest Medical University and the Affiliated Traditional Chinese Medicine Hospital of Southwest Medical University cohorts, the areas under the curve were 0.89 and 0.72, respectively, indicating effective classification performances. This study provided new insights and tools for the early identification of CRLM and demonstrated the potential of combining multiomics with deep learning-based pathomics in cancer research.

Humans

RNA sequencing offers new diagnostic opportunities in neurodevelopmental disorders: A systematic review.

PURPOSE: Transcriptomics by way of RNA sequencing (RNAseq) has emerged as a means to increase the diagnostic yield in genetic conditions. In this systematic review, we focus on the contribution of transcriptomics to improve the diagnostic yield in neurodevelopmental disorders. METHODS: We performed a systematic literature search in PubMed until January 2024, including articles describing diagnostic RNAseq on at least 1 individual with a primary neurodevelopmental phenotype. We extracted data on cohort size, phenotype, sample tissue, previously used diagnostic methods, added diagnostic yield of RNAseq, the use of control samples, and technical aspects of the RNA sequencing methodology. RESULTS: A total of 17 articles were eligible for inclusion in the systematic review. We found an average added diagnostic yield of 15.5% through RNA sequencing for individuals with neurodevelopmental disorders. There is heterogeneity in the tissue type, reported quality measures, and the computational pipeline. CONCLUSION: The significantly increased diagnostic yield demonstrates the value of this novel tool in the diagnostic setting of neurodevelopmental disorders. Our results offer an overview of common methodologies for RNAseq and allow us to formulate recommendations for genetic labs and clinicians when implementing RNAseq as a diagnostic tool. Lastly, we provide recommendations for future publications to increase transparency and reproducibility.

Humans

Spatial transcriptomics-aided localization for single-cell transcriptomics with STALocator.

Single-cell RNA-sequencing (scRNA-seq) techniques can measure gene expression at single-cell resolution but lack spatial information. Spatial transcriptomics (ST) techniques simultaneously provide gene expression data and spatial information. However, the data quality of the spatial resolution or gene coverage is still much lower than the quality of the single-cell transcriptomics data. To this end, we develop a ST-Aided Locator for single-cell transcriptomics (STALocator) to localize single cells to corresponding ST data. Applications on simulated data showed that STALocator performed better than other localization methods. When applied to the human brain and squamous cell carcinoma data, STALocator could robustly reconstruct the relative spatial organization of critical cell populations. Moreover, STALocator could enhance gene expression patterns for Slide-seqV2 data and predict genome-wide gene expression data for fluorescence in situ hybridization (FISH) and Xenium data, leading to the identification of more spatially variable genes and more biologically relevant Gene Ontology (GO) terms compared with the raw data. A record of this paper's transparent peer review process is included in the supplemental information.

Single-Cell Analysis

Integrated transcriptomic analysis of mRNA and miRNA in Brown adipose tissue of the greater horseshoe bats during hibernation.

Hibernation enables animals survive harsh environments by conserving energy through reduced metabolism and body temperature. Brown adipose tissue (BAT) plays a critical role in non-shivering thermogenesis, crucial for warming up during arousal phase. The greater horseshoe bats (Rhinolophus nippon) are typical hibernators and non-shivering thermogenesis in BAT tissue may persist throughout the arousal process in bats. This study examines gene expression and regulatory changes in BAT of these bats across active, hibernation, and arousal phases using transcriptome and miRNA sequencing. A total of 2721 differentially expressed mRNAs and 268 differentially expressed miRNAs were identified. The results reveal that the BAT transcriptome undergoes state-dependent remodeling throughout the hibernation process. The most pronounced divergence occurs between the active phase and torpor, involving cell cycle arrest, immunosuppression, thermogenic signal desensitization, and upregulation of lipid metabolism and autophagy pathways, reflecting the coordinated adaptation of energy conservation and thermogenic reserve. In contrast, transcriptional alterations between torpor and arousal are extremely limited, indicating that torpid BAT is already pre-primed for thermogenesis and requires only modest transcriptional adjustments to activate heat production. Notably, although body temperature recovers to active-phase levels during arousal, the molecular signature of BAT remains highly similar to that of the torpid state. Furthermore, the core thermogenic gene UCP1 showed no significant expression differences across the three groups. In conclusion, this study systematically delineates the miRNA-mRNA regulatory landscape of bat BAT across the hibernation process, and deepens our understanding of the thermoregulatory mechanisms underlying mammalian hibernation.

BAT

Systematic Dissection of Key Driver Perturbation Signatures in Single Cells via ECCITE-seq.

CRISPR screens, such as expanded CRISPR-compatible cellular indexing of transcriptomes and epitopes by sequencing (ECCITE-seq), enable the simultaneous measurement of transcriptomes, gRNA identity, and cell-surface protein expression at single-cell resolution to systematically interrogate gene function. This platform provides a powerful and scalable experimental approach for validating disease-associated regulators identified by large-scale association studies and other computational methods, including network-based analyses of multi-omics data. Here, as an example application, we describe an ECCITE-seq framework to characterize the transcriptomic consequences of perturbing multiple neuronal key driver genes associated with Alzheimer's disease (AD) in human-induced pluripotent stem cell (hiPSC)-derived neurons. More broadly, by integrating customized pooled gRNA libraries with different CRISPR effectors across multiple cell types, this approach allows for the assessment of the regulatory impact of candidate genes implicated in development and disease processes.

Humans

Perplexity as a Metric for Isoform Diversity in the Human Transcriptome.

Long-read sequencing (LRS) has revealed a far greater diversity of RNA isoforms than earlier technologies, increasing the critical need to determine which, and how many, isoforms per gene are biologically meaningful. To define the space of relevant isoforms from LRS, many existing analysis pipelines rely on arbitrary expression cutoffs, but a single threshold cannot accommodate the broad variability in isoform complexity across genes, cell-types, and disease states captured by LRS. To address this, we propose using perplexity-an interpretable measure derived from entropy-that quantifies the effective number of isoforms per gene based on the full, unfiltered isoform ratio distribution. Calculating perplexity for 124 ENCODE4 PacBio LRS datasets spanning 55 human cell types, we show that it provides intuitive assessments of isoform diversity and captures uncertainty across genes with varying complexity. Perplexity can be calculated at multiple gene regulatory levels-from transcript to protein-to compare how isoform diversity is reduced across stages of gene expression. On average, genes have an ORF-level perplexity of 2.1, indicating production of two distinct protein isoforms. We extended this analysis to evaluate expression variation across tissues and identified 4,593 ORFs across 3,102 genes with moderate to extreme tissue-specificity. We propose perplexity as a consistent, quantitative metric for interpreting isoform diversity across genes, cell types, and disease states. All results are compiled into a community resource to enable cross-study comparisons of novel isoforms.

Journal Article