Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

A Comprehensive Bioinformatics Approach to Analysis of Variants: Variant Calling, Annotation, and Prioritization.

Next-Generation Sequencing (NGS), also known as high-throughput sequencing technologies, has enabled rapid and efficient sequencing of large amounts of DNA and RNA. These technologies have revolutionized the field of genomics, transcriptomics, and proteomics and have been widely used in cancer research, leading to advances in clinical diagnosis and treatment. Improvements in the NGS technologies enabled millions of fragments to be sequenced simultaneously in a time- and cost-effective manner and resulted in large amount of genomic data which require efficient analysis methods. Analysis of the genomic data requires both efficient computer resources and bioinformatics approaches. This chapter details a comprehensive computational approach and analysis steps for genomic data analysis.

Computational Biology↗

The transcriptome and its translation during recovery from cell cycle arrest in Saccharomyces cerevisiae.

Complete genome sequences together with high throughput technologies have made comprehensive characterizations of gene expression patterns possible. While genome-wide measurement of mRNA levels was one of the first applications of these advances, other important aspects of gene expression are also amenable to a genomic approach, for example, the translation of message into protein. Earlier we reported a high throughput technology for simultaneously studying mRNA level and translation, which we termed translation state array analysis, or TSAA. The current studies test the proposition that TSAA can identify novel instances of translation regulation at the genome-wide level. As a biological model, cultures of Saccharomyces cerevisiae were cell cycle-arrested using either alpha-factor or the temperature-sensitive cdc15-2 allele. Forty-eight mRNAs were found to change significantly in translation state following release from alpha-factor arrest, including genes involved in pheromone response and cell cycle arrest such as BAR1, SST2, and FAR1. After the shift of the cdc15-2 strain from 37 degrees C to 25 degrees C, 54 mRNAs were altered in translation state, including the products of the stress genes HSP82, HSC82, and SSA2. Thus, regulation at the translational level seems to play a significant role in the response of yeast cells to external physical or biological cues. In contrast, surprisingly few genes were found to be translationally controlled as cells progressed through the cell cycle. Additional refinements of TSAA should allow characterization of both transcriptional and translational regulatory networks on a genomic scale, providing an additional layer of information that can be integrated into models of system biology and function.

Cell Cycle↗

Renal transcriptomes: segmental analysis of differential expression.

BACKGROUND/AIMS: Progress accomplished by complete genomes and cDNA-sequencing projects calls for methods that fully use these resources to study gene expression patterns in characterized cell populations. However, since the number of functional genes cannot be readily inferred from the genomic sequence, it is highly desirable to make use of methods enabling to study both known and unknown genes. METHODS: The method of serial analysis of gene expression provides short diagnostic cDNA tags without bias towards known genes. In addition, the frequency of each tag in the library conveys quantitative information on gene expression. A microassay was set-up to perform serial analysis of gene expression in minute samples such as those obtained by microdissecting nephron segments. RESULTS: Studies carried out in the thick ascending limb of Henle's loop and the collecting duct of the mouse kidney provided expression data for several thousand genes. Known markers were found appropriately enriched, and several of the thick ascending limb or collecting duct specific transcripts had no database match. CONCLUSIONS: The microassay for serial analysis of gene expression makes possible large-scale quantitative measurements of mRNA levels in nephron segments. The comprehensive picture generated by analyzing both known and unknown transcripts in defined cell populations should help to discover genes with dedicated functions.

Animals↗

The genexpress IMAGE knowledge base of the human muscle transcriptome: a resource of structural, functional, and positional candidate genes for muscle physiology and pathologies.

Sequence, gene mapping, and expression data corresponding to 910 genes transcribed in human skeletal muscle have been integrated to form the muscle module of the Genexpress IMAGE Knowledge Base. Based on cDNA array hybridization, a set of 14 transcripts preferentially or specifically expressed in muscle have been selected and characterized in more detail: Their pattern of expression was confirmed by Northern blot analysis; their structure was further characterized by full-insert cDNA sequencing and cDNA extension; the map location of the corresponding genes was refined by radiation hybrid mapping. Five of the 14 selected genes appear as interesting positional and functional candidate genes to study in relation with muscle physiology and/or specific orphan muscular pathologies. One example is discussed in more detail. The expression profiling data and the associated Genexpress Index2 entries for the 910 genes and the detailed characterization of the 14 selected transcripts are available from a dedicated Web server at. The database has been organized to provide the users with a working space where they can find curated, annotated, integrated data for their genes of interest. Different navigation routes to exploit the resource are discussed.

Base Sequence↗

Genomic and Immune Landscape of Pancreatic Ductal Adenocarcinoma Associated with Germline Pathogenic Variants in ATM.

PURPOSE: Germline pathogenic variants (PV) in ATM increase the risk of pancreatic ductal adenocarcinoma (PDAC), but the underlying tumor biology of PDAC associated with germline PV in ATM has not been adequately explored. EXPERIMENTAL DESIGN: Whole-genome, whole-exome, and RNA sequencing were performed on PDAC tumors from 25 germline ATM PV carriers diagnosed at Mayo Clinic between 2007 and 2017. Somatic and copy-number alterations, mutational signatures, transcriptomic subtypes, and the immune landscape were evaluated. RESULTS: High-quality whole-exome and whole-genome sequencing were obtained from 21 and 15 tumors, respectively. Biallelic inactivation of ATM was observed in 87%, KRAS PV in 90%, CDKN2A homozygous loss in 60%, and TP53 alterations in <10% of these tumors. A predominant clock-like mutational signature was present in all samples. Whole-transcriptome analysis identified that the aberrantly differentiated endocrine exocrine subtype accounted for 18% of PDAC and was consistently associated with >5-year overall survival. In addition, a 28-gene expression-based signature associated with overall survival was identified and further validated in The Cancer Genome Atlas cohort. Immune landscape analysis through CODEX identified enriched CD4 T-helper cell/tumor interactions and reduced B7H3-high cell/tumor interactions in ATM PV carriers compared with noncarriers. CONCLUSIONS: The observed absence of TP53 PV and enrichment for CDKN2A alterations in ATM tumors, along with differences in the mutational signatures, transcriptomic subtypes and immune landscape, improve our understanding of the mechanistic pathways involved in PDAC development in germline ATM PV carriers and help identify potential targeted therapeutic strategies.

Humans↗

Expression regulation network in papillae of sea cucumbers: Whole-transcriptome and DNA methylation datasets.

To elucidate the expression regulation network of papilla size of sea cucumbers (Apostichopus japonicus), the whole-transcriptome and DNA methylome datasets of different sizes of papillae in sea cucumbers were generated. Average clean bases of whole-transcriptome (16.35&#x2009;G) and DNA methylome (28.92&#x2009;G) were obtained using RNA sequencing and whole-genome bisulfite sequencing techniques. A total of 3,188 ceRNA networks were also identified including 3,081 long non-coding RNAs (lncRNA)/microRNAs (miRNA)/mRNA networks and 107 circular RNA (circRNA)/miRNA/mRNA networks. Methylome data indicate that there were 3,307 and 3,776 differentially methylated regions (DMRs) with high-level methylation as well as 3,125 and 3,016 DMRs with low-level methylation in big papillae compared to small papillae. The identified DMRs were mainly distributed in introns, promotors, or exons. The whole-transcriptome and DNA methylome datasets generated from this study not only established a robust theoretical foundation (especially from the epigenetic aspect) for elucidating expression regulation network determining papilla size in sea cucumbers but also can be a valuable resource of biomarker mining for papilla appearance-based selective breeding in sea cucumbers.

DNA Methylation↗

Comparative transcriptomic analysis of the gills and hepatopancreas of freshwater-cultured Litopenaeus vannamei under chronic nitrite stress.

To investigate the differences in molecular responses between the gills and hepatopancreas of freshwater-cultured Litopenaeus vannamei under chronic nitrite stress, a 30-day chronic stress experiment was conducted with a control group and a stress group. Transcriptomic analysis of the gills and hepatopancreas was performed using Illumina sequencing; differentially expressed genes (DEGs) were identified, and GO, KEGG, GSEA, PPI, and RT-qPCR validation were carried out. The results showed that 196 DEGs (161 up-regulated and 35 down-regulated) were identified in the gills, and 287 DEGs (199 up-regulated and 88 down-regulated) in the hepatopancreas, with only 18 DEGs shared between the two tissues. DEGs in the gills were enriched in oxidoreductase activity, glycerophospholipid metabolism, and tyrosine metabolism; DEGs in the hepatopancreas were enriched in lipid transporter activity, phagosome, ECM-receptor interaction, and riboflavin metabolism. GSEA revealed significant suppression of the mTOR pathway in the gills and the Polycomb complex pathway in the hepatopancreas. PPI network analysis identified hub genes P5CS and eEF2 in the gills, and PER, TUBB1, SHMT, and TUBB4B in the hepatopancreas. RT-qPCR validation was consistent with the RNA-seq results (R2&#xa0;=&#xa0;0.764). This study indicates that, under chronic nitrite stress, the gill response is centered on redox regulation and inhibition of growth metabolism, whereas the hepatopancreas response primarily involves lipid transport, cytoskeletal remodeling, and phagosome activation. The two tissues synergistically adapt through fundamental biosynthetic and motor protein pathways. This research provides molecular evidence for deciphering the nitrite tolerance mechanisms in freshwater-cultured shrimp.

Animals↗

A multi-modal survival prediction framework with group-based batch training and structural consistency alignment.

OBJECTIVE: Integrating whole-slide images (WSIs) with transcriptomic profiles is pivotal for enhancing cancer survival prediction. However, the intrinsic gigapixel resolution and variable sequence lengths of WSIs create a fundamental trade-off between training efficiency and the preservation of data heterogeneity in existing frameworks. Furthermore, substantial statistical and structural discrepancies between histological and genomic modalities often impede effective cross-modal alignment and fusion, thereby limiting prognostic accuracy. METHODS: We propose PRISM, an efficient multi-modal learning framework for integrating WSIs with transcriptomic profiles. To reconcile training efficiency with full data heterogeneity, PRISM first stochastically partitions variable-length WSI sequences into a main subset and a complementary residual subset, both of which are packed into fixed-length groups for batch training. The main subset is processed in the main branch, utilizing isolation masking to maintain intra-group sequence independence. Simultaneously, the residual subset is consolidated into "hyperslides" within a residual branch that leverages tailored supervision, effectively capturing inter-slide correlations. Furthermore, PRISM integrates an Informative Token Aggregation (ITA) module to reduce redundancy in WSIs and employs Cross-batch Structural Consistency Alignment (CBSCA) mechanism to enhance inter-modal structural connectivity. Finally, efficient cross-modal feature interaction is achieved through a Low-rank Bilinear Gated Fusion (LBGF) module. Code is available at https://github.com/Alisa2080/PRISM. RESULTS: Compared with existing methods, PRISM achieves the best overall C-index across five TCGA cohorts. On the larger TCGA-BRCA dataset, PRISM requires only 6&#xa0;hours of training time, substantially reducing computational cost relative to strong multimodal baselines. Furthermore, comprehensive evaluations demonstrate that PRISM achieves the best overall IBS ranking and favorable time-dependent AUC performance at 1, 3, and 5&#xa0;years, thereby delivering a more favorable trade-off between prognostic performance and computational efficiency. CONCLUSION: PRISM provides a favorable balance between predictive performance, calibration quality, and computational efficiency, highlighting its potential for practical deployment in multimodal survival modeling for computational pathology.

Humans↗

High MGMT expression identifies aggressive colorectal cancer with distinct genomic features and immune evasion properties.

INTRODUCTION: The epigenetic silencing of O6-methylguanine DNA methyltransferase (MGMT) is associated with reduced DNA repair capacity, carcinogenesis and increased sensitivity to alkylating chemotherapy. However, the biological role and clinical significance of MGMT overexpression in cancer remains poorly understood. METHODS: Using multiplexed quantitative immunofluorescence we measured the localized levels of MGMT protein, &#x3b3;H2AX and CD8+ T&#x2009;cells in multiple retrospective colorectal cancer (CRC) cohorts. Genomic and transcriptomic features of selected cases were also studied with whole exome DNA sequencing and genome-wide methylation analysis. MGMT-methylated human CRC cells SW620 were transfected with an MGMT-containing plasmid and co-cultured with allogeneic peripheral blood mononuclear cells. RESULTS: A subset of CRCs showed MGMT protein upregulation associated with lower &#x3b3;H2AX, reduced CD8+ tumor infiltrating lymphocytes (TILs), mismatch repair proficient (pMMR) status and shorter survival. CD8+ TILs were more distant from MGMT-expressing cells than MGMT-negative cells and the MGMT promoter methylation status did not highly correlate with MGMT protein levels in CRC. In genomic/transcriptomic analysis, high MGMT expression was associated with a lower nonsynonymous somatic mutational burden, higher transition-to-transversion mutation ratio, increased deleterious TP53 variants and distinct transcriptomic profiles. The exogenous expression of MGMT in SW620 CRC cells reduced the number of spontaneous nonsynonymous mutations, reproduced mutational features of MGMT-high CRC and limited the in vitro T-cell-mediated killing of malignant cells induced by proinflammatory cytokines in tumor/immune cell co-cultures. CONCLUSIONS: MGMT overexpression identifies a previously undescribed subset of CRCs with distinct biological and clinical properties including reduced mutagenesis, adaptive immune evasion, predominantly pMMR phenotype and aggressive clinical course. Direct, quantitative assessment of MGMT protein expression using spatially resolved analysis is more reliable than inference of MGMT expression by promoter methylation status in CRC.

Humans↗

Ulmus minor response to Dutch elm disease: de novo transcriptome assembly and annotation.

Dutch elm disease (DED), caused by Ophiostoma novo-ulmi (ONU), has devastated elm populations across Europe and North America since the 20th century. In this work, a de novo transcriptome assembly of Ulmus minor in response to ONU is presented. We used two DED-resistant genotypes, MDV2.3 and VAD2, and one DED-susceptible genotype, MDV1, to capture responses to ONU at four time points post-inoculation (6, 24, 72, and 144&#x2009;hours). RNA from collected samples was isolated and sequenced producing 60.88&#x2009;M 100&#x2009;bp paired-end reads per sample. We performed a de novo transcriptome assembly combining data from the three genotypes. The assembly was functionally annotated and validated through differential gene expression analysis of the response. This dataset provides a valuable resource for studying molecular mechanisms of DED resistance in elms, contributing to broadening our understanding of tree immunity and facilitating potential applications in functional annotation of future genome assemblies.

Transcriptome↗

New trends in bioinformatics: from genome sequence to personalized medicine.

Molecular medicine requires the integration and analysis of genomic, molecular, cellular, as well as clinical data and it thus offers a remarkable set of challenges to bioinformatics. Bioinformatics nowadays has an essential role both, in deciphering genomic, transcriptomic, and proteomic data generated by high-throughput experimental technologies, and in organizing information gathered from traditional biology and medicine. The evolution of bioinformatics, which started with sequence analysis and has led to high-throughput whole genome or transcriptome annotation today, is now going to be directed towards recently emerging areas of integrative and translational genomics, and ultimately personalized medicine.Therefore considerable efforts are required to provide the necessary infrastructure for high-performance computing, sophisticated algorithms, advanced data management capabilities, and-most importantly-well trained and educated personnel to design, maintain and use these environments. This review outlines the most promising trends in bioinformatics, which may play a major role in the pursuit of future biological discoveries and medical applications.

Computational Biology↗

Longitudinal Multi-Organ Transcriptomic Atlas of Salt-Induced Hypertension.

BACKGROUND: Salt-sensitive hypertension is a prevalent and clinically significant subtype of hypertension, where increased dietary salt intake elevates blood pressure and causes injury to multiple organ systems. Despite extensive research, dynamic molecular changes and conserved versus organ-specific transcriptional programs in hypertensive multi-organ damage remain poorly understood. Defining complex molecular pathways both in a temporal sequence and in an organ-specific manner is essential for developing targeted, precision therapies to mitigate hypertensive disease burden. METHODS: We generated a longitudinal multi-organ transcriptomic atlas of salt-sensitive hypertension using RNA sequencing of kidney cortex, kidney medulla, heart, and liver from Dahl salt-sensitive rats across four disease stages. A comprehensive bioinformatic analysis mapped dynamic transcriptional programs, evaluated 50 biological pathways, and defined upstream regulators. Histological and biochemical assays complemented transcriptomic analysis, while integration with human genome-wide association studies (GWAS) and compound-transcriptome analysis provided translational insights and identified candidate therapeutics. RESULTS: Salt-induced hypertension elicited both shared and tissue-specific transcriptional programs that evolved with disease progression. The kidney medulla showed robust early immune activation with metabolic suppression, while the cortex exhibited transient metabolic activation before declining and initiating immune activation. The liver and heart showed time-dependent metabolic and inflammatory remodeling. Cross-organ comparisons revealed a shared early proliferative response that converged on proinflammatory and fibrotic signatures. Upstream regulator analysis identified 79 time- and tissue-specific transcription factors associated with gene expression dynamics. GWAS integration analysis revealed endocrine signaling, ion transport, lipid metabolism, and detoxification as conserved pathways across species, underscoring the translational relevance of the model and study. Predictive compound-transcriptome analyses identified kinase inhibitors targeting phosphoinositide 3-kinase, mechanistic target of rapamycin and cyclin-dependent kinases as top candidates to counteract maladaptive transcriptional programs. CONCLUSIONS: This study defines temporal and tissue-specific transcriptomic remodeling in salt-sensitive hypertension and highlights the need for precision interventions to prevent progressive organ damage.

Journal Article↗

Functional characterization of the 9q34.13 locus identifies RAPGEF1 as a candidate gene modulating risk for melanoma and nevi via RAS activation.

Genome-wide association studies identified a melanoma- and nevus count-associated locus on chromosome band 9q34.13. Fine-mapping and melanocyte expression data collectively suggest two potential risk genes with opposite associations with risk: higher levels of Rap guanine nucleotide exchange factor 1 (RAPGEF1) and lower levels of uridine-cytidine kinase 1 (UCK1). Colocalization analyses and conditional transcriptome-wide association studies (TWASs) suggest multiple causal cis-regulatory sequence variants in partial linkage disequilibrium (LD) to each other. Melanocyte capture-HiC and CRISPR inhibition demonstrated regulatory interactions between fine-mapped variants and the RAPGEF1 and UCK1 promoters. Focusing on RAPGEF1, we demonstrate that RAPGEF1 expression promotes melanocyte growth and drives colony formation of human immortalized melanocytes. Following treatment with human epidermal growth factor (EGF), RAPGEF1 overexpression activated both RAP1 and RAS. Further, we show that RAPGEF1 expression is significantly enriched in melanomas that lack strongly activating RAS-MAPK pathway mutations, which suggests that RAPGEF1 may promote oncogenic RAS-MAPK pathway signaling in melanomas. Furthermore, in these tumors, we provide preliminary evidence to support the prognostic relevance of RAPGEF1 expression in individuals whose melanomas lack RAS or BRAF mutations. Together with other recent studies, these data suggest that germline variation influencing RAS activation may play a key role in nevus development and melanoma risk.

GWAS↗

Multiomics Integration Identifies a Molecular Subtype of Intrahepatic Cholangiocarcinoma With Enhanced Benefit From Adjuvant Therapy.

Intrahepatic cholangiocarcinoma (iCCA) is a molecularly heterogeneous liver cancer with a poor prognosis. Improved stratification is needed to guide postoperative therapy. In this study, we applied integrative multiomics analysis to classify iCCA and identify biomarkers predictive of adjuvant treatment benefit. Using publicly available datasets (including whole exome sequencing, RNA sequencing, proteomics, and phosphoproteomics from FU-iCCA cohort and a transcriptomic cohort GSE244807), we defined 3 robust molecular subtypes of iCCA. These subtypes exhibited distinct genomic alterations, pathway activation, and immune microenvironments, with significant differences in overall survival (OS). Through protein-protein interaction network analysis and consensus feature selection using 10 clustering algorithms, we prioritized 8 marker genes distinguishing the subtypes. A Cox proportional-hazards model constructed from these markers stratified patients into high- and low-risk groups. High-risk iCCA, characterized by elevated expression of markers such as CLDN18, MUC1, and MUC5AC, had significantly worse OS in the absence of adjuvant therapy. Notably, in an independent validation of 174 patients with iCCA who underwent resection (single-center cohort), high expression of any of these 3 markers were associated with markedly prolonged OS in patients who received adjuvant chemotherapy or chemoembolization, compared with those who did not. In contrast, marker-negative patients showed no clear benefit from adjuvant therapy. In conclusion, our multiomics approach identified a high-risk, mucin-enriched subtype of iCCA. CLDN18, MUC1, and MUC5AC emerge as candidate predictive biomarkers for adjuvant chemotherapy benefit in iCCA, warranting prospective validation to improve personalized postoperative management.

Humans↗

Hembase: browser and genome portal for hematology and erythroid biology.

Hembase (http://hembase.niddk.nih.gov) is an integrated browser and genome portal designed for web-based examination of the human erythroid transcriptome. To date, Hembase contains 15,752 entries from erythroblast Expressed Sequenced Tags (ESTs) and 380 referenced genes relevant for erythropoiesis. The database is organized to provide a cytogenetic band position, a unique name as well as a concise annotation for each entry. Search queries may be performed by name, keyword or cytogenetic location. Search results are linked to primary sequence data and three major human genome browsers for access to information considered current at the time of each search. Hembase provides interested scientists and clinical hematologists with a genome-based approach toward the study of erythroid biology.

Computational Biology↗

Heart transplantation changes the expression of distinct gene families.

We took advantage of the combination of a rat heart transplantation model with a modified differential display RT-PCR method to identify transcriptome changes in the right atria from transplanted compared with native hearts. Based on sequence homology search, the 37 cDNAs differentially displayed both 2 and 7 days posttransplantation were categorized into 7 unknown transcripts, 16 expressed sequence tags (ESTs), and 14 partially or completely characterized genes. The last group cDNAs, validated by relative RT-PCR, belonged to diverse gene families involved in specific metabolisms, protein synthesis, cell signaling, and transcription. Furthermore, we identified differential transcripts corresponding to denervation and fetal gene reexpression. We found coordinate downregulation of genes involved in energy metabolism and protein synthesis regulation, similar to that reported for senescent skeletal muscle. From these transcriptome changes, we propose that heart transplants and senescent muscles share common molecular mechanisms.

Animals↗

Novel Genetic Loci in Early-Onset Gout Derived From Whole-Genome Sequencing of an Adolescent Gout Cohort.

OBJECTIVE: Mechanisms underlying the adolescent-onset and early-onset gout are unclear. This study aimed to discover variants associated with early-onset gout. METHODS: We conducted whole-genome sequencing in a discovery adolescent-onset gout cohort of 905 individuals (gout onset 12 to 19 years) to discover common and low-frequency single-nucleotide variants (SNVs) associated with gout. Candidate common SNVs were genotyped in an early-onset gout cohort of 2,834 individuals (gout onset &#x2264;30 years old), and meta-analysis was performed with the discovery and replication cohorts to identify loci associated with early-onset gout. Transcriptome and epigenomic analyses, quantitative real-time polymerase chain reaction and RNA sequencing in human peripheral blood leukocytes, and knock-down experiments in human THP-1 macrophage cells investigated the regulation and function of candidate gene RCOR1. RESULTS: In addition to ABCG2, a urate transporter previously linked to pediatric-onset and early-onset gout, we identified two novel loci (Pmeta < 5.0 &#xd7; 10-8): rs12887440 (RCOR1) and rs35213808 (FSTL5-MIR4454). Additionally, we found associations at ABCG2 and SLC22A12 that were driven by low-frequency SNVs. SNVs in RCOR1 were linked to elevated blood leukocyte messenger RNA levels. THP-1 macrophage culture studies revealed the potential of decreased RCOR1 to suppress gouty inflammation. CONCLUSION: This is the first comprehensive genetic characterization of adolescent-onset gout. The identified risk loci of early-onset gout mediate inflammatory responsiveness to crystals that could mediate gouty arthritis. This study will contribute to risk prediction and therapeutic interventions to prevent adolescent-onset gout.

Humans↗

Impact of alternative initiation, splicing, and termination on the diversity of the mRNA transcripts encoded by the mouse transcriptome.

We analyzed the FANTOM2 clone set of 60,770 RIKEN full-length mouse cDNA sequences and 44,122 public mRNA sequences. We developed a new computational procedure to identify and classify the forms of splice variation evident in this data set and organized the results into a publicly accessible database that can be used for future expression array construction, structural genomics, and analyses of the mechanism and regulation of alternative splicing. Statistical analysis shows that at least 41% and possibly as much as 60% of multiexon genes in mouse have multiple splice forms. Of the transcription units with multiple splice forms, 49% contain transcripts in which the apparent use of an alternative transcription start (stop) is accompanied by alternative splicing of the initial (terminal) exon. This implies that alternative transcription may frequently induce alternative splicing. The fact that 73% of all exons with splice variation fall within the annotated coding region indicates that most splice variation is likely to affect the protein form. Finally, we compared the set of constitutive (present in all transcripts) exons with the set of cryptic (present only in some transcripts) exons and found statistically significant differences in their length distributions, the nucleotide distributions around their splice junctions, and the frequencies of occurrence of several short sequence motifs.

Alternative Splicing↗