Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Transcriptome analysis of channel catfish (Ictalurus punctatus): genes and expression profile from the brain.

Expressed sequence tag (EST) analysis was conducted using a complementary DNA (cDNA) library made from the brain mRNA of channel catfish (Ictalurus punctatus). As part of our transcriptome analysis in catfish to develop molecular reagents for comparative functional genomics, here we report analysis of 1201 brain cDNA clones. Of the 1201 clones, 595 clones (49.5%) were identified as known genes by BLAST searches and 606 clones (50.5%) as unknown genes. The 595 clones of known gene products represent transcripts of 251 genes. These known genes were categorized into 15 groups according to their biological functions. The largest group of known genes was the genes involved in translational machinery (21.4%) followed by mitochondrial genes (6.2%), structural genes (3.1%), genes homologous to sequences of unknown functions (2.3%), enzymes (2.7%), hormone and regulatory proteins (2.5%), genes involved in immune systems (2.1%), genes involved in sorting, transport, and metal metabolism (1.8%), transcriptional factors and DNA repair proteins (1.6%), proto-oncogenes (1.2%), lipid binding proteins (1.2%), stress-induced genes (0.7%), genes homologous to human genes involved in mental diseases (0.6%), and development or differentiation-related genes (0.3%). The number of genes represented by the 606 clones of unknown genes is not known at present, but the high percentage of clones showing no homology to any known genes in the GenBank databases may indicate that a great number of novel genes exist in teleost brain.

Animals↗

Bioinformatic approaches for accurate assessment of A-to-I editing in complete transcriptomes.

A-to-I RNA editing is an RNA modification that alters the RNA sequence relative to the its genomic blueprint. It is catalyzed by double-stranded RNA-specific adenosine deaminase (ADAR) enzymes, and contributes to the complexity and diversification of the proteome. Advancement in the study of A-to-I RNA editing has been facilitated by computational approaches for accurate mapping and quantification of A-to-I RNA editing based on sequencing data. In this chapter we review some of the main computational approaches currently used, describe potential hurdles, challenges and pitfalls, and discuss possible ways to mitigate them.

RNA Editing↗

MAO: a Multiple Alignment Ontology for nucleic acid and protein sequences.

The application of high-throughput techniques such as genomics, proteomics or transcriptomics means that vast amounts of heterogeneous data are now available in the public databases. Bioinformatics is responding to the challenge with new integrated management systems for data collection, validation and analysis. Multiple alignments of genomic and protein sequences provide an ideal environment for the integration of this mass of information. In the context of the sequence family, structural and functional data can be evaluated and propagated from known to unknown sequences. However, effective integration is being hindered by syntactic and semantic differences between the different data resources and the alignment techniques employed. One solution to this problem is the development of an ontology that systematically defines the terms used in a specific domain. Ontologies are used to share data from different resources, to automatically analyse information and to represent domain knowledge for non-experts. Here, we present MAO, a new ontology for multiple alignments of nucleic and protein sequences. MAO is designed to improve interoperation and data sharing between different alignment protocols for the construction of a high quality, reliable multiple alignment in order to facilitate knowledge extraction and the presentation of the most pertinent information to the biologist.

Databases, Genetic↗

Protocol to decode the role of transcriptionally active microbes in SARS-CoV-2-positive patients using an RNA-seq-based approach.

The elucidation of the role of microorganisms in human infections has been hindered by difficulties using conventional culture-based techniques. Here, we present a protocol for the investigation of transcriptionally active microbes (TAMs) using an RNA sequencing (RNA-seq)-based approach. We describe the steps for RNA isolation, viral genome sequencing, RNA-seq library preparation, and metatranscriptomic and transcriptomic analysis. This protocol permits a comprehensive evaluation of TAMs' contributions to the differential severity of infectious diseases, with a particular focus on diseases such as COVID-19. For complete details on the use and execution of this protocol, please refer to Devi et al.1.

Humans↗

Single-cell profiling of trabecular meshwork identifies mitochondrial dysfunction in a glaucoma model that is protected by vitamin B3 treatment.

Since the trabecular meshwork (TM) is central to intraocular pressure (IOP) regulation and glaucoma, a deeper understanding of its genomic landscape is needed. We present a multimodal, single-cell resolution analysis of mouse limbal cells (includes TM). In total, we sequenced 9,394 wild-type TM cell transcriptomes. We discovered three TM cell subtypes with characteristic signature genes validated by immunofluorescence on tissue sections and whole-mounts. The subtypes are robust, being detected in datasets for two diverse mouse strains and in independent data from two institutions. Results show compartmentalized enrichment of critical pathways in specific TM cell subtypes. Distinctive signatures include increased expression of genes responsible for 1) extracellular matrix structure and metabolism (TM1 subtype), 2) secreted ligand signaling to support Schlemm's canal cells (TM2), and 3) contractile and mitochondrial/metabolic activity (TM3). ATAC-sequencing data identified active transcription factors in TM cells, including LMX1B. Mutations in LMX1B cause high IOP and glaucoma. LMX1B is emerging as a key transcription factor for normal mitochondrial function and its expression is much higher in TM3 cells than other limbal cells. To understand the role of LMX1B in TM function and glaucoma, we single-cell sequenced limbal cells from Lmx1b V265D/+ mutant mice (2,491 TM cells). In V265D/+ mice, TM3 cells were uniquely affected by pronounced mitochondrial pathway changes. Mitochondria in TM cells of V265D/+ mice are swollen with a reduced cristae area, further supporting a role for mitochondrial dysfunction in the initiation of IOP elevation in these mice. Importantly, treatment with vitamin B3 (nicotinamide), to enhance mitochondrial function and metabolic resilience, significantly protected Lmx1b mutant mice from IOP elevation.

Journal Article↗

H-DBAS: alternative splicing database of completely sequenced and manually annotated full-length cDNAs based on H-Invitational.

The Human-transcriptome DataBase for Alternative Splicing (H-DBAS) is a specialized database of alternatively spliced human transcripts. In this database, each of the alternative splicing (AS) variants corresponds to a completely sequenced and carefully annotated human full-length cDNA, one of those collected for the H-Invitational human-transcriptome annotation meeting. H-DBAS contains 38,664 representative alternative splicing variants (RASVs) in 11,744 loci, in total. The data is retrievable by various features of AS, which were annotated according to manual annotations, such as by patterns of ASs, consequently invoked alternations in the encoded amino acids and affected protein motifs, GO terms, predicted subcellular localization signals and transmembrane domains. The database also records recently identified very complex patterns of AS, in which two distinct genes seemed to be bridged, nested or degenerated (multiple CDS): in all three cases, completely unrelated proteins are encoded by a single locus. By using AS Viewer, each AS event can be analyzed in the context of full-length cDNAs, enabling the user's empirical understanding of the relation between AS event and the consequent alternations in the encoded amino acid sequences together with various kinds of affected protein motifs. H-DBAS is accessible at http://jbirc.jbic.or.jp/h-dbas/.

Alternative Splicing↗

Cracking the egg: molecular dynamics and evolutionary aspects of the transition from the fully grown oocyte to embryo.

Fully grown oocytes (FGOs) contain all the necessary transcripts to activate molecular pathways underlying the oocyte-to-embryo transition (OET). To elucidate this critical period of development, an extensive survey of the FGO transcriptome was performed by analyzing 19,000 expressed sequence tags of the Mus musculus FGO cDNA library. Expression of 5400 genes and transposable elements is reported. For a majority of genes expressed in mouse FGOs, homologs transcribed in eggs of Xenopus laevis or Ciona intestinalis were found, pinpointing evolutionary conservation of most regulatory cascades underlying the OET in chordates. A large proportion of identified genes belongs to several gene families with oocyte-restricted expression, a likely result of lineage-specific genomic duplications. Gene loss by mutation and expression in female germline of retrotransposed genes specific to M. musculus is documented. These findings indicate rapid diversification of genes involved in female reproduction. Comparison of the FGO and two-cell embryo transcriptomes demarcated the processes important for oogenesis from those involved in OET and identified novel motifs in maternal mRNAs associated with transcript stability. Discovery of oocyte-specific eukaryotic translation initiation factor 4E distinguishes a novel system of translational regulation. These results implicate conserved pathways underlying transition from oogenesis to initiation of development and illustrate how genes acquire and lose reproductive functions during evolution, a potential mechanism for reproductive isolation.

Amino Acid Sequence↗

Establishment of the epithelial-specific transcriptome of normal and malignant human breast cells based on MPSS and array expression data.

INTRODUCTION: Diverse microarray and sequencing technologies have been widely used to characterise the molecular changes in malignant epithelial cells in breast cancers. Such gene expression studies to identify markers and targets in tumour cells are, however, compromised by the cellular heterogeneity of solid breast tumours and by the lack of appropriate counterparts representing normal breast epithelial cells. METHODS: Malignant neoplastic epithelial cells from primary breast cancers and luminal and myoepithelial cells isolated from normal human breast tissue were isolated by immunomagnetic separation methods. Pools of RNA from highly enriched preparations of these cell types were subjected to expression profiling using massively parallel signature sequencing (MPSS) and four different genome wide microarray platforms. Functional related transcripts of the differential tumour epithelial transcriptome were used for gene set enrichment analysis to identify enrichment of luminal and myoepithelial type genes. Clinical pathological validation of a small number of genes was performed on tissue microarrays. RESULTS: MPSS identified 6,553 differentially expressed genes between the pool of normal luminal cells and that of primary tumours substantially enriched for epithelial cells, of which 98% were represented and 60% were confirmed by microarray profiling. Significant expression level changes between these two samples detected only by microarray technology were shown by 4,149 transcripts, resulting in a combined differential tumour epithelial transcriptome of 8,051 genes. Microarray gene signatures identified a comprehensive list of 907 and 955 transcripts whose expression differed between luminal epithelial cells and myoepithelial cells, respectively. Functional annotation and gene set enrichment analysis highlighted a group of genes related to skeletal development that were associated with the myoepithelial/basal cells and upregulated in the tumour sample. One of the most highly overexpressed genes in this category, that encoding periostin, was analysed immunohistochemically on breast cancer tissue microarrays and its expression in neoplastic cells correlated with poor outcome in a cohort of poor prognosis estrogen receptor-positive tumours. CONCLUSION: Using highly enriched cell populations in combination with multiplatform gene expression profiling studies, a comprehensive analysis of molecular changes between the normal and malignant breast tissue was established. This study provides a basis for the identification of novel and potentially important targets for diagnosis, prognosis and therapy in breast cancer.

Biomarkers, Tumor↗

REACTOR: REgulon Activity analysis and Comparison Tool for single-cell transcriptOmics Research.

SUMMARY: We introduce REACTOR, a computational tool designed to detect differential activity of transcriptional regulators and their target genes (regulons) in single-cell RNA-sequencing data. It expands the currently available framework for regulon analysis by introducing a robust statistical test to detect differential regulon activity between conditions, such as disease versus control, with multiple replicates. By contrasting different conditions, REACTOR enables identification of key condition- and cell type-specific regulons. To demonstrate the use of REACTOR, we illustrate its performance in a publicly available COVID-19 dataset. AVAILABILITY: REACTOR R-package together with an implementation vignette are available at https://www.github.com/elolab/REACTOR.

Regulon↗

Construction of representative transcript and protein sets of human, mouse, and rat as a platform for their transcriptome and proteome analysis.

The number of mammalian transcripts identified by full-length cDNA projects and genome sequencing projects is increasing remarkably. Clustering them into a strictly nonredundant and comprehensive set provides a platform for functional analysis of the transcriptome and proteome, but the quality of the clustering and predictive usefulness have previously required manual curation to identify truncated transcripts and inappropriate clustering of closely related sequences. A Representative Transcript and Protein Sets (RTPS) pipeline was previously designed to identify the nonredundant and comprehensive set of mouse transcripts based on clustering of a large mouse full-length cDNA set (FANTOM2). Here we propose an alternative method that is more robust, requires less manual curation, and is applicable to other organisms in addition to mouse. RTPSs of human, mouse, and rat have been produced by this method and used for validation. Their comprehensiveness and quality are discussed by comparison with other clustering approaches. The RTPSs are available at .

Animals↗

Multi-omics approaches in idiopathic pulmonary fibrosis: from molecular mechanisms to therapeutic targets and precision medicine.

Idiopathic pulmonary fibrosis (IPF) is a progressive interstitial lung disease with limited therapeutic options and marked molecular heterogeneity. Despite available antifibrotic therapies, disease progression remains poorly predictable, highlighting the need for improved mechanistic understanding and therapeutic targeting. This review summarizes recent advances in multi-omics research to elucidate the molecular mechanisms underlying IPF and to identify potential biomarkers and pharmacological targets. Multi-omics studies, including genomics, epigenomics, transcriptomics, proteomics, metabolomics, microbiome profiling, and single-cell sequencing, have revealed key pathogenic mechanisms in IPF. Genetic susceptibility factors such as MUC5B promoter variants and telomere-related genes contribute to disease risk. Epigenetic regulation, including DNA methylation, histone modifications, and non-coding RNAs, plays a central role in fibrotic remodeling. Transcriptomic and proteomic analyses have identified dysregulated signaling pathways, including TGF-β, mTOR, cellular senescence, and extracellular matrix remodeling. Metabolomic alterations indicate disrupted lipid and amino acid metabolism. Importantly, integration of multi-omics datasets enables the identification of molecular endotypes, candidate biomarkers, and potential therapeutic targets. However, challenges including data integration, tissue heterogeneity, limited cohort size, and the need for functional validation remain important barriers to clinical translation. Continued development of multi-omics approaches may facilitate more accurate disease classification and support the development of personalized therapeutic strategies for IPF.

biomarkers↗

LINNAEUS: Simultaneous Single-Cell Lineage Tracing and Cell Type Identification.

A key goal of biology is to understand the origin of the many cell types that can be observed during diverse processes such as development, regeneration, and disease. Single-cell RNA-sequencing (scRNA-seq) is commonly used to identify cell types in a tissue or organ. However, organizing the resulting taxonomy of cell types into lineage trees to understand the origins of cell states and relationships between cells remains challenging. Here we present LINNAEUS (Spanjaard et al, Nat Biotechnol 36:469-473. https://doi.org/10.1038/nbt.4124 , 2018; Hu et al, Nat Genet 54:1227-1237. https://doi.org/10.1038/s41588-022-01129-5 , 2022) (LINeage tracing by Nuclease-Activated Editing of Ubiquitous Sequences)-a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA-seq with computational analysis of lineage barcodes, generated by genome editing of transgenic reporter genes, LINNAEUS can be used to reconstruct organism-wide single-cell lineage trees. LINNAEUS provides a systematic approach for tracing the origin of novel cell types, or known cell types under different conditions.

Single-Cell Analysis↗

ForestTreeDB: a database dedicated to the mining of tree transcriptomes.

ForestTreeDB is intended as a resource that centralizes large-scale expressed sequence tag (EST) sequencing results from several tree species (http://foresttree.org/ftdb). It currently encompasses 344,878 quality sequences from 68 libraries, from diverse organs of conifer and hybrid poplar trees. It utilizes the Nimbus data model to provide a hosting system for multiple projects, and uses object-relational mapping APIs in Java and Perl for data accesses within an Oracle database designed to be scalable, maintainable and extendable. Transcriptome builds or unigene sets occupy the focal point of the system. Several of the five current species-specific unigenes were used to design microarrays and SNP resources. The ForestTreeDB web application provides the means for multiple combination database queries. It presents the user with a list of discrete queries to retrieve and download large EST datasets or sequences from precompiled unigene assemblies. Functional annotation assignment is not trivial in conifers which are distantly related to angiosperm model plants. Optimal annotations are achieved through database queries that integrate results from several procedures based open-source tools. ForestTreeDB aims to facilitate sequence mining of coherent annotations in multiple species to support comparative genomic approaches. We plan to continuously enrich ForestTreeDB with other resources through collaborations with other genomic projects.

Databases, Nucleic Acid↗

Expressed sequence tags from loblolly pine embryos reveal similarities with angiosperm embryogenesis.

The process of embryogenesis in gymnosperms differs in significant ways from the more widely studied process in angiosperms. To further our understanding of embryogenesis in gymnosperms, we have generated Expressed Sequence Tags (ESTs) from four cDNA libraries constructed from un-normalized, normalized, and subtracted RNA populations of zygotic and somatic embryos of loblolly pine (Pinus taeda L.). A total of 68,721 ESTs were generated from 68,131 cDNA clones. Following clustering and assembly, these sequences collapsed into 5,274 contigs and 6,880 singleton sequences for a total of 12,154 non-redundant sequences. Searches of a non-identical amino acid database revealed a putative homolog for 9,189 sequences, leaving 2,965 sequences with no known function. More extensive searches of additional plant sequence data sets revealed a putative homolog for all but 1,388 (11.4%) of the sequences. Using gene ontologies, a known function could be assigned for 5,495 of the 12,154 total non-redundant sequences with 13,633 associations in total assigned. When compared to approximately 72,000 sequences in a collated P. taeda transcript assembly derived from >245,000 ESTs derived from root, xylem, stem, needles, pollen cone, and shoot ESTs, 3,458 (28.5%) of the non-redundant embryo sequences were unique and thereby provide a valuable addition to development of a complete loblolly pine transcriptome. To assess similarities between angiosperm and gymnosperm embryo development, we examined our EST collection for putative homologs of angiosperm genes implicated in embryogenesis. Out of 108 angiosperm embryogenesis-related genes, homologs were present for 83 of these genes suggesting that pine contains similar genes for embryogenesis and that our RNA sampling methods were successful. We also identified sequences from the pine embryo transcriptome that have no known function and may contribute to the programming of gene expression and embryo development.

Amino Acid Sequence↗

Beyond the gene: isoform diversity as a key contributor to human brain disorders.

The human brain exhibits exceptional transcriptomic complexity, with alternative splicing, promoter usage, and polyadenylation generating extensive transcript-isoform diversity. Isoform dysregulation is increasingly implicated in neurodevelopmental and psychiatric disorders (NPDs), yet the landscape, function, and genetic regulation of brain isoforms remain poorly understood due to limitations of short-read RNA sequencing. Advances in long-read sequencing (LR-seq) enable scalable full-length transcriptome profiling with single-cell and spatial resolution across developmental stages. Here, we review recent progress in isoform discovery, quantification, functional annotation, and genetic regulation, highlighting emerging links to human neurodevelopment and disease. LR-seq studies have uncovered tens of thousands of previously unannotated brain isoforms, with neuronal maturation characterized by increased exon inclusion and progressive 3' untranslated region (3' UTR) lengthening. Isoform-resolved genetic mapping outperforms gene-level analyses for NPD gene discovery and mechanistic interpretation. We argue that a shift from gene-centric to isoform-centric frameworks is essential to fully capture regulatory complexity in human neurogenetics. Together, these advances establish isoform diversity as a fundamental yet underappreciated axis of brain gene regulation and a key entry point for dissecting NPD biology.

Humans↗

Deletion in a (T)8 microsatellite abrogates expression regulation by 3'-UTR.

A high level of genetic instability might cause mutations to accumulate in tumours. Microsatellite instability (MSI), due to defects of the DNA mismatch repair system, affects in particular repeat sequences (microsatellites) scattered throughout the genome. By scanning transcriptome databases, we found that microsatellites in the human genome are less numerous in coding DNA than in the 3'-untranslated region (UTR), known to mediate control of gene expression. By mutation analysis, we identified a 1 bp deletion in a (T)(8) microsatellite embedded in the 1801 nucleotide long 3'-UTR of CEACAM1 gene, thought to be involved in tumour onset and progression. By Lentiviral Vector- mediated gene transfer, we showed that the wild-type but not the mutated CEACAM1 3'-UTR greatly decreased transgene expression at both mRNA and protein level. Messenger RNA abundance was fully regulated by the most 3' region of CEACAM1 3'-UTR. This region includes the (T)(8) microsatellite but not any known classified regulatory element. These data show that CEACAM1 3'-UTR contains non-canonical elements contributing to mRNA regulation, among which a short repeat sequence could play a critical regulatory function. This suggests that, in cancer cells, a single mutation in a 3'-UTR short microsatellite might strongly affect gene expression.

3' Untranslated Regions↗

Identification and analysis of chromodomain-containing proteins encoded in the mouse transcriptome.

The chromodomain is 40-50 amino acids in length and is conserved in a wide range of chromatic and regulatory proteins involved in chromatin remodeling. Chromodomain-containing proteins can be classified into families based on their broader characteristics, in particular the presence of other types of domains, and which correlate with different subclasses of the chromodomains themselves. Hidden Markov model (HMM)-generated profiles of different subclasses of chromodomains were used here to identify sequences encoding chromodomain-containing proteins in the mouse transcriptome and genome. A total of 36 different loci encoding proteins containing chromodomains, including 17 novel loci, were identified. Six of these loci (including three apparent pseudogenes, a novel HP1 ortholog, and two novel Msl-3 transcription factor-like proteins) are not present in the human genome, whereas the human genome contains four loci (two CDY orthologs and two apparent CDY pseudogenes) that are not present in mouse. A number of these loci exhibit alternative splicing to produce different isoforms, including 43 novel variants, some of which lack the chromodomain. The likely functions of these proteins are discussed in relation to the known functions of other chromodomain-containing proteins within the same family.

Acetyltransferases↗

Molting in Pancrustacea Is Characterized by Both Deeply Conserved and Recently Evolved Gene Modules.

Arthropods such as insects and crustaceans, which together form the monophyletic group Pancrustacea, possess a rigid chitinous exoskeleton that must be periodically shed through molting to allow growth and morphological change. Although molting is a deeply conserved developmental process across Arthropoda, our understanding of its molecular mechanisms is still largely derived from insect model species. Lineage-specific innovations and losses of molting-related genes raise fundamental questions about the extent of its conservation outside noninsect arthropods. Here, we investigate the evolutionary conservation of molting gene expression across five representative pancrustacean species using publicly available transcriptomic datasets. Changes in gene expression during molting are characterized by both deeply conserved and lineage-specific gene modules. Temporal gene expression analyses reveal that these lineage-specific signatures are not uniformly distributed across the molting process: the middle transitional phase is more lineage-specific, thereby exhibiting an inverse hourglass pattern. This is likely due to life-history-specific processes, development of the cuticle, and specialized structures of the exoskeleton. Overall, this study provides evidence for both the evolutionary conservation and divergence of this key postembryonic developmental process and highlights the modular architecture of the molting program.

Animals↗