Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “transcriptome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Bioinformatics of the Paracoccidioides brasiliensis EST Project.

Paracoccidioides brasiliensis is the etiological agent of paracoccidioidomycosis, an endemic mycosis of Latin America. This fungus presents a dimorphic character; it grows as a mycelium at room temperature, but it is isolated as yeast from infected individuals. It is believed that the transition from mycelium to yeast is important for the infective process. The Functional and Differential Genome of Paracoccidioides brasiliensis Project--PbGenome Project was developed to study the infection process by analyzing expressed sequence tags--ESTs, isolated from both mycelial and yeast forms. The PbGenome Project was executed by a consortium that included 70 researchers (professors and students) from two sequencing laboratories of the midwest region of Brazil; this project produced 25,741 ESTs, 19,718 of which with sufficient quality to be analyzed. We describe the computational procedures used to receive process, analyze these ESTs, and help with their functional annotations; we also detail the services that were used for sequence data exploration. Various programs were compared for filtering and grouping the sequences, and they were adapted to a user-friendly interface. This system made the analysis of the differential transcriptome of P. brasiliensis possible.

Brazil↗

Novel algorithm for transcriptome analysis.

A growing body of evidence implicates the oocyte as a key regulator of ovarian folliculogenesis and early embryonic development. We have screened bovine cDNA microarrays (containing expressed sequence tags representing >15,000 unique genes) with Cy3- and Cy5-labeled cDNA derived from bovine oocyte samples collected at two different stages of meiotic maturation (germinal vesicle vs. metaphase II; n = 3 samples per group). Here, we present a novel data analysis approach that uses all available information from above experiments to obtain and index the transcriptome of bovine oocytes and changes in transcriptome composition in response to meiotic maturation. Signal intensities (Fg) for all housekeeping genes were omitted prior to analysis. A local threshold for gene expression was computed as background intensity (Bg) plus 2 times the standard deviation of background and foreground signals. Within each array, data were normalized by the LOWESS procedure. Subsequently, a two-stage mixed model was fitted to remove systematic variations. In the first stage, the response was the LOWESS normalized Fg with treatment as a fixed effect. In stage 2, the residuals from stage 1 were analyzed in a gene-specific model that included treatment group and spots nested within patch and array. A test for the difference between least squares means for the treatment effect was performed. A false discovery rate (FDR) adjustment on the p values for the difference was carried out. This novel algorithm was compared with approaches that ignore the FDR and the threshold described herein and stark differences obtained.

Algorithms↗

GDR (Genome Database for Rosaceae): integrated web resources for Rosaceae genomics and genetics research.

BACKGROUND: Peach is being developed as a model organism for Rosaceae, an economically important family that includes fruits and ornamental plants such as apple, pear, strawberry, cherry, almond and rose. The genomics and genetics data of peach can play a significant role in the gene discovery and the genetic understanding of related species. The effective utilization of these peach resources, however, requires the development of an integrated and centralized database with associated analysis tools. DESCRIPTION: The Genome Database for Rosaceae (GDR) is a curated and integrated web-based relational database. GDR contains comprehensive data of the genetically anchored peach physical map, an annotated peach EST database, Rosaceae maps and markers and all publicly available Rosaceae sequences. Annotations of ESTs include contig assembly, putative function, simple sequence repeats, and anchored position to the peach physical map where applicable. Our integrated map viewer provides graphical interface to the genetic, transcriptome and physical mapping information. ESTs, BACs and markers can be queried by various categories and the search result sites are linked to the integrated map viewer or to the WebFPC physical map sites. In addition to browsing and querying the database, users can compare their sequences with the annotated GDR sequences via a dedicated sequence similarity server running either the BLAST or FASTA algorithm. To demonstrate the utility of the integrated and fully annotated database and analysis tools, we describe a case study where we anchored Rosaceae sequences to the peach physical and genetic map by sequence similarity. CONCLUSIONS: The GDR has been initiated to meet the major deficiency in Rosaceae genomics and genetics research, namely a centralized web database and bioinformatics tools for data storage, analysis and exchange. GDR can be accessed at http://www.genome.clemson.edu/gdr/.

Computer Graphics↗

A panoramic view of gene expression in the human kidney.

To gain a molecular understanding of kidney functions, we established a high-resolution map of gene expression patterns in the human kidney. The glomerulus and seven different nephron segments were isolated by microdissection from fresh tissue specimens, and their transcriptome was characterized by using the serial analysis of gene expression (SAGE) method. More than 400,000 mRNA SAGE tags were sequenced, making it possible to detect in each structure transcripts present at 18 copies per cell with a 95% confidence level. Expression of genes responsible for nephron transport and permeability properties was evidenced through transcripts for 119 solute carriers, 84 channels, 43 ion-transport ATPases, and 12 claudins. Searching for differences between the transcriptomes, we found 998 transcripts greatly varying in abundance from one nephron portion to another. Clustering analysis of these transcripts evidenced different extents of similarity between the nephron portions. Approximately 75% of the differentially distributed transcripts corresponded to cDNAs of known or unknown function that are accurately mapped in the human genome. This systematic large-scale analysis of individual structures of a complex human tissue reveals sets of genes underlying the function of well-defined nephron portions. It also provides quantitative expression data for a variety of genes mutated in hereditary diseases and helps in sorting candidate genes for renal diseases that affect specific portions of the human nephron.

Cluster Analysis↗

An updated catalogue of salivary gland transcripts in the adult female mosquito, Anopheles gambiae.

Salivary glands of blood-sucking arthropods contain a variety of compounds that prevent platelet and clotting functions and modify inflammatory and immunological reactions in the vertebrate host. In mosquitoes, only the adult female takes blood meals, while both sexes take sugar meals. With the recent description of the Anopheles gambiae genome, and with a set of approximately 3000 expressed sequence tags from a salivary gland cDNA library from adult female mosquitoes, we attempted a comprehensive description of the salivary transcriptome of this most important vector of malaria transmission. In addition to many transcripts associated with housekeeping functions, we found an active transposable element, a set of Wolbachia-like proteins, several transcription factors, including Forkhead, Hairy and doublesex, extracellular matrix components and 71 genes coding for putative secreted proteins. Fourteen of these 71 proteins had matching Edman degradation sequences obtained from SDS-PAGE experiments. Overall, 33 transcripts are reported for the first time as coding for salivary proteins. The tissue and sex specificity of these protein-coding transcripts were analyzed by RT-PCR and microarray experiments for insight into their possible function. Notably, two gene products appeared to be differentially spliced in the adult female salivary glands, whereas 13 contigs matched predicted intronic regions and may include additional alternatively spliced transcripts. Most An. gambiae salivary proteins represent novel protein families of unknown function, potentially coding for pharmacologically or microbiologically active substances. Supplemental data to this work can be found at http://www.ncbi.nlm.nih.gov/projects/omes/index.html#Ag2.

Amino Acid Sequence↗

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32 Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59 Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant↗

Correction of sequence-based artifacts in serial analysis of gene expression.

MOTIVATION: Serial Analysis of Gene Expression (SAGE) is a powerful technology for measuring global gene expression, through rapid generation of large numbers of transcript tags. Beyond their intrinsic value in differential gene expression analysis, SAGE tag collections afford abundant information on the size and shape of the sample transcriptome and can accelerate novel gene discovery. These latter SAGE applications are facilitated by the enhanced method of Long SAGE. A characteristic of sequencing-based methods, such as SAGE and Long SAGE is the unavoidable occurrence of artifact sequences resulting from sequencing errors. By virtue of their low-random incidence, such tag errors have minimal impact on differential expression analysis. However, to fully exploit the value of large SAGE tag datasets, it is desirable to account for and correct tag artifacts. RESULTS: We present estimates for occurrences of tag errors, and an efficient error correction algorithm. Error rate estimates are based on a stochastic model that includes the Polymerase chain reaction and sequencing error contributions. The correction algorithm, SAGEScreen, is a multi-step procedure that addresses ditag processing, estimation of empirical error rates from highly abundant tags, grouping of similar-sequence tags and statistical testing of observed counts. We apply SAGEScreen to Long SAGE libraries and compare error rates for several processing scenarios. Results with simulated tag collections indicate that SAGEScreen corrects 78% of recoverable tag errors and reduces the occurrences of singleton tags. AVAILABILITY: The SAGEScreen software is available for academic users from the first author.

Algorithms↗

Natterins, a new class of proteins with kininogenase activity characterized from Thalassophryne nattereri fish venom.

A novel family of proteins with kininogenase activity and unique primary structure was characterized using combined pharmacological, proteomic and transcriptomic approaches of Thalassophryne nattereri fish venom. The major venom components were isolated and submitted to bioassays corresponding to its main effects: nociception and edema. These activities were mostly located in one fraction (MS3), which was further fractionated. The isolated protein, named natterin, was able to induce edema, nociception and cleave human kininogen and kininogen-derived synthetic peptides, releasing kallidin (Lys-bradykinin). The enzymatic digestion was inhibited by kallikrein inhibitors as Trasylol and TKI. Natterin N-terminal peptide showed no similarity with already known proteins present in databanks. Primary structure of natterin was obtained by a transcriptomic approach using a representative cDNA library constructed from T. nattereri venom glands. Several expressed sequence tags (ESTs) were obtained and processed by bioinformatics revealing a major group (18%) of related sequences unknown to gene or protein sequence databases. This group included sequences showing the N-terminus of isolated natterin and was named Natterin family. Analysis of this family allowed us to identify five related sequences, which we called natterin 1-4 and P. Natterin 1 and 2 sequences include the N-terminus of the isolated natterin. Furthermore, internal peptides of natterin 1-3 were found in major spots of whole venom submitted to mass spectrometry/2DGE. Similarly to the ESTs, the complete sequences of natterins did not show any significant similarity with already described tissue kallikreins, kininogenases or any proteinase, all being entirely new. These data present a new task for the knowledge of the action of kininogenases and may help in understanding the mechanisms of T. nattereri fish envenoming, which is an important medical problem in North and Northeast of Brazil.

Amino Acid Sequence↗

Plasmodium post-genomics: better the bug you know?

Since the publication of the sequence of the genome of Plasmodium falciparum, the major causative agent of human malaria, many post-genomic studies have been completed. Invaluably, these data can now be analysed comparatively owing to the availability of a significant amount of genome-sequence data from several closely related model species of Plasmodium and accompanying global proteome and transcriptome studies. This review summarizes our current knowledge and how this has already been--and will continue to be--exploited in the search for vaccines and drugs against this most significant infectious disease of the tropics.

Animals↗

Malaria and the red blood cell membrane.

Malaria is the most serious and widespread parasitic disease of humans and is arguably the commonest disease of red blood cells (RBCs). Malaria has exerted a powerful effect on human evolution and selection for resistance has led to the appearance and persistence of a number of inherited diseases. After parasite invasion, RBCs are progressively and dramatically modified. New structures appear inside the RBC and novel parasite proteins are exported to the erythrocyte cytoplasm and membrane skeleton. Radical biochemical, morphological, and rheological alterations manifest as increased membrane rigidity, reduced cell deformability, and greater adhesiveness for the vascular endothelium and other blood cells. Numerous protein-protein interactions between the malaria-parasite and the host RBC are important for many aspects of parasite biology and the pathogenesis of malaria. In addition, there are many other parasite proteins located within the infected red cell and at the membrane skeleton, for which no precise functional roles have yet been elucidated. Sequencing and annotation of the complete genome of Plasmodium falciparum, the production of proteomic and transcriptomic profiles of parasites, and the development of a transfection system for the asexual stage of the parasite are all recent achievements that should advance understanding of the molecular mechanisms that underlie the parasite-induced functional alterations in red cells.

Animals↗

High-throughput alternative splicing quantification by primer extension and matrix-assisted laser desorption/ionization time-of-flight mass spectrometry.

Alternative splicing is a significant contributor to transcriptome diversity, and a high-throughput experimental method to quantitatively assess predictions from expressed sequence tag and microarray analyses may help to answer questions about the extent and functional significance of these variants. Here, we describe a method for high-throughput analysis of known or suspected alternative splicing variants (ASVs) using PCR, primer extension and matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS). Reverse-transcribed mRNA is PCR amplified with primers surrounding the site of alternative splicing, followed by a primer extension reaction designed to target sequence disparities between two or more variants. These primer extension products are assayed on a MALDI-TOF mass spectrometer and analyzed automatically. This method is high-throughput, highly accurate and reproducible, allowing for the verification of the existence of splicing variants in a variety of samples. An example given also demonstrates how this method can eliminate potential pitfalls from ordinary gel electrophoretic analysis of splicing variants where heteroduplexes formed from different variants can produce erroneous results. The new method can be used to create alternative variant profiles for cancer markers, to study complex splicing regulation, or to screen potential splicing therapies.

Actinin↗

MitoScribe single-cell molecular recorder logs graded signaling dynamics into mitochondrial DNA.

Genetically encoded DNA recorders convert transient biological events into stable genomic mutations, offering a means to reconstruct past cellular states. However, current approaches to log historical events by modifying genomic DNA have limited capacity to record the magnitude of biological signals within individual cells. Here, we introduce MitoScribe, a mitochondrial DNA (mtDNA)-based recording platform that uses mtDNA base editors (DdCBEs) to write graded biological signals into mtDNA as neutral, single-nucleotide substitutions at a defined site. Taking advantage of the hundreds to thousands of mitochondrial genome copies per cell, we demonstrate MitoScribe enables reproducible, highly sensitive, non-destructive, durable, and high-throughput measurements of molecular signals, including hypoxia, NF-κB activity, BMP and Wnt signaling. We show multiple modes of operation, including multiplexed recordings of two independent signals, and coincidence detection of temporally overlapping signals. Coupling MitoScribe with single-cell RNA sequencing and mitochondrial transcript enrichment, we further reconstruct signaling dynamics at the single-cell transcriptome level. Applying this approach during the directed differentiation of human induced pluripotent stem cells (iPSCs) toward mesoderm, we show that early heterogeneity in response to a differentiation cue predicts the later cell state. Together, MitoScribe provides a scalable platform for high-resolution molecular recording in complex cellular contexts.

Journal Article↗

Regulation of the Arabidopsis transcriptome by oxidative stress.

Oxidative stress, resulting from an imbalance in the accumulation and removal of reactive oxygen species such as hydrogen peroxide (H(2)O(2)), is a challenge faced by all aerobic organisms. In plants, exposure to various abiotic and biotic stresses results in accumulation of H(2)O(2) and oxidative stress. Increasing evidence indicates that H(2)O(2) functions as a stress signal in plants, mediating adaptive responses to various stresses. To analyze cellular responses to H(2)O(2), we have undertaken a large-scale analysis of the Arabidopsis transcriptome during oxidative stress. Using cDNA microarray technology, we identified 175 non-redundant expressed sequence tags that are regulated by H(2)O(2). Of these, 113 are induced and 62 are repressed by H(2)O(2). A substantial proportion of these expressed sequence tags have predicted functions in cell rescue and defense processes. RNA-blot analyses of selected genes were used to verify the microarray data and extend them to demonstrate that other stresses such as wilting, UV irradiation, and elicitor challenge also induce the expression of many of these genes, both independently of, and, in some cases, via H(2)O(2).

Adaptation, Physiological↗

Wildlife Trade and Genetic Basis of Disease Susceptibility: A Review.

The surge in the trade of wildlife and wildlife products drives several species to extinction while coinciding with the increase in several zoonotic diseases. It is therefore essential to explore the roles of wildlife trade in disease transmission, and how the knowledge of genetics and immunogenetics can help in alleviating the attending challenges. Pathogen-driven selection plays a fundamental role in maintaining immune gene diversity, as individuals with alleles conferring resistance to endemic diseases have higher survival rate. However, anthropogenic disturbances, such as wildlife exploitation, can disrupt these evolutionary processes, leading to reduced genetic diversity and increased disease vulnerability. Advanced genomic tools, such as next-generation sequencing (NGS), whole-genome sequencing (WGS), CRISPR-Cas9 gene editing, genome-wide association studies (GWAS), epigenetics and transcriptomic analysis, can help identify immune gene variations and predict disease susceptibility in both wild and captive populations. Massive research targeting wildlife markets and the interface between the wild and the market players is necessary. It would be interesting to understand dynamics of pathogens and disease susceptibility, through the application of genetics and immunogenetics, thereby enhancing efforts to address the challenges posed by wildlife trade and zoonotic disease emergence.

Animals↗

Gene expression profiling of the rat superior olivary complex using serial analysis of gene expression.

The superior olivary complex (SOC) is an auditory brainstem region that represents a favourable system to study rapid neurotransmission and the maturation of neuronal circuits. Here we performed serial analysis of gene expression (SAGE) on the SOC in 60-day-old Sprague-Dawley rats to identify genes specifically important for its function and to create a transcriptome reference for the subsequent identification of age-related or disease-related changes. Sequencing of 31 035 tags identified 10 473 different transcripts. Fifty-seven per cent of the unique tags with a count greater than four were statistically more highly represented in the SOC than in the hippocampus. Among them were genes encoding proteins involved in energy supply, the glutamate/glutamine shuttle, and myelination. Approximately 80 plasma membrane transporters, receptors, channels, and vesicular transporters were identified, and 25% of them displayed a significantly higher expression level in the SOC than in the hippocampus. Some of the plasma membrane proteins were not previously characterized in the SOC, e.g. the purinergic receptor subunit P2X(6) and the metabotropic GABA receptor Gpr51. Differential gene expression between SOC and hippocampus was confirmed using RNA in situ hybridization or immunohistochemistry. The extensive gene inventory presented here will alleviate the dissection of the molecular mechanisms underlying specific SOC functions and the comparison with other SAGE libraries from brain will ease the identification of promoters to generate region-specific transgenic animals. The analysis will be part of the publicly available database ID-GRAB.

Animals↗

Computational function assignment for potential drug targets: from single genes to cellular systems.

Biomedical science is currently undergoing an epoch-marking transition from its classical phase to the post-genome era. The outstanding success of world-wide genome sequencing efforts, evidenced by the recent publication of the draft of the human genome, together with the completion of several genomes of eukaryotic model organisms and the availability of microbial genome sequences, is opening up data sources of unprecedented scale for drug discovery. Furthermore, the elucidation of genome expression states through transcriptomic and proteomic techniques is playing a crucial role in the characterisation of disease at the molecular level. At the same time, our still very limited knowledge of the biological functions of genes and proteins at different levels of cellular organisation is preventing full exploitation of the available data. This review will discuss current computational techniques for function prediction based on the sequence-structure-function paradigm. Newly emerging approaches aimed at gaining an expanded understanding of function through integration of data from various sources and modelling of complex cellular systems will also be highlighted.

Amino Acid Sequence↗

Unravelling the role of the ToxR-like transcriptional regulator WmpR in the marine antifouling bacterium Pseudoalteromonas tunicata.

The dark-green-pigmented marine bacterium Pseudoalteromonas tunicata produces several target-specific compounds that act against a range of common fouling organisms, including bacteria, fungi, protozoa, invertebrate larvae and algal spores. The ToxR-like regulator WmpR has previously been shown to regulate expression of bioactive compounds, type IV pili and biofilm formation phenotypes which all appear at the onset of stationary phase. In this study a comparison of survival under starvation or stress between the wild-type P. tunicata strain and a wmpR mutant (D2W2) does not suggest a role for WmpR in regulating starvation- and stress-resistant phenotypes such as those that may be required in stationary phase. Both proteomic [2-dimensional PAGE (2D-PAGE)] and transcriptomic (RNA arbitrarily primed PCR) studies were used to discover members of the WmpR regulon. 2D-PAGE identified 11 proteins that were differentially expressed by WmpR. Peptide sequence data were obtained for six of these proteins and identified using the draft P. tunicata genome as being involved in protein synthesis, amino acid transamination and ubiquinone biosynthesis, as well as hypothetical proteins. The transcriptomic analysis identified three genes significantly up-regulated by WmpR, including a TonB-dependent outer-membrane protein, a non-ribosomal peptide synthetase and a hypothetical protein. Under iron-limitation the wild-type showed greater survival than D2W2, indicating the importance of WmpR under these conditions. Results from these studies show that WmpR controls the expression of genes encoding proteins involved in iron acquisition and uptake, amino acid metabolism and ubiquinone biosynthesis in addition to a number of proteins with as yet unknown functions.

Amino Acids↗

Single-cell RNA sequencing defines developmental progression and reproductive transitions of Pneumocystis carinii.

UNLABELLED: Pneumocystis species are host-obligate fungal pathogens that cause severe pneumonia in immunocompromised individuals. Despite their clinical importance, their life cycle remains poorly understood, in part because Pneumocystis depends on the host environment for most nutrients and requires sexual reproduction for survival, which occurs exclusively in vivo. This study presents the first single-cell RNA sequencing (scRNA-seq) atlas of Pneumocystis carinii, generated from isolated organisms recovered from the bronchoalveolar lavage fluid of infected rats to map the life cycle of P. carinii. Transcriptomes from 87,716 cells were analyzed using the 10× Genomics platform, revealing 13 transcriptionally distinct clusters representing key developmental stages, including biosynthetically active trophic forms, mating-competent intermediates, and asci undergoing sporulation. These states were characterized by expression of MAPK signaling components, β-glucan-modifying enzymes, and spore-associated genes, respectively. The scRNA-seq data support previous evidence that these host-obligate fungi undergo sexual reproduction and provide new insights into the gene expression patterns associated with different life cycle phases. Biomarkers associated with ascus formation identified by scRNA-seq were validated by RT-qPCR, showing decreased expression levels in ascus-depleted populations treated with anidulafungin, a drug that halts ascus formation. More broadly, this approach provides a strategy for studying the full life cycles of fungal pathogens that cannot be continuously cultured. IMPORTANCE: Pneumocystis species (spp.) are clinically significant fungal pathogens that cannot be sustainably cultured in vitro due to their host-obligate nature. This longstanding limitation has impeded progress in understanding their life cycle and identifying therapeutic vulnerabilities. Here, we apply scRNA-seq to P. carinii isolated directly from infected rat lungs, generating the first transcriptional map of its developmental progression. Our results define discrete gene expression states associated with trophic growth, mating activation, and ascus formation and provide transcriptional evidence for a structured life cycle, clarifying key developmental transitions and identifying potential regulatory targets for therapeutic intervention. Importantly, this study demonstrates that scRNA-seq can resolve the developmental biology of host-restricted fungal pathogens that cannot be cultured in vitro. This approach offers a generalizable framework for investigating other unculturable or obligate microbial pathogens directly within their native host environments, where traditional experimental tools are limited.

Pneumocystis carinii↗