Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Massively parallel sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Karyotypic stability, genotyping, differentiation, feeder-free maintenance, and gene expression sampling in three human embryonic stem cell lines derived prior to August 9, 2001.

The number of human embryonic stem cell (hESC) lines available to federally funded U.S. researchers is currently limited. Thus, determining their basic characteristics and disseminating these lines is important. In this report, we recovered and expanded the earliest available cryopreserved stocks of the BG01, BG02, and BG03 hESC lines. These cultures exhibited multiple definitive characteristics of undifferentiated cells, including long-term self-renewal, expression of markers of pluripotency, maintenance of a normal karyotype, and differentiation to mesoderm, endoderm, and ectoderm. Each cell line exhibited a unique genotype and human leukocyte antigen (HLA) isotype, confirming that they were isolated independently. BG01, BG02, and BG03 maintained in feederfree conditions demonstrated self-renewal, maintenance of normal karyotype, and gene expression indicative of undifferentiated pluripotent stem cells. A survey of gene expression in BG02 cells using massively parallel signature sequencing generated a digital read-out of transcript abundance and showed that this line was similar to other hESC lines. BG01, BG02, and BG03 hESCs are therefore independent, undifferentiated, and pluripotent lines that can be maintained without accumulation of karyotypic abnormalities.

Alkaline Phosphatase↗

The impact of SNPs on the interpretation of SAGE and MPSS experimental data.

Serial Analysis of Gene Expression (SAGE) and Massively Parallel Signature Sequencing (MPSS) are powerful techniques for gene expression analysis. A crucial step in analyzing SAGE and MPSS data is the assignment of experimentally obtained tags to a known transcript. However, tag to transcript assignment is not a straightforward process since alternative tags for a given transcript can also be experimentally obtained. Here, we have evaluated the impact of Single Nucleotide Polymorphisms (SNPs) on the generation of alternative SAGE and MPSS tags. This was achieved through the construction of a reference database of SNP-associated alternative tags, which has been integrated with SAGE Genie. A total of 2020 SNP-associated alternative tags were catalogued in our reference database and at least one SNP-associated alternative tag was observed for approximately 8.6% of all known human genes. A significant fraction (61.9%) of these alternative tags matched a list of experimentally obtained tags, validating their existence. In addition, the origin of four out of five SNP-associated alternative MPSS tags was experimentally confirmed through the use of the GLGI-MPSS protocol (Generation of Long cDNA fragments for Gene Identification). The availability of our SNP-associated alternative tag database will certainly improve the interpretation of SAGE and MPSS experiments.

DNA Restriction Enzymes↗

Integration of text- and data-mining using ontologies successfully selects disease gene candidates.

Genome-wide techniques such as microarray analysis, Serial Analysis of Gene Expression (SAGE), Massively Parallel Signature Sequencing (MPSS), linkage analysis and association studies are used extensively in the search for genes that cause diseases, and often identify many hundreds of candidate disease genes. Selection of the most probable of these candidate disease genes for further empirical analysis is a significant challenge. Additionally, identifying the genes that cause complex diseases is problematic due to low penetrance of multiple contributing genes. Here, we describe a novel bioinformatic approach that selects candidate disease genes according to their expression profiles. We use the eVOC anatomical ontology to integrate text-mining of biomedical literature and data-mining of available human gene expression data. To demonstrate that our method is successful and widely applicable, we apply it to a database of 417 candidate genes containing 17 known disease genes. We successfully select the known disease gene for 15 out of 17 diseases and reduce the candidate gene set to 63.3% (+/-18.8%) of its original size. This approach facilitates direct association between genomic data describing gene expression and information from biomedical texts describing disease phenotype, and successfully prioritizes candidate genes according to their expression in disease-affected tissues.

Anatomy↗

Analysis of the transcriptome of the protozoan Theileria parva using MPSS reveals that the majority of genes are transcriptionally active in the schizont stage.

Massively parallel signature sequencing (MPSS) was used to analyze the transcriptome of the intracellular protozoan Theileria parva. In total 1,095,000, 20 bp sequences representing 4371 different signatures were generated from T.parva schizonts. Reproducible signatures were identified within 73% of potentially detectable predicted genes and 83% had signatures in at least one MPSS cycle. A predicted leader peptide was detected on 405 expressed genes. The quantitative range of signatures was 4-52,256 transcripts per million (t.p.m.). Rare transcripts (<50 t.p.m.) were detected from 36% of genes. Sequence signatures approximated a lognormal distribution, as in microarray. Transcripts were widely distributed throughout the genome, although only 47% of 138 telomere-associated open reading frames exhibited signatures. Antisense signatures comprised 13.8% of the total, comparable with Plasmodium. Eighty five predicted genes with antisense signatures lacked a sense signature. Antisense transcripts were independently amplified from schizont cDNA and verified by sequencing. The MPSS transcripts per million for seven genes encoding schizont antigens recognized by bovine CD8 T cells varied 1000-fold. There was concordance between transcription and protein expression for heat shock proteins that were very highly expressed according to MPSS and proteomics. The data suggests a low level of baseline transcription from the majority of protein-coding genes.

Animals↗

Plant MPSS databases: signature-based transcriptional resources for analyses of mRNA and small RNA.

MPSS (massively parallel signature sequencing) is a sequencing-based technology that uses a unique method to quantify gene expression level, generating millions of short sequence tags per library. We have created a series of databases for four species (Arabidopsis, rice, grape and Magnaporthe grisea, the rice blast fungus). Our MPSS databases measure the expression level of most genes under defined conditions and provide information about potentially novel transcripts (antisense transcripts, alternative splice isoforms and regulatory intergenic transcripts). A modified version of MPSS has been used to perform deep profiling of small RNAs from Arabidopsis, and we have recently adapted our database to display these data. Interpretation of the small RNA MPSS data is facilitated by the inclusion of extensive repeat data in our genome viewer. All the data and the tools introduced in this article are available at http://mpss.udel.edu.

Arabidopsis↗

The expression profile of microRNAs in mouse embryos.

MicroRNAs (miRNAs), which are non-coding RNAs 18-25 nt in length, regulate a variety of biological processes, including vertebrate development. To identify new species of miRNA and to simultaneously obtain a comprehensive quantitative profile of small RNA expression in mouse embryos, we used the massively parallel signature sequencing technology that potentially identifies virtually all of the small RNAs in a sample. This approach allowed us to detect a total of 390 miRNAs, including 195 known miRNAs covering approximately 80% of previously registered mouse miRNAs as well as 195 new miRNAs, which are so far unknown in mouse. Some of these miRNAs showed temporal expression profiles during prenatal development (E9.5, E10.5 and E11.5). Several miRNAs were positioned in polycistron clusters, including one particular large transcription unit consisting of 16 known and 23 new miRNAs. Our results indicate existence of a significant number of new miRNAs expressed at specific stages of mammalian embryonic development and which were not detected by earlier methods.

Animals↗

Genome-wide in silico mapping of scaffold/matrix attachment regions in Arabidopsis suggests correlation of intragenic scaffold/matrix attachment regions with gene expression.

We carried out a genome-wide prediction of scaffold/matrix attachment regions (S/MARs) in Arabidopsis. Results indicate no uneven distribution on the chromosomal level but a clear underrepresentation of S/MARs inside genes. In cases where S/MARs were predicted within genes, these intragenic S/MARs were preferentially located within the 5'-half, most prominently within introns 1 and 2. Using Arabidopsis whole-genome expression data generated by the massively parallel signature sequencing methodology, we found a negative correlation between S/MAR-containing genes and transcriptional abundance. Expressed sequence tag data correlated the same way with S/MAR-containing genes. Thus, intragenic S/MARs show a negative correlation with transcription level. For various genes it has been shown experimentally that S/MARs can function as transcriptional regulators and that they have an implication in stabilizing expression levels within transgenic plants. On the basis of a genome-wide in silico S/MAR analysis, we found a significant correlation between the presence of intragenic S/MARs and transcriptional down-regulation.

Arabidopsis↗

Transcriptional similarities, dissimilarities, and conservation of cis-elements in duplicated genes of Arabidopsis.

In plants, duplication of individual genes, long chromosomal regions, and complete genomes provides a major source for evolutionary innovation. We investigated two different types of duplications, tandem and segmental duplications, in Arabidopsis for correlation, conservation, and differences of expression characteristics by making use of large genome-wide expression data as measured by the massively parallel signature sequencing method. Our analysis indicates that large fractions of duplicated gene pairs still share transcriptional characteristics. However, our results also indicate that expression divergence occurs frequently between duplicated gene pairs, a process which frequently might be employed for the retention of sequence redundant gene pairs. Preserved overall similarity between promoters of duplicated genes as well as preservation of individual cis-elements within the respective promoters indicates that the process of transcriptional neo- and subfunctionalization is restricted to only a fraction of cis-elements. We show that sequence similarities and shared regulatory properties within duplicated promoters provide a powerful means to undertake large-scale cis-regulatory element identification by applying an intragenomic phylogenetic footprinting approach. Our work lays a foundation for future comparative studies to elucidate the molecular manifestation of regulatory similarities and dissimilarities of duplicated genes.

Arabidopsis↗

Local coexpression domains of two to four genes in the genome of Arabidopsis.

Expression of genes in eukaryotic genomes is known to cluster, but cluster size is generally loosely defined and highly variable. We have here taken a very strict definition of cluster as sets of physically adjacent genes that are highly coexpressed and form so-called local coexpression domains. The Arabidopsis (Arabidopsis thaliana) genome was analyzed for the presence of such local coexpression domains to elucidate its functional characteristics. We used expression data sets that cover different experimental conditions, organs, tissues, and cells from the Massively Parallel Signature Sequencing repository and microarray data (Affymetrix) from a detailed root analysis. With these expression data, we identified 689 and 1,481 local coexpression domains, respectively, consisting of two to four genes with a pairwise Pearson's correlation coefficient larger than 0.7. This number is approximately 1- to 5-fold higher than the numbers expected by chance. A small (5%-10%) yet significant fraction of genes in the Arabidopsis genome is therefore organized into local coexpression domains. These local coexpression domains were distributed over the genome. Genes in such local domains were for the major part not categorized in the same functional category (GOslim). Neither tandemly duplicated genes nor shared promoter sequence nor gene distance explained the occurrence of coexpression of genes in such chromosomal domains. This indicates that other parameters in genes or gene positions are important to establish coexpression in local domains of Arabidopsis chromosomes.

Arabidopsis↗

Ubiquitous and endoplasmic reticulum-located lysophosphatidyl acyltransferase, LPAT2, is essential for female but not male gametophyte development in Arabidopsis.

Lysophosphatidyl acyltransferase (LPAT) is a pivotal enzyme controlling the metabolic flow of lysophosphatidic acid into different phosphatidic acids in diverse tissues. We examined putative LPAT genes in Arabidopsis thaliana and characterized two related genes that encode the cytoplasmic LPAT. LPAT2 is the lone gene that encodes the ubiquitous and endoplasmic reticulum (ER)-located LPAT. It could functionally complement a bacterial mutant with defective LPAT. LPAT2 and 3 synthesized in recombinant bacteria and yeast possessed in vitro enzyme activity higher on 18:1-CoA than on 16:0-CoA. LPAT2 was expressed ubiquitously in diverse tissues as revealed by RT-PCR, profiling with massively parallel signature sequencing, and promoter-driven beta-glucuronidase gene expression. LPAT2 was colocalized with calreticulin in the ER by immunofluorescence microscopy and subcellular fractionation. LPAT3 was expressed predominately but more actively than LPAT2 in pollen. A null allele (lpat2) having a T-DNA inserted into LPAT2 was identified. The heterozygous mutant (LPAT2/lpat2) had minimal altered vegetative phenotype but produced shorter siliques that contained normal seeds and remnants of aborted ovules in a 1:1 ratio. Results from selfing and crossing it with the wild type revealed that lpat2 caused lethality in the female gametophyte but not the male gametophyte, which had the redundant LPAT3. LPAT2-cDNA driven by an LPAT2 promoter functionally complemented lpat2 in transformed heterozygous mutants to produce the lpat2/lpat2 genotype. LPAT3-cDNA driven by the LPAT2 promoter could rescue the lpat2 female gametophytes to allow fertilization to occur but not to full embryo maturation. Two other related genes, putative LPAT4 and 5, were expressed ubiquitously albeit at low levels in diverse organs. When they were expressed in bacteria or yeast, the microbial extract did not contain LPAT activity higher than the endogenous LPAT activity. Whether LPAT4 and 5 encode LPATs remains to be elucidated.

1-Acylglycerophosphocholine O-Acyltransferase↗

Global analysis of the core cell cycle regulators of Arabidopsis identifies novel genes, reveals multiple and highly specific profiles of expression and provides a coherent model for plant cell cycle control.

Arabidopsis has over 80 genes encoding conserved and plant-specific core cell cycle regulators, but in most cases neither their timing of expression in the cell cycle is known nor whether they represent redundant and/or tissue-specific functions. Here we identify novel cell cycle regulators, including new cyclin-dependent kinases related to the mammalian galactosyltransferase-associated protein kinase p58, and new classes of cyclin-like and CDK-like proteins showing strong tissue specificity of expression. We analyse expression of all cell cycle regulators in synchronized Arabidopsis cell cultures using multiple approaches including Affymetrix microarrays, massively parallel signature sequencing and real-time reverse transcriptase polymerase chain reaction, and in plant material using the results of over 320 microarray experiments. These global analyses reveal that most core cell cycle regulators are expressed across almost all tissues and more than 85% are expressed at detectable levels in the cell suspension culture, allowing us to present a unified model of transcriptional regulation of the plant cell cycle. Characteristic patterns of D-cyclin expression in early and late G1 phase, either limited to the re-entry cycle or continuously oscillating, suggest that several CYCD genes with strong oscillatory regulation in late G1 may play the role of cyclin E in plants. Alone amongst the six groups of A and B type cyclins, members of CYCA3 peak in S-phase suggest it is a major component of S-phase kinases, whereas others show a peak in G2/M. 82 genes share this G2/M regulatory pattern, about half being new candidate mitotic genes of previously unknown function.

Arabidopsis↗

It's a small RNA world, after all.

Small RNAs (sRNAs) can regulate transcript and protein abundance. Previously, they have been identified using traditional cloning approaches, which has limited how many could be characterized. Now, the Meyers and Green laboratories have used massively parallel signature sequencing technology to find over 1.5 million sRNAs in Arabidopsis thaliana. These new sRNAs reveal a greater-than-expected potential role for sRNAs in gene regulation, preferential expression or usage of sRNAs in flowers, and the prospect of targeted sRNA-mediated regulation of pseudogenes. In addition, new plant microRNAs have been identified, some of which may be unique to Arabidopsis.

Arabidopsis↗

Current research priorities in chronic fatigue syndrome/myalgic encephalomyelitis: disease mechanisms, a diagnostic test and specific treatments.

Chronic fatigue syndrome (CFS) is an illness characterised by disabling fatigue of at least 6 months duration, which is accompanied by various rheumatological, infectious and neuropsychiatric symptoms. A collaborative study group has been formed to deal with the current areas for development in CFS research--namely, to develop an understanding of the molecular pathogenesis of CFS, to develop a diagnostic test and to develop specific and curative treatments. Various groups have studied the gene expression in peripheral blood of patients with CFS, and from those studies that have been confirmed using polymerase chain reaction (PCR), clearly, the most predominant functional theme is that of immunity and defence. However, we do not yet know the precise gene signature and metabolic pathways involved. Currently, this is being dealt with using a microarray representing 47,000 human genes and variants, massive parallel signature sequencing and real-time PCR. It will be important to ensure that once a gene signature has been identified, it is specific to CFS and does not occur in other diseases and infections. A diagnostic test is being developed using surface-enhanced, laser-desorption and ionisation-time-of-flight mass spectrometry based on a pilot study in which putative biomarkers were identified. Finally, clinical trials are being planned; novel treatments that we believe are important to trial in patients with CFS are interferon-beta and one of the anti-tumour necrosis factor-alpha drugs.

Fatigue Syndrome, Chronic↗

Evidence for the presence of disease-perturbed networks in prostate cancer cells by genomic and proteomic analyses: a systems approach to disease.

Prostate cancer is initially responsive to androgen ablation therapy and progresses to androgen-unresponsive states that are refractory to treatment. The mechanism of this transition is unknown. A systems approach to disease begins with the quantitative delineation of the informational elements (mRNAs and proteins) in various disease states. We employed two recently developed high-throughput technologies, massively parallel signature sequencing (MPSS) and isotope-coded affinity tag, to gain a comprehensive picture of the changes in mRNA levels and more restricted analysis of protein levels, respectively, during the transition from androgen-dependent LNCaP (model for early-stage prostate cancer) to androgen-independent CL1 cells (model for late-stage prostate cancer). We sequenced >5 million MPSS signatures, obtained >142,000 tandem mass spectra, and built comprehensive MPSS and proteomic databases. The integrated mRNA and protein expression data revealed underlying functional differences between androgen-dependent and androgen-independent prostate cancer cells. The high sensitivity of MPSS enabled us to identify virtually all of the expressed transcripts and to quantify the changes in gene expression between these two cell states, including functionally important low-abundance mRNAs, such as those encoding transcription factors and signal transduction molecules. These data enable us to map the differences onto extant physiologic networks, creating perturbation networks that reflect prostate cancer progression. We found 37 BioCarta and 14 Kyoto Encyclopedia of Genes and Genomes pathways that are up-regulated and 23 BioCarta and 22 Kyoto Encyclopedia of Genes and Genomes pathways that are down-regulated in LNCaP cells versus CL1 cells. Our efforts represent a significant step toward a systems approach to understanding prostate cancer progression.

Cell Line, Tumor↗

Global analysis of gene expression: methods, interpretation, and pitfalls.

Over the past 15 years, global analysis of mRNA expression has emerged as a powerful strategy for biological discovery. Using the power of parallel processing, robotics, and computer-based informatics, a number of high-throughput methods have been devised. These include DNA microarrays, serial analysis of gene expression, quantitative RT-PCR, differential-display RT-PCR, and massively parallel signature sequencing. Each of these methods has inherent advantages and disadvantages, often related to expense, technical difficulty, specificity, and reliability. Further, the ability to generate large data sets of gene expression has led to new challenges in bioinformatics. Nonetheless, this technological revolution is transforming disease classification, gene discovery, and our understanding of regulatory gene networks.

Computational Biology↗

Bayesian model accounting for within-class biological variability in Serial Analysis of Gene Expression (SAGE).

BACKGROUND: An important challenge for transcript counting methods such as Serial Analysis of Gene Expression (SAGE), "Digital Northern" or Massively Parallel Signature Sequencing (MPSS), is to carry out statistical analyses that account for the within-class variability, i.e., variability due to the intrinsic biological differences among sampled individuals of the same class, and not only variability due to technical sampling error. RESULTS: We introduce a Bayesian model that accounts for the within-class variability by means of mixture distribution. We show that the previously available approaches of aggregation in pools ("pseudo-libraries") and the Beta-Binomial model, are particular cases of the mixture model. We illustrate our method with a brain tumor vs. normal comparison using SAGE data from public databases. We show examples of tags regarded as differentially expressed with high significance if the within-class variability is ignored, but clearly not so significant if one accounts for it. CONCLUSION: Using available information about biological replicates, one can transform a list of candidate transcripts showing differential expression to a more reliable one. Our method is freely available, under GPL/GNU copyleft, through a user friendly web-based on-line tool or as R language scripts at supplemental web-site.

Astrocytoma↗

MPSS profiling of human embryonic stem cells.

BACKGROUND: Pooled human embryonic stem cells (hESC) cell lines were profiled to obtain a comprehensive list of genes common to undifferentiated human embryonic stem cells. RESULTS: Pooled hESC lines were profiled to obtain a comprehensive list of genes common to human ES cells. Massively parallel signature sequencing (MPSS) of approximately three million signature tags (signatures) identified close to eleven thousand unique transcripts, of which approximately 25% were uncharacterised or novel genes. Expression of previously identified ES cell markers was confirmed and multiple genes not known to be expressed by ES cells were identified by comparing with public SAGE databases, EST libraries and parallel analysis by microarray and RT-PCR. Chromosomal mapping of expressed genes failed to identify major hotspots and confirmed expression of genes that map to the X and Y chromosome. Comparison with published data sets confirmed the validity of the analysis and the depth and power of MPSS. CONCLUSIONS: Overall, our analysis provides a molecular signature of genes expressed by undifferentiated ES cells that can be used to monitor the state of ES cells isolated by different laboratories using independent methods and maintained under differing culture conditions

Cell Differentiation↗

Comparison of the gene expression profile of undifferentiated human embryonic stem cell lines and differentiating embryoid bodies.

BACKGROUND: The identification of molecular pathways of differentiation of embryonic stem cells (hESC) is critical for the development of stem cell based medical therapies. In order to identify biomarkers and potential regulators of the process of differentiation, a high quality microarray containing 16,659 seventy base pair oligonucleotides was used to compare gene expression profiles of undifferentiated hESC lines and differentiating embryoid bodies. RESULTS: Previously identified "stemness" genes in undifferentiated hESC lines showed down modulation in differentiated cells while expression of several genes was induced as cells differentiated. In addition, a subset of 194 genes showed overexpression of greater than > or = 3 folds in human embryoid bodies (hEB). These included 37 novel and 157 known genes. Gene expression was validated by a variety of techniques including another large scale array, reverse transcription polymerase chain reaction, focused cDNA microarrays, massively parallel signature sequencing (MPSS) analysis and immunocytochemisty. Several novel hEB specific expressed sequence tags (ESTs) were mapped to the human genome database and their expression profile characterized. A hierarchical clustering analysis clearly depicted a distinct difference in gene expression profile among undifferentiated and differentiated hESC and confirmed that microarray analysis could readily distinguish them. CONCLUSION: These results present a detailed characterization of a unique set of genes, which can be used to assess the hESC differentiation.

Biomarkers↗