Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Massively parallel sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Assessment of whole genome amplification-induced bias through high-throughput, massively parallel whole genome sequencing.

BACKGROUND: Whole genome amplification is an increasingly common technique through which minute amounts of DNA can be multiplied to generate quantities suitable for genetic testing and analysis. Questions of amplification-induced error and template bias generated by these methods have previously been addressed through either small scale (SNPs) or large scale (CGH array, FISH) methodologies. Here we utilized whole genome sequencing to assess amplification-induced bias in both coding and non-coding regions of two bacterial genomes. Halobacterium species NRC-1 DNA and Campylobacter jejuni were amplified by several common, commercially available protocols: multiple displacement amplification, primer extension pre-amplification and degenerate oligonucleotide primed PCR. The amplification-induced bias of each method was assessed by sequencing both genomes in their entirety using the 454 Sequencing System technology and comparing the results with those obtained from unamplified controls. RESULTS: All amplification methodologies induced statistically significant bias relative to the unamplified control. For the Halobacterium species NRC-1 genome, assessed at 100 base resolution, the D-statistics from GenomiPhi-amplified material were 119 times greater than those from unamplified material, 164.0 times greater for Repli-G, 165.0 times greater for PEP-PCR and 252.0 times greater than the unamplified controls for DOP-PCR. For Campylobacter jejuni, also analyzed at 100 base resolution, the D-statistics from GenomiPhi-amplified material were 15 times greater than those from unamplified material, 19.8 times greater for Repli-G, 61.8 times greater for PEP-PCR and 220.5 times greater than the unamplified controls for DOP-PCR. CONCLUSION: Of the amplification methodologies examined in this paper, the multiple displacement amplification products generated the least bias, and produced significantly higher yields of amplified DNA.

Bias↗

Signatures from tissue-specific MPSS libraries identify transcripts preferentially expressed in the mouse inner ear.

Specialization in cell function and morphology is influenced by the differential expression of mRNAs, many of which are expressed at low abundance and restricted to certain cell types. Detecting such transcripts in cDNA libraries may require sequencing millions of clones. Massively parallel signature sequencing (MPSS) is well suited to identifying transcripts that are expressed in discrete cell types and in low abundance. We have made MPSS libraries from microdissections of three inner ear tissues. By comparing these MPSS libraries to those of 87 other tissues included in the Mouse Reference Transcriptome online resource, we have identified genes that are highly enriched in, or specific to, the inner ear. We show by RT-PCR and in situ hybridization that signatures unique to the inner ear libraries identify transcripts with highly specific cell-type localizations. These transcripts serve to illustrate the utility of a resource that is available to the research community. Utilization of these resources will increase the number of known transcription units and expand our knowledge of the tissue-specific regulation of the transcriptome.

Animals↗

A transcriptomic and proteomic characterization of the Arabidopsis mitochondrial protein import apparatus and its response to mitochondrial dysfunction.

Mitochondria import hundreds of cytosolically synthesized proteins via the mitochondrial protein import apparatus. Expression analysis in various organs of 19 components of the Arabidopsis mitochondrial protein import apparatus encoded by 31 genes showed that although many were present in small multigene families, often only one member was prominently expressed. This was supported by comparison of real-time reverse transcriptase-polymerase chain reaction and microarray experimental data with expressed sequence tag numbers and massive parallel signature sequence data. Mass spectrometric analysis of purified mitochondria identified 17 import components, their mitochondrial sub-compartment, and verified the presence of TIM8, TIM13, TIM17, TIM23, TIM44, TIM50, and METAXIN proteins for the first time, to our knowledge. Mass spectrometry-detected isoforms correlated with the most abundant gene transcript measured by expression data. Treatment of Arabidopsis cell culture with mitochondrial electron transport chain inhibitors rotenone and antimycin A resulted in a significant increase in transcript levels of import components, with a greater increase observed for the minor isoforms. The increase was observed 12 h after treatment, indicating that it was likely a secondary response. Microarray analysis of rotenone-treated cells indicated the up-regulation of gene sets involved in mitochondrial chaperone activity, protein degradation, respiratory chain assembly, and division. The rate of protein import into isolated mitochondria from rotenone-treated cells was halved, even though rotenone had no direct effect on protein import when added to mitochondria isolated from untreated cells. These findings suggest that transcription of import component genes is induced when mitochondrial function is limited and that minor gene isoforms display a greater response than the predominant isoforms.

Arabidopsis↗

Sequence pattern matching on a massively parallel computer.

A method is described for finding all occurrences of a sequence pattern within a database of molecular sequences. Implementation of this on a massively parallel computer allows the user to perform very fast database searches using complex patterns. In particular, the software supports approximate pattern matching with score thresholds for either the entire pattern or specified elements thereof. Matches to individual elements can be linked by variable length gaps within user-specified limits.

Amino Acid Sequence↗

Genomic approaches to early embryogenesis and stem cell biology.

Large-scale systematic gene expression analyses of early embryos and stem cells provide useful information to identify genes expressed differentially or uniquely in these cells. We review the current status of various approaches applied to preimplantation embryos and stem cells: expressed sequence tag, serial analysis of gene expression, differential display, massively parallel signature sequencing, DNA microarray (DNA chip) analysis, and chromatin-immunoprecipitation microarrays. We also discuss the biological questions that can only be addressed by the analysis of global gene expression patterns, such as so-called stemness and developmental potency. As the emphasis now shifts from expression profiling to functional studies, we review the genome-scale functional studies of genes: expression cloning, gene trapping, RNA interference, and gene disruptions. Finally, we discuss the future clinical application of such methodologies.

Animals↗

Large-scale analysis of neural stem cells and progenitor cells.

The past few years have seen remarkable progress in our understanding of stem cell biology. The wealth of genomic data and the multiplicity of sources have enabled researchers to begin to profile stem cells in detail. In this paper we describe the biological and technical controls necessary to obtain reliable data and the relative merits of various large-scale analytical techniques including microarray, expressed sequence tag enumeration, serial analysis of gene expression and massively parallel signature sequencing. We suggest that while much has been learned, additional information remains to be gleaned by meta-analysis of existing data.

Animals↗

Robust analysis of 5'-transcript ends (5'-RATE): a novel technique for transcriptome analysis and genome annotation.

Complicated cloning procedures and the high cost of sequencing have inhibited the wide application of serial analysis of gene expression and massively parallel signature sequencing for genome-wide transcriptome profiling of complex genomes. Here we describe a new method called robust analysis of 5'-transcript ends (5'-RATE) for rapid and cost-effective isolation of long 5' transcript ends (approximately 80 bp). It consists of three major steps including 5'-oligocapping of mRNA, NlaIII tag and ditag generation, and pyrosequencing of NlaIII tags. Complicated steps, such as purification and cloning of concatemers, colony picking and plasmid DNA purification, are eliminated and the conventional Sanger sequencing method is replaced with the newly developed pyrosequencing method. Sequence analysis of a maize 5'-RATE library revealed complex alternative transcription start sites and a 5' poly(A) tail in maize transcripts. Our results demonstrate that 5'-RATE is a simple, fast and cost-effective method for transcriptome analysis and genome annotation of complex genomes.

5' Untranslated Regions↗

The use of MPSS for whole-genome transcriptional analysis in Arabidopsis.

We have generated 36,991,173 17-base sequence "signatures" representing transcripts from the model plant Arabidopsis. These data were derived by massively parallel signature sequencing (MPSS) from 14 libraries and comprised 268,132 distinct sequences. Comparable data were also obtained with 20-base signatures. We developed a method for handling these data and for comparing these signatures to the annotated Arabidopsis genome. As part of this procedure, 858,019 potential or "genomic" signatures were extracted from the Arabidopsis genome and classified based on the position and orientation of the signatures relative to annotated genes. A comparison of genomic and expressed signatures matched 67,735 signatures predicted to be derived from distinct transcripts and expressed at significant levels. Expressed signatures were derived from the sense strand of at least 19,088 of 29,084 annotated genes. A comparison of the genomic and expression signatures demonstrated that approximately 7.7% of genomic signatures were underrepresented in the expression data. These genomic signatures contained one of 20 four-base words that were consistently associated with reduced MPSS abundances. More than 89% of the sum of the expressed signature abundances matched the Arabidopsis genome, and many of the unmatched signatures found in high abundances were predicted to match to previously uncharacterized transcripts.

Arabidopsis↗

Conserved subgroups and developmental regulation in the monocot rop gene family.

Rop small GTPases are plant-specific signaling proteins with roles in pollen and vegetative cell growth, abscisic acid signal transduction, stress responses, and pathogen resistance. We have characterized the rop family in the monocots maize (Zea mays) and rice (Oryza sativa). The maize genome contains at least nine expressed rops, and the fully sequenced rice genome has seven. Based on phylogenetic analyses of all available Rops, the family can be subdivided into four groups that predate the divergence of monocots and dicots; at least three have been maintained in both lineages. However, the Rop family has evolved differently in the two lineages, with each exhibiting apparent expansion in different groups. These analyses, together with genetic mapping and identification of conserved non-coding sequences, predict orthology for specific rice and maize rops. We also identified consensus protein sequence elements specific to each Rop group. A survey of ROP-mRNA expression in maize, based on multiplex reverse transcriptase-polymerase chain reaction and a massively parallel signature sequencing database, showed significant spatial and temporal overlap of the nine transcripts, with high levels of all nine in tissues in which cells are actively dividing and expanding. However, only a subset of rops was highly expressed in mature leaves and pollen. Intriguingly, the grouping of maize rops based on hierarchical clustering of expression profiles was remarkably similar to that obtained by phylogenetic analysis. We hypothesize that the Rop groups represent classes with distinct functions, which are specified by the unique protein sequence elements in each group and by their distinct expression patterns.

Amino Acid Sequence↗

Handling calcium signaling: Arabidopsis CaMs and CMLs.

The Arabidopsis genome harbors seven calmodulin (CAM) and 50 CAM-like (CML) genes that encode potential calcium sensors. The CAMs encode only four protein isoforms. Selective pressure to maintain multiple CAMs indicates nonredundancy. Sequence divergence, even in the EF hand calcium-binding motif, exists among the CMLs and, therefore, divergent functions are likely to have evolved. Expression data recently available from Massively Parallel Signature Sequencing and Genevestigator compilation of microarrays are reviewed. The seven Arabidopsis CAMs are highly and relatively uniformly expressed. Differential expression is evident among the distinct CMLs over developmental stages, in various organs and in response to many different stimuli. In spite of the potential importance in mediating plant calcium signaling, the physiological functions of the Arabidopsis CaMs and CMLs remain largely unknown.

Amino Acid Sequence↗

Combining expression and comparative evolutionary analysis. The COBRA gene family.

Plant cell shape is achieved through a combination of oriented cell division and cell expansion and is defined by the cell wall. One of the genes identified to influence cell expansion in the Arabidopsis (Arabidopsis thaliana) root is the COBRA (COB) gene that belongs to a multigene family. Three members of the AtCOB gene family have been shown to play a role in specific types of cell expansion or cell wall biosynthesis. Functional orthologs of one of these genes have been identified in maize (Zea mays) and rice (Oryza sativa; Schindelman et al., 2001; Li et al., 2003; Brown et al., 2005; Persson et al., 2005; Ching et al., 2006; Jones et al., 2006). We present the maize counterpart of the COB gene family and the COB gene superfamily phylogeny. Most of the genes belong to a family with two main clades as previously identified by analysis of the Arabidopsis family alone. Within these clades, however, clear differences between monocot and eudicot family members exist, and these are analyzed in the context of Type I and Type II cell walls in eudicots and monocots. In addition to changes at the sequence level, gene regulation of this family in a eudicot, Arabidopsis, and a monocot, maize, is also characterized. Gene expression is analyzed in a multivariate approach, using data from a number of sources, including massively parallel signature sequencing libraries, transcriptional reporter fusions, and microarray data. This analysis has revealed that the expression of Arabidopsis and maize COB gene family members is highly developmentally and spatially regulated at the tissue and cell type-specific level, that gene superfamily members show overlapping and unique expression patterns, and that only a subset of gene superfamily members act in response to environmental stimuli. Regulation of expression of the Arabidopsis COB gene family members has highly diversified in comparison to that of the maize COB gene superfamily members. We also identify BRITTLE STALK 2-LIKE 3 as a putative ortholog of AtCOB.

Amino Acid Sequence↗

Gene expression analysis by signature pyrosequencing.

We describe a novel method for transcript profiling based on high-throughput parallel sequencing of signature tags using a non-gel-based microtiter plate format. The method relies on the identification of cDNA clones by pyrosequencing of the region corresponding to the 3'-end of the mRNA preceding the poly(A) tail. Simultaneously, the method can be used for gene discovery, since tags corresponding to unknown genes can be further characterized by extended sequencing. The protocol was validated using a model system for human atherosclerosis. Two 3'-tagged cDNA libraries, representing macrophages and foam cells, which are key components in the development of atherosclerotic plaques, were constructed using a solid phase approach. The libraries were analyzed by pyrosequencing, giving on average 25 bases. As a control, conventional expressed sequence tag (EST) sequencing using slab gel electrophoresis was performed. Homology searches were used to identify the genes corresponding to each tag. Comparisons with EST sequencing showed identical, unique matches in the majority of cases when the pyrosignature was at least 18 bases. A visualization tool was developed to facilitate differential analysis using a virtual chip format. The analysis resulted in identification of genes with possible relevance for development of atherosclerosis. The use of the method for automated massive parallel signature sequencing is discussed.

Base Sequence↗

Deep and comparative analysis of the mycelium and appressorium transcriptomes of Magnaporthe grisea using MPSS, RL-SAGE, and oligoarray methods.

BACKGROUND: Rice blast, caused by the fungal pathogen Magnaporthe grisea, is a devastating disease causing tremendous yield loss in rice production. The public availability of the complete genome sequence of M. grisea provides ample opportunities to understand the molecular mechanism of its pathogenesis on rice plants at the transcriptome level. To identify all the expressed genes encoded in the fungal genome, we have analyzed the mycelium and appressorium transcriptomes using massively parallel signature sequencing (MPSS), robust-long serial analysis of gene expression (RL-SAGE) and oligoarray methods. RESULTS: The MPSS analyses identified 12,531 and 12,927 distinct significant tags from mycelia and appressoria, respectively, while the RL-SAGE analysis identified 16,580 distinct significant tags from the mycelial library. When matching these 12,531 mycelial and 12,927 appressorial significant tags to the annotated CDS, 500 bp upstream and 500 bp downstream of CDS, 6,735 unique genes in mycelia and 7,686 unique genes in appressoria were identified. A total of 7,135 mycelium-specific and 7,531 appressorium-specific significant MPSS tags were identified, which correspond to 2,088 and 1,784 annotated genes, respectively, when matching to the same set of reference sequences. Nearly 85% of the significant MPSS tags from mycelia and appressoria and 65% of the significant tags from the RL-SAGE mycelium library matched to the M. grisea genome. MPSS and RL-SAGE methods supported the expression of more than 9,000 genes, representing over 80% of the predicted genes in M. grisea. About 40% of the MPSS tags and 55% of the RL-SAGE tags represent novel transcripts since they had no matches in the existing M. grisea EST collections. Over 19% of the annotated genes were found to produce both sense and antisense tags in the protein-coding region. The oligoarray analysis identified the expression of 3,793 mycelium-specific and 4,652 appressorium-specific genes. A total of 2,430 mycelial genes and 1,886 appressorial genes were identified by both MPSS and oligoarray. CONCLUSION: The comprehensive and deep transcriptome analysis by MPSS and RL-SAGE methods identified many novel sense and antisense transcripts in the M. grisea genome at two important growth stages. The differentially expressed transcripts that were identified, especially those specifically expressed in appressoria, represent a genomic resource useful for gaining a better understanding of the molecular basis of M. grisea pathogenicity. Further analysis of the novel antisense transcripts will provide new insights into the regulation and function of these genes in fungal growth, development and pathogenesis in the host plants.

DNA, Fungal↗

Sequence biases in large scale gene expression profiling data.

We present the results of a simple, statistical assay that measures the G+C content sensitivity bias of gene expression experiments without the requirement of a duplicate experiment. We analyse five gene expression profiling methods: Affymetrix GeneChip, Long Serial Analysis of Gene Expression (LongSAGE), LongSAGELite, 'Classic' Massively Parallel Signature Sequencing (MPSS) and 'Signature' MPSS. We demonstrate the methods have systematic and random errors leading to a different G+C content sensitivity. The relationship between this experimental error and the G+C content of the probe set or tag that identifies each gene influences whether the gene is detected and, if detected, the level of gene expression measured. LongSAGE has the least bias, while Signature MPSS shows a strong bias to G+C rich tags and Affymetrix data show different bias depending on the data processing method (MAS 5.0, RMA or GC-RMA). The bias in the Affymetrix data primarily impacts genes expressed at lower levels. Despite the larger sampling of the MPSS library, SAGE identifies significantly more genes (60% more RefSeq genes in a single comparison).

Animals↗

Genomic analysis of the 12-oxo-phytodienoic acid reductase gene family of Zea mays.

The 12-oxo-phytodienoic acid reductases (OPRs) are enzymes that catalyze the reduction of double bonds adjacent to an oxo group in alpha,beta-unsaturated aldehydes or ketones. Some of them have very high substrate specificity and are part of the octadecanoid pathway which convert linolenic acid to the phytohormone jasmonic acid (JA). Sequencing and analysis of ESTs and genomic sequences from available private and public databases revealed that the maize genome encodes eight OPR genes. Southern blot analysis and mapping of individual OPR genes to maize chromosomes using oat maize chromosome addition lines provides independent confirmation of this number of OPR genes in maize. A survey of massively parallel signature sequencing (MPSS) assays revealed that transcripts of each OPR gene accumulate differentially in diverse organs of maize plants suggesting distinct biological functions. Similarly, RNA blot analysis revealed that distinct OPR genes are differentially regulated in response to stress hormones, wounding or pathogen infection. ZmOPR1 and/or ZmOPR2 appear to function in defense responses to pathogens because they are transiently induced by salicylic acid (SA), chitooligosaccharides, and by infection with Cochliobolus carbonum, Cochliobolus heterostrophus and Fusarium verticillioides, but not by wounding. In contrast to these two genes, transcript levels of ZmOPR6 and ZmOPR7 and/or ZmOPR8 are highly induced by wounding or treatments with the wound-associated signaling molecules JA, ethylene and abscisic acid. However, accumulation of ZmOPR6 and ZmOPR7/8 mRNAs was not upregulated by SA treatments or by pathogen infection suggesting specific involvement in the wound-induced defense responses. None of the treatments induced transcripts of ZmOPR3, 4, or 5.

Abscisic Acid↗

Prediction and identification of Arabidopsis thaliana microRNAs and their mRNA targets.

BACKGROUND: A class of eukaryotic non-coding RNAs termed microRNAs (miRNAs) interact with target mRNAs by sequence complementarity to regulate their expression. The low abundance of some miRNAs and their time- and tissue-specific expression patterns make experimental miRNA identification difficult. We present here a computational method for genome-wide prediction of Arabidopsis thaliana microRNAs and their target mRNAs. This method uses characteristic features of known plant miRNAs as criteria to search for miRNAs conserved between Arabidopsis and Oryza sativa. Extensive sequence complementarity between miRNAs and their target mRNAs is used to predict miRNA-regulated Arabidopsis transcripts. RESULTS: Our prediction covered 63% of known Arabidopsis miRNAs and identified 83 new miRNAs. Evidence for the expression of 25 predicted miRNAs came from northern blots, their presence in the Arabidopsis Small RNA Project database, and massively parallel signature sequencing (MPSS) data. Putative targets functionally conserved between Arabidopsis and O. sativa were identified for most newly identified miRNAs. Independent microarray data showed that the expression levels of some mRNA targets anti-correlated with the accumulation pattern of their corresponding regulatory miRNAs. The cleavage of three target mRNAs by miRNA binding was validated in 5' RACE experiments. CONCLUSIONS: We identified new plant miRNAs conserved between Arabidopsis and O. sativa and report a wide range of transcripts as potential miRNA targets. Because MPSS data are generated from polyadenylated RNA molecules, our results suggest that at least some miRNA precursors are polyadenylated at certain stages. The broad range of putative miRNA targets indicates that miRNAs participate in the regulation of a variety of biological processes.

Arabidopsis↗