Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Massively parallel sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Whole genome shotgun sequencing of Brassica oleracea and its application to gene discovery and annotation in Arabidopsis.

Through comparative studies of the model organism Arabidopsis thaliana and its close relative Brassica oleracea, we have identified conserved regions that represent potentially functional sequences overlooked by previous Arabidopsis genome annotation methods. A total of 454,274 whole genome shotgun sequences covering 283 Mb (0.44 x) of the estimated 650 Mb Brassica genome were searched against the Arabidopsis genome, and conserved Arabidopsis genome sequences (CAGSs) were identified. Of these 229,735 conserved regions, 167,357 fell within or intersected existing gene models, while 60,378 were located in previously unannotated regions. After removal of sequences matching known proteins, CAGSs that were close to one another were chained together as potentially comprising portions of the same functional unit. This resulted in 27,347 chains of which 15,686 were sufficiently distant from existing gene annotations to be considered a novel conserved unit. Of 192 conserved regions examined, 58 were found to be expressed in our cDNA populations. Rapid amplification of cDNA ends (RACE) was used to obtain potentially full-length transcripts from these 58 regions. The resulting sequences led to the creation of 21 gene models at 17 new Arabidopsis loci and the addition of splice variants or updates to another 19 gene structures. In addition, CAGSs overlapping already annotated genes in Arabidopsis can provide guidance for manual improvement of existing gene models. Published genome-wide expression data based on whole genome tiling arrays and massively parallel signature sequencing were overlaid on the Brassica-Arabidopsis conserved sequences, and 1399 regions of intersection were identified. Collectively our results and these data sets suggest that several thousand new Arabidopsis genes remain to be identified and annotated.

Arabidopsis↗

Divergence of the Dof gene families in poplar, Arabidopsis, and rice suggests multiple modes of gene evolution after duplication.

It is widely accepted that gene duplication is a primary source of genetic novelty. However, the evolutionary fate of duplicated genes remains largely unresolved. The classical Ohno's Duplication-Retention-Non/Neofunctionalization theory, and the recently proposed alternatives such as subfunctionalization or duplication-degeneration-complementation, and subneofunctionalization, each can explain one or more aspects of gene fate after duplication. Duplicated genes are also affected by epigenetic changes. We constructed a phylogenetic tree using Dof (DNA binding with one finger) protein sequences from poplar (Populus trichocarpa) Torr. & Gray ex Brayshaw, Arabidopsis (Arabidopsis thaliana), and rice (Oryza sativa). From the phylogenetic tree, we identified 27 pairs of paralogous Dof genes in the terminal nodes. Analysis of protein motif structure of the Dof paralogs and their ancestors revealed six different gene fates after gene duplication. Differential protein methylation was revealed between a pair of duplicated poplar Dof genes, which have identical motif structure and similar expression pattern, indicating that epigenetics is involved in evolution. Analysis of reverse transcription-PCR, massively parallel signature sequencing, and microarray data revealed that the paralogs differ in expression pattern. Furthermore, analysis of nonsynonymous and synonymous substitution rates indicated that divergence of the duplicated genes was driven by positive selection. About one-half of the motifs in Dof proteins were shared by non-Dof proteins in the three plants species, indicating that motif co-option may be one of the forces driving gene diversification. We provided evidence that the Ohno's Duplication-Retention-Non/Neofunctionalization, subfunctionalization/duplication-degeneration-complementation, and subneofunctionalization hypotheses are complementary with, not alternative to, each other.

Amino Acid Sequence↗

Establishment of the epithelial-specific transcriptome of normal and malignant human breast cells based on MPSS and array expression data.

INTRODUCTION: Diverse microarray and sequencing technologies have been widely used to characterise the molecular changes in malignant epithelial cells in breast cancers. Such gene expression studies to identify markers and targets in tumour cells are, however, compromised by the cellular heterogeneity of solid breast tumours and by the lack of appropriate counterparts representing normal breast epithelial cells. METHODS: Malignant neoplastic epithelial cells from primary breast cancers and luminal and myoepithelial cells isolated from normal human breast tissue were isolated by immunomagnetic separation methods. Pools of RNA from highly enriched preparations of these cell types were subjected to expression profiling using massively parallel signature sequencing (MPSS) and four different genome wide microarray platforms. Functional related transcripts of the differential tumour epithelial transcriptome were used for gene set enrichment analysis to identify enrichment of luminal and myoepithelial type genes. Clinical pathological validation of a small number of genes was performed on tissue microarrays. RESULTS: MPSS identified 6,553 differentially expressed genes between the pool of normal luminal cells and that of primary tumours substantially enriched for epithelial cells, of which 98% were represented and 60% were confirmed by microarray profiling. Significant expression level changes between these two samples detected only by microarray technology were shown by 4,149 transcripts, resulting in a combined differential tumour epithelial transcriptome of 8,051 genes. Microarray gene signatures identified a comprehensive list of 907 and 955 transcripts whose expression differed between luminal epithelial cells and myoepithelial cells, respectively. Functional annotation and gene set enrichment analysis highlighted a group of genes related to skeletal development that were associated with the myoepithelial/basal cells and upregulated in the tumour sample. One of the most highly overexpressed genes in this category, that encoding periostin, was analysed immunohistochemically on breast cancer tissue microarrays and its expression in neoplastic cells correlated with poor outcome in a cohort of poor prognosis estrogen receptor-positive tumours. CONCLUSION: Using highly enriched cell populations in combination with multiplatform gene expression profiling studies, a comprehensive analysis of molecular changes between the normal and malignant breast tissue was established. This study provides a basis for the identification of novel and potentially important targets for diagnosis, prognosis and therapy in breast cancer.

Biomarkers, Tumor↗

Genome-wide prediction and identification of cis-natural antisense transcripts in Arabidopsis thaliana.

BACKGROUND: Natural antisense transcripts (NAT) are a class of endogenous coding or non-protein-coding RNAs with sequence complementarity to other transcripts. Several lines of evidence have shown that cis- and trans-NATs may participate in a broad range of gene regulatory events. Genome-wide identification of cis-NATs in human, mouse and rice has revealed their widespread occurrence in eukaryotes. However, little is known about cis-NATs in the model plant Arabidopsis thaliana. RESULTS: We developed a new computational method to predict and identify cis-encoded NATs in Arabidopsis and found 1,340 potential NAT pairs. The expression of both sense and antisense transcripts of 957 NAT pairs was confirmed using Arabidopsis full-length cDNAs and public massively parallel signature sequencing (MPSS) data. Three known or putative Arabidopsis imprinted genes have cis-antisense transcripts. Sequences and the genomic arrangement of two Arabidopsis NAT pairs are conserved in rice. CONCLUSION: We combined information from full-length cDNAs and Arabidopsis genome annotation in our NAT prediction work and reported cis-NAT pairs that could not otherwise be identified by using one of the two datasets only. Analysis of MPSS data suggested that for most Arabidopsis cis-NAT pairs, there is predominant expression of one of the two transcripts in a tissue-specific manner.

Alternative Splicing↗

AMPDB: the Arabidopsis Mitochondrial Protein Database.

The Arabidopsis Mitochondrial Protein Database is an Internet-accessible relational database containing information on the predicted and experimentally confirmed protein complement of mitochondria from the model plant Arabidopsis thaliana (http://www.ampdb.bcs.uwa.edu.au/). The database was formed using the total non-redundant nuclear and organelle encoded sets of protein sequences and allows relational searching of published proteomic analyses of Arabidopsis mitochondrial samples, a set of predictions from six independent subcellular-targeting prediction programs, and orthology predictions based on pairwise comparison of the Arabidopsis protein set with known yeast and human mitochondrial proteins and with the proteome of Rickettsia. A variety of precomputed physical-biochemical parameters are also searchable as well as a more detailed breakdown of mass spectral data produced from our proteomic analysis of Arabidopsis mitochondria. It contains hyperlinks to other Arabidopsis genomic resources (MIPS, TIGR and TAIR), which provide rapid access to changing gene models as well as hyperlinks to T-DNA insertion resources, Massively Parallel Signature Sequencing (MPSS) and Genome Tiling Array data and a variety of other Arabidopsis online resources. It also incorporates basic analysis tools built into the query structure such as a BLAST facility and tools for protein sequence alignments for convenient analysis of queried results.

Arabidopsis Proteins↗

Genomic and genetic characterization of rice Cen3 reveals extensive transcription and evolutionary implications of a complex centromere.

The centromere is the chromosomal site for assembly of the kinetochore where spindle fibers attach during cell division. In most multicellular eukaryotes, centromeres are composed of long tracts of satellite repeats that are recalcitrant to sequencing and fine-scale genetic mapping. Here, we report the genomic and genetic characterization of the complete centromere of rice (Oryza sativa) chromosome 3. Using a DNA fiber-fluorescence in situ hybridization approach, we demonstrated that the centromere of chromosome 3 (Cen3) contains approximately 441 kb of the centromeric satellite repeat CentO. Cen3 includes an approximately 1,881-kb domain associated with the centromeric histone CENH3. This CENH3-associated chromatin domain is embedded within a 3,113-kb region that lacks genetic recombination. Extensive transcription was detected within the CENH3 binding domain based on comprehensive annotation of protein-coding genes coupled with empirical measurements of mRNA levels using RT-PCR and massively parallel signature sequencing. Genes <10 kb from the CentO satellite array were expressed in several rice tissues and displayed histone modification patterns consistent with euchromatin, suggesting that rice centromeric chromatin accommodates normal gene expression. These results support the hypothesis that centromeres can evolve from gene-containing genomic regions.

Centromere↗

Evolution of DNA sequence nonhomologies among maize inbreds.

Allelic chromosomal regions totaling more than 2.8 Mb and located on maize (Zea mays) chromosomes 1L, 2S, 7L, and 9S have been sequenced and compared over distances of 100 to 350 kb between the two maize inbred lines Mo17 and B73. The alleles contain extended regions of nonhomology. On average, more than 50% of the compared sequence is noncolinear, mainly because of the insertion of large numbers of long terminal repeat (LTR)-retrotransposons. Only 27 LTR-retroelements are shared between alleles, whereas 62 are allele specific. The insertion of LTR-retrotransposons into the maize genome is statistically more recent for nonshared than shared ones. Most surprisingly, more than one-third of the genes (27/72) are absent in one of the inbreds at the loci examined. Such nonshared genes usually appear to be truncated and form clusters in which they are oriented in the same direction. However, the nonshared genome segments are gene-poor, relative to regions shared by both inbreds, with up to 12-fold difference in gene density. By contrast, miniature inverted terminal repeats (MITEs) occur at a similar frequency in the shared and nonshared fractions. Many times, MITES are present in an identical position in both LTRs of a retroelement, indicating that their insertion occurred before the replication of the retroelement in question. Maize ESTs and/or maize massively parallel signature sequencing tags were identified for the majority of the nonshared genes or homologs of them. In contrast with shared genes, which are usually conserved in gene order and location relative to rice (Oryza sativa), nonshared genes violate the maize colinearity with rice. Based on this, insertion by a yet unknown mechanism, rather than deletion events, seems to be the origin of the nonshared genes. The intergenic space between conserved genes is enlarged up to sixfold in maize compared with rice. Frequently, retroelement insertions create a different sequence environment adjacent to conserved genes.

Alleles↗

Nucleoside triphosphate diphosphohydrolase-2 is the ecto-ATPase of type I cells in taste buds.

The presence of one or more calcium-dependent ecto-ATPases (enzymes that hydrolyze extracellular 5'-triphosphates) in mammalian taste buds was first shown histochemically. Recent studies have established that dominant ecto-ATPases consist of enzymes now called nucleoside triphosphate diphosphohydrolases (NTPDases). Massively parallel signature sequencing (MPSS) from murine taste epithelium provided molecular evidence suggesting that NTPDase2 is the most likely member present in mouse taste papillae. Immunocytochemical and enzyme histochemical staining verified the presence of NTPDase2 associated with plasma membranes in a large number of cells within all mouse taste buds. To determine which of the three taste cell types expresses this enzyme, double-label assays were performed with antisera directed against the glial glutamate/aspartate transporter (GLAST), the transduction pathway proteins phospholipase Cbeta2 (PLCbeta2) or the G-protein subunit alpha-gustducin, and serotonin (5HT) as markers of type I, II, and III taste cells, respectively. Analysis of the double-labeled sections indicates that NTPDase2 immunoreactivity is found on cell processes that often envelop other taste cells, reminiscent of type I cells. In agreement with this observation, NTPDase2 was located to the same membrane as GLAST, indicating that this enzyme is present in type I cells. The presence of ecto-ATPase in taste buds likely reflects the importance of ATP as an intercellular signaling molecule in this system.

Adenosine Triphosphatases↗

KLK31P is a novel androgen regulated and transcribed pseudogene of kallikreins that is expressed at lower levels in prostate cancer cells than in normal prostate cells.

BACKGROUND: Fifteen human tissue kallikrein (KLK) genes have been identified as a cluster on chromosome 19. KLK expression is associated with various human diseases including cancers. Noncoding RNAs such as PCA3/DD3 and PCGEM1 have been identified in prostate cancer cells. METHODS: Using massively parallel signature sequencing (MPSS) technology, RT-PCR, and 5' rapid amplification of cDNA ends (RACE), we identified and cloned a novel gene that maps to the KLK locus. RESULTS: We have characterized this gene, named as KLK31P by the HUGO Gene Nomenclature Committee, as an unprocessed KLK pseudogene. It contains five exons, two of which are KLK-derived while the rest are "exonized" interspersed repeats. KLK31P is expressed abundantly in prostate tissues and is androgen regulated. KLK31P is expressed at lower levels in localized and metastatic prostate cancer cells than in normal prostate cells. CONCLUSIONS: KLK31P is a novel androgen regulated and transcribed pseudogene of kallikreins that may play a role in prostate carcinogenesis or maintenance.

Amino Acid Sequence↗

A genome-wide transcription analysis of a fungal riboflavin overproducer.

The production of many fine chemicals such as vitamins and amino acids is carried out in bioreactors using microorganisms. Usually, these strains are developed from wild-type organisms by classical mutation and selection. After several generations of strain improvement, no further enhancement can be achieved. Therefore, metabolic engineering (ME) is a rational approach to optimise such producer organisms beyond this point, or for starting all over from the beginning. Metabolic Engineering involves detailed analysis of the organism's metabolic and genetic properties, leading to the identification of new target genes. The fungal riboflavin overproducer Ashbya gossypii converts vegetable oil to vitamin B2 in a "one-step reaction". The productivity and selectivity of this microorganism have been optimised significantly over the years, first following a classical approach and now a rational one. The improvement is based on our understanding of vitamin B2 metabolism. We have been able to selectively enhance the pathways that are necessary for the formation of riboflavin and to inhibit those leading to unwanted side products. New targets for further improvements of this process have been found using a genome-wide transcript expression analysis; namely massive parallel signature sequencing (MPSS). With this analysis even completely unknown genes can be used for strain improvement.

Bioreactors↗

Gene expression profiling: methodological challenges, results, and prospects for addiction research.

This review describes the current methods used to profile gene expression. These methods include microarrays, spotted arrays, serial analysis of gene expression (SAGE), and massive parallel signature sequencing (MPSS). Methodological and statistical problems in interpreting microarray and spotted array experiments are also discussed. Methods and formats such as minimum information about microarray experiments (MIAME) needed to share gene expression data are described. The last part of the review provides an overview of the application of gene-expression profiling technology to substance abuse research and discusses future directions.

Animals↗

Monitoring genome-wide changes in gene expression in response to endogenous cytokinin reveals targets in Arabidopsis thaliana.

Cytokinins have been implicated in developmental and growth processes in plants including cell division, chloroplast biogenesis, shoot meristem initiation and senescence. The regulation of these processes requires changes in cytokinin-responsive gene expression. Here, we induced the expression of a bacterial isopentenyl transferase gene, IPT, in transgenic Arabidopsis thaliana seedlings to study the regulation of genome-wide gene expression in response to endogenous cytokinin. Using MPSS (massively parallel signature sequencing) we identified 823 and 917 genes that were up- and downregulated, respectively, following 24 h of IPT induction. When comparing the response to cytokinin after 6 and 24 h, we identified different clusters of genes showing a similar course of regulation. Our study provides researchers with the opportunity to rapidly assess whether genes of interest are regulated by cytokinins.

Alkyl and Aryl Transferases↗

New developments in microarray technology.

Microarrays have emerged as indispensable research tools for gene expression profiling and mutation analysis. New classification of cancer subtypes, dissecting the yeast metabolism and large-scale genotyping of human single nucleotide polymorphisms are important results being obtained with this technique. Realizing the microsphere-based massively parallel signature sequencing technique as fluid microarrays, building new types of protein arrays and constructing miniaturized flow-through systems, which can potentially take this technology from the research bench into industrial, clinical and other routine applications, exemplify the intense developments that are now ongoing in this field.

Biotechnology↗

TIR-X and TIR-NBS proteins: two new families related to disease resistance TIR-NBS-LRR proteins encoded in Arabidopsis and other plant genomes.

The Toll/interleukin-1 receptor (TIR) domain is found in one of the two large families of homologues of plant disease resistance proteins (R proteins) in Arabidopsis and other dicotyledonous plants. In addition to these TIR-NBS-LRR (TNL) R proteins, we identified two families of TIR-containing proteins encoded in the Arabidopsis Col-0 genome. The TIR-X (TX) family of proteins lacks both the nucleotide-binding site (NBS) and the leucine rich repeats (LRRs) that are characteristic of the R proteins, while the TIR-NBS (TN) proteins contain much of the NBS, but lack the LRR. In Col-0, the TX family is encoded by 27 genes and three pseudogenes; the TN family is encoded by 20 genes and one pseudogene. Using massively parallel signature sequencing (MPSS), expression was detected at low levels for approximately 85% of the TN-encoding genes. Expression was detected for only approximately 40% of the TX-encoding genes, again at low levels. Physical map data and phylogenetic analysis indicated that multiple genomic duplication events have increased the numbers of TX and TN genes in Arabidopsis. Genes encoding TX, TN and TNL proteins were demonstrated in conifers; TX and TN genes are present in very low numbers in grass genomes. The expression, prevalence, and diversity of TX and TN genes suggests that these genes encode functional proteins rather than resulting from degradation or deletions of TNL genes. These TX and TN proteins could be plant analogues of small TIR-adapter proteins that function in mammalian innate immune responses such as MyD88 and Mal.

Amino Acid Sequence↗

Kindling-induced overexpression of Homer 1A and its functional implications for epileptogenesis.

Despite an extensive research on the molecular basis of epilepsy, the essential players in the epileptogenic process leading to epilepsy are not known. Gene expression analysis is one strategy to enhance our understanding of the genes contributing to the functional neuronal changes underlying epileptogenesis. In the present study, we used the novel MPSS (massively parallel signature sequencing) method for analysis of gene expression in the rat kindling model of temporal lobe epilepsy. Kindling by repeated electrical stimulation of the amygdala resulted in the differential expression of 264 genes in the hippocampus compared to sham controls. The most strongly induced gene was Homer 1A, an immediate early gene involved in the modulation of glutamate receptor function. The overexpression of Homer 1A in the hippocampus of kindled rats was confirmed by RT-PCR. In order to evaluate the functional implications of Homer 1A overexpression for kindling, we used transgenic mice that permanently overexpress Homer 1A. Immunohistochemical characterization of these mice showed a marked Homer 1A overexpression in glutamatergic neurons of the hippocampus. Kindling of Homer 1A overexpressing mice resulted in a retardation of seizure generalization compared to wild-type controls. The data demonstrate that kindling-induced epileptogenesis leads to a striking overexpression of Homer 1A in the hippocampus, which may represent an intrinsic antiepileptogenic and anticonvulsant mechanism in the course of epileptogenesis that counteracts progression of the disease.

Animals↗

Statistical analysis of MPSS measurements: application to the study of LPS-activated macrophage gene expression.

Massively Parallel Signature Sequencing (MPSS), a recently developed high-throughput transcription profiling technology, has the ability to profile almost every transcript in a sample without requiring prior knowledge of the sequence of the transcribed genes. As is the case with DNA microarrays, effective data analysis depends crucially on understanding how noise affects measurements. We analyze the sources of noise in MPSS and present a quantitative model describing the variability between replicate MPSS assays. We use this model to construct statistical hypotheses that test whether an observed change in gene expression in a pair-wise comparison is significant. This analysis is then extended to the determination of the significance of changes in expression levels measured over the course of a time series of measurements. We apply these analytic techniques to the study of a time series of MPSS gene expression measurements on LPS-stimulated macrophages. To evaluate our statistical significance metrics, we compare our results with published data on macrophage activation measured by using Affymetrix GeneChips.

Base Sequence↗

Microbial diversity in the deep sea and the underexplored "rare biosphere".

The evolution of marine microbes over billions of years predicts that the composition of microbial communities should be much greater than the published estimates of a few thousand distinct kinds of microbes per liter of seawater. By adopting a massively parallel tag sequencing strategy, we show that bacterial communities of deep water masses of the North Atlantic and diffuse flow hydrothermal vents are one to two orders of magnitude more complex than previously reported for any microbial environment. A relatively small number of different populations dominate all samples, but thousands of low-abundance populations account for most of the observed phylogenetic diversity. This "rare biosphere" is very ancient and may represent a nearly inexhaustible source of genomic innovation. Members of the rare biosphere are highly divergent from each other and, at different times in earth's history, may have had a profound impact on shaping planetary processes.

Biodiversity↗

Identification of the maturation factor for dual oxidase. Evolution of an eukaryotic operon equivalent.

Dual oxidase 2 (DUOX2), an NADPH:O(2) oxidoreductase flavoprotein, is a component of the thyroid H(2)O(2) generator crucial for hormone synthesis at the apical membrane. Mutations in DUOX2 produce congenital hypothyroidism in humans. However, no functional DUOX-based NADPH oxidase has ever been reconstituted at the plasma membrane of transfected cells. It has been proposed that DUOX retention in the endoplasmatic reticulum (ER) of heterologous systems is due to the lack of an unidentified component required for functional maturation of the enzyme. By data mining of a massively parallel signature sequencing tissue expression data base, we identified an uncharacterized gene named DUOX maturation factor (DUOXA2) arranged head-to-head to and co-expressed with DUOX2. A paralog (DUOXA1) was similarly linked to DUOX1. The genomic rearrangement leading to linkage of ancient DUOX and DUOXA genes could be traced back before the divergence of echinoderms. We demonstrate that co-expression of DUOXA2, an ER-resident transmembrane protein, allows ER-to-Golgi transition, maturation, and translocation to the plasma membrane of functional DUOX2 in a heterologous system. The identification of DUOXA genes has important implications for studies of the molecular mechanisms controlling DUOX expression and the molecular genetics of congenital hypothyroidism.

Amino Acid Sequence↗