Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Massively parallel sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Genome wide profiling of human embryonic stem cells (hESCs), their derivatives and embryonal carcinoma cells to develop base profiles of U.S. Federal government approved hESC lines.

BACKGROUND: In order to compare the gene expression profiles of human embryonic stem cell (hESC) lines and their differentiated progeny and to monitor feeder contaminations, we have examined gene expression in seven hESC lines and human fibroblast feeder cells using Illumina bead arrays that contain probes for 24,131 transcript probes. RESULTS: A total of 48 different samples (including duplicates) grown in multiple laboratories under different conditions were analyzed and pairwise comparisons were performed in all groups. Hierarchical clustering showed that blinded duplicates were correctly identified as the closest related samples. hESC lines clustered together irrespective of the laboratory in which they were maintained. hESCs could be readily distinguished from embryoid bodies (EB) differentiated from them and the karyotypically abnormal hESC line BG01V. The embryonal carcinoma (EC) line NTera2 is a useful model for evaluating characteristics of hESCs. Expression of subsets of individual genes was validated by comparing with published databases, MPSS (Massively Parallel Signature Sequencing) libraries, and parallel analysis by microarray and RT-PCR. CONCLUSION: we show that Illumina's bead array platform is a reliable, reproducible and robust method for developing base global profiles of cells and identifying similarities and differences in large number of samples.

Carcinoma, Embryonal↗

Analysis of tag-position bias in MPSS technology.

BACKGROUND: Massively Parallel Signature Sequencing (MPSS) technology was recently developed as a high-throughput technology for measuring the concentration of mRNA transcripts in a sample. It has previously been observed that the position of the signature tag in a transcript (distance from 3' end) can affect the measurement, but this effect has not been studied in detail. RESULTS: We quantify the effect of tag-position bias in Classic and Signature MPSS technology using published data from Arabidopsis, rice and human. We investigate the relationship between measured concentration and tag-position using nonlinear regression methods. The observed relationship is shown to be broadly consistent across different data sets. We find that there exist different and significant biases in both Classic and Signature MPSS data. For Classic MPSS data, genes with tag-position in the middle-range have highest measured abundance on average while genes with tag-position in the high-range, far from the 3' end, show a significant decrease. For Signature MPSS data, high-range tag-position genes tend to have a flatter relationship between tag-position and measured abundance. Thus, our results confirm that the Signature MPSS method fixes a substantial problem with the Classic MPSS method. For both Classic and Signature MPSS data there is a positive correlation between measured abundance and tag-position for low-range tag-position genes. Compared with the effects of mRNA length and number of exons, tag-position bias seems to be more significant in Arabadopsis. The tag-position bias is reflected both in the measured abundance of genes with a significant tag count and in the proportion of unexpressed genes identified. CONCLUSION: Tag-position bias should be taken into consideration when measuring mRNA transcript abundance using MPSS technology, both in Classic and Signature MPSS methods.

Arabidopsis↗

Genome-wide gene expression profiling in Arabidopsis thaliana reveals new targets of abscisic acid and largely impaired gene regulation in the abi1-1 mutant.

The phytohormone abscisic acid (ABA) plays important regulatory roles in many plant developmental processes including seed dormancy, germination, growth, and stomatal movements. These physiological responses to ABA are in large part brought about by changes in gene expression. To study genome-wide ABA-responsive gene expression we applied massively parallel signature sequencing (MPSS) to samples from Arabidopsis thaliana wildtype (WT) and abi1-1 mutant seedlings. We identified 1354 genes that are either up- or downregulated following ABA treatment of WT seedlings. Among these ABA-responsive genes, many encode signal transduction components. In addition, we identified novel ABA-responsive gene families including those encoding ribosomal proteins and proteins involved in regulated proteolysis. In the ABA-insensitive mutant abi1-1, ABA regulation of about 84.5% and 6.9% of the identified genes was impaired or strongly diminished, respectively; however, 8.6% of the genes remained appropriately regulated. Compared to other methods of gene expression analysis, the high sensitivity and specificity of MPSS allowed us to identify a large number of ABA-responsive genes in WT Arabidopsis thaliana. The database given in our supplementary material (http://jcs.biologists.org/supplemental) provides researchers with the opportunity to rapidly assess whether genes of interest may be regulated by ABA. Regulation of the majority of the genes by ABA was impaired in the ABA-insensitive mutant abi1-1. However, a subset of genes continued to be appropriately regulated by ABA, which suggests the presence of at least two ABA signaling pathways, only one of which is blocked in abi1-1.

Abscisic Acid↗

Transcriptome profiling of human and murine ESCs identifies divergent paths required to maintain the stem cell state.

Human embryonic stem cells (hESCs) are an important source of stem cells in regenerative medicine, and much remains unknown about their molecular characteristics. To develop a detailed genomic profile of ESC lines in two different species, we compared transcriptomes of one murine and two different hESC lines by massively parallel signature sequencing (MPSS). Over 2 million signature tags from each line and their differentiating embryoid bodies were sequenced. Major differences and conserved similarities between species identified by MPSS were validated by reverse transcription polymerase chain reaction (RT-PCR) and microarray. The two hESC lines were similar overall, with differences that are attributable to alleles and propagation. Human-mouse comparisons, however, identified only a small (core) set of conserved genes that included genes known to be important in ESC biology, as well as additional novel genes. Identified were major differences in leukemia inhibitory factor, transforming growth factor-beta, and Wnt and fibroblast growth factor signaling pathways, as well as the expression of genes encoding metabolic, cytoskeletal, and matrix proteins, many of which were verified by RT-PCR or by comparing them with published databases. The study reported here underscores the importance of cross-species comparisons and the versatility and sensitivity of MPSS as a powerful complement to current array technology.

Animals↗

Assessing self-renewal and differentiation in human embryonic stem cell lines.

Like other cell populations, undifferentiated human embryonic stem cells (hESCs) express a characteristic set of proteins and mRNA that is unique to the cells regardless of culture conditions, number of passages, and methods of propagation. We sought to identify a small set of markers that would serve as a reliable indicator of the balance of undifferentiated and differentiated cells in hESC populations. Markers of undifferentiated cells should be rapidly downregulated as the cells differentiate to form embryoid bodies (EBs), whereas markers that are absent or low during the undifferentiated state but that are induced as hESCs differentiate could be used to assess the presence of differentiated cells in the cultures. In this paper, we describe a list of markers that reliably distinguish undifferentiated and differentiated cells. An initial list of approximately 150 genes was generated by scanning published massively parallel signature sequencing, expressed sequence tag scan, and microarray datasets. From this list, a subset of 109 genes was selected that included 55 candidate markers of undifferentiated cells, 46 markers of hESC derivatives, four germ cell markers, and four trophoblast markers. Expression of these candidate marker genes was analyzed in undifferentiated hESCs and differentiating EB populations in four different lines by immunocytochemistry, reverse transcription-polymer-ase chain reaction (RT-PCR), microarray analysis, and quantitative RT-PCR (qPCR). We show that qPCR, with as few as 12 selected genes, can reliably distinguish differentiated cells from undifferentiated hESC populations.

Animals↗

Digital expression profiles of human endogenous retroviral families in normal and cancerous tissues.

Human endogenous retroviruses (HERVs) are remnants of ancient retroviral infections that became fixed in the germ line DNA millions of years ago. The fact that humoral and cellular immune responses against HERV-encoded proteins have been identified in cancer patients suggests that these antigens might be used in cancer immunotherapy or diagnosis. We analyzed the digital expression patterns of the HERV-K (HML-2), -W, -H and -E families in normal and cancerous tissues. Thirty-one proviral members of the HERV-K family and one representative each for the other HERV families were used as probes to search human EST data. Matching of HERV proviruses to ESTs was HERV family-specific and the expression profiles of the HERV families distinct. The HERV-K family was expressed in normal tissues such as muscle, skin and brain, as well as in germ cell tumors and other cancerous tissues. HERV-H was the only family expressed in cancers of the intestine, bone marrow, bladder and cervix, and was more highly expressed than the other families in cancers of the stomach, colon and prostate. In contrast, HERV-W was predominantly expressed in normal placenta. Expression patterns were confirmed by MPSS (massively parallel signature sequencing) data where available. For the HERV-K family, we mapped most ESTs to their corresponding proviruses and assessed the coding capacities of the matched proviruses. This study shows that HERV families are more widely expressed than originally thought and that some members of the HERV-K and -H families could encode targets for cancer immunotherapy.

Chromosome Mapping↗

Ligation errors in DNA computing.

DNA computing is a novel method of computing proposed by Adleman (1994), in which the data is encoded in the sequences of oligonucleotides. Massively parallel reactions between oligonucleotides are expected to make it possible to solve huge problems. In this study, reliability of the ligation process employed in the DNA computing is tested by estimating the error rate at which wrong oligonucleotides are ligated. Ligation of wrong oligonucleotides would result in a wrong answer in the DNA computing. The dependence of the error rate on the number of mismatches between oligonucleotides and on the combination of bases is investigated.

Animals↗

Long-range mRNA folding shapes expression and sequence of bacterial genes.

Bacterial gene expression is strongly influenced by local mRNA secondary structure, yet the impact of long-range folding remains poorly understood. Here, we show that sequences hundreds of nucleotides from the mRNA 5' end can act as potent repressors of gene expression through long-range base pairing to the ribosome binding site (RBS), subjecting anti-RBS sequences to negative selection. Using massively parallel reporter assays in Bacillus subtilis, we identify anti-RBS sequences as among the strongest determinants of reduced mRNA abundance across the transcript body. We demonstrate that distal anti-RBS elements engage in long-range folding with the Shine-Dalgarno sequence, blocking ribosome entry and promoting mRNA decay. Consistent with these repressive effects, anti-RBS-like sequences are depleted throughout diverse bacterial coding sequences but not from leaderless transcripts, and introducing distal anti-RBS to native genes reduces expression. Our findings establish that long-range mRNA folding is a conserved force shaping gene expression and constrains coding sequence evolution.

Bacillus subtilis↗

Molecular engineering approaches for DNA sequencing and analysis.

High-throughput DNA sequencing development for mutation screening and identification is essential to realize the goal of pharmacogenomics and personalized medicine, which will lead to a new era in clinical medicine and healthcare. Molecular engineering approaches to modify the building blocks of DNA by introducing functional groups for purification and detection has led to the development of high-throughput genetic analysis technologies. This review is focused on the following two DNA sequencing approaches. The first approach is based on the use of molecular affinity and mass spectrometry to perform quick and highly accurate mutation screening, heterozygote identification and insertion/deletion detection. The second approach is based on a sequencing-by-synthesis platform that has the potential for generating DNA sequencing data in a massive, parallel manner. The basic principles, fundamental challenges and methods of implementation of these exciting new technologies will be discussed.

Animals↗

Allele Level Sequencing of Killer Cell Immunoglobulin-Like Receptor Genes Using Oxford Nanopore Long Read Sequencing.

The human Killer cell Immunoglobulin-like Receptor (KIR) genes, found on chromosome 19, encode for cell surface protein receptors that, through interaction with their ligand, modulate the action of Natural Killer (NK) cells and some subsets of T lymphocytes. KIR genes exhibit extensive variation through variable gene content, copy number, and allele polymorphism. The combination of KIR genes and their ligands is implicated in various clinical settings including haematopoietic stem cell and solid organ transplant, and infectious disease progression. KIR gene content has been used in the selection of optimal stem cell donors with haplotype variations in recipient and donor giving differential clinical outcomes. With the introduction of massively parallel clonal next generation sequencing and single molecule long read third generation sequencing, allele level determination of KIR genotypes has become feasible. We describe a method for amplicon-based long read sequencing on the Oxford Nanopore Technologies platform that provides largely unambiguous allele level typing of KIR genes. The method was validated using DNA extracted from 48 10th International Histocompatibility Workshop (IHWS) cell lines with previously published allele level KIR genotypes and 176 Western Australian samples previously tested for the presence or absence of KIR genes. Our long-read sequencing method was able to accurately determine KIR alleles with an overall concordance of 97%-99% with the published data. Importantly, phasing ambiguity caused by the inability to phase heterozygous base positions over long stretches of gene sequence was resolved in several samples. Thus, our long read PCR sequencing strategy can be used to determine KIR genotypes at allele resolution level.

Humans↗

Single-cell transcriptomic atlas of Alzheimer's disease middle temporal gyrus reveals region, cell type and sex specificity of gene expression with novel genetic risk for MERTK in female.

Alzheimer's disease, the most common age-related neurodegenerative disease, is closely associated with both amyloid-ß plaque and neuroinflammation. Two thirds of Alzheimer's disease patients are females and they have a higher disease risk. Moreover, women with Alzheimer's disease have more extensive brain histological changes than men along with more severe cognitive symptoms and neurodegeneration. To identify how sex difference induces structural brain changes, we performed unbiased massively parallel single nucleus RNA sequencing on Alzheimer's disease and control brains focusing on the middle temporal gyrus, a brain region strongly affected by the disease but not previously studied with these methods. We identified a subpopulation of selectively vulnerable layer 2/3 excitatory neurons that that were RORB-negative and CDH9-expressing. This vulnerability differs from that reported for other brain regions, but there was no detectable difference between male and female patterns in middle temporal gyrus samples. Disease-associated, but sex-independent, reactive astrocyte signatures were also present. In clear contrast, the microglia signatures of diseased brains differed between males and females. Combining single cell transcriptomic data with results from genome-wide association studies (GWAS), we identified MERTK genetic variation as a risk factor for Alzheimer's disease selectively in females. Taken together, our single cell dataset revealed a unique cellular-level view of sex-specific transcriptional changes in Alzheimer's disease, illuminating GWAS identification of sex-specific Alzheimer's risk genes. These data serve as a rich resource for interrogation of the molecular and cellular basis of Alzheimer's disease.

Journal Article↗

Photolithographic synthesis of high-density oligonucleotide probe arrays.

High-density DNA probe arrays provide a massively parallel approach to nucleic acid sequence analysis that is transforming gene-based biomedical research and diagnostics. Light-directed combinatorial oligonucleotide synthesis has enabled the large-scale production of GeneChip probe arrays which contain several hundred of thousand oligonucleotide sequences on glass "chips" about one cm2 in size. Due to their very high information content, GeneChip probe arrays are finding widespread use in the hybridization-based detection and analysis of mutations and polymorphisms ("genotyping"), and in a wide range of gene expression studies. The manufacturing process integrates solid-phase photochemical oligonucleotide synthesis with lithographic techniques adapted from the microelectronics industry. The present-generation methodology employs MeNPOC photo-activatable nucleoside monomers with proximity photolithography, and is currently capable of printing individual 10 microns 2 probe features at a density of 10(6) probes/cm2.

DNA Probes↗

Sequencing single molecules of DNA.

In 2004, the NIH set a remarkable challenge: the 1000 dollars genome. Roughly speaking, success would provide, by 2015, the ability to sequence the complete genome of an individual human, quickly and at an accessible price. An intermediate goal of a 100,000 dollars genome was set for 2010. While the cost of Sanger sequencing has dropped dramatically over the past two decades, it is unlikely that the 100,000 dollars genome will be achieved by this means. New massively parallel technologies will push the cost of sequencing towards this mark, but it is doubtful whether these efforts will match the 1000 dollars goal. The best bets for ultrarapid, low-cost sequencing are single-molecule approaches.

DNA↗

Spatially resolved in vitro molecular ecology.

Sensitive CCD-based fluorescence detection has made spatially resolved studies of evolving cell-free molecular systems possible. In recent years our attention has focussed on making the transition to open and interacting spatially-resolved amplification systems using silicon microreactor technology and on providing a hardware platform for individual based simulation of such systems. Significant progress has been achieved in this direction. Open microflow reactors have been realized in zero (well-mixed), one and two dimensions with volumes small enough to allow long-time studies with limited biochemical materials. The primer directed 3SR reaction (amplifying DNA and RNA) has been used as a basis for constructing interacting model systems with both predator-prey and cooperative amplification character. Theoretical work has demonstrated the need for individual based modeling of such systems: a significant fraction of the population consists of distinct sequence polymers in any case. A massively parallel processor-configurable computer NGEN has been designed and constructed which allows the high speed simulation in hardware of relatively large populations of locally interacting individual strings of chosen length (e.g. up to 2000*2000 for 64 bases), in addition to its application as an evolvable hardware machine. Simulations show self-replicating spots to stabilize the cooperative amplification in evolving systems (a mechanism proposed by the author in 1994). Both oscillatory kinetics and pattern formation are expected in the experimental model systems under investigation which profoundly affect the course of evolution. Such in vitro model systems serve both to test current theories of cooperative evolution and provide clues for optimisation strategies in molecular biotechnology.

Biotechnology↗

Protein design by optimization of a sequence-structure quality function.

An automated procedure for protein design by optimization of a sequence-structure quality has been developed. The method selects a statistically optimal sequence for a particular structure, on the assumption that such a protein will adopt the desired structure. We present two optimization algorithms: one provides an exact optimization while the other uses a combinatorial technique for comparatively rapid results. Both are suitable for massively parallel computers. A prototype system was used to design sequences which should adopt the four-helix bundle conformation of myohemerythrin. These appear satisfactory to secondary structure and profile analysis. Detailed inspection reveals that the sequences are generally plausible but, as expected, lack some specific structural features. The design parameters provide some insight into the general determinants of protein structure.

Algorithms↗

Massive parallel analysis of DNA-Hoechst 33258 binding specificity with a generic oligodeoxyribonucleotide microchip.

A generic oligodeoxyribonucleotide microchip was used to determine the sequence specificity of Hoechst 33258 binding to double-stranded DNA. The generic microchip contained 4096 oxctadeoxynucleo-tides in which all possible 4(6)= 4096 hexadeoxy-nucleotide sequences are flanked on both the 3'- and 5'-ends with equimolar mixtures of four bases. The microchip was manufactured by chemical immobilization of presynthesized 8mers within polyacrylamide gel pads. A selected set of immobilized 8mers was converted to double-stranded form by hybridization with a mixture of fluorescently labeled complementary 8mers. Massive parallel measurements of melting curves were carried out for the majority of 2080 6mer duplexes, in both the absence and presence of the Hoechst dye. The sequence-specific affinity for Hoechst 33258 was calculated as the increase in melting temperature caused by ligand binding. The dye exhibited specificity for A:T but not G:C base pairs. The affinity is low for two A:T base pairs, increases significantly for three, and reaches a plateau for four A:T base pairs. The relative ligand affinity for all trinucleotide and tetranucleotide sequences (A/T)(3)and (A/T)(4)was estimated. The free energy of dye binding to several duplexes was calculated from the equilibrium melting curves of the duplexes formed on the oligonucleotide microchips. This method can be used as a general approach for massive screening of the sequence specificity of DNA-binding compounds.

Bisbenzimidazole↗

Massively parallel approaches for characterizing noncoding functional variation in human evolution.

The genetic differences underlying unique phenotypes in humans compared to our closest primate relatives have long remained a mystery. Similarly, the genetic basis of adaptations between human groups during our expansion across the globe is poorly characterized. Uncovering the downstream phenotypic consequences of these genetic variants has been difficult, as a substantial portion lies in noncoding regions, such as cis-regulatory elements (CREs). Here, we review recent high-throughput approaches to measure the functions of CREs and the impact of variation within them. CRISPR screens can directly perturb CREs in the genome to understand downstream impacts on gene expression and phenotypes, while massively parallel reporter assays can decipher the regulatory impact of sequence variants. Machine learning has begun to be able to predict regulatory function from sequence alone, further scaling our ability to characterize genome function. Applying these tools across diverse phenotypes, model systems, and ancestries is beginning to revolutionize our understanding of noncoding variation underlying human evolution.

Humans↗

The massively parallel genetic algorithm for RNA folding: MIMD implementation and population variation.

A massively parallel Genetic Algorithm (GA) has been applied to RNA sequence folding on three different computer architectures. The GA, an evolution-like algorithm that is applied to a large population of RNA structures based on a pool of helical stems derived from an RNA sequence, evolves this population in parallel. The algorithm was originally designed and developed for a 16384 processor SIMD (Single Instruction Multiple Data) MasPar MP-2. More recently it has been adapted to a 64 processor MIMD (Multiple Instruction Multiple Data) SGI ORIGIN 2000, and a 512 processor MIMD CRAY T3E. The MIMD version of the algorithm raises issues concerning RNA structure data-layout and processor communication. In addition, the effects of population variation on the predicted results are discussed. Also presented are the scaling properties of the algorithm from the perspective of the number of physical processors utilized and the number of virtual processors (RNA structures) operated upon.

Algorithms↗