Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Massively parallel sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Integrating microarrays into disease-gene identification strategies.

Positional cloning represents one of the most successful paradigm shifts in identifying the underlying patho-mechanisms in human disease. While traditional discovery tools focused on identifying defects at the tissue or cellular level, positional cloning identifies the damaged region of the genome as the preliminary step. While a large number of inherited single gene disorders have been mapped using this approach, a bottleneck still exists in combing through the genomic interval, often millions of nucleotides in length, to identify the nucleotide changes which result in a defective protein and subsequent disease. Along with the recent unravelling of the human genetic code, the development of massively parallel tools, such as microarrays, represent an equally important step forward in unraveling pathogenic genome dysfunctions. There are many emerging variants on microarray technology, such as expression arrays, exon arrays, array-based comparative genomic hybridization and sequencing arrays. Several of these platforms, if used properly, can accelerate the positional cloning process. The proper use of the platform is driven by knowledge of the underlying molecular defect being searched for and the operating characteristics of the array. The resultant insight forms the basis for improved molecular diagnostics and novel therapeutic targets.

Computational Biology↗

Massively parallel characterization of adolescent idiopathic scoliosis risk variants.

Adolescent idiopathic scoliosis (AIS) is a common pediatric musculoskeletal disorder characterized by lateral spinal curvature, often leading to chronic pain and deformity. Although a significant genetic component to AIS is recognized, the functional impact of most associated genetic variants, particularly those in noncoding regions, remains largely unknown. Using massively parallel reporter assays, we characterize 1664 variant positions in linkage disequilibrium with 26 AIS lead variants identified by genome-wide association studies (GWASs) in chondrocytes, a major cell type implicated in AIS pathogenesis. Using a library of 7173 candidate regulatory sequences, we compare the 1664 reference alleles against 4708 alternate alleles in two human chondrocyte cell lines (TC28a2 and SW1353). Our analysis identifies 92 variants that exhibit significant differential regulatory activity between their reference and alternate alleles, 79 of which are predicted to disrupt transcription factor binding sites, often correlating with their observed regulatory effect. Notably, we validate rs9496392, a single-nucleotide variant near the ADGRG6 locus, which shows consistent differential regulatory activity in both cell lines. ADGRG6 is a key regulator of cartilage homeostasis, and its cartilage-specific knockout in mice results in a scoliosis-like phenotype. The AIS risk allele of rs9496392 (T) is predicted to strongly disrupt several TFBSs, including SP1. This study provides a foundational catalog of functional AIS-associated regulatory variants active in chondrocytes, offering crucial insights into the perturbed gene regulatory networks in AIS. These findings lay the groundwork for identifying biomarkers and potential therapeutic targets for this complex childhood disease.

Journal Article↗

Chromosome Conformation Capture Carbon Copy (5C): a massively parallel solution for mapping interactions between genomic elements.

Physical interactions between genetic elements located throughout the genome play important roles in gene regulation and can be identified with the Chromosome Conformation Capture (3C) methodology. 3C converts physical chromatin interactions into specific ligation products, which are quantified individually by PCR. Here we present a high-throughput 3C approach, 3C-Carbon Copy (5C), that employs microarrays or quantitative DNA sequencing using 454-technology as detection methods. We applied 5C to analyze a 400-kb region containing the human beta-globin locus and a 100-kb conserved gene desert region. We validated 5C by detection of several previously identified looping interactions in the beta-globin locus. We also identified a new looping interaction in K562 cells between the beta-globin Locus Control Region and the gamma-beta-globin intergenic region. Interestingly, this region has been implicated in the control of developmental globin gene switching. 5C should be widely applicable for large-scale mapping of cis- and trans- interaction networks of genomic elements and for the study of higher-order chromosome structure.

Base Sequence↗

Clinical applications of microarray-based diagnostic tests.

Nearly 15 years have passed since the possibility of analyzing nucleic acid analytes in a massively parallel fashion was proposed using the then new concept of microarrays. A decade ago, proof of principle demonstration projects established the use of high density microarrays to genotype multiple polymorphisms within a large gene [cystic fibrosis transmembrance regulator (CFTR)], to rapidly analyze DNA sequences by hybridization and to ascertain differential gene expression of the entire genome of an organism. The use of microarrays has had an explosive influence on the rate at which new biological information can be learned, including in a nonhypothesis driven manner. The past decade has also seen these research tools applied increasingly to questions of clinical and medical relevance. Genotyping drug metabolizing enzyme genes, resequencing important tumor suppressor genes, and classifying neoplastic disease by differential gene expression profiles are but a few of the many possibilities to provide clinically useful information using microarray-based diagnostic tests.

Base Sequence↗

Secondary structure computer prediction of the poliovirus 5' non-coding region is improved by a genetic algorithm.

Comparison of the secondary structure of the 5' non-coding region of poliovirus 3 RNA derived from the genetic algorithm with the model of Skinner et al. (J. Mol. Biol., 207, 379-392, 1989) demonstrates many of the confirmed structural elements. The genetic algorithm (Shapiro and Navetta, J. Supercomput., 8, 195-201, 1994) generates a population of all possible stems, then mixes, combines, and recombines these stems in multiple iterations on a massively parallel computer, ultimately selecting a most fit structure based on its energy. The secondary structure of the region containing the determinants of neurovirulence was better predicted using the genetic algorithm, whereas the dynamic programming algorithm (Zuker, Science, 244, 48-52, 1989) required phylogenetic comparative sequence analysis to arrive at the correct conclusion. In addition, artificial mutations were introduced throughout this region of the genome and although rearrangements in structure may occur, many structures persisted, suggesting that the given structures thus selected may have evolved to withstand isolated mutations. The genetic algorithm-derived structure for the 5' non-coding region compares favorably with the biological data and functions previously described, and contains all of the 'persistent' structures, suggesting also that the persistence factor may be an aid to validating structures.

Algorithms↗

Ultrathin-layer gel electrophoresis of biopolymers.

Emerging need for large-scale, high-resolution analysis of biopolymers, such as DNA sequencing polymerase chain reaction, (PCR) product sizing, single nucleotide polymorphism (SNP) hunting and analysis of protein molecules necessitated the development of automated and high-throughput gel electrophoresis based methods enabling rapid, high-performance separations in a wide molecular weight range. Scaling down electric field mediated separation processes supports higher throughput due to the applicability of higher voltages, thus speeding up analysis time. Indeed, efforts in miniaturization resulted in faster, easier, less costly and more convenient analyses, fulfilling the needs of the emerging biotechnology industry for microscale and massively parallel assays. The two primary approaches in miniaturizing electrophoresis dimensions are the capillary and microslab formats. This latter one evolved towards ultrathin-layer gel electrophoresis which is, except from the thickness of the separation platform, slightly in the upper side of the scale, resulting in considerably easier handling. Ultrathin-layer gel electrophoresis combines the advantages of conventional slab-gel electrophoresis (multilane format) and capillary gel electrophoresis (rapid, high-efficiency separations). It is readily automated, automatic versions of it have been extensively used for large-scale DNA sequencing in the Human Genome Project and more recently became popular in high throughput DNA fragment analysis. Ultrathin-layer techniques are the first step towards the wider use of electrophoresis microchips in perfecting a user-friendly interface between the user and the microdevice.

Animals↗

PrimerStation: a highly specific multiplex genomic PCR primer design server for the human genome.

PrimerStation (http://ps.cb.k.u-tokyo.ac.jp) is a web service that calculates primer sets guaranteeing high specificity against the entire human genome. To achieve high accuracy, we used the hybridization ratio of primers in liquid solution. Calculating the status of sequence hybridization in terms of the stringent hybridization ratio is computationally costly, and no web service checks the entire human genome and returns a highly specific primer set calculated using a precise physicochemical model. To shorten the response time, we precomputed candidates for specific primers using a massively parallel computer with 100 CPUs (SunFire 15 K) about 3 months in advance. This enables PrimerStation to search and output qualified primers interactively. PrimerStation can select highly specific primers suitable for multiplex PCR by seeking a wider temperature range that minimizes the possibility of cross-reaction. It also allows users to add heuristic rules to the primer design, e.g. the exclusion of single nucleotide polymorphisms (SNPs) in primers, the avoidance of poly(A) and CA-repeats in the PCR products, and the elimination of defective primers using the secondary structure prediction. We performed several tests to verify the PCR amplification of randomly selected primers for ChrX, and we confirmed that the primers amplify specific PCR products perfectly.

DNA Primers↗

Detection of sequences in the cerebellar cortex: numerical estimate of the possible number of tidal-wave inducing sequences represented.

The two major cortices of the brain--the cerebral and cerebellar cortex--are massively connected through intercalated nuclei (pontine, cerebellar and thalamic nuclei). We suggest that the two cortices co-operate by generating precise temporal patterns in the cerebral cortex that are detected in the cerebellar cortex as temporal patterns assembled spatially in the mossy fibers. We will begin by showing that the tidal-wave mechanism works in the cerebellar cortex as a read-out mechanism for such spatio-temporal patterns due to the synchronous activity they generate in the parallel fiber system which drives the Purkinje cells--the output neurons of the cerebellar cortex--to fire action potentials. We will review the anatomy of the mossy fibers and show that within a "beam", or "row" of cerebellar cortex the mossy fibers in principle could embed a vast number of tidal-wave generating sequences. Based on anatomical data we will argue that the cerebellar mossy fiber-granule cell-Purkinje cell system can potentially detect and--through learning--select from an enormous number of spatio-temporal patterns.

Animals↗

Impact of massively parallel computation on protein structure determination.

For the past two decades, an important paradigm in protein chemistry has been the assertion that a biologically active protein is at thermodynamic equilibrium and therefore adopts its minimum free energy structure. Although some evidence now suggests that not all proteins conform to this notion, it is true often enough to remain an important guiding principle in structure determination, whether by direct computation or by the computationally assisted approaches of diffraction and resonance. Among the difficulties in predicting structure from sequence are the lack of a useful potential function incorporating the influence of solvent and the inability to sample the phase space efficiently or even to determine whether a free energy minimum is, in fact, the global minimum. These problems are general. Although they are greatly mitigated by experimental information, they become increasingly severe as empirical constraints are reduced. We review the difficulties involved in the general problem of protein structure prediction and discuss the impact of increased computer power in the context of new approaches to solvation and parallel algorithm design. A general focus of our discussion is the need to understand the theoretical basis for effective theories and to accommodate in numerical methods the interplay of different temporal and spatial scales.

Algorithms↗

An algorithm for assembly of ordered restriction maps from single DNA molecules.

The restriction mapping of a massive number of individual DNA molecules by optical mapping enables assembly of physical maps spanning mammalian and plant genomes; however, not through computational means permitting completely de novo assembly. Existing algorithms are not practical for genomes larger than lower eukaryotes due to their high time and space complexity. In many ways, sequence assembly parallels map assembly, so that the overlap-layout-consensus strategy, recently shown effective in assembling very large genomes in feasible time, sheds new light on solving map construction issues associated with single molecule substrates. Accordingly, we report an adaptation of this approach as the formal basis for de novo optical map assembly and demonstrate its computational feasibility for assembly of very large genomes. As such, we discuss assembly results for a series of genomes: human, plant, lower eukaryote and bacterial. Unlike sequence assembly, the optical map assembly problem is actually more complex because restriction maps from single molecules are constructed, manifesting errors stemming from: missing cuts, false cuts, and high variance of estimated fragment sizes; chimeric maps resulting from artifactually merged molecules; and true overlap scores that are "in the noise" or "slightly above the noise." We address these problems, fundamental to many single molecule measurements, by an effective error correction method using global overlap information to eliminate spurious overlaps and chimeric maps that are otherwise difficult to identify.

Algorithms↗

Computing with DNA.

We consider molecular models for computing and derive a DNA-based mechanism for solving intractable problems through massive parallelism. In principle, such methods might reduce the effort needed to solve otherwise difficult tasks, such as factoring large numbers, a computationally intensive task whose intractability forms the basis for much of modern cryptography.

Base Composition↗

Large-scale molecular dynamics simulation of DNA: implementation and validation of the AMBER98 force field in LAMMPS.

Molecular modelling played a central role in the discovery of the structure of DNA by Watson and Crick. Today, such modelling is done on computers: the more powerful these computers are, the more detailed and extensive can be the study of the dynamics of such biological macromolecules. To fully harness the power of modern massively parallel computers, however, we need to develop and deploy algorithms which can exploit the structure of such hardware. The Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) is a scalable molecular dynamics code including long-range Coulomb interactions, which has been specifically designed to function efficiently on parallel platforms. Here we describe the implementation of the AMBER98 force field in LAMMPS and its validation for molecular dynamics investigations of DNA structure and flexibility against the benchmark of results obtained with the long-established code AMBER6 (Assisted Model Building with Energy Refinement, version 6). Extended molecular dynamics simulations on the hydrated DNA dodecamer d(CTTTTGCAAAAG)(2), which has previously been the subject of extensive dynamical analysis using AMBER6, show that it is possible to obtain excellent agreement in terms of static, dynamic and thermodynamic parameters between AMBER6 and LAMMPS. In comparison with AMBER6, LAMMPS shows greatly improved scalability in massively parallel environments, opening up the possibility of efficient simulations of order-of-magnitude larger systems and/or for order-of-magnitude greater simulation times.

Algorithms↗

Gene expression analysis with universal n-mer arrays.

Gene expression profiling is one of the many applications that have benefited from the massively parallel nucleic acid detection capability of DNA microarrays. Current expression arrays, however, are expensive and inflexible. They are custom-designed for each organism and they do not offer the possibility of incorporating updated genomic information without production of a new chip. One possible solution is the development of a universal chip, consisting of all 4n possible DNA sequences of length n. Studying different organisms or new genes would simply require modifications to the hybridization pattern analysis software. The key problem is to find a value of n that is large enough to afford sufficient specificity, yet is small enough for practical fabrication and readout. We developed an analytical model, supported by computer-assisted calculation with yeast and mouse transcript data, to argue that it is both practical and useful to fabricate n-mer arrays with 10 < or = n < or = 16.

Alternative Splicing↗

A systematic strategy for identifying causal single nucleotide polymorphisms and their target genes on Juvenile arthritis risk haplotypes.

BACKGROUND: Although genome-wide association studies (GWAS) have identified multiple regions conferring genetic risk for juvenile idiopathic arthritis (JIA), we are still faced with the task of identifying the single nucleotide polymorphisms (SNPs) on the disease haplotypes that exert the biological effects that confer risk. Until we identify the risk-driving variants, identifying the genes influenced by these variants, and therefore translating genetic information to improved clinical care, will remain an insurmountable task. We used a function-based approach for identifying causal variant candidates and the target genes on JIA risk haplotypes. METHODS: We used a massively parallel reporter assay (MPRA) in myeloid K562 cells to query the effects of 5,226 SNPs in non-coding regions on JIA risk haplotypes for their ability to alter gene expression when compared to the common allele. The assay relies on 180&#xa0;bp oligonucleotide reporters ("oligos") in which the allele of interest is flanked by its cognate genomic sequence. Barcodes were added randomly by PCR to each oligo to achieve&#x2009;>&#x2009;20 barcodes per oligo to provide a quantitative read-out of gene expression for each allele. Assays were performed in both unstimulated K562 cells and cells stimulated overnight with interferon gamma (IFNg). As proof of concept, we then used CRISPRi to demonstrate the feasibility of identifying the genes regulated by enhancers harboring expression-altering SNPs. RESULTS: We identified 553 expression-altering SNPs in unstimulated K562 cells and an additional 490 in cells stimulated with IFNg. We further filtered the SNPs to identify those plausibly situated within functional chromatin, using open chromatin and H3K27ac ChIPseq peaks in unstimulated cells and open chromatin plus H3K4me1 in stimulated cells. These procedures yielded 42 unique SNPs (total&#x2009;=&#x2009;84) for each set. Using CRISPRi, we demonstrated that enhancers harboring MPRA-screened variants in the TRAF1 and LNPEP/ERAP2 loci regulated multiple genes, suggesting complex influences of disease-driving variants. CONCLUSION: Using MPRA and CRISPRi, JIA risk haplotypes can be queried to identify plausible candidates for disease-driving variants. Once these candidate variants are identified, target genes can be identified using CRISPRi informed by the 3D chromatin structures that encompass the risk haplotypes.

Humans↗

Cancer cell-based genomic and small molecule screens.

This chapter focuses on the promising post-genomic technologies being used for discovery of new, safer, and better cancer drugs and drug targets. Since cancer is largely a disease of the cell, usually involving unrestricted cell proliferation as a result of heritable genetic changes such as mutation, this chapter will focus on cell-centric technologies and their utility in addressing major questions in cancer biology. Recent advances in cell-based technology, including phenotypic assays, image-based readouts, primary tumor cell growth and maintenance in vitro, gene and small molecule delivery tools, and automated systems for cell manipulation, provide a novel means to understand the etiology and mechanisms of cancer as never before. In addition to the abundant tool sophistication, many aspects of cancer can be emulated and monitored in cell systems, which makes them ideal vehicles for exploitation to discover new targets and drugs. This chapter will first handle nomenclature and provide a context for a "good drug target" within the framework of the human genome, then overview functional genomic gene-based library screening approaches with specific applications to cancer target discovery. Second, small molecule screening applications will be handled, with an emphasis on the new paradigm of massively parallel screening and resultant multidimensional dataset analysis approaches to identify drug candidates, assign mechanism of action, and address problems in deriving selective and safe chemical entities.

Animals↗

Allosteric selection of ribozymes that respond to the second messengers cGMP and cAMP.

RNA transcripts containing the hammerhead ribozyme have been engineered to self-destruct in the presence of specific nucleoside 3',5'-cyclic monophosphate compounds. These RNA molecular switches were created by a new combinatorial strategy termed 'allosteric selection,' which favors the emergence of ribozymes that rapidly self-cleave only when incubated with their corresponding effector compounds. Representative RNAs exhibit 5,000-fold activation upon cGMP or cAMP addition, display precise molecular recognition characteristics, and operate with catalytic rates that match those exhibited by unaltered ribozymes. These findings demonstrate that a vast number of ligand-responsive ribozymes with dynamic structural characteristics can be generated in a massively parallel fashion. Moreover, optimized allosteric ribozymes could serve as highly selective sensors of chemical agents or as unique genetic control elements for the programmed destruction of cellular RNAs.

Acids↗

RNA folding pathway functional intermediates: their prediction and analysis.

The massively parallel genetic algorithm (GA) for RNA structure prediction uses the concepts of mutation, recombination, and survival of the fittest to evolve a population of thousands of possible RNA structures toward a solution structure. As described below, the properties of the algorithm are ideally suited to use in the prediction of possible folding pathways and functional intermediates of RNA molecules given their sequences. Utilizing Stem Trace, an interactive visualization tool for RNA structure comparison, analysis of not only the solution ensembles developed by the algorithm, but also the stages of development of each of these solutions, can give strong insight into these folding pathways. The GA allows the incorporation of information from biological experiments, making it possible to test the influence of particular interactions between structural elements on the dynamics of the folding pathway. These methods are used to reveal the folding pathways of the potato spindle tuber viroid (PSTVd) and the host killing mechanism of Escherichia coli plasmid R1, both of which are successfully explored through the combination of the GA and Stem Trace. We also present novel intermediate folds of each molecule, which appear to be phylogenetically supported, as determined by use of the methods described below.

Algorithms↗