Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Massively parallel sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Genomics and proteomics tools for the clinic.

The sequencing of the human genome was only made possible by the massively parallel use of automated high-throughput technologies. These technologies have required the development and interfacing of new hardware and software in a way which would have been hardly conceivable only ten years ago. As a consequence, an unbroken trend to more 'industrialized science' is apparent. The wealth of information generated by intensive sequencing efforts is now being further exploited by using complex tools to comprehensively analyze complex systems at the DNA, RNA and protein level. A landmark innovation was the introduction of the biochip principle, best exemplified with the development of the DNA chip, mainly used for RNA expression profiling. The chip principle, together with miniaturization, has now become the dominating theme for a number of new genomics and proteomics technologies, culminating in the lab-on-a-chip concept which, in the next five to ten years, could advance at a comparable rate to that of computers over the last 50 years. Some of the new technologies are already used for comprehensive analysis of clinical samples in an attempt to describe disease and disease risk at the molecular level. However, all of these technologies are far from routine in clinical use and it is also too early to decide whether molecular fingerprints or signature profiles will have the diagnostic and prognostic power currently predicted.

Gene Expression Profiling↗

Increased detection of structural templates using alignments of designed sequences.

Protein structure prediction by comparative modeling benefits greatly from the use of multiple sequence alignment information to improve the accuracy of structural template identification and the alignment of target sequences to structural templates. Unfortunately, this benefit is limited to those protein sequences for which at least several natural sequence homologues exist. We show here that the use of large diverse alignments of computationally designed protein sequences confers many of the same benefits as natural sequences in identifying structural templates for comparative modeling targets. A large-scale massively parallelized application of an all-atom protein design algorithm, including a simple model of peptide backbone flexibility, has allowed us to generate 500 diverse, non-native, high-quality sequences for each of 264 protein structures in our test set. PSI-BLAST searches using the sequence profiles generated from the designed sequences ("reverse" BLAST searches) give near-perfect accuracy in identifying true structural homologues of the parent structure, with 54% coverage. In 41 of 49 genomes scanned using reverse BLAST searches, at least one novel structural template (not found by the standard method of PSI-BLAST against PDB) is identified. Further improvements in coverage, through optimizing the scoring function used to design sequences and continued application to new protein structures beyond the test set, will allow this method to mature into a useful strategy for identifying distantly related structural templates.

Algorithms↗

Associative database of protein sequences.

MOTIVATION: We present a new concept that combines data storage and data analysis in genome research, based on an associative network memory. As an illustration, 115 000 conserved regions from over 73 000 published sequences (i.e. from the entire annotated part of the SWISSPROT sequence database) were identified and clustered by a self-organizing network. Similarity and kinship, as well as degree of distance between the conserved protein segments, are visualized as neighborhood relationship on a two-dimensional topographical map. RESULTS: Such a display overcomes the restrictions of linear list processing and allows local and global sequence relationships to be studied visually. Families are memorized as prototype vectors of conserved regions. On a massive parallel machine, clustering and updating of the database take only a few seconds; a rapid analysis of incoming data such as protein sequences or ESTs is carried out on present-day workstations. AVAILABILITY: Access to the database is available at http://www.bioinf.mdc-berlin.de/unter2.html++ + CONTACT: (hanke,lehmann,reich)@mdc-berlin.de; bork@embl-heidelberg.de

Amino Acid Sequence↗

Protein sequencing experiment planning using analogy.

Experiment design and execution is a central activity in the natural sciences. The SeqER system provides a general architecture for the integration of automated planning techniques with a variety of domain knowledge in order to plan scientific experiments. These planning techniques include rule-based methods and, especially, the use of derivational analogy. Derivational analogy allows planning experience, captured as cases, to be reused. Analogy also allows the system to function in the absence of strong domain knowledge. Cases are efficiently and flexibly retrieved from a large casebase using massively parallel methods. The SeqER system is initially configured to plan protein sequencing experiments. Planning is interleaved with experiment execution, simulated using the SequenceIt program. SeqER interacts with a human user who analyzes the data generated by laboratory procedures and supplies SeqER with new hypotheses. SeqER is a vehicle in which to test theories about how scientists reason about experimental design.

Decision Making, Computer-Assisted↗

Amplification and assembly of chip-eluted DNA (AACED): a method for high-throughput gene synthesis.

A basic problem in gene synthesis is the acquisition of many short oligonucleotide sequences needed for the assembly of genes. Photolithographic methods for the massively parallel synthesis of high-density oligonucleotide arrays provides a potential source, once appropriate methods have been devised for their elution in forms suitable for enzyme-catalyzed assembly. Here, we describe a method based on the photolithographic synthesis of long (>60mers) single-stranded oligonucleotides, using a modified maskless array synthesizer. Once the covalent bond between the DNA and the glass surface is cleaved, the full-length oligonucleotides are selected and amplified using PCR. After cleavage of flanking primer sites, a population of unique, internal 40mer dsDNA sequences are released and are ready for use in biological applications. Subsequent gene assembly experiments using this DNA pool were performed and were successful in creating longer DNA fragments. This is the first report demonstrating the use of eluted chip oligonucleotides in biological applications such as PCR and assembly PCR.

Genes↗

Predicting dynamic expression patterns in budding yeast with a fungal DNA language model.

Predicting gene expression from DNA sequence remains challenging due to complex regulatory codes. We introduce a masked DNA language model pretrained on 165 fungal genomes closely related to budding yeast that captures conserved regulatory grammar. Fine-tuning the LM on yeast RNA-seq data-including high-resolution transcriptional regulator induction time courses generated in this study-yielded Shorkie, a model that substantially improves gene expression prediction compared to baselines trained without self-supervision. Shorkie identified canonical transcription factor (TF) binding motifs and tracked their usage across induction experiments. Furthermore, Shorkie accurately predicted variant effects, outperforming leading sequence-to-expression models in cis-eQTL classification and achieving high concordance with massively parallel reporter assays. Interpretability analyses revealed Shorkie's ability to resolve promoter dynamics, splicing signals, and temporal changes in regulatory motif usage. This framework demonstrates that evolutionary-scale pretraining combined with transfer learning substantially improves our ability to decode gene regulation from sequence, providing insights into noncoding variants and regulatory networks.

Journal Article↗

Characterization of synthetic DNA bar codes in Saccharomyces cerevisiae gene-deletion strains.

Incorporation of strain-specific synthetic DNA tags into yeast Saccharomyces cerevisiae gene-deletion strains has enabled identification of gene functions by massively parallel growth rate analysis. However, it is important to confirm the sequences of these tags, because mutations introduced during construction could lead to significant errors in hybridization performance. To validate this experimental system, we sequenced 11,812 synthetic 20-mer molecular bar codes and adjacent sequences (>1.8 megabases synthetic DNA) by pyrosequencing and Sanger methods. At least 31% of the genome-integrated 20-mer tags contain differences from those originally synthesized. However, these mutations result in anomalous hybridization in only a small subset of strains, and the sequence information enables redesign of hybridization probes for arrays. The robust performance of the yeast gene-deletion dual oligonucleotide bar-code design in array hybridization validates the use of molecular bar codes in living cells for tracking their growth phenotype.

DNA Primers↗

Prediction of RNA base pairing probabilities on massively parallel computers.

We present an implementation of McCaskill's algorithm for computing the base pair probabilities of an RNA molecule for massively parallel message passing architectures. The program can be used to routinely fold RNA sequences of more than 10,000 nucleotides. Applications to complete viral genomes are discussed.

Algorithms↗

Active learning of enhancers and silencers in the developing neural retina.

Deep learning is a promising strategy for modeling cis-regulatory elements. However, models trained on genomic sequences often fail to explain why the same transcription factor can activate or repress transcription in different contexts. To address this limitation, we developed an active learning approach to train models that distinguish between enhancers and silencers composed of binding sites for the photoreceptor transcription factor cone-rod homeobox (CRX). After training the model on nearly all bound CRX sites from the genome, we coupled synthetic biology with uncertainty sampling to generate additional rounds of informative training data. This allowed us to iteratively train models on data from multiple rounds of massively parallel reporter assays. The ability of the resulting models to discriminate between CRX sites with identical sequence but opposite functions establishes active learning as an effective strategy to train models of regulatory DNA. A record of this paper's transparent peer review process is included in the supplemental information.

Retina↗

fastDNAmL: a tool for construction of phylogenetic trees of DNA sequences using maximum likelihood.

We have developed a new tool, called fastDNAml, for constructing phylogenetic trees from DNA sequences. The program can be run on a wide variety of computers ranging from Unix workstations to massively parallel systems, and is available from the Ribosomal Database Project (RDP) by anonymous FTP. Our program uses a maximum likelihood approach and is based on version 3.3 of Felsenstein's dnaml program. Several enhancements, including algorithmic changes, significantly improve performance and reduce memory usage, making it feasible to construct even very large trees. Trees containing 40-100 taxa have been easily generated, and phylogenetic estimates are possible even when hundreds of sequences exist. We are currently using the tool to construct a phylogenetic tree based on 473 small subunit rRNA sequences from prokaryotes.

Algorithms↗

[Requirements for the development of molecularly defined targeted therapy in oncology].

BACKGROUND: The revolutionary development in the biosciences during the last 15 years--including the sequencing of the human genome and multiple other genomes as well the possibility of massive parallel analysis of thousands of data points within a single experiment--have greatly improved our understanding of the pathophysiology of tumor development and progression as well as metastasis. THE CURRENT SITUATION: The further development of genomic technology for routine clinical use is slowly but steadily leading to a paradigm shift: the currently used tumor definition based on pathologic and histological assessment is replaced by molecularly defined tumor subtypes. This is an essential prerequisite for the development of molecularly defined individualized therapies in oncology leading to more effective treatments with less severe side effects. The first molecularly defined drugs (targeted drugs) have already entered the clinical arena. However, not only diagnostics and therapy on oncology will change dramatically in the future, but also the way we have to organize oncology itself. OUTLOOK: The integration of genomics and information technologies will require the establishment of interdisciplinary teams beyond the already established collaborations between surgeons, radiotherapists and medical oncologists.

Algorithms↗

Regulatory Evolution and the Genetic Basis of Human Brain Expansion.

The evolution of the human brain is characterized by profound changes in structure and function, despite relatively limited divergence in protein-coding genes compared to other primates. This paradox has led to increasing recognition of gene regulatory elements (GREs) as primary drivers of evolutionary innovation. In this review, we synthesize current knowledge on the role of conserved noncoding elements (CNEs), human accelerated regions (HARs), and transposable element (TE)-derived sequences in shaping gene regulatory networks (GRNs) underlying brain development. Comparative analyses across humans and closely related primates, including the chimpanzee, gorilla, and orangutan, reveal that while core regulatory architectures are highly conserved, subtle changes in regulatory elements drive species-specific gene expression patterns. We highlight how CNEs provide a stable regulatory framework, whereas HARs and TE-derived elements introduce lineage-specific modifications that fine-tune neurodevelopmental processes. Advances in functional genomics, including CRISPR-based perturbations, massively parallel reporter assays, and single-cell multi-omics, have enabled direct interrogation of regulatory function, linking sequence variation to cellular phenotypes. Furthermore, we discuss how regulatory evolution contributes to both cognitive innovation and susceptibility to neurological disorders. Despite significant progress, challenges remain in establishing causal relationships between regulatory variation and phenotypic outcomes. Future integration of multi-omics data and comparative models will be essential for resolving these complexities. Together, this review provides a comprehensive framework for understanding the molecular basis of primate brain evolution through the lens of gene regulation.

Brain evolution↗

Group testing with DNA chips: generating designs and decoding experiments.

DNA microarrays are a valuable tool for massively parallel DNA-DNA hybridization experiments. Currently, most applications rely on the existence of sequence-specific oligonucleotide probes. In large families of closely related target sequences, such as different virus subtypes, the high degree of similarity often makes it impossible to find a unique probe for every target. Fortunately, this is unnecessary. We propose a microarray design methodology based on a group testing approach. While probes might bind to multiple targets simultaneously, a properly chosen probe set can still unambiguously distinguish the presence of one target set from the presence of a different target set. Our method is the first one that explicitly takes cross-hybridization and experimental errors into account while accommodating several targets. The approach consists of three steps: (1) Pre-selection of probe candidates, (2) Generation of a suitable group testing design, and (3) Decoding of hybridization results to infer presence or absence of individual targets. Our results show that this approach is very promising, even for challenging data sets and experimental error rates of up to 5%. On a data set of 28S rDNA sequences we were able to identify 660 sequences, a substantial improvement over a prior approach using unique probes which only identified 408 sequences.

Algorithms↗

Protein Threading Based on Multiple Protein Structure Alignment.

Protein threading, a method employed in protein three-dimensional (3D) structure prediction was only proposed in the early 1990's although predicting protein 3D structure from its given amino acid sequence has been around since 1970's. Here we describe a protein threading method/system that we have developed based on multiple protein structure alignment. In order to compute multiple structure alignments, we developed a similar structure search program on massive parallel computers and a program for constructing a multiple structure alignment from pairwise structure alignments, where the latter is based on the center star method for sequence alignment. A simple dynamic-programming based algorithm which uses a profile matrix obtained from the result of multiple structure alignment was also developed to compute a threading (i.e., an alignment between a target sequence and a known structure). Using this system, we participated in the threading category (category AL) of CASP3 (Third Community Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction). The results are discussed.

Journal Article↗

[DNA-arrays, a breakthrough in bacterial identification?].

DNA-arrays are mainly known for their application in transcriptome analysis leading for instance to the discovery of new marker genes for diagnostics and prognostics in oncology. However, DNA arrays are also used for massively parallel analysis of DNA molecules allowing their quantification, the detection of single nucleotide polymorphisms and re-sequencing. This multi detection system is now applied to the << old >> problems of detecting and identifying bacteria in a biological sample and for the fine molecular characterization of a bacterial isolate. This new tool should serve for the diagnostic of an infection and for epidemiological studies such as those performed for the control of nosocomial infections or for the surveillance of bioterrorism attacks. DNA arrays carrying probes for 16S RNA specific of hundreds of bacterial species allow the identification of bacteria within a community by a single hybridization of amplified 16S rDNAs with universal primers and re-sequencing DNA arrays are used for multi locus sequence typing in a single step. Finally, the genome of an isolate could be characterized by DNA-arrays focused on a specific question like presence of toxin or antibiotic resistance genes. Up to now, DNA arrays are used in research laboratories for the rapid characterization at the genomic level of a strain collection, for evolutionary and population genetics studies and for the characterization of bacterial communities. Industrializing the process of DNA-array construction and hybridization is now needed in order to transfer this technology to hospitals and diagnostic laboratories.

Bacteria↗

Ultra high-speed sorting.

BACKGROUND: Cell sorting has a history dating back approximately 40 years. The main limitation has been that, although flow cytometry is a science, cell sorting has been an art during most of this time. Recent advances in assisting technologies have helped to decrease the amount of expertise necessary to perform sorting. METHODS: Droplet-based sorting is based on a controlled disturbance of a jet stream dependent on surface tension. Sorting yield and purity are highly dependent on stable jet break-off position. System pressures and orifice diameters dictate the number of droplets per second, which is the sort rate limiting step because modern electronics can more than handle the higher cell signal processing rates. RESULTS: Cell sorting still requires considerable expertise. Complex multicolor sorting also requires new and more sophisticated sort decisions, especially when cell subpopulations are rare and need to be extracted from background. High-speed sorting continues to pose major problems in terms of biosafety due to the aerosols generated. CONCLUSIONS: Cell sorting has become more stable and predictable and requires less expertise to operate. However, the problems of aerosol containment continue to make droplet-based cell sorting problematical. Fluid physics and cell viability restraints pose practical limits for high-speed sorting that have almost been reached. Over the next 5 years there may be advances in fluidic switching sorting in lab-on-a-chip microfluidic systems that could not only solve the aerosol and viability problems but also make ultra high-speed sorting possible and practical through massively parallel and exponential staging microfluidic architectures.

Aerosols↗

Exploring glycopeptide-resistance in Staphylococcus aureus: a combined proteomics and transcriptomics approach for the identification of resistance-related markers.

BACKGROUND: To unravel molecular targets involved in glycopeptide resistance, three isogenic strains of Staphylococcus aureus with different susceptibility levels to vancomycin or teicoplanin were subjected to whole-genome microarray-based transcription and quantitative proteomic profiling. Quantitative proteomics performed on membrane extracts showed exquisite inter-experimental reproducibility permitting the identification and relative quantification of >30% of the predicted S. aureus proteome. RESULTS: In the absence of antibiotic selection pressure, comparison of stable resistant and susceptible strains revealed 94 differentially expressed genes and 178 proteins. As expected, only partial correlation was obtained between transcriptomic and proteomic results during stationary-phase. Application of massively parallel methods identified one third of the complete proteome, a majority of which was only predicted based on genome sequencing, but never identified to date. Several over-expressed genes represent previously reported targets, while series of genes and proteins possibly involved in the glycopeptide resistance mechanism were discovered here, including regulators, global regulator attenuator, hyper-mutability factor or hypothetical proteins. Gene expression of these markers was confirmed in a collection of genetically unrelated strains showing altered susceptibility to glycopeptides. CONCLUSION: Our proteome and transcriptome analyses have been performed during stationary-phase of growth on isogenic strains showing susceptibility or intermediate level of resistance against glycopeptides. Altered susceptibility had emerged spontaneously after infection with a sensitive parental strain, thus not selected in vitro. This combined analysis allows the identification of hundreds of proteins considered, so far as hypothetical protein. In addition, this study provides not only a global picture of transcription and expression adaptations during a complex antibiotic resistance mechanism but also unravels potential drug targets or markers that are constitutively expressed by resistant strains regardless of their genetic background, amenable to be used as diagnostic targets.

Anti-Bacterial Agents↗

Predicting RNA H-type pseudoknots with the massively parallel genetic algorithm.

MOTIVATION: Using the genetic algorithm (GA) for RNA folding on a massively parallel supercomputer, MasPar MP-2 with 16,384 processors, we successfully predicted the existence of H-type pseudoknots in several sequences. RESULTS: The GA is applied to folding the tRNA-like 3' end of turnip yellow mosaic virus (TYMV) RNA sequence with 82 nucleotides, the 3' UTRs of satellite tobacco necrosis virus (STNV)-2 RNA sequence with 619 nucleotides and STNV-I RNA sequence with 622 nucleotides, and the bacteriophage T2, T4 and T6 gene 32 mRNA sequences with 946, 1340 and 946 nucleotides, respectively. The GA's results match the phylogenetically supported tertiary structures of these sequences.

Algorithms↗