Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

A new solid-phase chemical DNA sequencing method which uses streptavidin-coated magnetic beads.

To simplify the chemical DNA sequencing protocol, we developed a new solid-phase method which uses streptavidin-coated magnetic beads. This method is based on the finding that the biotinylated DNA-streptavidin complex was stable under the conditions for some chemical sequencing reactions. The 5'-biotinylated DNA generated by the polymerase chain reaction was first captured by streptavidin-coated magnetic beads and then subjected to a set of simplified chemical sequencing reactions on the beads at room temperature. Followed by the piperidine cleavage reaction, the products were resolved by gel electrophoresis, transferred onto a nylon membrane and visualized by chemiluminescent detection. As a consequence, high-quality sequencing ladders were obtained, due to complete removal of contaminating chemicals, without the time-consuming precipitation/centrifugation steps used in the conventional chemical sequencing protocol.

Bacterial Proteins↗

ATAC-seq in Emerging Model Organisms: Challenges and Strategies.

The Assay for Transposase-Accessible Chromatin with sequencing (ATAC-seq) is a versatile and widely utilized method for identifying potential regulatory regions, such as promoters and enhancers, within a genome. ATAC-seq has been successfully applied to a wide range of established and emerging model organisms. However, implementing this method in emerging model systems, such as arthropods, can be challenging due to several factors that influence data quality. These factors include the availability of a sufficient amount and quality of tissue or cells, the need for species- and tissue-specific protocol optimization, the completeness and accuracy of the reference genome, and the quality of the genome annotation. In this article, we emphasize the key steps in the ATAC-seq protocol that, based on our experience, have the greatest impact on data quality when adapting this method for emerging model organisms. Specifically, we discuss the importance of nuclei isolation, the incubation conditions of the Tn5 transposase, and PCR amplification of the library. Furthermore, we outline essential quality checkpoints during the bioinformatic analysis of ATAC-seq data to assist in assessing data integrity and consistency. Given that many emerging model organisms may not be readily available in laboratory cultures, we also emphasize the importance of evaluating how different preservation methods affect ATAC-seq data quality. Based on examples in one spider and one ant species, we demonstrate that replication and thorough quality controls at all steps of the protocol and data analysis are essential to assess the usability of ATAC-seq data. Our data highlights the importance of isolating the right number of intact nuclei, as well as ensuring optimal amplification conditions during library preparation to obtain good-quality sequence data for downstream analyses. We recommend using fresh tissue samples if possible because we show that direct cryopreservation of the tissue may affect chromatin integrity. This effect could be avoided or reduced by preserving the homogenate in cell culture medium. Overall, we explain the ATAC-seq protocol and downstream analyses in detail and give step-by-step advice to researchers who are new to the field and want to implement this method. With careful planning and validation, ATAC-seq can reveal the regulatory landscape of a genome and aid in identifying elements that govern gene expression.

Animals↗

Mulan: multiple-sequence local alignment and visualization for studying function and evolution.

Multiple-sequence alignment analysis is a powerful approach for understanding phylogenetic relationships, annotating genes, and detecting functional regulatory elements. With a growing number of partly or fully sequenced vertebrate genomes, effective tools for performing multiple comparisons are required to accurately and efficiently assist biological discoveries. Here we introduce Mulan (http://mulan.dcode.org/), a novel method and a network server for comparing multiple draft and finished-quality sequences to identify functional elements conserved over evolutionary time. Mulan brings together several novel algorithms: the TBA multi-aligner program for rapid identification of local sequence conservation, and the multiTF program for detecting evolutionarily conserved transcription factor binding sites in multiple alignments. In addition, Mulan supports two-way communication with the GALA database; alignments of multiple species dynamically generated in GALA can be viewed in Mulan, and conserved transcription factor binding sites identified with Mulan/multiTF can be integrated and overlaid with extensive genome annotation data using GALA. Local multiple alignments computed by Mulan ensure reliable representation of short- and large-scale genomic rearrangements in distant organisms. Mulan allows for interactive modification of critical conservation parameters to differentially predict conserved regions in comparisons of both closely and distantly related species. We illustrate the uses and applications of the Mulan tool through multispecies comparisons of the GATA3 gene locus and the identification of elements that are conserved in a different way in avians than in other genomes, allowing speculation on the evolution of birds. Source code for the aligners and the aligner-evaluation software can be freely downloaded from http://www.bx.psu.edu/miller_lab/.

Animals↗

Targeted oligonucleotide-mediated microsatellite identification (TOMMI) from large-insert library clones.

BACKGROUND: In the last few years, microsatellites have become the most popular molecular marker system and have intensively been applied in genome mapping, biodiversity and phylogeny studies of livestock. Compared to single nucleotide polymorphism (SNP) as another popular marker system, microsatellites reveal obvious advantages. They are multi-allelic, possibly more polymorphic and cheaper to genotype. Calculations showed that a multi-allelic marker system always has more power to detect Linkage Disequilibrium (LD) than does a di-allelic marker system. Traditional isolation methods using partial genomic libraries are time-consuming and cost-intensive. In order to directly generate microsatellites from large-insert libraries a sequencing approach with repeat-containing oligonucleotides is introduced. RESULTS: Seventeen porcine microsatellite markers were isolated from eleven PAC clones by targeted oligonucleotide-mediated microsatellite identification (TOMMI), an improved efficient and rapid flanking sequence-based approach for the isolation of STS-markers. With the application of TOMMI, an average of 1.55 (CA/GT) microsatellites per PAC clone was identified. The number of alleles, allele size distribution, polymorphism information content (PIC), average heterozygosity (HT), and effective allele number (NE) for the STS-markers were calculated using a sampling of 336 unrelated animals representing fifteen pig breeds (nine European and six Chinese breeds). Sixteen of the microsatellite markers proved to be polymorphic (2 to 22 alleles) in this heterogeneous sampling. Most of the publicly available (porcine) microsatellite amplicons range from approximately 80 bp to 200 bp. Here, we attempted to utilize as much sequence information as possible to develop STS-markers with larger amplicons. Indeed, fourteen of the seventeen STS-marker amplicons have minimal allele sizes of at least 200 bp. Thus, most of the generated STS-markers can easily be integrated into multilocus assays covering a broader separation spectrum. Linkage mapping results of the markers indicate their potential immediate use in QTL studies to further dissect trait associated chromosomal regions. CONCLUSION: The sequencing strategy described in this study provides a targeted, inexpensive and fast method to develop microsatellites from large-insert libraries. It is well suited to generate polymorphic markers for selected chromosomal regions, contigs of overlapping clones and yields sufficient high quality sequence data to develop amplicons greater than 250 bases.

Alleles↗

Active deep brain stimulation during MRI: a feasibility study.

The goal of this study was to evaluate the feasibility of active deep brain stimulation (DBS) during the application of standard clinical sequences for functional MRI (fMRI) in phantom measurements. During active DBS, we investigated induced voltage, temperature at the electrode tips and lead, forces on the electrode and lead, consequences of defective leads and loose connections, proper operation of the neurostimulator, and image quality. Sequences for diffusion- and perfusion-weighted imaging, fMRI, and morphologic MRI were used. The DBS electrode and lead were placed in a NaCl solution-filled phantom. The results indicate that there are severe potential hazards for patients. Strong heating, high induced voltage, and even sparking at defects in the connecting cable could be observed. However, it was demonstrated that under certain conditions, safe MR examinations during active DBS are feasible. Certain safety precautions are recommended in this report.

Body Temperature↗

PEDE (Pig EST Data Explorer): construction of a database for ESTs derived from porcine full-length cDNA libraries.

We generated the PEDE (Pig EST Data Explorer; http://pede.dna.affrc.go.jp/) database using sequences assembled from porcine 5' ESTs from oligo-capped full-length cDNA libraries. Thus far we have performed EST analysis of various organs (thymus, spleen, uterus, lung, liver, ovary and peripheral blood mononuclear cells) and assembled 68,076 high-quality sequences into 5546 contigs and 28,461 singlets. PEDE provides a search interface for getting results of homology searches and enables users to obtain information on sequence data and cDNA clones of interest. Single-nucleotide polymorphisms detected through comparison of the EST sequences are classified by origin (western and oriental breeds) and are searchable in the database. This database system can accelerate analyses of livestock traits and yields information that can lead to new applications in pigs as model systems for medical research.

Animals↗

Genomic DNA sequencing methods.

Sequence analysis of cosmids from C. elegans and other organisms currently is best done using the random or "shotgun" strategy (Wilson et al., 1994). After shearing by sonication, DNA is used to prepare M13 subclone libraries which provide good coverage and high-quality sequence data. The subclones are assembled and the data edited using software tools developed especially for C. elegans genomic sequencing. These same tools facilitate much of the subsequent work to complete both strands of the sequence and resolve any remaining ambiguities. Analysis of the finished sequence is then accomplished using several additional computer tools including Genefinder and ACeDB. Taken together, these methods and tools provide a powerful means for genome analysis in the nematode.

Animals↗

Transition of Staphylococcus aureus tetracycline resistance plasmid pT181 from independent multicopy replicon to predominantly integrated chromosomal element over 65 years.

Mobile genetic elements (MGEs), including plasmids, phages and genome islands, are major sources of bacterial genetic diversity. The small plasmid pT181 confers tetracycline resistance in bacterial pathogen Staphylococcus aureus via an efflux pump, TetK. pT181 was one of the earliest sequenced S. aureus plasmids, and has been isolated in both clinical and livestock-associated strains for decades, both as an independent replicon and integrated in the chromosome as part of staphylococcal cassette chromosome mec (SCCmec). Bacterial genome analysis tools and high-quality sequences with metadata are publicly available, but these resources remain underleveraged for examining historical data, especially when studying the spread of MGEs across a species and over time. Using publicly available reads and metadata, we explored the evolution of pT181 over almost seven decades of samples to identify temporal trends in sequence evolution, copy number changes, and spread across S. aureus and beyond. pT181 was prevalent across S. aureus (found in 9.5% of 83,366 genomes tested), with a conserved sequence outside of three hypervariable regions. The history of pT181 since 1954 is characterized by spread across strains, significant variation in plasmid copy number of the independent replicon, and increasing frequency of integration of the plasmid into the S. aureus chromosome. We have identified multiple chromosomal integration locations of the plasmid, including outside of the previously characterized SCCmec. We find that pT181 has been transferred across staphylococcaceae and into a Gram-negative species. The repeated integration of pT181 into the chromosome may indicate co-evolution of the plasmid and the host, potentially to facilitate increased antibiotic resistance.

Journal Article↗

Increased detection of structural templates using alignments of designed sequences.

Protein structure prediction by comparative modeling benefits greatly from the use of multiple sequence alignment information to improve the accuracy of structural template identification and the alignment of target sequences to structural templates. Unfortunately, this benefit is limited to those protein sequences for which at least several natural sequence homologues exist. We show here that the use of large diverse alignments of computationally designed protein sequences confers many of the same benefits as natural sequences in identifying structural templates for comparative modeling targets. A large-scale massively parallelized application of an all-atom protein design algorithm, including a simple model of peptide backbone flexibility, has allowed us to generate 500 diverse, non-native, high-quality sequences for each of 264 protein structures in our test set. PSI-BLAST searches using the sequence profiles generated from the designed sequences ("reverse" BLAST searches) give near-perfect accuracy in identifying true structural homologues of the parent structure, with 54% coverage. In 41 of 49 genomes scanned using reverse BLAST searches, at least one novel structural template (not found by the standard method of PSI-BLAST against PDB) is identified. Further improvements in coverage, through optimizing the scoring function used to design sequences and continued application to new protein structures beyond the test set, will allow this method to mature into a useful strategy for identifying distantly related structural templates.

Algorithms↗

ExoMeth sequencing of DNA: eliminating the need for subcloning and oligonucleotide primers.

A method is reported for sequencing DNA based on exonuclease III digestion and strand protection by using modified nucleoside triphosphates. Up to 10 kilobases of sequence information may be obtained from each strand of a given template without subcloning. Prior knowledge of the restriction map is not important; prior knowledge of any of the sequence is not required. Nor are oligonucleotide primers needed. Double-stranded cosmids, plasmids, lambda phage, or linear molecules (including amplified molecules) may be used as starting material. The method creates a single-stranded template from these starting molecules, thus generating high-quality sequence ladders. Most commonly used DNA polymerases may be utilized, including reverse transcriptase and T7 DNA polymerase. The approach is "ordered", so little time is wasted on redundant sequencing.

Base Sequence↗

An integrated computational pipeline and database to support whole-genome sequence annotation.

We describe here our experience in annotating the Drosophila melanogaster genome sequence, in the course of which we developed several new open-source software tools and a database schema to support large-scale genome annotation. We have developed these into an integrated and reusable software system for whole-genome annotation. The key contributions to overall annotation quality are the marshalling of high-quality sequences for alignments and the design of a system with an adaptable and expandable flexible architecture.

Animals↗

Assign 2.0: software for the analysis of Phred quality values for quality control of HLA sequencing-based typing.

As improvements to DNA sequencing technology have resulted in increasing the throughput of DNA sequencing, the bottleneck for high throughput DNA sequencing-based typing (SBT) has shifted to sequence analysis, genotyping and quality control (QC). Consistent high-quality DNA sequence is required in order to reduce manual verification and editing of sequence electropherograms. However, identifying systematic changes in quality is difficult to achieve without the aid of sophisticated sequence analysis programs dedicated to this purpose. We describe a computer software program called Assign 2.0, which integrates sequence QC analysis and genotyping in order to facilitate high-throughput SBT. Assign 2.0 performs an analysis of Phred quality values in order to produce quality scores for a sample and a sequencing run. This enables sample-to-sample and run-to-run QC monitoring and provides a mechanism for the comparison of sequence quality between various genes, various reagents and various protocols with the aim of improving the overall quality of DNA sequence data. This, in turn, will result in reducing sequence analysis as a bottleneck for high-throughput SBT.

Alleles↗

CR-EST: a resource for crop ESTs.

The crop expressed sequence tag database, CR-EST (http://pgrc.ipk-gatersleben.de/cr-est/), is a publicly available online resource providing access to sequence, classification, clustering and annotation data of crop EST projects. CR-EST currently holds more than 200,000 sequences derived from 41 cDNA libraries of four species: barley, wheat, pea and potato. The barley section comprises approximately one-third of all publicly available ESTs. CR-EST deploys an automatic EST preparation pipeline that includes the identification of chimeric clones in order to transparently display the data quality. Sequences are clustered in species-specific projects to currently generate a non-redundant set of approximately 22,600 consensus sequences and approximately 17,200 singletons, which form the basis of the provided set of unigenes. A web application allows the user to compute BLAST alignments of query sequences against the CR-EST database, query data from Gene Ontology and metabolic pathway annotations and query sequence similarities from stored BLAST results. CR-EST also features interactive JAVA-based tools, allowing the visualization of open reading frames and the explorative analysis of Gene Ontology mappings applied to ESTs.

Crops, Agricultural↗

Sequencing and analysis of common bean ESTs. Building a foundation for functional genomics.

Although common bean (Phaseolus vulgaris) is the most important grain legume in the developing world for human consumption, few genomic resources exist for this species. The objectives of this research were to develop expressed sequence tag (EST) resources for common bean and assess nodule gene expression through high-density macroarrays. We sequenced a total of 21,026 ESTs derived from 5 different cDNA libraries, including nitrogen-fixing root nodules, phosphorus-deficient roots, developing pods, and leaves of the Mesoamerican genotype, Negro Jamapa 81. The fifth source of ESTs was a leaf cDNA library derived from the Andean genotype, G19833. Of the total high-quality sequences, 5,703 ESTs were classified as singletons, while 10,078 were assembled into 2,226 contigs producing a nonredundant set of 7,969 different transcripts. Sequences were grouped according to 4 main categories, metabolism (34%), cell cycle and plant development (11%), interaction with the environment (19%), and unknown function (36%), and further subdivided into 15 subcategories. Comparisons to other legume EST projects suggest that an entirely different repertoire of genes is expressed in common bean nodules. Phaseolus-specific contigs, gene families, and single nucleotide polymorphisms were also identified from the EST collection. Functional aspects of individual bean organs were reflected by the 20 contigs from each library composed of the most redundant ESTs. The abundance of transcripts corresponding to selected contigs was evaluated by RNA blots to determine whether gene expression determined by laboratory methods correlated with in silico expression. Evaluation of root nodule gene expression by macroarrays and RNA blots showed that genes related to nitrogen and carbon metabolism are integrated for ureide production. Resources developed in this project provide genetic and genomic tools for an international consortium devoted to bean improvement.

Carbon↗

Cloning and sequence analysis of homeobox transcription factor cDNAs with an inosine-containing probe.

Much effort has been directed toward the isolation and characterization of homeobox cDNAs from numerous cell types because they encode transcription factors important to many cellular processes, including pattern formation in the embryo, cell growth and cell differentiation. Many novel homeobox cDNAs have been isolated by screening libraries by hybridization with degenerate oligonucleotides designed from conserved amino acid sequences in the third helix of the homeodomain. However, the degeneracy of the genetic code necessitates that these oligonucleotides be highly degenerate, often precluding their use as sequencing primers to rapidly determine clone identity. Here we describe a screening protocol for homeobox cDNAs that utilizes a short oligonucleotide probe with inosine residues incorporated at positions of maximum codon degeneracy. This probe specifically hybridizes to many classes of homeobox transcription factor cDNAs, but its primary advantage is that it also serves as an effective sequencing primer, which allows the investigator to rapidly determine whether the clones encode a protein of interest. In a screen of 500,000 plaques of a rat aorta cDNA library by this method, we identified 13 positive plaques of which 12 were found to contain homeobox cDNAs representing 5 distinct genes, and, using this probe, it was possible to obtain initial high-quality sequence information from every clone isolated that contained a homeodomain.

Animals↗

New approaches to the analysis of palindromic sequences from the human genome: evolution and polymorphism of an intronic site at the NF1 locus.

The nature of any long palindrome that might exist in the human genome is obscured by the instability of such sequences once cloned in Escherichia coli. We describe and validate a practical alternative to the analysis of naturally-occurring palindromes based upon cloning and propagation in Saccharomyces cerevisiae. With this approach we have investigated an intronic sequence in the human Neurofibromatosis 1 (NF1) locus that is represented by multiple conflicting versions in GenBank. We find that the site is highly polymorphic, exhibiting different degrees of palindromy in different individuals. A side-by-side comparison of the same plasmids in E.coli versus. S.cerevisiae demonstrated that the more palindromic alleles were inevitably corrupted upon cloning in E.coli, but could be propagated intact in yeast. The high quality sequence obtained from the yeast-based approach provides insight into the various mechanisms that destabilize a palindrome in E.coli, yeast and humans, into the diversification of a highly polymorphic site within the NF1 locus during primate evolution, and into the association between palindromy and chromosomal translocation.

Alleles↗

A novel phosphoramidite method for automated synthesis of oligonucleotides on glass supports for biosensor development.

Two protocols for functionalization of glass supports with hexaethylene glycol (HEG)-linked oligonucleotides were developed. The first method (standard amidite protocol) made use of the 2-cyanoethyl-phosphoramidite derivative of 4,4'-dimethoxytrityl-protected HEG. This was first coupled to the support by standard solid-phase phosphoramidite chemistry followed by extension with a thymidylic acid icosanucleotide. Stepwise addition of the linker phosphoramidite graduated at 1% (relative to the total sites available) per step at 50 degrees C resulted in an optimal yield of immobilized oligonucleotides at a density of 2.24 x 10(10) strands/mm2. This observed loading maximum lies well below the theoretical maximum loading owing to nonspecific adsorption of HEG on the glass and subsequent blocking of reactive sites. Surface loadings as high as 3.73 x 10(10)/mm2 and of excellent sequence quality were achieved with a reverse amidite protocol. The support was first modified into a 2-cyanoethyl-N,N-diisopropylphosphoramidite analog followed by coupling with 4,4'-dimethoxytrityl-protected HEG. This protocol is conveniently available when using a conventional DNA synthesizer. The reverse amidite protocol allowed for control of the surface loading at values suitable for subsequent analytical applications that make use of immobilized oligonucleotides as probes for selective hybridization of sample nucleic acids of unknown sequence and concentration.

Amides↗

Paradigms for computational nucleic acid design.

The design of DNA and RNA sequences is critical for many endeavors, from DNA nanotechnology, to PCR-based applications, to DNA hybridization arrays. Results in the literature rely on a wide variety of design criteria adapted to the particular requirements of each application. Using an extensively studied thermodynamic model, we perform a detailed study of several criteria for designing sequences intended to adopt a target secondary structure. We conclude that superior design methods should explicitly implement both a positive design paradigm (optimize affinity for the target structure) and a negative design paradigm (optimize specificity for the target structure). The commonly used approaches of sequence symmetry minimization and minimum free-energy satisfaction primarily implement negative design and can be strengthened by introducing a positive design component. Surprisingly, our findings hold for a wide range of secondary structures and are robust to modest perturbation of the thermodynamic parameters used for evaluating sequence quality, suggesting the feasibility and ongoing utility of a unified approach to nucleic acid design as parameter sets are refined further. Finally, we observe that designing for thermodynamic stability does not determine folding kinetics, emphasizing the opportunity for extending design criteria to target kinetic features of the energy landscape.

Algorithms↗