Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

From mapping to sequencing, post-sequencing and beyond.

The Rice Genome Research Program (RGP) in Japan has been collaborating with the international community in elucidating a complete high-quality sequence of the rice genome. As the pioneer in large-scale analysis of the rice genome, the RGP has successfully established the fundamental tools for genome research such as a genetic map, a yeast artificial chromosome (YAC)-based physical map, a transcript map and a phage P1 artificial chromosome (PAC)/bacterial artificial chromosome (BAC) sequence-ready physical map, which serve as common resources for genome sequencing. Among the 12 rice chromosomes, the RGP is in charge of sequencing six chromosomes covering 52% of the 390 Mb total length of the genome. The contribution of the RGP to the realization of decoding the rice genome sequence with high accuracy and deciphering the genetic information in the genome will have a great impact in understanding the biology of the rice plant that provides a major food source for almost half of the world's population. A high-quality draft sequence (phase 2) was completed in December 2002. Since then, much of the finished quality sequence (phase 3) has become available in public databases. With the completion of sequencing in December 2004, it is expected that the genome sequence would facilitate innovative research in functional and applied genomics. A map-based genome sequence is indispensable for further improvement of current rice varieties and for development of novel varieties carrying agronomically important traits such as high yield potential and tolerance to both biotic and abiotic stresses. In addition to genome sequencing, various related projects have been initiated to generate valuable resources, which could serve as indispensable tools in clarifying the structure and function of the rice genome. These resources have been made available to the scientific community through the Rice Genome Resource Center (RGRC) of the National Institute of Agrobiological Sciences (NIAS) to enable rapid progress in research that will lead to thorough understanding of the rice plant. As the next trend in rice genome research will focus on determining the function of about 40,000-50,000 genes predicted in the genome as well as applying various genomics tools in rice breeding, an unlimited access to rice DNA and seed stocks will provide a broad community of scientists with the necessary materials for formulating new concepts, developing innovative research and making new scientific discoveries in rice genomics.

Centromere↗

PET-Tool: a software suite for comprehensive processing and managing of Paired-End diTag (PET) sequence data.

BACKGROUND: We recently developed the Paired End diTag (PET) strategy for efficient characterization of mammalian transcriptomes and genomes. The paired end nature of short PET sequences derived from long DNA fragments raised a new set of bioinformatics challenges, including how to extract PETs from raw sequence reads, and correctly yet efficiently map PETs to reference genome sequences. To accommodate and streamline data analysis of the large volume PET sequences generated from each PET experiment, an automated PET data process pipeline is desirable. RESULTS: We designed an integrated computation program package, PET-Tool, to automatically process PET sequences and map them to the genome sequences. The Tool was implemented as a web-based application composed of four modules: the Extractor module for PET extraction; the Examiner module for analytic evaluation of PET sequence quality; the Mapper module for locating PET sequences in the genome sequences; and the Project Manager module for data organization. The performance of PET-Tool was evaluated through the analyses of 2.7 million PET sequences. It was demonstrated that PET-Tool is accurate and efficient in extracting PET sequences and removing artifacts from large volume dataset. Using optimized mapping criteria, over 70% of quality PET sequences were mapped specifically to the genome sequences. With a 2.4 GHz LINUX machine, it takes approximately six hours to process one million PETs from extraction to mapping. CONCLUSION: The speed, accuracy, and comprehensiveness have proved that PET-Tool is an important and useful component in PET experiments, and can be extended to accommodate other related analyses of paired-end sequences. The Tool also provides user-friendly functions for data quality check and system for multi-layer data management.

Animals↗

An initial strategy for the systematic identification of functional elements in the human genome by low-redundancy comparative sequencing.

With the recent completion of a high-quality sequence of the human genome, the challenge is now to understand the functional elements that it encodes. Comparative genomic analysis offers a powerful approach for finding such elements by identifying sequences that have been highly conserved during evolution. Here, we propose an initial strategy for detecting such regions by generating low-redundancy sequence from a collection of 16 eutherian mammals, beyond the 7 for which genome sequence data are already available. We show that such sequence can be accurately aligned to the human genome and used to identify most of the highly conserved regions. Although not a long-term substitute for generating high-quality genomic sequences from many mammalian species, this strategy represents a practical initial approach for rapidly annotating the most evolutionarily conserved sequences in the human genome, providing a key resource for the systematic study of human genome function.

Animals↗

End sequence determination from large insert clones using energy transfer fluorescent primers.

Genome mapping strategies depend heavily on confirmatory data of several types to establish overlaps between contiguous stretches of cloned DNA derived from genomic regions. One type of ancillary data that can contribute to establishing these overlaps is DNA sequence data derived from the ends of large (> 30 kb) inserts in genomic clones. This type of data can be difficult to obtain routinely, because large clones are often unstable and microgram quantities of highly purified DNA are required in each sequencing reaction to obtain sufficient signal for accurate base calling and maximum read length. Recently, we have been experimenting with methods to consistently obtain up to 800 bases of high-quality sequence data from the ends of large insert clones using ThermoSequenase DNA polymerase and Energy Transfer fluorescent primers. Our experimental approach and results, described in this paper, indicate that routinely obtaining high-quality sequence data from the ends of large insert genomic clones is feasible. Such data can contribute to the assessment of common regions between large insert clones, to the establishment of conservation of synteny between closely related species, and to the detection of additional contiguous clones.

Chromosome Mapping↗

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score ≥ 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S↗

Octamer-primed sequencing technology: development of primer identification software.

Octamer sequencing technology (OST) is a primer-directed sequencing strategy in which an individual octamer primer is selected from a pre-synthesized octamer primer library and used to sequence a DNA fragment. However, selecting candidate primers from such a library is time consuming and can be a bottleneck in the sequencing process. To accelerate the sequencing process and to obtain high quality sequencing data, a computer program, electronic OST or eOST, was developed to automatically identify candidate primers from an octamer primer library. eOST integrates the base calling software PHRED to provide a quality assessment for target sequences and identifies potential primer binding sites located within a high quality target region. To increase the sequencing success rate, eOST includes a simple dynamic folding algorithm to automatically calculate the free energy and predict the secondary structure within the template in the vicinity of the octamer-binding site. Several parameters were found to be important, including base quality threshold, the window size of the template sequence segment, and the threshold [Delta] G value. OST, coupled with the eOST software, can be used to sequence short DNA fragments or in the finishing assembly stage of large-scale sequencing of genomic DNA.

Algorithms↗

Automated template quantification for DNA sequencing facilities.

The quantification of plasmid DNA by the PicoGreen dye binding assay has been automated, and the effect of quantification of user-submitted templates on DNA sequence quality in a core laboratory has been assessed. The protocol pipets, mixes and reads standards, blanks and up to 88 unknowns, generates a standard curve, and calculates template concentrations. For pUC19 replicates at five concentrations, coefficients of variance were 0.1, and percent errors were from 1% to 7% (n=198). Standard curves with pUC19 DNA were nonlinear over the 1 to 1733 ng/microL concentration range required to assay the majority (98.7%) of user-submitted templates. Over 35,000 templates have been quantified using the protocol. For 1350 user-submitted plasmids, 87% deviated by >or=20% from the requested concentration (500 ng/microL). Based on data from 418 sequencing reactions, quantification of user-submitted templates was shown to significantly improve DNA sequence quality. The protocol is applicable to all types of double-stranded DNA, is unaffected by primer (1 pmol/microL), and is user modifiable. The protocol takes 30 min, saves 1 h of technical time, and costs approximately $0.20 per unknown.

Plasmids↗

Method for 96-well M13 DNA template preparations for large-scale sequencing.

Efficient preparation of DNA templates is an important step in large-scale DNA sequencing. The ensure high-quality sequence data, we have prepared M13 phage DNA templates using a glass fiber-filtration method. We present the adaptation of this protocol to a 96-well format using commercially available filter plates. Two variations are described: one using polyethylene glycol precipitation and a second where the phage particles are disrupted before filtration, thus eliminating the need for precipitation. Using either of these protocols, 96 templates can be prepared in less than 2 h. Sufficient DNA for 1-2 dye primer sequencing reactions is routinely obtained from 1 mL of culture, and the resulting sequence data are of high quality.

Bacteriophage M13↗

From the double-helix to novel approaches to the sequencing of large genomes.

Elucidation of the structure of DNA by Watson and Crick [Nature 171 (1953) 737-738] has led to many crucial molecular experiments, including studies on DNA replication, transcription, physical mapping, and most recently to serious attempts directed toward the sequencing of large genomes [Watson, Science 248 (1990) 44-49]. I am totally convinced of the great importance of the Human Genome Project, and toward achieving this goal I strongly favor 'top-down' approaches consisting of the physical mapping and preparation of contiguous 50-100-kb fragments directly from the genome, followed by their automated sequencing based on the rapid assembly of primers by hexamer ligation together with primer walking. Our 'top-down' procedures totally avoids conventional cloning, subcloning and random sequencing, which are the elements of the present 'bottom-up' procedures. Fragments of 50-100 kb are prepared in sufficient quantities either by in vitro excision with rare-cutting restriction systems (including Achilles' heel cleavage [AC] or the RecA-AC procedures of Koob et al. [Nucleic Acids Res. 20 (1992) 5831-5836]) or by in vivo excision and amplification using the yeast FRT/Flp system or the phage lambda att/Int system. Such fragments, when derived directly from the Escherichia coli genome, are arranged in consecutive order, so that 50 specially constructed strains of E. coli would supply 50 end-to-end arranged approx. 100-kb fragments, which will cover the entire approx. 5-Mb E. coli genome. For the 150-Mb Drosophila melanogaster genome, 1500 of such consecutive 100-kb fragments (supplied by 1500 strains) are required to cover the entire genome. The fragments will be sequenced by the SPEL-6 method involving hexamer ligation [Szybalski, Gene 90 (1990) 177-178; Fresenius J. Anal. Chem. 4 (1992) 343] and primer walking. The 18-mer primers are synthesized in only a few minutes from three contiguous hexamers annealed to the DNA strand to be sequenced when using an over 100-fold excess of hexamers and T4 DNA ligase at room temperature, preferably in the presence of the single-strand-binding (SSB) protein of E. coli. These 18-nt primers are immediately extended by the DNA polymerase, Sequenase 2.0, in the dideoxy sequencing reaction. Very high quality sequencing ladders are obtained for single-stranded DNA or denatured double-stranded approx. 50-kb fragments, as exemplified by phage lambda DNA.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

SNPsFinder--a web-based application for genome-wide discovery of single nucleotide polymorphisms in microbial genomes.

UNLABELLED: Single nucleotide polymorphisms (SNPs) are the most abundant form of genetic variations in closely related microbial species, strains or isolates. Some SNPs confer selective advantages for microbial pathogens during infection and many others are powerful genetic markers for distinguishing closely related strains or isolates that could not be distinguished otherwise. To facilitate SNP discovery in microbial genomes, we have developed a web-based application, SNPsFinder, for genome-wide identification of SNPs. SNPsFinder takes multiple genome sequences as input to identify SNPs within homologous regions. It can also take contig sequences and sequence quality scores from ongoing sequencing projects for SNP prediction. SNPsFinder will use genome sequence annotation if available and map the predicted SNP regions to known genes or regions to assist further evaluation of the predicted SNPs for their functional significance. SNPsFinder can generate PCR primers for all predicted SNP regions according to user's input parameters to facilitate experimental validation. The results from SNPsFinder analysis are accessible through the World Wide Web. AVAILABILITY: The SNPsFinder program is available at http://snpsfinder.lanl.gov/. SUPPLEMENTARY INFORMATION: The user's manual is available at http://snpsfinder.lanl.gov/UsersManual/

Algorithms↗

Marine genomics: a clearing-house for genomic and transcriptomic data of marine organisms.

BACKGROUND: The Marine Genomics project is a functional genomics initiative developed to provide a pipeline for the curation of Expressed Sequence Tags (ESTs) and gene expression microarray data for marine organisms. It provides a unique clearing-house for marine specific EST and microarray data and is currently available at http://www.marinegenomics.org. DESCRIPTION: The Marine Genomics pipeline automates the processing, maintenance, storage and analysis of EST and microarray data for an increasing number of marine species. It currently contains 19 species databases (over 46,000 EST sequences) that are maintained by registered users from local and remote locations in Europe and South America in addition to the USA. A collection of analysis tools are implemented. These include a pipeline upload tool for EST FASTA file, sequence trace file and microarray data, an annotative text search, automated sequence trimming, sequence quality control (QA/QC) editing, sequence BLAST capabilities and a tool for interactive submission to GenBank. Another feature of this resource is the integration with a scientific computing analysis environment implemented by MATLAB. CONCLUSION: The conglomeration of multiple marine organisms with integrated analysis tools enables users to focus on the comprehensive descriptions of transcriptomic responses to typical marine stresses. This cross species data comparison and integration enables users to contain their research within a marine-oriented data management and analysis environment.

Animals↗

Generation and performance of an equine-specific large-scale gene expression microarray.

OBJECTIVE: To create high-quality sequence data for the generation of an equine gene expression microarray and evaluate array performance by use of lipopolysaccharide (LPS) exposure of synoviocytes. SAMPLE POPULATION: Public nucleotide sequence database from Equus caballus and synoviocytes from clinically normal adult horses. PROCEDURE: Computer procurement of equine gene sequences, probe design, and manufacture of an oligomicroarray were performed. Array performance was evaluated by use of patterns for equine synoviocytes in response to LPS. RESULTS: Starting with 18,924 equine gene sequences, 3,098 equine 3' sequences were annotated and met the inclusion criteria for an expression microarray. An equine oligonucleotide expression microarray was created by use of 68,266 of the 25-oligomer probes to uniquely identify each gene. Most genes in the array (68%) were expressed in equine synoviocytes. Repeatability of the array was high (r, > 0.99), and LPS upregulated (> 5-fold change) 84 genes, many of which were inflammatory mediators, and downregulated (> 5-fold change) 14 genes. An initial pattern of gene expression for effects of LPS on synoviocytes consisted of 102 genes. CONCLUSIONS AND CLINICAL RELEVANCE: Use of a computer algorithm to curate an equine sequence database generated high-quality annotated species-specific gene sequences and probe sets for a gene expression oligomicroarray, which was used to document changes in gene expression associated with LPS exposure of equine synoviocytes. The equine public database was expanded from 290 annotated genes to > 3,000 provisionally annotated genes. Similar curation and annotation of public databases could be used to create other species-specific microarrays.

Animals↗

Global Profiling and Analysis of 5' Monophosphorylated mRNA Decay Intermediates.

During RNA turnover, the action of endo- and exo-ribonucleases can yield RNA decay intermediates with specific 5' ends. These RNA decay intermediates have been demonstrated to be the outcome of decapping, microRNA-directed endo-cleavage, or the protected fragments of ribosomes and exon-junction complexes. Therefore, global analysis of RNA decay intermediates can facilitate studies of many RNA decay pathways. In this chapter, we describe a high-throughput sequencing protocol named parallel analysis of RNA ends (PARE), which allows genome-wide profiling of 5' monophosphorylated mRNA decay intermediates from plants or other eukaryotes. Also, we present the tools and scripts necessary for the proper analysis of RNA degradome data obtained from the PARE method. Details and modifications of library construction procedures and bioinformatic analyses to optimize sequencing quality and cope with emerging sequencing platforms and findings are highlighted.

RNA Stability↗

Novel retinal genes discovered by mining the mouse embryonic RetinalExpress database.

PURPOSE: Bioinformatics has emerged as a powerful tool for identifying novel genes and pathways associated with retinal biology and disease. The developing mouse retina expresses an exceedingly large and complex variety of genes. Many of these genes have not been characterized but nevertheless are likely to have important developmental or physiological functions. The purpose of this study was to use an in silico approach with a mouse embryonic retinal database of cDNAs/expressed sequence tags (ESTs) named RetinalExpress to identify previously uncharacterized genes that are represented in the developing retina. METHODS: cDNA clones unique to the RetinalExpress database were identified by comparing clones in the RetinalExpress database with those in other cDNA/EST databases. We used a hierarchical filtering procedure with high stringency criteria that included sequence quality, colinearity with hypothetical gene sequences, and absence of any substantial existing annotation to select clones that were likely to represent novel genes. Selected clones were located on mouse chromosomes using National Center for Biotechnology Informatics Map Viewer software and the database from the University of California at Santa Cruz Genome Bioinformatics Web browser. The expression of selected retinal transcripts was determined using reverse transcriptase (RT)-PCR. In situ hybridization of sectioned embryonic and postnatal retinas was performed to determine spatial expression patterns of selected transcripts. RESULTS: Of the 27,765 cDNA clones from RetinalExpress that we filtered through several public cDNA/EST databases, 26 cDNA/EST sequences were identified that, at the time of the analysis, were unique to RetinalExpress. Seventeen clones were selected for RT-PCR analysis, and retinal transcripts corresponding to previously uncharacterized genes were unambiguously detected for six clones. Three genes encoded open reading frames containing putative functional domains; one sequence contained an HMG DNA binding domain, another, an RFX DNA binding domain, and another, a phospholipase C catalytic domain X. Transcripts from the genes encoding DNA binding domains were expressed in embryonic and postnatal retinas with distinct spatial patterns. CONCLUSIONS: The characterization of 26 mouse genes whose partial nucleotide sequences were uniquely represented in the RetinalExpress cDNA/EST database demonstrated the feasibility of retinal gene discovery using in silico analysis. Two of these genes had distinctive spatial expression patterns in the retina and one was likely to function as a DNA binding protein in embryonic and postnatal retinas. The gene identification approach described here demonstrates the usefulness of establishing large cDNA/EST databases from highly specialized neuronal tissues such as the retina to find novel genes.

Animals↗

Relationship between multiple sequence alignments and quality of protein comparative models.

Comparative modeling is the method of choice, whenever applicable, for protein structure prediction, not only because of its higher accuracy compared to alternative methods, but also because it is possible to estimate a priori the quality of the models that it can produce, thereby allowing the usefulness of a model for a given application to be assessed beforehand. By and large, the quality of a comparative model depends on two factors: the extent of structural divergence between the target and the template and the quality of the sequence alignment between the two protein sequences. The latter is usually derived from a multiple sequence alignment (MSA) of as many proteins of the family as possible, and its accuracy depends on the number and similarity distribution of the sequences of the protein family. Here we describe a method to evaluate the expected difficulty, and by extension accuracy, of a comparative model on the basis of the MSA used to build it. The parameter that we derive is used to compare the results obtained in the last two editions of the Critical Assessment of Methods for Structure Prediction (CASP) experiment as a function of the difficulty of the modeling exercise. Our analysis demonstrates that the improvement in the scope and quality of comparative models between the two experiments is largely due to the increased number of available protein sequences and to the consequent increased chance that a large and appropriately spaced set of protein sequences homologous to the proteins of interest is available.

Amino Acid Sequence↗

Wurst: a protein threading server with a structural scoring function, sequence profiles and optimized substitution matrices.

Wurst is a protein threading program with an emphasis on high quality sequence to structure alignments (http://www.zbh.uni-hamburg.de/wurst). Submitted sequences are aligned to each of about 3000 templates with a conventional dynamic programming algorithm, but using a score function with sophisticated structure and sequence terms. The structure terms are a log-odds probability of sequence to structure fragment compatibility, obtained from a Bayesian classification procedure. A simplex optimization was used to optimize the sequence-based terms for the goal of alignment and model quality and to balance the sequence and structural contributions against each other. Both sequence and structural terms operate with sequence profiles.

Amino Acid Substitution↗

Computer-based methods for the mouse full-length cDNA encyclopedia: real-time sequence clustering for construction of a nonredundant cDNA library.

We developed computer-based methods for constructing a nonredundant mouse full-length cDNA library. Our cDNA library construction process comprises assessment of library quality, sequencing the 3' ends of inserts and clustering, and completing a re-array to generate a nonredundant library from a redundant one. After the cDNA libraries are generated, we sequence the 5' ends of the inserts to check the quality of the library; then we determine the sequencing priority of each library. Selected libraries undergo large-scale sequencing of the 3' ends of the inserts and clustering of the tag sequences. After clustering, the nonredundant library is constructed from the original libraries, which have redundant clones. All libraries, plates, clones, sequences, and clusters are uniquely identified, and all information is saved in the database according to this identifier. At press time, our system has been in place for the past two years; we have clustered 939,725 3' end sequences into 127,385 groups from 227 cDNA libraries/sublibraries (see http://genome.gse.riken.go.jp/).

5' Untranslated Regions↗