Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

A team quality improvement sequence for complex problems.

To solve complex quality problems teams need to follow a systematic sequence of inquiry and action. In this article a practical description of a team quality improvement sequence (TQIS) is given based on the experience of the more successful teams in the Norwegian total quality management experiment. There are nine phases in the sequence and teams have the flexibility to choose the best quality methods for completing each phase. The strengths of the framework are in ensuring that personnel time is used cost effectively and that changes are made which result in measurable improvement. One limitation is that the framework has not been as widely tested as FOCUS-PDCA (find, organise, clarify, understand, select-plan, do, check, act) and other frameworks to which the TQIS framework is compared. It is proposed that if team projects are to be the main vehicle for quality improvement, then their work must be made more cost effective. The article aims to stimulate research into the conditions necessary for different quality teams to be successful in health care, and draws on the research to propose a "risk of team failure index" to improve the management of such teams.

Norway↗

Capillary DNA sequencing: maximizing the sequence output.

Like most other DNA sequencing core facilities, one of our continuing goals is to improve our sequence output without substantially adding to cost. To minimize sample-to-sample variability in template DNA concentration, we implemented the rolling circle amplification (RCA) procedure for preparing our DNA templates. In addition to saving time and reducing the number of steps in template DNA preparation, the RCA method has the potential to normalize the DNA concentration in samples that can be sequenced directly without additional purification. In the present study, we used RCA-generated templates to test a recently reported procedure that increased sequence quality by resuspending the sequenced products in low concentrations of agarose before capillary electrophoresis (CE) on a MegaBACE 1000 platform. Although we did not obtain the expected result using the specified procedure, a modification resulted in up to 60% increase in total sequence yield per sample plate. A combination of agarose and formamide-EDTA in the resuspension solution enabled us to generate long-read and high-quality sequences for more than 38,000 templates with minimal additional cost.

DNA, Bacterial↗

Protein design by optimization of a sequence-structure quality function.

An automated procedure for protein design by optimization of a sequence-structure quality has been developed. The method selects a statistically optimal sequence for a particular structure, on the assumption that such a protein will adopt the desired structure. We present two optimization algorithms: one provides an exact optimization while the other uses a combinatorial technique for comparatively rapid results. Both are suitable for massively parallel computers. A prototype system was used to design sequences which should adopt the four-helix bundle conformation of myohemerythrin. These appear satisfactory to secondary structure and profile analysis. Detailed inspection reveals that the sequences are generally plausible but, as expected, lack some specific structural features. The design parameters provide some insight into the general determinants of protein structure.

Algorithms↗

Multiple sequence alignments of partially coding nucleic acid sequences.

BACKGROUND: High quality sequence alignments of RNA and DNA sequences are an important prerequisite for the comparative analysis of genomic sequence data. Nucleic acid sequences, however, exhibit a much larger sequence heterogeneity compared to their encoded protein sequences due to the redundancy of the genetic code. It is desirable, therefore, to make use of the amino acid sequence when aligning coding nucleic acid sequences. In many cases, however, only a part of the sequence of interest is translated. On the other hand, overlapping reading frames may encode multiple alternative proteins, possibly with intermittent non-coding parts. Examples are, in particular, RNA virus genomes. RESULTS: The standard scoring scheme for nucleic acid alignments can be extended to incorporate simultaneously information on translation products in one or more reading frames. Here we present a multiple alignment tool, codaln, that implements a combined nucleic acid plus amino acid scoring model for pairwise and progressive multiple alignments that allows arbitrary weighting for almost all scoring parameters. Resource requirements of codaln are comparable with those of standard tools such as ClustalW. CONCLUSION: We demonstrate the applicability of codaln to various biologically relevant types of sequences (bacteriophage Levivirus and Vertebrate Hox clusters) and show that the combination of nucleic acid and amino acid sequence information leads to improved alignments. These, in turn, increase the performance of analysis tools that depend strictly on good input alignments such as methods for detecting conserved RNA secondary structure elements.

Algorithms↗

Natural variation among human adenoviruses: genome sequence and annotation of human adenovirus serotype 1.

The 36,001 base pair DNA sequence of human adenovirus serotype 1 (HAdV-1) has been determined, using a 'leveraged primer sequencing strategy' to generate high quality sequences economically. This annotated genome (GenBank AF534906) confirms anticipated similarity to closely related species C (formerly subgroup), human adenoviruses HAdV-2 and -5, and near identity with earlier reports of sequences representing parts of the HAdV-1 genome. A first round of HAdV-1 sequence data acquisition used PCR amplification and sequencing primers from sequences common to the genomes of HAdV-2 and -5. The subsequent rounds of sequencing used primers derived from the newly generated data. Corroborative re-sequencing with primers selected from this HAdV-1 dataset generated sparsely tiled arrays of high quality sequencing ladders spanning both complementary strands of the HAdV-1 genome. These strategies allow for rapid and accurate low-pass sequencing of genomes. Such rapid genome determinations facilitate the development of specific probes for differentiation of family, serotype, subtype and strain (e.g. pathogen genome signatures). These will be used to monitor epidemic outbreaks of acute respiratory disease in a defined test bed by the Epidemic Outbreak Surveillance (EOS) project.

Adenovirus Infections, Human↗

A study on protein sequence alignment quality.

One of the most central methods in bioinformatics is the alignment of two protein or DNA sequences. However, so far large-scale benchmarks examining the quality of these alignments are scarce. On the other hand, recently several large-scale studies of the capacity of different methods to identify related sequences has led to new insights about the performance of fold recognition methods. To increase our understanding about fold recognition methods, we present a large-scale benchmark of alignment quality. We compare alignments from several different alignment methods, including sequence alignments, hidden Markov models, PSI-BLAST, CLUSTALW, and threading methods. For most methods, the alignment quality increases significantly at about 20% sequence identity. The difference in alignment quality between different methods is quite small, and the main difference can be seen at the exact positioning of the sharp rise in alignment quality, that is, around 15-20% sequence identity. The alignments are improved by using structural information. In general, the best alignments are obtained by methods that use predicted secondary structure information and sequence profiles obtained from PSI-BLAST. One interesting observation is that for different pairs many different methods create the best alignments. This finding implies that if a method that could select the best alignment method for each pair existed, a significant improvement of the alignment quality could be gained.

Computational Biology↗

TargetQC: A targeted quality control framework for clinical genomic testing.

Reliable genetic testing depends on accurate assessment of sequencing quality in clinically relevant genomic regions that directly influence variant interpretation. We developed TargetQC, a flexible quality control framework that supports user-defined gene sets, coverage thresholds, and variant sets for evaluating sequencing performance across exome sequencing (ES) and genome sequencing (GS) platforms. TargetQC assesses exon and gene coverage, identifies regions meeting predefined coverage thresholds, evaluates variant detection accuracy, and measures sequencing quality at pathogenic variant sites. We applied TargetQC to the reference sample NA12878 and 665 clinical samples across five ES platforms and one GS platform. ES-VendorB and ES-VendorE achieved the most complete coverage of OMIM coding regions in NA12878, whereas ES-VendorD and ES-VendorE showed the highest coverage compliance in clinical samples. ES-VendorB and GS demonstrated the highest variant detection accuracy. TargetQC provides a practical framework for benchmarking sequencing performance and informing platform selection in clinical genomics.

exome sequencing↗

Benchmark for evaluating the quality of DNA sequencing: proposal from an international external quality assessment scheme.

BACKGROUND: In the past 15 years, clinical laboratory science has been transformed by the use of technologies that cross the traditional boundaries between laboratory disciplines. However, during this period, issues of quality have not always been given adequate attention. The European Molecular Genetics Quality Network (EMQN) has developed a novel external quality assessment scheme for evaluation of DNA sequencing. We report the results of an international survey of the quality of DNA sequencing among 64 laboratories from 21 countries. METHODS: Current practice for DNA sequence analysis was established by use of an online questionnaire. Participating laboratories were provided with 4 DNA samples of validated genotype. Evaluation of the results included assessing the quality of sequence data, variant genotypes, and mutation nomenclature. To accommodate variations in mutation nomenclature, variants indicated by participants were scored for compliance with 3 acceptable marking schemes. RESULTS: A total of 346 genotypes were analyzed. Of these, 19 (5%) genotyping errors were made. Of these, 10 (53%) were false-negative and 9 (47%) were false-positive results. A further 27 (8%) errors were made in naming mutations. Results were analyzed for 3 indicators of data quality: PHRED quality scores, Quality Read Length, and Quality Read Overlap. Most laboratories produced results of acceptable diagnostic quality as judged by these indicators. The results were used to calculate a consensus benchmark for DNA sequencing against which individual laboratories could rank their performance. CONCLUSIONS: We propose that the consensus benchmark can be used as a baseline against which the aggregate and individual laboratory standard of DNA sequencing may be tracked from year to year.

Benchmarking↗

Finishing a whole-genome shotgun: release 3 of the Drosophila melanogaster euchromatic genome sequence.

BACKGROUND: The Drosophila melanogaster genome was the first metazoan genome to have been sequenced by the whole-genome shotgun (WGS) method. Two issues relating to this achievement were widely debated in the genomics community: how correct is the sequence with respect to base-pair (bp) accuracy and frequency of assembly errors? And, how difficult is it to bring a WGS sequence to the accepted standard for finished sequence? We are now in a position to answer these questions. RESULTS: Our finishing process was designed to close gaps, improve sequence quality and validate the assembly. Sequence traces derived from the WGS and draft sequencing of individual bacterial artificial chromosomes (BACs) were assembled into BAC-sized segments. These segments were brought to high quality, and then joined to constitute the sequence of each chromosome arm. Overall assembly was verified by comparison to a physical map of fingerprinted BAC clones. In the current version of the 116.9 Mb euchromatic genome, called Release 3, the six euchromatic chromosome arms are represented by 13 scaffolds with a total of 37 sequence gaps. We compared Release 3 to Release 2; in autosomal regions of unique sequence, the error rate of Release 2 was one in 20,000 bp. CONCLUSIONS: The WGS strategy can efficiently produce a high-quality sequence of a metazoan genome while generating the reagents required for sequence finishing. However, the initial method of repeat assembly was flawed. The sequence we report here, Release 3, is a reliable resource for molecular genetic experimentation and computational analysis.

Animals↗

preAssemble: a tool for automatic sequencer trace data processing.

BACKGROUND: Trace or chromatogram files (raw data) are produced by automatic nucleic acid sequencing equipment or sequencers. Each file contains information which can be interpreted by specialised software to reveal the sequence (base calling). This is done by the sequencer proprietary software or publicly available programs. Depending on the size of a sequencing project the number of trace files can vary from just a few to thousands of files. Sequencing quality assessment on various criteria is important at the stage preceding clustering and contig assembly. Two major publicly available packages--Phred and Staden are used by preAssemble to perform sequence quality processing. RESULTS: The preAssemble pre-assembly sequence processing pipeline has been developed for small to large scale automatic processing of DNA sequencer chromatogram (trace) data. The Staden Package Pregap4 module and base-calling program Phred are utilized in the pipeline, which produces detailed and self-explanatory output that can be displayed with a web browser. preAssemble can be used successfully with very little previous experience, however options for parameter tuning are provided for advanced users. preAssemble runs under UNIX and LINUX operating systems. It is available for downloading and will run as stand-alone software. It can also be accessed on the Norwegian Salmon Genome Project web site where preAssemble jobs can be run on the project server. CONCLUSION: preAssemble is a tool allowing to perform quality assessment of sequences generated by automatic sequencing equipment. preAssemble is flexible since both interactive jobs on the preAssemble server and the stand alone downloadable version are available. Virtually no previous experience is necessary to run a default preAssemble job, on the other hand options for parameter tuning are provided. Consequently preAssemble can be used as efficiently for just several trace files as for large scale sequence processing.

Algorithms↗

Gene prediction and verification in a compact genome with numerous small introns.

The genomes of clusters of related eukaryotes are now being sequenced at an increasing rate, creating a need for accurate, low-cost annotation of exon-intron structures. In this paper, we demonstrate that reverse transcription-polymerase chain reaction (RT-PCR) and direct sequencing based on predicted gene structures satisfy this need, at least for single-celled eukaryotes. The TWINSCAN gene prediction algorithm was adapted for the fungal pathogen Cryptococcus neoformans by using a precise model of intron lengths in combination with ungapped alignments between the genome sequences of the two closely related Cryptococcus varieties. This approach resulted in approximately 60% of known genes being predicted exactly right at every coding base and splice site. When previously unannotated TWINSCAN predictions were tested by RT-PCR and direct sequencing, 75% of targets spanning two predicted introns were amplified and produced high-quality sequence. When targets spanning the complete predicted open reading frame were tested, 72% of them amplified and produced high-quality sequence. We conclude that sequencing a small number of expressed sequence tags (ESTs) to provide training data, running TWINSCAN on an entire genome, and then performing RT-PCR and direct sequencing on all of its predictions would be a cost-effective method for obtaining an experimentally verified genome annotation.

Algorithms↗

18S ribosomal RNA and tetrapod phylogeny.

Previous phylogenetic analyses of tetrapod 18S ribosomal RNA (rRNA) sequences support the grouping of birds with mammals, whereas other molecular data, and morphological and paleontological data favor the grouping of birds with crocodiles. The 18S rRNA gene has consequently been considered odd, serving as "definitive evidence of different genes providing significantly different estimates of phylogeny in higher organisms" (p. 156; Huelsenbeck et al., 1996, Trends Ecol. Evol. 11:152-158). Our research indicates that the previous discrepancy of phylogenetic results between the 18S rRNA gene and other genes is caused mainly by (1) the misalignment of the sequences, (2) the inappropriate use of the frequency parameters, and (3) poor sequence quality. When the sequences are aligned with the aide of the secondary structure of the 18S rRNA molecule and when the frequency parameters are estimated either from all sites or from the variable domains where substitutions have occurred, the 18S rRNA sequences no longer support the grouping of the avian species with the mammalian species.

Algorithms↗

Heterochromatic sequences in a Drosophila whole-genome shotgun assembly.

BACKGROUND: Most eukaryotic genomes include a substantial repeat-rich fraction termed heterochromatin, which is concentrated in centric and telomeric regions. The repetitive nature of heterochromatic sequence makes it difficult to assemble and analyze. To better understand the heterochromatic component of the Drosophila melanogaster genome, we characterized and annotated portions of a whole-genome shotgun sequence assembly. RESULTS: WGS3, an improved whole-genome shotgun assembly, includes 20.7 Mb of draft-quality sequence not represented in the Release 3 sequence spanning the euchromatin. We annotated this sequence using the methods employed in the re-annotation of the Release 3 euchromatic sequence. This analysis predicted 297 protein-coding genes and six non-protein-coding genes, including known heterochromatic genes, and regions of similarity to known transposable elements. Bacterial artificial chromosome (BAC)-based fluorescence in situ hybridization analysis was used to correlate the genomic sequence with the cytogenetic map in order to refine the genomic definition of the centric heterochromatin; on the basis of our cytological definition, the annotated Release 3 euchromatic sequence extends into the centric heterochromatin on each chromosome arm. CONCLUSIONS: Whole-genome shotgun assembly produced a reliable draft-quality sequence of a significant part of the Drosophila heterochromatin. Annotation of this sequence defined the intron-exon structures of 30 known protein-coding genes and 267 protein-coding gene models. The cytogenetic mapping suggests that an additional 150 predicted genes are located in heterochromatin at the base of the Release 3 euchromatic sequence. Our analysis suggests strategies for improving the sequence and annotation of the heterochromatic portions of the Drosophila and other complex genomes.

Algorithms↗

Amplification and direct sequence analysis of the 23S rRNA gene from thermophilic bacteria.

We present a simplified and fast method to obtain high-quality sequences directly from PCRs without the traditional gel purification. We also report on an improved method to obtain sequence-quality PCR products from microorganisms that are difficult to lyse with no need for DNA extraction. The technique uses exonuclease 1 and shrimp alkaline phosphatase to degrade residual dNTPs and primers. Our technique is shown to work on both Gram-positive and Gram-negative bacteria.

Bacteria↗

Automated identification of single nucleotide polymorphisms from sequencing data.

Single nucleotide polymorphisms (SNPs) provide abundant information about genetic variation. Large scale discovery of high frequency SNPs is being undertaken using various methods. However, the publicly available SNP data are not always accurate, and therefore should be verified. If only a particular gene locus is concerned,locus-specific polymerase chain reaction amplification may be useful. Problem of this method is that the secondary peak has to be measured. We have analyzed trace data from conventional sequencing equipment and found an applicable rule to discern SNPs from noise. We have developed software that integrates this function to automatically identify SNPs. The software works accurately for high quality sequences and also can detect SNPs in low quality sequences. Further, it can determine allele frequency, display this information as a bar graph and assign corresponding nucleotide combinations. It is very useful for identifying de novo SNPs in a DNA fragment of interest.

Algorithms↗

Automated identification of single nucleotide polymorphisms from sequencing data.

The single nucleotide polymorphism (SNP) is the difference of the DNA sequence between individuals and provides abundant information about genetic variation. Large scale discovery of high frequency SNPs is being undertaken using various methods. However, the publicly available SNP data sometimes need to be verified. If only a particular gene locus is concerned, locus-specific polymerase chain reaction amplification may be useful. Problem of this method is that the secondary peak has to be measured. We have analyzed trace data from conventional sequencing equipment and found an applicable rule to discern SNPs from noise. The rule is applied to multiply aligned sequences with a trace and the peak height of the traces are compared between samples. We have developed software that integrates this function to automatically identify SNPs. The software works accurately for high quality sequences and also can detect SNPs in low quality sequences. Further, it can determine allele frequency, display this information as a bar graph and assign corresponding nucleotide combinations. It is also designed for a person to verify and edit sequences easily on the screen. It is very useful for identifying de novo SNPs in a DNA fragment of interest.

Algorithms↗