Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

The rice mitochondrial genomes and their variations.

Based on highly redundant and high-quality sequences, we assembled rice (Oryza sativa) mitochondrial genomes for two cultivars, 93-11 (an indica variety) and PA64S (an indica-like variety with maternal origin of japonica), which are paternal and maternal strains of an elite superhybrid rice Liang-You-Pei-Jiu (LYP-9), respectively. Following up with a previous analysis on rice chloroplast genomes, we divided mitochondrial sequence variations into two basic categories, intravarietal and intersubspecific. Intravarietal polymorphisms are variations within mitochondrial genomes of an individual variety. Intersubspecific polymorphisms are variations between subspecies among their major genotypes. In this study, we identified 96 single nucleotide polymorphisms (SNPs), 25 indels, and three segmental sequence variations as intersubspecific polymorphisms. A signature sequence fragment unique to indica varieties was confirmed experimentally and found in two wild rice samples, but absent in japonica varieties. The intersubspecific polymorphism rate for mitochondrial genomes is 0.02% for SNPs and 0.006% for indels, nearly 2.5 and 3 times lower than that of their chloroplast counterparts and 21 and 38 times lower than corresponding rates of the rice nuclear genome, respectively. The intravarietal polymorphism rates among analyzed mitochondrial genomes, such as 93-11 and PA64S, are 1.26% and 1.38% for SNPs and 1.13% and 1.09% for indels, respectively. Based on the total number of SNPs between the two mitochondrial genomes, we estimate that the divergence of indica and japonica mitochondrial genomes occurred approximately 45,000 to 250,000 years ago.

Base Sequence↗

Identification of Escherichia coli O172 O-antigen gene cluster and development of a serogroup-specific PCR assay.

AIM: To characterize the locus for O-antigen biosynthesis from Escherichia coli O172 type strain and to develop a rapid, specific and sensitive PCR-based method for identification and detection of E. coli O172. METHODS AND RESULTS: DNA of O-antigen gene cluster of E. coli O172 was amplified by long-range PCR method using primers based on housekeeping genes galF and gnd Shot gun bank was constructed and high quality sequencing was performed. The putative genes for synthesis of UDP-FucNAc, O-unit flippase, O-antigen polymerase and glycosyltransferases were assigned by the homology search. The evolutionary relationship between O-antigen gene clusters of E. coli O172 and E. coli O26 is shown by sequence comparison. Genes specific to E. coli O172 strains were identified by PCR assays using primers based on genes for O-unit flippase, O-antigen polymerase and glycosyltransferases. The specificity of PCR assays was tested using all E. coli and Shigella O-antigen type strains, as well as 24 clinical E. coli isolates. The sensitivity of PCR assays was determined, and the detection limits were 1 pg microl(-1) chromosomal DNA, 0.2 CFU g(-1) pork and 0.2 CFU ml(-1) water. The total time required from beginning to end of the procedure was within 16 h. CONCLUSION: The O-antigen gene cluster of E. coli O172 was identified and PCR assays based on O-antigen specific genes showed high specificity and sensitivity. SIGNIFICANCE AND IMPACT OF THE STUDY: An O-antigen gene cluster was identified by sequencing. The specific genes were determined for E. coli O172. The sensitivity of O-antigen specific PCR assay was tested. Although Shiga toxin-producing O172 strains were not yet isolated from clinical specimens, they may emerge as pathogens.

Base Sequence↗

Human BAC ends quality assessment and sequence analyses.

End sequences from bacterial artificial chromosomes (BACs) provide highly specific sequence markers in large-scale sequencing projects. To date, we have generated >300,000 end sequences from >186,000 human BAC clones with an average read length of >460 bp for a total of 141 Mb covering approximately 4.7% of the genome. Over 60% of the clones have BAC end sequences (BESs) from both ends representing more than fivefold coverage of the human genome by the paired-end clones. Our quality assessments and sequence analyses indicate that BESs from human BAC libraries developed at The California Institute of Technology (CalTech) and Roswell Park Cancer Institute have similar properties. The analyses have highlighted differences in insert size for different segments of the CalTech library. Problems with the fidelity of tracking of sequence data back to physical clones have been observed in some subsets of the overall BES dataset. The annotation results of BESs for the contents of available genomic sequences, sequence tagged sites, expressed sequence tags, protein encoding regions, and repeats indicate that this resource will be valuable in many areas of genome research.

Chromosome Mapping↗

Improvement of technical and analytical performance in DNA sequencing by external quality assessment-based molecular training.

BACKGROUND: From 2003 to 2005, the European Union supported the EQUAL-initiative to develop methodological external quality assessment (EQA) schemes for genotyping (EQUALqual), quantitative PCR (EQUALquant), and sequencing (EQUALseq). As a relevant part of the EQUALseq program, a training course was held subsequent to the first EQA Program (EQAP1). The success of this course was reassessed in a 2nd EQUALseq round (EQAP2). METHODS: In September 2005, a 3-day training course took place. We invited 8 laboratories with below-average performance in EQAP1 to improve their methodological and analytical/proficiency skills by lectures and practical work. To compare the results of the pretraining and posttraining EQUALseq rounds, we distributed 2 samples used in the first EQUAL round, but this time we provided different oligonucleotide sets. We evaluated the results by means of a previously described scoring system. RESULTS: In EQAP2, 6 laboratories returned complete data sets, corresponding to an overall 14% of the 43 laboratories that had finished EQAP1. The scoring results for samples A (P=0.0025) and B (P=0.0125) demonstrated a significant improvement in EQAP2. Overall, a substantial improvement of technical and interpretative skills was demonstrated (P=0.0051). In general, the workshop experience was highly rated by the participants. CONCLUSIONS: Methodologic EQAPs in DNA sequencing are appropriate tools to uncover strengths and weaknesses in both technique and proficiency, emphasizing the need for mandatory EQAPs. Training courses, together with 2nd-round reiterations, should be implemented into methodological EQAPs in molecular diagnostics to improve technical performance and proficiency in genetic testing.

European Union↗

Development and evaluation of a quality-controlled ribosomal sequence database for 16S ribosomal DNA-based identification of Staphylococcus species.

To establish an improved ribosomal gene sequence database as part of the Ribosomal Differentiation of Microorganisms (RIDOM) project and to overcome the drawbacks of phenotypic identification systems and publicly accessible sequence databases, both strands of the 5' end of the 16S ribosomal DNA (rDNA) of 81 type and reference strains comprising all validly described staphylococcal (sub)species were sequenced. Assuming a normal distribution for pairwise distances of all unique staphylococcal sequences and choosing a reporting criterion of > or =98.7% similarity for a "distinct species," a statistical error probability of 1.0% was calculated. To evaluate this database, a 16S rDNA fragment (corresponding to Escherichia coli positions 54 to 510) of 55 clinical Staphylococcus isolates (including those of the small-colony variant phenotype) were sequenced and analyzed by the RIDOM approach. Of these isolates, 54 (98.2%) had a similarity score above the proposed threshold using RIDOM; 48 (87.3%) of the sequences gave a perfect match, whereas 83.6% were found by searching National Center for Biotechnology Information (NCBI) database entries. In contrast to RIDOM, which showed four ambiguities at the species level (mainly concerning Staphylococcus intermedius versus Staphylococcus delphini), the NCBI database search yielded 18 taxon-related ambiguities and showed numerous matches exhibiting redundant or unspecified entries. Comparing molecular results with those of biochemical procedures, ID 32 Staph (bioMerieux, Marcy I'Etoile, France) and VITEK 2 (bioMerieux) failed to identify 13 (23.6%) and 19 (34.5%) isolates, respectively, due to incorrect identification and/or categorization below acceptable values. In contrast to phenotypic methods and the NCBI database, the novel high-quality RIDOM sequence database provides excellent identification of staphylococci, including rarely isolated species and phenotypic variants.

Bacterial Typing Techniques↗

Improving quality of expressed sequence tag (EST) databases: recovery of reversed, antisense cDNA sequences.

Expressed sequence tag (EST) databases contain a significant number (5-20%) of reversed, antisense, cDNA sequences that can be recognized by the label "reversed clone: similarity on wrong strand" in the annotations to the sequence. Despite this high number of altered sequences, no attempt has been made to explain the alteration in molecular terms, or to evaluate their effect on the quality of the information curated in EST databases. In this paper we try to explain the way these altered sequences are originated, and propose a plausible mechanism: a "double priming" of the first strand oligo-dT primer at both ends of nascent cDNAs. In this way, a symmetrical cDNA intermediate is generated, an intermediate that can be cloned after partial digestion with the restriction enzyme used for the directional cloning. Furthermore, when "secondary" priming takes place inside the cDNA, the chain synthesized is prone to be truncated prematurely, with the subsequent loss of upstream information. One of the most subtle effects of this cloning alteration is the generation of virtual open reading frames (ORFs) in sequences with no homologues available for comparison. Nevertheless, and according to our model and our data, the "double priming mechanism" does not shift the ORF effected, so antisense sequences should be considered as normal ones after a simple transformation in their inverse-complementary forms.

Artifacts↗

Mouse BAC ends quality assessment and sequence analyses.

A large-scale BAC end-sequencing project at The Institute for Genomic Research (TIGR) has generated one of the most extensive sets of sequence markers for the mouse genome to date. With a sequencing success rate of >80%, an average read length of 485 bp, and ABI3700 capillary sequencers, we have generated 449,234 nonredundant mouse BAC end sequences (mBESs) with 218 Mb total from 257,318 clones from libraries RPCI-23 and RPCI-24, representing 15x clone coverage, 7% sequence coverage, and a marker every 7 kb across the genome. A total of 191,916 BACs have sequences from both ends providing 12x genome coverage. The average Q20 length is 406 bp and 84% of the bases have phred quality scores > or = 20. RPCI-24 mBESs have more Q20 bases and longer reads on average than RPCI-23 sequences. ABI3700 sequencers and the sample tracking system ensure that > 95% of mBESs are associated with the right clone identifiers. We have found that a significant fraction of mBESs contains L1 repeats and approximately 48% of the clones have both ends with > or = 100 bp contiguous unique Q20 bases. About 3% mBESs match ESTs and > 70% of matches were conserved between the mouse and the human or the rat. Approximately 0.1% mBESs contain STSs. About 0.2% mBESs match human finished sequences and > 70% of these sequences have EST hits. The analyses indicate that our high-quality mouse BAC end sequences will be a valuable resource to the community.

Animals↗

Evaluation of window cohabitation of DNA sequencing errors and lowest PHRED quality values.

When analyzing sequencing reads, it is important to distinguish between putative correct and wrong bases. An open question is how a PHRED quality value is capable of identifying the miscalled bases and if there is a quality cutoff that allows mapping of most errors. Considering the fact that a low quality value does not necessarily indicate a miscalled position, we decided to investigate if window-based analyses of quality values might better predict errors. There are many reasons to look for a perfect window in DNA sequences, such as when using SAGE technique, looking for BLAST seeding and clustering sequences. Thus, we set out to find a quality cutoff value that would distinguish non-perfect windows from perfect ones. We produced and compared 846 reads of pUC18 with the published pUC consensus, by local alignment. We then generated a database containing all mismatches, insertions and gaps in order to map real perfect windows. An investigation was made to find the potential to predict perfect windows when all bases in the window show quality values over a given cutoff. We conclude that, in window-based applications, a PHRED quality value cutoff of 7 masks most of the errors without masking real correct windows. We suggest that the putative wrong bases be indicated in lower case, increasing the information on the sequence databases without increasing the size the files.

Algorithms↗

ESTScan: a program for detecting, evaluating, and reconstructing potential coding regions in EST sequences.

One of the problems associated with the large-scale analysis of unannotated, low quality EST sequences is the detection of coding regions and the correction of frameshift errors that they often contain. We introduce a new type of hidden Markov model that explicitly deals with the possibility of errors in the sequence to analyze, and incorporates a method for correcting these errors. This model was implemented in an efficient and robust program, ESTScan. We show that ESTScan can detect and extract coding regions from low-quality sequences with high selectivity and sensitivity, and is able to accurately correct frameshift errors. In the framework of genome sequencing projects, ESTScan could become a very useful tool for gene discovery, for quality control, and for the assembly of contigs representing the coding regions of genes.

Algorithms↗

EDITtoTrEMBL: a distributed approach to high-quality automated protein sequence annotation.

SUMMARY: Many databases in molecular biology face the problem that the ever increasing rate of data production can no longer be handled by traditional methods, especially human curation. Therefore, a number of projects are currently investigating methods for automated sequence annotation. This paper describes the EBI's approach to this problem for protein sequences by integration of arbitrary analysis programs into a distributed and highly flexible environment. Our software framework allows an individual treatment of sequences depending on their particular properties, which is achieved through a high-level description of the preconditions and capabilities of analysing modules. This not only improves the overall performance of the annotation process, as unnecessary steps are avoided, but also enhances its quality since dependencies between different modules are taken into account. We have implemented a prototype and use it in the production of TrEMBL releases. AVAILABILITY: Upon request.

Algorithms↗

3.0 T high-resolution MR imaging of carpal ligaments and TFCC.

PURPOSE: To determine the diagnostic value of 3.0 Tesla MRI for imaging carpal ligaments and triangular fibrocartilage complex (TFCC). Image quality of different optimized MRI sequences is evaluated for high resolution wrist anatomy. MATERIALS AND METHODS: Ten healthy volunteers were examined at 3.0 T and 1.5 T using following sequences: T1 SE, fat-saturated PD-/T2-TSE, TIRM, 3D T1/T2* DESS, 3D-CISS, 2D and 3D T2* MEDIC. Voxel size varied from 0.2 x 0.2 x 1.5 mm (2D sequences) to 0.33 mm (3) and 0.26 mm (3) (3D sequences). Image quality (signal-to-noise-ratio, contrast-to-noise-ratio, artifacts) and carpal ligament/TFCC detection rate were judged by a score. The results obtained from the 3.0 T and 1.5 T devices were compared. RESULTS: With identical voxel size, image matrix and FOV, 3.0 T MRI provided significantly better image quality and ligament detection rates for all sequences in comparison with 1.5 T. The 2D and 3D MEDIC sequences yielded best image quality and detection rates. Excellent image quality and visualization of ligament structures by the fat-suppressed PD-TSE sequence were compromised by a relatively high susceptibility to pulsation and motion artifacts. T1 SE and 3D DESS sequences gave moderate image quality and allowed only partial differentiation between ligament structures. TIRM, T2-TSE and 3D-CISS sequence proved to be unsuitable for examining ligaments at 3.0 T due to their poor image quality and detection rate. CONCLUSION: 3.0 T MRI of the wrist proved to be superior to 1.5 T MRI for high-resolution imaging of carpal ligaments and TFCC using 2D and 3D T2* MEDIC sequences. Clinical studies investigating ligament injuries or carpal instability are recommended for evaluating clinical relevance of high-resolution MRI of the wrist.

Cartilage, Articular↗

Shedding genomic light on Aristotle's lantern.

Sea urchins have proved fascinating to biologists since the time of Aristotle who compared the appearance of their bony mouth structure to a lantern in The History of Animals. Throughout modern times it has been a model system for research in developmental biology. Now, the genome of the sea urchin Strongylocentrotus purpuratus is the first echinoderm genome to be sequenced. A high quality draft sequence assembly was produced using the Atlas assembler to combine whole genome shotgun sequences with sequences from a collection of BACs selected to form a minimal tiling path along the genome. A formidable challenge was presented by the high degree of heterozygosity between the two haplotypes of the selected male representative of this marine organism. This was overcome by use of the BAC tiling path backbone, in which each BAC represents a single haplotype, as well as by improvements in the Atlas software. Another innovation introduced in this project was the sequencing of pools of tiling path BACs rather than individual BAC sequencing. The Clone-Array Pooled Shotgun Strategy greatly reduced the cost and time devoted to preparing shotgun libraries from BAC clones. The genome sequence was analyzed with several gene prediction methods to produce a comprehensive gene list that was then manually refined and annotated by a volunteer team of sea urchin experts. This latter annotation community edited over 9000 gene models and uncovered many unexpected aspects of the sea urchin genetic content impacting transcriptional regulation, immunology, sensory perception, and an organism's development. Analysis of the basic deuterostome genetic complement supports the sea urchin's role as a model system for deuterostome and, by extension, chordate development.

Animals↗

High-quality automated DNA sequencing primed with hexamer strings.

The finishing phase of genome sequencing projects is expensive, in part, because of the cost of de novo synthesis of custom primers and the management burden associated with obtaining and using them for primer walking. One approach to reduce these high costs is the use of a presynthesized library of short oligonucleotides (8-10 bases) rather than long primers. The use of such a library eliminates the need for custom synthesis of oligonucleotides, providing the convenience of priming from any site by combining two to three short oligonucleotides to form a string with the required specificity. The first practical implementation of this strategy presented a robust protocol for using hexamer strings with radioisotopic labelling. Whereas versions of this technique have subsequently been implemented on fluorescent sequencers we felt that there was a need to develop and extensively test a protocol that consistently gave read lengths comparable to dye-terminator sequencing with longer primers. We have developed a new two-cycle fluorescent Sequenase terminator procedure for using hexamer strings. We tested this procedure using a set of 32 different 3 hexamer primer strings, each known to be functional to some degree in radioisotopic sequencing, on single-stranded M13mp18 template and ABI 373 DNA sequencers. The overall success rate of priming with these hexamer primer strings is 97% with the failure of only one string. In this case, the corresponding 18-mer primer also failed to produce usable sequence from M13mp18 template. The average read length from reactions successfully primed with the 31 different hexamer strings was 461 bases with > 99% base-calling accuracy. The current protocol is robust enough to be used in virtually any situation where primer walking on single-stranded templates is used. The success rate and read lengths make it universally applicable to the sequencing of single-stranded templates on automated sequencers. It is also amenable to automation.

Automation↗

Effect of the sampling sequence on the quality of Papanicolaou smear.

The aim of the study was to determine whether the order of cell collection (ie, obtaining either endocervical first or ectocervical cells first) has an effect on the quality of the Papanicolaou smear. 1129 smears were obtained using an Ayre spatula and an endocervical brush. In 564 cases, the endocervical brush was used first, and in 565 cases, the spatula was used first. The number of smears obscured by blood, the smears without endocervical component, and the smears with poor fixation were compared between the two groups. More smears were partially obscured by blood when brush was used first (78, 13.8% compared with 48, 8.5%, P = 0.004). No endocervical component was found in seven (1.2%) smears from the brush-first group compared with five (0.9%) of the spatula-first group, which is an insignificant difference. There were no significant differences in the number of poor-fixated smears, too-thick smears, and satisfactory smears but limited by inflammation between the two methods. The quality of the Papanicolaou smear can be improved by using the Ayre spatula first followed by the endocervical brush. Fewer smears will be contaminated by blood which may result in more squamous intraepithelial lesions being detected.

Female↗

The reverse DNA sequencing using Bst DNA polymerase.

The reverse DNA sequencing (RDS) [1] is a rapid method used to check the DNA sequences by sequencing them from the opposite orientation. Because the RDS is basically a double stranded sequencing, the quality of the sequence patterns so obtained generally is not as good as those obtained by the single stranded sequencing, and extra bands and higher background are produced more frequently. This paper shows that the RDS could now generate as good sequence patterns as those obtained by sequencing on the single stranded DNA template if Bst DNA polymerase instead of the conventional enzymes, such as the Klenow enzyme, was used in the RDS. Bst DNA polymerase is heat stable (optimum reaction temperature 65 degrees C) and has recently been successfully used in the conventional DNA sequencing. The RDS has recently been further simplified to meet the need of large DNA sequencing projects such as the human genome project. The combination of the simplified RDS and the use of Bst polymerase should be expected to facilitate greatly the work on sequence confirmation and correction.

Base Sequence↗

Sequence verification as quality-control step for production of cDNA microarrays.

To generate cDNA arrays in our core laboratory, we amplified about 2300 PCR products from a human, sequence-verified cDNA clone library. As a quality-control step, we sequenced the PCR products immediately before printing. The sequence information was used to search the GenBank database to confirm the identities. Although these clones were previously sequence verified by the company, we found that only 79% of the clones matched the original database after handling. Our experience strongly indicates the necessity to sequence verify the clones at the final stage before printing on microarray slides and to modify the gene list accordingly.

Gene Library↗

Direct sequencing of terminal regions of genomic P1 clones. A general strategy for the design of sequence-tagged site markers.

A method for the preparation of P1 DNA is presented, which allows the direct sequencing of ends of inserts in genomic P1 clones using the Applied Biosystems 373A DNA Sequencer and the Dye Terminator sequencing methodology. We surveyed several common methods of DNA preparation including alkaline lysis, Triton-lysozyme lysis, CsCl density-gradient purification, and a commercial column matrix DNA purification kit manufactured by Qiagen. We found that a modified alkaline lysis preparation of P1 DNA was most successful for generating P1 DNA that could be sequenced directly. We also noted that the host bacterial strain from which the P1 DNA was purified dramatically affected the quality of sequencing templates. The bacterial strains NS3145 and NS3529, in which the Drosophila melanogaster and human P1 genomic libraries are harbored, routinely yielded poor-quality sequencing templates. However, the bacterial strain DH10B routinely yielded P1 DNA that was sequenced successfully. A bacterial mating scheme is presented that exploits gamma delta transposition events to allow the transfer of P1 clones from the library host strain to DH10B. Using either an SP6 or a T7 primer, an average of 350 base pairs of DNA sequence was obtained with an uncalled base frequency of approximately 2%. About 4% of P1 end sequences generated corresponded to unique Drosophila loci present in the Genbank database. These single-pass DNA sequences were used to design sequence-tagged site markers for physical mapping studies in both humans and Drosophila.

Animals↗