Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

CAFTAN: a tool for fast mapping, and quality assessment of cDNAs.

BACKGROUND: The German cDNA Consortium has been cloning full length cDNAs and continued with their exploitation in protein localization experiments and cellular assays. However, the efficient use of large cDNA resources requires the development of strategies that are capable of a speedy selection of truly useful cDNAs from biological and experimental noise. To this end we have developed a new high-throughput analysis tool, CAFTAN, which simplifies these efforts and thus fills the gap between large-scale cDNA collections and their systematic annotation and application in functional genomics. RESULTS: CAFTAN is built around the mapping of cDNAs to the genome assembly, and the subsequent analysis of their genomic context. It uses sequence features like the presence and type of PolyA signals, inner and flanking repeats, the GC-content, splice site types, etc. All these features are evaluated in individual tests and classify cDNAs according to their sequence quality and likelihood to have been generated from fully processed mRNAs. Additionally, CAFTAN compares the coordinates of mapped cDNAs with the genomic coordinates of reference sets from public available resources (e.g., VEGA, ENSEMBL). This provides detailed information about overlapping exons and the structural classification of cDNAs with respect to the reference set of splice variants. The evaluation of CAFTAN showed that is able to correctly classify more than 85% of 5950 selected "known protein-coding" VEGA cDNAs as high quality multi- or single-exon. It identified as good 80.6 % of the single exon cDNAs and 85 % of the multiple exon cDNAs. The program is written in Perl and in a modular way, allowing the adoption of this strategy to other tasks like EST-annotation, or to extend it by adding new classification rules and new organism databases as they become available. We think that it is a very useful program for the annotation and research of unfinished genomes. CONCLUSION: CAFTAN is a high-throughput sequence analysis tool, which performs a fast and reliable quality prediction of cDNAs. Several thousands of cDNAs can be analyzed in a short time, giving the curator/scientist a first quick overview about the quality and the already existing annotation of a set of cDNAs. It supports the rejection of low quality cDNAs and helps in the selection of likely novel splice variants, and/or completely novel transcripts for new experiments.

Chromosome Mapping↗

MTT: a software tool for quality control in sequence assembly.

A large-scale sequencing project requires a tool to control the quality of the input data because a sizable number of trace data may be of low quality. If these data are allowed to enter the sequence assembly pipeline, harm will be done. Hence, it is important to detect such data as soon as possible. MTT (Move-Track-Trim) is a software package analyzing the quality of the lanes. It subjects each lane to a series of tests, and if a lane does not pass all tests, it is flagged as a "bad" lane. The use has a chance to examine both the "good" and the "bad" lanes and reclassify a "bad" lane as "good," or vice versa. Alternatively, the user may decide to retrack the gel or get rid of some lanes altogether. As a by-product of the analysis, MTT performs other useful functions. It trims the lanes and compresses the lane files and moves them to the directories where assembly is carried out. It also generates some useful statistics describing the quality of the gel.

Base Sequence↗

ESTWeb: bioinformatics services for EST sequencing projects.

ESTWeb is an internet based software package designed for uniform data processing and storage for large-scale EST sequencing projects. The package provides for: (a) reception of sequencing chromatograms; (b) sequence processing such as base-calling, vector screening, comparison with public databases; (c) storage of data and analysis in a relational database, (d) generation of a graphical report of individual sequence quality; and (e) issuing of reports with statistics of productivity and redundancy. The software facilitates real-time monitoring and evaluation of EST sequence acquisition progress along an EST sequencing project.

Computational Biology↗

Microarray-based resequencing of multiple Bacillus anthracis isolates.

We used custom-designed resequencing arrays to generate 3.1 Mb of genomic sequence from a panel of 56 Bacillus anthracis strains. Sequence quality was shown to be very high by replication (discrepancy rate of 7.4 x 10(-7)) and by comparison to independently generated shotgun sequence (discrepancy rate < 2.5 x 10(-6)). Population genomics studies of microbial pathogens using rapid resequencing technologies such as resequencing arrays are critical for recognizing newly emerging or genetically engineered strains.

Bacillus anthracis↗

Sequencing-based typing of HLA-A locus using mRNA and a single locus-specific PCR followed by cycle-sequencing with AmpliTaq DNA polymerase, FS.

The large number (59) of alleles now known at the HLA-A locus is a serious challenge to the existing methods for HLA typing, including many of the DNA based methods. Here, we describe a sequencing-based typing (SBT) protocol for typing of HLA-A alleles using a single A-locus-specific PCR. This reaction amplifies an 824 base pair product from cDNA, prepared from mRNA, covering exons 1-3 and most of exon 4. This product allows identification of all possible combinations of two alleles from this locus. The sequencing strategy used for allele assignment contains several improvements compared to those previously published. The enzyme AmpliTaq DNA Polymerase, FS, used, combines high-quality sequencing, i.e. long reads, low background, and uniform peak heights making the identification of heterozygous positions very reliable in a fast and easy protocol developed by determining the optima for a number of variables. Thus, this strategy meets most of the requirements for the use of sequencing in HLA typing. Furthermore, this method is very flexible. The use of a PCR primer-pair tailed with the recognition sites for two different sequencing primers allows the application of the same sets of fluorescent-labelled sequencing primers regardless of the amplified locus. Thus, the protocol can very easily be extended to cover the B- and C-locus too, simply by adding PCR reactions specific for these loci to the protocol. Using this protocol, we investigated a total of 65 cell lines and clinical samples, many of the latter chosen from samples difficult to type by serology. Our method gave in all cases unambiguous results and proved functional for work requiring the highest resolution.

Alleles↗

Advances in the Exon-Intron Database (EID).

Investigation of exon-intron gene structures is a non-trivial task due to enormous expansions of the eukaryotic genomes, great variety of gene forms, and the imperfectness in sequence data. A number of available informational systems on various gene characteristics complement each other and are indispensable for many genomic studies. Among them, the Exon-Intron Database (EID) is a good choice for large-scale computational examination of exon/intron structure and splicing. It has many internal filters that control for sequence quality, consistency of gene descriptions, accordance to standards, and possible errors. New innovations in EID are described. The collection of exons and introns has been extended beyond coding regions and current versions of EID contain data on untranslated regions of gene sequences as well. Intron-less genes are included as a special part of EID. For species with entirely sequenced genomes, species-specific databases have been generated. A novel Mammalian Orthologous Intron Database (MOID) has been introduced which includes the full set of introns that come from orthologous genes that have the same positions relative to the reading frames. Examples of statistical analyses of gene sequences using EID are provided. We present the latest data on our comparison of intron positions in 11,025 orthologous genes of human, mouse and rat, and find no convincing cases of intron gain. We discuss relevant data-quality issues of genomic databases. In particular, 5% of genes in genomic databases contain internal stop codons. This fact is due to a combination of biological reasons and also to errors in sequence annotations. The EID is freely available at www.meduohio.edu/bioinfo/eid/.

Base Sequence↗

Detecting and analyzing DNA sequencing errors: toward a higher quality of the Bacillus subtilis genome sequence.

During the determination of a DNA sequence, the introduction of artifactual frameshifts and/or in-frame stop codons in putative genes can lead to misprediction of gene products. Detection of such errors with a method based on protein similarity matching is only possible when related sequences are available in databases. Here, we present a method to detect frameshift errors in DNA sequences that is based on the intrinsic properties of the coding sequences. It combines the results of two analyses, the search for translational initiation/termination sites and the prediction of coding regions. This method was used to screen the complete Bacillus subtilis genome sequence and the regions flanking putative errors were resequenced for verification. This procedure allowed us to correct the sequence and to analyze in detail the nature of the errors. Interestingly, in several cases in-frame termination codons or frameshifts were not sequencing errors but confirmed to be present in the chromosome, indicating that the genes are either nonfunctional (pseudogenes) or subject to regulatory processes such as programmed translational frameshifts. The method can be used for checking the quality of the sequences produced by any prokaryotic genome sequencing project.

Bacillus subtilis↗

The maize genome as a model for efficient sequence analysis of large plant genomes.

The genomes of flowering plants vary in size from about 0.1 to over 100 gigabase pairs (Gbp), mostly because of polyploidy and variation in the abundance of repetitive elements in intergenic regions. High-quality sequences of the relatively small genomes of Arabidopsis (0.14 Gbp) and rice (0.4 Gbp) have now been largely completed. The sequencing of plant genomes that have a more representative size (the mean for flowering plant genomes is 5.6 Gbp) has been seen as a daunting task, partly because of their size and partly because of the numerous highly conserved repeats. Nevertheless, creative strategies and powerful new tools have been generated recently in the plant genetics community, so that sequencing large plant genomes is now a realistic possibility. Maize (2.4-2.7 Gbp) will be the first gigabase-size plant genome to be sequenced using these novel approaches. Pilot studies on maize indicate that the new gene-enrichment, gene-finishing and gene-orientation technologies are efficient, robust and comprehensive. These strategies will succeed in sequencing the gene-space of large genome plants, and in locating all of these genes and adjacent sequences on the genetic and physical maps.

DNA, Plant↗

The map-based sequence of the rice genome.

Rice, one of the world's most important food plants, has important syntenic relationships with the other cereal species and is a model plant for the grasses. Here we present a map-based, finished quality sequence that covers 95% of the 389 Mb genome, including virtually all of the euchromatin and two complete centromeres. A total of 37,544 non-transposable-element-related protein-coding genes were identified, of which 71% had a putative homologue in Arabidopsis. In a reciprocal analysis, 90% of the Arabidopsis proteins had a putative homologue in the predicted rice proteome. Twenty-nine per cent of the 37,544 predicted genes appear in clustered gene families. The number and classes of transposable elements found in the rice genome are consistent with the expansion of syntenic regions in the maize and sorghum genomes. We find evidence for widespread and recurrent gene transfer from the organelles to the nuclear chromosomes. The map-based sequence has proven useful for the identification of genes underlying agronomic traits. The additional single-nucleotide polymorphisms and simple sequence repeats identified in our study should accelerate improvements in rice production.

Cell Nucleus↗

Nematode.net: a tool for navigating sequences from parasitic and free-living nematodes.

Nematode.net (www.nematode.net) is a web- accessible resource for investigating gene sequences from nematode genomes. The database is an outgrowth of the parasitic nematode EST project at Washington University's Genome Sequencing Center (GSC), St Louis. A sister project at the University of Edinburgh and the Sanger Institute is also underway. More than 295,000 ESTs have been generated from >30 nematodes other than Caenorhabditis elegans including key parasites of humans, animals and plants. Nematode.net currently provides NemaGene EST cluster consensus sequence, enhanced online BLAST search tools, functional classifications of cluster sequences and comprehensive information concerning the ongoing generation of nematode genome data. The long-term goal of nematode.net is to provide the scientific community with the highest quality sequence information and tools for studying these diverse species.

Animals↗

A cSNP map and database for human chromosome 21.

Single nucleotide polymorphisms (SNPs) are likely to contribute to the study of complex genetic diseases. The genomic sequence of human chromosome 21q was recently completed with 225 annotated genes, thus permitting efficient identification and precise mapping of potential cSNPs by bioinformatics approaches. Here we present a human chromosome 21 (HC21) cSNP database and the first chromosome-specific cSNP map. Potential cSNPs were generated using three approaches: (1) Alignment of the complete HC21 genomic sequence to cognate ESTs and mRNAs. Candidate cSNPs were automatically extracted using a novel program for context-dependent SNP identification that efficiently discriminates between true variation, poor quality sequencing, and paralogous gene alignments. (2) Multiple alignment of all known HC21 genes to all other human database entries. (3) Gene-targeted cSNP discovery. To date we have identified 377 cSNPs averaging ~1 SNP per 1.5 kb of transcribed sequence, covering 65% of known genes in the chromosome. Validation of our bioinformatics approach was demonstrated by a confirmation rate of 78% for the predicted cSNPs, and in total 32% of the cSNPs in our database have been confirmed. The database is publicly available at http://csnp.unige.ch or http://csnp.isb-sib.ch. These SNPs provide a tool to study the contribution of HC21 loci to complex diseases such as bipolar affective disorder and allele-specific contributions to Down syndrome phenotypes.

Base Composition↗

A statistical score for assessing the quality of multiple sequence alignments.

BACKGROUND: Multiple sequence alignment is the foundation of many important applications in bioinformatics that aim at detecting functionally important regions, predicting protein structures, building phylogenetic trees etc. Although the automatic construction of a multiple sequence alignment for a set of remotely related sequences cause a very challenging and error-prone task, many downstream analyses still rely heavily on the accuracy of the alignments. RESULTS: To address the need for an objective evaluation framework, we introduce a statistical score that assesses the quality of a given multiple sequence alignment. The quality assessment is based on counting the number of significantly conserved positions in the alignment using importance sampling method in conjunction with statistical profile analysis framework. We first evaluate a novel objective function used in the alignment quality score for measuring the positional conservation. The results for the Src homology 2 (SH2) domain, Ras-like proteins, peptidase M13, subtilase and beta-lactamase families demonstrate that the score can distinguish sequence patterns with different degrees of conservation. Secondly, we evaluate the quality of the alignments produced by several widely used multiple sequence alignment programs using a novel alignment quality score and a commonly used sum of pairs method. According to these results, the Mafft strategy L-INS-i outperforms the other methods, although the difference between the Probcons, TCoffee and Muscle is mostly insignificant. The novel alignment quality score provides similar results than the sum of pairs method. CONCLUSION: The results indicate that the proposed statistical score is useful in assessing the quality of multiple sequence alignments.

Algorithms↗

Large-scale simulation of coverage and error rate tradeoffs for cancer detection in cell-free DNA whole-genome sequencing.

MOTIVATION: Cell-free DNA (cfDNA) whole-genome sequencing (WGS) is a promising approach for detecting cancer recurrence. It enables cancer detection by identifying all tumor-derived cfDNA (ctDNA) molecules carrying somatic single nucleotide variants (sSNVs). While ideally, a sequencing platform should be highly accurate for reliable ctDNA detection, in reality, all sequencing platforms introduce sequencing errors that generate false positives indistinguishable from true SNVs. Understanding how sequencing parameters influence ctDNA detection sensitivity at low tumor fractions (TFs) in cfDNA samples is essential for guiding sequencing strategies in clinical contexts. To model cfDNA sequencing for tumor detection, which contains asymmetric noise and multiple interacting parameters, analytical modeling is intractable, motivating large-scale parallelized simulation. RESULTS: We developed a simulation framework to generate in silico cfDNA data across 10 cancer types. In total, 480 million cfDNA samples were simulated from tumor WGS profiles. Overall, the lowest detectable TF differs substantially between cancer types under identical sequencing conditions due to variations in mutational load. For cancers with high mutational load, 3&#xd7; coverage with low-error techniques reliably detects TFs below 0.1%. In contrast, cancers with low mutational load require at least six-fold higher coverage to achieve comparable detection thresholds. Increasing sequencing quality scores from Q30 to Q55 at 30&#xd7; coverage further enhances sensitivity, enabling detection of TFs as low as 1&#x2009;&#xd7;&#x2009;10-5. This study provides a comprehensive framework for optimizing sequencing parameters, offering valuable guidance for tailoring future technology development for specific cancer types and clinical applications. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/UMCUGenetics/cfdetect/tree/main.

Whole Genome Sequencing↗

Exoquence DNA sequencing.

We have developed a strategy for DNA sequencing based on exonuclease III digestion followed by double strand specific endonuclease digestion and direct dideoxynucleotide sequencing reaction. This strategy eliminates the need for subcloning, oligonucleotide primers, and prior knowledge of the DNA to be sequenced. All template and primer duplexes needed for sequencing a complete insert can be prepared in one day from uncharacterized starting DNA. Sequence information can be obtained from different regions of the DNA simultaneously. The method uses double-stranded DNA to generate single-stranded template and primer, and thus produces high quality sequence results. Commercially available dideoxy-sequencing kits are well suited for this method. The strategy should be applicable for both automatic and routine laboratory DNA sequencing.

Exodeoxyribonucleases↗

Apparent homology of expressed genes from wood-forming tissues of loblolly pine (Pinus taeda L.) with Arabidopsis thaliana.

Pinus taeda L. (loblolly pine) and Arabidopsis thaliana differ greatly in form, ecological niche, evolutionary history, and genome size. Arabidopsis is a small, herbaceous, annual dicotyledon, whereas pines are large, long-lived, coniferous forest trees. Such diverse plants might be expected to differ in a large number of functional genes. We have obtained and analyzed 59,797 expressed sequence tags (ESTs) from wood-forming tissues of loblolly pine and compared them to the gene sequences inferred from the complete sequence of the Arabidopsis genome. Approximately 50% of pine ESTs have no apparent homologs in Arabidopsis or any other angiosperm in public databases. When evaluated by using contigs containing long, high-quality sequences, we find a higher level of apparent homology between the inferred genes of these two species. For those contigs 1,100 bp or longer, approximately 90% have an apparent Arabidopsis homolog (E value < 10-10). Pines and Arabidopsis last shared a common ancestor approximately 300 million years ago. Few genes would be expected to retain high sequence similarity for this time if they did not have essential functions. These observations suggest substantial conservation of gene sequence in seed plants.

3' Untranslated Regions↗

Generation and analysis of large-scale expressed sequence tags (ESTs) from a full-length enriched cDNA library of porcine backfat tissue.

BACKGROUND: Genome research in farm animals will expand our basic knowledge of the genetic control of complex traits, and the results will be applied in the livestock industry to improve meat quality and productivity, as well as to reduce the incidence of disease. A combination of quantitative trait locus mapping and microarray analysis is a useful approach to reduce the overall effort needed to identify genes associated with quantitative traits of interest. RESULTS: We constructed a full-length enriched cDNA library from porcine backfat tissue. The estimated average size of the cDNA inserts was 1.7 kb, and the cDNA fullness ratio was 70%. In total, we deposited 16,110 high-quality sequences in the dbEST division of GenBank (accession numbers: DT319652-DT335761). For all the expressed sequence tags (ESTs), approximately 10.9 Mb of porcine sequence were generated with an average length of 674 bp per EST (range: 200-952 bp). Clustering and assembly of these ESTs resulted in a total of 5,008 unique sequences with 1,776 contigs (35.46%) and 3,232 singleton (65.54%) ESTs. From a total of 5,008 unique sequences, 3,154 (62.98%) were similar to other sequences, and 1,854 (37.02%) were identified as having no hit or low identity (<95%) and 60% coverage in The Institute for Genomic Research (TIGR) gene index of Sus scrofa. Gene ontology (GO) annotation of unique sequences showed that approximately 31.7, 32.3, and 30.8% were assigned molecular function, biological process, and cellular component GO terms, respectively. A total of 1,854 putative novel transcripts resulted after comparison and filtering with the TIGR SsGI; these included a large percentage of singletons (80.64%) and a small proportion of contigs (13.36%). CONCLUSION: The sequence data generated in this study will provide valuable information for studying expression profiles using EST-based microarrays and assist in the condensation of current pig TCs into clusters representing longer stretches of cDNA sequences. The isolation of genes expressed in backfat tissue is the first step toward a better understanding of backfat tissue on a genomic basis.

Adipose Tissue↗

The Aggregated Gut Viral Catalogue (AVrC): A unified resource for exploring the viral diversity of the human gut.

The growing interest in the role of the gut virome in human health and disease, has led to several recent large-scale viral catalogue projects mining human gut metagenomes each using varied computational tools and quality control criteria. Importantly, there has been to date no consistent comparison of these catalogues' quality, diversity, and overlap. In this project, we therefore systematically surveyed nine previously published human gut viral catalogues. While these catalogues collectively screened >40,000 human fecal metagenomes, 82% of the recovered 345,613 viral sequences were unique to one catalogue, highlighting limited redundancy between the ressources and suggesting the need for an aggregated resource bringing these viral sequences together. We further expanded these viral catalogues by mining 7,867 infant gut metagenomes from 12 large-scale infant studies collected in 9 different countries. From these datasets, we constructed the Aggregated Gut Viral Catalogue (AVrC), a unified modular resource containing 1,018,941 dereplicated viral sequences (449,859 species-level vOTUs). Using computational inference tools, annotations were obtained for each vOTU representative sequence quality, viral taxonomy, predicted viral lifestyle, and putative host. This project aims to facilitate the reuse of previously published viral catalogues by the research community and follows a modular framework to enable future expansions as novel data becomes available.

Humans↗

DNA sequencing by indexer walking.

BACKGROUND: There is a need for DNA sequencing methods that are faster, more accurate, and less expensive than existing techniques. Here we present a new method for DNA analysis by means of indexer walking. METHODS: For DNA sequencing by indexer walking, we ligated double-stranded synthetic oligonucleotides (indexers) to DNA fragments that were produced by type IIS restriction endonucleases, which generate nonidentical 4-nucleotide 5' overhangs. The subsequent amplification (30 thermal cycles) of indexed DNA provided a template for automated DNA sequencing with fluorescent dideoxy terminators. The data gathered in the first sequencing reaction permitted further movement into the unknown nucleotide sequence by digestion of analyzed DNA with selected type IIS restriction endonuclease followed by ligation of the next indexer. A library of presynthesized indexers consisting of 256 oligonucleotides was used for bidirectional analysis of DNA molecules and provided universal primers for sequencing. RESULTS: The proposed protocol was successfully applied to sequencing of cryptic plasmids isolated from pathogenic strains of Escherichia coli. The overall error rate for base-calling was 0.5%, with a mean read length of 550 nucleotides. Approximately 1000 nucleotides of high-quality sequence could be obtained per day from a single clone. CONCLUSIONS: Indexer walking can be used as a low-cost procedure for nucleotide sequence determination of DNA molecules, such as natural plasmids, cDNA clones, and longer DNA fragments. It can also serve as an alternative method for gap filling at the final stage of genome sequencing projects.

DNA Primers↗