Base qualities help sequencing software.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to S Dear.
Explore the source record for details and available documents.
Large-scale genomic sequencing requires a software infrastructure to support and integrate applications that are not directly compatible. We describe a suite of software tools built around the Common Assembly Format (CAF), a comprehensive representation of a sequence assembly as a text file. These tools form the backbone of sequencing informatics at the Sanger Centre and the Genome Sequencing Center. The CAF format is intentionally flexible, and our Perl and C libraries, which parse and manipulate it, provide powerful tools for creating new applications as well as wrappers to incorporate other software. The tools are available free by anonymous FTP from ftp://ftp.sanger.ac.uk/pub/badger/.
A software system for transforming fragments from four-color fluorescence-based gel electrophoresis experiments into assembled sequence is described. It has been developed for large-scale processing of all trace data, including shotgun and finishing reads, regardless of clone origin. Design considerations are discussed in detail, as are programming implementation and graphic tools. The importance of input validation, record tracking, and use of base quality values is emphasized. Several quality analysis metrics are proposed and applied to sample results from recently sequenced clones. Such quantities prove to be a valuable aid in evaluating modifications of sequencing protocol. The system is in full production use at both the Genome Sequencing Center and the Sanger Centre, for which combined weekly production is approximately 100, 000 sequencing reads per week.
Explore the source record for details and available documents.
Cystic fibrosis patients referred to two genetics centres in southern England and not found to carry common CF-associated mutations in one or both of their CFTR genes have been subjected to an extensive mutation search. The whole of the coding region of the CFTR gene, all intron-exon boundaries and 5' and 3' untranslated regions have been examined by a combination of single stranded conformational polymorphism analysis and chemical mismatch detection; 48 chromosomes with rare mutations have been identified, including 7 novel mutations, 182delT in exon 1, G27X in exon 2, Q151X in exon 4, Q220X in exon 6a, Q525X in exon 10, 3041delG in exon 16, and 4271delC in exon 23.
Explore the source record for details and available documents.
We describe a set of programs for creating and using indexes for the distributed forms of the major sequence libraries. The indexes conform to the specification of those distributed on cd-rom by the EMBL sequence library. The programs create entry name, accession number, author and freetext indexes and a brief directory index. If a suitable application program is given an entry name or accession number these indexes allow rapid retrieval of sequences or annotation. Similarly the author and freetext indexes provide the data for extremely fast searching on author names and "keywords". The indexing programs can create indexes for EMBL, Swiss-Prot, GenBANK, PIR and NRL3d libraries. We also describe the organisation and use of the different sequence libraries and their index files.
There are now a number of machines for determining DNA sequences. These devices are currently of two types: those such as the Applied Biosystems 373A and the Pharmacia A.L.F. which interpret the sequences of samples as they run on gels within the machine, and those, such as the Bio-Rad and Amersham readers that scan and analyse conventional autoradiographs. Both types of machine can produce their data in the form of traces which represent the band intensity of each of the four base types at each position in the sequence. At present all the machines write files in different formats. We describe a machine independent format for storing data derived from automatic sequencing machines. Files in this format can store the derived sequence, the traces and a set of confidence measures for each base. We have adopted the format as the standard for our sequence handling software.
We describe a sequence assembly and editing program for managing large and small projects. It is being used to sequence complete cosmids and has substantially reduced the time taken to process the data. In addition to handling conventionally derived sequences it can use data obtained from Applied Biosystems,Inc. 373A and Pharmacia A.L.F. fluorescent sequencing machines. Readings are assembled automatically. All editing is performed using a mouse operated contig editor that displays aligned sequences and their traces together on the screen. The editor, which can be used on single contigs or for joining contigs, permits rapid movement along the aligned sequences. Insertions, deletions and replacements can be made in individual aligned readings and global changes can be made by editing the consensus. All changes are recorded. A click on a mouse button will display the traces covering the current cursor position, hence allowing quick resolution of problems. Another function automatically moves the cursor to the next unresolved character. The editor also provides facilities for annotating the sequences. Typical annotations include flagging the positions of primers used for walking, or for marking sites, such as compressions, that have caused problems during sequencing. Graphical displays aid the assessment of progress.
Explore the source record for details and available documents.