Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Library Automation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

A knowledge base for predicting protein localization sites in eukaryotic cells.

To automate examination of massive amounts of sequence data for biological function, it is important to computerize interpretation based on empirical knowledge of sequence-function relationships. For this purpose, we have been constructing a knowledge base by organizing various experimental and computational observations as a collection of if-then rules. Here we report an expert system, which utilizes this knowledge base, for predicting localization sites of proteins only from the information on the amino acid sequence and the source origin. We collected data for 401 eukaryotic proteins with known localization sites (subcellular and extracellular) and divided them into training data and testing data. Fourteen localization sites were distinguished for animal cells and 17 for plant cells. When sorting signals were not well characterized experimentally, various sequence features were computationally derived from the training data. It was found that 66% of the training data and 59% of the testing data were correctly predicted by our expert system. This artificial intelligence approach is powerful and flexible enough to be used in genome analyses.

Algorithms↗

Rapid identification of mycolic acid patterns of mycobacteria by high-performance liquid chromatography using pattern recognition software and a Mycobacterium library.

Current methods for identifying mycobacteria by high-performance liquid chromatography (HPLC) require a visual assessment of the generated chromatographic data, which often involves time-consuming hand calculations and the use of flow charts. Our laboratory has developed a personal computer-based file containing patterns of mycolic acids detected in 45 species of Mycobacterium, including both slowly and rapidly growing species, as well as Tsukamurella paurometabolum and members of the genera Corynebacterium, Nocardia, Rhodococcus, and Gordona. The library was designed to be used in conjunction with a commercially available pattern recognition software package, Pirouette (Infometrix, Seattle, Wash.). Pirouette uses the K-nearest neighbor algorithm, a similarity-based classification method, to categorize unknown samples on the basis of their multivariate proximities to samples of a preassigned category. Multivariate proximity is calculated from peak height data, while peak heights are named by retention time matching. The system was tested for accuracy by using 24 species of Mycobacterium. Of the 1,333 strains evaluated, > or = 97% were correctly identified. Identification of M. tuberculosis (n = 649) was 99.85% accurate, and identification of the M. avium complex (n = 211) was > or = 98% accurate; > or = 95% of strains of both double-cluster and single-cluster M. gordonae (n = 47) were correctly identified. This system provides a rapid, highly reliable assessment of HPLC-generated chromatographic data for the identification of mycobacteria.

Algorithms↗

Development of a quantitative, high-throughput cell-based enzyme-linked immunosorbent assay for detection of colony-stimulating factor-1 receptor tyrosine kinase inhibitors.

Inhibitors of receptor tyrosine kinases are implicated as therapeutic agents for the treatment of many human diseases including cancer, inflammation and diabetes. Cell-based assays to examine inhibition of receptor tyrosine kinase mediated intracellular signaling are often laborious and not amenable to high-throughput cell-based screening of compound libraries. Here we describe the development of a nonradioactive, sandwich enzyme-linked immunosorbent assay (ELISA) to quantify the activation and inhibition of ligand-induced phosphorylation of the colony-stimulating factor-1 receptor (CSF-1R) in 96-well microtiter plate format. The assay involves the capture of the Triton X-100 solubilized human CSF-1R, from HEK293E cells overexpressing histidine epitope-tagged CSF-1R (CSF-1R/HEK293E), with immobilized CSF-1R antibody and detection of phosphosphorylation of the activated receptor with a phosphotyrosine specific antibody. The assay exhibited a 5-fold increase in phosphorylated CSF-1R signal from CSF-1R/HEK293E cells treated with colony-stimulating factor (CSF-1) relative to treated vector control cells. Additionally, using a histidine epitope-specific capture antibody, this method can also be adapted to quantify the phosphorylation state of any recombinantly expressed, histidine-tagged receptor tyrosine kinase. This method is a substantial improvement in throughput and quantitation of CSF-1R phosphorylation over conventional immunoblotting techniques.

Automation↗

Content-based retrieval of historical Ottoman documents stored as textual images.

There is an accelerating demand to access the visual content of documents stored in historical and cultural archives. Availability of electronic imaging tools and effective image processing techniques makes it feasible to process the multimedia data in large databases. In this paper, a framework for content-based retrieval of historical documents in the Ottoman Empire archives is presented. The documents are stored as textual images, which are compressed by constructing a library of symbols occurring in a document, and the symbols in the original image are then replaced with pointers into the codebook to obtain a compressed representation of the image. The features in wavelet and spatial domain based on angular and distance span of shapes are used to extract the symbols. In order to make content-based retrieval in historical archives, a query is specified as a rectangular region in an input image and the same symbol-extraction process is applied to the query region. The queries are processed on the codebook of documents and the query images are identified in the resulting documents using the pointers in textual images. The querying process does not require decompression of images. The new content-based retrieval framework is also applicable to many other document archives using different scripts.

Abstracting and Indexing↗

Small changes in the polymer structure influence the adsorption behavior of fibrinogen on polymer surfaces: validation of a new rapid screening technique.

Numerous studies conclude that the selective adsorption of plasma proteins on materials contacting blood or tissue affects all subsequent interactions related to the biocompatibility of artificial surfaces. However, there are only a few studies available, which clearly demonstrate that there is a correlation between surface chemistry and selective protein adsorption. Detailed knowledge of such correlations would facilitate the design of biocompatible materials. In this study, a rapid, fluorescence-based, screening technique using a 384-well format for polymer-protein interactions was developed. The screening assay was used to measure the adsorption of human fibrinogen on 46 test polymers (44 polyarylates selected from a combinatorial library of tyrosine-derived polyarylates, and two lactide-based polymers). In this library of polyarylates, structural changes are generated by variations in either the polymer backbone or the polymer pendent chain. Although no overall trend between polymer hydrophobicity and fibrinogen adsorption could be identified using the entire set of test polymers (R(2) = 0.43), fibrinogen adsorption was clearly correlated with variations in the pendent chain structure. Thus, when the test polymers were grouped by backbone composition, increased hydrophobicity of the pendent chain was significantly correlated with reduced fibrinogen adsorption. The following R(2) coefficients within the polymer backbone groups were determined: 0.87 (diglycolates); 0.98 (glutarates); 0.73 (adipates); 0.87 (suberates); 0.67 (3-methyl-adipates). Our results demonstrate that it is possible to screen for protein-material interactions in a cost-effective fashion using a miniaturized immunofluorescence technique. Further, we demonstrate that small changes in chemical composition can significantly influence the adsorption of human fibrinogen on polymer surfaces. The lactide-based polymers were among those polymers exhibiting the highest tendency to adsorb fibrinogen. This information may be useful when polymers have to be selected for specific biomaterial applications.

Adsorption↗

Integrating segmentation methods from the Insight Toolkit into a visualization application.

The Insight Toolkit (ITK) initiative from the National Library of Medicine has provided a suite of state-of-the-art segmentation and registration algorithms ideally suited to volume visualization and analysis. A volume visualization application that effectively utilizes these algorithms provides many benefits: it allows access to ITK functionality for non-programmers, it creates a vehicle for sharing and comparing segmentation techniques, and it serves as a visual debugger for algorithm developers. This paper describes the integration of image processing functionalities provided by the ITK into VolView, a visualization application for high performance volume rendering. A free version of this visualization application is publicly available and is available in the online version of this paper. The process for developing ITK plugins for VolView according to the publicly available API is described in detail, and an application of ITK VolView plugins to the segmentation of Abdominal Aortic Aneurysms (AAAs) is presented. The source code of the ITK plugins is also publicly available and it is included in the online version.

Algorithms↗

NIPALSTREE: a new hierarchical clustering approach for large compound libraries and its application to virtual screening.

A hierarchical clustering algorithm--NIPALSTREE--was developed that is able to analyze large data sets in high-dimensional space. The result can be displayed as a dendrogram. At each tree level the algorithm projects a data set via principle component analysis onto one dimension. The data set is sorted according to this one dimension and split at the median position. To avoid distortion of clusters at the median position, the algorithm identifies a potentially more suited split point left or right of the median. The procedure is recursively applied on the resulting subsets until the maximal distance between cluster members exceeds a user-defined threshold. The approach was validated in a retrospective screening study for angiotensin converting enzyme (ACE) inhibitors. The resulting clusters were assessed for their purity and enrichment in actives belonging to this ligand class. Enrichment was observed in individual branches of the dendrogram. In further retrospective virtual screening studies employing the MDL Drug Data Report (MDDR), COBRA, and the SPECS catalog, NIPALSTREE was compared with the hierarchical k-means clustering approach. Results show that both algorithms can be used in the context of virtual screening. Intersecting the result lists obtained with both algorithms improved enrichment factors while losing only few chemotypes.

Algorithms↗

Profiling of generic anti-phosphopeptide antibodies and kinases with peptide microarrays using radioactive and fluorescence-based assays.

Kinases represent one of the largest enzyme families and key regulatory proteins in the cell. Only a small subset of these enzymes has been characterised so far. We have prepared different types of phosphopeptide and peptide microarrays displaying peptides deduced from annotated human phosphorylation sites and cytoplasmic domains of all annotated human membrane proteins. This approach was enabled by fully-automated high throughput micro-scale synthesis of peptides by the SPOT technology combined with chemo-selective immobilisation on modified glass slides. The phosphopeptide microarrays displaying 2923 peptides in total have been used for the characterisation of commercially available generic anti-phosphopeptide antibodies. This enabled us to detect Abl kinase activity on a microarray with anti-phosphotyrosine antibodies yielding results comparable to those obtained from a radioactive assay. More than 13 000 peptides deposited on six glass slides were used to profile casein kinase 2 (CK2) using a radioactive assay, since no generic antibody for the reliable detection of serine or threonine phosphorylation could be identified. All previously identified substrates were detected in the microarray experiment. In order to confirm whether substrates on the microarray are substrates in solution phase assays, more than 700 peptides were synthesised and tested with CK2 in a solution phase assay. All substrates identified in the solution phase assay were also detected on the microarray.

Amino Acid Sequence↗

Autoantibody testing in children with newly diagnosed type 1 diabetes mellitus.

OBJECTIVES: To determine the role of autoantibody tests for autoimmune diseases in children with newly diagnosed type 1 diabetes mellitus. DATA SOURCES: MEDLINE, EMBASE and the Cochrane Library. Citation lists of included studies were scanned and relevant professional and patient websites reviewed. Laboratories and manufacturers were contacted to identify ongoing or unpublished research. REVIEW METHODS: Following scoping searches on thyroid and coeliac autoantibodies, a systematic review of autoantibody tests for diagnosis of coeliac disease was carried out. Studies were included where cohorts of untreated patients with unknown disease status were included, all patients had undergone the reference test (biopsy) and antibody tests, and sensitivity and specificity were reported or calculable. Selected studies were then evaluated against a quality checklist. Summary statistics of diagnostic accuracy, i.e. sensitivity, specificity, positive and negative likelihood ratios and diagnostic odds ratios, were calculated for all studies. A decision analytic model was developed to evaluate the cost utility of screening for coeliac disease at diagnosis of diabetes. RESULTS: All antibody tests for diagnosis of coeliac disease showed reasonably good diagnostic test accuracy. Studies reported variable measures of test accuracy, which may be due to aspects of study quality, differences in the tests and their execution in the laboratories, different populations and reference standards. The decision analytic model indicated screening for coeliac disease at diagnosis of diabetes was cost-effective. Sensitivity analyses exploring variations in the cost and disutility of gluten-free diet, the utilities attached to treated and untreated coeliac disease and the decrease in life expectancy associated with treated and untreated coeliac disease did substantially affect the cost-effectiveness of the screening strategies considered. CONCLUSIONS: In terms of test accuracy in testing for coeliac disease, immunoglobulin A (IgA) anti-endomysium is the most accurate test. If an enzyme-linked immunoassay test was required, which may be more suitable for screening purposes as it can be semi-automated, testing for IgA tissue transglutaminase is likely to be most accurate. The decision analytic model shows that the most accurate tests combined with confirmatory biopsy are the most cost-effective, whilst combinations of tests add little or no further value. There is limited information regarding test accuracy in screening populations with diabetes, and there is some uncertainty over whether the test characteristics would remain the same. Further research is required regarding the role of screening in silent coeliac disease and regarding long-term outcomes and complications of untreated coeliac disease.

Adolescent↗

Application of a self-consistent mean field theory to predict protein side-chains conformation and estimate their conformational entropy.

Understanding the relations between the conformation of the side-chains and the backbone geometry is crucial for structure prediction as well as for homology modelling. To attempt to unravel these rules, we have developed a method which allows us to predict the position of the side-chains from the co-ordinates of the main-chain atoms. This method is based on a rotamer library and refines iteratively a conformational matrix of the side-chains of a protein, CM, such that its current element at each cycle CM (ij) gives the probability that side-chain i of the protein adopts the conformation of its possible rotamer j. Each residue feels the average of all possible environments, weighted by their respective probabilities. The method converges in only a few cycles, thereby deserving the name of self consistent mean field method. Using the rotamer with the highest probability in the optimized conformational matrix to define the conformation of the side-chain leads to the result that on average 72% of chi 1, 75% of chi 2 and 62% of chi 1 + 2 are correctly predicted for a set of 30 proteins. Tests with six pairs of homologous proteins have shown that the method is quite successful even when the protein backbone deviates from the correct conformation. The second application of the optimized conformational matrix was to provide estimates of the conformational entropy of the side-chains in the folded state of the protein. The relevance of this entropy is discussed.

Amino Acids↗

An optimal structure-discriminative amino acid index for protein fold recognition.

Identifying the fold class of a protein sequence of unknown structure is a fundamental problem in modern biology. We apply a supervised learning algorithm to the classification of protein sequences with low sequence identity from a library of 174 structural classes created with the Combinatorial Extension structural alignment methodology. A class of rules is considered that assigns test sequences to structural classes based on the closest match of an amino acid index profile of the test sequence to a profile centroid for each class. A mathematical optimization procedure is applied to determine an amino acid index of maximal structural discriminatory power by maximizing the ratio of between-class to within-class profile variation. The optimal index is computed as the solution to a generalized eigenvalue problem, and its performance for fold classification is compared to that of other published indices. The optimal index has significantly more structural discriminatory power than all currently known indices, including average surrounding hydrophobicity, which it most closely resembles. It demonstrates >70% classification accuracy over all folds and nearly 100% accuracy on several folds with distinctive conserved structural features. Finally, there is a compelling universality to the optimal index in that it does not appear to depend strongly on the specific structural classes used in its computation.

Algorithms↗

Adaptation and evaluation of a personal electronic nose for selective multivapor analysis.

The evaluation of a commercial, belt-mountable "electronic nose" modified for the rapid recognition and quantification of individual solvent vapors and simple vapor mixtures at low ppm concentrations is described. Marketed under the name VaporLab this direct-reading instrument was designed for qualitative determinations of the presence or absence of selected individual vapors and was adapted in this study for quantitative determinations of vapors and vapor-mixture components. Vapor samples are concentrated on a small adsorbent bed and then thermally desorbed for analysis by an array of four polymer-coated surface acoustic wave sensors. Tests were performed with 13 organic solvent vapors individually and in selected binary, ternary, and quaternary mixtures at concentrations ranging from 0.1 to 12 times the respective American Conference of Governmental Industrial Hygienists' (ACGIH) threshold limit value (TLV). Pattern recognition analyses yielded a library of response patterns to which subsequent actual and virtual (i.e., Monte-Carlo simulated) samples were compared to assess performance. Limits of detection >0.025 x TLV are achieved (based on the most sensitive sensors) for 0.25 L of preconcentrated air samples collected over a 2-min period. Individual vapors from different functional group classes can be recognized, quantified, and discriminated from other vapors with little error, and discrimination of the components of binary mixtures is possible where the component vapor response patterns are sufficiently different. Within-class individual vapor and binary-mixture discriminations are more difficult and most ternary and higher-order mixtures could not be analyzed with acceptable accuracy. Changes in ambient humidity have no effect on responses and changes in temperature lead to well-behaved and compensable changes in responses. Tests of fluctuating concentrations demonstrate the capability for accurately tracking short-term variations in exposure. Overall, results suggest that this instrument could serve effectively as a personal exposure monitor in previously characterized occupational environments with proper revisions in design.

Air Pollution, Indoor↗

Transcriptome analysis of the salivary glands of Dermacentor andersoni Stiles (Acari: Ixodidae).

Amongst blood-feeding arthropods, ticks of the family Ixodidae (hard ticks) are vectors and reservoirs of a greater variety of infectious agents than any other ectoparasite. Salivary glands of ixodid ticks secrete a large number of pharmacologically active molecules that not only facilitate feeding but also promote establishment of infectious agents. Genomic, proteomic and immunologic characterization of bioactive salivary gland molecules are, therefore, important as they offer new insights into molecular events occurring at the tick-host interface and they have implications for development of novel control strategies. The present work uses complementary DNA (cDNA) sequence analysis to identify salivary gland transcripts expressed by the Rocky Mountain wood tick, Dermacentor andersoni, a vector of the human pathogens causing Rocky Mountain spotted fever, Colorado tick fever, tularemia, and Powassan encephalitis as well as the veterinary pathogen Anaplasma marginale. Dermacentor andersoni is also capable of inducing tick paralysis. Automated single-pass DNA sequencing was conducted on 1440 randomly selected cDNA clones from the salivary glands of adult female D. andersoni collected during the early stages of feeding (18-24h). Analysis of the expressed sequence tags (ESTs) resulted in 544 singletons and 218 clusters with more than one quality read and attempts were made to assign putative functions to tick genes based on amino acid identity to published protein databases. Approximately 25.6% (195) of the sequences showed limited or no homology to previously identified gene products. A number of novel sequences were identified which presented significant sequence similarity to mammalian genes normally associated with extracellular matrix (ECM), regulation of immune responses, tumor suppression, and wound healing. Several coding sequences possessed various degrees of homology to previously described proteins from other tick species. Preliminary nucleotide variation analysis of these and other tick sequences suggests extensive nucleotide diversity, which has implications for evolution of tick feeding. Intra-species diversity studies can be a promising tool for identifying sequence variations potentially associated with phenotypic traits affecting vector-host-pathogen interactions.

Amino Acid Sequence↗

Fluorescence in situ hybridization mapping of human chromosome 19: mapping and verification of cosmid contigs formed by random restriction enzyme fingerprinting.

Automated restriction enzyme fingerprinting of 7900 cosmids from chromosome 19 and calculation of the likelihood of their overlap based on shared fragments have resulted in the assembly of 743 sets of overlapping cosmids (contigs). We have mapped 22% of the formed contigs (n = 165) and all of the contigs with minimal tiling paths exceeding 6 members (n = 50) to chromosomal bands by fluorescence in situ hybridization using DNA from at least one member cosmid. The estimated average size of the formed contigs is 60-70 kb. Thus, members of a correctly formed contig are expected to lie close to each other in metaphase and interphase chromatin. Therefore, we tested the contig assembly process by comparing the band assignment of two or more members selected from each of 97 contigs. Forty-two of these contigs were further characterized for valid assembly by determining the proximity of members in interphase chromatin. Using these tests, we surveyed a total of 431 joins counted along the minimal tiling path (280 in interphase as well as metaphase) and found 6 erroneous joins, one in each of 6 contigs (6% of tested).

Chromosomes, Human, Pair 19↗

Computer-assisted prediction, classification, and delimitation of protein binding sites in nucleic acids.

We present a method to determine the location and extent of protein binding regions in nucleic acids by computer-assisted analysis of sequence data. The program ConsIndex establishes a library of consensus descriptions based on sequence sets containing known regulatory elements. These defined consensus descriptions are used by the program ConsInspector to predict binding sites in new sequences. We show the programs to correctly determine the significant regions involved in transcriptional control of seven sequence elements. The internal profile of relative variability of individual nucleotide positions within these regions paralleled experimental profiles of biological significance. Consensus descriptions are determined by employing an anchored alignment scheme, the results of which are then evaluated by a novel method which is superior to cluster algorithms. The alignment procedure is able to include several closely related sequences without biasing the consensus description. Moreover, the algorithm detects additional elements on the basis of a moderate distance correlation and is capable of discriminating between real binding sites and false positive matches. The software is well suited to cope with the frequent phenomenon of optional elements present in a subset of functionally similar sequences, while taking maximal advantage of the existing sequence data base. Since it requires only a minimum of seven sequences for a single element, it is applicable to a wide range of binding sites.

Algorithms↗

A method to improve visual similarity of breast masses for an interactive computer-aided diagnosis environment.

The purpose of this study was to develop and test a method for selecting "visually similar" regions of interest depicting breast masses from a reference library to be used in an interactive computer-aided diagnosis (CAD) environment. A reference library including 1000 malignant mass regions and 2000 benign and CAD-generated false-positive regions was established. When a suspicious mass region is identified, the scheme segments the region and searches for similar regions from the reference library using a multifeature based k-nearest neighbor (KNN) algorithm. To improve selection of reference images, we added an interactive step. All actual masses in the reference library were subjectively rated on a scale from 1 to 9 as to their "visual margins speculations". When an observer identifies a suspected mass region during a case interpretation he/she first rates the margins and the computerized search is then limited only to regions rated as having similar levels of spiculation (within +/-1 scale difference). In an observer preference study including 85 test regions, two sets of the six "similar" reference regions selected by the KNN with and without the interactive step were displayed side by side with each test region. Four radiologists and five nonclinician observers selected the more appropriate ("similar") reference set in a two alternative forced choice preference experiment. All four radiologists and five nonclinician observers preferred the sets of regions selected by the interactive method with an average frequency of 76.8% and 74.6%, respectively. The overall preference for the interactive method was highly significant (p < 0.001). The study demonstrated that a simple interactive approach that includes subjectively perceived ratings of one feature alone namely, a rating of margin "spiculation," could substantially improve the selection of "visually similar" reference images.

Algorithms↗

A method for the prediction of GPCRs coupling specificity to G-proteins using refined profile Hidden Markov Models.

BACKGROUND: G- Protein coupled receptors (GPCRs) comprise the largest group of eukaryotic cell surface receptors with great pharmacological interest. A broad range of native ligands interact and activate GPCRs, leading to signal transduction within cells. Most of these responses are mediated through the interaction of GPCRs with heterotrimeric GTP-binding proteins (G-proteins). Due to the information explosion in biological sequence databases, the development of software algorithms that could predict properties of GPCRs is important. Experimental data reported in the literature suggest that heterotrimeric G-proteins interact with parts of the activated receptor at the transmembrane helix-intracellular loop interface. Utilizing this information and membrane topology information, we have developed an intensive exploratory approach to generate a refined library of statistical models (Hidden Markov Models) that predict the coupling preference of GPCRs to heterotrimeric G-proteins. The method predicts the coupling preferences of GPCRs to Gs, Gi/o and Gq/11, but not G12/13 subfamilies. RESULTS: Using a dataset of 282 GPCR sequences of known coupling preference to G-proteins and adopting a five-fold cross-validation procedure, the method yielded an 89.7% correct classification rate. In a validation set comprised of all receptor sequences that are species homologues to GPCRs with known coupling preferences, excluding the sequences used to train the models, our method yields a correct classification rate of 91.0%. Furthermore, promiscuous coupling properties were correctly predicted for 6 of the 24 GPCRs that are known to interact with more than one subfamily of G-proteins. CONCLUSION: Our method demonstrates high correct classification rate. Unlike previously published methods performing the same task, it does not require any transmembrane topology prediction in a preceding step. A web-server for the prediction of GPCRs coupling specificity to G-proteins available for non-commercial users is located at http://bioinformatics.biol.uoa.gr/PRED-COUPLE.

Algorithms↗

Automated finishing with autofinish.

Currently, the genome sequencing community is producing shotgun sequence data at a very high rate, but finishing (collecting additional directed sequence data to close gaps and improve the quality of the data) is not matching that rate. One reason for the difference is that shotgun sequencing is highly automated but finishing is not: Most finishing decisions, such as which directed reads to obtain and which specialized sequencing techniques to use, are made by people. If finishing rates are to increase to match shotgun sequencing rates, most finishing decisions also must be automated. The Autofinish computer program (which is part of the computer software package) does this by automatically choosing finishing reads. Autofinish is able to suggest most finishing reads required for completion of each sequencing project, greatly reducing the amount of human attention needed. sometimes completely finishes the project, with no human decisions required. It cannot solve the most complex problems, so we recommend that Autofinish be allowed to suggest reads for the first three rounds of finishing, and if the project still is not finished completely, a human finisher complete the work. We compared this Autofinish-Hybrid method of finishing against a human finisher in five different projects with a variety of shotgun depths by finishing each project twice--once with each method. This comparison shows that the Autofinish-Hybrid method saves many hours over a human finisher alone, while using roughly the same number and type of reads and closing gaps at roughly the same rate. Autofinish currently is in production use at several large sequencing centers. It is designed to be adaptable to the finishing strategy of the lab--it can finish using some or all of the following: resequencing reads, reverses, custom primer walks on either subclone templates or whole clone templates, PCR, or minilibraries. Autofinish has been used for finishing cDNA, genomic clones, and whole bacterial genomes (see http://www.phrap.org).

Contig Mapping↗