Search PubMedSearch

Biomedical subjects

D J States

Publications and source records attributed to D J States.

At least 19 recordsLinked to original sources

Comparative accuracy of methods for protein sequence similarity search.

MOTIVATION: Searching a protein sequence database for homologs is a powerful tool for discovering the structure and function of a sequence. Two new methods for searching sequence databases have recently been described: Probabilistic Smith-Waterman (PSW), which is based on Hidden Markov models for a single sequence using a standard scoring matrix, and a new version of BLAST (WU-BLAST2), which uses Sum statistics for gapped alignments. RESULTS: This paper compares and contrasts the effectiveness of these methods with three older methods (Smith-Waterman: SSEARCH, FASTA and BLASTP). The analysis indicates that the new methods are useful, and often offer improved accuracy. These tools are compared using a curated (by Bill Pearson) version of the annotated portion of PIR 39. Three different statistical criteria are utilized: equivalence number, minimum errors and the receiver operating characteristic. For complete-length protein query sequences from large families, PSW's accuracy is superior to that of the other methods, but its accuracy is poor when used with partial-length query sequences. False negatives are twice as common as false positives irrespective of the search methods if a family-specific threshold score that minimizes the total number of errors (i.e. the most favorable threshold score possible) is used. Thus, sensitivity, not selectivity, is the major problem. Among the analyzed methods using default parameters, the best accuracy was obtained from SSEARCH and PSW for complete-length proteins, and the two BLAST programs, plus SSEARCH, for partial-length proteins.

Databases, Factual

Filter matrix estimation in automated DNA sequencing.

In four-color fluourescence-based automated DNA sequencing, a 4 x 4 filter matrix parameterizes the relationship between the dye-intensity signals of interest and the data collected by an optical imaging system. The filter matrix is important because the estimated DNA sequence is based on the dye intensities that can only be recovered via inversion of the matrix. In this paper, we present a calibration method for the estimation of the columns of this matrix, using data generated through a special experiment in which DNA samples are labeled with only one fluorescent dye at a time. Simulations and applications of the method to real data are provided, with promising results.

Algorithms

Sequence assembly validation by multiple restriction digest fragment coverage analysis.

DNA sequence analysis depends on the accurate assembly of fragment reads for the determination of a consensus sequence. This report examines the possibility of analyzing multiple, independent restriction digests as a method for testing the fidelity of sequence assembly. A dynamic programming algorithm to determine the maximum likelihood alignment of error prone electrophoretic mobility data to the expected fragment mobilities given the consensus sequence and restriction enzymes is derived and used to assess the likelihood of detecting rearrangements in genomic sequencing projects. The method is shown to reliably detect errors in sequence fragment assembly without the necessity of making reference to an overlying physical map. An html form-based interface is available at http:/(/)www.ibc.wustl.edu/services/validate. html.

Algorithms

A method to determine the filter matrix in four-dye fluorescence-based DNA sequencing.

In a previous paper (Yin et al., Electrophoresis 1996, 17, 1143-1150), an automated method for matrix determination in four-dye fluorescence-based DNA sequencing was presented. As a continuation of that work, we have developed an alternative method to estimate the matrix from raw sequence data. The method uses an iterative clustering technique to associate each 4 x 1 data vector with one column of the desired filter matrix, using Kullback's I-divergence as a distance measure. The method requires less preprocessing of the data and less computation than the approach described by Yin et al. (Electrophoresis 1996, 17, 1143-1150). An example demonstrating applicability of the proposed method to Applied Biosystems sequencer data is given.

Algorithms

A Bayesian evolutionary distance for parametrically aligned sequences.

There is an inherent relationship between the process of pairwise sequence alignment and the estimation of evolutionary distance. This relationship is explored and made explicit. Assuming an evolutionary model and given a specific pattern of observed base mismatches, the relative probabilities of evolution at each evolutionary distance are computed using a Bayesian framework. The mean or the median of this probability distribution provides a robust estimate of the central value. The evolutionary distance has traditionally been computed as zero for an observed homology of 20 bases with no mismatches; we prove that it is highly probable that the distance is greater than 0.01. The mean of the distribution is 0.047, which is a better estimate of the evolutionary distance. Bayesian estimates of the evolutionary distance incorporate arbitrary prior information about variable mutation rates both over time and along sequence position, thus requiring only a weak form of the molecular-clock hypothesis. The endpoints of the similarity between genomic DNA sequences are often ambiguous. The probability of evolution at each evolutionary distance can be estimated over the entire set of alignments by choosing the best alignment at each distance and the corresponding probability of duplication at that evolutionary distance. A central value of this distribution provides a robust evolutionary distance estimate. We provide an efficient algorithm for computing the parametric alignment, considering evolutionary distance as the only parameter. These techniques and estimates are used to infer the duplication history of the genomic sequence in C. elegans and in S. cerevisiae. Our results indicate that repeats discovered using a single scoring matrix show a considerable bias in subsequent evolutionary distance estimates.

Animals

Lane tracking software for four-color fluorescence-based electrophoretic gel images.

Software to track sample lanes automatically in four-color, fluorescence-based, electrophoretic gel images has been developed for application in large-scale DNA sequencing projects. Lanes and lane boundaries are tracked by analyzing a first difference approximation to the gradient of a vertically integrated and processed "brightness" profile. Initially lanes are located in a region of the gel image selected for good horizontal lane spacing and signal strength. The software uses models of expected lane and interlane spacing and lateral lane behavior to maintain accurate tracking on imperfect gels. In areas where intensity-based tracking is difficult, interpixel column correlation is also used to locate and define lane features. Summary statistics and compressed-in-time images are generated for user evaluation of tracking performance. The software developed has been tested successfully on gel images with degradations including significant horizontal lane motion (curving) and image artifacts, and is now in full-scale use in our sequencing projects.

Algorithms

Compact encoding strategies for DNA sequence similarity search.

Determining whether two DNA sequences are similar is an essential component of DNA sequence analysis. Dynamic programming is the algorithm of choice if computational time is not the most important consideration. Heuristic search tools, such as BLAST, are computationally more efficient, but they may miss some of the sequence similarities (Altschul et al., 1990). These tools often use common k-tuples (words) between the two sequences to determine anchor points for the alignment, and spend most of their computational time extending the alignment beyond these anchor points. We discuss and provide a DNA sequence similarity search implementation (called SENSEI) that improves upon the performance of BLASTN by almost an order of magnitude for comparable sensitivity. This improvement is a result of using compactly encoded scoring tables for k-tuples, encoding bases with a single bit, filtering the sequence to remove the simple sequence repeats using XNUN, and masking the known species-specific repeats in the query sequence. To reduce memory requirements, especially for large genomic DNA query sequences, we recommend generating the neighborhood words from the target sequence at run-time, instead of generating them by preprocessing the query sequence.

Base Sequence

Combined use of sequence similarity and codon bias for coding region identification.

A computer program called BLASTX was previously shown to be effective in identifying and assigning putative function to likely protein coding regions by detecting significant similarity between a conceptually translated nucleotide query sequence and members of a protein sequence database. We present and assess the sensitivity of a new option to this software tool, herein called BLASTC, which employs information obtained from biases in codon utilization, along with the information obtained from sequence similarity. A rationale for combining these diverse information sources was derived, and analyses of the information available from codon utilization in several species were performed, with wide variation seen. Codon bias information was found on average to improve the sensitivity of detection of short coding regions of human origin by about a factor of 5. The implications of combining information sources on the interpretation of positive findings are discussed.

Algorithms

The Repeat Pattern Toolkit (RPT): analyzing the structure and evolution of the C. elegans genome.

Over 3.6 million bases of DNA sequence from chromosome III of the C. elegans have been determined. The availability of this extended region of contiguous sequence has allowed us to analyze the nature and prevalence of repetitive sequences in the genome of a eukaryotic organism with a high gene density. We have assembled a Repeat Pattern Toolkit (RPT) to analyze the patterns of repeats occurring in DNA. The tools include identifying significant local alignments (utilizing both two-way and three-way alignments), dividing the set of alignments into connected components (signifying repeat families), computing evolutionary distance between repeat family members, constructing minimum spanning trees from the connected components, and visualizing the evolution of the repeat families. Over 7000 families of repetitive sequences were identified. The size of the families ranged from isolated pairs to over 1600 segments of similar sequence. Approximately 12.3% of the analyzed sequence participates in a repeat element.

Animals

Identification of protein coding regions by database similarity search.

Sequence similarity between a translated nucleotide sequence and a known biological protein can provide strong evidence for the presence of a homologous coding region, even between distantly related genes. The computer program BLASTX performed conceptual translation of a nucleotide query sequence followed by a protein database search in one programmatic step. We characterized the sensitivity of BLASTX recognition to the presence of substitution, insertion and deletion errors in the query sequence and to sequence divergence. Reading frames were reliably identified in the presence of 1% query errors, a rate that is typical for primary sequence data. BLASTX is appropriate for use in moderate and large scale sequencing projects at the earliest opportunity, when the data are most prone to containing errors.

Algorithms

Computationally efficient cluster representation in molecular sequence megaclassification.

Molecular sequence megaclassification is a technique for automated protein sequence analysis and annotation. Implementation of the method has been limited by the need to store and randomly access a database of all the sequence pair similarities. More than 80,000 protein sequences are now present in the public databases, and the pair similarity data table for the full protein sequence database requires over 1 gigabyte of storage. In this paper we present a computationally efficient representation of groups based on a graph theory approach where sequence clusters are described by a minimal spanning tree of highest scoring similarity pairs. This representation allows a classification of N proteins to be stored in order(N) memory. The use of this minimal spanning tree representation simplifies analysis of groups, the description of group characteristics and the manual correction of artifacts resulting from false hits. The new tree representation also introduces new possibilities for artifact generation in sequence classification. Methods for detecting and removing these artifacts are discussed.

Algorithms

Molecular sequence accuracy: analysing imperfect data.

Molecular sequences are experimentally derived data that can be expected to contain errors as a result of diverse phenomena such as biological variation, molecular cloning artifacts, imperfect sequence determination, and data handling during contig assembly. Errors will affect the reliability of database searches and sequence alignments, but their impact may be minimized by the use of analytical techniques that anticipate that the data will be imperfect.

Amino Acid Sequence

Molecular sequence accuracy and the analysis of protein coding regions.

Molecular sequences, like all experimental data, have finite error rates. The impact of errors on the information content of molecular sequence data is dependent on the analytic paradigm used to interpret the data. We studied the impact of nucleic acid sequence errors on the ability to align predicted amino acid sequences with the sequences of related proteins. We found that with a simultaneous translation and alignment algorithm, identification of sequence homologies is resilient to the introduction of random errors. Proteins with greater than 30% sequence identity can be reliably recognized even in the presence of 1% frameshifting (insertion or deletion) error rates and 5% base substitution rates. Incorporation of prior knowledge about the location and characteristics of errors improves tolerance to error of amino acid sequence alignments. Similarly, inclusion of prior knowledge of biased codon utilization by yeast (Saccharomyces cerevisiae) allows reliable detection of correct reading frames in yeast sequences even in the presence of 5% substitution and 1% frameshift errors.

Algorithms

Structure of the human neutrophil elastase gene.

The gene for human neutrophil elastase (NE), a powerful serine protease carried by blood neutrophils and capable of destroying most connective tissue proteins, was cloned from a genomic DNA library of a normal individual. The NE gene consists of 5 exons and 4 introns included in a single copy 4-kilobase segment of chromosome 11 at q14. The coding exons of the NE gene predict a primary translation product of 267 residues including a 29-residue N-terminal precursor peptide and a 20-residue C-terminal precursor peptide. Analysis of the N-terminal peptide sequence suggests it contains a 27-residue "pre" signal peptide followed by a "proN" dipeptide, similar to that of other blood cell lysosomal proteases. The sequences for the mature 218-residue NE protein are included in exons II-V. The 5'-flanking region of the gene includes typical TATA, CAAT, and GC sequences within 61 base pairs (bp) of the cap site. The sequence 1.5 kilobases 5' to exon I contains several interesting repetitive sequences including six tandem repeats of unique 52- or 53-bp sequences. The 5'-flanking region also contains a 19-bp segment with 90% homology to a segment of the 5'-flanking region of the human myeloperoxidase (MPO) gene, a gene also expressed in bone marrow precursor cells and a protein stored in the same neutrophil granules as NE. In addition, like the MPO gene, the NE 5'-flanking region has several regions with greater than or equal to 75% homology to sequences 5' to c-myc, but there is no overlap between the NE-c-myc and MPO-c-myc homologous sequences.

Amino Acid Sequence

Electrostatic effects and hydrogen exchange behaviour in proteins. The pH dependence of exchange rates in lysozyme.

The pH dependence of the exchange rates for a number of tryptophan and amide hydrogen atoms in hen egg-white lysozyme has been determined at temperatures well below the thermal denaturation temperature. The pH behaviour of each hydrogen is unique and can differ markedly from that of simple compounds. A model for electrostatic effects in proteins is described and used to explain a number of the features of the pH dependence of the exchange rates of certain hydrogens. The results indicate that exchange takes place from a conformation of the protein closely similar to that of the native protein, with local fluctuations providing the mechanism for exchange. For the more-buried hydrogens at low pH values there is a general increase in the exchange rates caused by the decreasing stability of the protein as calculated from the electrostatic model. The analysis shows how evidence from hydrogen exchange studies can be used to provide information about electrostatic interactions in localized regions of proteins. A description of the electrostatic model and some applications are given in the Appendix.

Amino Acid Sequence

Conformations of intermediates in the folding of the pancreatic trypsin inhibitor.

Intermediates in the folding pathway of the bovine pancreatic trypsin inhibitor (PTI) have been examined by 1H nuclear magnetic resonance (n.m.r.). The intermediates were trapped during the reoxidation and consequent refolding of reduced PTI by alkylating free thiols; each intermediate contained different disulphide linkages. The n.m.r. spectra reveal that conformational features of the native protein are present in the intermediate containing just one of the three normal disulphide linkages (30-51). As additional normal disulphide bonds are formed, the conformation becomes more similar to that of the native protein. Introduction of additional but incorrect disulphide bonds does not lead to an increase in observable globular structure. A description of the folding process in terms of the conformations of the different intermediates is proposed. The significance of these results for the general mechanism of protein folding is outlined.

Amino Acid Sequence