Search PubMed⌕ Search

PubMed · 10869016

Fast probabilistic analysis of sequence function using scoring matrices.

Abstract

MOTIVATION: We present techniques for increasing the speed of sequence analysis using scoring matrices. Our techniques are based on calculating, for a given scoring matrix, the quantile function, which assigns a probability, or p, value to each segmental score. Our techniques also permit the user to specify a p threshold to indicate the desired trade-off between sensitivity and speed for a particular sequence analysis. The resulting increase in speed should allow scoring matrices to be used more widely in large-scale sequencing and annotation projects. RESULTS: We develop three techniques for increasing the speed of sequence analysis: probability filtering, lookahead scoring, and permuted lookahead scoring. In probability filtering, we compute the score threshold that corresponds to the user-specified p threshold. We use the score threshold to limit the number of segments that are retained in the search process. In lookahead scoring, we test intermediate scores to determine whether they will possibly exceed the score threshold. In permuted lookahead scoring, we score each segment in a particular order designed to maximize the likelihood of early termination. Our two lookahead scoring techniques reduce substantially the number of residues that must be examined. The fraction of residues examined ranges from 62 to 6%, depending on the p threshold chosen by the user. These techniques permit sequence analysis with scoring matrices at speeds that are several times faster than existing programs. On a database of 12 177 alignment blocks, our techniques permit sequence analysis at a speed of 225 residues/s for a p threshold of 10-6, and 541 residues/s for a p threshold of 10-20. In order to compute the quantile function, we may use either an independence assumption or a Markov assumption. We measure the effect of first- and second-order Markov assumptions and find that they tend to raise the p value of segments, when compared with the independence assumption, by average ratios of 1.30 and 1.69, respectively. We also compare our technique with the empirical 99. 5th percentile scores compiled in the BLOCKSPLUS database, and find that they correspond on average to a p value of 1.5 x 10-5. AVAILABILITY: The techniques described above are implemented in a software package called EMATRIX. This package is available from the authors for free academic use or for licensed commercial use. The EMATRIX set of programs is also available on the Internet at http://motif.stanford.edu/ematrix.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

T D Wu, C G Nevill-Manning, D L Brutlag. 2000. Fast probabilistic analysis of sequence function using scoring matrices.. https://doi.org/10.1093/bioinformatics%2F16.3.233

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Memory effects and macroscopic manifestation of randomness.

It is shown that due to memory effects the complex behavior of components in a stochastic system can be transmitted to macroscopic evolution of the system as a whole. Within the Markov approximation widely used in ordinary statistical mechanics, memory effects are neglected. As a result, a time-scale separation between the macroscopic and the microscopic level of description exists, the macroscopic differential picture is not a consequence of microscopic nondifferentiable dynamics. On the other hand, the presence of complete memory in a system means that all its components have the same behavior. If the memory function has no characteristic time scales, the correct description of the macroscopic evolution of such systems has to be in terms of the fractional calculus.

Markov Chains↗

Prediction of protein subcellular locations using Markov chain models.

A novel method was introduced to predict protein subcellular locations from sequences. Using sequence data, this method achieved a prediction accuracy higher than previous methods based on the amino acid composition. For three subcellular locations in a prokaryotic organism, the overall prediction accuracy reached 89.1%. For eukaryotic proteins, prediction accuracies of 73.0% and 78.7% were attained within four and three location categories, respectively. These results demonstrate the applicability of this relative simple method and possible improvement of prediction for the protein subcellular location.

Markov Chains↗

A free energy analysis by unfolding applied to 125-mers on a cubic lattice.

BACKGROUND: A common approach to the protein folding problem involves computer simulation of folding using lattice models of amino acid sequences. Key factors for good performance in such models are the correct choice of the temperature and the average interaction energy between residues. In order to push the lattice approach to its limit it is important to have a method to adjust these parameters for optimal folding that is not limited by our ability to successfully simulate folding in a reasonable time. RESULTS: In this study, we adopt a simple cubic-lattice model and present a method for calculating the free energy of a chain as a function of the number of native contacts. This does not require that we are able to fold the sequence by simulation and it provides a method of estimating the folding transition temperature. For a given set of parameters, the free energy analysis also allows an estimate of foldability. By applying the method to sequences with 27 and 125 residues, we show that optimal folding occurs near the folding transition temperature and at either zero or small negative average interaction energy. We find ourselves able to fold only 125-mers that have significant short-range native contacts. CONCLUSIONS: A free energy analysis during unfolding is a useful tool for the study of foldability and should be applicable to a variety of folding models. In this way we are able to fold some 125-mer designed sequences and our results confirm the finding that short-range contacts contribute to foldability.

Markov Chains↗