Search PubMed⌕ Search

PubMed · 15838592

Using string kernel to predict signal peptide cleavage site based on subsite coupling model.

Abstract

Owing to the importance of signal peptides for studying the molecular mechanisms of genetic diseases, reprogramming cells for gene therapy, and finding new drugs for healing a specific defect, it is in great demand to develop a fast and accurate method to identify the signal peptides. Introduction of the so-called {-3,-1, +1} coupling model (Chou, K. C.: Protein Engineering, 2001, 14-2, 75-79) has made it possible to take into account the coupling effect among some key subsites and hence can significantly enhance the prediction quality of peptide cleavage site. Based on the subsite coupling model, a kind of string kernels for protein sequence is introduced. Integrating the biologically relevant prior knowledge, the constructed string kernels can thus be used by any kernel-based method. A Support vector machines (SVM) is thus built to predict the cleavage site of signal peptides from the protein sequences. The current approach is compared with the classical weight matrix method. At small false positive ratios, our method outperforms the classical weight matrix method, indicating the current approach may at least serve as a powerful complemental tool to other existing methods for predicting the signal peptide cleavage site. The software that generated the results reported in this paper is available upon requirement, and will appear at http://www.pami.sjtu.edu.cn/wm.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

M Wang, J Yang, K-C Chou. 2005-04-21. Using string kernel to predict signal peptide cleavage site based on subsite coupling model.. https://doi.org/10.1007/s00726-005-0189-6

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Comparing ARG Inference Methods Under Transmission of Reproductive Success: Tree Imbalance Matters.

Inferring coalescent trees from genomic data has become a major subject in population genetics, particularly with the recent advances in tree sequence reconstruction methods. However, it remains unclear how well these methods perform for imbalanced genealogies. Such imbalances can arise from processes such as cultural transmission of reproductive success (CTRS) or positive selection. Using simulated genomic data, we benchmarked three major software packages, SINGER, Relate, and tsinfer, by comparing the imbalance of reconstructed trees by these methods with that of the true simulated trees, for three indices that quantify this imbalance. The three methods performed well under scenarios yielding balanced trees. However, their accuracy declined as imbalance increased. Performances also varied with mutation rate, recombination rate, and sample size. This study opens possibilities for applying these methods to infer CTRS or positive selection in large-scale genomic datasets, using simulation-based inference such as approximate Bayesian computation.

Models, Genetic↗

Diffusive Noise Controls Early Stages of Genetic Demixing.

Theoretical descriptions of the stepping-stone model, a cornerstone of spatial population genetics, have long overlooked diffusive noise arising from migration dynamics. We derive a fluctuating hydrodynamic description of this model from microscopic rules, which we then use to demonstrate that diffusive noise significantly alters early-time genetic demixing, which we characterize through heterozygosity, a key measure of diversity. Combining macroscopic fluctuation theory and microscopic simulations, we demonstrate that the scaling of allele-number fluctuations in a spatial domain displays an early-time behavior dominated by diffusive noise. Our results underscore the need for additional terms in existing continuum theories and highlight the necessity of including diffusive noise in models of spatially structured populations.

Models, Genetic↗

The relationship between the pleiotropic phenotypic effects of a mutation fixed by selection.

A pleiotropic model of mutation is presented that allows for correlations between the effects of a new mutation and for the distribution of mutational effects to vary from being leptokurtic to normally distributed. Using this model I quantify how selection transforms the correlation between the effects of a new (random) mutation into the correlation between the effects of a mutation that is fixed by selection and contributes to an adaptation. Results suggest that under most conditions the correlation between the effects of a fixed mutation is less than the correlation between the effects of a new mutation. I also generalize previous results that quantified the expected size of a fixed mutation's effect on a character given an observed effect of that mutation on another character. In agreement with previous results, work here suggests that as the observed effect becomes large and beneficial the expected effect on another character approaches the expected effect of a new (random) mutation given the observed effect. Lastly, these theoretical results are related to recent empirical work that found beneficial mutations had a positive correlation in their pleiotropic effects.

Models, Genetic↗