Search PubMed⌕ Search

Biomedical subjects

Kuo-Chen Chou

Publications and source records attributed to Kuo-Chen Chou.

At least 19 recordsLinked to original sources

SLLE for predicting membrane protein types.

Introduction of the concept of pseudo amino acid composition (PROTEINS: Structure, Function, and Genetics 43 (2001) 246; Erratum: ibid. 44 (2001) 60) has made it possible to incorporate a considerable amount of sequence-order effects by representing a protein sample in terms of a set of discrete numbers, and hence can significantly enhance the prediction quality of membrane protein type. As a continuous effort along such a line, the Supervised Locally Linear Embedding (SLLE) technique for nonlinear dimensionality reduction is introduced (Science 22 (2000) 2323). The advantage of using SLLE is that it can reduce the operational space by extracting the essential features from the high-dimensional pseudo amino acid composition space, and that the cluster-tolerant capacity can be increased accordingly. As a consequence by combining these two approaches, high success rates have been observed during the tests of self-consistency, jackknife and independent data set, respectively, by using the simplest nearest neighbour classifier. The current approach represents a new strategy to deal with the problems of protein attribute prediction, and hence may become a useful vehicle in the area of bioinformatics and proteomics.

Algorithms↗

Using GO-PseAA predictor to predict enzyme sub-class.

Enzyme function is much less conserved than anticipated, i.e., the requirement for sequence similarity that implies similarity in enzymatic function is much higher than the requirement that implies similarity in protein structure. This is because the function of an enzyme is an extremely complicated problem that may involve very subtle structural details as well as many other physical chemistry factors. Accordingly, if simply based on the sequence similarity approach, it would hardly get a decent success rate in predicting enzyme sub-class even for a dataset consisting of samples with 50% sequence identity. To cope with such a situation, the GO-PseAA predictor was adopted to identify the sub-class for each of the six main enzyme families. It has been observed that, even for the much more stringent datasets in which none of the enzymes has 25% sequence identity to any others, the overall success rates are 73-95%, suggesting that the GO-PseAA predictor can catch the core features of the statistical samples concerned and may become a useful high throughput tool in proteomics and bioinformatics.

Algorithms↗

Predicting protein localization in budding yeast.

MOTIVATION: Most of the existing methods in predicting protein subcellular location were used to deal with the cases limited within the scope from two to five localizations, and only a few of them can be effectively extended to cover the cases of 12-14 localizations. This is because the more the locations involved are, the poorer the success rate would be. Besides, some proteins may occur in several different subcellular locations, i.e. bear the feature of 'multiplex locations'. So far there is no method that can be used to effectively treat the difficult multiplex location problem. The present study was initiated in an attempt to address (1) how to efficiently identify the localization of a query protein among many possible subcellular locations, and (2) how to deal with the case of multiplex locations. RESULTS: By hybridizing gene ontology, functional domain and pseudo amino acid composition approaches, a new method has been developed that can be used to predict subcellular localization of proteins with multiplex location feature. A global analysis of the proteins in budding yeast classified into 22 locations was performed by jack-knife cross-validation with the new method. The overall success identification rate thus obtained is 70%. In contrast to this, the corresponding rates obtained by some other existing methods were only 13-14%, indicating that the new method is very powerful and promising. Furthermore, predictions were made for the four proteins whose localizations could not be determined by experiments, as well as for the 236 proteins whose localizations in budding yeast were ambiguous according to experimental observations. However, according to our predicted results, many of these 'ambiguous proteins' were found to have the same score and ranking for several different subcellular locations, implying that they may simultaneously exist, or move around, in these locations. This finding is intriguing because it reflects the dynamic feature of these proteins in a cell that may be associated with some special biological functions.

Algorithms↗

Predicting 22 protein localizations in budding yeast.

According to the recent experiments, proteins in budding yeast can be distinctly classified into 22 subcellular locations. Of these proteins, some bear the multi-locational feature, i.e., occur in more than one location. However, so far all the existing methods in predicting protein subcellular location were developed to deal with only the mono-locational case where a query protein is assumed to belong to one, and only one, subcellular location. To stimulate the development of subcellular location prediction, an augmentation procedure is formulated that will enable the existing methods to tackle the multi-locational problem as well. It has been observed thru a jackknife cross-validation test that the success rate obtained by the augmented GO-FnD-PseAA algorithm [BBRC 320 (2004) 1236] is overwhelmingly higher than those by the other augmented methods. It is anticipated that the augmented GO-FunD-PseAA predictor will become a very useful tool in predicting protein subcellular localization for both basic research and practical application.

Algorithms↗

Predicting protein structural class by functional domain composition.

The functional domain composition is introduced to predict the structural class of a protein or domain according to the following classification: all-alpha, all-beta, alpha/beta, alpha+beta, micro (multi-domain), sigma (small protein), and rho (peptide). The advantage by doing so is that both the sequence-order-related features and the function-related features are naturally incorporated in the predictor. As a demonstration, the jackknife cross-validation test was performed on a dataset that consists of proteins and domains with only less than 20% sequence identity to each other in order to get rid of any homologous bias. The overall success rate thus obtained was 98%. In contrast to this, the corresponding rates obtained by the simple geometry approaches based on the amino acid composition were only 36-39%. This indicates that using the functional domain composition to represent the sample of a protein for statistical prediction is very promising, and that the functional type of a domain is closely correlated with its structural class.

Databases, Protein↗

Weighted-support vector machines for predicting membrane protein types based on pseudo-amino acid composition.

Membrane proteins are generally classified into the following five types: (1) type I membrane proteins, (2) type II membrane proteins, (3) multipass transmembrane proteins, (4) lipid chain-anchored membrane proteins and (5) GPI-anchored membrane proteins. Prediction of membrane protein types has become one of the growing hot topics in bioinformatics. Currently, we are facing two critical challenges in this area: first, how to take into account the extremely complicated sequence-order effects, and second, how to deal with the highly uneven sizes of the subsets in a training dataset. In this paper, stimulated by the concept of using the pseudo-amino acid composition to incorporate the sequence-order effects, the spectral analysis technique is introduced to represent the statistical sample of a protein. Based on such a framework, the weighted support vector machine (SVM) algorithm is applied. The new approach has remarkable power in dealing with the bias caused by the situation when one subset in the training dataset contains many more samples than the other. The new method is particularly useful when our focus is aimed at proteins belonging to small subsets. The results obtained by the self-consistency test, jackknife test and independent dataset test are encouraging, indicating that the current approach may serve as a powerful complementary tool to other existing methods for predicting the types of membrane proteins.

Algorithms↗

Using amphiphilic pseudo amino acid composition to predict enzyme subfamily classes.

MOTIVATION: With protein sequences entering into databanks at an explosive pace, the early determination of the family or subfamily class for a newly found enzyme molecule becomes important because this is directly related to the detailed information about which specific target it acts on, as well as to its catalytic process and biological function. Unfortunately, it is both time-consuming and costly to do so by experiments alone. In a previous study, the covariant-discriminant algorithm was introduced to identify the 16 subfamily classes of oxidoreductases. Although the results were quite encouraging, the entire prediction process was based on the amino acid composition alone without including any sequence-order information. Therefore, it is worthy of further investigation. RESULTS: To incorporate the sequence-order effects into the predictor, the 'amphiphilic pseudo amino acid composition' is introduced to represent the statistical sample of a protein. The novel representation contains 20 + 2lambda discrete numbers: the first 20 numbers are the components of the conventional amino acid composition; the next 2lambda numbers are a set of correlation factors that reflect different hydrophobicity and hydrophilicity distribution patterns along a protein chain. Based on such a concept and formulation scheme, a new predictor is developed. It is shown by the self-consistency test, jackknife test and independent dataset tests that the success rates obtained by the new predictor are all significantly higher than those by the previous predictors. The significant enhancement in success rates also implies that the distribution of hydrophobicity and hydrophilicity of the amino acid residues along a protein chain plays a very important role to its structure and function.

Algorithms↗

Prediction of protein subcellular locations by GO-FunD-PseAA predictor.

The localization of a protein in a cell is closely correlated with its biological function. With the explosion of protein sequences entering into DataBanks, it is highly desired to develop an automated method that can fast identify their subcellular location. This will expedite the annotation process, providing timely useful information for both basic research and industrial application. In view of this, a powerful predictor has been developed by hybridizing the gene ontology approach [Nat. Genet. 25 (2000) 25], functional domain composition approach [J. Biol. Chem. 277 (2002) 45765], and the pseudo-amino acid composition approach [Proteins Struct. Funct. Genet. 43 (2001) 246; Erratum: ibid. 44 (2001) 60]. As a showcase, the recently constructed dataset [Bioinformatics 19 (2003) 1656] was used for demonstration. The dataset contains 7589 proteins classified into 12 subcellular locations: chloroplast, cytoplasmic, cytoskeleton, endoplasmic reticulum, extracellular, Golgi apparatus, lysosomal, mitochondrial, nuclear, peroxisomal, plasma membrane, and vacuolar. The overall success rate of prediction obtained by the jackknife cross-validation was 92%. This is so far the highest success rate performed on this dataset by following an objective and rigorous cross-validation procedure.

Algorithms↗

Insights from modelling the 3D structure of the extracellular domain of alpha7 nicotinic acetylcholine receptor.

Based on the crystal structure of acetylcholine-binding protein, the three-dimensional structures of the extracellular domain, or the ligand-binding domains, of the monomer, homodimer, and homopentamer of the alpha7 nicotinic acetylcholine receptor were derived. The interface between two subunits, where the ligand-binding site is located, was investigated. Furthermore, an explicit definition of the ligand-binding pocket was illustrated that might provide useful clues for conducting various mutagenesis studies for finding drugs against schizophrenia and Alzheimer's disease.

Amino Acid Sequence↗

Identify catalytic triads of serine hydrolases by support vector machines.

The core of an enzyme molecule is its active site from the viewpoints of both academic research and industrial application. To reveal the structural and functional mechanism of an enzyme, one needs to know its active site; to conduct structure-based drug design by regulating the function of an enzyme, one needs to know the active site and its microenvironment as well. Given the atomic coordinates of an enzyme molecule, how can we predict its active site? To tackle such a problem, a distance group approach was proposed and the support vector machine algorithm applied to predict the catalytic triad of serine hydrolase family. The success rate by jackknife test for the 139 serine hydrolases was 85%, implying that the method is quite promising and may become a useful tool in structural bioinformatics.

Algorithms↗

Predicting subcellular localization of proteins by hybridizing functional domain composition and pseudo-amino acid composition.

Recent advances in large-scale genome sequencing have led to the rapid accumulation of amino acid sequences of proteins whose functions are unknown. Since the functions of these proteins are closely correlated with their subcellular localizations, many efforts have been made to develop a variety of methods for predicting protein subcellular location. In this study, based on the strategy by hybridizing the functional domain composition and the pseudo-amino acid composition (Cai and Chou [2003]: Biochem. Biophys. Res. Commun. 305:407-411), the Intimate Sorting Algorithm (ISort predictor) was developed for predicting the protein subcellular location. As a showcase, the same plant and non-plant protein datasets as investigated by the previous investigators were used for demonstration. The overall success rate by the jackknife test for the plant protein dataset was 85.4%, and that for the non-plant protein dataset 91.9%. These are so far the highest success rates achieved for the two datasets by following a rigorous cross validation test procedure, further confirming that such a hybrid approach may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology.

Algorithms↗

Modelling extracellular domains of GABA-A receptors: subtypes 1, 2, 3, and 5.

GABA is the main inhibitory neurotransmitter in the mammalian central nervous system. When GABA binds to the ubiquitous GABA-A receptors on neurons, chloride channels are activated leading to a rapid increase in chloride conductance that depresses excitatory depolarization. The GABA-A receptors are targets for many clinically important drugs, such as the benzodiazepines, general anaesthetics, and barbiturates. All of these drugs enhance the chloride current activated by GABA. Of the GABA-A receptor family, the subtype 2 is critical for the treatment of anxiety spectrum disorders. To avoid unwanted side effects, it is necessary to find highly selective drugs that interact only with subtype 2 but not with the related receptors such as subtypes 1, 3, and 5. To realize such a goal, it is important to have not only the 3D (dimensional) structure of subtype 2 but also the 3D structures of subtypes 1, 3, and 5. In this study, the 3D structures of all the four subtypes of GABA-A receptors have been derived. The computer-modeled heteropentameric structures bear the following features: (1) each of the five subunits in the pentamer has an intrachain disulfide bond, a hallmark of ligand-gated pentameric channels; (2) those residues which are sensitive to the binding of the benzodiazepine site ligands are grouped around the alpha1,2,3,5/gamma2 interfaces; and (3) those residues which are sensitive to the binding of GABA molecules are grouped around the alpha1,2,3,5/beta2 interfaces. All these findings are fully consistent with experimental observations. Meanwhile, for those sensitive or key residues, a close look at their subtle difference among the four subtypes has been provided through a highlighted superposition picture. In addition to providing the atomic coordinates, the predicted structures have further clarified some ambiguities that could not been uniquely determined by the existing experimental data, such as the directionality of the subunit arrangement in the heteropentamers. The 3D models may provide a reasonable structural frame or footing for designing highly selective drugs. The present models might be also useful in understanding the basic mechanism of operation of the GABA-A receptors, stimulating novel strategies for developing more specific drugs and better treatments.

Amino Acid Sequence↗

A novel approach to predict active sites of enzyme molecules.

Enzymes are critical in many cellular signaling cascades. With many enzyme structures being solved, there is an increasing need to develop an automated method for identifying their active sites. However, given the atomic coordinates of an enzyme molecule, how can we predict its active site? This is a vitally important problem because the core of an enzyme molecule is its active site from the viewpoints of both pure scientific research and industrial application. In this article, a topological entity was introduced to characterize the enzymatic active site. Based on such a concept, the covariant discriminant algorithm was formulated for identifying the active site. As a paradigm, the serine hydrolase family was demonstrated. The overall success rate by jackknife test for a data set of 88 enzyme molecules was 99.92%, and that for a data set of 50 independent enzyme molecules was 99.91%. Meanwhile, it was shown through an example that the prediction algorithm can also be used to find any typographic error of a PDB file in annotating the constituent amino acids of catalytic triad and to suggest a possible correction. The very high success rates are due to the introduction of a covariance matrix in the prediction algorithm that makes allowance for taking into account the coupling effects among the key constituent atoms of active site. It is anticipated that the novel approach is quite promising and may become a useful high throughput tool in enzymology, proteomics, and structural bioinformatics.

Algorithms↗

Application of SVM to predict membrane protein types.

As a continuous effort to develop automated methods for predicting membrane protein types that was initiated by Chou and Elrod (PROTEINS: Structure, Function, and Genetics, 1999, 34, 137-153), the support vector machine (SVM) is introduced. Results obtained through re-substitution, jackknife, and independent data set tests, respectively, have indicated that the SVM approach is quite a promising one, suggesting that the covariant discriminant algorithm (Chou and Elrod, Protein Eng. 12 (1999) 107) and SVM, if effectively complemented with each other, will become a powerful tool for predicting membrane protein types and the other protein attributes as well.

Algorithms↗

Predicting subcellular localization of proteins in a hybridization space.

MOTIVATION: The localization of a protein in a cell is closely correlated with its biological function. With the number of sequences entering into databanks rapidly increasing, the importance of developing a powerful high-throughput tool to determine protein subcellular location has become self-evident. In view of this, the Nearest Neighbour Algorithm was developed for predicting the protein subcellular location using the strategy of hybridizing the information derived from the recent development in gene ontology with that from the functional domain composition as well as the pseudo amino acid composition. RESULTS: As a showcase, the same plant and non-plant protein datasets as investigated by the previous investigators were used for demonstration. The overall success rate of the jackknife test for the plant protein dataset was 86%, and that for the non-plant protein dataset 91.2%. These are the highest success rates achieved so far for the two datasets by following a rigorous cross-validation test procedure, suggesting that such a hybrid approach (particularly by incorporating the knowledge of gene ontology) may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology. AVAILABILITY: The software would be made available on sending a request to the authors.

Algorithms↗

Bio-support vector machines for computational proteomics.

MOTIVATION: One of the most important issues in computational proteomics is to produce a prediction model for the classification or annotation of biological function of novel protein sequences. In order to improve the prediction accuracy, much attention has been paid to the improvement of the performance of the algorithms used, few is for solving the fundamental issue, namely, amino acid encoding as most existing pattern recognition algorithms are unable to recognize amino acids in protein sequences. Importantly, the most commonly used amino acid encoding method has the flaw that leads to large computational cost and recognition bias. RESULTS: By replacing kernel functions of support vector machines (SVMs) with amino acid similarity measurement matrices, we have modified SVMs, a new type of pattern recognition algorithm for analysing protein sequences, particularly for proteolytic cleavage site prediction. We refer to the modified SVMs as bio-support vector machine. When applied to the prediction of HIV protease cleavage sites, the new method has shown a remarkable advantage in reducing the model complexity and enhancing the model robustness.

Algorithms↗

Predicting the linkage sites in glycoproteins using bio-basis function neural network.

MOTIVATION: Although, it is known that O-glycosidically linked oligosaccharides are commonly conjugated to a serine, threonine or hydroxylysine residue of the polypeptide, the chemical nature of the anchoring monosaccharide and the size of the oligosaccharide unit varies. Among different types, O-linked or mucin-type oligosaccharides are intimately involved in the secretion of proteins, be they enzymes, hormones or structural glycoproteins. Knowledge of the linkage sites in glycoproteins is critical to the design of specific and efficient inhibitors against the enzyme to catalyse the formation of the carbohydrate-peptide linkage. RESULTS: We present a method for predicting the linkage sites in O-linked glycoproteins using bio-basis function neural networks. The mean prediction accuracy of this method is 91.15 +/- 2.75% while it is 82.28 +/- 6.45% using back-propagation neural networks. Importantly, this method has significantly reduced the CPU time for modelling.

Algorithms↗

Polyprotein cleavage mechanism of SARS CoV Mpro and chemical modification of the octapeptide.

The cleavage mechanism of severe acute respiratory syndrome (SARS) coronavirus main proteinase (M(pro) or 3CL(pro)) for the octapeptide AVLQSGFR is studied using molecular mechanics (MM) and quantum mechanics (QM). The catalytic dyad His-41 and Cys-145 in the active pocket between domain I and II seem to polarize the pi-electron density of the peptide bond between Gln and Ser in the octapeptide, leading to an increase of positive charge on C(CO) of Gln and negative charge on N(NH) of Ser. The possibility of enhancing the chemical bond between Gln and Ser based on the "distorted key" theory [Anal. Biochem. 233 (1996) 1] is examined. The scissile peptide bond between Gln and Ser is found to be solidified through "hybrid peptide bond" by changing the carbonyl group CO of Gln to CH(2) or CF(2). This leads to a break of the pi-bond system for the peptide bond, making the octapeptide (AVLQSGFR) a "distorted key" and a potential starting system for the design of anti SARS drugs.

Amino Acid Sequence↗