Search PubMed⌕ Search

Biomedical subjects

Kuo-Chen Chou

Publications and source records attributed to Kuo-Chen Chou.

At least 73 records · Page 4Linked to original sources

Identify catalytic triads of serine hydrolases by support vector machines.

The core of an enzyme molecule is its active site from the viewpoints of both academic research and industrial application. To reveal the structural and functional mechanism of an enzyme, one needs to know its active site; to conduct structure-based drug design by regulating the function of an enzyme, one needs to know the active site and its microenvironment as well. Given the atomic coordinates of an enzyme molecule, how can we predict its active site? To tackle such a problem, a distance group approach was proposed and the support vector machine algorithm applied to predict the catalytic triad of serine hydrolase family. The success rate by jackknife test for the 139 serine hydrolases was 85%, implying that the method is quite promising and may become a useful tool in structural bioinformatics.

Algorithms↗

Predicting subcellular localization of proteins by hybridizing functional domain composition and pseudo-amino acid composition.

Recent advances in large-scale genome sequencing have led to the rapid accumulation of amino acid sequences of proteins whose functions are unknown. Since the functions of these proteins are closely correlated with their subcellular localizations, many efforts have been made to develop a variety of methods for predicting protein subcellular location. In this study, based on the strategy by hybridizing the functional domain composition and the pseudo-amino acid composition (Cai and Chou [2003]: Biochem. Biophys. Res. Commun. 305:407-411), the Intimate Sorting Algorithm (ISort predictor) was developed for predicting the protein subcellular location. As a showcase, the same plant and non-plant protein datasets as investigated by the previous investigators were used for demonstration. The overall success rate by the jackknife test for the plant protein dataset was 85.4%, and that for the non-plant protein dataset 91.9%. These are so far the highest success rates achieved for the two datasets by following a rigorous cross validation test procedure, further confirming that such a hybrid approach may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology.

Algorithms↗

Modelling extracellular domains of GABA-A receptors: subtypes 1, 2, 3, and 5.

GABA is the main inhibitory neurotransmitter in the mammalian central nervous system. When GABA binds to the ubiquitous GABA-A receptors on neurons, chloride channels are activated leading to a rapid increase in chloride conductance that depresses excitatory depolarization. The GABA-A receptors are targets for many clinically important drugs, such as the benzodiazepines, general anaesthetics, and barbiturates. All of these drugs enhance the chloride current activated by GABA. Of the GABA-A receptor family, the subtype 2 is critical for the treatment of anxiety spectrum disorders. To avoid unwanted side effects, it is necessary to find highly selective drugs that interact only with subtype 2 but not with the related receptors such as subtypes 1, 3, and 5. To realize such a goal, it is important to have not only the 3D (dimensional) structure of subtype 2 but also the 3D structures of subtypes 1, 3, and 5. In this study, the 3D structures of all the four subtypes of GABA-A receptors have been derived. The computer-modeled heteropentameric structures bear the following features: (1) each of the five subunits in the pentamer has an intrachain disulfide bond, a hallmark of ligand-gated pentameric channels; (2) those residues which are sensitive to the binding of the benzodiazepine site ligands are grouped around the alpha1,2,3,5/gamma2 interfaces; and (3) those residues which are sensitive to the binding of GABA molecules are grouped around the alpha1,2,3,5/beta2 interfaces. All these findings are fully consistent with experimental observations. Meanwhile, for those sensitive or key residues, a close look at their subtle difference among the four subtypes has been provided through a highlighted superposition picture. In addition to providing the atomic coordinates, the predicted structures have further clarified some ambiguities that could not been uniquely determined by the existing experimental data, such as the directionality of the subunit arrangement in the heteropentamers. The 3D models may provide a reasonable structural frame or footing for designing highly selective drugs. The present models might be also useful in understanding the basic mechanism of operation of the GABA-A receptors, stimulating novel strategies for developing more specific drugs and better treatments.

Amino Acid Sequence↗

A novel approach to predict active sites of enzyme molecules.

Enzymes are critical in many cellular signaling cascades. With many enzyme structures being solved, there is an increasing need to develop an automated method for identifying their active sites. However, given the atomic coordinates of an enzyme molecule, how can we predict its active site? This is a vitally important problem because the core of an enzyme molecule is its active site from the viewpoints of both pure scientific research and industrial application. In this article, a topological entity was introduced to characterize the enzymatic active site. Based on such a concept, the covariant discriminant algorithm was formulated for identifying the active site. As a paradigm, the serine hydrolase family was demonstrated. The overall success rate by jackknife test for a data set of 88 enzyme molecules was 99.92%, and that for a data set of 50 independent enzyme molecules was 99.91%. Meanwhile, it was shown through an example that the prediction algorithm can also be used to find any typographic error of a PDB file in annotating the constituent amino acids of catalytic triad and to suggest a possible correction. The very high success rates are due to the introduction of a covariance matrix in the prediction algorithm that makes allowance for taking into account the coupling effects among the key constituent atoms of active site. It is anticipated that the novel approach is quite promising and may become a useful high throughput tool in enzymology, proteomics, and structural bioinformatics.

Algorithms↗

Application of SVM to predict membrane protein types.

As a continuous effort to develop automated methods for predicting membrane protein types that was initiated by Chou and Elrod (PROTEINS: Structure, Function, and Genetics, 1999, 34, 137-153), the support vector machine (SVM) is introduced. Results obtained through re-substitution, jackknife, and independent data set tests, respectively, have indicated that the SVM approach is quite a promising one, suggesting that the covariant discriminant algorithm (Chou and Elrod, Protein Eng. 12 (1999) 107) and SVM, if effectively complemented with each other, will become a powerful tool for predicting membrane protein types and the other protein attributes as well.

Algorithms↗

Predicting subcellular localization of proteins in a hybridization space.

MOTIVATION: The localization of a protein in a cell is closely correlated with its biological function. With the number of sequences entering into databanks rapidly increasing, the importance of developing a powerful high-throughput tool to determine protein subcellular location has become self-evident. In view of this, the Nearest Neighbour Algorithm was developed for predicting the protein subcellular location using the strategy of hybridizing the information derived from the recent development in gene ontology with that from the functional domain composition as well as the pseudo amino acid composition. RESULTS: As a showcase, the same plant and non-plant protein datasets as investigated by the previous investigators were used for demonstration. The overall success rate of the jackknife test for the plant protein dataset was 86%, and that for the non-plant protein dataset 91.2%. These are the highest success rates achieved so far for the two datasets by following a rigorous cross-validation test procedure, suggesting that such a hybrid approach (particularly by incorporating the knowledge of gene ontology) may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology. AVAILABILITY: The software would be made available on sending a request to the authors.

Algorithms↗

Bio-support vector machines for computational proteomics.

MOTIVATION: One of the most important issues in computational proteomics is to produce a prediction model for the classification or annotation of biological function of novel protein sequences. In order to improve the prediction accuracy, much attention has been paid to the improvement of the performance of the algorithms used, few is for solving the fundamental issue, namely, amino acid encoding as most existing pattern recognition algorithms are unable to recognize amino acids in protein sequences. Importantly, the most commonly used amino acid encoding method has the flaw that leads to large computational cost and recognition bias. RESULTS: By replacing kernel functions of support vector machines (SVMs) with amino acid similarity measurement matrices, we have modified SVMs, a new type of pattern recognition algorithm for analysing protein sequences, particularly for proteolytic cleavage site prediction. We refer to the modified SVMs as bio-support vector machine. When applied to the prediction of HIV protease cleavage sites, the new method has shown a remarkable advantage in reducing the model complexity and enhancing the model robustness.

Algorithms↗

Predicting the linkage sites in glycoproteins using bio-basis function neural network.

MOTIVATION: Although, it is known that O-glycosidically linked oligosaccharides are commonly conjugated to a serine, threonine or hydroxylysine residue of the polypeptide, the chemical nature of the anchoring monosaccharide and the size of the oligosaccharide unit varies. Among different types, O-linked or mucin-type oligosaccharides are intimately involved in the secretion of proteins, be they enzymes, hormones or structural glycoproteins. Knowledge of the linkage sites in glycoproteins is critical to the design of specific and efficient inhibitors against the enzyme to catalyse the formation of the carbohydrate-peptide linkage. RESULTS: We present a method for predicting the linkage sites in O-linked glycoproteins using bio-basis function neural networks. The mean prediction accuracy of this method is 91.15 +/- 2.75% while it is 82.28 +/- 6.45% using back-propagation neural networks. Importantly, this method has significantly reduced the CPU time for modelling.

Algorithms↗

Polyprotein cleavage mechanism of SARS CoV Mpro and chemical modification of the octapeptide.

The cleavage mechanism of severe acute respiratory syndrome (SARS) coronavirus main proteinase (M(pro) or 3CL(pro)) for the octapeptide AVLQSGFR is studied using molecular mechanics (MM) and quantum mechanics (QM). The catalytic dyad His-41 and Cys-145 in the active pocket between domain I and II seem to polarize the pi-electron density of the peptide bond between Gln and Ser in the octapeptide, leading to an increase of positive charge on C(CO) of Gln and negative charge on N(NH) of Ser. The possibility of enhancing the chemical bond between Gln and Ser based on the "distorted key" theory [Anal. Biochem. 233 (1996) 1] is examined. The scissile peptide bond between Gln and Ser is found to be solidified through "hybrid peptide bond" by changing the carbonyl group CO of Gln to CH(2) or CF(2). This leads to a break of the pi-bond system for the peptide bond, making the octapeptide (AVLQSGFR) a "distorted key" and a potential starting system for the design of anti SARS drugs.

Amino Acid Sequence↗

Predicting enzyme family class in a hybridization space.

Given the sequence of a protein, how can we predict whether it is an enzyme or a non-enzyme? If it is, what enzyme family class it belongs to? Because these questions are closely relevant to the biological function of a protein and its acting object, their importance is self-evident. Particularly with the explosion of protein sequences entering into data banks and the relatively much slower progress in using biochemical experiments to determine their functions, it is highly desired to develop an automated method that can be used to give fast answers to these questions. By hybridizing the gene ontology and pseudo-amino-acid composition, we have introduced a new method that is called GO-PseAA predictor and operate it in a hybridization space. To avoid redundancy and bias, demonstrations were performed on a data set in which none of the proteins in an individual class has > or =40% sequence identity to any other. The overall success rate thus obtained by the jackknife cross-validation test in identifying enzyme and non-enzyme was 93%, and that in identifying the enzyme family was 94% for the following six main Enzyme Commission (EC) classes: (1) oxidoreductase, (2) transferase, (3) hydrolase, (4) lyase, (5) isomerase, and (6) ligase. The corresponding rates by the independent data set test were 98% and 97%, respectively.

Amino Acid Sequence↗

Structural bioinformatics and its impact to biomedical science.

During the last two decades, the number of sequence-known proteins has increased rapidly. In contrast, the corresponding increment for structure-known proteins is much slower. The unbalanced situation has critically limited our ability to understand the molecular mechanism of proteins and conduct structure-based drug design by timely using the updated information of newly found sequences. Therefore, it is highly desired to develop an automated method for fast deriving the 3D (3-dimensional) structure of a protein from its sequence. Under such a circumstance, the structural bioinformatics was emerging naturally as the time required. In this review, three main strategies developed in structural bioinformatics, i.e., pure energetic approach, heuristic approach, and homology modeling approach, as well as their underlying principles, are briefly introduced. Meanwhile, a series of demonstrations are presented to show how the structural bioinformatics has been applied to timely derive the 3D structures of some functionally important proteins, helping to understand their action mechanisms and stimulating the course of drug discovery. Also, the limitation of these approaches and the future challenges of structural bioinformatics are briefly addressed.

Amino Acid Sequence↗

P-selectin cell adhesion molecule in inflammation, thrombosis, cancer growth and metastasis.

P-selectin (CD62P) is a member of the selectin family of cell adhesion molecules. It is expressed on stimulated endothelial cells and activated platelets and mediates leukocyte rolling on stimulated endothelial cells and heterotypic aggregation of activated platelets onto leukocytes. It also mediates heterotypic aggregation of activated platelets to cancer cells and adhesion of cancer cells to stimulated endothelial cells. Using P-selectin knockout mice, the importance of P-selectin-mediated cell adhesive interactions in the pathogeneses of inflammation, thrombosis, growth and metastasis of cancers has been clearly demonstrated. Here we will summarize the current knowledge and highlight the important progress in the biomedical research of P-selectin biology, providing a suitable target for therapeutic interventions developed through both experimental and bioinformatic approaches.

Animals↗

Prediction and classification of protein subcellular location-sequence-order effect and pseudo amino acid composition.

Given a protein sequence, how to identify its subcellular location? With the rapid increase in newly found protein sequences entering into databanks, the problem has become more and more important because the function of a protein is closely correlated with its localization. To practically deal with the challenge, a dataset has been established that allows the identification performed among the following 14 subcellular locations: (1) cell wall, (2) centriole, (3) chloroplast, (4) cytoplasm, (5) cytoskeleton, (6) endoplasmic reticulum, (7) extracellular, (8) Golgi apparatus, (9) lysosome, (10) mitochondria, (11) nucleus, (12) peroxisome, (13) plasma membrane, and (14) vacuole. Compared with the datasets constructed by the previous investigators, the current one represents the largest in the scope of localizations covered, and hence many proteins which were totally out of picture in the previous treatments, can now be investigated. Meanwhile, to enhance the potential and flexibility in taking into account the sequence-order effect, the series-mode pseudo-amino-acid-composition has been introduced as a representation for a protein. High success rates are obtained by the re-substitution test, jackknife test, and independent dataset test, respectively. It is anticipated that the current automated method can be developed to a high throughput tool for practical usage in both basic research and pharmaceutical industry.

Algorithms↗

A new hybrid approach to predict subcellular localization of proteins by incorporating gene ontology.

Based on the recent development in the gene ontology and functional domain databases, a new hybridization approach is developed for predicting protein subcellular location by combining the gene product, functional domain, and quasi-sequence-order effects. As a showcase, the same prokaryotic and eukaryotic datasets, which were studied by many previous investigators, are used for demonstration. The overall success rate by the jackknife test for the prokaryotic set is 94.7% and that for the eukaryotic set 92.9%. These are so far the highest success rates achieved for the two datasets by following a rigorous cross-validation test procedure, suggesting that such a hybrid approach may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology. The very high success rates also reflect the fact that the subcellular localization of a protein is closely correlated with: (1). the biological objective to which the gene or gene product contributes, (2). the biochemical activity of a gene product, and (3). the place in the cell where a gene product is active.

Algorithms↗

Predicting protein quaternary structure by pseudo amino acid composition.

In the protein universe, many proteins are composed of two or more polypeptide chains, generally referred to as subunits, that associate through noncovalent interactions and, occasionally, disulfide bonds. With the number of protein sequences entering into data banks rapidly increasing, we are confronted with a challenge: how to develop an automated method to identify the quaternary attribute for a new polypeptide chain (i.e., whether it is formed just as a monomer, or as a dimer, trimer, or any other oligomer). This is important, because the functions of proteins are closely related to their quaternary attribute. For example, some critical ligands only bind to dimers but not to monomers; some marvelous allosteric transitions only occur in tetramers but not other oligomers; and some ion channels are formed by tetramers, whereas others are formed by pentamers. To explore this problem, we adopted the pseudo amino acid composition originally proposed for improving the prediction of protein subcellular location (Chou, Proteins, 2001; 43:246-255). The advantage of using the pseudo amino acid composition to represent a protein is that it has paved a way that can take into account a considerable amount of sequence-order effects to significantly improve prediction quality. Results obtained by resubstitution, jack-knife, and independent data set tests, have indicated that the current approach might be quite promising in dealing with such an extremely complicated and difficult problem.

Algorithms↗

Binding mechanism of coronavirus main proteinase with ligands and its implication to drug design against SARS.

In order to stimulate the development of drugs against severe acute respiratory syndrome (SARS), based on the atomic coordinates of the SARS coronavirus main proteinase determined recently [Science 13 (May) (2003) (online)], studies of docking KZ7088 (a derivative of AG7088) and the AVLQSGFR octapeptide to the enzyme were conducted. It has been observed that both the above compounds interact with the active site of the SARS enzyme through six hydrogen bonds. Also, a clear definition of the binding pocket for KZ7088 has been presented. These findings may provide a solid basis for subsite analysis and mutagenesis relative to rational design of highly selective inhibitors for therapeutic application. Meanwhile, the idea of how to develop inhibitors of the SARS enzyme based on the knowledge of its own peptide substrates (the so-called "distorted key" approach) was also briefly elucidated.

Amino Acid Sequence↗

Nearest neighbour algorithm for predicting protein subcellular location by combining functional domain composition and pseudo-amino acid composition.

In this paper, based on the approach by combining the "functional domain composition" [K.C. Chou, Y. D. Cai, J. Biol. Chem. 277 (2002) 45765] and the pseudo-amino acid composition [K.C. Chou, Proteins Struct. Funct. Genet. 43 (2001) 246; Correction Proteins Struct. Funct. Genet. 2044 (2001) 2060], the Nearest Neighbour Algorithm (NNA) was developed for predicting the protein subcellular location. Very high success rates were observed, suggesting that such a hybrid approach may become a useful high-throughput tool in the area of bioinformatics and proteomics.

Algorithms↗

Prediction of protein secondary structure content by artificial neural network.

The neural network method was applied to the prediction of the content of protein secondary structure elements, including alpha-helix, beta-strand, beta-bridge, 3(10)-helix, pi-helix, H-bonded turn, bend, and random coil. The "pair-coupled amino acid composition" originally proposed by K. C. Chou [J Protein Chem 1999, 18, 473] was adopted as the input. Self-consistency and independent-dataset tests were used to appraise the performance of the neural network. Results of both tests indicated high performance of the method.

Algorithms↗