Search PubMed⌕ Search

Biomedical subjects

Kuo-Chen Chou

Publications and source records attributed to Kuo-Chen Chou.

At least 37 records · Page 2Linked to original sources

A novel fingerprint map for detecting SARS-CoV.

Spike (S) protein is the most important membrane protein on the surface of severe acute respiratory syndrome coronavirus (SARS-CoV). It associates with cellular receptors to mediate infection of their target cells. Inspired by such a mechanism, an in-depth investigation into the genome sequences of S protein of SARS-CoV and its receptor are conducted thru a mathematical transformation and graphic approach. As an outcome, a novel method for visualizing the characteristic of SARS-CoV is suggested. An extensive comparison among a large number of genome sequences has proved that the characteristic thus revealed is unique for SARS-CoV. As such, the characteristic can be regarded as the fingerprint map of SARS-CoV for diagnostic usage. Moreover, the conclusion has been further supported in a real case in Guangdong province of China. The fingerprint map proposed here has the merits of clear visibility and reliability that can serve as a complementary clinical tool for detecting SARS-CoV, particularly for the cases where the results obtained by the conventional methods are uncertain or conflicted with each other.

Algorithms↗

Prediction of protease types in a hybridization space.

Regulating most physiological processes by controlling the activation, synthesis, and turnover of proteins, proteases play pivotal regulatory roles in conception, birth, digestion, growth, maturation, ageing, and death of all organisms. Different types of proteases have different functions and biological processes. Therefore, it is important for both basic research and drug discovery to consider the following two problems. (1) Given the sequence of a protein, can we identify whether it is a protease or non-protease? (2) If it is, what protease type does it belong to? Although the two problems can be solved by various experimental means, it is both time-consuming and costly to do so. The avalanche of protein sequences generated in the post-genetic era has challenged us to develop an automated method for making a fast and reliable identification. By hybridizing the functional domain composition and pseudo-amino acid composition, we have introduced a new method called "FunD-PseAA predictor" that is operated in a hybridization space. To avoid redundancy and bias, demonstrations were performed on a dataset where none of the proteins has >or=25% sequence identity to any other. The overall success rate thus obtained by the jackknife cross-validation test in identifying protease and non-protease was 92.95%, and that in identifying the protease type was 94.75% among the following six types: (1) aspartic, (2) cysteine, (3) glutamic, (4) metallo, (5) serine, and (6) threonine. Demonstration was also made on an independent dataset, and the corresponding overall success rates were 98.36% and 97.11%, respectively, suggesting the FunD-PseAA predictor is very powerful and may become a useful tool in bioinformatics and proteomics.

Algorithms↗

Low-frequency Fourier spectrum for predicting membrane protein types.

Cell membranes are vitally important to living cells. Although the infrastructure of biological membrane is provided by the lipid bilayer, membrane proteins perform most of the specific functions. Knowledge of membrane protein types often provides crucial hints toward determining the function of an uncharacterized membrane protein. With the avalanche of new protein sequences generated in the post-genomic era, it is highly demanded to develop a high throughput tool in identifying the type of newly found membrane proteins according to their primary sequences, so as to timely annotate them for reference usage in both basic research and drug discovery. To realize this, the key is to establish a powerful identifier that can catch their characteristic sequence patterns for different membrane protein types. However, it is not easy because they are buried in a pile of long and complicated sequences. In this paper, based on the concept of the pseudo-amino acid composition [K.C. Chou, PROTEINS: Struct., Funct., Genet. 43 (2001) 246-255], the low-frequency Fourier spectrum analysis is introduced. The merits by doing so are that the sequence pattern information can be more effectively incorporated into a set of discrete components, and that all the existing prediction algorithms can be straightforwardly used on such a formulation for protein samples. High success rates were observed by the re-substitution test, jackknife test, and independent dataset test, indicating that the low-frequency Fourier spectrum approach may become a very useful tool for membrane protein type prediction. The novel approach also holds a high potential for predicting many other attributes of proteins.

Fourier Analysis↗

Prediction of protein signal sequences and their cleavage sites by statistical rulers.

Functioning as an "address tag" or "zip code" that guides nascent proteins (newly synthesized proteins in the cytosol) to wherever they are needed, signal peptides (also called targeting signals or signal sequences) have become a crucial tool in finding new drugs or reprogramming cells for gene therapy. To effectively and timely use such a tool, however, the first important thing is to develop an automated method for quickly and accurately identifying the signal peptide for a given nascent protein. With the avalanche of new protein sequences generated in the post-genomic era, the challenge has become even more urgent and critical. In this paper, five statistical rulers were derived via performing a mutual information analysis. By combining these statistical rulers, a new prediction algorithm was established and high success prediction rates were observed. The new algorithm may play a complementary role to the existing algorithms in this area. It is anticipated that the mutual information approach introduced here may be very useful for studying many other sequence-coupling problems in molecular biology as well.

Algorithms↗

Theoretical studies of Alzheimer's disease drug candidate 3-[(2,4-dimethoxy)benzylidene]-anabaseine (GTS-21) and its derivatives.

Theoretical and molecular modeling studies have been conducted for understanding the details of how 3-[(2,4-dimethoxy)benzylidene]-anabaseine dihydrochloride (GTS-21) and its metabolism derivatives bind with the receptor of alpha7 nicotinic acetylcholine dimer. Good accordance with experimental results has been achieved. It was found that the van der Waals repulsion makes the dominant contribution to the binding energy. GTS-21 and its metabolites are apparently too large for the binding sites of the alpha7 dimer. To improve the effectiveness of the drug, a possible approach is to reduce its volume while maintaining the presence of the active groups. Our studies, in combination with experimental studies, will lead to a promising basis for practical drug design against Alzheimer's disease.

Alzheimer Disease↗

Predicting protein subnuclear location with optimized evidence-theoretic K-nearest classifier and pseudo amino acid composition.

The nucleus is the brain of eukaryotic cells that guides the life processes of the cell by issuing key instructions. For in-depth understanding of the biochemical process of the nucleus, the knowledge of localization of nuclear proteins is very important. With the avalanche of protein sequences generated in the post-genomic era, it is highly desired to develop an automated method for fast annotating the subnuclear locations for numerous newly found nuclear protein sequences so as to be able to timely utilize them for basic research and drug discovery. In view of this, a novel approach is developed for predicting the protein subnuclear location. It is featured by introducing a powerful classifier, the optimized evidence-theoretic K-nearest classifier, and using the pseudo amino acid composition [K.C. Chou, PROTEINS: Structure, Function, and Genetics, 43 (2001) 246], which can incorporate a considerable amount of sequence-order effects, to represent protein samples. As a demonstration, identifications were performed for 370 nuclear proteins among the following 9 subnuclear locations: (1) Cajal body, (2) chromatin, (3) heterochromatin, (4) nuclear diffuse, (5) nuclear pore, (6) nuclear speckle, (7) nucleolus, (8) PcG body, and (9) PML body. The overall success rates thus obtained by both the re-substitution test and jackknife cross-validation test are significantly higher than those by existing classifiers on the same working dataset. It is anticipated that the powerful approach may also become a useful high throughput vehicle to bridge the huge gap occurring in the post-genomic era between the number of gene sequences in databases and the number of gene products that have been functionally characterized. The OET-KNN classifier will be available at www.pami.sjtu.edu.cn/people/hbshen.

Amino Acid Sequence↗

Fuzzy KNN for predicting membrane protein types from pseudo-amino acid composition.

Cell membranes are vitally important to the life of a cell. Although the basic structure of biological membrane is provided by the lipid bilayer, membrane proteins perform most of the specific functions. Membrane proteins are putatively classified into five different types. Identification of their types is currently an important topic in bioinformatics and proteomics. In this paper, based on the concept of representing protein samples in terms of their pseudo-amino acid composition, the fuzzy K-nearest neighbors (KNN) algorithm has been introduced to predict membrane protein types, and high success rates were observed. It is anticipated that, the current approach, which is based on a branch of fuzzy mathematics and represents a new strategy, may play an important complementary role to the existing methods in this area. The novel approach may also have notable impact on prediction of the other attributes, such as protein structural class, protein subcellular localization, and enzyme family class, among many others.

Algorithms↗

Using supervised fuzzy clustering to predict protein structural classes.

Prediction of protein classification is both an important and a tempting topic in protein science. This is because of not only that the knowledge thus obtained can provide useful information about the overall structure of a query protein, but also that the practice itself can technically stimulate the development of novel predictors that may be straightforwardly applied to many other relevant areas. In this paper, a novel approach, the so-called "supervised fuzzy clustering approach" is introduced that is featured by utilizing the class label information during the training process. Based on such an approach, a set of "if-then" fuzzy rules for predicting the protein structural classes are extracted from a training dataset. It has been demonstrated through two different working datasets that the overall success prediction rates obtained by the supervised fuzzy clustering approach are all higher than those by the unsupervised fuzzy c-means introduced by the previous investigators [C.T. Zhang, K.C. Chou, G.M. Maggiora. Protein Eng. (1995) 8, 425-435]. It is anticipated that the current predictor may play an important complementary role to other existing predictors in this area to further strengthen the power in predicting the structural classes of proteins and their other characteristic attributes.

Algorithms↗

Boosting classifier for predicting protein domain structural class.

A novel classifier, the so-called "LogitBoost" classifier, was introduced to predict the structural class of a protein domain according to its amino acid sequence. LogitBoost is featured by introducing a log-likelihood loss function to reduce the sensitivity to noise and outliers, as well as by performing classification via combining many weak classifiers together to build up a very strong and robust classifier. It was demonstrated thru jackknife cross-validation tests that LogitBoost outperformed other classifiers including "support vector machine," a very powerful classifier widely used in biological literatures. It is anticipated that LogitBoost can also become a useful vehicle in classifying other attributes of proteins according to their sequences, such as subcellular localization and enzyme family class, among many others.

Algorithms↗

Using optimized evidence-theoretic K-nearest neighbor classifier and pseudo-amino acid composition to predict membrane protein types.

Knowledge of membrane protein type often provides crucial hints toward determining the function of an uncharacterized membrane protein. With the avalanche of new protein sequences emerging during the post-genomic era, it is highly desirable to develop an automated method that can serve as a high throughput tool in identifying the types of newly found membrane proteins according to their primary sequences, so as to timely make the relevant annotations on them for the reference usage in both basic research and drug discovery. Based on the concept of pseudo-amino acid composition [K.C. Chou, Proteins: Struct. Funct. Genet. 43 (2001) 246-255; Erratum: Proteins: Struct. Funct. Genet. 44 (2001) 60] that has made it possible to incorporate a considerable amount of sequence-order effects by representing a protein sample in terms of a set of discrete numbers, a novel predictor, the so-called "optimized evidence-theoretic K-nearest neighbor" or "OET-KNN" classifier, was proposed. It was demonstrated via the self-consistency test, jackknife test, and independent dataset test that the new predictor, compared with many previous ones, yielded higher success rates in most cases. The new predictor can also be used to improve the prediction quality for, among many other protein attributes, structural class, subcellular localization, enzyme family class, and G-protein coupled receptor type. The OET-KNN classifier will be available as a web-server at http://www.pami.sjtu.edu.cn/kcchou.

Algorithms↗

Using LogitBoost classifier to predict protein structural classes.

Prediction of protein classification is an important topic in molecular biology. This is because it is able to not only provide useful information from the viewpoint of structure itself, but also greatly stimulate the characterization of many other features of proteins that may be closely correlated with their biological functions. In this paper, the LogitBoost, one of the boosting algorithms developed recently, is introduced for predicting protein structural classes. It performs classification using a regression scheme as the base learner, which can handle multi-class problems and is particularly superior in coping with noisy data. It was demonstrated that the LogitBoost outperformed the support vector machines in predicting the structural classes for a given dataset, indicating that the new classifier is very promising. It is anticipated that the power in predicting protein structural classes as well as many other bio-macromolecular attributes will be further strengthened if the LogitBoost and some other existing algorithms can be effectively complemented with each other.

Animals↗

Predicting membrane protein type by functional domain composition and pseudo-amino acid composition.

Given the sequence of a protein, how can we predict whether it is a membrane protein or non-membrane protein? If it is, what membrane protein type it belongs to? Since these questions are closely relevant to the function of an uncharacterized protein, their importance is self-evident. Particularly, with the explosion of protein sequences entering into databanks and the relatively much slower progress in using biochemical experiments to determine their functions, it is highly desired to develop an automated method that can be used to give a fast answers to these questions. By hybridizing the functional domain (FunD) and pseudo-amino acid composition (PseAA), a new strategy called FunD-PseAA predictor was introduced. To test the power of the predictor, a highly non-homologous data set was constructed where none of proteins has 25% sequence identity to any other. The overall success rates obtained with the FunD-PseAA predictor on such a data set by the jackknife cross-validation test was 85% for the case in identifying membrane protein and non-membrane protein, and 91% in identifying the membrane protein type among the following 5 categories: (1) type-1 membrane protein, (2) type-2 membrane protein, (3) multipass transmembrane protein, (4) lipid chain-anchored membrane protein, and (5) GPI-anchored membrane protein. These rates are much higher than those obtained by the other methods on the same stringent data set, indicating that the FunD-PseAA predictor may become a useful high throughput tool in bioinformatics and proteomics.

Amino Acid Sequence↗

Modeling the tertiary structure of human cathepsin-E.

Cathepsin-E is an endolysosomal aspartic proteinase and is predominantly expressed in immune system cells. Deficiency of cathepsin-E is associated with the development of atopic dermatitis, a pruritic inflammatory skin disease, which has put us to face a high selectivity challenge in the development of drugs for the therapy of Alzheimer's disease or breast cancer. This is because BACE1 (also known as beta-secretase) and cathepsin-D, both belonging to the family of aspartic proteinases, might interact with the same compound as cathepsin-E does. BACE1 is a putative prime therapeutic target for the treatment of Alzheimer's disease, and cathepsin-D a potential target for breast cancer. Accordingly, in the course of finding drugs against Alzheimer's disease or breast cancer by inhibiting BACE1 or cathepsin-D, the desired drugs should selectively inhibit only BACE1 or cathepsin-D, but definitely not cathepsin-E. To realize this, it is indispensable to find out the structural difference of the three enzymes. Since the crystal structures of BACE1 and cathepsin-D are already known, the lack of three-dimensional structure of cathepsin-E has become the bottleneck in this regard. In view of this, the three-dimensional structure of cathepsin-E has been developed. Although the overall structures of the three enzymes are quite similar to each other, some subtle difference around their active sites that distinguishes cathepsin-E from cathepsin-D and BACE1 has been revealed through an analysis of hydrogen bond network and microenvironment. The computed three-dimensional structure of cathepsin-E and the relevant findings might provide useful insights for designing inhibitors with the desired selectivity.

Amino Acid Sequence↗

Heuristic molecular lipophilicity potential (HMLP): a 2D-QSAR study to LADH of molecular family pyrazole and derivatives.

The quantum chemical and structure-based technique heuristic molecular lipophilicity potential (HMLP) is used in the liver alcohol dehydrogenase (LADH) study of molecular family pyrazole and derivatives. The molecular lipophilic index LM, molecular hydrophilic index HM, lipophilic indices lss, and hydrophilic indices hss of the substitutes (fragments), and atomic lipophilicity indices las are constructed and used in QSAR study. The HMLP indices are correlated with bioactivities of 18 pyrazole derivatives according to the 2D QSAR procedure. The multiple linear regression equation between the bioactivities of pyrazole derivatives and HMLP indices are built using partial least square (PLS) with the optimal statistical quantity (r=0.987, s=0.479, F=47.19). The inhibition mechanism of LADH of the pyrazole derivatives is explained according to the physical meaning of HMLP indices. During the HMLP calculations for the 2D QSAR, the only input parameters are the atomic van der Waals radius without the need to resort to any empirical parameters. Accordingly, HMLP can provide a rigorous theoretical approach with a crystal clear physical meaning for the 2D QSAR.

Alcohol Dehydrogenase↗

An application of gene comparative image for predicting the effect on replication ratio by HBV virus gene missense mutation.

Hepatitis B viruses (HBVs) show instantaneous and high-ratio mutations when they are replicated, some sorts of which significantly affect the efficiency of virus replication through enhancing or depressing the viral replication, while others have no influence at all. The mechanism of gene expression is closely correlated with its gene sequence. With the rapid increase in the number of newly found sequences entering into data banks, it is highly desirable to develop an automated method for simulating the gene regulating function. The establishment of such a predictor will no doubt expedite the process of prioritizing genes and proteins identified by genomics efforts as potential molecular targets for drug design. Based on the power of cellular automata (CA) in treating complex systems with simple rules, a novel method to present HBV gene image has been introduced. The results show that the images thus obtained can very efficiently simulate the effects of the gene missense mutation on the virus replication. It is anticipated that CA may also serve as a useful vehicle for many other studies on complicated biological systems.

Computational Biology↗

Using GO-PseAA predictor to identify membrane proteins and their types.

Cell membranes are crucial to the life of a cell. Although the basic structure of biological membrane is provided by the lipid bilayer, most of the specific functions are carried out by membrane proteins. Knowledge of membrane protein type often offers important clues toward determining the function of an uncharacterized protein. Therefore, predicting the type of a membrane protein from its primary sequence, or even just identifying whether the uncharacterized protein belongs to a membrane protein or not, is an important and challenging problem in bioinformatics and proteomics. To deal with these problems, the GO-PseAA predictor is introduced that is operated in a hybridization space by combining the gene ontology and pseudo amino acid composition. Meanwhile, to test the prediction quality, a dataset was constructed that contains 6476 non-membrane proteins and 5122 membrane proteins classified into five different types. To avoid redundancy and bias, none of the proteins included has > or = 40% sequence identity to any other. It has been observed that the overall success rate by the jackknife cross-validation test in identifying non-membrane proteins and membrane proteins was 94.76%, and that in identifying the five membrane protein types was 95.84%. The high success rates suggest that the GO-PseAA predictor can catch the core feature of the statistical samples concerned and may become an automated high throughput toll in molecular and cell biology.

Algorithms↗

Molecular modeling and chemical modification for finding peptide inhibitor against severe acute respiratory syndrome coronavirus main proteinase.

Severe acute respiratory syndrome (SARS) is a respiratory disease caused by a newly found virus, called SARS coronavirus. In this study, the cleavage mechanism of the SARS coronavirus main proteinase (Mpro or 3CLpro) on the octapeptide NH2-AVLQ downward arrowSGFR-COOH was investigated using molecular mechanics and quantum mechanics simulations based on the experimental structure of the proteinase. It has been observed that the catalytic dyad (His-41/Cys-145) site between domains I and II attracts the pi electron density from the peptide bond Gln-Ser, increasing the positive charge on C(CO) of Gln and the negative charge on N(NH) of Ser, so as to weaken the Gln-Ser peptide bond. The catalytic functional group is the imidazole group of His-41 and the S in Cys-145. Ndelta1 on the imidazole ring plays the acid-base catalytic role. Based on the "distorted key theory" [K.C. Chou, Anal. Biochem. 233 (1996) 1-14], the possibility to convert the octapeptide to a competent inhibitor has been studied. It has been found that the chemical bond between Gln and Ser will become much stronger and no longer cleavable by the SARS enzyme after either changing the carbonyl group CO of Gln to CH2 or CF2 or changing the NH of Ser to CH2 or CF2. The octapeptide thus modified might become an effective inhibitor or a potential drug candidate against SARS.

Acute Disease↗

Predicting enzyme family classes by hybridizing gene product composition and pseudo-amino acid composition.

A new method has been developed to predict the enzymatic attribute of proteins by hybridizing the gene product composition and pseudo amino acid composition. As a demonstration, a working dataset was generated with a cutoff of 60% sequence identity to avoid redundancy and bias in statistical prediction. The dataset thus constructed contains 39989 protein sequences, of which 27469 are non-enzymes and 12520 enzymes that were further classified into 6 enzyme family classes according to their 6 main EC (Enzyme Commission) numbers (2314 are oxidoreductases, 3653 transferases, 3246 hydrolases, 1307 lyases, 676 isomerases, and 1324 ligases). The overall success rate by the jackknife test for the identification between enzyme and non-enzyme was 94%, and that for the identification among the 6 enzyme family classes was 98%. It is anticipated that, with the rapid increase of protein sequences entering into databanks, the current method will become a useful automated tool in identifying the enzymatic attribute of a newly found protein sequence.

Algorithms↗