Search PubMed⌕ Search

Biomedical subjects

Kuo-Chen Chou

Publications and source records attributed to Kuo-Chen Chou.

At least 55 records · Page 3Linked to original sources

SLLE for predicting membrane protein types.

Introduction of the concept of pseudo amino acid composition (PROTEINS: Structure, Function, and Genetics 43 (2001) 246; Erratum: ibid. 44 (2001) 60) has made it possible to incorporate a considerable amount of sequence-order effects by representing a protein sample in terms of a set of discrete numbers, and hence can significantly enhance the prediction quality of membrane protein type. As a continuous effort along such a line, the Supervised Locally Linear Embedding (SLLE) technique for nonlinear dimensionality reduction is introduced (Science 22 (2000) 2323). The advantage of using SLLE is that it can reduce the operational space by extracting the essential features from the high-dimensional pseudo amino acid composition space, and that the cluster-tolerant capacity can be increased accordingly. As a consequence by combining these two approaches, high success rates have been observed during the tests of self-consistency, jackknife and independent data set, respectively, by using the simplest nearest neighbour classifier. The current approach represents a new strategy to deal with the problems of protein attribute prediction, and hence may become a useful vehicle in the area of bioinformatics and proteomics.

Algorithms↗

Using fourier spectrum analysis and pseudo amino acid composition for prediction of membrane protein types.

Membrane proteins are generally classified into the following five types: (1) type I membrane protein, (2) type II membrane protein, (3) multipass transmembrane proteins, (4) lipid chain-anchored membrane proteins, and (5) GPI-anchored membrane proteins. Given the sequence of an uncharacterized membrane protein, how can we identify which one of the above five types it belongs to? This is important because the biological function of a membrane protein is closely correlated with its type. Particularly, with the explosion of protein sequences entering into databanks, it is in high demand to develop an automated method to address this problem. To realize this, the key is to catch the statistical characteristics for each of the five types. However, it is not easy because they are buried in a pile of long and complicated sequences. In this paper, based on the concept of the pseudo amino acid composition (Chou, K. C. (2001). PROTEINS: Structure, Function, and Genetics 43: 246-255), the technique of Fourier spectrum analysis is introduced. By doing so, the sample of a protein is represented by a set of discrete components that can incorporate a considerable amount of the sequence order effects as well as its amino acid composition information. On the basis of such a statistical frame, the support vector machine (SVM) is introduced to perform predictions. High success rates were yielded by the self-consistency test, jackknife test, and independent dataset test, suggesting that the current approach holds a promising potential to become a high throughput tool for membrane protein type prediction as well as other related areas.

Amino Acids↗

Assessment of chemical libraries for their druggability.

High throughput virtual screening is acknowledged as the initial means for identifying hit compounds that will be eventually transformed to leads or drug candidates. To improve quality of screening, it is essential to have powerful methods for the analysis of the compound databases. For this purpose, we have developed a novel and practical scoring function to assess the druggability of compounds. The proposed function consists of 12 metrics that take into account physical, chemical and structural properties as well as the presence of undesirable functional groups. We have applied this 12-metric scoring function to 44 different databases that include more than 3.8 million compounds, which are commercially available. The overall quality of each database was evaluated according to the score and rank measured by our 12-metric function. Our findings suggest that, the majority of compounds that do not satisfy druggable rules do so due to high molecular weight, high logP values and the presence of reactive functional groups.

Chemistry, Pharmaceutical↗

Computational methods for protein-protein interaction and their application.

Protein-protein interactions play a central role in numerous processes in cell and are one of the main research fields in current functional proteomics. The increase of finished genomic sequences has greatly stimulated the progress for detecting the functions of the genes and their encoded proteins. As complementary ways to the high through-put experimental methods, various methods of bioinformatics have been developed for the study of the protein-protein interaction. These methods range from the sequence homology-based to the genomic-context based. Recently, it tends to integrate the data from different methods to build the protein-protein interaction network, and to predict the protein function from the analysis of the network structure. Efforts are ongoing to improve these methods and to search for novel aspects in genomes that could be exploited for function prediction. This review highlights the recent advances of the bioinformatics methods in protein-protein interaction researches. In the end, the application of the protein-protein interaction has also been discussed.

Computational Biology↗

Pattern recognition methods for protein functional site prediction.

Protein functional site prediction is closely related to drug design, hence to public health. In order to save the cost and the time spent on identifying the functional sites in sequenced proteins in biology laboratory, computer programs have been widely used for decades. Many of them are implemented using the state-of-the-art pattern recognition algorithms, including decision trees, neural networks and support vector machines. Although the success of this effort has been obvious, advanced and new algorithms are still under development for addressing some difficult issues. This review will go through the major stages in developing pattern recognition algorithms for protein functional site prediction and outline the future research directions in this important area.

Algorithms↗

HIV-1 gp120 V3 loop for structure-based drug design.

HIV-1 cell entry is mediated by sequential interactions of the envelope protein gp120 with the receptor CD4 and a coreceptor, usually CCR5 or CXCR4, depending on the individual virion. Considerable efforts on exploiting the HIV coreceptors as drug targets have led to the new class of coreceptor antagonists. While these antiretroviral drugs aim at preventing virus/coreceptor interaction by binding to host proteins, neutralizing antibodies directed against the coreceptor-binding sites on gp120 have attracted attention as possible vaccine candidates. However, both approaches are complicated by the multiple protective mechanisms of gp120 which allow for rapid escape from selective pressures exerted by drugs or antibodies. Thus, advances in rational drug and vaccine design rely heavily on improved insights into the relation between genotype and phenotype, the evolution of coreceptor usage, and, ultimately the structural biology of coreceptor usage and inhibition. The third variable (V3) loop of gp120, crucially involved in all these aspects, will be a major focus of this review.

Antibodies↗

Progress in protein structural class prediction and its impact to bioinformatics and proteomics.

The structural class is an important attribute used to characterize the overall folding type of a protein or its domain. Since the concept of protein structural class was developed about 3 decades ago based on a visual inspection of polypeptide chain topologies in a dataset of only 31 gloular proteins, the number of structure-known proteins has been increased rapidly. For example, as of 12-July-2005, the entries deposited into RCSB PDB Protein Data Bank for proteins, peptides, and viruses whose 3-dimensional structures were determined by X-ray and NMR techniques have been increased to 28,920. To properly cover more and more structure-known proteins, some modification and expansion from the original structural classification scheme have been developed. Meanwhile, many different approaches have been proposed for predicting the structural class of proteins. In this review, the new classification schemes are briefly introduced. The attention is focused on the progress in structural class prediction and its impact in stimulating the development of identifying the other attributes of proteins. It is interesting to point out that the development of the latter has actually in turn greatly enriched the power of the former. Also, some promising approaches for the further development of protein structural class prediction are also addressed.

Computational Biology↗

Selection of molecular descriptors with artificial intelligence for the understanding of HIV-1 protease peptidomimetic inhibitors-activity.

Quantitative Structure Activity Relationship (QSAR) techniques are used routinely by computational chemists in drug discovery and development to analyze datasets of compounds. Quantitative numerical methods like Partial Least Squares (PLS) and Artificial Neural Networks (ANN) have been used on QSAR to establish correlations between molecular properties and bioactivity. However, ANN may be advantageous over PLS because it considers the interrelations of the modeled variables. This study focused on the HIV-1 Protease (HIV-1 Pr) inhibitors belonging to the peptidomimetic class of compounds. The main objective was to select molecular descriptors with the best predictive value for antiviral potency (Ki). PLS and ANN were used to predict Ki activity of HIV-1 Pr inhibitors and the results were compared. To address the issue of dimensionality reduction, Genetic Algorithms (GA) were used for variable selection and their performance was compared against that of ANN. Finally, the structure of the optimum ANN achieving the highest Pearson's-R coefficient was determined. On the basis of Pearson's-R, PLS and ANN were compared to determine which exhibits maximum performance. Training and validation of models was performed on 15 random split sets of the master dataset consisted of 231 compounds. For each compound 192 molecular descriptors were considered. The molecular structure and constant of inhibition (Ki) were selected from the NIAID database. Study findings suggested that non-covalent interactions such as hydrophobicity, shape and hydrogen bonding describe well the antiviral activity of the HIV-1 Pr compounds. The significance of lipophilicity and relationship to HIV-1 associated hyperlipidemia and lipodystrophy syndrome warrant further investigation.

Algorithms↗

A new nucleotide-composition based fingerprint of SARS-CoV with visualization analysis.

It has been observed by conducting an extensive analysis of the two-dimensional cellular automata images of known SARS-CoV genome sequences that the V-shaped cross-lines only exist in some special locations, and hence can be used as a fingerprint to identify the SARS sequences. Such a discovery can be used to rapidly and reliably diagnose SARS coronavirus for both basic research in laboratories and practical application in clinics.

Algorithms↗

Application of bioinformatics in search for cleavable peptides of SARS-CoV M(pro) and chemical modification of octapeptides.

According to the "distorted key" theory as elaborated in a review article years ago (Chou, K.C.: Analytical Biochemistry, 1996, 233, 1-14), the knowledge of the cleavable peptides by SARS-CoV M(pro) (severe acute respiratory syndrome coronavirus main proteinase) can provide very useful insights on developing drugs against SARS. In view of this, the softwares, ZCURVE_CoV 1.0 and ZCURVE_CoV 2.0 (http://tubic.tju.edu.cn/sars/), developed recently for SARS-Coronavirus are used to analyze the 36 complete SARS-Coronavirus RNA sequences in the gene bank NCBI (http://www.ncbi.nlm.nih.gov/) from different sources for protein coding genes, and to search for the cleavage sites of SARS-CoV M(pro) in polyproteins pp1a and pp1ab. A total of 396 cleavage points are found in the 36 SARS-Coronavirus and 11 cleavable octapeptides abstracted from the 396 cleavage sites. The statistical distributions of amino acids for the cleavable octapeptides at the subsites R4, R3, R2, R1, R1', R2', R3' and R4' are calculated. The cleavage-specific positions are on R2, R1 and R1', and the positions R3 and R4 are featured by some certain specificity for SARS-CoV M(pro). The structural characters of amino acid residues around the cleavage-specific positions are discussed. Two most promising octapeptides, i.e., NH(2)-ATLQ downward arrowAIAS-COOH and NH(2)-ATLQ downward arrowAENV-COOH, are selected to be the candidates for chemical modification, converting into the inhibitors of SARS-CoV M(pro). A possible strategy to convert a cleavable octapeptide by SARS enzyme into a drug candidate against SARS is elucidated.

Amino Acid Sequence↗

Using GO-PseAA predictor to predict enzyme sub-class.

Enzyme function is much less conserved than anticipated, i.e., the requirement for sequence similarity that implies similarity in enzymatic function is much higher than the requirement that implies similarity in protein structure. This is because the function of an enzyme is an extremely complicated problem that may involve very subtle structural details as well as many other physical chemistry factors. Accordingly, if simply based on the sequence similarity approach, it would hardly get a decent success rate in predicting enzyme sub-class even for a dataset consisting of samples with 50% sequence identity. To cope with such a situation, the GO-PseAA predictor was adopted to identify the sub-class for each of the six main enzyme families. It has been observed that, even for the much more stringent datasets in which none of the enzymes has 25% sequence identity to any others, the overall success rates are 73-95%, suggesting that the GO-PseAA predictor can catch the core features of the statistical samples concerned and may become a useful high throughput tool in proteomics and bioinformatics.

Algorithms↗

Predicting protein localization in budding yeast.

MOTIVATION: Most of the existing methods in predicting protein subcellular location were used to deal with the cases limited within the scope from two to five localizations, and only a few of them can be effectively extended to cover the cases of 12-14 localizations. This is because the more the locations involved are, the poorer the success rate would be. Besides, some proteins may occur in several different subcellular locations, i.e. bear the feature of 'multiplex locations'. So far there is no method that can be used to effectively treat the difficult multiplex location problem. The present study was initiated in an attempt to address (1) how to efficiently identify the localization of a query protein among many possible subcellular locations, and (2) how to deal with the case of multiplex locations. RESULTS: By hybridizing gene ontology, functional domain and pseudo amino acid composition approaches, a new method has been developed that can be used to predict subcellular localization of proteins with multiplex location feature. A global analysis of the proteins in budding yeast classified into 22 locations was performed by jack-knife cross-validation with the new method. The overall success identification rate thus obtained is 70%. In contrast to this, the corresponding rates obtained by some other existing methods were only 13-14%, indicating that the new method is very powerful and promising. Furthermore, predictions were made for the four proteins whose localizations could not be determined by experiments, as well as for the 236 proteins whose localizations in budding yeast were ambiguous according to experimental observations. However, according to our predicted results, many of these 'ambiguous proteins' were found to have the same score and ranking for several different subcellular locations, implying that they may simultaneously exist, or move around, in these locations. This finding is intriguing because it reflects the dynamic feature of these proteins in a cell that may be associated with some special biological functions.

Algorithms↗

Predicting 22 protein localizations in budding yeast.

According to the recent experiments, proteins in budding yeast can be distinctly classified into 22 subcellular locations. Of these proteins, some bear the multi-locational feature, i.e., occur in more than one location. However, so far all the existing methods in predicting protein subcellular location were developed to deal with only the mono-locational case where a query protein is assumed to belong to one, and only one, subcellular location. To stimulate the development of subcellular location prediction, an augmentation procedure is formulated that will enable the existing methods to tackle the multi-locational problem as well. It has been observed thru a jackknife cross-validation test that the success rate obtained by the augmented GO-FnD-PseAA algorithm [BBRC 320 (2004) 1236] is overwhelmingly higher than those by the other augmented methods. It is anticipated that the augmented GO-FunD-PseAA predictor will become a very useful tool in predicting protein subcellular localization for both basic research and practical application.

Algorithms↗

Predicting protein structural class by functional domain composition.

The functional domain composition is introduced to predict the structural class of a protein or domain according to the following classification: all-alpha, all-beta, alpha/beta, alpha+beta, micro (multi-domain), sigma (small protein), and rho (peptide). The advantage by doing so is that both the sequence-order-related features and the function-related features are naturally incorporated in the predictor. As a demonstration, the jackknife cross-validation test was performed on a dataset that consists of proteins and domains with only less than 20% sequence identity to each other in order to get rid of any homologous bias. The overall success rate thus obtained was 98%. In contrast to this, the corresponding rates obtained by the simple geometry approaches based on the amino acid composition were only 36-39%. This indicates that using the functional domain composition to represent the sample of a protein for statistical prediction is very promising, and that the functional type of a domain is closely correlated with its structural class.

Databases, Protein↗

Weighted-support vector machines for predicting membrane protein types based on pseudo-amino acid composition.

Membrane proteins are generally classified into the following five types: (1) type I membrane proteins, (2) type II membrane proteins, (3) multipass transmembrane proteins, (4) lipid chain-anchored membrane proteins and (5) GPI-anchored membrane proteins. Prediction of membrane protein types has become one of the growing hot topics in bioinformatics. Currently, we are facing two critical challenges in this area: first, how to take into account the extremely complicated sequence-order effects, and second, how to deal with the highly uneven sizes of the subsets in a training dataset. In this paper, stimulated by the concept of using the pseudo-amino acid composition to incorporate the sequence-order effects, the spectral analysis technique is introduced to represent the statistical sample of a protein. Based on such a framework, the weighted support vector machine (SVM) algorithm is applied. The new approach has remarkable power in dealing with the bias caused by the situation when one subset in the training dataset contains many more samples than the other. The new method is particularly useful when our focus is aimed at proteins belonging to small subsets. The results obtained by the self-consistency test, jackknife test and independent dataset test are encouraging, indicating that the current approach may serve as a powerful complementary tool to other existing methods for predicting the types of membrane proteins.

Algorithms↗

Using amphiphilic pseudo amino acid composition to predict enzyme subfamily classes.

MOTIVATION: With protein sequences entering into databanks at an explosive pace, the early determination of the family or subfamily class for a newly found enzyme molecule becomes important because this is directly related to the detailed information about which specific target it acts on, as well as to its catalytic process and biological function. Unfortunately, it is both time-consuming and costly to do so by experiments alone. In a previous study, the covariant-discriminant algorithm was introduced to identify the 16 subfamily classes of oxidoreductases. Although the results were quite encouraging, the entire prediction process was based on the amino acid composition alone without including any sequence-order information. Therefore, it is worthy of further investigation. RESULTS: To incorporate the sequence-order effects into the predictor, the 'amphiphilic pseudo amino acid composition' is introduced to represent the statistical sample of a protein. The novel representation contains 20 + 2lambda discrete numbers: the first 20 numbers are the components of the conventional amino acid composition; the next 2lambda numbers are a set of correlation factors that reflect different hydrophobicity and hydrophilicity distribution patterns along a protein chain. Based on such a concept and formulation scheme, a new predictor is developed. It is shown by the self-consistency test, jackknife test and independent dataset tests that the success rates obtained by the new predictor are all significantly higher than those by the previous predictors. The significant enhancement in success rates also implies that the distribution of hydrophobicity and hydrophilicity of the amino acid residues along a protein chain plays a very important role to its structure and function.

Algorithms↗

Prediction of protein subcellular locations by GO-FunD-PseAA predictor.

The localization of a protein in a cell is closely correlated with its biological function. With the explosion of protein sequences entering into DataBanks, it is highly desired to develop an automated method that can fast identify their subcellular location. This will expedite the annotation process, providing timely useful information for both basic research and industrial application. In view of this, a powerful predictor has been developed by hybridizing the gene ontology approach [Nat. Genet. 25 (2000) 25], functional domain composition approach [J. Biol. Chem. 277 (2002) 45765], and the pseudo-amino acid composition approach [Proteins Struct. Funct. Genet. 43 (2001) 246; Erratum: ibid. 44 (2001) 60]. As a showcase, the recently constructed dataset [Bioinformatics 19 (2003) 1656] was used for demonstration. The dataset contains 7589 proteins classified into 12 subcellular locations: chloroplast, cytoplasmic, cytoskeleton, endoplasmic reticulum, extracellular, Golgi apparatus, lysosomal, mitochondrial, nuclear, peroxisomal, plasma membrane, and vacuolar. The overall success rate of prediction obtained by the jackknife cross-validation was 92%. This is so far the highest success rate performed on this dataset by following an objective and rigorous cross-validation procedure.

Algorithms↗

Insights from modelling the 3D structure of the extracellular domain of alpha7 nicotinic acetylcholine receptor.

Based on the crystal structure of acetylcholine-binding protein, the three-dimensional structures of the extracellular domain, or the ligand-binding domains, of the monomer, homodimer, and homopentamer of the alpha7 nicotinic acetylcholine receptor were derived. The interface between two subunits, where the ligand-binding site is located, was investigated. Furthermore, an explicit definition of the ligand-binding pocket was illustrated that might provide useful clues for conducting various mutagenesis studies for finding drugs against schizophrenia and Alzheimer's disease.

Amino Acid Sequence↗