Search PubMed⌕ Search

Biomedical subjects

K C Chou

Publications and source records attributed to K C Chou.

At least 55 records · Page 3Linked to original sources

Prediction of the tertiary structure of the complement control protein module.

Complement control protein (CCP) modules, or short consensus repeats (SCR), exist in a wide variety of complement and adhesion proteins, principally the selectins. We have predicted the three-dimensional structure of a CCP module based upon secondary structural information derived by two-dimensional NMR [Barlow et al. (1991), Biochemistry 30, 997-1004]. Accordingly, the CCP is predicted to contain seven beta-strands with extensive hydrogen-bonding interactions, and shows a compact, globular structure. Comparison of this model to the X-ray structure of a kringle domain suggests that the CCP unit is more compact than a kringle structure, and that despite their similarities in size and disulfide bond format, the two are not homologous. Although the function of CCP domains is unknown, it is hoped that the structural model presented herein will facilitate further inquiry into how they contribute to so many systems of biological importance.

Amino Acid Sequence↗

Classification and prediction of beta-turn types.

Although a beta-turn consists of only four amino acids, it assumes many different types in proteins. Is this basically dependent on the tetrapeptide sequence alone or is it due to a variety of interactions with the other part of a protein? To answer this question, a residue-coupled model is proposed that can reflect the sequence-coupling effect for a tetrapeptide in not only a beta-turn or non-beta-turn, but also different types of a beta-turn. The predicted results by the model for 6022 tetrapeptides indicate that the rates of correct prediction for beta-turn types I, I', II, II', VI, and VIII and non-beta-turns are 68.54%, 93.60%, 85.19%, 97.75%, 100%, 88.75%, and 61.02%, respectively. Each of these seven rates is significantly higher than 1/7 = 14.29%, the completely randomized rate, implying that the formation of different beta-turn types or non-beta-turns is considerably correlated with the sequences of a tetrapeptide.

Algorithms↗

Disposition of amphiphilic helices in heteropolar environments.

It is known that alpha helices in globular proteins usually consist of two types of residues, hydrophobic and hydrophilic, with the number of each type being roughly equal. Except for many transmembrane helices, alpha-helices are generally amphiphilic to some degree. This is not entirely surprising because alpha-helices typically reside in heteropolar environments that arise from the polar aqueous solution that surrounds a protein and the apolar "hydrophobic core" located at its center. The packing of alpha-helices in such heteropolar environments is driven by the minimization of free energy brought about by placing hydrophobic sidechains into apolar environments and hydrophilic sidechains into polar environments. The interface between the two environments can be characterized by an interfacial plane, called the demarcation plane, that optimally separates the two classes of residues. The inclination angle omega between the axis of the helix and the demarcation plane provides a measure of the degree of amphiphilicity of an alpha-helix. For highly amphiphilic helices, omega approximately 0. The inclination angle provides a new measure of amphiphilicity that complements the hydrophobic moments of Eisenberg et al. Based on the simple physical model described above, an algorithm is developed for predicting the helix inclination angle. The calculated results show that the inclination angle for most alpha-helices extracted from globular proteins is less than 25 degrees in magnitude. This suggests that helices found in globular proteins tend to be reasonably amphiphilic with half their face dominated by hydrophobic residues and the other half by hydrophilic residues. A new two-dimensional representation that characterizes the disposition of hydrophobic and hydrophilic residues in alpha-helices, called a "wenxiang diagram," is presented. The wenxiang diagram can also be used as an important element to represent a protein molecule.

Algorithms↗

Prediction of beta-turns.

A residue-coupled model is proposed to predict the beta-turns in proteins. The rates of correct prediction for the 455 beta-turn tetrapeptides and 3807 non-beta-turn tetrapeptides in the training database are 94.7 and 81.3%, respectively. The rates of correct prediction for the 110 beta-turn tetrapeptides and 30,229 non-beta-turn tetrapeptides in the testing database are 80.0 and 80.2%, respectively. Compared with the rates of correct prediction based on the residue-independent model reported previously, the quality of prediction is significantly improved by the new model, implying that the residue-coupled effect along a polypeptide chain is important for the formation of reversal turns, such as beta-turns, during the process of protein folding.

Amino Acid Sequence↗

The benzylthio-pyrimidine U-31,355, a potent inhibitor of HIV-1 reverse transcriptase.

U-31,355, or 4-amino-2-(benzylthio)-6-chloropyrimidine is an inhibitor of human immunodeficiency virus type 1 (HIV-1) reverse transcriptase (RT) and possesses anti-HIV activity in HIV-1-infected lymphocytes grown in tissue culture. The compound acts as a specific inhibitor of the RNA-directed DNA polymerase function of HIV-1RT and does not impair the functions of the DNA-catalyzed DNA polymerase or the Rnase H of the enzyme. Kinetic studies were carried out to elucidate the mechanism of RT inhibition by U-31,355. The data were analyzed using Briggs-Haldane kinetics, assuming that the reaction is ordered in that the template:primer binds to the enzyme first, followed by the addition of dNTP, and that the polymerase is a processive enzyme. Based on these assumptions, a velocity equation was derived that allows the calculation of all the essential forward and backward rate constants for the reactions occurring between the enzyme, its substrates, and the inhibitor. The results obtained indicate that U-31,355 acts as a mixed inhibitor with respect to the template:primer and dNTP binding sites associated with the RNA-directed DNA polymerase domain of the enzyme. The inhibitor possessed a significantly higher binding affinity for the enzyme-substrate complexes, than for the free enzyme and consequently did not directly affect the functions of the substrate binding sites. Therefore, U-31,355 appears to impair an event occurring after the formation of the enzyme-substrate complexes, which involves either inhibition of the phosphoester bond formation or translocation of the enzyme relative to its template:primer following the formation of the ester bond. Moreover, the potency of U-31,355 depends on the base composition of the template:primer in that the inhibitor showed a much higher binding affinity for the enzyme-poly (rC):(dG)10 complexes than for the poly (rA):(dT)10 complexes.

Animals↗

Prediction of human immunodeficiency virus protease cleavage sites in proteins.

Knowledge of the polyprotein cleavage sites by HIV protease will refine our understanding of its specificity and the information thus acquired is useful for designing specific and efficient HIV protease inhibitors. The pace in searching for the proper inhibitors of HIV protease will be greatly expedited if one can find an accurate and rapid method for predicting the cleavage sites in proteins by HIV protease. Various prediction models or algorithms have been developed during the past 5 years. This Review is devoted to addressing the following problems: (1) Why is it important to predict the cleavability of a peptide by HIV protease? (2) What progresses have been made in developing the prediction methods, and what merits and weakness does each of these methods carry? The attention is focused on the state-of-the-art, which is featured by a discriminant function algorithm developed very recently as well as an improved database (the program and database are available upon request) established according to new experimental results.

Algorithms↗

Predicting human immunodeficiency virus protease cleavage sites in proteins by a discriminant function method.

Based on the sequence-coupled (Markov chain) model and vector-projection principle, a discriminant function method is proposed to predict sites in protein substrates that should be susceptible to cleavage by the HIV-1 protease. The discriminant function is defined by delta = phi+ - phi-, where phi+ and phi- are the cleavable and noncleavable attributes for a given peptide, and they can be derived from two complementary sets of peptides, S+ and S-, known to be cleavable and noncleavable, respectively, by the enzyme. The rate of correct prediction by the method for the 62 cleavable peptides and 239 noncleavable peptides in the training set are 100 and 96.7%, respectively. Application of the method to the 55 sequences which are outside the training set and known to be cleaved by the HIV-1 protease accurately predicted 100% of the peptides as substrates of the enzyme. The method also predicted all but one of the sites hydrolyzed by the protease in native HIV-1 and HIV-2 reverse transcriptases, where the HIV-1 protease discriminates between nearly identical sequences in a very subtle fashion. Finally, the algorithm predicts correctly all of the HIV-1 protease processing sites in the native gag and gag/pol HIV-1 polyproteins, and all of the cleavage sites identified in denatured protease and reverse transcriptase. The new predictive algorithm provides a novel route toward understanding the specificity of this important therapeutic target.

Algorithms↗

Do "antisense proteins" exist?

A DNA double helix consists of two complementary strands antiparallel with each other. One of them is the sense chain, while the other is an antisense chain which does not directly involve the protein-encoding process. The reason that an antisense chain cannot encode for a protein is generally attributed to the lack of certain preconditions such as a promotor and some necessary sequence segments. Suppose it were provided with all these preconditions, could an antisense chain encode for an "antisense protein"? To answer this question, an analysis has been performed based on the existing database. Nine proteins have been found that have a 100% sequence match with the hypothetical antisense proteins derived from the known Escherichia coli antisense chains.

Base Sequence↗

Knowledge-based model building of the tertiary structures for lectin domains of the selectin family.

A combination of a knowledge-based approach and energy minimization was used to predict the three-dimensional structures of the lectin domains of P-selectin, E-selectin, and L-selectin, respectively. Each of these domains contains 118 amino acids. The starting points for energy minimization were generated based on a framework that consists of a number of separated segments derived from the structure-known carbohydrate-recognition domain of the mannose-binding protein (MBP), which belongs to the same C-type lectin family as the selectin molecules do. The structures thus found for P-, L-, and E-selectin lectin domains share a common feature, i.e., they all contain two alpha-helices, and two antiparallel beta-sheets of which one is formed by two strands (strands 1 and 5) and the other by three (strands 2, 3, and 4). Besides, they all possess two intact disulfide bonds formed by the pair of Cys-19 and Cys-117, and the pair of Cys-90 and Cys-109. The root-mean-square deviations calculated over the set of backbone atoms between P- and L-selectin lectin domains is 3.10 A, that between P- and E-selectin lectin domains 2.48 A, and that between L- and E-selectin lectin domains 3.07 A. A notable feature is the convergence-divergence duality of the 77-107 polypeptide in the three domains; i.e., part of the peptide is folded into a closely similar conformation, and part of it into a highly different one.

Amino Acid Sequence↗

Neural network prediction of the HIV-1 protease cleavage sites.

A back propagation neural network method has been developed to study the pattern of polypeptides that can be cleaved by the HIV-1 protease. This method can incorporate many characteristics of the peptides, such as hydrophobicity, beta-sheet and alpha-helix propensities. Mutations can also be applied to probe the most important factors that influence the cleavage.

Amino Acid Sequence↗

The convergence-divergence duality in lectin domains of selectin family and its implications.

A comparison of the three-dimensional structures of P-, L-, and E-selectin lectin domains reveals that there is a convergence-divergence duality for the 77-107 polypeptide in the three domains; i.e. part of the peptide is folded into a closely similar conformation, and part of it into a highly different one. Since the 77-107 residues are associated with the putative binding sites of the selectin family for ligands, this kind of duality might well reflect the common character of ligands to the selectin family as well as the specificity to each of their respective receptors. The finding may be of use for rationally designing selectin inhibitors with a given specificity and possible antiadhesion drugs.

Amino Acid Sequence↗

Does the folding type of a protein depend on its amino acid composition?

Proteins of known structures are generally classified into one of the following four folding types: alpha, beta, alpha + beta, and alpha/beta proteins. Recent findings [Muskal and Kim (1992) J. Mol. Biol. 225, 713-727] suggested that the folding type of a protein might basically depend on its amino acid composition. If this is true, why is that the predicted results of the protein folding type from amino acid composition always failed to reach the desired accuracy? An examination of the prediction approach indicates that none of the previous algorithms has ever taken into account the coupling effect among different amino acid components. In view of this, a new algorithm has been developed which distinguishes itself from the previous ones by incorporating such a coupling effect. The very high rates, 99.2% and 95.3%, of correct predictions thus obtained for a recently constructed training set of 120 proteins and testing set of 64 proteins, respectively, provide confirmation of the above suggestion.

Algorithms↗

A sequence-coupled vector-projection model for predicting the specificity of GalNAc-transferase.

The specificity of GalNAc-transferase is consistent with the existence of an extended site composed of nine subsites, denoted by R4, R3, R2, R1, R0, R1', R2', R3', and R4', where the acceptor at R0 is either Ser or Thr to which the reducing monosaccharide is being anchored. To predict whether a peptide will react with the enzyme to form a Ser- or Thr-conjugated glycopeptide, a new method has been proposed based on the vector-projection approach as well as the sequence-coupled principle. By incorporating the sequence-coupled effect among the subsites, the interaction mechanism among subsites during glycosylation can be reflected and, by using the vector projection approach, arbitrary assignment for insufficient experimental data can be avoided. The very high ratio of correct predictions versus total predictions for the data in both the training and the testing sets indicates that the method is self-consistent and efficient. It provides a rapid means for predicting O-glycosylation and designing effective inhibitors of GalNAc-transferase, which might be useful for targeting drugs to specific sites in the body and for enzyme replacement therapy for the treatment of genetic disorders.

Amino Acid Sequence↗

A vector projection method for predicting the specificity of GalNAc-transferase.

The specificity of UDP-GalNAc:polypeptide N-acetylgalactosaminytransferase (GalNAc-transferase) is consistent with the existence of an extended site composed of nine subsites, denoted by P4, P3, P2, P1, P0, P1', P2', P3', P4', where the acceptor at P0 is being either Ser or Thr. To predict whether a peptide will react with the enzyme to form a Ser- or Thr-conjugated glycopeptide, a vector projection method is proposed which uses a training set of amino acid sequences surrounding 90 Ser and 106 Thr O-glycosylation sites extracted from the National Biomedical Research Foundation Protein Database. The model postulates independent interactions of the 9 amino acid moieties with their respective binding sites. The high ratio of correct predictions vs. total predictions for the data in both the training and the testing sets indicates that the method is self-consistent and efficient. It provides a rapid means for predicting O-glycosylation and designing effective inhibitors of GalNAc-transferase.

Amino Acid Sequence↗

A novel approach to predicting protein structural classes in a (20-1)-D amino acid composition space.

The development of prediction methods based on statistical theory generally consists of two parts: one is focused on the exploration of new algorithms, and the other on the improvement of a training database. The current study is devoted to improving the prediction of protein structural classes from both of the two aspects. To explore a new algorithm, a method has been developed that makes allowance for taking into account the coupling effect among different amino acid components of a protein by a covariance matrix. To improve the training database, the selection of proteins is carried out so that they have (1) as many non-homologous structures as possible, and (2) a good quality of structure. Thus, 129 representative proteins are selected. They are classified into 30 alpha, 30 beta, 30 alpha + beta, 30 alpha/beta, and 9 zeta (irregular) proteins according to a new criterion that better reflects the feature of the structural classes concerned. The average accuracy of prediction by the current method for the 4 x 30 regular proteins is 99.2%, and that for 64 independent testing proteins not included in the training database is 95.3%. To further validate its efficiency, a jackknife analysis has been performed for the current method as well as the previous ones, and the results are also much in favor of the current method. To complete the mathematical basis, a theorem is presented and proved in Appendix A that is instructive for understanding the novel method at a deeper level.

Amino Acids↗

Monte Carlo simulation studies on the prediction of protein folding types from amino acid composition. II. Correlative effect.

A number of methods to predicting the folding type of a protein based on its amino acid composition have been developed during the past few years. In order to perform an objective and fair comparison of different prediction methods, a Monte Carlo simulation method was proposed to calculate the asymptotic limit of the prediction accuracy [Zhang and Chou (1992), Biophys. J. 63, 1523-1529, referred to as simulation method I]. However, simulation method I was based on an oversimplified assumption, i.e., there are no correlations between the compositions of different amino acids. By taking into account such correlations, a new method, referred to as simulation method II, has been proposed to recalculate the objective accuracy of prediction for the least Euclidean distance method [Nakashima et al. (1986), J. Bochem. 99, 152-162] and the least Minkowski distance method [Chou (1989), Prediction in Protein Structure and the Principles of Protein Conformation, Plenum Press, New York, pp. 549-586], respectively. The results show that the prediction accuracy of the former is still better than that of the latter, as found by simulation method I; however, after incorporating the correlative effect, the objective prediction accuracies become lower for both methods. The reason for this phenomenon is discussed in detail. The simulation method and the idea developed in this paper can be applied to examine any other statistical prediction method, including the computer-simulated neural network method.

Amino Acids↗

An eigenvalue-eigenvector approach to predicting protein folding types.

The accuracy of predicting protein folding types can be significantly enhanced by a recently developed algorithm in which the coupling effect among different amino acid components is taken into account [Chou and Zhang (1994) J. Biol. Chem. 269, 22014-22020]. However, in practical calculations using this powerful algorithm, one may sometimes face ill-conditioned matrices. To overcome such a difficulty, an effective eigenvalue-eigenvector approach is proposed. Furthermore, the new approach has been used to predict a recently constructed set of 76 proteins not included in the training set, and the accuracy of prediction is also much higher than those of other methods.

Algorithms↗