Search PubMed⌕ Search

Biomedical subjects

C T Zhang

Publications and source records attributed to C T Zhang.

At least 19 recordsLinked to original sources

Recognition of protein coding genes in the yeast genome at better than 95% accuracy based on the Z curve.

The Z curve is a three-dimensional space curve constituting the unique representation of a given DNA sequence in the sense that each can be uniquely reconstructed from the other. Based on the Z curve, a new protein coding gene-finding algorithm specific for the yeast genome at better than 95% accuracy has been proposed. Six cross-validation tests were performed to confirm the above accuracy. Using the new algorithm, the number of protein coding genes in the yeast genome is re-estimated. The estimate is based on the assumption that the unknown genes have similar statistical properties to the known genes. It is found that the number of protein coding genes in the 16 yeast chromosomes is </=5645, significantly smaller than the 5800-6000 which is widely accepted, and much larger than the 4800 estimated by another group recently. The mitochondrial genes were not included into the above estimate. A codingness index called the YZ score (YZ OE [0,1]) is proposed to recognize protein coding genes in the yeast genome. Among the ORFs annotated in the MIPS (Munich Information Centre for Protein Sequences) database, those recognized as non-coding by the present algorithm are listed in this paper in detail. The criterion for a coding or non-coding ORF is simply decided by YZ > 0.5 or YZ < 0.5, respectively. The YZ scores for all the ORFs annotated in the MIPS database have been calculated and are available on request by sending e-mail to the corresponding author.

Algorithms↗

S curve, a graphic representation of protein secondary structure sequence and its applications.

A secondary structure sequence is a symbolic string composed of three kinds of letters, indicating the helix, strand, and coil (including turns), respectively. A graphic representation for this abstract symbolic sequence is proposed here, called the S curve. The S curve is the unique representation for a given secondary structure sequence in the sense that the sequence and the S curve can be uniquely determined from the other. Therefore, the S curve contains all the information that the secondary structure sequence contains. Different geometrical properties of the S curve are studied in details, which reflect the basic characteristics of the secondary structure sequences. The S curves are used to display, analyze, and compare the secondary structure sequences. Detailed application examples are presented. One advantage of the S curve methodology is that the main patterns of a given secondary structure sequence can be grasped quickly in a perceivable form. This is particularly useful in the cases in which longer sequences are involved and structures of proteins are unknown.

Amino Acid Sequence↗

A graphic approach to evaluate algorithms of secondary structure prediction.

Algorithms of secondary structure prediction have undergone the developments of nearly 30 years. However, the problem of how to appropriately evaluate and compare algorithms has not yet completely solved. A graphic method to evaluate algorithms of secondary structure prediction has been proposed here. Traditionally, the performance of an algorithm is evaluated by a number, i.e., accuracy of various definitions. Instead of a number, we use a graph to completely evaluate an algorithm, in which the mapping points are distributed in a three-dimensional space. Each point represents the predictive result of the secondary structure of a protein. Because the distribution of mapping points in the 3D space generally contains more information than a number or a set of numbers, it is expected that algorithms may be evaluated and compared by the proposed graphic method more objectively. Based on the point distribution, six evaluation parameters are proposed, which describe the overall performance of the algorithm evaluated. Furthermore, the graphic method is simple and intuitive. As an example of application, two advanced algorithms, i.e., the PHD and NNpredict methods, are evaluated and compared. It is shown that there is still much room for further improvement for both algorithms. It is pointed out that the accuracy for predicting either the alpha-helix or beta-strand in proteins with higher alpha-helix or beta-strand content, respectively, should be greatly improved for both algorithms.

Algorithms↗

A quadratic discriminant analysis of protein structure classification based on the Helix/Strand content.

Based on the 210 non-homologous proteins (domains) classified manually by Michie et al. (J. Mol. Biol. 262, 168-185, 1996), a new structure classification criterion of globular proteins relying on the content of helix/strand has been proposed, using a quadratic discriminant method. Each protein is classified into one of the three classes, i.e. those of alpha class, beta class and alphabeta class (including alpha/beta and alpha+beta classes). According to the new structure classification criterion, of the 210 proteins in the training set, 207 are correctly classified and thus the accuracy is 207/210=98.57%. Multiple cross-validation tests are performed. The jackknife test shows that of the 210 proteins 207 are correctly classified with an accuracy of 98.57%. To test the method further, of 3577 proteins (domains) extracted from SCOP, 91.39% of them are correctly reclassified by the new classification criterion. On average, the accuracy of the new criterion is about 8 percentage points higher than that of the criterion proposed by Nakashima et al. (J. Biochem. 99, 153-162, 1986). Our result shows that the classification based solely on structures is basically consistent with that combining both structural and evolutionary information. Further complete automated classification scheme should consider both structures and evolutionary relationship. The methodology presented provides an appropriate mathematical format to reach this goal.

Animals↗

Prediction of protein (domain) structural classes based on amino-acid index.

A protein (domain) is usually classified into one of the following four structural classes: all-alpha, all-beta, alpha/beta and alpha + beta. In this paper, a new formulation is proposed to predict the structural class of a protein (domain) from its primary sequence. Instead of the amino-acid composition used widely in the previous structural class prediction work, the auto-correlation functions based on the profile of amino-acid index along the primary sequence of the query protein (domain) are used for the structural class prediction. Consequently, the overall predictive accuracy is remarkably improved. For the same training database consisting of 359 proteins (domains) and the same component-coupled algorithm [Chou, K.C. & Maggiora, G.M. (1998) Protein Eng. 11, 523-538], the overall predictive accuracy of the new method for the jackknife test is 5-7% higher than the accuracy based only on the amino-acid composition. The overall predictive accuracy finally obtained for the jackknife test is as high as 90.5%, implying that a significant improvement has been achieved by making full use of the information contained in the primary sequence for the class prediction. This improvement depends on the size of the training database, the auto-correlation functions selected and the amino-acid index used. We have found that the amino-acid index proposed by Oobatake and Ooi, i.e. the average nonbonded energy per residue, leads to the optimal predictive result in the case for the database sets studied in this paper. This study may be considered as an alternative step towards making the structural class prediction more practical.

Algorithms↗

Skewed distribution of protein secondary structure contents over the conformational triangle.

A conformational triangle method is presented to analyze the secondary structure contents of 1028 structurally known proteins in the non-redundant data set of the recent 25% PDB_SELECT. The secondary structure contents of each protein are mapped on to a point in the triangle. It was found that the distribution of the 1028 points is strongly skewed in the triangle and about 42% of the whole area is empty, which is called the forbidden area. The detailed border between the allowable and forbidden areas was calculated. The possible explanation of the skewed distribution is discussed. The distributions of the mapping points for enzymes and non-enzymes in this non-redundant data set are compared. It was found that a necessary rather than a sufficient condition for an enzyme molecule is that its coil content must be >/=0.223. It is hoped that the skewed distribution observed here could be used to test the secondary structure and threading predictions.

Computational Biology↗

A new quantitative criterion to distinguish between alpha/beta and alpha+beta proteins (domains).

According to the statistical analysis, it is shown that the differences of the content of alpha-helix and beta-strand between alpha/beta and alpha+beta proteins are of statistical significance. Based on the secondary structure content and the percentage of parallel or anti-parallel strands, any mixed alphabeta protein can be represented by a point in a three-dimensional prism. The distribution of the mapping points for 79 mixed alphabeta proteins (domains), of which 26 are class alpha/beta and 53 are class alpha+beta, shows that the two kinds of points are situated at distinct regions roughly. A new quantitative criterion based on the Fisher discriminant algorithm is proposed to distinguish between the alpha/beta and alpha+beta proteins (domains). Of the 79 proteins 77 are correctly classified (97.5%). As a stringent cross-validation test, the jackknife test shows that of the 79 proteins 77 are correctly classified. The jackknife test accuracy is still 97.5%. These figures indicate the self-consistence and the extrapolating effectiveness of the new quantitative criterion. Applying the new criterion to reclassify the alpha/beta and alpha+beta proteins (domains) in SCOP is also discussed. It is hoped that the new quantitative criterion will be useful for the development of protein classification databases.

Algorithms↗

A novel approach to distinguish between intron-containing and intronless genes based on the format of Z curves.

A novel method to distinguish between intron-containing and intronless DNA sequences has been proposed, based on different statistic behaviors between them. In this method, DNA sequences are first represented as Z curves. Three exponents alpha, beta and gamma for each given sequence are calculated based on the format of the Z curve for the DNA sequence. A three-dimensional space is spanned by the three exponents. Each DNA sequence may be represented by a point in this space. One hundred intronless and intron-containing genes, respectively, were selected randomly from the GenBank or EMBL database. It is shown that the 200 points are roughly distributed in different regions. The best separating plane to separate the two regions is obtained by using Fisher's discriminant algorithm. For any given sequence to be discriminated, calculate three exponents alpha, beta and gamma, corresponding to a point in the three-dimensional space. If the point is situated at the upper region of the separating plane, the sequence is discriminated as an intronless one; otherwise, the sequence is an intron-containing one. A test of the method for the sequences in an independent test set shows that the discriminant accuracy reaches as high as 89.0%.

Animals↗

Prediction and classification of domain structural classes.

Can the coupling effect among different amino acid components be used to improve the prediction of protein structural classes? The answer is yes according to the study by Chou and Zhang (Crit. Rev. Biochem. Mol. Biol. 30:275-349, 1995), but a completely opposite conclusion was drawn by Eisenhaber et al. when using a different dataset constructed by themselves (Proteins 25:169-179, 1996). To resolve such a perplexing problem, predictions were performed by various approaches for the datasets from an objective database, the SCOP database (Murzin, Brenner, Hubbard, and Chothia. J. Mol. Biol. 247:536-540, 1995). According to SCOP, the classification of structural classes for protein domains is based on the evolutionary relationship and on the principles that govern the 3D structure of proteins, and hence is more natural and reliable. The results from both resubstitution tests and jackknife tests indicate that the overall rates of correct prediction by the algorithm incorporated with the coupling effect among different amino acid components are significantly higher than those by the algorithms without using such an effect. It is elucidated through an analysis that the main reasons for Eisenhaber et al. to have reached an opposite conclusion are the result of (1) misusing the component-coupled algorithm, and (2) using a conceptually incorrect rule to classify protein structural classes. The formulation and analysis presented in this article are conducive to clarify these problems, helping correctly to apply the prediction algorithm and interpret the results.

Algorithms↗

Prediction of the secondary structure contents of globular proteins based on three structural classes.

The prediction of the secondary structural contents (those of alpha-helix and beta-strand) of a globular protein is of great use in the prediction of protein structure. In this paper, a new prediction algorithm has been proposed based on Chou's database [Chou (1995), Proteins 21, 319-344]. The new algorithm is an improved multiple linear regression method, taking into account the nonlinear and coupling terms of the frequencies of different amino acids and the length of the protein. The prediction is also based on the structural classes of proteins, but instead of four classes, only three classes are considered, the alpha class, beta class, and the mixed alpha+beta and alpha/beta class or simply the alphabeta class. Thus the ambiguity that usually occurs between alpha+beta proteins and alpha/beta proteins is eliminated. A resubstitution examination for the algorithm shows that the average absolute errors are 0.040 and 0.035 for the prediction of alpha-helix content and beta-strand content, respectively. An examination of cross-validation, the jackknife analysis, shows that the average absolute errors are 0.051 and 0.045 for the prediction of alpha-helix content and beta-strand content, respectively. Both examinations indicate the self-consistency and the extrapolating effectiveness of the new algorithm. Compared with other methods, ours has the merits of simplicity and convenience for use, as well as high prediction accuracy. By incorporating the prediction of the structural classes, the only input of our method is the amino acid composition and the length of the protein to be predicted.

Algorithms↗

A new criterion to classify globular proteins based on their secondary structure contents.

MOTIVATION: With the enlargement of protein structure databases, it is hoped that a method to classify proteins automatically will be developed. Although the classification criterion proposed by Nakashima et al. ( J. Biochem., 1986, 99, 153-162) was widely used in the literature, it leads to some inconsistencies with the classification databases currently available in the class assignment of protein structures. To improve their work, a new classification criterion is proposed relying on statistical analysis of the secondary structure contents of more than 200 proteins with well-known structural classes. The Fisher linear discriminant algorithm is used to derive the new classification criterion. RESULTS: Three cross-validation tests are performed to evaluate the new criterion. In the jackknife test, of the 210 proteins used to derive the criterion, 206 are correctly classified with an accuracy of 98.10%. Of the 16 proteins of purely intermediate structure (i.e. structures lying near borderlines between two classes) in the first test set, 15 are correctly classified with an accuracy of 93.75%. For the second test set which consists of 200 proteins selected randomly from SCOP, a testing accuracy of 94.00% is obtained. For comparison, the criterion of Nakashima et al. is also used to classify the 210, 16 and 200 proteins, respectively. Consequently, accuracies of 94.76%, 62.50% and 91.50% are obtained, respectively. On average, the accuracy of the new classification criterion is 4% higher than that of Nakashima et al. AVAILABILITY: The program is available on request from the first author. CONTACT: ctzhang@tju.edu.cn

Algorithms↗

A new fourier transform approach for protein coding measure based on the format of the Z curve.

MOTIVATION: At the core of most protein gene-finding algorithms are the coding measures used to make a decision on coding/non-coding. Of the protein coding measures, the Fourier measure is one of the most important. However, due to the limited length of the windows usually used, the accuracy of the measure is not satisfactory. This paper is devoted to improving the accuracy by lengthening the sequence to amplify the periodicity of 3 in the coding regions. RESULTS: A new algorithm is presented called the lengthen-shuffle Fourier transform algorithm. For the same window length, the percentage accuracy of the new algorithm is 6-7% higher than that of the ordinary Fourier transform algorithm. The resulting percentage accuracy (average of specificity and sensitivity) of the new measure is 84.9% for the window length 162 bp. AVAILABILITY: The program is available on request fromC.-T. Zhang. CONTACT: ctzhang@tju.edu.cn

Algorithms↗

Prediction of the helix/strand content of globular proteins based on their primary sequences.

An improved multiple linear regression method has been proposed to predict the content of alpha-helix and beta-strand of a globular protein based on its primary sequence. The amino acid composition and the auto-correlation functions based on the hydrophobicity profile of the primary sequence have been taken into account in the algorithm. The resubstitution test shows that the average absolute errors are 0.077 and 0.073 with the standard deviations 0.059 and 0.057 for the prediction of the content of alpha-helix and beta-strand, respectively. A stringent cross-validation test, i.e., the jackknife test, shows that the average absolute errors are 0.087 and 0.081 with the standard deviations 0.067 and 0.065 for the prediction of the content of alpha-helix and beta-strand, respectively. Both tests indicate the self-consistency and the extrapolating effectiveness of the new algorithm. This greatly improves on previous results (Eisenhaber,F., Imperiale,F., Argos,P. and Frommel,C., 1996, Proteins, 25, 157-168). Compared with other methods currently available, our method has the merits of simplicity and ease-of-use as well as a higher prediction accuracy. The only input of the method is the primary sequence of the query protein to be predicted. The program is available on request via e-mail: ctzhang@tju.edu.cn.

Amino Acid Sequence↗

A symmetrical theory of DNA sequences and its applications.

A unified symmetrical theory of DNA sequences has been established based on the basic symmetry of the DNA bases. It is shown that the symmetry of DNA sequences is inherently related to that of a cube and its inscribed regular tetrahedron. A DNA group is defined as a particular alternating group of order 4, in which the permuted objects are four bases. The symmetry of DNA sequences is described by the DNA group which is isomorphic to the tetrahedral group. The matrix representation for the DNA group has been obtained, and used to establish the relationships between the transforms of bases and the rotations of the tetrahedron. It is found that any DNA sequence can be uniquely described by three independent distributions, i.e., the distributions of the bases of purine/pyrimidine, of amino group/keto group and of strong/weak hydrogen bonds along the sequence. The three distributions are invariant in some sense under the transforms of the DNA group, indicating that the three distributions are inherent for the sequence. The mathematical format of the theory lays a foundation for further development. The applications of the theory to analyse some DNA sequences are presented.

Animals↗

Relations of the numbers of protein sequences, families and folds.

The relations among the numbers of protein sequences, families and folds have been studied theoretically. It is found that the number of families is related to the natural logarithm of the number of sequences. The logarithmic relation should not be changed regardless of what value of the homology threshold is applied in the protein sequence comparison routines. To study the relation between the numbers of families and folds, the degenerate degree of a fold has been introduced. The degenerate degree of a fold is the number of protein families which adopt the same fold. The distribution of the degenerate degrees of folds has been found to be very likely exponential. Based on the distribution, the average degenerate degree d is calculated. The number of folds is simply equal to that of families divided by the average degenerate degree of folds. It is shown that d is an increasing function of time. The current value of d is about 2. It will continue to increase and reach the value of at least 3.3 in some years. By using the above result, the numbers of protein folds for four species have been estimated. In particular, the number of folds for human proteins is estimated to be < or =5200.

Animals↗

Disposition of amphiphilic helices in heteropolar environments.

It is known that alpha helices in globular proteins usually consist of two types of residues, hydrophobic and hydrophilic, with the number of each type being roughly equal. Except for many transmembrane helices, alpha-helices are generally amphiphilic to some degree. This is not entirely surprising because alpha-helices typically reside in heteropolar environments that arise from the polar aqueous solution that surrounds a protein and the apolar "hydrophobic core" located at its center. The packing of alpha-helices in such heteropolar environments is driven by the minimization of free energy brought about by placing hydrophobic sidechains into apolar environments and hydrophilic sidechains into polar environments. The interface between the two environments can be characterized by an interfacial plane, called the demarcation plane, that optimally separates the two classes of residues. The inclination angle omega between the axis of the helix and the demarcation plane provides a measure of the degree of amphiphilicity of an alpha-helix. For highly amphiphilic helices, omega approximately 0. The inclination angle provides a new measure of amphiphilicity that complements the hydrophobic moments of Eisenberg et al. Based on the simple physical model described above, an algorithm is developed for predicting the helix inclination angle. The calculated results show that the inclination angle for most alpha-helices extracted from globular proteins is less than 25 degrees in magnitude. This suggests that helices found in globular proteins tend to be reasonably amphiphilic with half their face dominated by hydrophobic residues and the other half by hydrophilic residues. A new two-dimensional representation that characterizes the disposition of hydrophobic and hydrophilic residues in alpha-helices, called a "wenxiang diagram," is presented. The wenxiang diagram can also be used as an important element to represent a protein molecule.

Algorithms↗

Study on the isentropic equations of nucleotide sequences and their application.

The tetrahedral representation of DNA sequences and its applications have been studied by many authors. In this paper we study the isentropic equations of DNA sequences and their application. First, the DNA sequence entropy is introduced, and the entropy current and divergence are defined. Second, the isentropic equations are deduced and the isentropic curves on the three coordinate planes are respectively drawn by a computer. Third, an analysis is given on the entropy distribution on the coordinate planes. Finally, we make use of the results to discuss the relationship of the fastest increasing directions of the entropy for the cytochrome c genes of seven species.

Animals↗

Do "antisense proteins" exist?

A DNA double helix consists of two complementary strands antiparallel with each other. One of them is the sense chain, while the other is an antisense chain which does not directly involve the protein-encoding process. The reason that an antisense chain cannot encode for a protein is generally attributed to the lack of certain preconditions such as a promotor and some necessary sequence segments. Suppose it were provided with all these preconditions, could an antisense chain encode for an "antisense protein"? To answer this question, an analysis has been performed based on the existing database. Nine proteins have been found that have a 100% sequence match with the hypothetical antisense proteins derived from the known Escherichia coli antisense chains.

Base Sequence↗