Search PubMed⌕ Search

Biomedical subjects

Yaoqi Zhou

Publications and source records attributed to Yaoqi Zhou.

At least 19 recordsLinked to original sources

Achieving 80% ten-fold cross-validated accuracy for secondary structure prediction by large-scale training.

An integrated system of neural networks, called SPINE, is established and optimized for predicting structural properties of proteins. SPINE is applied to three-state secondary-structure and residue-solvent-accessibility (RSA) prediction in this paper. The integrated neural networks are carefully trained with a large dataset of 2640 chains, sequence profiles generated from multiple sequence alignment, representative amino acid properties, a slow learning rate, overfitting protection, and an optimized sliding-widow size. More than 200,000 weights in SPINE are optimized by maximizing the accuracy measured by Q(3) (the percentage of correctly classified residues). SPINE yields a 10-fold cross-validated accuracy of 79.5% (80.0% for chains of length between 50 and 300) in secondary-structure prediction after one-month (CPU time) training on 22 processors. An accuracy of 87.5% is achieved for exposed residues (RSA >95%). The latter approaches the theoretical upper limit of 88-90% accuracy in assigning secondary structures. An accuracy of 73% for three-state solvent-accessibility prediction (25%/75% cutoff) and 79.3% for two-state prediction (25% cutoff) is also obtained.

Algorithms↗

Protein binding site prediction using an empirical scoring function.

Most biological processes are mediated by interactions between proteins and their interacting partners including proteins, nucleic acids and small molecules. This work establishes a method called PINUP for binding site prediction of monomeric proteins. With only two weight parameters to optimize, PINUP produces not only 42.2% coverage of actual interfaces (percentage of correctly predicted interface residues in actual interface residues) but also 44.5% accuracy in predicted interfaces (percentage of correctly predicted interface residues in the predicted interface residues) in a cross validation using a 57-protein dataset. By comparison, the expected accuracy via random prediction (percentage of actual interface residues in surface residues) is only 15%. The binding sites of the 57-protein set are found to be easier to predict than that of an independent test set of 68 proteins. The average coverage and accuracy for this independent test set are 30.5 and 29.4%, respectively. The significant gain of PINUP over expected random prediction is attributed to (i) effective residue-energy score and accessible-surface-area-dependent interface-propensity, (ii) isolation of functional constraints contained in the conservation score from the structural constraints through the combination of residue-energy score (for structural constraints) and conservation score and (iii) a consensus region built on top-ranked initial patches.

Algorithms↗

QBES: predicting real values of solvent accessibility from sequences by efficient, constrained energy optimization.

Solvent accessibility, one of the key properties of amino acid residues in proteins, can be used to assist protein structure prediction. Various approaches such as neural network, support vector machines, probability profiles, information theory, Bayesian theory, logistic function, and multiple linear regression have been developed for solvent accessibility prediction. In this article, a much simpler quadratic programming method based on the buriability parameter set of amino acid residues is developed. The new method, called QBES (Quadratic programming and Buriability Energy function for Solvent accessibility prediction), is reasonably accurate for predicting the real value of solvent accessibility. By using a dataset of 30 proteins to optimize three parameters, the average correlation coefficients between the predicted and actual solvent accessibility are about 0.5 for all four independent test sets ranging from 126 to 513 proteins. The method is efficient. It takes only 20 min for a regular PC to obtain results of 30 proteins with an average length of 263 amino acids. Although the proposed method is less accurate than a few more sophisticated methods based on neural network or support vector machines, this is the first attempt to predict solvent accessibility by energy optimization with constraints. Possible improvements and other applications of the method are discussed.

Protein Conformation↗

Uneven size distribution of mammalian genes in the number of tissues expressed and in the number of co-expressed genes.

Tissue specificity, the traditional predictor of gene function, has recently been used to interpret the selective pressure associated with gene architecture. In this work, we examine gene structures and their relation to the number of tissues expressed and to the number of co-expressed genes, using a recent atlas of microarray-based mouse gene expression in 55 normal tissues. We define tissue specificity and expression-pattern specificity according to the number of tissues expressed and the number of co-expressed genes, respectively. We find that, consistent with previous findings, tissue non-specific (housekeeping) genes are short in all gene regions (coding regions, intron, 5' and 3' untranslated regions). However, in contrast to previous suggestion that tissue-specific genes are long, the genes that are the most tissue-specific (expressed only in one tissue) are also short. We further show that both expression-pattern-specific and non-specific genes are long in coding and non-coding regions. The origins for short tissue-specific genes and long expression-pattern-specific genes are not clear. Genes with highly non-specific expression patterns (i.e. genes with a large number of co-expressed genes) are composed of genes that spread all tissues but are overwhelmingly enriched in the central nervous system (e.g. brain). Thus, the large sizes of these genes are possibly related to the functional complexity and/or accelerated evolutions of the central nervous system.

Animals↗

Fast and accurate method for identifying high-quality protein-interaction modules by clique merging and its application to yeast.

Molecular networks in cells are organized into functional modules, where genes in the same module interact densely with each other and participate in the same biological process. Thus, identification of modules from molecular networks is an important step toward a better understanding of how cells function through the molecular networks. Here, we propose a simple, automatic method, called MC(2), to identify functional modules by enumerating and merging cliques in the protein-interaction data from large-scale experiments. Application of MC(2) to the S. cerevisiae protein-interaction data produces 84 modules, whose sizes range from 4 to 69 genes. The majority of the discovered modules are significantly enriched with a highly specific process term (at least 4 levels below root) and a specific cellular component in Gene Ontology (GO) tree. The average fraction of genes with the most enriched GO term for all modules is 82% for specific biological processes and 78% for specific cellular components. In addition, the predicted modules are enriched with coexpressed proteins. These modules are found to be useful for annotating unknown genes and uncovering novel functions of known genes. MC(2) is efficient, and takes only about 5 min to identify modules from the current yeast gene interaction network with a typical PC (Intel Xeon 2.5 GHz CPU and 512 MB memory). The CPU time of MC(2) is affordable (12 h) even when the number of interactions is increased by a factor of 10. MC(2) and its results are publicly available on http://theory.med.buffalo.edu/MC2.

Algorithms↗

What is a desirable statistical energy function for proteins and how can it be obtained?

Can one obtain a physical energy function for proteins from statistical analysis of protein structures? A direct answer to this question is likely "no." Aless demanding question is whether one can produce a statistical energy function that has the desirable features of a physical-based energy function. Such a desirable energy function would be founded on a physical basis with few or no adjustable parameters, reproduce the known physical characters of amino acid residues, be mostly database independent and transferable, and, more importantly, reasonably accurate in various applications. In this review, we show how such a desirable energy function can be obtained via introducing a simple physical-based reference state called DFIRE (Distance-scaled, Finite, Ideal-gas REference state).

Computer Simulation↗

Design and folding of a multidomain protein.

To test whether the folding process of a large protein can be understood on the basis of the folding behavior of the domains that constitute it, we coupled two well-studied small -helical proteins, the B-domain of protein A (60 amino acids) and Rd-apocytochrome b562 (Rd-apocyt b562, 106 amino acids), by fusing the C-terminal helix of the B-domain of protein A with the N-terminal helix of Rd-apocyt b562 without changing their hydrophobic core residues. The success of the design was confirmed by determining the structure of the engineered protein with multidimensional NMR methods. Kinetic studies showed that the logarithms of the folding/unfolding rate constants of the engineered protein are linearly dependent on concentrations of guanidinium chloride in the measurable range from 1.7 to 4 M. Their slopes (m-values) are close to those of Rd-apocyt b562. In addition, the 1H-15N HSQC spectrum taken at 1.5 M guanidinium chloride reveals that only the Rd-apocyt b562 domain in the designed protein remained folded. These results suggest that the two domains have weak energetic coupling. Interestingly, the redesigned protein folds faster than Rd-apocyt b562, suggesting that the fused helix stabilizes the rate-limiting transition state.

Cytochrome b Group↗

Docking prediction using biological information, ZDOCK sampling technique, and clustering guided by the DFIRE statistical energy function.

We entered the CAPRI experiment during the middle of Round 4 and have submitted predictions for all 6 targets released since then. We used the following procedures for docking prediction: (1) the identification of possible binding region(s) of a target based on known biological information, (2) rigid-body sampling around the binding region(s) by using the docking program ZDOCK, (3) ranking of the sampled complex conformations by employing the DFIRE-based statistical energy function, (4) clustering based on pairwise root-mean-square distance and the DFIRE energy, and (5) manual inspection and relaxation of the side-chain conformations of the top-ranked structures by geometric constraint. Reasonable predictions were made for 4 of the 6 targets. The best fraction of native contacts within the top 10 models are 89.1% for Target 12, 54.3% for Target 13, 29.3% for Target 14, and 94.1% for Target 18. The origin of successes and failures is discussed. .

Algorithms↗

SPEM: improving multiple sequence alignment with sequence profiles and predicted secondary structures.

MOTIVATION: Multiple sequence alignment is an essential part of bioinformatics tools for a genome-scale study of genes and their evolution relations. However, making an accurate alignment between remote homologs is challenging. Here, we develop a method, called SPEM, that aligns multiple sequences using pre-processed sequence profiles and predicted secondary structures for pairwise alignment, consistency-based scoring for refinement of the pairwise alignment and a progressive algorithm for final multiple alignment. RESULTS: The alignment accuracy of SPEM is compared with those of established methods such as ClustalW, T-Coffee, MUSCLE, ProbCons and PRALINE(PSI) in easy (homologs) and hard (remote homologs) benchmarks. Results indicate that the average sum of pairwise alignment scores given by SPEM are 7-15% higher than those of the methods compared in aligning remote homologs (sequence identity <30%). Its accuracy for aligning homologs (sequence identity >30%) is statistically indistinguishable from those of the state-of-the-art techniques such as ProbCons or MUSCLE 6.0. AVAILABILITY: The SPEM server and its executables are available on http://theory.med.buffalo.edu.

Algorithms↗

Web-based toolkits for topology prediction of transmembrane helical proteins, fold recognition, structure and binding scoring, folding-kinetics analysis and comparative analysis of domain combinations.

We have developed the following web servers for protein structural modeling and analysis at http://theory.med.buffalo.edu: THUMBUP, UMDHMM(TMHP) and TUPS, predictors of transmembrane helical protein topology based on a mean-burial-propensity scale of amino acid residues (THUMBUP), hidden Markov model (UMDHMM(TMHP)) and their combinations (TUPS); SPARKS 2.0 and SP3, two profile-profile alignment methods, that match input query sequence(s) to structural templates by integrating sequence profile with knowledge-based structural score (SPARKS 2.0) and structure-derived profile (SP3); DFIRE, a knowledge-based potential for scoring free energy of monomers (DMONOMER), loop conformations (DLOOP), mutant stability (DMUTANT) and binding affinity of protein-protein/peptide/DNA complexes (DCOMPLEX & DDNA); TCD, a program for protein-folding rate and transition-state analysis of small globular proteins; and DOGMA, a web-server that allows comparative analysis of domain combinations between plant and other 55 organisms. These servers provide tools for prediction and/or analysis of proteins on the secondary structure, tertiary structure and interaction levels, respectively.

Amino Acids↗

A knowledge-based energy function for protein-ligand, protein-protein, and protein-DNA complexes.

We developed a knowledge-based statistical energy function for protein-ligand, protein-protein, and protein-DNA complexes by using 19 atom types and a distance-scale finite ideal-gas reference (DFIRE) state. The correlation coefficients between experimentally measured protein-ligand binding affinities and those predicted by the DFIRE energy function are around 0.63 for one training set and two testing sets. The energy function also makes highly accurate predictions of binding affinities of protein-protein and protein-DNA complexes. Correlation coefficients between theoretical and experimental results are 0.73 for 82 protein-protein (peptide) complexes and 0.83 for 45 protein-DNA complexes, despite the fact that the structures of protein-protein (peptide) and protein-DNA complexes were not used in training the energy function. The results of the DFIRE energy function on protein-ligand complexes are compared to the published results of 12 other scoring functions generated from either physical-based, knowledge-based, or empirical methods. They include AutoDock, X-Score, DrugScore, four scoring functions in Cerius 2 (LigScore, PLP, PMF, and LUDI), four scoring functions in SYBYL (F-Score, G-Score, D-Score, and ChemScore), and BLEEP. While the DFIRE energy function is only moderately successful in ranking native or near native conformations, it yields the strongest correlation between theoretical and experimental binding affinities of the testing sets and between rmsd values and energy scores of docking decoys in a benchmark of 100 protein-ligand complexes. The parameters and the program of the all-atom DFIRE energy function are freely available for academic users at http://theory.med.buffalo.edu.

Algorithms↗

Fold recognition by combining sequence profiles derived from evolution and from depth-dependent structural alignment of fragments.

Recognizing structural similarity without significant sequence identity has proved to be a challenging task. Sequence-based and structure-based methods as well as their combinations have been developed. Here, we propose a fold-recognition method that incorporates structural information without the need of sequence-to-structure threading. This is accomplished by generating sequence profiles from protein structural fragments. The structure-derived sequence profiles allow a simple integration with evolution-derived sequence profiles and secondary-structural information for an optimized alignment by efficient dynamic programming. The resulting method (called SP(3)) is found to make a statistically significant improvement in both sensitivity of fold recognition and accuracy of alignment over the method based on evolution-derived sequence profiles alone (SP) and the method based on evolution-derived sequence profile and secondary structure profile (SP(2)). SP(3) was tested in SALIGN benchmark for alignment accuracy and Lindahl, PROSPECTOR 3.0, and LiveBench 8.0 benchmarks for remote-homology detection and model accuracy. SP(3) is found to be the most sensitive and accurate single-method server in all benchmarks tested where other methods are available for comparison (although its results are statistically indistinguishable from the next best in some cases and the comparison is subjected to the limitation of time-dependent sequence and/or structural library used by different methods.). In LiveBench 8.0, its accuracy rivals some of the consensus methods such as ShotGun-INBGU, Pmodeller3, Pcons4, and ROBETTA. SP(3) fold-recognition server is available on http://theory.med.buffalo.edu.

Algorithms↗

SCUD: fast structure clustering of decoys using reference state to remove overall rotation.

We developed a method for fast decoy clustering by using reference root-mean-squared distance (rRMSD) rather than commonly used pairwise RMSD (pRMSD) values. For 41 proteins with 2000 decoys each, the computing efficiency increases nine times without a significant change in the accuracy of near-native selections. Tests on additional protein decoys based on different reference conformations confirmed this result. Further analysis indicates that the pRMSD and rRMSD values are highly correlated (with an average correlation coefficient of 0.82) and the clusters obtained from pRMSD and rRMSD values are highly similar (the representative structures of the top five largest clusters from the two methods are 74% identical). SCUD (Structure ClUstering of Decoys) with an automatic cutoff value is available at http://theory.med.buffalo.edu.

Cluster Analysis↗

SPARKS 2 and SP3 servers in CASP6.

Two single-method servers, SPARKS 2 and SP3, participated in automatic-server predictions in CASP6. The overall results for all as well as detailed performance in comparative modeling targets are presented. It is shown that both SPARKS 2 and SP3 are able to recognize their corresponding best templates for all easy comparative modeling targets. The alignment accuracy, however, is not always the best among all the servers. Possible factors are discussed. SPARKS 2 and SP3 fold recognition servers, as well as their executables, are freely available for all academic users on http://theory.med.buffalo.edu.

Algorithms↗

Protein flexibility prediction by an all-atom mean-field statistical theory.

We extended a mean-field model to proteins with all atomic detail. The all-atom mean-field model was used to calculate the dynamic and thermodynamic properties of a three-helix bundle fragment of Staphylococcal protein A (Protein Data Bank [PDB] ID 1BDD) and alpha-spectrin SH3 domain protein (PDB ID 1SHG). We show that a model with all-atomic detail provides a significantly more accurate prediction of flexibility of residues in proteins than does a coarse-grained residue-level model. The accuracy of flexibility prediction is further confirmed by application of the method to 18 additional proteins with the largest size of 224 residues.

Mathematics↗

Fold helical proteins by energy minimization in dihedral space and a DFIRE-based statistical energy function.

Statistical energy functions are discrete (or stepwise) energy functions that lack van der Waals repulsion. As a result, they are often applied directly to a given structure (native or decoy) without further energy minimization being performed to the structure. However, the full benefit (or hidden defect) of an energy function cannot be revealed without energy minimization. This paper tests a recently developed, all-atom statistical energy function by energy minimization with a fixed secondary helical structure in dihedral space. This is accomplished by combining the statistical energy function based on a distance-scaled finite ideal-gas reference (DFIRE) state with a simple repulsive interaction and an improper torsion energy function. The energy function was used to minimize 2000 random initial structures of 41 small and medium-sized helical proteins in a dihedral space with a fixed helical region. Results indicate that near-native structures for most studied proteins can be obtained by minimization alone. The average minimum root-mean-squared distance (rmsd) from the native structure for all 41 proteins is 4.1 A. The energy function (together with a simple clustering of similar structures) also makes a reasonable selection of near-native structures from minimized structures. The average rmsd value and the average rank for the best structure in the top five is 6.8 A and 2.4, respectively. The accuracy of the structures sampled and the structure selections can be improved significantly with the removal of flexible terminal regions in rmsd calculations and in minimization and with the increase in the number of minimizations. The minimized structures form an excellent decoy set for testing other energy functions because most structures are well-packed with minimum hard-core overlaps with correct hydrophobic/hydrophilic partitioning. They are available online at http://theory.med.buffalo.edu.

Algorithms↗

A physical reference state unifies the structure-derived potential of mean force for protein folding and binding.

Extracting knowledge-based statistical potential from known structures of proteins is proved to be a simple, effective method to obtain an approximate free-energy function. However, the different compositions of amino acid residues at the core, the surface, and the binding interface of proteins prohibited the establishment of a unified statistical potential for folding and binding despite the fact that the physical basis of the interaction (water-mediated interaction between amino acids) is the same. Recently, a physical state of ideal gas, rather than a statistically averaged state, has been used as the reference state for extracting the net interaction energy between amino acid residues of monomeric proteins. Here, we find that this monomer-based potential is more accurate than an existing all-atom knowledge-based potential trained with interfacial structures of dimers in distinguishing native complex structures from docking decoys (100% success rate vs. 52% in 21 dimer/trimer decoy sets). It is also more accurate than a recently developed semiphysical empirical free-energy functional enhanced by an orientation-dependent hydrogen-bonding potential in distinguishing native state from Rosetta docking decoys (94% success rate vs. 74% in 31 antibody-antigen and other complexes based on Z score). In addition, the monomer potential achieved a 93% success rate in distinguishing true dimeric interfaces from artificial crystal interfaces. More importantly, without additional parameters, the potential provides an accurate prediction of binding free energy of protein-peptide and protein-protein complexes (a correlation coefficient of 0.87 and a root-mean-square deviation of 1.76 kcal/mol with 69 experimental data points). This work marks a significant step toward a unified knowledge-based potential that quantitatively captures the common physical principle underlying folding and binding. A Web server for academic users, established for the prediction of binding free energy and the energy evaluation of the protein-protein complexes, may be found at http://theory.med.buffalo.edu.

Antigen-Antibody Complex↗

Single-body residue-level knowledge-based energy score combined with sequence-profile and secondary structure information for fold recognition.

An elaborate knowledge-based energy function is designed for fold recognition. It is a residue-level single-body potential so that highly efficient dynamic programming method can be used for alignment optimization. It contains a backbone torsion term, a buried surface term, and a contact-energy term. The energy score combined with sequence profile and secondary structure information leads to an algorithm called SPARKS (Sequence, secondary structure Profiles and Residue-level Knowledge-based energy Score) for fold recognition. Compared with the popular PSI-BLAST, SPARKS is 21% more accurate in sequence-sequence alignment in ProSup benchmark and 10%, 25%, and 20% more sensitive in detecting the family, superfamily, fold similarities in the Lindahl benchmark, respectively. Moreover, it is one of the best methods for sensitivity (the number of correctly recognized proteins), alignment accuracy (based on the MaxSub score), and specificity (the average number of correctly recognized proteins whose scores are higher than the first false positives) in LiveBench 7 among more than twenty servers of non-consensus methods. The simple algorithm used in SPARKS has the potential for further improvement. This highly efficient method can be used for fold recognition on genomic scales. A web server is established for academic users on http://theory.med.buffalo.edu.

Algorithms↗