Search PubMed⌕ Search

Biomedical subjects

Hongyi Zhou

Publications and source records attributed to Hongyi Zhou.

At least 19 recordsLinked to original sources

What is a desirable statistical energy function for proteins and how can it be obtained?

Can one obtain a physical energy function for proteins from statistical analysis of protein structures? A direct answer to this question is likely "no." Aless demanding question is whether one can produce a statistical energy function that has the desirable features of a physical-based energy function. Such a desirable energy function would be founded on a physical basis with few or no adjustable parameters, reproduce the known physical characters of amino acid residues, be mostly database independent and transferable, and, more importantly, reasonably accurate in various applications. In this review, we show how such a desirable energy function can be obtained via introducing a simple physical-based reference state called DFIRE (Distance-scaled, Finite, Ideal-gas REference state).

Computer Simulation↗

Design and folding of a multidomain protein.

To test whether the folding process of a large protein can be understood on the basis of the folding behavior of the domains that constitute it, we coupled two well-studied small -helical proteins, the B-domain of protein A (60 amino acids) and Rd-apocytochrome b562 (Rd-apocyt b562, 106 amino acids), by fusing the C-terminal helix of the B-domain of protein A with the N-terminal helix of Rd-apocyt b562 without changing their hydrophobic core residues. The success of the design was confirmed by determining the structure of the engineered protein with multidimensional NMR methods. Kinetic studies showed that the logarithms of the folding/unfolding rate constants of the engineered protein are linearly dependent on concentrations of guanidinium chloride in the measurable range from 1.7 to 4 M. Their slopes (m-values) are close to those of Rd-apocyt b562. In addition, the 1H-15N HSQC spectrum taken at 1.5 M guanidinium chloride reveals that only the Rd-apocyt b562 domain in the designed protein remained folded. These results suggest that the two domains have weak energetic coupling. Interestingly, the redesigned protein folds faster than Rd-apocyt b562, suggesting that the fused helix stabilizes the rate-limiting transition state.

Cytochrome b Group↗

SPEM: improving multiple sequence alignment with sequence profiles and predicted secondary structures.

MOTIVATION: Multiple sequence alignment is an essential part of bioinformatics tools for a genome-scale study of genes and their evolution relations. However, making an accurate alignment between remote homologs is challenging. Here, we develop a method, called SPEM, that aligns multiple sequences using pre-processed sequence profiles and predicted secondary structures for pairwise alignment, consistency-based scoring for refinement of the pairwise alignment and a progressive algorithm for final multiple alignment. RESULTS: The alignment accuracy of SPEM is compared with those of established methods such as ClustalW, T-Coffee, MUSCLE, ProbCons and PRALINE(PSI) in easy (homologs) and hard (remote homologs) benchmarks. Results indicate that the average sum of pairwise alignment scores given by SPEM are 7-15% higher than those of the methods compared in aligning remote homologs (sequence identity <30%). Its accuracy for aligning homologs (sequence identity >30%) is statistically indistinguishable from those of the state-of-the-art techniques such as ProbCons or MUSCLE 6.0. AVAILABILITY: The SPEM server and its executables are available on http://theory.med.buffalo.edu.

Algorithms↗

Web-based toolkits for topology prediction of transmembrane helical proteins, fold recognition, structure and binding scoring, folding-kinetics analysis and comparative analysis of domain combinations.

We have developed the following web servers for protein structural modeling and analysis at http://theory.med.buffalo.edu: THUMBUP, UMDHMM(TMHP) and TUPS, predictors of transmembrane helical protein topology based on a mean-burial-propensity scale of amino acid residues (THUMBUP), hidden Markov model (UMDHMM(TMHP)) and their combinations (TUPS); SPARKS 2.0 and SP3, two profile-profile alignment methods, that match input query sequence(s) to structural templates by integrating sequence profile with knowledge-based structural score (SPARKS 2.0) and structure-derived profile (SP3); DFIRE, a knowledge-based potential for scoring free energy of monomers (DMONOMER), loop conformations (DLOOP), mutant stability (DMUTANT) and binding affinity of protein-protein/peptide/DNA complexes (DCOMPLEX & DDNA); TCD, a program for protein-folding rate and transition-state analysis of small globular proteins; and DOGMA, a web-server that allows comparative analysis of domain combinations between plant and other 55 organisms. These servers provide tools for prediction and/or analysis of proteins on the secondary structure, tertiary structure and interaction levels, respectively.

Amino Acids↗

Catalytic amination and dechlorination of para-nitrochlorobenzene (p-NCB) in water over palladium-iron bimetallic catalyst.

Chemical treatment of para-nitrochlorobenzene (p-NCB) by palladium/iron (Pd/Fe) bimetallic particles represents one of the latest innovative technologies for the remediation of contaminated soil and groundwater. The amination and dechlorination reaction is believed to take place predominantly on the surface site of the Pd/Fe catalysts. The p-NCB was first transformed to p-chloroaniline (p-CAN) then quickly reduced to aniline. 100% of p-NCB was removed in 30 min when bimetallic Pd/Fe particles with 0.03% Pd at the Pd/Fe mass concentration of 3g 75 ml(-1) were used. The p-NCB removal efficiency and the subsequent dechlorination rate increased with the increase of bulk loading of palladium and Pd/Fe. As expected, p-NCB removal efficiency increased with temperature as well. In particular, the removal efficiency of p-NCB was measured to be 67%, 79%, 80%, 90% and 100% for reaction temperature 20, 25, 30, 35 and 40 degrees C, respectively. Our results show that no other intermediates were generated besides Cl(-), p-CAN and aniline during the catalytic amination and dechlorination of p-NCB.

Amination↗

Fold recognition by combining sequence profiles derived from evolution and from depth-dependent structural alignment of fragments.

Recognizing structural similarity without significant sequence identity has proved to be a challenging task. Sequence-based and structure-based methods as well as their combinations have been developed. Here, we propose a fold-recognition method that incorporates structural information without the need of sequence-to-structure threading. This is accomplished by generating sequence profiles from protein structural fragments. The structure-derived sequence profiles allow a simple integration with evolution-derived sequence profiles and secondary-structural information for an optimized alignment by efficient dynamic programming. The resulting method (called SP(3)) is found to make a statistically significant improvement in both sensitivity of fold recognition and accuracy of alignment over the method based on evolution-derived sequence profiles alone (SP) and the method based on evolution-derived sequence profile and secondary structure profile (SP(2)). SP(3) was tested in SALIGN benchmark for alignment accuracy and Lindahl, PROSPECTOR 3.0, and LiveBench 8.0 benchmarks for remote-homology detection and model accuracy. SP(3) is found to be the most sensitive and accurate single-method server in all benchmarks tested where other methods are available for comparison (although its results are statistically indistinguishable from the next best in some cases and the comparison is subjected to the limitation of time-dependent sequence and/or structural library used by different methods.). In LiveBench 8.0, its accuracy rivals some of the consensus methods such as ShotGun-INBGU, Pmodeller3, Pcons4, and ROBETTA. SP(3) fold-recognition server is available on http://theory.med.buffalo.edu.

Algorithms↗

SPARKS 2 and SP3 servers in CASP6.

Two single-method servers, SPARKS 2 and SP3, participated in automatic-server predictions in CASP6. The overall results for all as well as detailed performance in comparative modeling targets are presented. It is shown that both SPARKS 2 and SP3 are able to recognize their corresponding best templates for all easy comparative modeling targets. The alignment accuracy, however, is not always the best among all the servers. Possible factors are discussed. SPARKS 2 and SP3 fold recognition servers, as well as their executables, are freely available for all academic users on http://theory.med.buffalo.edu.

Algorithms↗

Catalytic dechlorination kinetics of p-dichlorobenzene over Pd/Fe catalysts.

p-Dichlorobenzene (p-DCB) was dechlorinated using Pd/Fe bimetallic catalytic reductants synthesized by chemical deposition. Batch experiments demonstrated that the Pd/Fe bimetallic particles could effectively dechlorinate p-DCB, p-DCB and its intermediate chlorobenzene were removed completely at a Pd loading of 0.02% (weight ratio of Pd to Fe) and Pd/Fe power to solution ratio about 4g 75 ml-1 in 90 min. Dechlorination was affected by various factors such as the reaction temperature, pH, Pd loading percentage over Fe and the introduction of Pd/Fe catalysts et al. Chlorobenzene represents partially stable dechlorinated intermediates in the generation of benzene and part of p-DCB was dechlorinated to benzene indirectly on the surface of Pd/Fe. The dechlorination of p-DCB took place on the surface of the Pd/Fe bimetallic particles in a pseudo-first-order reaction, the activation energy of the dechlorination reaction was determined to be 80.0 kJ mol-1 at the temperature range of 287-313 K.

Catalysis↗

Structure relationship for catalytic dechlorination rate of dichlorobenzenes in water.

Three isomers of dichlorobenzene (o-, m- and p-DCB) were dechlorinated by Pd/Fe catalyst in aqueous solutions through catalytic reduction. The dechlorination reaction took place on the surface site of the catalyst via a pseudo-first-order kinetics, and resulted in benzene as the final reduction product. The rate constants of the reductive dechlorination for the three dichlorobenzenes (DCBs) in the presence of Pd/Fe as a catalyst were measured experimentally. In all cases, the reaction rate constants were found to increase with the decrease in the Gibbs free energy of the formation of DCBs. The reaction rate constant for o-, m- and p-DCBs in the presence of 0.020% (w/w) Pd/Fe at 25 degrees C was determined to be 0.0213, 0.0223, and 0.0254 min(-1), respectively. While the activation energy of each dechlorination reaction was measured to be 102.5, 96.6 and 80.0 kJ mol(-1) for o-, m- and p-DCBs, respectively. The results demonstrated that p-DCBs were reduced more easily than o- or m-DCBs, and the order of the tendency of the dechlorination was p-DCB>m-DCB>o-DCB. The presented data show the catalytic reduction using Pd/Fe as a catalyst is a fast and easy approach for the dechlorination of DCBs.

Catalysis↗

A physical reference state unifies the structure-derived potential of mean force for protein folding and binding.

Extracting knowledge-based statistical potential from known structures of proteins is proved to be a simple, effective method to obtain an approximate free-energy function. However, the different compositions of amino acid residues at the core, the surface, and the binding interface of proteins prohibited the establishment of a unified statistical potential for folding and binding despite the fact that the physical basis of the interaction (water-mediated interaction between amino acids) is the same. Recently, a physical state of ideal gas, rather than a statistically averaged state, has been used as the reference state for extracting the net interaction energy between amino acid residues of monomeric proteins. Here, we find that this monomer-based potential is more accurate than an existing all-atom knowledge-based potential trained with interfacial structures of dimers in distinguishing native complex structures from docking decoys (100% success rate vs. 52% in 21 dimer/trimer decoy sets). It is also more accurate than a recently developed semiphysical empirical free-energy functional enhanced by an orientation-dependent hydrogen-bonding potential in distinguishing native state from Rosetta docking decoys (94% success rate vs. 74% in 31 antibody-antigen and other complexes based on Z score). In addition, the monomer potential achieved a 93% success rate in distinguishing true dimeric interfaces from artificial crystal interfaces. More importantly, without additional parameters, the potential provides an accurate prediction of binding free energy of protein-peptide and protein-protein complexes (a correlation coefficient of 0.87 and a root-mean-square deviation of 1.76 kcal/mol with 69 experimental data points). This work marks a significant step toward a unified knowledge-based potential that quantitatively captures the common physical principle underlying folding and binding. A Web server for academic users, established for the prediction of binding free energy and the energy evaluation of the protein-protein complexes, may be found at http://theory.med.buffalo.edu.

Antigen-Antibody Complex↗

Single-body residue-level knowledge-based energy score combined with sequence-profile and secondary structure information for fold recognition.

An elaborate knowledge-based energy function is designed for fold recognition. It is a residue-level single-body potential so that highly efficient dynamic programming method can be used for alignment optimization. It contains a backbone torsion term, a buried surface term, and a contact-energy term. The energy score combined with sequence profile and secondary structure information leads to an algorithm called SPARKS (Sequence, secondary structure Profiles and Residue-level Knowledge-based energy Score) for fold recognition. Compared with the popular PSI-BLAST, SPARKS is 21% more accurate in sequence-sequence alignment in ProSup benchmark and 10%, 25%, and 20% more sensitive in detecting the family, superfamily, fold similarities in the Lindahl benchmark, respectively. Moreover, it is one of the best methods for sensitivity (the number of correctly recognized proteins), alignment accuracy (based on the MaxSub score), and specificity (the average number of correctly recognized proteins whose scores are higher than the first false positives) in LiveBench 7 among more than twenty servers of non-consensus methods. The simple algorithm used in SPARKS has the potential for further improvement. This highly efficient method can be used for fold recognition on genomic scales. A web server is established for academic users on http://theory.med.buffalo.edu.

Algorithms↗

Critical nucleation size in the folding of small apparently two-state proteins.

For apparently two-state proteins, we found that the size (number of folded residues) of a transition state is mostly encoded by the topology, defined by total contact distance (TCD) of the native state, and correlates with its folding rate. This is demonstrated by using a simple procedure to reduce the native structures of the 41 two-state proteins with native TCD as a constraint, and is further supported by analyzing the results of eight proteins from protein engineering studies. These results support the hypothesis that the major rate-limiting process in the folding of small apparently two-state proteins is the search for a critical number of residues with the topology close to that of the native state.

Kinetics↗

Quantifying the effect of burial of amino acid residues on protein stability.

The average contribution of individual residue to folding stability and its dependence on buried accessible surface area (ASA) are obtained by two different approaches. One is based on experimental mutation data, and the other uses a new knowledge-based atom-atom potential of mean force. We show that the contribution of a residue has a significant correlation with buried ASA and the regression slopes of 20 amino acid residues (called the buriability) are all positive (pro-burial). The buriability parameter provides a quantitative measure of the driving force for the burial of a residue. The large buriability gap observed between hydrophobic and hydrophilic residues is responsible for the burial of hydrophobic residues in soluble proteins. Possible factors that contribute to the buriability gap are discussed.

Amino Acids↗

An accurate, residue-level, pair potential of mean force for folding and binding based on the distance-scaled, ideal-gas reference state.

Structure prediction on a genomic scale requires a simplified energy function that can efficiently sample the conformational space of polypeptide chains. A good energy function at minimum should discriminate native structures against decoys. Here, we show that a recently developed, residue-specific, all-atom knowledge-based potential (167 atomic types) based on distance-scaled, finite ideal-gas reference state (DFIRE-all-atom) can be substantially simplified to 20 residue types located at side-chain center of mass (DFIRE-SCM) without a significant change in its capability of structure discrimination. Using 96 standard multiple decoy sets, we show that there is only a small reduction (from 80% to 78%) in success rate of ranking native structures as the top 1. The success rate is higher than two previously developed, all-atom distance-dependent statistical pair potentials. Applied to structure selections of 21 docking decoys without modification, the DFIRE-SCM potential is 29% more successful in recognizing native complex structures than an all-atom statistical potential trained by a database of dimeric interfaces. The potential also achieves 92% accuracy in distinguishing true dimeric interfaces from artificial crystal interfaces. In addition, the DFIRE potential with the C(alpha) positions as the interaction centers recognizes 123 native structures out of a comprehensive 125-protein TOUCHSTONE decoy set in which each protein has 24,000 decoys with only C(alpha) positions. Furthermore, the performance by DFIRE-SCM on newly established 25 monomeric and 31 docking Rosetta-decoy sets is comparable to (or better than in the case of monomeric decoy sets) that of a recently developed, all-atom Rosetta energy function enhanced with an orientation-dependent hydrogen bonding potential.

Amino Acids↗

The dependence of all-atom statistical potentials on structural training database.

An accurate statistical energy function that is suitable for the prediction of protein structures of all classes should be independent of the structural database used for energy extraction. Here, two high-resolution, low-sequence-identity structural databases of 333 alpha-proteins and 271 beta-proteins were built for examining the database dependence of three all-atom statistical energy functions. They are RAPDF (residue-specific all-atom conditional probability discriminatory function), atomic KBP (atomic knowledge-based potential), and DFIRE (statistical potential based on distance-scaled finite ideal-gas reference state). These energy functions differ in the reference states used for energy derivation. The energy functions extracted from the different structural databases are used to select native structures from multiple decoys of 64 alpha-proteins and 28 beta-proteins. The performance in native structure selections indicates that the DFIRE-based energy function is mostly independent of the structural database whereas RAPDF and KBP have a significant dependence. The construction of two additional structural databases of alpha/beta and alpha + beta-proteins further confirmed the weak dependence of DFIRE on the structural databases of various structural classes. The possible source for the difference between the three all-atom statistical energy functions is that the physical reference state of ideal gas used in the DFIRE-based energy function is least dependent on the structural database.

Algorithms↗

Cooperativity in Scapharca dimeric hemoglobin: simulation of binding intermediates and elucidation of the role of interfacial water.

Cooperative binding of ligands to proteins can serve to increase their efficiency and to regulate their activity. Thus, understanding of the mechanism of cooperativity is one of the central concerns of molecular biology. For the tetrameric human hemoglobin (HbA), the cooperative mechanism involves a reasonably well understood combination of tertiary and quaternary changes that occur during the binding process. The dimeric hemoglobin of Scapharca (HbI), which is composed of subunits with the same fold as in HbA, is also highly cooperative but the structural changes on ligand binding are small. A re-orientation of Phe97 in the binding pocket and changes in the number of interfacial water molecules have been implicated in the cooperative mechanism. To explore the role of these factors, we have investigated models of partially liganded intermediate states of HbI with molecular dynamics simulation methods. Since, unlike HbA, no structures for intermediates are available, they were constructed by combining subunits from the unliganded and liganded dimers. Two structurally distinct intermediates were examined, and it was shown that the transition between the two intermediates is directly coupled to the number of interfacial water molecules. Further, it was found that there is a well-defined water channel that connects the interface between the subunits to bulk water. The bottleneck (gate) of the channel, which can be open or closed, is made of hydrophilic residues. The implication of the present results for the cooperative mechanism of HbI is discussed.

Allosteric Regulation↗

Predicting the topology of transmembrane helical proteins using mean burial propensity and a hidden-Markov-model-based method.

Helices in membrane spanning regions are more tightly packed than the helices in soluble proteins. Thus, we introduce a method that uses a simple scale of burial propensity and a new algorithm to predict transmembrane helical (TMH) segments and a positive-inside rule to predict amino-terminal orientation. The method (the topology predictor of transmembrane helical proteins using mean burial propensity [THUMBUP]) correctly predicted the topology of 55 of 73 proteins (or 75%) with known three-dimensional structures (the 3D helix database). This level of accuracy can be reached by MEMSAT 1.8 (a 200-parameter model-recognition method) and a new HMM-based method (a 111-parameter hidden Markov model, UMDHMM(TMHP)) if they were retrained with the 73-protein database. Thus, a method based on a physiochemical property can provide topology prediction as accurate as those methods based on more complicated statistical models and learning algorithms for the proteins with accurately known structures. Commonly used HMM-based methods and MEMSAT 1.8 were trained with a combination of the partial 3D helix database and a 1D helix database of TMH proteins in which topology information were obtained by gene fusion and other experimental techniques. These methods provide a significantly poorer prediction for the topology of TMH proteins in the 3D helix database. This suggests that the 1D helix database, because of its inaccuracy, should be avoided as either a training or testing database. A Web server of THUMBUP and UMDHMM(TMHP) is established for academic users at http://www.smbs.buffalo.edu/phys_bio/service.htm. The 3D helix database is also available from the same Web site.

Algorithms↗

Stability scale and atomic solvation parameters extracted from 1023 mutation experiments.

The stability scale of 20 amino acid residues is derived from a database of 1023 mutation experiments on 35 proteins. The resulting scale of hydrophobic residues has an excellent correlation with the octanol-to-water transfer free energy corrected with an additional Flory-Huggins molar-volume term (correlation coefficient r = 0.95, slope = 1.05, and a near zero intercept). Thus, hydrophobic contribution to folding stability is characterized remarkably well by transfer experiments. However, no corresponding correlation is found for hydrophilic residues. Both the hydrophilic portion and the entire scale, however, correlate strongly with average burial accessible surface (r = 0.76 and 0.97, respectively). Such a strong correlation leads to a near uniform value of the atomic solvation parameters for atoms C, S, O/N, O(-0.5), and N(+0.5,1). All are in the range of 12-28 cal x mol(-1) A(-2), close to the original estimate of hydrophobic contribution of 25-30 cal x mol(-1) A(-2) to folding stability. Without any adjustable parameters, the new stability scale and new atomic solvation parameters yielded an accurate prediction of protein-protein binding free energy for a separate database of 21 protein-protein complexes (r = 0.80 and slope = 1.06, and r = 0.83 and slope = 0.93, respectively).

Amino Acids↗