Search PubMed⌕ Search

Biomedical subjects

Zhirong Sun

Publications and source records attributed to Zhirong Sun.

At least 19 recordsLinked to original sources

NBA-Palm: prediction of palmitoylation site implemented in Naïve Bayes algorithm.

BACKGROUND: Protein palmitoylation, an essential and reversible post-translational modification (PTM), has been implicated in cellular dynamics and plasticity. Although numerous experimental studies have been performed to explore the molecular mechanisms underlying palmitoylation processes, the intrinsic feature of substrate specificity has remained elusive. Thus, computational approaches for palmitoylation prediction are much desirable for further experimental design. RESULTS: In this work, we present NBA-Palm, a novel computational method based on Naïve Bayes algorithm for prediction of palmitoylation site. The training data is curated from scientific literature (PubMed) and includes 245 palmitoylated sites from 105 distinct proteins after redundancy elimination. The proper window length for a potential palmitoylated peptide is optimized as six. To evaluate the prediction performance of NBA-Palm, 3-fold cross-validation, 8-fold cross-validation and Jack-Knife validation have been carried out. Prediction accuracies reach 85.79% for 3-fold cross-validation, 86.72% for 8-fold cross-validation and 86.74% for Jack-Knife validation. Two more algorithms, RBF network and support vector machine (SVM), also have been employed and compared with NBA-Palm. CONCLUSION: Taken together, our analyses demonstrate that NBA-Palm is a useful computational program that provides insights for further experimentation. The accuracy of NBA-Palm is comparable with our previously described tool CSS-Palm. The NBA-Palm is freely accessible from: http://www.bioinfo.tsinghua.edu.cn/NBA-Palm.

Acyltransferases↗

Effects of potassium alkalis and sodium alkalis on the dechlorination of o-chlorophenol in supercritical water.

Effects of potassium alkalis and sodium alkalis on the dechlorination of o-chlorophenol (o-CP) in supercritical water (SCW) were studied in this paper under the conditions of 450 degrees C and 25 MPa. Experimental results indicated that the dechlorination of o-CP can be accelerated significantly by all alkalis investigated. The dechlorination of o-CP proceeded mainly via two pathways: hydrodechlorination and hydrolysis. Both of the two pathways can be promoted by alkalis, and the dechlorination of o-CP can be accelerated by both the cations and hydroxide ion dissociated from alkalis. The overall dechlorination of o-CP can be accelerated by cations via promoting the hydrodechlorination pathway, while, hydroxide ion via promoting the hydrolysis pathway. In addition, the hydrodechlorination can be accelerated faster by sodium alkalis than that by potassium ones, while, the hydrolysis can be promoted faster by potassium alkalis. This difference may be caused by the different charge density between potassium ion and sodium ion, and the different solubility and dissociation constant between potassium alkalis and sodium alkalis in SCW. Dechlorination of o-CP with addition of alkalis prior to supercritical water oxidation (SCWO) process not only can avoid the reactor corrosion caused by the generated hydrochloric acid in direct SCWO of o-CP, but also can reduce the formation of toxic chlorinated byproducts compared with direct SCWO process or SCWO of o-CP with addition of alkali.

Alkalies↗

Funneled landscape leads to robustness of cell networks: yeast cell cycle.

We uncovered the underlying energy landscape for a cellular network. We discovered that the energy landscape of the yeast cell-cycle network is funneled towards the global minimum (G0/G1 phase) from the experimentally measured or inferred inherent chemical reaction rates. The funneled landscape is quite robust against random perturbations. This naturally explains robustness from a physical point of view. The ratio of slope versus roughness of the landscape becomes a quantitative measure of robustness of the network. The funneled landscape can be seen as a possible realization of the Darwinian principle of natural selection at the cellular network level. It provides an optimal criterion for network connections and design. Our approach is general and can be applied to other cellular networks.

Cell Cycle↗

Preferential duplication in the sparse part of yeast protein interaction network.

Gene duplication is an important mechanism driving the evolution of biomolecular network. Thus, it is expected that there should be a strong relationship between a gene's duplicability and the interactions of its protein product with other proteins in the network. We studied this question in the context of the protein interaction network (PIN) of Saccharomyces cerevisiae. We found that duplicates have, on average, significantly lower clustering coefficient (CC) than singletons, and the proportion of duplicates (PD) decreases steadily with CC. Furthermore, using functional annotation data, we observed a strong negative correlation between PD and the mean CC for functional categories. By partitioning the network into modules and assigning each protein a modularity measure Q(n), we found that CC of a protein is a reflection of its modularity. Moreover, the core components of complexes identified in a recent high-throughput experiment, characterized by high CC, have lower PD than that of the attachments. Subsequently, 2 types of hub were identified by their degree, CC and Q(n). Although PD of intramodular hubs is much less than the network average, PD of intermodular hubs is comparable to, or even higher than, the network average. Our results suggest that high CC, and thus high modularity, pose strong evolutionary constraints on gene duplicability, and gene duplication prefers to happen in the sparse part of PINs.

Cluster Analysis↗

MeMo: a web tool for prediction of protein methylation modifications.

Protein methylation is an important and reversible post-translational modification of proteins (PTMs), which governs cellular dynamics and plasticity. Experimental identification of the methylation site is labor-intensive and often limited by the availability of reagents, such as methyl-specific antibodies and optimization of enzymatic reaction. Computational analysis may facilitate the identification of potential methylation sites with ease and provide insight for further experimentation. Here we present a novel protein methylation prediction web server named MeMo, protein methylation modification prediction, implemented in Support Vector Machines (SVMs). Our present analysis is primarily focused on methylation on lysine and arginine, two major protein methylation sites. However, our computational platform can be easily extended into the analyses of other amino acids. The accuracies for prediction of protein methylation on lysine and arginine have reached 67.1 and 86.7%, respectively. Thus, the MeMo system is a novel tool for predicting protein methylation and may prove useful in the study of protein methylation function and dynamics. The MeMo web server is available at: http://www.bioinfo.tsinghua.edu.cn/~tigerchen/memo.html.

Arginine↗

Funneled landscape leads to robustness of cellular networks: MAPK signal transduction.

We uncover the underlying potential energy landscape for a cellular network. We find that the potential energy landscape of the mitogen-activated protein-kinase signal transduction network is funneled toward the global minimum. The funneled landscape is quite robust against random perturbations. This naturally explains robustness from a physical point of view. The ratio of slope versus roughness of the landscape becomes a quantitative measure of robustness of the network. Funneled landscape is a realization of the Darwinian principle of natural selection at the cellular network level. It provides an optimal criterion for network connections and design. Our approach is general and can be applied to other cellular networks.

Animals↗

Inferring functional linkages between proteins from evolutionary scenarios.

Identifying potential protein interactions is of great importance in understanding the topologies of cellular networks, which is much needed and valued in current systematic biological studies. The development of our computational methods to predict protein-protein interactions have been spurred on by the massive sequencing efforts of the genomic revolution. Among these methods is phylogenetic profiling, which assumes that proteins under similar evolutionary pressures with similar phylogenetic profiles might be functionally related. Here, we introduce a method for inferring functional linkages between proteins from their evolutionary scenarios. The term evolutionary scenario refers to a series of events that occurred in speciation over time, which can be reconstructed given a phylogenetic profile and a species tree. Common evolutionary pressures on two proteins can then be inferred by comparing their evolutionary scenarios, which is a direct indication of their functional linkage. This scenario method has proven to have better performance compared with the classical phylogenetic profile method, when applied to the same test set. In addition, predicted results of the two methods are found to be fairly different, suggesting the possibility of merging them in order to achieve a better performance. We analyzed the influence of the topology of the phylogenetic tree on the performance of this method, and found it to be robust to perturbations in the topology of the tree. However, if a completely random tree is incorporated, performance will decline significantly. The evolutionary scenario method was used for inferring functional linkages in 67 species, and 40,006 linkages were predicted. We examine our prediction for budding yeast and find that almost all predicted linkages are supported by further evidence.

Computational Biology↗

CTKPred: an SVM-based method for the prediction and classification of the cytokine superfamily.

Cell proliferation, differentiation and death are controlled by a multitude of cell-cell signals and loss of this control has devastating consequences. Prominent among these regulatory signals is the cytokine superfamily, which has crucial functions in the development, differentiation and regulation of immune cells. In this study, a support vector machine (SVM)-based method was developed for predicting families and subfamilies of cytokines using dipeptide composition. The taxonomy of the cytokine superfamily with which our method complies was described in the Cytokine Family cDNA Database (dbCFC) and the dataset used in this study for training and testing was obtained from the dbCFC and Structural Classification of Proteins (SCOP). The method classified cytokines and non-cytokines with an accuracy of 92.5% by 7-fold cross-validation. The method is further able to predict seven major classes of cytokine with an overall accuracy of 94.7%. A server for recognition and classification of cytokines based on multi-class SVMs has been set up at http://bioinfo.tsinghua.edu.cn/~huangni/CTKPred/.

Algorithms↗

A novel statistical ligand-binding site predictor: application to ATP-binding sites.

Structural genomics initiatives are leading to rapid growth in newly determined protein 3D structures, the functional characterization of which may still be inadequate. As an attempt to provide insights into the possible roles of the emerging proteins whose structures are available and/or to complement biochemical research, a variety of computational methods have been developed for the screening and prediction of ligand-binding sites in raw structural data, including statistical pattern classification techniques. In this paper, we report a novel statistical descriptor (the Oriented Shell Model) for protein ligand-binding sites, which utilizes the distance and angular position distribution of various structural and physicochemical features present in immediate proximity to the center of a binding site. Using the support vector machine (SVM) as the classifier, our model identified 69% of the ATP-binding sites in whole-protein scanning tests and in eukaryotic proteins the accuracy is particularly high. We propose that this feature extraction and machine learning procedure can screen out ligand-binding-capable protein candidates and can yield valuable biochemical information for individual proteins.

Adenosine Triphosphate↗

In silico identification of the key components and steps in IFN-gamma induced JAK-STAT signaling pathway.

Systems biology efforts are increasingly adopting quantitative, mechanistic modeling to study cellular signal transduction pathways and other networks. However, it is uncertain whether the particular set of kinetic parameter values of the model closely approximates the corresponding biological system. We propose that the parameters be assigned statistical distributions that reflect the degree of uncertainty for a comprehensive simulation analysis. From this analysis, we globally identify the key components and steps in signal transduction networks at a systems level. We investigated a recent mathematical model of interferon gamma induced Janus kinase-signal transducers and activators of transcription (JAK-STAT) signaling pathway by applying multi-parametric sensitivity analysis that is based on simultaneous variation of the parameter values. We find that suppressor of cytokine signaling-1, nuclear phosphatases, cytoplasmic STAT1, and the corresponding reaction steps are sensitive perturbation points of this pathway.

Computational Biology↗

A novel method for protein secondary structure prediction using dual-layer SVM and profiles.

A high-performance method was developed for protein secondary structure prediction based on the dual-layer support vector machine (SVM) and position-specific scoring matrices (PSSMs). SVM is a new machine learning technology that has been successfully applied in solving problems in the field of bioinformatics. The SVM's performance is usually better than that of traditional machine learning approaches. The performance was further improved by combining PSSM profiles with the SVM analysis. The PSSMs were generated from PSI-BLAST profiles, which contain important evolution information. The final prediction results were generated from the second SVM layer output. On the CB513 data set, the three-state overall per-residue accuracy, Q3, reached 75.2%, while segment overlap (SOV) accuracy increased to 80.0%. On the CB396 data set, the Q3 of our method reached 74.0% and the SOV reached 78.1%. A web server utilizing the method has been constructed and is available at http://www.bioinfo.tsinghua.edu.cn/pmsvm.

Computational Biology↗

HMMGEP: clustering gene expression data using hidden Markov models.

SUMMARY: The package HMMGEP performs cluster analysis on gene expression data using hidden Markov models. AVAILABILITY: HMMGEP, including the source code, documentation and sample data files, is available at http://www.bioinfo.tsinghua.edu.cn:8080/~rich/hmmgep_download/index.html.

Algorithms↗

DBSubLoc: database of protein subcellular localization.

We have built a protein subcellular localization annotation database, the DBSubLoc database, which is available at http://www.bioinfo.tsinghua. edu.cn/dbsubloc.html. Annotations were taken from primary protein databases, model organism genome projects and literature texts, and then were analyzed to dig out the subcellular localization features of the proteins. The proteins are also classified into different categories. Based on sequence alignment, non-redundant subsets of the database have been built, which may provide useful information for subcellular localization prediction. The database now contains >60,000 protein sequences including approximately 30,000 protein sequences in the non-redundant data sets. Online download, search and Blast tools are also available.

Animals↗

Structural and functional characterization of the human CCR5 receptor in complex with HIV gp120 envelope glycoprotein and CD4 receptor by molecular modeling studies.

The entry of human immunodeficiency virus (HIV) into cells depends on a sequential interaction of the gp120 envelope glycoprotein with the cellular receptors CD4 and members of the chemokine receptor family. The CC chemokine receptor CCR5 is such a receptor for several chemokines and a major coreceptor for the entry of R5 HIV type-1 (HIV-1) into cells. Although many studies focus on the interaction of CCR5 with HIV-1, the corresponding interaction sites in CCR5 and gp120 have not been matched. Here we used an approach combining protein structure modeling, docking and molecular dynamics simulation to build a series of structural models of the CCR5 in complexes with gp120 and CD4. Interactions such as hydrogen bonds, salt bridges and van der Waals contacts between CCR5 and gp120 were investigated. Three snapshots of CCR5-gp120-CD4 models revealed that the initial interactions of CCR5 with gp120 are involved in the negatively charged N-terminus (Nt) region of CCR5 and positively charged bridging sheet region of gp120. Further interactions occurred between extracellular loop2 (ECL2) of CCR5 and the base of V3 loop regions of gp120. These interactions may induce the conformational changes in gp120 and lead to the final entry of HIV into the cell. These results not only strongly support the two-step gp120-CCR5 binding mechanism, but also rationalize extensive biological data about the role of CCR5 in HIV-1 gp120 binding and entry, and may guide efforts to design novel inhibitors.

CD4 Antigens↗

Mining gene expression data using a novel approach based on hidden Markov models.

In this work we have developed a new framework for microarray gene expression data analysis. This framework is based on hidden Markov models. We have benchmarked the performance of this probability model-based clustering algorithm on several gene expression datasets for which external evaluation criteria were available. The results showed that this approach could produce clusters of quality comparable to two prevalent clustering algorithms, but with the major advantage of determining the number of clusters. We have also applied this algorithm to analyze published data of yeast cell cycle gene expression and found it able to successfully dig out biologically meaningful gene groups. In addition, this algorithm can also find correlation between different functional groups and distinguish between function genes and regulation genes, which is helpful to construct a network describing particular biological associations. Currently, this method is limited to time series data. Supplementary materials are available at http://www.bioinfo.tsinghua.edu.cn/~rich/hmmgep_supp/.

Algorithms↗

An approach to identify over-represented cis-elements in related sequences.

Computational identification of transcription factor binding sites is an important research area of computational biology. Positional weight matrix (PWM) is a model to describe the sequence pattern of binding sites. Usually, transcription factor binding sites prediction methods based on PWMs require user-defined thresholds. The arbitrary threshold and also the relatively low specificity of the algorithm prevent the result of such an analysis from being properly interpreted. In this study, a method was developed to identify over-represented cis-elements with PWM-based similarity scores. Three sets of closely related promoters were analyzed, and only over- represented motifs with high PWM similarity scores were reported. The thresholds to evaluate the similarity scores to the PWMs of putative transcription factors binding sites can also be automatically determined during the analysis, which can also be used in further research with the same PWMs. The online program is available on the website: http://www.bioinfo.tsinghua.edu.cn/- zhengjsh/OTFBS/.

Actins↗

Proteins with class alpha/beta fold have high-level participation in fusion events.

Now that complete genome sequences are available for a variety of organisms, the elucidation of potential gene products function is a central goal in the post-genome era. Domain fusion analysis has been proposed recently to infer the functional association of the component proteins. Here, we took a new approach to the analysis of the structural features of the proteins involved in fusion events. An exhaustive survey of fusion events within 30 completely sequenced genomes and subsequent structure annotations to the component proteins at a SCOP superfamily level with hidden Markov models was carried out. A domain fusion map was then constructed. The results revealed that proteins with the class alpha/beta fold are frequently involved in fusion events, around 86% of the total 676 assigned single-domain fusion pairs including at least one component protein belonging to the alpha/beta fold class. Moreover, the domain fusion map in our work may offer an attractive framework for designing chimeric enzymes following Nature's lead, and may give useful hints for exploring the evolutionary history of proteins. (c) 2002 Elsevier Science Ltd.

Archaeal Proteins↗