Search PubMed⌕ Search

Biomedical subjects

Sukjoon Yoon

Publications and source records attributed to Sukjoon Yoon.

10 recordsLinked to original sources

Large scale data mining approach for gene-specific standardization of microarray gene expression data.

MOTIVATION: The identification of the change of gene expression in multifactorial diseases, such as breast cancer is a major goal of DNA microarray experiments. Here we present a new data mining strategy to better analyze the marginal difference in gene expression between microarray samples. The idea is based on the notion that the consideration of gene's behavior in a wide variety of experiments can improve the statistical reliability on identifying genes with moderate changes between samples. RESULTS: The availability of a large collection of array samples sharing the same platform in public databases, such as NCBI GEO, enabled us to re-standardize the expression intensity of a gene using its mean and variation in the wide variety of experimental conditions. This approach was evaluated via the re-identification of breast cancer-specific gene expression. It successfully prioritized several genes associated with breast tumor, for which the expression difference between normal and breast cancer cells was marginal and thus would have been difficult to recognize using conventional analysis methods. Maximizing the utility of microarray data in the public database, it provides a valuable tool particularly for the identification of previously unrecognized disease-related genes. AVAILABILITY: A user friendly web-interface (http://compbio.sookmyung.ac.kr/~lage/) was constructed to provide the present large-scale approach for the analysis of GEO microarray data (GS-LAGE server).

Algorithms↗

Tumor necrosis factor-alpha and interleukin-1beta increases CTRP1 expression in adipose tissue.

CTRP1, a member of the CTRP superfamily, consists of an N-terminal signal peptide sequence followed by a variable region, a collagen repeat domain, and a C-terminal globular domain. CTRP1 is expressed at high levels in adipose tissues of LPS-stimulated Sprague-Dawley rats. The LPS-induced increase in CTRP1 gene expression was found to be mediated by TNF-alpha and IL-1beta. Also, a high level of expression of CTRP1 mRNA was observed in adipose tissues of Zucker diabetic fatty (fa/fa) rats, compared to Sprague-Dawley rats in the absence of LPS stimulation. These findings indicate that CTRP1 expression may be associated with a low-grade chronic inflammation status in adipose tissues.

Adipokines↗

Analysis of chameleon sequences by energy decomposition on a pairwise per-residue basis.

The conversion from alpha-helix to beta-strand that has been widely observed in so-called chameleon sequences has received considerable attention since such a structural change may induce many amyloidogenic proteins to self-assemble into fibrils thus causing fatal diseases. Here we report a large scale-analysis of the energetics of secondary structural conversions in a collection of chameleon sequences retrieved from the Protein Data Bank. Major energetic contributions to the secondary structural conversion were analyzed by carrying out energy decomposition on a pairwise per-residue basis, i.e., (i,i), (i,i +/- 1), (i,i +/- 2), (i,i +/- 3), (i,i +/- 4) and > (i,i +/- 4) intra-/inter-residual interactions. While the overall potential energy differences were subtle, individual residue-based interacting energy differences were observed to vary significantly depending on the specific type of secondary structural conversion. The average energy difference between alpha-helix and beta-strand, beta)>, in the chameleon sequences varied significantly in (i,i), (i,i +/- 1) and > (i,i +/- 4) interactions. The major energetic factors in secondary structure conversions were electrostatic interactions and the polar term for solvation energy. In addition, residue-based average energy differences in alpha-helix --> beta-strand conversions were well-correlated to those in alpha-helix --> random coil --> beta-strand conversions (R2 = 0.92). Assuming that three secondary structural elements can transform in either direction, this strong correlation indicates that the present energy decomposition method using database structures of chameleon sequences provides a reliable tool for the characterization of secondary structure fluctuations in amino acid sequences.

Amino Acid Motifs↗

RASSF1A suppresses oncogenic H-Ras-induced c-Jun N-terminal kinase activation.

The constitutive activation of JNK has been implicated in Ras-induced cellular transformation and activated JNK is down-regulated by the tumor suppressor protein, RASSF1A. In this study, we examined whether RASSF1A blocked oncogenic Ras-induced JNK activation. Exogenous expression of H-RasG12V induced JNK phosphorylation and RASSF1A co-transfected with H-RasG12V efficiently suppressed Ras-triggered JNK activation in various cancer cell lines. RASSF1A expression revived the H-RasG12V-induced p27Kip1 down-regulation. JNK siRNA treatment also promoted recovery from the H-RasG12V-induced p27Kip1 down-regulation. These results demonstrate that RASSF1A inhibited H-RasG12V-induced JNK activation and JNK-mediated p27Kip1 down-regulation. From these results, we propose that RASSF1A exerts a tumor-suppressing effect by blocking oncogenic Ras-induced JNK activation.

Cell Line, Tumor↗

Surrogate docking: structure-based virtual screening at high throughput speed.

Structure-based screening using fully flexible docking is still too slow for large molecular libraries. High quality docking of a million molecule library can take days even on a cluster with hundreds of CPUs. This performance issue prohibits the use of fully flexible docking in the design of large combinatorial libraries. We have developed a fast structure-based screening method, which utilizes docking of a limited number of compounds to build a 2D QSAR model used to rapidly score the rest of the database. We compare here a model based on radial basis functions and a Bayesian categorization model. The number of compounds that need to be actually docked depends on the number of docking hits found. In our case studies reasonable quality models are built after docking of the number of molecules containing approximately 50 docking hits. The rest of the library is screened by the QSAR model. Optionally a fraction of the QSAR-prioritized library can be docked in order to find the true docking hits. The quality of the model only depends on the training set size - not on the size of the library to be screened. Therefore, for larger libraries the method yields higher gain in speed no change in performance. Prioritizing a large library with these models provides a significant enrichment with docking hits: it attains the values of approximately 13 and approximately 35 at the beginning of the score-sorted libraries in our two case studies: screening of the NCI collection and a combinatorial libraries on CDK2 kinase structure. With such enrichments, only a fraction of the database must actually be docked to find many of the true hits. The throughput of the method allows its use in screening of large compound collections and in the design of large combinatorial libraries. The strategy proposed has an important effect on efficiency but does not affect retrieval of actives, the latter being determined by the quality of the docking method itself.

Bayes Theorem↗

Rapid assessment of contact-dependent secondary structure propensity: relevance to amyloidogenic sequences.

We have previously demonstrated that calculation of contact-dependent secondary structure propensity (CSSP) is highly sensitive in detecting non-native beta-strand propensities in the core sequences of known amyloidogenic proteins. Here we describe a CSSP method based on an artificial neural network that rapidly and accurately quantifies the influence of tertiary contacts (TCs) on secondary structure propensity in local regions of protein sequences. The present method exhibited 72% accuracy in predicting the alternate secondary structure adopted by chameleon sequences located in highly disparate TC regions. Analysis of 1930 nonhomologous protein domains reveals that the alpha-helix and the beta-strand largely share the same sequence context, and that tertiary context is a major determinant of the native conformation. Conversely, it appears that the propensity of random coils for either the alpha-helix or the beta-strand is largely invariant to tertiary effects. The present CSSP method successfully reproduced the amyloidogenic character observed in local regions of the human islet amyloid polypeptide (hIAPP). Furthermore, CSSP profiles were strongly correlated (r = 0.76) with the observed mutational effects on the aggregation rate of acylphosphatase. Taken together, these results provide compelling evidence in support of the present CSSP approach as a sensitive probe useful for analysis of full-length proteins and for detection of core sequences that may trigger amyloid fibril formation. The combined speed and simplicity of the CSSP method lends itself to proteome-wide analysis of the amyloidogenic nature of common proteins.

Algorithms↗

Computational identification of proteins for selectivity assays.

At the stage of optimization of a chemical series the compounds are normally assayed for binding or inhibition on the target protein as well as on several proteins from a selectivity panel. These proteins are normally identified on the basis of sequence homology to the target protein. Experimental selectivity data are also taken into account if available. Cases when a nonhomologous protein has a significant affinity to the compound series are going to be missed if the selectivity panel is identified by homology. Experimental data is usually either unavailable or limited to a small fraction of proteins that should be considered. We have developed a computational method of identification of selectivity panel proteins. It is based on the evaluation of binding site similarity to the target protein using docking scores of target-selected molecular probes. These probes are obtained by docking a large library of drug-like compounds to the target protein followed by selecting a diverse subset from the best virtual binders. Docking scores of these probes to other proteins measure binding site similarity to the target. Because the method does not require prior knowledge of either affinities or structures of inhibitors for the target, it can be applied to any protein with known 3D structure. Validation of the method includes rediscovery of nonhomologous proteins that bind common ligands: estradiol, tamoxifen, and riboflavin. Given 3D structures, the method can effectively discriminate proteins with similar binding sites from random proteins independent of sequence homology.

Algorithms↗

Improved method for predicting beta-turn using support vector machine.

MOTIVATION: Numerous methods for predicting beta-turns in proteins have been developed based on various computational schemes. Here, we introduce a new method of beta-turn prediction that uses the support vector machine (SVM) algorithm together with predicted secondary structure information. Various parameters from the SVM have been adjusted to achieve optimal prediction performance. RESULTS: The SVM method achieved excellent performance as measured by the Matthews correlation coefficient (MCC = 0.45) using a 7-fold cross validation on a database of 426 non-homologous protein chains. To our best knowledge, this MCC value is the highest achieved so far for predicting beta-turn. The overall prediction accuracy Qtotal was 77.3%, which is the best among the existing prediction methods. Among its unique attractive features, the present SVM method avoids overtraining and compresses information and provides a predicted reliability index.

Algorithms↗

Detecting hidden sequence propensity for amyloid fibril formation.

The preponderance of evidence implicates protein misfolding in many unrelated human diseases. In all cases, normal correctly folded proteins transform from their proper native structure into an abnormal beta-rich structure known as amyloid fibril. Here we introduce a computational algorithm to detect nonnative (hidden) sequence propensity for amyloid fibril formation. Analyzing sequence-structure relationships in terms of tertiary contact (TC), we find that the hidden beta-strand propensity of a query local sequence can be quantitatively estimated from the secondary structure preferences of template sequences of known secondary structure found in regions of high TC. The present method correctly pinpoints the minimal peptide fragment shown experimentally as the likely local mediator of amyloid fibril formation in beta-amyloid peptide, islet amyloid polypeptide (hIAPP), alpha-synuclein, and human acetylcholinesterase (AChE). It also found previously unrecognized beta-strand propensities in the prototypical helical protein myoglobin that has been reported as amyloidogenic. Analysis of 2358 nonhomologous protein domains provides compelling evidence that most proteins contain sequences with significant hidden beta-strand propensity. The present method may find utility in many medically relevant applications, such as the engineering of protein sequences and the discovery of therapeutic agents that specifically target these sequences for the prevention and treatment of amyloid diseases.

Acetylcholinesterase↗

Identification of a minimal subset of receptor conformations for improved multiple conformation docking and two-step scoring.

Docking and scoring are critical issues in virtual drug screening methods. Fast and reliable methods are required for the prediction of binding affinity especially when applied to a large library of compounds. The implementation of receptor flexibility and refinement of scoring functions for this purpose are extremely challenging in terms of computational speed. Here we propose a knowledge-based multiple-conformation docking method that efficiently accommodates receptor flexibility thus permitting reliable virtual screening of large compound libraries. Starting with a small number of active compounds, a preliminary docking operation is conducted on a large ensemble of receptor conformations to select the minimal subset of receptor conformations that provides a strong correlation between the experimental binding affinity (e.g., Ki, IC50) and the docking score. Only this subset is used for subsequent multiple-conformation docking of the entire data set of library (test) compounds. In conjunction with the multiple-conformation docking procedure, a two-step scoring scheme is employed by which the optimal scoring geometries obtained from the multiple-conformation docking are re-scored by a molecular mechanics energy function including desolvation terms. To demonstrate the feasibility of this approach, we applied this integrated approach to the estrogen receptor alpha (ERalpha) system for which published binding affinity data were available for a series of structurally diverse chemicals. The statistical correlation between docking scores and experimental values was significantly improved from those of single-conformation dockings. This approach led to substantial enrichment of the virtual screening conducted on mixtures of active and inactive ERalpha compounds.

Protein Binding↗