Search PubMed⌕ Search

Biomedical subjects

Hyunju Lee

Publications and source records attributed to Hyunju Lee.

4 recordsLinked to original sources

CGI: a new approach for prioritizing genes by combining gene expression and protein-protein interaction data.

MOTIVATION: Identifying candidate genes associated with a given phenotype or trait is an important problem in biological and biomedical studies. Prioritizing genes based on the accumulated information from several data sources is of fundamental importance. Several integrative methods have been developed when a set of candidate genes for the phenotype is available. However, how to prioritize genes for phenotypes when no candidates are available is still a challenging problem. RESULTS: We develop a new method for prioritizing genes associated with a phenotype by Combining Gene expression and protein Interaction data (CGI). The method is applied to yeast gene expression data sets in combination with protein interaction data sets of varying reliability. We found that our method outperforms the intuitive prioritizing method of using either gene expression data or protein interaction data only and a recent gene ranking algorithm GeneRank. We then apply our method to prioritize genes for Alzheimer's disease. AVAILABILITY: The code in this paper is available upon request.

Algorithms↗

Induction of atopic eczema/dermatitis syndrome-like skin lesions by repeated topical application of a crude extract of Dermatophagoides pteronyssinus in NC/Nga mice.

Mite antigen has been considered to play important roles in the development of atopic eczema/dermatitis syndrome (AEDS). In the present study, we attempted to induce an AEDS-like skin lesion in mice using Dermatophagoides pteronyssinus crude extract (DPE) as an antigen and performed pathophysiological evaluations. Ears of mice were tape-stripped and DPE was painted 3 times a week. Eczematous skin lesion and ear swelling were apparent in NC/Nga mice treated with DPE after 2 weeks, whereas neither skin lesion nor ear swelling were observed in BALB/c mice even after 30 days. Histological evaluation demonstrated that edema, epidermal hyperplasia and the accumulation of inflammatory cells were apparent in the ears of DPE-treated NC/Nga mice. In contrast to skin lesion and ear swelling, total serum IgE levels were increased in both NC/Nga and BALB/c mice. Treatment with DPE also increased auricular lymph node weight in both NC/Nga mice and BALB/c mice. To further characterize, we analyzed cytokine mRNA expression in ears and lymph nodes of DPE-treated NC/Nga mice. Increased expression of IL-4 and TNF-alpha mRNA was observed in both ears and lymph nodes of NC/Nga mice treated with DPE. Additionally, there was no change in the responsiveness of BALB/c mice to DPE treatment by adaptive transfer of serum from DPE-treated NC/Nga mice to BALB/c mice. Taken together, our results indicate that eczematous skin lesion and ear swelling caused by repeated application of DPE in NC/Nga mice has a Th2-dominant background and that inflammation is involved in this process. The animal model of AEDS established in this report may be used to investigate the pathogenesis of AEDS and evaluate the potential therapeutic agents for AEDS.

Administration, Topical↗

An integrated approach to the prediction of domain-domain interactions.

BACKGROUND: The development of high-throughput technologies has produced several large scale protein interaction data sets for multiple species, and significant efforts have been made to analyze the data sets in order to understand protein activities. Considering that the basic units of protein interactions are domain interactions, it is crucial to understand protein interactions at the level of the domains. The availability of many diverse biological data sets provides an opportunity to discover the underlying domain interactions within protein interactions through an integration of these biological data sets. RESULTS: We combine protein interaction data sets from multiple species, molecular sequences, and gene ontology to construct a set of high-confidence domain-domain interactions. First, we propose a new measure, the expected number of interactions for each pair of domains, to score domain interactions based on protein interaction data in one species and show that it has similar performance as the E-value defined by Riley et al. Our new measure is applied to the protein interaction data sets from yeast, worm, fruitfly and humans. Second, information on pairs of domains that coexist in known proteins and on pairs of domains with the same gene ontology function annotations are incorporated to construct a high-confidence set of domain-domain interactions using a Bayesian approach. Finally, we evaluate the set of domain-domain interactions by comparing predicted domain interactions with those defined in iPfam database that were derived based on protein structures. The accuracy of predicted domain interactions are also confirmed by comparing with experimentally obtained domain interactions from H. pylori. As a result, a total of 2,391 high-confidence domain interactions are obtained and these domain interactions are used to unravel detailed protein and domain interactions in several protein complexes. CONCLUSION: Our study shows that integration of multiple biological data sets based on the Bayesian approach provides a reliable framework to predict domain interactions. By integrating multiple data sources, the coverage and accuracy of predicted domain interactions can be significantly increased.

Algorithms↗

Diffusion kernel-based logistic regression models for protein function prediction.

Assigning functions to unknown proteins is one of the most important problems in proteomics. Several approaches have used protein-protein interaction data to predict protein functions. We previously developed a Markov random field (MRF) based method to infer a protein's functions using protein-protein interaction data and the functional annotations of its protein interaction partners. In the original model, only direct interactions were considered and each function was considered separately. In this study, we develop a new model which extends direct interactions to all neighboring proteins, and one function to multiple functions. The goal is to understand a protein's function based on information on all the neighboring proteins in the interaction network. We first developed a novel kernel logistic regression (KLR) method based on diffusion kernels for protein interaction networks. The diffusion kernels provide means to incorporate all neighbors of proteins in the network. Second, we identified a set of functions that are highly correlated with the function of interest, referred to as the correlated functions, using the chi-square test. Third, the correlated functions were incorporated into our new KLR model. Fourth, we extended our model by incorporating multiple biological data sources such as protein domains, protein complexes, and gene expressions by converting them into networks. We showed that the KLR approach of incorporating all protein neighbors significantly improved the accuracy of protein function predictions over the MRF model. The incorporation of multiple data sets also improved prediction accuracy. The prediction accuracy is comparable to another protein function classifier based on the support vector machine (SVM), using a diffusion kernel. The advantages of the KLR model include its simplicity as well as its ability to explore the contribution of neighbors to the functions of proteins of interest.

Databases, Protein↗