Search PubMed⌕ Search

Biomedical subjects

George Karypis

Publications and source records attributed to George Karypis.

11 recordsLinked to original sources

Building multiclass classifiers for remote homology detection and fold recognition.

BACKGROUND: Protein remote homology detection and fold recognition are central problems in computational biology. Supervised learning algorithms based on support vector machines are currently one of the most effective methods for solving these problems. These methods are primarily used to solve binary classification problems and they have not been extensively used to solve the more general multiclass remote homology prediction and fold recognition problems. RESULTS: We present a comprehensive evaluation of a number of methods for building SVM-based multiclass classification schemes in the context of the SCOP protein classification. These methods include schemes that directly build an SVM-based multiclass model, schemes that employ a second-level learning approach to combine the predictions generated by a set of binary SVM-based classifiers, and schemes that build and combine binary classifiers for various levels of the SCOP hierarchy beyond those defining the target classes. CONCLUSION: Analyzing the performance achieved by the different approaches on four different datasets we show that most of the proposed multiclass SVM-based classification approaches are quite effective in solving the remote homology prediction and fold recognition problems and that the schemes that use predictions from binary models constructed for ancestral categories within the SCOP hierarchy tend to not only lead to lower error rates but also reduce the number of errors in which a superfamily is assigned to an entirely different fold and a fold is predicted as being from a different SCOP class. Our results also show that the limited size of the training data makes it hard to learn complex second-level models, and that models of moderate complexity lead to consistently better results.

Algorithms↗

YASSPP: better kernels and coding schemes lead to improvements in protein secondary structure prediction.

The accurate prediction of a protein's secondary structure plays an increasingly critical role in predicting its function and tertiary structure, as it is utilized by many of the current state-of-the-art methods for remote homology, fold recognition, and ab initio structure prediction. We developed a new secondary structure prediction algorithm called YASSPP, which uses a pair of cascaded models constructed from two sets of binary SVM-based models. YASSPP uses an input coding scheme that combines both position-specific and nonposition-specific information, utilizes a kernel function designed to capture the sequence conservation signals around the local window of each residue, and constructs a second-level model by incorporating both the three-state predictions produced by the first-level model and information about the original sequence. Experiments on three standard datasets (RS126, CB513, and EVA common subset 4) show that YASSPP is capable of producing the highest Q3 and SOV scores than that achieved by existing widely used schemes such as PSIPRED, SSPro 4.0, SAM-T99sec, as well as previously developed SVM-based schemes. On the EVA dataset it achieves a Q3 and SOV score of 79.34 and 78.65%, which are considerably higher than the best reported scores of 77.64 and 76.05%, respectively.

Algorithms↗

QCRNA 1.0: a database of quantum calculations for RNA catalysis.

This work outlines a new on-line database of quantum calculations for RNA catalysis (QCRNA) available via the worldwide web at http://theory.chem.umn.edu/QCRNA. The database contains high-level density functional calculations for a large range of molecules, complexes and chemical mechanisms important to phosphoryl transfer reactions and RNA catalysis. Calculations are performed using a strict, consistent protocol such that a wealth of cross-comparisons can be made to elucidate meaningful trends in biological phosphate reactivity. Currently, around 2000 molecules have been collected in varying charge states in the gas phase and in solution. Solvation was treated with both the PCM and COSMO continuum solvation models. The data can be used to study important trends in reactivity of biological phosphates, or used as benchmark data for the design of new semiempirical quantum models for hybrid quantum mechanical/molecular mechanical simulations.

Computer Graphics↗

Profile-based direct kernels for remote homology detection and fold recognition.

MOTIVATION: Protein remote homology detection is a central problem in computational biology. Supervised learning algorithms based on support vector machines are currently one of the most effective methods for remote homology detection. The performance of these methods depends on how the protein sequences are modeled and on the method used to compute the kernel function between them. RESULTS: We introduce two classes of kernel functions that are constructed by combining sequence profiles with new and existing approaches for determining the similarity between pairs of protein sequences. These kernels are constructed directly from these explicit protein similarity measures and employ effective profile-to-profile scoring schemes for measuring the similarity between pairs of proteins. Experiments with remote homology detection and fold recognition problems show that these kernels are capable of producing results that are substantially better than those produced by all of the existing state-of-the-art SVM-based methods. In addition, the experiments show that these kernels, even when used in the absence of profiles, produce results that are better than those produced by existing non-profile-based schemes. AVAILABILITY: The programs for computing the various kernel functions are available on request from the authors.

Algorithms↗

Data clustering in life sciences.

Clustering has a wide range of applications in life sciences and over the years has been used in many areas ranging from the analysis of clinical information, phylogeny, genomics, and proteomics. The primary goal of this article is to provide an overview of the various issues involved in clustering large biological datasets, describe the merits and underlying assumptions of some of the commonly used clustering approaches, and provide insights on how to cluster datasets arising in various areas within life sciences. We also provide a brief introduction to CLUTO, a general purpose toolkit for clustering various datasets, with an emphasis on its applications to problems and analysis requirements within life sciences.

Algorithms↗

Macromolecule mass spectrometry: citation mining of user documents.

Identifying research users, applications, and impact is important for research performers, managers, evaluators, and sponsors. Identification of the user audience and the research impact is complex and time consuming due to the many indirect pathways through which fundamental research can impact applications. This paper identified the literature pathways through which two highly-cited papers of 2002 Chemistry Nobel Laureates Fenn and Tanaka impacted research, technology development, and applications. Citation Mining, an integration of citation bibliometrics and text mining, was applied to the >1600 first generation Science Citation Index (SCI) citing papers to Fenn's 1989 Science paper on Electrospray Ionization for Mass Spectrometry, and to the >400 first generation SCI citing papers to Tanaka's 1988 Rapid Communications in Mass Spectrometry paper on Laser Ionization Time-of-Flight Mass Spectrometry. Bibliometrics was performed on the citing papers to profile the user characteristics. Text mining was performed on the citing papers to identify the technical areas impacted by the research, and the relationships among these technical areas.

Authorship↗

A Boolean algorithm for reconstructing the structure of regulatory networks.

Advances in transcriptional analysis offer great opportunities to delineate the structure and hierarchy of regulatory networks in biochemical systems. We present an approach based on Boolean analysis to reconstruct a set of parsimonious networks from gene disruption and over expression data. Our algorithms, Causal Predictor (CP) and Relaxed Causal Predictor (RCP) distinguish the direct and indirect causality relations from the non-causal interactions, thus significantly reducing the number of miss-predicted edges. The algorithms also yield substantially fewer plausible networks. This greatly reduces the number of experiments required to deduce a unique network from the plausible network structures. Computational simulations are presented to substantiate these results. The algorithms are also applied to reconstruct the entire network of galactose utilization pathway in Saccharomyces cerevisiae. These algorithms will greatly facilitate the elucidation of regulatory networks using large scale gene expression profile data.

Algorithms↗

Interferon-inducible gene expression signature in peripheral blood cells of patients with severe lupus.

Systemic lupus erythematosus (SLE) is a complex, inflammatory autoimmune disease that affects multiple organ systems. We used global gene expression profiling of peripheral blood mononuclear cells to identify distinct patterns of gene expression that distinguish most SLE patients from healthy controls. Strikingly, about half of the patients studied showed dysregulated expression of genes in the IFN pathway. Furthermore, this IFN gene expression "signature" served as a marker for more severe disease involving the kidneys, hematopoetic cells, and/or the central nervous system. These results provide insights into the genetic pathways underlying SLE, and identify a subgroup of patients who may benefit from therapies targeting the IFN pathway.

Down-Regulation↗

wCLUTO: a Web-enabled clustering toolkit.

As structural and functional genomics efforts provide the biological community with ever-broadening sets of interrelated data, the need to explore such complex information for subtle relationships expands. We present wCLUTO, a Web-enabled version of the stand-alone application CLUTO, designed to apply clustering methods to genomic information. Its first application is focused on the clustering transcriptome data from microarrays. Data can be uploaded by the user into the clustering tool, a choice of several clustering methods can be made and configured, and data are presented to the user in a variety of visual formats, including a three-dimensional "mountain" view of the clusters. Parameters can be explored to rapidly examine a variety of clustering results, and the resulting clusters can be downloaded either for manipulation by other programs or to be saved in a format for publication.

Algorithms↗