Search PubMed⌕ Search

Biomedical subjects

Krzysztof Fidelis

Publications and source records attributed to Krzysztof Fidelis.

12 recordsLinked to original sources

Using local gene expression similarities to discover regulatory binding site modules.

BACKGROUND: We present an approach designed to identify gene regulation patterns using sequence and expression data collected for Saccharomyces cerevisae. Our main goal is to relate the combinations of transcription factor binding sites (also referred to as binding site modules) identified in gene promoters to the expression of these genes. The novel aspects include local expression similarity clustering and an exact IF-THEN rule inference algorithm. We also provide a method of rule generalization to include genes with unknown expression profiles. RESULTS: We have implemented the proposed framework and tested it on publicly available datasets from yeast S. cerevisae. The testing procedure consists of thorough statistical analyses of the groups of genes matching the rules we infer from expression data against known sets of co-regulated genes. For this purpose we have used published ChIP-Chip data and Gene Ontology annotations. In order to make these tests more objective we compare our results with recently published similar studies. CONCLUSION: Results we obtain show that local expression similarity clustering greatly enhances overall quality of the derived rules, both in terms of enrichment of Gene Ontology functional annotation and coherence with ChIP-Chip binding data. Our approach thus provides reliable hypotheses on co-regulation that can be experimentally verified. An important feature of the method is its reliance only on widely accessible sequence and expression data. The same procedure can be easily applied to other microbial organisms.

Binding Sites↗

Generalized modeling of enzyme-ligand interactions using proteochemometrics and local protein substructures.

Modeling and understanding protein-ligand interactions is one of the most important goals in computational drug discovery. To this end, proteochemometrics uses structural and chemical descriptors from several proteins and several ligands to induce interaction-models. Here, we present a new and generalized approach in which proteins varying greatly in terms of sequence and structure are represented by a library of local substructures. Using linear regression and rule-based learning, we combine such local substructures with chemical descriptors from the ligands to model binding affinity for a training set of hydrolase and lyase enzymes. We evaluate the predictive performance of these models using cross validation and sets of unseen ligand with unknown three-dimensional structure. The models are shown to generalize by outperforming models using descriptors from only proteins or only ligands, or models using global structure similarities rather than local similarities. Thus, we demonstrate that this approach is capable of describing dependencies between local structural properties and ligands in otherwise dissimilar protein structures. These dependencies are often, but not always, associated with local substructures that are in contact with the ligands. Finally, we show that strongly bound enzyme-ligand complexes require the presence of particular local substructures, while weakly bound complexes may be described by the absence of certain properties. The results demonstrate that the alignment-independent approach using local substructures is capable of describing protein-ligand interaction for largely different proteins and hence opens up for proteochemometrics-analysis of the interaction-space of entire proteomes. Current approaches are limited to families of closely related proteins. families of closely related proteins.

Algorithms↗

Critical assessment of methods of protein structure prediction (CASP)--round 6.

This article is an introduction to the special issue of the journal Proteins, dedicated to the sixth CASP experiment to assess the state of the art in protein structure prediction. The article describes the conduct of the experiment and the categories of prediction included, and outlines the evaluation and assessment procedures. A brief summary of progress over the decade of CASP experiments is also provided.

Algorithms↗

CASP6 data processing and automatic evaluation at the protein structure prediction center.

We present a short overview of the system governing data processing and automatic evaluation of predictions in CASP6, implemented at the Livermore Protein Structure Prediction Center. The system incorporates interrelated facilities for registering participants, collecting prediction targets from crystallographers and NMR spectroscopists and making them available to the CASP6 participants, accepting predictions and providing their preliminary evaluation, and finally, storing and visualizing results. We have automatically evaluated predictions submitted to CASP6 using criteria and methods developed over the successive CASP experiments. Also, we have tested a new evaluation technique based on non-rigid-body type superpositions. Approximately the same number of predictions has been submitted to CASP6 as to all previous CASPs combined, making navigation through and understanding of the data particularly challenging. To facilitate this, we have substantially modernized all data handling procedures, including implementation of a dedicated relational database. An overview of our redesigned website is also presented (http://predictioncenter.org/casp6/).

Algorithms↗

System for accepting server predictions in CASP6.

We describe the new CASP system for collecting and verifying predictions generated by servers. The system was developed to ensure reliable execution of the server assessment part of CASP, with particular emphasis on data consistency. Following the principle that predictions should not be modified by anyone but their authors and to allow a later meaningful assessment, submissions are now verified for correctness of format and contents within the strict 48 hour CASP deadlines for this type of submission. This article also provides an overview of the rules governing server participation in CASP6 and some statistics pertaining to servers in CASP6.

Automation↗

Progress over the first decade of CASP experiments.

CASP has now completed a decade of monitoring the state of the art in protein structure prediction. The quality of structure models produced in the latest experiment, CASP6, has been compared with that in earlier CASPs. Significant although modest progress has again been made in the fold recognition regime, and cumulatively, progress in this area is impressive. Models of previously unknown folds again appear to have modestly improved, and several mixed alpha/beta structures have been modeled in a topologically correct manner. Progress remains hard to detect in high sequence identity comparative modeling, but server performance in this area has moved forward.

Algorithms↗

Discovering regulatory binding-site modules using rule-based learning.

Transcription factors regulate expression by binding selectively to sequence sites in cis-regulatory regions of genes. It is therefore reasonable to assume that genes regulated by the same transcription factors should all contain the corresponding binding sites in their regulatory regions and exhibit similar expression profiles as measured by, for example, microarray technology. We have used this assumption to analyze genome-wide yeast binding-site and microarray expression data to reveal the combinatorial nature of gene regulation. We obtained IF-THEN rules linking binding-site combinations (binding-site modules) to genes with particular expression profiles, and thereby provided testable hypotheses on the combinatorial coregulation of gene expression. We showed that genes associated with such rules have a significantly higher probability of being bound by the same transcription factors, as indicated by a genome-wide location analysis, than genes associated with only common binding sites or similar expression. Furthermore, we also found that such genes were significantly more often biologically related in terms of Gene Ontology annotations than genes only associated with common binding sites or similar expression. We analyzed expression data collected under different sets of stress conditions and found many binding-site modules that are conserved over several of these condition sets, as well as modules that are specific to particular biological responses. Our results on the reoccurrence of binding sites in different modules provide specific data on how binding sites may be combined to allow a large number of expression outcomes using relatively few transcription factors.

Algorithms↗

Assessment of progress over the CASP experiments.

The quality of structure models produced in the CASP5 experiment has been compared with that in earlier CASPs. The most significant progress is in the fold recognition regime, where the development of meta-servers has allowed more accurate consensus models to be generated. In contrast to this, there is little evidence of progress in producing more accurate comparative models, particularly those based on sequence identities > 30%. For comparative models based on low-sequence identity and for fold recognition models, accuracy depends primarily on the fraction of the target structure that is similar to an available template, and the quality of the alignment. Overall, these results indicate that there are still no effective methods of improving model quality beyond that obtained by successfully copying a template structure. For models of proteins with previously unknown folds, there appears to be a pause in the previous consistent improvement. There is some evidence that more groups are producing top-quality models, however. Although specific progress between successive experiments is sometimes difficulty to identify, over the history of all the CASPs there has been steady, if sometimes slow, progress in all modeling regimes.

Algorithms↗

Critical assessment of methods of protein structure prediction (CASP)-round V.

This article provides an introduction to the special issue of the journal Proteins dedicated to the fifth CASP experiment to assess the state of the art in protein structure prediction. The article describes the conduct, the categories of prediction, and the evaluation and assessment procedures of the experiment. A brief summary of progress over the five CASP experiments is provided. Related developments in the field are also described.

Computational Biology↗

A novel approach to fold recognition using sequence-derived properties from sets of structurally similar local fragments of proteins.

Comparative modeling methods can consistently produce reliable structural models for protein sequences with more than 25% sequence identity to proteins with known structure. However, there is a good chance that also sequences with lower sequence identity have their structural components represented in structural databases. To this end, we present a novel fragment-based method using sets of structurally similar local fragments of proteins. The approach differs from other fragment-based methods that use only single backbone fragments. Instead, we use a library of groups containing sets of sequence fragments with geometrically similar local structures and extract sequence related properties to assign these specific geometrical conformations to target sequences. We test the ability of the approach to recognize correct SCOP folds for 273 sequences from the 49 most popular folds. 49% of these sequences have the correct fold as their top prediction, while 82% have the correct fold in one of the top five predictions. Moreover, the approach shows no performance reduction on a subset of sequence targets with less than 10% sequence identity to any protein used to build the library.

Amino Acid Sequence↗