Search PubMed⌕ Search

Biomedical subjects

A Godzik

Publications and source records attributed to A Godzik.

At least 37 records · Page 2Linked to original sources

From fold to function predictions: an apoptosis regulator protein BID.

With the rapidly increasing pace of genome sequencing projects and the resulting flood of predicted amino acid sequences of uncharacterized proteins, protein sequence analysis, and in particular, protein structure prediction is quickly gaining in importance. Prediction algorithms can be used for preliminary annotation of newly sequenced proteins and, at least in some cases, provide insights into their function and specific mode of action. Such annotations for several microbial genomes were performed by several groups and placed in public domain for evaluation. An example presented in this work comes from a related project of structural and functional predictions for proteins involved in the process of controlled cell death (apoptosis). The BID protein belongs to an important class of regulators of apoptosis identified by short sequence motifs. Here, several fold prediction methods are used to build a series of three-dimensional models. Structure analysis of the models with reference to the biological data available allows selection of the most appropriate model. It is found that the most likely structural model of BID is built on the structure of Bcl-X(L). The model is discussed in terms of experimental data on specific proteolytic cleavage of BID and its effect on BID interactions with other proteins and membranes.

Algorithms↗

Saturated BLAST: an automated multiple intermediate sequence search used to detect distant homology.

MOTIVATION: Two proteins can have a similar 3-dimensional structure and biological function, but have sequences sufficiently different that traditional protein sequence comparison algorithms do not identify their relationship. The desire to identify such relations has led to the development of more sensitive sequence alignment strategies. One such strategy is the Intermediate Sequence Search (ISS), which connects two proteins through one or more intermediate sequences. In its brute-force implementation, ISS is a strategy that repetitively uses the results of the previous query as new search seeds, making it time-consuming and difficult to analyze. RESULTS: Saturated BLAST is a package that performs ISS in an efficient and automated manner. It was developed using Perl and Perl/Tk and implemented on the LINUX operating system. Starting with a protein sequence, Saturated BLAST runs a BLAST search and identifies representative sequences for the next generation of searches. The procedure is run until convergence or until some predefined criteria are met. Saturated BLAST has a friendly graphic user interface, a built-in BLAST result parser, several multiple alignment tools, clustering algorithms and various filters for the elimination of false positives, thereby providing an easy way to edit, visualize, analyze, monitor and control the search. Besides detecting remote homologies, Saturated BLAST can be used to maintain protein family databases and to search for new genes in genomic databases.

Algorithms↗

Comparison of sequence profiles. Strategies for structural predictions using sequence information.

Distant homologies between proteins are often discovered only after three-dimensional structures of both proteins are solved. The sequence divergence for such proteins can be so large that simple comparison of their sequences fails to identify any similarity. New generation of sensitive alignment tools use averaged sequences of entire homologous families (profiles) to detect such homologies. Several algorithms, including the newest generation of BLAST algorithms and BASIC, an algorithm used in our group to assign fold predictions for proteins from several genomes, are compared to each other on the large set of structurally similar proteins with little sequence similarity. Proteins in the benchmark are classified according to the level of their similarity, which allows us to demonstrate that most of the improvement of the new algorithms is achieved for proteins with strong functional similarities, with almost no progress in recognizing distant fold similarities. It is also shown that details of profile calculation strongly influence its sensitivity in recognizing distant homologies. The most important choice is how to include information from diverging members of the family, avoiding generating false predictions, while accounting for entire sequence divergence within a family. PSI-BLAST takes a conservative approach, deriving a profile from core members of the family, providing a solid improvement without almost any false predictions. BASIC strives for better sensitivity by increasing the weight of divergent family members and paying the price in lower reliability. A new FFAS algorithm introduced here uses a new procedure for profile generation that takes into account all the relations within the family and matches BASIC sensitivity with PSI-BLAST like reliability.

Algorithms↗

Improving the quality of twilight-zone alignments.

Several recent publications illustrated advantages of using sequence profiles in recognizing distant homologies between proteins. At the same time, the practical usefulness of distant homology recognition depends not only on the sensitivity of the algorithm, but also on the quality of the alignment between a prediction target and the template from the database of known proteins. Here, we study this question for several supersensitive protein algorithms that were previously compared in their recognition sensitivity (Rychlewski et al., 2000). A database of protein pairs with similar structures, but low sequence similarity is used to rate the alignments obtained with several different methods, which included sequence-sequence, sequence-profile, and profile-profile alignment methods. We show that incorporation of evolutionary information encoded in sequence profiles into alignment calculation methods significantly increases the alignment accuracy, bringing them closer to the alignments obtained from structure comparison. In general, alignment quality is correlated with recognition and alignment score significance. For every alignment method, alignments with statistically significant scores correlate with both correct structural templates and good quality alignments. At the same time, average alignment lengths differ in various methods, making the comparison between them difficult. For instance, the alignments obtained by FFAS, the profile-profile alignment algorithm developed in our group are always longer that the alignments obtained with the PSI-BLAST algorithms. To address this problem, we develop methods to truncate or extend alignments to cover a specified percentage of protein lengths. In most cases, the elongation of the alignment by profile-profile methods is reasonable, adding fragments of similar structure. The examples of erroneous alignment are examined and it is shown that they can be identified based on the model quality.

Algorithms↗

Sensitive sequence comparison as protein function predictor.

Protein function assignments based on postulated homology as recognized by high sequence similarity are used routinely in genome analysis. Improvements in sensitivity of sequence comparison algorithms got to the point, that proteins with previously undetectable sequence similarity, such as for instance 10-15% of identical residues, sometimes can be classified as similar. What is the relation between such proteins? Is it possible that they are homologous? What is the practical significance of detecting such similarities? A simplified analysis of the relation between sequence similarity and function similarity is presented here for the well-characterized proteins from the E. coli genome. Using a simple measure of functional similarity based on E.C. classification of enzymes, it is shown that it correlates well with sequence similarity measured by statistical significance of the alignment score. Proteins, similar by this standard, even in cases of low sequence identity, have a much larger chance of having similar function than the randomly chosen protein pairs. Interesting exceptions to these rules are discussed.

Algorithms↗

Protein structure prediction in biology and medicine

Protein structure prediction from sequence has made rapid advances in the last few years, gaining increased attention after the success of the CASP meetings. Specifically, threading predictions were shown to work for many cases previously thought to be outside of the scope of prediction methods. The papers in this session address a wide scope of topics, ranging from techniques for validation of prediction methods and further improvements of threading algorithms, to specific applications of protein structure predictions in biology.

Journal Article↗

Search for a new description of protein topology and local structure.

A novel description of protein structure in terms of the generalized secondary structure elements (GSSE) is proposed. GSSE's are defined as fragments of the protein structure where the chain doesn't radically change its direction. In this new language, global protein topology becomes a particular arrangement of the relatively small number of large, rod like GSSE's. Protein topology can be described by an adjacency matrix giving information, which GSSE's are close in space to each other and defining a graph, where GSSE's are equivalent to vertices and interactions between them to edges. The information about the local structure is translated into the local density of pseudo-Calpha atoms along the chain and the curvature of the chain. This new description has a number of interesting and useful features. For instance, enumeration theorems of graph theory can be used to estimate a number of possible topologies for a protein built from a given number of elements. Different topologies, including novel ones, can be generated from the known by various permutations of elements. Many new regularities in protein structures become suddenly visible in a new description. A new local structure description is more amenable to predictions and easier to use in fold predictions.

Animals↗

Ion channel activity of the BH3 only Bcl-2 family member, BID.

BID is a member of the BH3-only subgroup of Bcl-2 family proteins that displays pro-apoptotic activity. The NH(2)-terminal region of BID contains a caspase-8 (Casp-8) cleavage site and the cleaved form of BID translocates to mitochondrial membranes where it is a potent inducer of cytochrome c release. Secondary structure and fold predictions suggest that BID has a high degree of alpha-helical content and structural similarity to Bcl-X(L), which itself is highly similar to bacterial pore-forming toxins. Moreover, circular dichroism analysis confirmed a high alpha-helical content of BID. Amino-terminal truncated BIDDelta1-55, mimicking the Casp-8-cleaved molecule, formed channels in planar bilayers at neutral pH and in liposomes at acidic pH. In contrast, full-length BID displayed channel activity only at nonphysiological pH 4.0 (but not at neutral pH) in planar bilayers and failed to form channels in liposomes even under acidic conditions. On a single channel level, BIDDelta1-55 channels were voltage-gated and exhibited multiconductance behavior at neutral pH. When full-length BID was cleaved by Casp-8, it too demonstrated channel activity similar to that seen with BIDDelta1-55. Thus, BID appears to share structural and functional similarity with other Bcl-2 family proteins known to have channel-forming activity, but its activity exhibits a novel form of activation: proteolytic cleavage.

Animals↗

The Helicobacter pylori genome: from sequence analysis to structural and functional predictions.

Fold assignments for proteins from the Helicobacter pylori genome are carried out using BASIC, a profile-profile alignment algorithm recently tested on the Mycoplasma genitalium and Escherichia coli genomes. The fold assignments are followed by automated function evaluation, based on the multilevel description of functional sites in proteins. Over 40% of the proteins encoded in the H. pylori genome can be recognized as belonging to a protein family with known structure. Previous estimates suggested that only 10-15% of genome proteins could be characterized this way. This dramatic increase in the number of recognized homologies between H. pylori proteins and structurally characterized protein families is partly due to the rapid increase of the database of known protein structures, but mostly it is due to the significant improvement in prediction algorithms. Knowledge of a protein fold adds a new dimension to our understanding of its function and, similarly, structure prediction can also add to understanding, verification, and/or prediction of function for uncharacterized proteins. Several examples analyzed in more detail in this article illustrate insights that can be achieved from structure and detailed function prediction.

Algorithms↗

CAFASP-1: critical assessment of fully automated structure prediction methods.

The results of the first Critical Assessment of Fully Automated Structure Prediction (CAFASP-1) are presented. The objective was to evaluate the success rates of fully automatic web servers for fold recognition which are available to the community. This study was based on the targets used in the third meeting on the Critical Assessment of Techniques for Protein Structure Prediction (CASP-3). However, unlike CASP-3, the study was not a blind trial, as it was held after the structures of the targets were known. The aim was to assess the performance of methods without the user intervention that several groups used in their CASP-3 submissions. Although it is clear that "human plus machine" predictions are superior to automated ones, this CAFASP-1 experiment is extremely valuable for users of our methods; it provides an indication of the performance of the methods alone, and not of the "human plus machine" performance assessed in CASP. This information may aid users in choosing which programs they wish to use and in evaluating the reliability of the programs when applied to their specific prediction targets. In addition, evaluation of fully automated methods is particularly important to assess their applicability at genomic scales. For each target, groups submitted the top-ranking folds generated from their servers. In CAFASP-1 we concentrated on fold-recognition web servers only and evaluated only recognition of the correct fold, and not, as in CASP-3, alignment accuracy. Although some performance differences appeared within each of the four target categories used here, overall, no single server has proved markedly superior to the others. The results showed that current fully automated fold recognition servers can often identify remote similarities when pairwise sequence search methods fail. Nevertheless, in only a few cases outside the family-level targets has the score of the top-ranking fold been significant enough to allow for a confident fully automated prediction. Because the goals, rules, and procedures of CAFASP-1 were different from those used at CASP-3, the results reported here are not comparable with those reported in CASP-3. Nevertheless, it is clear that current automated fold recognition methods can not yet compete with "human-expert plus machine" predictions. Finally, CAFASP-1 has been useful in identifying the requirements for a future blind trial of automated served-based protein structure prediction.

Algorithms↗

Functional insights from structural predictions: analysis of the Escherichia coli genome.

Fold assignments for proteins from the Escherichia coli genome are carried out using BASIC, a profile-profile alignment algorithm, recently tested on fold recognition benchmarks and on the Mycoplasma genitalium genome and PSI BLAST, the newest generation of the de facto standard in homology search algorithms. The fold assignments are followed by automated modeling and the resulting three-dimensional models are analyzed for possible function prediction. Close to 30% of the proteins encoded in the E. coli genome can be recognized as homologous to a protein family with known structure. Most of these homologies (23% of the entire genome) can be recognized both by PSI BLAST and BASIC algorithms, but the latter recognizes an additional 260 homologies. Previous estimates suggested that only 10-15% of E. coli proteins can be characterized this way. This dramatic increase in the number of recognized homologies between E. coli proteins and structurally characterized protein families is partly due to the rapid increase of the database of known protein structures, but mostly it is due to the significant improvement in prediction algorithms. Knowing protein structure adds a new dimension to our understanding of its function and the predictions presented here can be used to predict function for uncharacterized proteins. Several examples, analyzed in more detail in this paper, include the DPS protein protecting DNA from oxidative damage (predicted to be homologous to ferritin with iron ion acting as a reducing agent) and the ahpC/tsa family of proteins, which provides resistance to various oxidating agents (predicted to be homologous to glutathione peroxidase).

Algorithms↗

From fold predictions to function predictions: automation of functional site conservation analysis for functional genome predictions.

A database of functional sites for proteins with known structures, SITE, is constructed and used in conjunction with a simple pattern matching program SiteMatch to evaluate possible function conservation in a recently constructed database of fold predictions for Escherichia coli proteins (Rychlewski L et al., 1999, Protein Sci 8:614-624). In this and other prediction databases, fold predictions are based on algorithms that can recognize weak sequence similarities and putatively assign new proteins into already characterized protein families. It is not clear whether such sequence similarities arise from distant homologies or general similarity of physicochemical features along the sequence. Leaving aside the important question of nature of relations within fold superfamilies, it is possible to assess possible function conservation by looking at the pattern of conservation of crucial functional residues. SITE consists of a multilevel function description based on structure annotations and structure analyses. In particular, active site residues, ligand binding residues, and patterns of hydrophobic residues on the protein surface are used to describe different functional features. SiteMatch, a simple pattern matching program, is designed to check the conservation of residues involved in protein activity in alignments generated by any alignment method. Here, this procedure is used to study conservation of functional features in alignments between protein sequences from the E. coli genome and their optimal structural templates. The optimal templates were identified and alignments taken from the database of genomic structural predictions was described in a previous publication (Rychlewski L et al., 1999, Protein Sci 8:614-624). An automated assessment of function conservation is used to analyze the relation between fold and function similarity for a large number of fold predictions. For instance, it is shown that identifying low significance predictions with a high level of functional residue conservations can be used to extend the prediction sensitivity for fold prediction methods. Over 100 new fold/function predictions in this class were obtained in the E. coli genome. At the same time, about 30% of our previous fold predictions are not confirmed as function predictions, further highlighting the problem of function divergence in fold superfamilies.

Algorithms↗

Functional analysis of the Escherichia coli genome using the sequence-to-structure-to-function paradigm: identification of proteins exhibiting the glutaredoxin/thioredoxin disulfide oxidoreductase activity.

The application of an automated method for the screening of protein activity based on the sequence-to-structure-to-function paradigm is presented for the complete Escherichia coli genome. First, the structure of the protein is identified from its sequence using a threading algorithm, which aligns the sequences to the best matching structure in a structural database and extends sequence analysis well beyond the limits of local sequence identity. Then, the active site is identified in the resulting sequence-to-structure alignment using a "fuzzy functional form" (FFF), a three-dimensional descriptor of the active site of a protein. Here, this sequence-to-structure-to-function concept is applied to analysis of the complete E. coli genome, i.e. all E. coli open reading frames (ORFs) are screened for the thiol-disulfide oxidoreductase activity of the glutaredoxin/thioredoxin protein family. We show that the method can identify the active sites in ten sequences that are known to or proposed to exhibit this activity. Furthermore, oxidoreductase activity is predicted in two other sequences that have not been identified previously. This method distinguishes protein pairs with similar active sites from proteins pairs that are just topological cousins, i.e. those having similar global folds, but not necessarily similar active sites. Thus, this method provides a novel approach for extraction of active site and functional information based on three-dimensional structures, rather than simple sequence analysis. Prediction of protein activity is fully automated and easily extendible to new functions. Finally, it is demonstrated here that the method can be applied to complete genome database analysis.

Algorithms↗

Fold prediction by a hierarchy of sequence, threading, and modeling methods.

Several fold recognition algorithms are compared to each other in terms of prediction accuracy and significance. It is shown that on standard benchmarks, hybrid methods, which combine scoring based on sequence-sequence and sequence-structure matching, surpass both sequence and threading methods in the number of accurate predictions. However, the sequence similarity contributes most to the prediction accuracy. This strongly argues that most examples of apparently nonhomologous proteins with similar folds are actually related by evolution. While disappointing from the perspective of the fundamental understanding of protein folding, this adds a new significance to fold recognition methods as a possible first step in function prediction. Despite hybrid methods being more accurate at fold prediction than either the sequence or threading methods, each of the methods is correct in some cases where others have failed. This partly reflects a different perspective on sequence/structure relationship embedded in various methods. To combine predictions from different methods, estimates of significance of predictions are made for all methods. With the help of such estimates, it is possible to develop a "jury" method, which has accuracy higher than any of the single methods. Finally, building full three-dimensional models for all top predictions helps to eliminate possible false positives where alignments, which are optimal in the one-dimensional sequences, lead to unsolvable sterical conflicts for the full three-dimensional models.

Algorithms↗

Fold and function predictions for Mycoplasma genitalium proteins.

BACKGROUND: Uncharacterized proteins from newly sequenced genomes provide perfect targets for fold and function prediction. RESULTS: For 38% of the entire genome of Mycoplasma genitalium, sequence similarity to a protein with a known structure can be recognized using a new sequence alignment algorithm. When comparing genomes of M. genitalium and Escherichia coli, > 80% of M. genitalium proteins have a significant sequence similarity to a protein in E. coli and there are > 40 examples that have not been recognized before. For all cases of proteins with significant profile similarities, there are strong analogies in their functions, if the functions of both proteins are known. The results presented here and other recent results strongly support the argument that such proteins are actually homologous. Assuming this homology allows one to make tentative functional assignments for > 50 previously uncharacterized proteins, including such intriguing cases as the putative beta-lactam antibiotic resistance protein in M. gentalium. CONCLUSIONS: Using a new profile-to-profile alignment algorithm, the three-dimensional fold can be predicted for almost 40% of proteins from a genome of the small bacterium M. genitalium, and tentative function can be assigned to almost 80% of the entire genome. Some predictions lead to new insights about known functions or point to hitherto unexpected features of M. genitalium.

Bacterial Proteins↗

Functional analysis of the Escherichia coli genome for members of the alpha/beta hydrolase family.

BACKGROUND: Database-searching methods based on sequence similarity have become the most commonly used tools for characterizing newly sequenced proteins. Due to the often underestimated functional diversity in protein families and superfamilies, however, it is difficult to make the characterization specific and accurate. In this work, we have extended a method for active-site identification from predicted protein structures. RESULTS: The structural conservation and variation of the active sites of the alpha/beta hydrolases with known structures were studied. The similarities were incorporated into a three-dimensional motif that specifies essential requirements for the enzymatic functions. A threading algorithm was used to align 651 Escherichia coli open reading frames (ORFs) to one of the members of the alpha/beta hydrolase fold family. These ORFs were then screened according to our three-dimensional motif and with an extra requirement that demands conservation of the key active-site residues among the proteins that bear significant sequence similarity to the ORFs. 17 ORFs from E. coli were predicted to have hydrolase activity and their putative active-site residues were identified. Most were in agreement with the experiments and results of other database-searching methods. The study further suggests that YHET_ECOLI, a hypothetical protein classified as a member of the UPF0017 family (an uncharacterized protein family), bears all the hallmarks of the alpha/beta hydrolase family. CONCLUSIONS: The novel feature of our method is that it uses three-dimensional structural information for function prediction. The results demonstrate the importance and necessity of such a method to fill the gap between sequence alignment and function prediction; furthermore, the method provides a way to verify the structure predictions, which enables an expansion of the applicable scope of the threading algorithms.

Amino Acid Sequence↗

Function driven protein evolution. A possible proto-protein for the RNA-binding proteins.

We introduce a hypothesis that present day proteins evolved from "proto-proteins," small 15-20 residue peptides with some elements of secondary structure and primitive function. Increasingly stable and functional proteins arose by adding structural elements to produce the small domains or protein modules that we would recognize today. From this point of view, the surprising similarities between small structural fragments of large proteins, that are usually taken as examples of convergent, function-driven evolution, are interpreted in exactly the opposite way--as traces of common evolutionary origin. As an example, a hypothetical evolutionary tree for two families of RNA binding proteins, the OB fold, a family of all beta proteins, and RBD fold, an alpha/beta protein family is presented. We argue that both protein families could have evolved from the same RNA-binding proto-protein, which had a form of beta-loop-beta RNA binding motif.

Amino Acid Sequence↗

Derivation and testing of pair potentials for protein folding. When is the quasichemical approximation correct?

Many existing derivations of knowledge-based statistical pair potentials invoke the quasichemical approximation to estimate the expected side-chain contact frequency if there were no amino acid pair-specific interactions. At first glance, the quasichemical approximation that treats the residues in a protein as being disconnected and expresses the side-chain contact probability as being proportional to the product of the mole fractions of the pair of residues would appear to be rather severe. To investigate the validity of this approximation, we introduce two new reference states in which no specific pair interactions between amino acids are allowed, but in which the connectivity of the protein chain is retained. The first estimates the expected number of side-chain contracts by treating the protein as a Gaussian random coil polymer. The second, more realistic reference state includes the effects of chain connectivity, secondary structure, and chain compactness by estimating the expected side-chain contrast probability by placing the sequence of interest in each member of a library of structures of comparable compactness to the native conformation. The side-chain contact maps are not allowed to readjust to the sequence of interest, i.e., the side chains cannot repack. This situation would hold rigorously if all amino acids were the same size. Both reference states effectively permit the factorization of the side-chain contact probability into sequence-dependent and structure-dependent terms. Then, because the sequence distribution of amino acids in proteins is random, the quasichemical approximation to each of these reference states is shown to be excellent. Thus, the range of validity of the quasichemical approximation is determined by the magnitude of the side-chain repacking term, which is, at present, unknown. Finally, the performance of these two sets of pair interaction potentials as well as side-chain contact fraction-based interaction scales is assessed by inverse folding tests both without and with allowing for gaps.

Models, Chemical↗