Three sisters, different names.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to C Sander.
Explore the source record for details and available documents.
By the middle of 1993, > 30,000 protein sequences has been listed. For 1000 of these, the three-dimensional (tertiary) structure has been experimentally solved. Another 7000 can be modelled by homology. For the remaining 21,000 sequences, secondary structure prediction provides a rough estimate of structural features. Predictions in three states range between 35% (random) and 88% (homology modelling) overall accuracy. Using information about evolutionary conservation as contained in multiple sequence alignments, the secondary structure of 4700 protein sequences was predicted by the automatic e-mail server PHD. For proteins with at least one known homologue, the method has an expected overall three-state accuracy of 71.4% for proteins with at least one known homologue (evaluated on 126 unique protein chains).
In protein engineering and design it is very important that residues can be inspected in their specific environment. A standard relational database system cannot serve this purpose adequately because it cannot handle relations between individual residues. With SCAN3D we introduce a new database system for integrated sequence and structure analysis of proteins. It uses the relational paradigm wherever possible. Its main power, however, stems from the ability to retrieve stretches of consecutive residues with certain properties by comparing a property profile with all stretches of residues in the database, exploiting the ordered character of proteins. In doing so, it bypasses the large number of join operations that would be required by relational database systems. An additional advantage of using property profile matching is that searches can be carried out allowing a pre-set number of mismatches. Also, as the database is read-only, SCAN3D does not need interactive data update mechanisms. Queries typical of a molecular engineering environment are demonstrated with specific examples: analysis of peptides that induce local structure, analysis of site-dependent rotamers and residue--residue contact analysis.
Point mutations are frequently used to explore the structure and/or function of proteins. The ability to predict the structural effects of point mutations would make the planning of such experiments more reliable. We have now derived a set of detailed predictive rules based on the comparison of crystal structures of point mutants and wild types in 83 cases. Despite the surprising simplicity of these rules, they describe well the conformational changes in 85% of all point mutant structures available at present.
Comparison of structures can reveal surprising connections between protein families and provide new insights into the relationship between sequence, structure and function. The solution structure of LexA repressor from Escherichia coli reveals an unexpected structural similarity to a widespread class of prokaryotic and eukaryotic regulatory proteins, which is typified by catabolite gene activator protein (CAP). The use of combined sequence profiles allows the identification of two new prokaryotic members of the superfamily: listeriolysin regulatory protein (PrfA) and ferric uptake regulatory protein (Fur). LexA, PrfA and Fur are the first examples of prokaryotic regulatory proteins in which DNA recognition is mediated by a variant of the classical helix-turn-helix motif, with an insertion in the turn region.
A method has been developed to detect pairs of positions with correlated mutations in protein multiple sequence alignments. The method is based on reconstruction of the phylogenetic tree for a set of sequences and statistical analysis of the distribution of mutations in the branches of the tree. The database of homology-derived protein structures (HSSP) is used as the source of multiple sequence alignments for proteins of known three-dimensional structure. We analyse pairs of positions with correlated mutations in 67 protein families and show quantitatively that the presence of such positions is a typical feature of protein families. A significant but weak tendency is observed for correlated residue pairs to be close in the three-dimensional structure. With further improvements, methods of this type may be useful for the prediction of residue--residue contacts and subsequent prediction of protein structure using distance geometry algorithms. In conclusion, we suggest a new experimental approach to protein structure determination in which selection of functional mutants after random mutagenesis and analysis of correlated mutations provide sufficient proximity constraints for calculation of the protein fold.
We present the prototype of a software system, called GeneQuiz, for large-scale biological sequence analysis. The system was designed to meet the needs that arise in computational sequence analysis and our past experience with the analysis of 171 protein sequences of yeast chromosome III. We explain the cognitive challenges associated with this particular research activity and present our model of the sequence analysis process. The prototype system consists of two parts: (i) the database update and search system (driven by perl programs and rdb, a simple relational database engine also written in perl) and (ii) the visualization and browsing system (developed under C++/ET++). The principal design requirement for the first part was the complete automation of all repetitive actions: database updates, efficient sequence similarity searches and sampling of results in a uniform fashion. The user is then presented with "hit-lists" that summarize the results from heterogeneous database searches. The expert's primary task now simply becomes the further analysis of the candidate entries, where the problem is to extract adequate information about functional characteristics of the query protein rapidly. This second task is tremendously accelerated by a simple combination of the heterogeneous output into uniform relational tables and the provision of browsing mechanisms that give access to database records, sequence entries and alignment views. Indexing of molecular sequence databases provides fast retrieval of individual entries with the use of unique identifiers as well as browsing through databases using pre-existing cross-references. The presentation here covers an overview of the architecture of the system prototype and our experiences on its applicability in sequence analysis.(ABSTRACT TRUNCATED AT 250 WORDS)
HSSP (homology-derived structures of proteins) is a derived database merging structural (2-D and 3-D) and sequence information (1-D). For each protein of known 3D structure from the Protein Data Bank, the database has a file with all sequence homologues, properly aligned to the PDB protein. Homologues are very likely to have the same 3D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of sequence aligned sequence families, but it is also a database of implied secondary and tertiary structures.
FSSP (families of structurally similar proteins) is a database of structural alignments of proteins in the Protein Data Bank (PDB). The database currently contains an extended structural family for each of 330 representative protein chains. Each data set contains structural alignments of one search structure with all other structurally significantly similar proteins in the representative set (remote homologs, < 30% sequence identity), as well as all structures in the Protein Data Bank with 70-30% sequence identity relative to the search structure (medium homologs). Very close homologs (above 70% sequence identity) are excluded as they rarely have marked structural differences. The alignments of remote homologs are the result of pairwise all-against-all structural comparisons in the set of 330 representative protein chains. All such comparisons are based purely on the 3D co-ordinates of the proteins and are derived by automatic (objective) structure comparison programs. The significance of structural similarity is estimated based on statistical criteria. The FSSP database is available electronically from the EMBL file server and by anonymous ftp (file transfer protocol).
The sequence of RNase L has been re-examined by computer analysis. We propose a molecular architecture of RNase L, with an unusual combination, in one protein chain, of 9 ankyrin-like repeats, a functional active protein kinase and a C-terminal catalytic RNase similar to the yeast protein, IRE1. The protein kinase may be involved in a new signal transduction pathway which remains to be discovered.
The sequence of the HIV Nef protein has no significant homology to other proteins in the SwissProt database, and experimental data concerning its function are sparse and contradictory. Using a novel protein sequence comparison method, we find similarities between different Nef sequences and the alpha chain of human MHC class I proteins. The possible biological implications of this finding are discussed.
With a rapidly growing pool of known tertiary structures, the importance of protein structure comparison parallels that of sequence alignment. We have developed a novel algorithm (DALI) for optimal pairwise alignment of protein structures. The three-dimensional co-ordinates of each protein are used to calculate residue-residue (C alpha-C alpha) distance matrices. The distance matrices are first decomposed into elementary contact patterns, e.g. hexapeptide-hexapeptide submatrices. Then, similar contact patterns in the two matrices are paired and combined into larger consistent sets of pairs. A Monte Carlo procedure is used to optimize a similarity score defined in terms of equivalent intramolecular distances. Several alignments are optimized in parallel, leading to simultaneous detection of the best, second-best and so on solutions. The method allows sequence gaps of any length, reversal of chain direction and free topological connectivity of aligned segments. Sequential connectivity can be imposed as an option. The method is fully automatic and identifies structural resemblances and common structural cores accurately and sensitively, even in the presence of geometrical distortions. An all-against-all alignment of over 200 representative protein structures results in an objective classification of known three-dimensional folds in agreement with visual classifications. Unexpected topological similarities of biological interest have been detected, e.g. between the bacterial toxin colicin A and globins, and between the eukaryotic POU-specific DNA-binding domain and the bacterial lambda repressor.
The explosive accumulation of protein sequences in the wake of large-scale sequencing projects is in stark contrast to the much slower experimental determination of protein structures. Improved methods of structure prediction from the gene sequence alone are therefore needed. Here, we report a substantial increase in both the accuracy and quality of secondary-structure predictions, using a neural-network algorithm. The main improvements come from the use of multiple sequence alignments (better overall accuracy), from "balanced training" (better prediction of beta-strands), and from "structure context training" (better prediction of helix and strand lengths). This method, cross-validated on seven different test sets purged of sequence similarity to learning sets, achieves a three-state prediction accuracy of 69.7%, significantly better than previous methods. In addition, the predicted structures have a more realistic distribution of helix and strand segments. The predictions may be suitable for use in practice as a first estimate of the structural type of newly sequenced proteins.
The problem of protein structure prediction is formulated here as that of evaluating how well an amino acid sequence fits a hypothetical structure. The simplest and most complicated approaches, secondary structure prediction and all-atom free energy calculations, can be viewed as sequence-structure fitness problems. Here, an approach of intermediate complexity is described, which involves; (1) description of a protein structure in terms of contact interface vectors, with both intra-protein and protein-solvent contacts counted, (2) derivation of sequence preferences for 2 up to 29 contact interface types, (3) generation of numerous hypothetical model structures by placing the input sequence into a large set of known three-dimensional structures in all possible alignments, (4) evaluation of these models by summing the sequence preferences over all structural positions and (5) choice of predicted three-dimensional structure as that with the best sequence-structure fitness. Evolutionary information is incorporated by using position-dependent core weights derived from multiple sequence alignments. A number of tests of the method are performed: (1) evaluation of cyclic shifts of a sequence in its native structure; (2) alignment of a sequence in its native structure, allowing gaps; (3) alignment search with a sequence or sequence fragment in a database of structures; and (4) alignment search with a structure in a database of sequences. The main results are: (1) a native sequence can very well find its native structure among a large number of alternatives, in correct alignment; (2) substructures, such as (beta alpha)n units, can be detected in spite of very low sequence similarity; (3) remote homologous can be detected, with some dependence on the set of parameters used; (4) contact interface parameters are clearly superior to classical secondary structure parameters; (5) a simple interface description in terms of just two states, protein-protein and protein-water contacts, performs surprisingly well; (6) the use of core weights considerably improves accuracy in detection of remote homologues; (7) based on a sequence database search with a myoglobin contact profile, the C-terminal domain of a viral origin of replication binding protein is predicted to have an all-helical fold. The sequence-structure fitness concept is sufficiently general to accommodate a large variety of protein structure prediction methods, including new models of intermediate complexity currently being developed.
We have trained a two-layered feed-forward neural network on a non-redundant data base of 130 protein chains to predict the secondary structure of water-soluble proteins. A new key aspect is the use of evolutionary information in the form of multiple sequence alignments that are used as input in place of single sequences. The inclusion of protein family information in this form increases the prediction accuracy by six to eight percentage points. A combination of three levels of networks results in an overall three-state accuracy of 70.8% for globular proteins (sustained performance). If four membrane protein chains are included in the evaluation, the overall accuracy drops to 70.2%. The prediction is well balanced between alpha-helix, beta-strand and loop: 65% of the observed strand residues are predicted correctly. The accuracy in predicting the content of three secondary structure types is comparable to that of circular dichroism spectroscopy. The performance accuracy is verified by a sevenfold cross-validation test, and an additional test on 26 recently solved proteins. Of particular practical importance is the definition of a position-specific reliability index. For half of the residues predicted with a high level of reliability the overall accuracy increases to better than 82%. A further strength of the method is the more realistic prediction of segment length. The protein family prediction method is available for testing by academic researchers via an electronic mail server.
Explore the source record for details and available documents.
Iterative profile sequence analysis reveals a remote homology of peroxisomal serine-pyruvate aminotransferases from mammals to the small subunit of soluble hydrogenases from cyanobacteria, an isopenicillin N epimerase, the NifS gene products from bacteria and yeast, and the phosphoserine aminotransferase family. All members of this new class whose function is known are pyridoxal phosphate-dependent enzymes, yet they have distinct catalytic activities. Upon alignment, a lysine around position 200 remains invariant and is predicted to be the pyridoxal phosphate-binding residue. Based on the detected homology, it is predicted that NifS has also a pyridoxal phosphate-dependent serine (or related) aminotransferase function associated with nitrogen economy and/or protection during nitrogen fixation.
The transition of guanine nucleotide binding proteins between the 'on' (GTP-bound) and 'off' (GDP-bound) states has become a paradigm of molecular switching after a chemical reaction. The mechanism by which the switch signal is transmitted to the downstream recipients in the intracellular signal pathway has been extensively studied by biochemical, biophysical and genetic methods, but a clear picture of this process has yet to emerge. Based on the similarities of ras-p21 and elongation factor Tu we propose here a model of the GDP state of ras-p21 that is in agreement with all relevant experimental evidence. The model provides important clues about: (1) a possible molecular mechanism for signal transmission from the site of GTP hydrolysis to downstream effectors; (2) a major conformational change during signal generation and a key residue involved in this process (Tyr-64); and (3) regions in ras-p21 that can be differentially recognized by binding to external partners in a GTP/GDP state dependent fashion, most notably residues D69, Q70, R73, T74, R102, K104, D105 at the end of the alpha-helices 2 and 3.