Search PubMed⌕ Search

Biomedical subjects

C Sander

Publications and source records attributed to C Sander.

At least 109 records · Page 6Linked to original sources

The immunoglobulin fold. Structural classification, sequence patterns and common core.

Since the first crystal structure of an immunoglobulin revealed a modular architecture, the characteristic beta-sheet fold of the immunoglobulin domain has been found in many other proteins of diverse biological function. Here, a systematic comparison of 23 Ig domain structures with less than 25% pairwise residue identity was performed using automatic structural alignment and analysis of beta-sheet and loop topology. Sequence consensus patterns were identified for nine distinct families with at most marginal similarity to each other. The analysis reveals a common structural core of only four beta-strands (b, c, e and f), embedded in an antiparallel curled beta-sheet sandwich with a total of three to five additional strands (a, c', c'', d, g) and a characteristic intersheet angle. The variation in the position of the edge strands (a, c', c'', d and g) relative to the common core defines four different topological subtypes that correlate with the length of the intervening sequence between strands c and e, the most variable region in sequence. The switch of strand c' from one sheet to the other in seven-stranded domains appears to result from short c-e segments, rather than being a major structural discriminator. The high degree of structural flexibility outside the common core and the extreme variability of side-chain packing inside the core do not support a protein folding pathway common to all members of the structural class. Mutation rates of immunoglobulin-like domains in different proteins vary considerably. Disulfide bridges, thought to contribute to structural stability, are not necessarily invariant in number and location within a subclass.

Amino Acid Sequence↗

A novel RNA-binding motif in omnipotent suppressors of translation termination, ribosomal proteins and a ribosome modification enzyme?

Using computer methods for database search, multiple alignment, protein sequence motif analysis and secondary structure prediction, a putative new RNA-binding motif was identified. The novel motif is conserved in yeast omnipotent translation termination suppressor SUP1, the related DOM34 protein and its pseudogene homologue; three groups of eukaryotic and archaeal ribosomal proteins, namely L30e, L7Ae/S6e and S12e; an uncharacterized Bacillus subtilis protein related to the L7A/S6e group; and Escherichia coli ribosomal protein modification enzyme RimK. We hypothesize that a new type of RNA-binding domain may be utilized to deliver additional activities to the ribosome.

Amino Acid Sequence↗

The chaperone function of DnaK requires the coupling of ATPase activity with substrate binding through residue E171.

Central to the chaperone function of Hsp70 stress proteins including Escherichia coli DnaK is the ability of Hsp70 to bind unfolded protein substrates in an ATP-dependent manner. Mg2+/ATP dissociates bound substrates and, furthermore, substrate binding stimulates the ATPase of Hsp70. This coupling is proposed to require a glutamate residue, E175 of bovine Hsc70, that is entirely conserved within the Hsp70 family, as it contacts bound Mg2+/ATP and is part of a hinge required for a postulated ATP-dependent opening/closing movement of the nucleotide binding cleft which then triggers substrate release. We analyzed the effects of dnaK mutations which alter the corresponding glutamate-171 of DnaK to alanine, leucine or lysine. In vivo, the mutated dnaK alleles failed to complement the delta dnaK52 mutation and were dominant negative in dnaK+ cells. In vitro, all three mutant DnaK proteins were inactive in known DnaK-dependent reactions, including refolding of denatured luciferase and initiation of lambda DNA replication. The mutant proteins retained ATPase activity, as well as the capacity to bind peptide substrates. The intrinsic ATPase activities of the mutant proteins, however, did exhibit increased Km and Vmax values. More importantly, these mutant proteins showed no stimulation of ATPase activity by substrates and no substrate dissociation by Mg2+/ATP. Thus, glutamate-171 is required for coupling of ATPase activity with substrate binding, and this coupling is essential for the chaperone function of DnaK.

Adenosine Triphosphatases↗

Structural similarity of plant chitinase and lysozymes from animals and phage. An evolutionary connection.

A search in the database of known three-dimensional protein structures with the structure of a plant endochitinase revealed a subtle but unambiguous similarity to lysozymes from animals and phages. An evolutionary connection between plant endochitinases and lysozymes is supported by similar overall topology of fold, overlapping substrate specificities and remarkable conservation of some sequence and architectural detail around the active site. Much of the knowledge about lysozyme can now be extended by analogy to endochitinase. New insights into the mechanism of endochitinase are expected to stimulate genetic engineering studies into plant defense mechanisms against pests and pathogens.

Amino Acid Sequence↗

Yeast chromosome III: new gene functions.

One year after the release of the sequence of yeast chromosome III, we have re-examined its open reading frames (ORFs) by computer methods. More than 61% of the 171 probable gene products have significant sequence similarities in the current databases; as many as 54% have already known functions or are related to functionally characterized proteins, allowing partial prediction of protein function, 11 percentage points more than reported a year ago; 19% are similar to proteins of known three-dimensional structure, allowing model building by homology. The most interesting new identifications include a sugar kinase distantly related to ribokinases, a phosphatidyl serine synthetase, a putative transcription regulator, a flavodoxin-like protein, and a zinc finger protein belonging to a distinct subfamily. Several ORFs have similarities to uncharacterized proteins, resulting in new families in search of a function'. About 54% of ORFs match sequences from other phyla, including numerous fragments in the database of expressed sequence tags (ESTs). Most significant similarities to ESTs are with proteins in conserved families widely represented in the databases. About 30% of ORFs contain one or more predicted transmembrane segments. The increase in the power of functional and structural prediction comes from improvements in sequence analysis and from richer databases and is expected to facilitate substantially the experimental effort in characterizing the function of new gene products.

Amino Acid Sequence↗

Redefining the goals of protein secondary structure prediction.

Secondary structure prediction recently has surpassed the 70% level of average accuracy, evaluated on the single residue states helix, strand and loop (Q3). But the ultimate goal is reliable prediction of tertiary (three-dimensional, 3D) structure, not 100% single residue accuracy for secondary structure. A comparison of pairs of structurally homologous proteins with divergent sequences reveals that considerable variation in the position and length of secondary structure segments can be accommodated within the same 3D fold. It is therefore sufficient to predict the approximate location of helix, strand, turn and loop segments, provided they are compatible with the formation of 3D structure. Accordingly, we define here a measure of segment overlap (Sov) that is somewhat insensitive to small variations in secondary structure assignments. The new segment overlap measure ranges from an ignorance level of 37% (random protein pairs) via a current level of 72% for a prediction method based on sequence profile input to neural networks (PHD) to an average 90% level for homologous protein pairs. We conclude that the highest scores one can reasonably expect for secondary structure prediction are a single residue accuracy of Q3 > 85% and a fractional segment overlap of Sov > 90%.

Amino Acid Sequence↗

Enlarged representative set of protein structures.

To reduce redundancy in the Protein Data Bank of 3D protein structures, which is caused by many homologous proteins in the data bank, we have selected a representative set of structures. The selection algorithm was designed to (1) select as many nonhomologous structures as possible, and (2) to select structures of good quality. The representative set may reduce time and effort in statistical analyses.

Amino Acid Sequence↗

Correlated mutations and residue contacts in proteins.

The maintenance of protein function and structure constrains the evolution of amino acid sequences. This fact can be exploited to interpret correlated mutations observed in a sequence family as an indication of probable physical contact in three dimensions. Here we present a simple and general method to analyze correlations in mutational behavior between different positions in a multiple sequence alignment. We then use these correlations to predict contact maps for each of 11 protein families and compare the result with the contacts determined by crystallography. For the most strongly correlated residue pairs predicted to be in contact, the prediction accuracy ranges from 37 to 68% and the improvement ratio relative to a random prediction from 1.4 to 5.1. Predicted contact maps can be used as input for the calculation of protein tertiary structure, either from sequence information alone or in combination with experimental information.

Amino Acid Sequence↗

Combining evolutionary information and neural networks to predict protein secondary structure.

Using evolutionary information contained in multiple sequence alignments as input to neural networks, secondary structure can be predicted at significantly increased accuracy. Here, we extend our previous three-level system of neural networks by using additional input information derived from multiple alignments. Using a position-specific conservation weight as part of the input increases performance. Using the number of insertions and deletions reduces the tendency for overprediction and increases overall accuracy. Addition of the global amino acid content yields a further improvement, mainly in predicting structural class. The final network system has sustained overall accuracy of 71.6% in a multiple cross-validation test on 126 unique protein chains. A test on a new set of 124 recently solved protein structures that have no significant sequence similarity to the learning set confirms the high level of accuracy. The average cross-validated accuracy for all 250 sequence-unique chains is above 72%. Using various data sets, the method is compared to alternative prediction methods, some of which also use multiple alignments: the performance advantage of the network system is at least 6 percentage points in three-state accuracy. In addition, the network estimates secondary structure content from multiple sequence alignments about as well as circular dichroism spectroscopy on a single protein and classifies 75% of the 250 proteins correctly into one of four protein structural classes. Of particular practical importance is the definition of a position-specific reliability index. For 40% of all residues the method has a sustained three-state accuracy of 88%, as high as the overall average for homology modelling. A further strength of the method is greatly increased accuracy in predicting the placement of secondary structure segments.

Amino Acid Sequence↗

Searching protein structure databases has come of age.

The number of protein structures known in atomic detail has increased from one in 1960 (Kendrew, J.C., Strandberg, B.E., Hart, R.G., Davies, D.R., Phillips, D.C., Shore, V.C. Nature (London) 185:422-427, 1960) to more than 1000 in 1994. The rate at which new structures are being published exceeds one a day as a result of recent advances in protein engineering, crystallography, and spectroscopy. More and more frequently, a newly determined structure is similar in fold to a known one, even when no sequence similarity is detectable. A new generation of computer algorithms has now been developed that allows routine comparison of a protein structure with the database of all known structures. Such structure database searches are already used daily and they are beginning to rival sequence database searches as a tool for discovering biologically interesting relationships.

Algorithms↗

Parser for protein folding units.

General patterns of protein structural organization have emerged from studies of hundreds of structures elucidated by X-ray crystallography and nuclear magnetic resonance. Structural units are commonly identified by visual inspection of molecular models using qualitative criteria. Here, we propose an algorithm for identification of structural units by objective, quantitative criteria based on atomic interactions. The underlying physical concept is maximal interactions within each unit and minimal interaction between units (domains). In a simple harmonic approximation, interdomain dynamics is determined by the strength of the interface and the distribution of masses. The most likely domain decomposition involves units with the most correlated motion, or largest interdomain fluctuation time. The decomposition of a convoluted 3-D structure is complicated by the possibility that the chain can cross over several times between units. Grouping the residues by solving an eigenvalue problem for the contact matrix reduces the problem to a one-dimensional search for all reasonable trial bisections. Recursive bisection yields a tree of putative folding units. Simple physical criteria are used to identify units that could exist by themselves. The units so defined closely correspond to crystallographers' notion of structural domains. The results are useful for the analysis of folding principles, for modular protein design and for protein engineering.

Actins↗

Conservation and prediction of solvent accessibility in protein families.

Currently, the prediction of three-dimensional (3D) protein structure from sequence alone is an exceedingly difficult task. As an intermediate step, a much simpler task has been pursued extensively: predicting 1D strings of secondary structure. Here, we present an analysis of another 1D projection from 3D structure: the relative solvent accessibility of each residue. We show that solvent accessibility is less conserved in 3D homologues than is secondary structure, and hence is predicted less accurately from automatic homology modeling; the correlation coefficient of relative solvent accessibility between 3D homologues is only 0.77, and the average accuracy of predictions based on sequence alignments is only 0.68. The latter number provides an effective upper limit on the accuracy of predicting accessibility from sequence when homology modeling is not possible. We introduce a neural network system that predicts relative solvent accessibility (projected onto ten discrete states) using evolutionary profiles of amino acid substitutions derived from multiple sequence alignments. Evaluated in a cross-validation test on 238 unique proteins, the correlation between predicted and observed relative accessibility is 0.54. Interpreted in terms of a three-state (buried, intermediate, exposed) description of relative accessibility, the fraction of correctly predicted residue states is about 58%. In absolute terms this accuracy appears poor, but given the relatively low conservation of accessibility in 3D families, the network system is not far from its likely optimal performance. The most reliably predicted fraction of the residues (50%) is predicted as accurately as by automatic homology modeling. Prediction is best for buried residues, e.g., 86% of the completely buried sites are correctly predicted as having 0% relative accessibility.

Biological Evolution↗

Amino acid analysis and protein database compositional search as a rapid and inexpensive method to identify proteins.

The identification of protein samples in minute quantities of protein samples, e.g., from two-dimensional polyacrylamide gel electrophoresis analysis, is an everyday problem in biology laboratories. Here we show that computer-assisted amino acid analysis can fulfill this task. Amino acid analysis data can be used to compare the amino acid composition of an unknown protein with protein compositions in a database (compositional search). Routine amino acid analysis data can, despite a certain margin of error, be used to identify a protein. Compared to protein sequencing, amino analysis is much cheaper, faster, and allows higher sample throughput. Thus, the method may replace protein sequencing as a first attempt in identification, provided a homolog can be found in the database.

Amino Acid Sequence↗

A human cDNA coding for the Leydig insulin-like peptide (Ley I-L).

cDNA clones for the human Leydig insulin-like peptide (Ley I-L) have been isolated and characterized. The nucleotide sequence of the 743-bp cDNA includes an incomplete 7-bp 5'-noncoding region, an open reading frame of 393 bp, and a 343-bp 3'-noncoding region. By primer extension analysis, the transcription start site was determined as being 14-bp upstream of the translation start site. The underlying gene is expressed in the testis but not in other organs. From the cDNA sequence, it can be deduced that the Ley I-L protein is synthesized as a 131-amino-acid (aa) preproprotein and that it contains a 24-aa signal peptide. Comparison of the pro Ley I-L protein with members of the insulin-like hormone superfamily predicts that the biologically active hormone, after proteolytic processing of the C peptide, consists of a 31-aa long B chain and a 26-aa long A chain, and that it has a molecular weight of 6.25 kDa.

Amino Acid Sequence↗

Design of protein structures: helix bundles and beyond.

The design of proteins or peptides with novel functions can be achieved either by modifying existing molecules or by inventing entirely new structures and sequences that are unknown in nature. Combinatorial-design strategies have led to the first de novo proteins, but these still lack some of the desired attributes. The most promising practical strategies for developing proteins with useful biological or chemical function combine theoretical design with experimental screening or selection systems.

Amino Acid Sequence↗

Structure prediction of proteins--where are we now?

Although the 'structure from sequence' prediction problem remains fundamentally unsolved, new and promising methods in one, two and three dimensions have reopened the field. Significantly improved one-dimensional prediction of secondary structure from multiple sequence alignments is now in routine use. In the two-dimensional approach, inter-residue contacts can be detected by analysis of correlated mutations, albeit with low accuracy. Finally, three-dimensional methods, in which pseudopotentials or information values are derived from the databases, are proving their value for distinguishing between correct and incorrect models.

Amino Acid Sequence↗