Search PubMed⌕ Search

Biomedical subjects

C Sander

Publications and source records attributed to C Sander.

At least 163 records · Page 9Linked to original sources

The internalization signal in the cytoplasmic tail of lysosomal acid phosphatase consists of the hexapeptide PGYRHV.

Lysosomal acid phosphatase (LAP) is rapidly internalized from the cell surface due to a tyrosine-containing internalization signal in its 19 amino acid cytoplasmic tail. Measuring the internalization of a series of LAP cytoplasmic tail truncation and substitution mutants revealed that the N-terminal 12 amino acids of the cytoplasmic tail are sufficient for rapid endocytosis and that the hexapeptide 411-PGYRHV-416 is the tyrosine-containing internalization signal. Truncation and substitution mutants of amino acid residues following Val416 can prevent internalization even though these residues do not belong to the internalization signal. It was shown recently that part of the LAP cytoplasmic tail peptide corresponding to 410-PPGY-413 forms a well-ordered beta turn structure in solution. Two-dimensional NMR spectroscopy of two modified LAP tail peptides, in which the single tyrosine was substituted either by phenylalanine or by alanine, revealed that the tendency to form a beta turn is reduced by 25% in the phenylalanine-containing peptide and by approximately 50% in the alanine-containing mutant peptide. Our results suggest, that in the short cytoplasmic tail of LAP tyrosine is required for stabilization of the right turn and that the aromatic ring system of the tyrosine residue is a contact point to the putative cytoplasmic receptor.

Acid Phosphatase↗

Selection of representative protein data sets.

The Protein Data Bank currently contains about 600 data sets of three-dimensional protein coordinates determined by X-ray crystallography or NMR. There is considerable redundancy in the data base, as many protein pairs are identical or very similar in sequence. However, statistical analyses of protein sequence-structure relations require nonredundant data. We have developed two algorithms to extract from the data base representative sets of protein chains with maximum coverage and minimum redundancy. The first algorithm focuses on optimizing a particular property of the selected proteins and works by successive selection of proteins from an ordered list and exclusion of all neighbors of each selected protein. The other algorithm aims at maximizing the size of the selected set and works by successive thinning out of clusters of similar proteins. Both algorithms are generally applicable to other data bases in which criteria of similarity can be defined and relate to problems in graph theory. The largest nonredundant set extracted from the current release of the Protein Data Bank has 155 protein chains. In this set, no two proteins have sequence similarity higher than a certain cutoff (30% identical residues for aligned subsequences longer than 80 residues), yet all structurally unique protein families are represented. Periodically updated lists of representative data sets are available by electronic mail from the file server "netserv@embl-heidelberg.de." The selection may be useful in statistical approaches to protein folding as well as in the analysis and documentation of the known spectrum of three-dimensional protein structures.

Algorithms↗

Comprehensive sequence analysis of the 182 predicted open reading frames of yeast chromosome III.

With the completion of the first phase of the European yeast genome sequencing project, the complete DNA sequence of chromosome III of Saccharomyces cerevisiae has become available (Oliver, S. G., et al., 1992, Nature 357, 38-46). We have tested the predictive power of computer sequence analysis of the 176 probable protein products of this chromosome, after exclusion of six problem cases. When the results of database similarity searches are pooled with prior knowledge, a likely function can be assigned to 42% of the proteins, and a predicted three-dimensional structure to a third of these (14% of the total). The function of the remaining 58% remains to be determined. Of these, about one-third have one or more probable transmembrane segments. Among the most interesting proteins with predicted functions are a new member of the type X polymerase family, a transcription factor with an N-terminal DNA-binding domain related to GAL4, a "fork head" DNA-binding domain previously known only in Drosophila and in mammals, and a putative methyltransferase. Our analysis increased the number of known significant sequence similarities on chromosome III by 13, to now 67. Although the near 40% success rate of identifying unknown protein function by sequence analysis is surprisingly high, the information gap between known protein sequences and unknown function is expected to widen and become a major bottleneck of genome projects in the near future. Based on the experience gained in this test study, we suggest that the development of an automated computer workbench for protein sequence analysis must be an important item in genome projects.

Acetolactate Synthase↗

A database of protein structure families with common folding motifs.

The availability of fast and robust algorithms for protein structure comparison provides an opportunity to produce a database of three-dimensional comparisons, called families of structurally similar proteins (FSSP). The database currently contains an extended structural family for each of 154 representative (below 30% sequence identity) protein chains. Each data set contains: the search structure; all its relatives with 70-30% sequence identity, aligned structurally; and all other proteins from the representative set that contain substructures significantly similar to the search structure. Very close relatives (above 70% sequence identity) rarely have significant structural differences and are excluded. The alignments of remote relatives are the result of pairwise all-against-all structural comparisons in the set of 154 representative protein chains. The comparisons were carried out with each of three novel automatic algorithms that cover different aspects of protein structure similarity. The user of the database has the choice between strict rigid-body comparisons and comparisons that take into account interdomain motion or geometrical distortions; and, between comparisons that require strictly sequential ordering of segments and comparisons, which allow altered topology of loop connections or chain reversals. The data sets report the structurally equivalent residues in the form of a multiple alignment and as a list of matching fragments to facilitate inspection by three-dimensional graphics. If substructures are ignored, the result is a database of structure alignments of full-length proteins, including those in the twilight zone of sequence similarity.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

Protein design on computers. Five new proteins: Shpilka, Grendel, Fingerclasp, Leather, and Aida.

What is the current state of the art in protein design? This question was approached in a recent two-week protein design workshop sponsored by EMBO and held at the EMBL in Heidelberg. The goals were to test available design tools and to explore new design strategies. Five novel proteins were designed: Shpilka, a sandwich of two four-stranded beta-sheets, a scaffold on which to explore variations in loop topology; Grendel, a four-helical membrane anchor, ready for fusion to water-soluble functional domains; Finger-clasp, a dimer of interdigitating beta-beta-alpha units, the simplest variant of the "handshake" structural class; Aida, an antibody binding surface intended to be specific for flavodoxin; Leather--a minimal NAD binding domain, extracted from a larger protein. Each design is available as a set of three-dimensional coordinates, the corresponding amino acid sequence and a set of analytical results. The designs are placed in the public domain for scrutiny, improvement, and possible experimental verification.

Algorithms↗

Fast and simple Monte Carlo algorithm for side chain optimization in proteins: application to model building by homology.

An unknown protein structure can be predicted with fair accuracy once an evolutionary connection at the sequence level has been made to a protein of known 3-D structure. In model building by homology, one typically starts with a backbone framework, rebuilds new loop regions, and replaces nonconserved side chains. Here, we use an extremely efficient Monte Carlo algorithm in rotamer space with simulated annealing and simple potential energy functions to optimize the packing of side chains on given backbone models. Optimized models are generated within minutes on a workstation, with reasonable accuracy (average of 81% side chain chi 1 dihedral angles correct in the cores of proteins determined at better than 2.5 A resolution). As expected, the quality of the models decreases with decreasing accuracy of backbone coordinates. If the back-bone was taken from a homologous rather than the same protein, about 70% side chain chi 1 angles were modeled correctly in the core in a case of strong homology and about 60% in a case of medium homology. The algorithm can be used in automated, fast, and reproducible model building by homology.

Algorithms↗

GCI: a network server for interactive 3D graphics.

The Graphics Command Interpreter (GCI) is an independent server module that can be interfaced to any program that needs interactive three-dimensional (3D) graphics capabilities. The principal advantage of GCI is its simplicity. Only a limited set of powerful features have been implemented, including object management, global and local transformations, rotation, translation, clipping, scaling, viewport operations, window management, menu handling and picking. GCI and the master (client) program it serves run concurrently, communicating over a local or remote TCP/IP network. GCI sets up socket communication and provides a 3D graphics window and a terminal emulator for the master program. Communication between the two programs is via ASCII strings over standard I/O channels. The implied language for messages is very simple. GCI interprets messages from the master program and implements them as changes of graphical objects or as text messages to the user. GCI provides the user with facilities to manipulate the view of the displayed 3D objects interactively, independently of the master program, and to communicate mouse-controlled selection of menu items or 3D points as well as keyboard strings to the master program. The program is written in C and initially implemented using the Silicon Graphics GL graphics library. As the need to link special libraries to the master program is completely avoided, GCI can very easily be interfaced to existing programs written in any language and running on any operating system capable of TCP/IP communication. The program is freely available.

Computer Communication Networks↗

The essential tyrosine of the internalization signal in lysosomal acid phosphatase is part of a beta turn.

For rapid endocytosis lysosomal acid phosphatase requires a Tyr-containing signal in its cytoplasmic domain, as do cell surface receptors mediating endocytosis and clustering in coated pits. To determine the structure of the internalization signal an 18 amino acid peptide representing the cytoplasmic tail of lysosomal acid phosphatase was analyzed by two-dimensional nuclear magnetic resonance spectroscopy. Part of the peptide, 5-PPGY-8, forms a well-ordered beta turn of type I in solution. Our result and data on the structure of the endocytosis signal of the low density lipoprotein receptor reported by Bansal and Gierasch in the accompanying paper represent experimental determinations of the three-dimensional structure of protein transport signals and suggest that the essential aromatic amino acid of internalization signals is recognized by a putative cytoplasmic receptor in the structural context of a tight turn.

Acid Phosphatase↗

GTPase domains of ras p21 oncogene protein and elongation factor Tu: analysis of three-dimensional structures, sequence families, and functional sites.

GTPase domains are functional and structural units employed as molecular switches in a variety of important cellular functions, such as growth control, protein biosynthesis, and membrane traffic. Amino acid sequences of more than 100 members of different subfamilies are known, but crystal structures of only mammalian ras p21 and bacterial elongation factor Tu have been determined. After optimal superposition of these remarkably similar structures, careful multiple sequence alignment, and calculation of residue-residue interactions, we analyzed the two subfamilies in terms of structural conservation, sequence conservation, and residue contact strength. There are three main results. (i) A structure-based alignment of p21 and elongation factor Tu. (ii) The definition of a common conserved structural core that may be useful as the basis of model building by homology of the three-dimensional structure of any GTPase domain. (iii) Identification of sequence regions, other than the effector loop and the nucleotide binding site, that may be involved in the functional cycle: they are loop L4, known to change conformation after GTP hydrolysis; helix alpha 2, especially Arg-73 and Met-67 in ras p21; loops L8 and L10, including ras p21 Arg-123, Lys-147, and Leu-120; and residues located spatially near the N and C termini. These regions are candidate sites for interaction either with the GTP/GDP exchange factor, with a GTPase-affected function, or with a molecule delivered to a destination site with the aid of the GTPase domain.

Amino Acid Sequence↗

Identification by computer sequence analysis of transcriptional regulator proteins in Dictyostelium discoideum and Serratia marcescens.

We have performed computer searches in the database of known protein sequences for proteins similar in sequence to bacteriophage regulatory proteins of known 3-D structure. The searches are more selective than other methods due to the use of a length-dependent threshold in sequence similarity, above which structural homology is implied with high certainty. Two probable DNA binding proteins were identified which are predicted to have a three-dimensional structure very similar to bacteriophage cro and repressor proteins. Approximate three-dimensional model coordinates are available from the authors. Both proteins contain the helix-turn-helix sequence motif typical of a wide class of DNA binding proteins and their function is deduced by analogy to sequence-similar proteins of known function. We predict that the Y.Smal protein in the restriction-modification enzyme gene locus of the enterobacterium serratia marcescens is a regulator of endonuclease expression; and, that the vegetative specific gene VSH7 of the slime mold dictyostelium discoideum codes for a regulator of gene expression specific for the slime mold growth phase before the onset of the developmental program. Point mutations that would have a strong effect on growth regulation phenotype are suggested. The VSH7 protein would be the first eukaryotic representative of the cro/phage repressor class.

Amino Acid Sequence↗

Database algorithm for generating protein backbone and side-chain co-ordinates from a C alpha trace application to model building and detection of co-ordinate errors.

The problem of constructing all-atom model co-ordinates of a protein from an outline of the polypeptide chain is encountered in protein structure determination by crystallography or nuclear magnetic resonance spectroscopy, in model building by homology and in protein design. Here, we present an automatic procedure for generating full protein co-ordinates (backbone and, optionally, side-chains) given the C alpha trace and amino acid sequence. To construct backbones, a protein structure database is first scanned for fragments that locally fit the chain trace according to distance criteria. A best path algorithm then sifts through these segments and selects an optimal path with minimal mismatch at fragment joints. In blind tests, using fully known protein structures, backbones (C alpha, C, N, O) can be reconstructed with a reliability of 0.4 to 0.6 A root-mean-square position deviation and not more than 0 to 5% peptide flips. This accuracy is sufficient to identify possible errors in protein co-ordinate sets. To construct full co-ordinates, side-chains are added from a library of frequently occurring rotamers using a simple and fast Monte Carlo procedure with simulated annealing. In tests on X-ray structures determined at better than 2.5 A resolution, the positions of side-chain atoms in the protein core (less than 20% relative accessibility) have an accuracy of 1.6 A (r.m.s. deviation) and 70% of chi 1 angles are within 30 degrees of the X-ray structure. The computer program MaxSprout is available on request.

Algorithms↗

A structure-derived sequence pattern for the detection of type I copper binding domains in distantly related proteins.

A structure-based approach to the definition of sequence patterns characteristic of protein domains is presented by example. The approach requires a multiple sequence alignment of a family (or set of related families) as well as at least one three-dimensional structure. The pattern derived does not merely summarize the information in the known sequences but attempts to generalize the pattern specifications based on structural insight. In this example, the pattern-driven database search identified correctly most of the known type I copper-binding domains and detected the presence of a homologous domain in a previously unknown case (CopA protein). The significance of these results is discussed.

Amino Acid Sequence↗

Database of homology-derived protein structures and the structural meaning of sequence alignment.

The database of known protein three-dimensional structures can be significantly increased by the use of sequence homology, based on the following observations. (1) The database of known sequences, currently at more than 12,000 proteins, is two orders of magnitude larger than the database of known structures. (2) The currently most powerful method of predicting protein structures is model building by homology. (3) Structural homology can be inferred from the level of sequence similarity. (4) The threshold of sequence similarity sufficient for structural homology depends strongly on the length of the alignment. Here, we first quantify the relation between sequence similarity, structure similarity, and alignment length by an exhaustive survey of alignments between proteins of known structure and report a homology threshold curve as a function of alignment length. We then produce a database of homology-derived secondary structure of proteins (HSSP) by aligning to each protein of known structure all sequences deemed homologous on the basis of the threshold curve. For each known protein structure, the derived database contains the aligned sequences, secondary structure, sequence variability, and sequence profile. Tertiary structures of the aligned sequences are implied, but not modeled explicitly. The database effectively increases the number of known protein structures by a factor of five to more than 1800. The results may be useful in assessing the structural significance of matches in sequence database searches, in deriving preferences and patterns for structure prediction, in elucidating the structural role of conserved residues, and in modeling three-dimensional detail by homology.

Amino Acid Sequence↗

Detection of common three-dimensional substructures in proteins.

We present a fully automatic algorithm for three-dimensional alignment of protein structures and for the detection of common substructures and structural repeats. Given two proteins, the algorithm first identifies all pairs of structurally similar fragments and subsequently clusters into larger units pairs of fragments that are compatible in three dimensions. The detection of similar substructures is independent of insertion/deletion penalties and can be chosen to be independent of the topology of loop connections and to allow for reversal of chain direction. Using distance geometry filters and other approximations, the algorithm, implemented in the WHAT IF program, is so fast that structural comparison of a single protein with the entire database of known protein structures can be performed routinely on a workstation. The method reproduces known non-trivial superpositions such as plastocyanin on azurin. In addition, we report surprising structural similarity between ubiquitin and a (2Fe-2S) ferredoxin.

Algorithms↗

The structure of ColE1 rop in solution.

The structure of the ColE1 repressor of primer (rop) protein in solution was determined from the proton nuclear magnetic resonance data by a combined use of distance geometry and restrained molecular dynamics calculations. A set of structures was determined with low internal energy and virtually no violations of the experimental distance restraints. Rop forms homodimers: Two helical hairpins are arranged as an antiparallel four helix bundle with a left-handed rope-like twist of the helix axes and with left-handed bundle topology. The very compact packing of the side chains in the helix interfaces of the rop coiled-coil structure may well account for its high stability. Overall, the solution structure is highly similar to the recently determined X-ray structure (Banner, D.W., Kokkinidis, M. and Tsernoglou, D. (1987) J. Mol. Biol., 196, 657-675), although there are minor differences in regions where packing forces appear to influence the crystal structure.

Bacterial Proteins↗