Search PubMed⌕ Search

Biomedical subjects

C Sander

Publications and source records attributed to C Sander.

At least 73 records · Page 4Linked to original sources

Sequencing and analysis of a 35.4 kb region on the right [corrected] arm of chromosome IV from Saccharomyces cerevisiae reveal 23 open reading frames.

The complete DNA sequence of cosmid clone 31A5 containing a 35 452 bp segment from the right [corrected] arm of chromosome IV from Saccharomyces cerevisiae, was determined from an ordered set of subclones in combination with primer walking on the cosmid. The sequence contains 23 open reading frames (ORFs) of more than 100 amino acid residues and the tRNA-Va12a gene. Five ORFs corresponded to the known yeast genes SNQ2, SES1, GCV1, RPL2B and RPS18A. The DNA sequence for RPS18A is interrupted by an intron. One ORF corresponded to a part of the yeast gene HEX2 at the end of the cosmid insert. Four ORFs encoded putative proteins which showed strong homologies to other previously known proteins, three of yeast origin and one of non-yeast origin. Two ORFs were classified as having borderline homologies: one had similarity to two protein families and another to two protein products of unknown function from other species. The remaining 11 ORFs bore no significant similarity to any published protein.

Amino Acid Sequence↗

Positioning hydrogen atoms by optimizing hydrogen-bond networks in protein structures.

A method is presented that positions polar hydrogen atoms in protein structures by optimizing the total hydrogen bond energy. For this goal, an empirical hydrogen bond force field was derived from small molecule crystal structures. Bifurcated hydrogen bonds are taken into account. The procedure also predicts ionization states of His, Asp, and Glu residues. During optimization, side-chain conformations of His, Gln, and Asn residues are allowed to change their last chi angle by 180 degrees to compensate for crystallographic misassignments. Crystal structure symmetry is taken into account where appropriate. The results can have significant implications for molecular dynamics simulations, protein engineering, and docking studies. The largest impact, however, is in protein structure verification: over 85% of protein structures tested can be improved by using our procedure.

Amino Acids↗

Computational comparisons of model genomes.

Complete genomes from model organisms provide new challenges for computational molecular biology. Novel questions emerge from the genome data obtained from the functional prediction of thousands of gene products. In this review, we present some approaches to the computational comparison of genomes, based on sequence and text analysis, and comparisons of genome composition and gene order.

Biotechnology↗

A comparison of structural and dynamic properties of different simulation methods applied to SH3.

The dynamic and static properties of molecular dynamics simulations using various methods for treating solvent were compared. The SH3 protein domain was chosen as a test case because of its small size and high surface-to-volume ratio. The simulations were analyzed in structural terms by examining crystal packing, distribution of polar residues, and conservation of secondary structure. In addition, the "essential dynamics" method was applied to compare each of the molecular dynamics trajectories with a full solvent simulation. This method proved to be a powerful tool for the comparison of large concerted atomic motions in SH3. It identified methods of simulation that yielded significantly different dynamic properties compared to the full solvent simulation. Simulating SH3 using the stochastic dynamics algorithm with a vacuum (reduced charge) force field produced properties close to those of the full solvent simulation. The application of a recently described solvation term did not improve the dynamic properties. The large concerted atomic motions in the full solvent simulation as revealed by the essential dynamics method were analyzed for possible biological implications. Two loops, which have been shown to be involved in ligand binding, were seen to move in concert to open and close the ligand-binding site.

Algorithms↗

The PDBFINDER database: a summary of PDB, DSSP and HSSP information with added value.

MOTIVATION: The Protein Data Bank currently contains more than 4700 protein coordinate sets. It is often desirable to make a selection from these files based on a criterion like R-factor, experimental method, length of the amino acid sequence, or the number of homologous sequences in SWISSPROT. Doing this using the distributed form of the Protein Data Bank can be a tedious task, because (1) this requires reading one file for every single entry, and (2) not all of the information is present in a consistent computer readable way in all of the entries. RESULTS: The PDBFINDER database provides an easy to interpret file containing summary information about all Protein Data Bank files. Summary information from the DSSP (Definition of Secondary Structure of Proteins) and HSSP (Homology derived Secondary Structure of Proteins) databases is also included. Furthermore, where essential data were missing from the Protein Data Bank file, this information has been retrieved from the original literature. AVAILABILITY: The latest version of the PDBFINDER database can be downloaded by anonymous ftp from swift.embl-heidelberg.de, directory:/pdbfinder. CONTACT: E-mail address hooft@embl-heidelberg.de.

Amino Acid Sequence↗

The prediction of protein contacts from multiple sequence alignments.

We have studied the question of how much extra predictive power the correlated mutational behaviour of pairs of amino acid residues separated along a sequence has concerning the likelihood of those residues being in contact in the folded protein. The mutational behaviour is deduced from multiple sequence alignments. Our findings are that there is, indeed, some valuable information available from this source and that it is sufficient to make a significant improvement in our ability to predict contacts, when compared with earlier methods that do not take into account the correlations between the mutations. This improvement is approximately twice as large as can be obtained by the more economical method of simply averaging pair preferences over the same sequence alignment. Even when using a method based on pair preferences, a further significant improvement can be made by penalizing more variable regions (on the reasonable assumption that invariant residues are relatively more likely to be in contact), though we have found no way of improving the pair preference method to the extent that it matches the method based on correlated behaviour. Our new method is thought to be the best data-based method of contact prediction developed so far, achieving, on average, an improvement over a random (i.e. information-free) prediction of a factor of five when the number of contacts predicted is chosen to match the number that actually occur.

Algorithms↗

Bridging the protein sequence-structure gap by structure predictions.

The problem of accurately predicting protein three-dimensional structure from sequence has yet to be solved. Recently, several new and promising methods that work in one, two, or three dimensions have invigorated the field. Modeling by homology can yield fairly accurate three-dimensional structures for approximately 25% of the currently known protein sequences. Techniques for cooperatively fitting sequences into known three-dimensional folds, called threading methods, can increase this rate by detecting very remote homologies in favorable cases. Prediction of protein structure in two dimensions, i.e. prediction of interresidue contacts, is in its infancy. Prediction tools that work in one dimension are both mature and generally applicable; they predict secondary structure, residue solvent accessibility, and the location of transmembrane helices with reasonable accuracy. These and other prediction methods have gained immensely from the rapid increase of information in publicly accessible databases. Growing databases will lead to further improvements of prediction methods and, thus, to narrowing the gap between the number of known protein sequences and known protein structures.

Amino Acid Sequence↗

Mutation of the ras genes is a rare genetic event in the histologic transformation of follicular lymphoma.

The role of ras gene mutations in the progression of follicular lymphoma has been ascertained by SSCP-PCR and sequencing. A total of 40 transformed lymphomas were studied, 16 of which had a matched preceding low-grade biopsy. Only one transformed lymphoma was found to have a missense mutation at codon 12 of N-ras, resulting in an amino acid change of glycine to serine. We conclude that mutation within the ras gene family is a rare event in the transformation of follicular lymphoma.

Base Sequence↗

Macromolecular structure information and databases. The EU BRIDGE Database Project Consortium.

The current status and future outlook of macromolecular structure databases and information handling, with particular reference to European databases, are reviewed. Issues concerning the efficiency with which data are represented, validated, archived and accessed are discussed in view of the fast growing body of information on structures of biological macromolecules.

Databases, Factual↗

A sequence property approach to searching protein databases.

Currently available sequence alignment programs are generally not capable of detecting functional and structural homologs in the twilight zone of sequence similarity, i.e. when the sequence identity falls below about 25%. Here we attempt to detect such weak similarities using an approach based on a notion of protein sequence similarity radically different from that used in sequential alignment. The approach defines protein sequence dissimilarity (or distance) as a weighted sum of differences of compositional properties such as singlet and doublet amino acid composition, molecular weight, isoelectric point (protein property search or PropSearch). With PropSearch, either single sequences can be used for a database query, or multiple sequences can be merged into an "average" sequence reflecting the average composition of a protein family. First, we show that members of structural protein families have a low mutual PropSearch distance when the weights are optimized to discriminate maximally between structural families. Second, we demonstrate the results of database searches using the PropSearch method. Such searches are very rapid when scanning a preprocessed database and do not require alignments. In cases in which conventional alignment tools fail to detect similarities, PropSearch can be used to generate hypotheses about possible structural or functional relationships between a new sequence and sequences in the database.

Algorithms↗

Investigating the structural determinants of the p21-like triphosphate and Mg2+ binding site.

Amongst the superfamily of nucleotide binding proteins, the classical mononucleotide binding fold (CMBF), is the one that has been best characterized structurally. The common denominator of all the members is the triphosphate/Mg2+ binding site, whose signature has been recognized as two structurally conserved stretches of residues: the Kinase 1 and 2 motifs that participate in triphosphate and Mg2+ binding, respectively. The Kinase 1 motif is borne by a loop (the P-loop), whose structure is conserved throughout the whole CMBF family. The low sequence similarity between the different members raises questions about which interactions are responsible for the active structure of the P-loop. What are the minimal requirements for the active structure of the P-loop? Why is the P-loop structure conserved despite the diverse environments in which it is found? To address this question, we have engineered the Kinase 1 and 2 motifs into a protein that has the CMBF and no nucleotide binding activity, the chemotactic protein from Escherichia coli, CheY. The mutant does not exhibit any triphosphate/Mg2+ binding activity. The crystal structure of the mutant reveals that the engineered P-loop is in a different conformation than that found in the CMBF. This demonstrates that the native structure of the P-loop requires external interactions with the rest of the protein. On the basis of an analysis of the conserved tertiary contacts of the P-loop in the mononucleotide binding superfamily, we propose a set of residues that could play an important role in the acquisition of the active structure of the P-loop.

Amino Acid Sequence↗

Evolutionary link between glycogen phosphorylase and a DNA modifying enzyme.

We report here an unexpected similarity in three-dimensional structure between glucosyltransferases involved in very different biochemical pathways, with interesting evolutionary and functional implications. One is the DNA modifying enzyme beta-glucosyltransferase from bacteriophage T4, alias UDP-glucose:5-hydroxymethyl-cytosine beta-glucosyltransferase. The other is the metabolic enzyme glycogen phosphorylase, alias 1.4-alpha-D-glucan:orthophosphate alpha-glucosyltransferase. Structural alignment revealed that the entire structure of beta-glucosyltransferase is topographically equivalent to the catalytic core of the much larger glycogen phosphorylase. The match includes two domains in similar relative orientation and connecting helices, with a positional root-mean-square deviation of only 3.4 A for 256 C alpha atoms. An interdomain rotation seen in the R- to T-state transition of glycogen phosphorylase is similar to that observed in beta-glucosyltransferase on substrate binding. Although not a single functional residue is identical, there are striking similarities in the spatial arrangement and in the chemical nature of the substrates. The functional analogies are (beta-glucosyltransferase-glycogen phosphorylase): ribose ring of UDP-pyridoxal ring of pyridoxal phosphate co-enzyme; phosphates of UDP-phosphate of co-enzyme and reactive orthophosphate; glucose unit transferred to DNA-terminal glucose unit extracted from glycogen. We anticipate the discovery of additional structurally conserved members of the emerging glucosyltransferase superfamily derived from a common ancient evolutionary ancestor of the two enzymes.

Amino Acid Sequence↗

Novel protein families in archaean genomes.

In a quest for novel functions in archaea, all archaean hypothetical open reading frames (ORFs), as annotated in the Swiss-Prot protein sequence database, were used to search the latest databases for the identification of characterized homologues. Of the 95 hypothetical archaean ORFs, 25 were found to be homologous to another hypothetical archaean ORF, while 36 were homologous to non-archaean proteins, of which as many as 30 were homologous to a characterized protein family. Thus the level of sequence similarity in this set reaches 64%, while the level of function assignment is only 32%. Of the ORFs with predicted functions, 12 homologies are reported here for the first time and represent nine new functions and one gene duplication at an acetyl-coA synthetase locus. The novel functions include components of the transcriptional and translational apparatus, such as ribosomal proteins, modification enzymes and a translation initiation factor. In addition, new enzymes are identified in archaea, such as cobyric acid synthase, dCTP deaminase and the first archaean homologues of a new subclass of ATP binding proteins found in fungi. Finally, it is shown that the putative laminin receptor family of eukaryotes and an archaean homologue belong to the previously characterized ribosomal protein family S2 from eubacteria. From the present and previous work, the major implication is that archaea seem to have a mode of expression of genetic information rather similar to eukaryotes, while eubacteria may have proceeded into unique ways of transcription and translation. In addition, with the detection of proteins in various metabolic and genetic processes in archaea, we can further predict the presence of additional proteins involved in these processes.

Animal Population Groups↗