Search PubMedSearch

Biomedical subjects

C Sander

Publications and source records attributed to C Sander.

At least 19 recordsLinked to original sources

Protein folds and families: sequence and structure alignments.

Dali and HSSP are derived databases organizing protein space in the structurally known regions. We use an automatic structure alignment program (Dali) for the classification of all known 3D structures based on all-against-all comparison of 3D structures in the Protein Data Bank. The HSSP database associates 1D sequences with known 3D structures using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). As a result, the HSSP database not only provides aligned sequence families, but also implies secondary and tertiary structures covering 36% of all sequences in Swiss-Prot. The structure classification by Dali and the sequence families in HSSP can be browsed jointly from a web interface providing a rich network of links between neighbours in fold space, between domains and proteins, and between structures and sequences. In particular, this results in a database of explicit multiple alignments of protein families in the twilight zone of sequence similarity. The organization of protein structures and families provides a map of the currently known regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination. The databases are available from http://www.embl-ebi.ac.uk/dali/

Databases, Factual

Dictionary of recurrent domains in protein structures.

The rapid growth in the number of experimentally determined three-dimensional protein structures has sharpened the need for comprehensive and up-to-date surveys of known structures. Classic work on protein structure classification has made it clear that a structural survey is best carried out at the level of domains, i.e., substructures that recur in evolution as functional units in different protein contexts. We present a method for automated domain identification from protein structure atomic coordinates based on quantitative measures of compactness and, as the new element, recurrence. Compactness criteria are used to recursively divide a protein into a series of successively smaller and smaller substructures. Recurrence criteria are used to select an optimal size level of these substructures, so that many of the chosen substructures are common to different proteins at a high level of statistical significance. The joint application of these criteria automatically yields consistent domain definitions between remote homologs, a result difficult to achieve using compactness criteria alone. The method is applied to a representative set of 1,137 sequence-unique protein families covering 6,500 known structures. Clustering of the resulting set of domains (substructures) yields 594 distinct fold classes (types of substructures). The Dali Domain Dictionary (http://www.embl-ebi.ac.uk/dali/) not only provides a global structural classification, but also a comprehensive description of families of protein sequences grouped around representative proteins of known structure. The classification will be continuously updated and can serve as a basis for improving our understanding of protein evolution and function and for evolving optimal strategies to complete the map of all natural protein structures.

Amino Acid Sequence

Updated catalogue of homologues to human disease-related proteins in the yeast genome.

The recent availability of the full Saccharomyces cerevisiae genome offers a perfect opportunity for revising the number of homologues to human disease-related proteins. We carried out automatic analysis of the complete S. cerevisiae genome and of the set of human disease-related proteins as identified in the SwissProt sequence data base. We identified 285 yeast proteins similar to 155 human disease-related proteins, including 239 possible cases of human-yeast direct functional equivalence (orthology). Of these, 40 cases are suggested as new, previously undiscovered relationships. Four of them are particularly interesting, since the yeast sequence is the most phylogenetically distant member of the protein family, including proteins related to diseases such as phenylketonuria, lupus erythematosus, Norum and fish eye disease and Wiskott-Aldrich syndrome.

Amino Acid Sequence

The HSSP database of protein structure-sequence alignments and family profiles.

HSSP (http: //www.sander.embl-ebi.ac.uk/hssp/) is a derived database merging structure (3-D) and sequence (1-D) information. For each protein of known 3D structure from the Protein Data Bank (PDB), we provide a multiple sequence alignment of putative homologues and a sequence profile characteristic of the protein family, centered on the known structure. The list of homologues is the result of an iterative database search in SWISS-PROT using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed putative homologues are very likely to have the same 3D structure as the PDB protein to which they have been aligned. As a result, the database not only provides aligned sequence families, but also implies secondary and tertiary structures covering 33% of all sequences in SWISS-PROT.

Computer Communication Networks

Touring protein fold space with Dali/FSSP.

The FSSP database and its new supplement, the Dali Domain Dictionary, present a continuously updated classification of all known 3D protein structures. The classification is derived using an automatic structure alignment program (Dali) for the all-against-all comparison of structures in the Protein Data Bank. From the resulting enumeration of structural neighbours (which form a surprisingly continuous distribution in fold space) we derive a discrete fold classification in three steps: (i) sequence-related families are covered by a representative set of protein chains; (ii) protein chains are decomposed into structural domains based on the recurrence of structural motifs; (iii) folds are defined as tight clusters of domains in fold space. The fold classification, domain definitions and test sets for sequence-structure alignment (threading) are accessible on the web at www.embl-ebi.ac.uk/dali . The web interface provides a rich network of links between neighbours in fold space, between domains and proteins, and between structures and sequences leading, for example, to a database of explicit multiple alignments of protein families in the twilight zone of sequence similarity. The Dali/FSSP organization of protein structures provides a map of the currently known regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination.

Computer Communication Networks

Frame: detection of genomic sequencing errors.

MOTIVATION: The underlying error rate for genomic sequencing sometimes results in the introduction of artificial frameshifts and in-frame stop codons into putative protein encoding genes. Severe errors are then introduced into the inferred transcripts through mis-translation or premature termination. RESULTS: We describe a system for screening segments of DNA for frameshift and in-frame stop errors in coding regions. The method is based on homology matching using blastx to compare all six reading frames of the query nucleotide sequence against selected protein sequence databases. Fragments of protein matching neighbouring regions of the query DNA are united and extended laterally to define candidate open reading frames, within which, frameshifts and stops are identified. Suitable targets include prokaryotic or other intron-free genomic sequence and complementary DNAs. As an example of its use, we report here two frameshifted ORFs that deviate from the original TIGR sequence annotations for the recently released Helicobacter pylori genome. AVAILABILITY: The tool is accessible via the URL http://www.sander.ebi.ac.uk/frame/. CONTACT: brown@ebi.ac.uk.

Amino Acid Sequence

MView: a web-compatible database search or multiple alignment viewer.

UNLABELLED: MView is a tool for converting the results of a sequence database search into the form of a coloured multiple alignment of hits stacked against the query. Alternatively, an existing multiple alignment can be processed. In either case, the output is simply HTML, so the result is platform independent and does not require a separate application or applet to be loaded. AVAILABILITY: Free from http://www.sander.ebi.ac.uk/mview/ subject to copyright restrictions. CONTACT: brown@ebi.ac.uk

Computer Communication Networks

Removing near-neighbour redundancy from large protein sequence collections.

MOTIVATION: To maximize the chances of biological discovery, homology searching must use an up-to-date collection of sequences. However, the available sequence databases are growing rapidly and are partially redundant in content. This leads to increasing strain on CPU resources and decreasing density of first-hand annotation. RESULTS: These problems are addressed by clustering closely similar sequences to yield a covering of sequence space by a representative subset of sequences. No pair of sequences in the representative set has >90% mutual sequence identity. The representative set is derived by an exhaustive search for close similarities in the sequence database in which the need for explicit sequence alignment is significantly reduced by applying deca- and pentapeptide composition filters. The algorithm was applied to the union of the Swissprot, Swissnew, Trembl, Tremblnew, Genbank, PIR, Wormpep and PDB databases. The all-against-all comparison required to generate a representative set at 90% sequence identity was accomplished in 2 days CPU time, and the removal of fragments and close similarities yielded a size reduction of 46%, from 260 000 unique sequences to 140 000 representative sequences. The practical implications are (i) faster homology searches using, for example, Fasta or Blast, and (ii) unified annotation for all sequences clustered around a representative. As tens of thousands of sequence searches are performed daily world-wide, appropriate use of the non-redundant database can lead to major savings in computer resources, without loss of efficacy. AVAILABILITY: A regularly updated non-redundant protein sequence database (nrdb90), a server for homology searches against nrdb90, and a Perl script (nrdb90.pl) implementing the algorithm are available for academic use from http://www.embl-ebi.ac. uk/holm/nrdb90. CONTACT: holm@embl-ebi.ac.uk

Algorithms

EUCLID: automatic classification of proteins in functional classes by their database annotations.

UNLABELLED: A tool is described for the automatic classification of sequences in functional classes using their database annotations. The Euclid system is based on a simple learning procedure from examples provided by human experts. AVAILABILITY: Euclid is freely available for academics at http://www.gredos.cnb.uam.es/EUCLID, with the corresponding dictionaries for the generation of three, eight and 14 functional classes. CONTACT: E-mail: valencia@cnb.uam.es SUPPLEMENTARY INFORMATION: The results of the EUCLID classification of different genomes are available at http://www.sander.ebi.ac. uk/genequiz/. A detailed description of the different applications mentioned in the text is available at http://www.gredos.cnb.uam. es/EUCLID/Full_Paper

Computational Biology

Dermatoscopy and high frequency sonography: two useful non-invasive methods to increase preoperative diagnostic accuracy in pigmented skin lesions.

Dermatoscopy and high frequency sonography have recently been combined to increase diagnostic preoperative accuracy in the treatment of pigmented skin lesions. In this monocentric study 80 patients with pigmented skin lesions were evaluated clinically, by dermatoscopy, and 20 MHz-sonography followed by dermatohistopathological evaluation; 39 malignant melanomas, 37 common nevi, 3 dysplastic nevi, and 1 nevus Spitz were diagnosed histologically. In 72 of the 80 cases (91.3%) dermatoscopical diagnoses were confirmed by histopathology, compared to only 79% correct clinical diagnoses. For the mere clinical diagnosis of melanoma sensitivity was 79%, specificity was 78% and diagnostic accuracy was 65%. All diagnostic values increased by dermatoscopy: sensitivity reached 90%, specificity was 93%, and diagnostic accuracy was 83%. In order to determine tumor thickness preoperatively tumor thickness was measured by 20 MHz sonography. The correlation of tumor thickness between histometric and sonographic results was determined for nevi (r = 0.93) and melanoma (r = 0.95); 74.3% of melanomas were diagnosed correctly within an 0.2 mm range. Regarding the clinical important limit of 1 mm tumor thickness, 87.2% were diagnosed in accordance with histometric evaluation. An increase of 18% in diagnostic accuracy by dermatoscopy and 87.2% of correctly diagnosed cases of tumor thickness of malignant melanoma by high frequency sonography clearly demonstrate that these methods should be considered standard procedures in the diagnosis of pigmented skin lesions and will facilitate the decision on necessary surgical treatment.

Humans

[Indirect blood pressure measurement in cats with diabetes mellitus, chronic nephropathy and hypertrophic cardiomyopathy].

In the present study blood pressure was measured in cats comparing two indirect methods (oscillometric versus Doppler-sonographic) over a wide pressure range. It was shown, that at the lower pressures Doppler and oscillometric measurements were basically equivalent. However for higher pressures oscillometric measurements were consistently lower than Doppler measurements. This difference became greater as blood pressure increased. The determination of blood pressure by the Doppler-sonographic method was always possible, whereas the measurement by the oscillometric method was often not possible, especially at higher blood pressure levels. In a second step, the frequency of hypertension was determined in cats with diabetes mellitus, chronic renal failure and hypertrophic cardiomyopathy. Eight cats with diabetes mellitus had oszillometric blood pressure values of 101-155 mmHg systolic, 42-105 mmHg diastolic and 65-125 mmHg mean arterial pressure determined at the front leg and 110-167 mmHg systolic, 44-98 mmHg diastolic, and 61-125 mean arterial pressure determined at the tail. The Doppler-sonographic values were 120-180 mmHg. Only the oscillometric measurement (at the tail) of the systolic pressure was significantly higher than that of normal cats. In 11 cats with chronic renal failure the following values were determined by the oszillometric method: at the front leg 137-182 mmHg systolic, 74-138 mmHg diastolic, 100-162 mmHg mean arterial pressure and at the tail 134-189 mmHg systolic, 53-109 mmHg diastolic, 80-135 mmHg mean arterial pressure. With the Doppler-sonographic technique the blood pressure was between 120 and 280 mmHg. All blood pressure measurements were significantly higher than those of healthy cats, except the oscillometric measurements of diastolic blood pressure. In 12 cats with hypertrophic cardiomyopathy systolic pressure was 108-179 mmHg, diastolic pressure was 64-135 mmHg, and mean arterial pressure was 89-154 mmHg at the front leg using the oscillometric method. At the tail results were as follows: 121-201 mmHg systolic, 61-141 mmHg diastolic, and 85-160 mmHg mean arterial pressure. By the Doppler-sonographic technique determined blood pressure was 110-260 mmHg. All oscillometric measurements except the diastolic pressure determined at the front leg were significantly higher than in normal cats. Four cats with chronic renal failure and five cats with hypertrophic cardiomyopathy showed retinal hemorrhages and/or detachments. Eight of this nine cats had blood pressure measurements above the normal range. We conclude that hypertension can be detected in cats with several diseases. In most cases reliable measurements can only be obtained by Doppler-sonographic methods.

Animals

The interaction of class B G protein-coupled receptors with their hormones.

In common with many G protein-coupled receptors, dysfunction in members of the Class B or glucagon-like receptors can elicit a wide spectrum of disease related activities. Consequently, they are potential targets in many different areas of pharmacological research. Unlike the class A or rhodopsin-like receptors, for which at least some structural similarity to bacteriorhodopsin has been detected, absolutely no structural information is available for the Class B G protein-coupled receptors. We present a computational study that exploits the experimental work performed by evolution to indicate residues that are potentially involved in ligand binding in the Class B G protein-coupled receptors. We perform an analysis of mutations that occurred in a correlated fashion between the receptors and their peptidic ligands. The inference that the residues detected in this manner are involved in a direct interaction between the receptor and the ligand is in good agreement with the mutation studies that have already been published.

Amino Acid Sequence

Are binding residues conserved?

We present our attempt to quantify the evolutionary dynamics of functional residues in a representative set of protein structures and their homologous sequences. Using the log-odds formalism, the preference for all twenty amino acids to be conserved or participate in binding (or active) sites is examined. It appears that while there is a tendency for functional residues to be conserved, the two preference scales do not coincide. Remarkable differences between amino acid types emerge from this comparative study. The current approach is expected to lead towards a better understanding of functional site architecture in proteins.

Amino Acid Sequence

Protein fold recognition by prediction-based threading.

In fold recognition by threading one takes the amino acid sequence of a protein and evaluates how well it fits into one of the known three-dimensional (3D) protein structures. The quality of sequence-structure fit is typically evaluated using inter-residue potentials of mean force or other statistical parameters. Here, we present an alternative approach to evaluating sequence-structure fitness. Starting from the amino acid sequence we first predict secondary structure and solvent accessibility for each residue. We then thread the resulting one-dimensional (1D) profile of predicted structure assignments into each of the known 3D structures. The optimal threading for each sequence-structure pair is obtained using dynamic programming. The overall best sequence-structure pair constitutes the predicted 3D structure for the input sequence. The method is fine-tuned by adding information from direct sequence-sequence comparison and applying a series of empirical filters. Although the method relies on reduction of 3D information into 1D structure profiles, its accuracy is, surprisingly, not clearly inferior to methods based on evaluation of residue interactions in 3D. We therefore hypothesise that existing 1D-3D threading methods essentially do not capture more than the fitness of an amino acid sequence for a particular 1D succession of secondary structure segments and residue solvent accessibility. The prediction-based threading method on average finds any structurally homologous region at first rank in 29% of the cases (including sequence information). For the 22% first hits detected at highest scores, the expected accuracy rose to 75%. However, the task of detecting entire folds rather than homologous fragments was managed much better; 45 to 75% of the first hits correctly recognised the fold.

Algorithms

DNA sequencing and analysis of 130 kb from yeast chromosome XV.

We have determined the nucleotide sequence of 129,524 bases of yeast (Saccharomyces cerevisiae) chromosome XV. Sequence analysis revealed the presence of 59 non-overlapping open reading frames (ORFs) of length > 300 bp, three tRNA genes, four delta elements and one Ty-element. Among the 21 previously known yeast genes (36% of all ORFs in this fragment) were nucleoporin (NUP1), ras protein (RAS1), RNA polymerase III (RPC1) and elongation factor 2 (EF2). Further, 31 ORFs (53% of the total) were found to be homologous to known protein or DNA sequences, or sequence patterns. For seven ORFs (11% of the total) no homology was found. Among the most interesting protein identification in this DNA fragment are an inositol polyphosphatase, the second gene of this type found in yeast (homologous to the human OCRL gene involved in Lowe's syndrome), a new ADP ribosylation factor of the arf6 subfamily, the first protein containing three C2 domains, and an ORF similar to a Bacillus subtilis cell-cycle related protein. For each ORF detailed sequence analysis was carried out, with a full consideration of its biological function and pointing out key regions of interest for further functional analysis.

ADP-Ribosylation Factors