Search PubMedSearch

Biomedical subjects

C Sander

Publications and source records attributed to C Sander.

At least 37 records · Page 2Linked to original sources

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural (3-D) and sequence (1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in SwissProt using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 29% of all SwissProt-stored sequences.

Amino Acid Sequence

Dali/FSSP classification of three-dimensional protein folds.

The FSSP database presents a continuously updated structural classification of three-dimensional protein folds. It is derived using an automatic structure comparison program (Dali) for the all-against-all comparison of over 6000 three-dimensional coordinate sets in the Protein Data Bank (PDB). Sequence-related protein families are covered by a representative set of 813 protein chains. Hierachical clustering based on structural similarities yields a fold tree that defines 253 fold classes. For each representative protein chain, there is a database entry containing structure-structure alignments with its structural neighbours in the PDB. The database is accessible online through World Wide Web browsers and by anonymous ftp (file transfer protocol). The overview of fold space and the individual data sets provide a rich source of information for the study of both divergent and convergent aspects of molecular evolution, and define useful test sets and a standard of truth for assessing the correctness of sequence-sequence or sequence-structure alignments.

Amino Acid Sequence

Characterization of new proteins found by analysis of short open reading frames from the full yeast genome.

We have analysed short open reading frames (between 150 and 300 base pairs long) of the yeast genome (Saccharomyces cerevisiae) with a two-step strategy. The first step selects a candidate set of open reading frames from the DNA sequence based on statistical evaluation of DNA and protein sequence properties. The second step filters the candidate set by selecting open reading frames with high similarity to other known sequences (from any organism). As a result, we report ten new predicted proteins not present in the current sequence databases. These include a new alcohol dehydrogenase, a protein probably related to the cell cycle, as well as a homolog of the prokaryotic ribosomal protein L36 likely to be a mitochondrial ribosomal protein coded in the nuclear genome. We conclude that the analysis of short open reading frames leads to biologically interesting discoveries, even though the quantitative yield of new proteins is relatively low.

Alcohol Dehydrogenase

Predicting protein structure using hidden Markov models.

We discuss how methods based on hidden Markov models performed in the fold-recognition section of the CASP2 experiment. Hidden Markov models were built for a representative set of just over 1,000 structures from the Protein Data Bank (PDB). Each CASP2 target sequence was scored against this library of HMMs. In addition, an HMM was built for each of the target sequences and all of the sequences in PDB were scored against that target model, with a good score on both methods indicating a high probability that the target sequence is homologous to the structure. The method worked well in comparison to other methods used at CASP2 for targets of moderate difficulty, where the closest structure in PDB could be aligned to the target with at least 15% residue identity.

Markov Chains

Classification of protein families and detection of the determinant residues with an improved self-organizing map.

Using a SOM (self-organizing map) we can classify sequences within a protein family into subgroups that generally correspond to biological subcategories. These maps tend to show sequence similarity as proximity in the map. Combining maps generated at different levels of resolution, the structure of relations in protein families can be captured that could not otherwise be represented in a single map. The underlying representation of maps enables us to retrieve characteristic sequence patterns for individual subgroups of sequences. Such patterns tend to correspond to functionally important regions. We present a modified SOM algorithm that includes a convergence test that dynamically controls the learning parameters to adapt them to the learning set instead of being fixed and externally optimized by trial and error. Given the variability of protein family size and distribution, the addition of this features is necessary. The method is successfully tested with a number of families. The rab family of small GTPases is used to illustrate the performance of the method.

Algorithms

Bioinformatics: from genome data to biological knowledge.

Recently, molecular biologists have sequenced about a dozen bacterial genomes and the first eukaryotic genome. We can now obtain answers to detailed questions about the complete set of genes of an organism. Bioinformatics methods are increasingly used for attaching biological knowledge to long lists of genes, assigning genes to biological pathways, comparing the gene sets of different species, identifying specificity factors, and describing sets of highly conserved proteins common to all domains of life. Substantial progress has recently been made in the availability of primary and added-value databases, in the development of algorithms and of network information services for genome analysis. The pharmaceutical industry has greatly benefited from the accumulation of sequence data through the identification of targets and candidates for the development of drugs, vaccines, diagnostic markers and therapeutic proteins.

Databases, Factual

Enzyme HIT.

Explore the source record for details and available documents.

Acid Anhydride Hydrolases

Objectively judging the quality of a protein structure from a Ramachandran plot.

MOTIVATION: Statistical methods that compare observed and expected distributions of experimental observables provide powerful tools for the quality control of protein structures. The distribution of backbone dihedral angles ('Ramachandran plot') has often been used for such quality control, but without a firm statistical foundation. RESULTS: A new and-simple method is presented for judging the quality of a protein structure based on the distribution of backbone dihedral angles. Inputs to the method are 60 torsion angle distributions extracted from protein structures solved at high resolution; one for each combination of residue type and tri-state secondary structure. Output for a protein is a Ramachandran Z-score, expressing the quality of the Ramachandran plot relative to current state-of-the-art structures.

Algorithms

Urticaria haemorrhagica profunda.

Substantial subcutaneous haemorrhage without preceding trauma or underlying bleeding disorder is a rare occurrence in dermatological practice, essentially restricted to early childhood (acute haemorrhagic oedema of childhood). We report an adolescent with a morphologically unique bleeding manifestation. A 16-year-old boy presented with two episodes of massive subcutaneous haemorrhage in association with urticarial vasculitis. There was no history of preceding trauma or haemorrhagic disorder. Haemorrhage was observed in areas typically affected by angioedema, such as the periorbital, perioral, lingual, sublingual and laryngeal areas. History revealed an atopic diathesis with hay fever and examination showed alopecia areata. An antinuclear antibody titre and the presence of lupus anticoagulant indicated transient antiphospholipid antibodies. As urticaria corresponds to urticaria profunda angioedema, we hypothesize a pathophysiological relationship between superficial urticarial vasculitis and the deep variant of urticarial vasculitic disease, leading to the unique morphology present in our patient.

Adolescent

An evolutionary treasure: unification of a broad set of amidohydrolases related to urease.

The recent determination of the three-dimensional structure of urease revealed striking similarities of enzyme architecture to adenosine deaminase and phosphotriesterase, evidence of a distant evolutionary relationship that had gone undetected by one-dimensional sequence comparisons. Here, based on an analysis of conservation patterns in three dimensions, we report the discovery of the same active-site architecture in an even larger set of enzymes involved primarily in nucleotide metabolism. As a consequence, we predict the three-dimensional fold and details of the active site architecture for dihydroorotases, allantoinases, hydantoinases, AMP-, adenine and cytosine deaminases, imidazolonepropionase, aryldialkylphosphatase, chlorohydrolases, formylmethanofuran dehydrogenases, and proteins involved in animal neuronal development. Two member families are common to archaea, eubacteria, and eukaryota. Thirteen other functions supported by the same structural motif and conserved chemical mechanism apparently represent later adaptations for different substrate specificities in different cellular contexts.

Amidohydrolases

Decision support system for the evolutionary classification of protein structures.

The structures of nearly a thousand sequence-unique proteins represent only 300 different 3D shapes. Is structural resemblance between proteins with little sequence similarity the result of physical convergence to favourable folding patterns, or does it reflect a memory of common evolutionary history? Separating these two processes is important for organizing genome data in terms of protein families and for theoretical approaches to protein structure prediction by fold recognition techniques. Achieving separation requires a combination of structure, sequence and functional analysis of proteins. For this purpose, we are developing a decision support system that scans heterogeneous protein sequence and structure related databases, and collects or calculates characters indicative of common functional constraints. The criteria include sequence homology, analysis of 3D clusters of conserved residues, conservation of active sites, and keyword analysis of biological function. Even without extensive refinement, application of a combination of these criteria to a test set representing all currently known protein structures yields 87% coverage with 7% false positives, compared to 53% coverage by only 1D sequence criteria. Thus, the semiautomatic prototype system significantly enhances the efficiency of unifying families of functionally related proteins in spite of long evolutionary distances.

Binding Sites

Absence of anaplastic lymphoma kinase (ALK) and Epstein-Barr virus gene products in primary cutaneous anaplastic large cell lymphoma and lymphomatoid papulosis.

The prevalence of the t(2;5)(p23;q35) and/or anaplastic lymphoma kinase (ALK) gene products in cutaneous anaplastic large cell (ALC) lymphomas and a potential precursor lesion, lymphomatoid papulosis (LyP), is controversial. ALK gene products, which are absent from normal lymphohaematopoietic cells, are a phenotypic marker of lymphomas carrying the t(2;5). We used in situ hybridization and immunohistology to screen 14 cutaneous ALC lymphomas, 21 cases of LyP, and one nodal ALC lymphoma associated with LyP for ALK gene products. ALK gene products were not detectable in these cases. In contrast, ALK gene products were found in a lymphonodal ALC lymphoma with subsequent extension to the skin and in t(2;5)-positive cell lines. Detection of the Epstein-Barr virus (EBV)-encoded small nuclear transcripts (EBER), and of immunoglobulin light chain transcripts served to check for the presence of cellular RNA in the tissue sections. EBER transcripts were found in scattered reactive lymphoid cells, but not in atypical or tumour cells. ALK gene expression and EBV infection seem to be a rare finding in cutaneous ALC lymphomas and LyP. This points to a molecular aetiology of primary cutaneous ALC lymphomas and LyP distinct from that of extracutaneous CD30+ lymphoproliferative disease. Detection of the t(2;5) or ALK gene products in cutaneous lymphoproliferative lesions therefore requires exclusion of extracutaneous ALC lymphoma in such patients.

Anaplastic Lymphoma Kinase

Mapping the protein universe.

The comparison of the three-dimensional shapes of protein molecules poses a complex algorithmic problem. Its solution provides biologists with computational tools to organize the rapidly growing set of thousands of known protein shapes, to identify new types of protein architecture, and to discover unexpected evolutionary relations, reaching back billions of years, between protein molecules. Protein shape comparison also improves tools for identifying gene functions in genome databases by defining the essential sequence-structure features of a protein family. Finally, an exhaustive all-on-all shape comparison provides a map of physical attractor regions in the abstract shape space of proteins, with implications for the processes of protein folding and evolution.

Algorithms

Genomes with distinct function composition.

The functional composition of organisms can be analysed for the first time with the appearance of complete or sizeable parts of various genomes. We have reduced the problem of protein function classification to a simple scheme with three classes of protein function: energy-, information- and communication-associated proteins. Finer classification schemes can be easily mapped to the above three classes. To deal with the vast amount of information, a system for automatic function classification using database annotations has been developed. The system is able to classify correctly about 80% of the query sequences with annotations. Using this system, we can analyse samples from the genomes of the most represented species in sequence databases and compare their genomic composition. The similarities and differences for different taxonomic groups are strikingly intuitive. Viruses have the highest proportion of proteins involved in the control and expression of genetic information. Bacteria have the highest proportion of their genes dedicated to the production of proteins associated with small molecule transformations and transport. Animals have a very large proportion of proteins associated with intra- and intercellular communication and other regulatory processes. In general, the proportion of communication-related proteins increases during evolution, indicating trends that led to the emergence of the eukaryotic cell and later the transition from unicellular to multicellular organisms.

Animals

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural three dimensional (3-D) and sequence one dimensional(1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in Swissprot using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 27% of all Swissprot-stored sequences.

Amino Acid Sequence