Search PubMed⌕ Search

Biomedical subjects

Francisco S Domingues

Publications and source records attributed to Francisco S Domingues.

8 recordsLinked to original sources

A new measure for functional similarity of gene products based on Gene Ontology.

BACKGROUND: Gene Ontology (GO) is a standard vocabulary of functional terms and allows for coherent annotation of gene products. These annotations provide a basis for new methods that compare gene products regarding their molecular function and biological role. RESULTS: We present a new method for comparing sets of GO terms and for assessing the functional similarity of gene products. The method relies on two semantic similarity measures; simRel and funSim. One measure (simRel) is applied in the comparison of the biological processes found in different groups of organisms. The other measure (funSim) is used to find functionally related gene products within the same or between different genomes. Results indicate that the method, in addition to being in good agreement with established sequence similarity approaches, also provides a means for the identification of functionally related proteins independent of evolutionary relationships. The method is also applied to estimating functional similarity between all proteins in Saccharomyces cerevisiae and to visualizing the molecular function space of yeast in a map of the functional space. A similar approach is used to visualize the functional relationships between protein families. CONCLUSION: The approach enables the comparison of the underlying molecular biology of different taxonomic groups and provides a new comparative genomics tool identifying functionally related gene products independent of homology. The proposed map of the functional space provides a new global view on the functional relationships between gene products or protein families.

Algorithms↗

NOXclass: prediction of protein-protein interaction types.

BACKGROUND: Structural models determined by X-ray crystallography play a central role in understanding protein-protein interactions at the molecular level. Interpretation of these models requires the distinction between non-specific crystal packing contacts and biologically relevant interactions. This has been investigated previously and classification approaches have been proposed. However, less attention has been devoted to distinguishing different types of biological interactions. These interactions are classified as obligate and non-obligate according to the effect of the complex formation on the stability of the protomers. So far no automatic classification methods for distinguishing obligate, non-obligate and crystal packing interactions have been made available. RESULTS: Six interface properties have been investigated on a dataset of 243 protein interactions. The six properties have been combined using a support vector machine algorithm, resulting in NOXclass, a classifier for distinguishing obligate, non-obligate and crystal packing interactions. We achieve an accuracy of 91.8% for the classification of these three types of interactions using a leave-one-out cross-validation procedure. CONCLUSION: NOXclass allows the interpretation and analysis of protein quaternary structures. In particular, it generates testable hypotheses regarding the nature of protein-protein interactions, when experimental results are not available. We expect this server will benefit the users of protein structural models, as well as protein crystallographers and NMR spectroscopists. A web server based on the method and the datasets used in this study are available at http://noxclass.bioinf.mpi-inf.mpg.de/.

Algorithms↗

Automated clustering of ensembles of alternative models in protein structure databases.

Experimentally determined protein structures have been classified in different public databases according to their structural and evolutionary relationships. Frequently, alternative structural models, determined using X-ray crystallography or NMR spectroscopy, are available for a protein. These models can present significant structural dissimilarity. Currently there is no classification available for these alternative structures. In order to classify them, we developed STRuster, an automated method for clustering ensembles of structural models according to their backbone structure. The method is based on the calculation of carbon alpha (Calpha) distance matrices. Two filters are applied in the calculation of the dissimilarity measure in order to identify both large and small (but significant) backbone conformational changes. The resulting dissimilarity value is used for hierarchical clustering and partitioning around medoids (PAM). Hierarchical clustering reflects the hierarchy of similarities between all pairs of models, while PAM groups the models into the 'optimal' number of clusters. The method has been applied to cluster the structures in each SCOP species level and can be easily applied to any other sets of conformers. The results are available at: http://bioinf.mpi-sb.mpg.de/projects/struster/.

Aldehyde-Lyases↗

Calculating the statistical significance of changes in pathway activity from gene expression data.

We present a statistical approach to scoring changes in activity of metabolic pathways from gene expression data. The method identifies the biologically relevant pathways with corresponding statistical significance. Based on gene expression data alone, only local structures of genetic networks can be recovered. Instead of inferring such a network, we propose a hypothesis-based approach. We use given knowledge about biological networks to improve sensitivity and interpretability of findings from microarray experiments. Recently introduced methods test if members of predefined gene sets are enriched in a list of top-ranked genes in a microarray study. We improve this approach by defining scores that depend on all members of the gene set and that also take pairwise co-regulation of these genes into account. We calculate the significance of co-regulation of gene sets with a nonparametric permutation test. On two data sets the method is validated and its biological relevance is discussed. It turns out that useful measures for co-regulation of genes in a pathway can be identified adaptively. We refine our method in two aspects specific to pathways. First, to overcome the ambiguity of enzyme-to-gene mappings for a fixed pathway, we introduce algorithms for selecting the best fitting gene for a specific enzyme in a specific condition. In selected cases, functional assignment of genes to pathways is feasible. Second, the sensitivity of detecting relevant pathways is improved by integrating information about pathway topology. The distance of two enzymes is measured by the number of reactions needed to connect them, and enzyme pairs with a smaller distance receive a higher weight in the score calculation.

Journal Article↗

WILMA-automated annotation of protein sequences.

Large-scale annotation of sets of proteins is a frequently occurring task in association with genome sequencing projects. Here, we present an automated platform for the functional annotation of large sets of protein sequences. Various bioinformatics tools are used to achieve a comprehensive description of protein sequences and to link these results to standard Gene Ontology descriptors for molecular function, biological processes and cellular components. Access to the annotation is provided via a web-interface and database queries. These interfaces allow to formulate proteome wide queries as well as the investigation of details of individual results. WILMA annotations of the proteomes of Homo sapiens, Mus musculus, Arabidopsis thaliana and Caenorhabditis elegans are accessible at http://www.came.sbg.ac.at/wilma/

Amino Acid Sequence↗

Structural localization of disease-associated sequence variations in the NACHT and LRR domains of PYPAF1 and NOD2.

Several autoinflammatory diseases with distinct clinical manifestations have been associated with sequence variations in the gene products PYPAF1/CIAS1 and NOD2/CARD15. Both proteins belong to the PYD/CARD-containing family of apoptosis regulators and activators of pro-inflammatory caspases. To gain insight into the dysfunctional role of sequence alterations, we assembled a structure-based multiple sequence alignment of family members and related proteins. This allowed us to analyze the putative effect of the alterations on the function of nucleotide-binding (NACHT) and leucine-rich repeat (LRR) domains shared by the family members. In support of this analysis, we carefully selected template structures for the NACHT and LRR domains and mapped the genetic variations onto 3D domain models. Additionally, we propose a model of the NACHT and LRR domain complex. Our study revealed that many of the disease-associated sequence variants are located close to highly conserved sequence regions of functional relevance and are spatially adjacent in the predicted 3D structure. The implications on the domain functions such as NTP-hydrolysis or oligomerization are discussed.

Amino Acid Sequence↗

Identification of mammalian orthologs associates PYPAF5 with distinct functional roles.

PYRIN- and CARD-containing proteins belong to a recently identified protein family involved in the regulation of apoptosis and inflammatory processes. Variations in the gene products of the family members PYPAF1 and NOD2/CARD15 have been associated with several autoinflammatory diseases. We could identify the mouse orthologs of PYPAF1, PYPAF5, NOD1, NOD2 and the rat ortholog of PYPAF5. Intriguingly, we found that PYPAF5 has been reported previously not only as regulator of NF-kappaB and caspase-1, but also as angiotensin II and vasopressin receptor. In particular, based on a comprehensive sequence analysis, we propose a structural model for this hormone receptor that is different from the model suggested previously.

Amino Acid Sequence↗

Protein function from sequence and structure data.

With the large amount of genomics and proteomics data that we are confronted with, computational support for the elucidation of protein function becomes more and more pressing. Many different kinds of biological data harbour signals of protein function, but these signals are often concealed. Computational methods that use protein sequence and structure data can be used for discovering these signals. They provide information that can substantially speed up experimental function elucidation. In this review we concentrate on such methods.

Amino Acid Sequence↗