Search PubMedSearch

Biomedical subjects

C Ouzounis

Publications and source records attributed to C Ouzounis.

At least 19 recordsLinked to original sources

RGD sequences in several receptor proteins: novel cell adhesion function of receptors?

In the process of homology modelling of the 3-dimensional structure of alleles of the human histocompatibility protein HLA-DQ, we discovered that its RGD tripeptide (beta 167-169) forms part of a loop. A search through protein sequence data bases, revealed this cell adhesion motif in 67 integral plasma membrane proteins (in 48 extracellularly, and in the remaining 19 intracellularly), which are bona fide receptors, and none of them has thus far been considered as a cell adhesion protein. The 3-dimensional structure of one of these, the rat neonatal Fc receptor, is known and its extracellular RGD sequence is in an adhesion-like loop, a fact that went unnoticed in the original papers. In a few other cases, e.g. rat and mouse growth hormone receptor, and mouse CD40 ligand, homology modelling by ourselves and others reveals that the said sequences are part of a loop, in similarity to all RGD sequences found in proteins with established adhesion function and known 3-dimensional structure. Likewise, inspection of all known protein 3-dimensional structures containing an RGD sequence, and not having a documented cell adhesion function (total of 65 separate entries) shows that such sequence is mostly (52/65 or 80% of cases) part of a loop. We therefore call attention to these surprising findings, discuss the possible cell adhesion role of these receptor proteins, and draw an analogy from the two well characterised examples, that of soluble IGF binding protein 1 and the transcriptional activator protein Tat of HIV, where their RGD sequences have been shown by site-directed mutagenesis to participate in cell-adhesion interactions, without prior knowledge of the location of the tripeptide, or the 3-dimensional structure of the respective protein.

Animals

EUCLID: automatic classification of proteins in functional classes by their database annotations.

UNLABELLED: A tool is described for the automatic classification of sequences in functional classes using their database annotations. The Euclid system is based on a simple learning procedure from examples provided by human experts. AVAILABILITY: Euclid is freely available for academics at http://www.gredos.cnb.uam.es/EUCLID, with the corresponding dictionaries for the generation of three, eight and 14 functional classes. CONTACT: E-mail: valencia@cnb.uam.es SUPPLEMENTARY INFORMATION: The results of the EUCLID classification of different genomes are available at http://www.sander.ebi.ac. uk/genequiz/. A detailed description of the different applications mentioned in the text is available at http://www.gredos.cnb.uam. es/EUCLID/Full_Paper

Computational Biology

Are binding residues conserved?

We present our attempt to quantify the evolutionary dynamics of functional residues in a representative set of protein structures and their homologous sequences. Using the log-odds formalism, the preference for all twenty amino acids to be conserved or participate in binding (or active) sites is examined. It appears that while there is a tendency for functional residues to be conserved, the two preference scales do not coincide. Remarkable differences between amino acid types emerge from this comparative study. The current approach is expected to lead towards a better understanding of functional site architecture in proteins.

Amino Acid Sequence

Conserved clusters of functionally related genes in two bacterial genomes.

An approach for genome comparison, combining function classification of gene products and sequence comparison, is presented. The genomes of Haemophilus influenzae and Escherichia coli are analyzed, and all genes are classified into nine major functional classes, corresponding to important cellular processes. To study gene order relationships and genome organization in the two bacteria, we performed statistics on neighboring pairs of genes. To estimate the significance of the observations, a statistical model based on binomial distributions has been developed. Significant patterns of gene order are observed within, as well as between, the two bacterial genomes: Functionally related genes tend to be neighbors more often than do unrelated genes. Some of these groups represent well-known operons, but additional gene clusters are identified. These clusters correspond to genomic elements that have been conserved during bacterial evolution. In addition to nearest-neighbor relationships, the method is also useful to study the relative direction of transcription in genomes, which is also highly conserved between homologous gene pairs. This new approach combines the high-level description of molecular function with pair statistics that express genome organization. It is expected to complement traditional methods of sequence analysis in the study of genomic structure, function, and evolution.

Binomial Distribution

Evolution of immunoglobulin-like modules in chitinases: their structural flexibility and functional implications.

BACKGROUND: Chitinase A from Serratia marcescens is a glycosyl hydrolase consisting of three distinct domains. The N-terminal domain (ChiN domain, amino acids 24-137) has an immunoglobulin-like fold. This ChiN domain is structurally similar to fibronectin type III domains (FnIII domains), which exist in other chitinases, but does not share any sequence similarity with them. RESULT: Structure comparisons of the ChiN domain and FnIII domains confirm the similar fold, but fail to establish any sequence similarity. Sequence searches and comparisons between ChiN and FnIII domain sequences show a remarkable difference between the two domains in chitinases from an evolutionary point of view. A low temperature structure of chitinase A shows that the ChiN module is flexible with respect to the catalytic body of the protein. CONCLUSIONS: We postulate that the ChiN and FnIII domains evolved independently in chitinases which share otherwise homologous catalytic domains. The flexibility of the ChiN domain, together with biochemical knowledge of the function of similar domains, leads us to propose that immunoglobulin-like folds in chitinases are involved in interactions with the chitin chain during catalysis.

Bacterial Proteins

The emergence of major cellular processes in evolution.

The phylogenetic distribution of divergently related protein families into the three domains of life (archaea, bacteria and eukaryotes) can signify the presence or absence of entire cellular processes in these domains and their ancestors. We can thus study the emergence of the major transitions during cellular evolution, and resolve some of the controversies surrounding the evolutionary status of archaea and the origins of the eukaryotic cell. In view of the ongoing projects that sequence the complete genomes of several Archaea, this work forms a testable prediction when the genome sequences become available. Using the presence of the protein families as taxonomic traits, and linking them to biochemical pathways, we are able to reason about the presence of the corresponding cellular processes in the last universal ancestor of contemporary cells. The analysis shows that metabolism was already a complex network of reactions which included amino acid, nucleotide, fatty acid, sugar and coenzyme metabolism. In addition, genetic processes such as translation are conserved and close to the original form. However, other processes such as DNA replication and repair or transcription are exceptional and seem to be associated with the structural changes that drove eukaryotes and bacteria away from their common ancestor. There are two major hypotheses in the present work: first, that archaea are probably closer to the last universal ancestor than any other extant life form, and second, that the major cellular processes were in place before the major splitting. The last universal ancestor had metabolism and translation very similar to the contemporary ones, while having an operonic genome organization and archaean-like transcription. Evidently, all cells today contain remnants of the primordial genome of the last universal ancestor.

Archaea

Genomes with distinct function composition.

The functional composition of organisms can be analysed for the first time with the appearance of complete or sizeable parts of various genomes. We have reduced the problem of protein function classification to a simple scheme with three classes of protein function: energy-, information- and communication-associated proteins. Finer classification schemes can be easily mapped to the above three classes. To deal with the vast amount of information, a system for automatic function classification using database annotations has been developed. The system is able to classify correctly about 80% of the query sequences with annotations. Using this system, we can analyse samples from the genomes of the most represented species in sequence databases and compare their genomic composition. The similarities and differences for different taxonomic groups are strikingly intuitive. Viruses have the highest proportion of proteins involved in the control and expression of genetic information. Bacteria have the highest proportion of their genes dedicated to the production of proteins associated with small molecule transformations and transport. Animals have a very large proportion of proteins associated with intra- and intercellular communication and other regulatory processes. In general, the proportion of communication-related proteins increases during evolution, indicating trends that led to the emergence of the eukaryotic cell and later the transition from unicellular to multicellular organisms.

Animals

Computational comparisons of model genomes.

Complete genomes from model organisms provide new challenges for computational molecular biology. Novel questions emerge from the genome data obtained from the functional prediction of thousands of gene products. In this review, we present some approaches to the computational comparison of genomes, based on sequence and text analysis, and comparisons of genome composition and gene order.

Biotechnology

HinCyc: a knowledge base of the complete genome and metabolic pathways of H. influenzae.

We present a methodology for predicting the metabolic pathways of an organism from its genomic sequence by reference to a knowledge base of known metabolic pathways. We applied these techniques to the genome of H. influenzae by reference to the EcoCyc knowledge base to predict which of 81 metabolic pathways of E. coli are found in H. influenzae. The resulting prediction is a complex hypothesis that is presented in computer form as HinCyc: an electronic encyclopedia of the genes and metabolic pathways of H. influenzae. HinCyc connects the predicted genes, enzymes, enzyme-catalyzed reactions, and biochemical pathways in a WWW-accessible knowledge base to allow scientists to explore this complex hypothesis.

Computer Communication Networks

New protein functions in yeast chromosome VIII.

The analysis of the 269 open reading frames of yeast chromosome VIII by computational methods has yielded 24 new significant sequence similarities to proteins of known function. The resulting predicted functions include three particularly interesting cases of translation-associated proteins: peptidyl-tRNA hydrolase, a ribosome recycling factor homologue, and a protein similar to cytochrome b translational activator CBS2. The methodological limits of the meaningful transfer of functional information between distant homologues are discussed.

Alcohol Oxidoreductases

Exploring the Mycoplasma capricolum genome: a minimal cell reveals its physiology.

We report on the analysis of 214kb of the parasitic eubacterium Mycoplasma capricolum sequenced by genomic walking techniques. The 287 putative proteins detected to date represent about half of the estimated total number of 500 predicted for this organism. A large fraction of these (75%) can be assigned a likely function as a result of similarity searches. Several important features of the functional organization of this small genome are already apparent. Among these are (i) the expected relatively large number of enzymes involved in metabolic transport and activation, for efficient use of host cell nutrients; (ii) the presence of anabolic enzymes; (iii) the unexpected diversity of enzymes involved in DNA replication and repair; and (iv) a sizeable number of orthologues (82 so far) in Escherichia coli. This survey is beginning to provide a detailed view of how M. capricolum manages to maintain essential cellular processes with a genome much smaller than that of its bacterial relatives.

Amino Acid Sequence

Barley beta-glucosidase: expression during seed germination and maturation and partial amino acid sequences.

Unlike most of the hydrolytic enzymes that participate in endosperm mobilization, beta-glucosidase of barley (Hordeum vulgare) seeds does not increase during germination, even in the presence of exogenously added gibberellic acid. However, the germination process affects the physical properties of beta-glucosidase in terms of charge and apparent molecular weight. Analysis of developing barley grains shows that the enzyme is synthesized two weeks before maturation and is stored in the endosperm of the dry dormant seed. Partial amino acid sequencing of the purified beta-glucosidase demonstrates significant similarity between the barley enzyme and beta-glycosidases that belong to family 1 of glycosyl hydrolases.

Amino Acid Sequence

GeneQuiz: a workbench for sequence analysis.

We present the prototype of a software system, called GeneQuiz, for large-scale biological sequence analysis. The system was designed to meet the needs that arise in computational sequence analysis and our past experience with the analysis of 171 protein sequences of yeast chromosome III. We explain the cognitive challenges associated with this particular research activity and present our model of the sequence analysis process. The prototype system consists of two parts: (i) the database update and search system (driven by perl programs and rdb, a simple relational database engine also written in perl) and (ii) the visualization and browsing system (developed under C++/ET++). The principal design requirement for the first part was the complete automation of all repetitive actions: database updates, efficient sequence similarity searches and sampling of results in a uniform fashion. The user is then presented with "hit-lists" that summarize the results from heterogeneous database searches. The expert's primary task now simply becomes the further analysis of the candidate entries, where the problem is to extract adequate information about functional characteristics of the query protein rapidly. This second task is tremendously accelerated by a simple combination of the heterogeneous output into uniform relational tables and the provision of browsing mechanisms that give access to database records, sequence entries and alignment views. Indexing of molecular sequence databases provides fast retrieval of individual entries with the use of unique identifiers as well as browsing through databases using pre-existing cross-references. The presentation here covers an overview of the architecture of the system prototype and our experiences on its applicability in sequence analysis.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals