Search PubMed⌕ Search

Biomedical subjects

Christophe Geourjon

Publications and source records attributed to Christophe Geourjon.

17 recordsLinked to original sources

euHCVdb: the European hepatitis C virus database.

The hepatitis C virus (HCV) genome shows remarkable sequence variability, leading to the classification of at least six major genotypes, numerous subtypes and a myriad of quasispecies within a given host. A database allowing researchers to investigate the genetic and structural variability of all available HCV sequences is an essential tool for studies on the molecular virology and pathogenesis of hepatitis C as well as drug design and vaccine development. We describe here the European Hepatitis C Virus Database (euHCVdb, http://euhcvdb.ibcp.fr), a collection of computer-annotated sequences based on reference genomes. The annotations include genome mapping of sequences, use of recommended nomenclature, subtyping as well as three-dimensional (3D) molecular models of proteins. A WWW interface has been developed to facilitate database searches and the export of data for sequence and structure analyses. As part of an international collaborative effort with the US and Japanese databases, the European HCV Database (euHCVdb) is mainly dedicated to HCV protein sequences, 3D structures and functional analyses.

Databases, Protein↗

Insights into early extracellular matrix evolution: spongin short chain collagen-related proteins are homologous to basement membrane type IV collagens and form a novel family widely distributed in invertebrates.

Collagens are thought to represent one of the most important molecular innovations in the metazoan line. Basement membrane type IV collagen is present in all Eumetazoa and was found in Homoscleromorpha, a sponge group with a well-organized epithelium, which may represent the first stage of tissue differentiation during animal evolution. In contrast, spongin seems to be a demosponge-specific collagenous protein, which can totally substitute an inorganic skeleton, such as in the well-known bath sponge. In the freshwater sponge Ephydatia mülleri, we previously characterized a family of short-chain collagens that are likely to be main components of spongins. Using a combination of sequence- and structure-based methods, we present evidence of remote homology between the carboxyl-terminal noncollagenous NC1 domain of spongin short-chain collagens and type IV collagen. Unexpectedly, spongin short-chain collagen-related proteins were retrieved in nonsponge animals, suggesting that a family related to spongin constitutes an evolutionary sister to the type IV collagen family. Formation of the ancestral NC1 domain and divergence of the spongin short-chain collagen-related and type IV collagen families may have occurred before the parazoan-eumetazoan split, the earliest divergence among extant animal phyla. Molecular phylogenetics based on NC1 domain sequences suggest distinct evolutionary histories for spongin short-chain collagen-related and type IV collagen families that include spongin short-chain collagen-related gene loss in the ancestors of Ecdyzosoa and of vertebrates. The fact that a majority of invertebrates encodes spongin short-chain collagen-related proteins raises the important question to the possible function of its members. Considering the importance of collagens for animal structure and substratum attachment, both families may have played crucial roles in animal diversification.

Amino Acid Sequence↗

Staphylococcus aureus operates protein-tyrosine phosphorylation through a specific mechanism.

Protein phosphorylation on tyrosine has been originally characterized in animal systems and has been shown to be involved in several fundamental processes including signal transduction, growth control, and malignancy. It has been later demonstrated to occur also in a number of bacteria, and recent data suggest that it may participate in the control of bacterial pathogenicity. In this work, we provide evidence that the gram-positive human pathogen Staphylococcus aureus harbors a protein-tyrosine kinase activity. This activity is borne by a protein, termed Cap5B2, whose phosphorylating capacity is expressed only in the presence of a stimulatory protein, either Cap5A1 or Cap5A2, that enhances its affinity for the phosphoryl donor ATP. In fact, the last 27/29 amino acids of the C-terminal domain of either polypeptide are sufficient for stimulating Cap5B2 activity. The stimulation of Cap5B2 by Cap5A1 involves essentially three amino acid residues in a helix of Cap5A1 (Asp202, Glu203, and Asp205) and three residues in a helix (helix 7) of Cap5B2 (Glu190, Lys192, and Lys193), thus suggesting helix-helix interaction between these two proteins. This type of helix-helix interaction resembles the interaction required for the activation of MinD ATPase by MinE protein in the process of septum-site determination, MinD sharing sequence similarity with Cap5B2. Such activation mechanism is described here in a gram-positive bacterial tyrosine kinase, and differs from the activation mechanism previously proposed for gram-negative bacteria. Therefore, it appears that S. aureus, and possibly other gram-positive bacteria, utilizes a specific molecular mechanism for triggering protein-tyrosine kinase activity.

Amino Acid Sequence↗

Hepatitis C databases, principles and utility to researchers.

Part of the effort to develop hepatitis C-specific drugs a nd vaccines is the study of genetic variability of allpublicly available HCV sequences. Three HCV databases are currently available to aid this effort and to provide additional insight into the basic biology, immunology, and evolution of the virus. The Japanese HCV database (http://s2as02.genes.nig.ac.jp) gives access to a genomic mapping of sequences as well as their phylogenetic relationships. The European HCV database (http://euhcvdb.ibcp.fr) offers access to a computer-annotated set of sequences and molecular models of HCV proteins and focuses on protein sequence, structure and function analysis. The HCV database at the Los Alamos National Laboratory in the United States (http://hcv.lanl.gov) provides access to a manually annotated sequence database and a database of immunological epitopes which contains concise descriptions of experimental results. In this paper, we briefly describe each of these databases and their associated websites and tools, and give some examples of their use in furthering HCV research.

Biomedical Research↗

The SuMo server: 3D search for protein functional sites.

UNLABELLED: We provide the scientific community with a web server which gives access to SuMo, a bioinformatic system for finding similarities in arbitrary 3D structures or substructures of proteins. SuMo is based on a unique representation of macromolecules using selected triplets of chemical groups having their own geometry and symmetry, regardless of the restrictive notions of main chain and lateral chains of amino acids. The heuristic for extracting similar sites was used to drive two major large-scale approaches. First, searching for ligand binding sites onto a query structure has been made possible by comparing the structure against each of the ligand binding sites found in the Protein Data Bank (PDB). Second, the reciprocal process, i.e. searching for a given 3D site of interest among the structures of the PDB is also possible and helps detect cross-reacting targets in drug design projects. AVAILABILITY: The web server is freely accessible to academia through http://sumo-pbil.ibcp.fr and full support is available from MEDIT (http://www.medit.fr). CONTACT: mjambon@burnham.org.

Amino Acid Sequence↗

The Q-loop disengages from the first intracellular loop during the catalytic cycle of the multidrug ABC transporter BmrA.

The ATP-binding cassette is the most abundant family of transporters including many medically relevant members and gathers both importers and exporters involved in the transport of a wide variety of substrates. Although three high resolution three-dimensional structures have been obtained for a prototypic exporter, MsbA, two have been subjected to much criticism. Here, conformational changes of BmrA, a multidrug bacterial transporter structurally related to MsbA, have been studied. A three-dimensional model of BmrA, based on the "open" conformation of Escherichia coli MsbA, was probed by simultaneously introducing two cysteine residues, one in the first intracellular loop of the transmembrane domain and the other in the Q-loop of the nucleotide-binding domain (NBD). Intramolecular disulfide bonds could be created in the absence of any effectors, which prevented both drug transport and ATPase activity. Interestingly, addition of ATP/Mg plus vanadate strongly prevented this bond formation in a cysteine double mutant, whereas ATP/Mg alone was sufficient when the ATPase-inactive E504Q mutation was also introduced, in agreement with additional BmrA models where the ATP-binding sites are positioned at the NBD/NBD interface. Furthermore, cross-linking between the two cysteine residues could still be achieved in the presence of ATP/Mg plus vanadate when homobifunctional cross-linkers separated by more than 13 Angstrom were added. Altogether, these results give support to the existence, in the resting state, of a monomeric conformation of BmrA similar to that found within the open MsbA dimer and show that a large motion is required between intracellular loop 1 and the nucleotide-binding domain for the proper functioning of a multidrug ATP-binding cassette transporter.

ATP-Binding Cassette Transporters↗

GeneFarm, structural and functional annotation of Arabidopsis gene and protein families by a network of experts.

Genomic projects heavily depend on genome annotations and are limited by the current deficiencies in the published predictions of gene structure and function. It follows that, improved annotation will allow better data mining of genomes, and more secure planning and design of experiments. The purpose of the GeneFarm project is to obtain homogeneous, reliable, documented and traceable annotations for Arabidopsis nuclear genes and gene products, and to enter them into an added-value database. This re-annotation project is being performed exhaustively on every member of each gene family. Performing a family-wide annotation makes the task easier and more efficient than a gene-by-gene approach since many features obtained for one gene can be extrapolated to some or all the other genes of a family. A complete annotation procedure based on the most efficient prediction tools available is being used by 16 partner laboratories, each contributing annotated families from its field of expertise. A database, named GeneFarm, and an associated user-friendly interface to query the annotations have been developed. More than 3000 genes distributed over 300 families have been annotated and are available at http://genoplante-info.infobiogen.fr/Genefarm/. Furthermore, collaboration with the Swiss Institute of Bioinformatics is underway to integrate the GeneFarm data into the protein knowledgebase Swiss-Prot.

Arabidopsis↗

HCVDB: hepatitis C virus sequences database.

UNLABELLED: To date, more than 30 000 hepatitis C virus (HCV) sequences have been deposited in the generalist databases DNA Data Bank of Japan (DDBJ), EMBL Nucleotide Sequence Database (EMBL) and GenBank. The main difficulties with HCV sequences in these databases are their retrieval, annotation and analyses. To help HCV researchers face the increasing needs of HCV sequence analyses, we developed a specialised database of computer-annotated HCV sequences, called HCVDB. HCVDB is re-built every month from an up-to-date EMBL database by an automated process. HCVDB provides key data about the HCV sequences (e.g. genotype, genomic region, protein names and functions, known 3-dimensional structures) and ensures consistency of the annotations, which enables reliable keyword queries. The database is highly integrated with sequence and structure analysis tools and the SRS (LION bioscience) keywords query system. Thus, any user can extract subsets of sequences matching particular criteria or enter their own sequences and analyse them with various bioinformatics programs available on the same server. AVAILABILITY: HCVDB is available from http://hepatitis.ibcp.fr.

Amino Acid Sequence↗

A new bioinformatic approach to detect common 3D sites in protein structures.

An innovative bioinformatic method has been designed and implemented to detect similar three-dimensional (3D) sites in proteins. This approach allows the comparison of protein structures or substructures and detects local spatial similarities: this method is completely independent from the amino acid sequence and from the backbone structure. In contrast to already existing tools, the basis for this method is a representation of the protein structure by a set of stereochemical groups that are defined independently from the notion of amino acid. An efficient heuristic for finding similarities that uses graphs of triangles of chemical groups to represent the protein structures has been developed. The implementation of this heuristic constitutes a software named SuMo (Surfing the Molecules), which allows the dynamic definition of chemical groups, the selection of sites in the proteins, and the management and screening of databases. To show the relevance of this approach, we focused on two extreme examples illustrating convergent and divergent evolution. In two unrelated serine proteases, SuMo detects one common site, which corresponds to the catalytic triad. In the legume lectins family composed of >100 structures that share similar sequences and folds but may have lost their ability to bind a carbohydrate molecule, SuMo discriminates between functional and non-functional lectins with a selectivity of 96%. The time needed for searching a given site in a protein structure is typically 0.1 s on a PIII 800MHz/Linux computer; thus, in further studies, SuMo will be used to screen the PDB.

Algorithms↗

Integrated databanks access and sequence/structure analysis services at the PBIL.

The World Wide Web server of the PBIL (Pôle Bioinformatique Lyonnais) provides on-line access to sequence databanks and to many tools of nucleic acid and protein sequence analyses. This server allows to query nucleotide sequence banks in the EMBL and GenBank formats and protein sequence banks in the SWISS-PROT and PIR formats. The query engine on which our data bank access is based is the ACNUC system. It allows the possibility to build complex queries to access functional zones of biological interest and to retrieve large sequence sets. Of special interest are the unique features provided by this system to query the data banks of gene families developed at the PBIL. The server also provides access to a wide range of sequence analysis methods: similarity search programs, multiple alignments, protein structure prediction and multivariate statistics. An originality of this server is the integration of these two aspects: sequence retrieval and sequence analysis. Indeed, thanks to the introduction of re-usable lists, it is possible to perform treatments on large sets of data. The PBIL server can be reached at: http://pbil.univ-lyon1.fr.

Databases, Genetic↗

Detection of unrelated proteins in sequences multiple alignments by using predicted secondary structures.

MOTIVATION: Multiple sequence alignments are essential tools for establishing the homology relations between proteins. Essential amino acids for the function and/or the structure are generally conserved, thus providing key arguments to help in protein characterization. However for distant proteins, it is more difficult to establish, in a reliable way, the homology relations that may exist between them. In this article, we show that secondary structure prediction is a valuable way to validate protein families at low identity rate. RESULTS: We show that the analysis of the secondary structures compatibility is a reliable way to discard non-related proteins in low identity multiple alignment. AVAILABILITY: This validation is possible through our NPS@ server (http://npsa-pbil.ibcp.fr)

Algorithms↗

Conservation of amino acids into multiple alignments involved in pairwise interactions in three-dimensional protein structures.

We present an original strategy, that involves a bioinformatic software structure, in order to perform an exhaustive and objective statistical analysis of three-dimensional structures of proteins. We establish the relationship between multiple sequences alignments and various structural features of proteins. We show that amino acids implied in disulfide bonds, salt bridges and hydrophobic interactions have been studied. Furthermore, we point out that the more variable the sequences within a multiple alignment, the more informative the multiple alignment. The results support multiple alignments usefulness for predictions of structural features.

Amino Acid Sequence↗

Low resolution structure determination shows procollagen C-proteinase enhancer to be an elongated multidomain glycoprotein.

Procollagen C-proteinase enhancer (PCPE) is an extracellular matrix glycoprotein that can stimulate the action of tolloid metalloproteinases, such as bone morphogenetic protein-1, on a procollagen substrate, by up to 20-fold. The PCPE molecule consists of two CUB domains followed by a C-terminal NTR (netrin-like) domain. In order to obtain structural insights into the function of PCPE, the recombinant protein was characterized by a range of biophysical techniques, including analytical ultracentrifugation, transmission electron microscopy, and small angle x-ray scattering. All three approaches showed PCPE to be a rod-like molecule, with a length of approximately 150 A. Homology modeling of both CUB domains and the NTR domain was consistent with the low-resolution structure of PCPE deduced from the small angle x-ray scattering data. Comparison with the low-resolution structure of the procollagen C-terminal region supports a recently proposed model (Ricard-Blum, S., Bernocco, S., Font, B., Moali, C., Eichenberger, D., Farjanel, J., Burchardt, E. R., van der Rest, M., Kessler, E., and Hulmes, D. J. S. (2002) J. Biol. Chem. 277, 33864-33869) for the mechanism of action of PCPE.

Amino Acid Sequence↗

Evidence for crucial electrostatic interactions between Bcl-2 homology domains BH3 and BH4 in the anti-apoptotic Nr-13 protein.

Nr-13 is an anti-apoptotic member of the Bcl-2 family previously shown to interact with Bax. The biological significance of this interaction was explored both in yeast and vertebrate cells and revealed that Nr-13 is able to counteract the pro-apoptotic activity of Bax. The Bax-interacting domain has been identified and corresponds to alpha-helices 5 and 6 in Nr-13. Site-directed mutagenesis has revealed that the N-terminal region of Nr-13 is essential for activity and corresponds to a genuine Bcl-2 homology domain (BH4). The modelling of Nr-13, based on its similarity with other Bcl-2 family proteins and energy minimization, suggests the possibility of electrostatic interactions between the two N-terminal-conserved domains BH4 and BH3. Disruption of these interactions severely affects Nr-13 anti-apoptotic activity. Together our results suggest that electrostatic interactions between BH4 and BH3 domains play a role in the control of activity of Nr-13 and a subset of Bcl-2 family members.

Amino Acid Sequence↗

A new family of phosphotransferases with a P-loop motif.

In most Gram-positive bacteria, catabolite repression is mediated by a bifunctional enzyme, the histidine-containing protein kinase/phosphatase (HprK/P). Based either on its primary sequence or on its recently solved three-dimensional structure, no straightforward homology with other known proteins was found. However, we showed here that HprK/P exhibits a restricted homology with an unrelated phosphotransferase, the phosphoenolpyruvate carboxykinase. This includes notably two consecutive Asp residues from the phosphoenolpyruvate carboxykinase active site, whose equivalent residues were mutated in Bacillus subtilis HprK/P. Characterization of the corresponding mutants emphasizes the crucial role of these Asp residues in the HprK/P functioning. Furthermore, superimposition of HprK/P and phosphoenolpyruvate carboxykinase active sites supports the view that both enzymes bear significant resemblance in their overall mechanism of functioning showing that these two enzymes constitute a new family of phosphotransferases.

Amino Acid Motifs↗

Geno3D: automatic comparative molecular modelling of protein.

Geno3D (http://geno3d-pbil.ibcp.fr) is an automatic web server for protein molecular modelling. Starting with a query protein sequence, the server performs the homology modelling in six successive steps: (i) identify homologous proteins with known 3D structures by using PSI-BLAST; (ii) provide the user all potential templates through a very convenient user interface for target selection; (iii) perform the alignment of both query and subject sequences; (iv) extract geometrical restraints (dihedral angles and distances) for corresponding atoms between the query and the template; (v) perform the 3D construction of the protein by using a distance geometry approach and (vi) finally send the results by e-mail to the user.

Algorithms↗

Selective recognition of enzymatically active prostate-specific antigen (PSA) by anti-PSA monoclonal antibodies.

Prostate-specific antigen (PSA) is widely used as a serum marker for the diagnosis of prostate cancer. To evaluate two anti-free PSA monoclonal antibodies (mAbs) as potential tools in new generations of more relevant PSA assays, we report here their properties towards the recognition of specific forms of free PSA in seminal fluids, LNCaP supernatants, 'non-binding' PSA and sera from cancer patients. PSA from these different origins was immunopurified by the two anti-free PSA mAbs (5D3D11 and 6C8D8) as well as by an anti-total PSA mAb. The composition of the different immunopurified PSA fractions was analysed and their respective enzymatic activities were determined. In seminal fluid, enzymatically active PSA was equally purified with the three mAbs. In LNCaP supernatants and human sera, 5D3D11 immunopurified active PSA mainly, whereas 6C8D8 immunopurified PSA with residual activity. In sera of prostate cancer patients, we identified the presence of a mature inactive PSA form which can be activated into active PSA by use of high saline concentration or capture by an anti-total PSA mAb capable of enhancing PSA activity. According to PSA models built by comparative modelling with the crystal structure of horse prostate kallikrein described previously, we assume that active and activable PSA could correspond to mature intact PSA with open and closed conformations of the kallikrein loop. The specificity of 5D3D11 was restricted to both active and activable PSA, whereas 6C8D8 recognized all free PSA including intact PSA, proforms and internally cleaved PSA.

Amino Acid Sequence↗