Search PubMed⌕ Search

Biomedical subjects

Patrice Koehl

Publications and source records attributed to Patrice Koehl.

16 recordsLinked to original sources

PDB_Hydro: incorporating dipolar solvents with variable density in the Poisson-Boltzmann treatment of macromolecule electrostatics.

We describe a new way to calculate the electrostatic properties of macromolecules which eliminates the assumption of a constant dielectric value in the solvent region, resulting in a Generalized Poisson-Boltzmann-Langevin equation (GPBLE). We have implemented a web server (http://lorentz.immstr.pasteur.fr/pdb_hydro.php) that both numerically solves this equation and uses the resulting water density profiles to place water molecules at preferred sites of hydration. Surface atoms with high or low hydration preference can be easily displayed using a simple PyMol script, allowing for the tentative prediction of the dimerization interface in homodimeric proteins, or lipid binding regions in membrane proteins. The web site includes options that permit mutations in the sequence as well as reconstruction of missing side chain and/or main chain atoms. These tools are accessible independently from the electrostatics calculation, and can be used for other modeling purposes. We expect this web server to be useful to structural biologists, as the knowledge of solvent density should prove useful to get better fits at low resolution for X-ray diffraction data and to computational biologists, for whom these profiles could improve the calculation of interaction energies in water between ligands and receptors in docking simulations.

Binding Sites↗

NOMAD-Ref: visualization, deformation and refinement of macromolecular structures based on all-atom normal mode analysis.

Normal mode analysis (NMA) is an efficient way to study collective motions in biomolecules that bypasses the computational costs and many limitations associated with full dynamics simulations. The NOMAD-Ref web server presented here provides tools for online calculation of the normal modes of large molecules (up to 100,000 atoms) maintaining a full all-atom representation of their structures, as well as access to a number of programs that utilize these collective motions for deformation and refinement of biomolecular structures. Applications include the generation of sets of decoys with correct stereochemistry but arbitrary large amplitude movements, the quantification of the overlap between alternative conformations of a molecule, refinement of structures against experimental data, such as X-ray diffraction structure factors or Cryo-EM maps and optimization of docked complexes by modeling receptor/ligand flexibility through normal mode motions. The server can be accessed at the URL http://lorentz.immstr.pasteur.fr/nomad-ref.php.

Computer Graphics↗

Plant NBS-LRR proteins: adaptable guards.

The majority of disease resistance genes in plants encode nucleotide-binding site leucine-rich repeat (NBS-LRR) proteins. This large family is encoded by hundreds of diverse genes per genome and can be subdivided into the functionally distinct TIR-domain-containing (TNL) and CC-domain-containing (CNL) subfamilies. Their precise role in recognition is unknown; however, they are thought to monitor the status of plant proteins that are targeted by pathogen effectors.

Amino Acid Sequence↗

Electrostatics calculations: latest methodological advances.

Electrostatics plays a major role in the stabilization and function of biomolecules; as such, it remains a major focus of theoretical and computational studies of macromolecules. Electrostatic interactions are long range, and strongly dependent on the solvent and ions surrounding the biomolecule under study. During the past year, progress has been reported in the treatment of electrostatics in explicit and implicit solvent models. Interesting new developments of explicit solvent models include more efficient Ewald summation methods, as well as alternative approaches based on reaction field theory, periodic images and Euler summations. Implicit solvent models remain divided into those that solve the Poisson-Boltzmann equation numerically and those based on the generalized Born formalism. Both approaches are now included in molecular dynamics simulations and their accuracies may be assessed by direct comparison against experimental data. It is worth mentioning the recent development of web interfaces that facilitate access to and usage of existing tools for computing electrostatic interactions.

Algorithms↗

BAliBASE 3.0: latest developments of the multiple sequence alignment benchmark.

Multiple sequence alignment is one of the cornerstones of modern molecular biology. It is used to identify conserved motifs, to determine protein domains, in 2D/3D structure prediction by homology and in evolutionary studies. Recently, high-throughput technologies such as genome sequencing and structural proteomics have lead to an explosion in the amount of sequence and structure information available. In response, several new multiple alignment methods have been developed that improve both the efficiency and the quality of protein alignments. Consequently, the benchmarks used to evaluate and compare these methods must also evolve. We present here the latest release of the most widely used multiple alignment benchmark, BAliBASE, which provides high quality, manually refined, reference alignments based on 3D structural superpositions. Version 3.0 of BAliBASE includes new, more challenging test cases, representing the real problems encountered when aligning large sets of complex sequences. Using a novel, semiautomatic update protocol, the number of protein families in the benchmark has been increased and representative test cases are now available that cover most of the protein fold space. The total number of proteins in BAliBASE has also been significantly increased from 1444 to 6255 sequences. In addition, full-length sequences are now provided for all test cases, which represent difficult cases for both global and local alignment programs. Finally, the BAliBASE Web site (http://www-bio3d-igbmc.u-strasbg.fr/balibase) has been completely redesigned to provide a more user-friendly, interactive interface for the visualization of the BAliBASE reference alignments and the associated annotations.

Amino Acid Sequence↗

MAO: a Multiple Alignment Ontology for nucleic acid and protein sequences.

The application of high-throughput techniques such as genomics, proteomics or transcriptomics means that vast amounts of heterogeneous data are now available in the public databases. Bioinformatics is responding to the challenge with new integrated management systems for data collection, validation and analysis. Multiple alignments of genomic and protein sequences provide an ideal environment for the integration of this mass of information. In the context of the sequence family, structural and functional data can be evaluated and propagated from known to unknown sequences. However, effective integration is being hindered by syntactic and semantic differences between the different data resources and the alignment techniques employed. One solution to this problem is the development of an ontology that systematically defines the terms used in a specific domain. Ontologies are used to share data from different resources, to automatically analyse information and to represent domain knowledge for non-experts. Here, we present MAO, a new ontology for multiple alignments of nucleic and protein sequences. MAO is designed to improve interoperation and data sharing between different alignment protocols for the construction of a high quality, reliable multiple alignment in order to facilitate knowledge extraction and the presentation of the most pertinent information to the biologist.

Databases, Genetic↗

Comprehensive evaluation of protein structure alignment methods: scoring by geometric measures.

We report the largest and most comprehensive comparison of protein structural alignment methods. Specifically, we evaluate six publicly available structure alignment programs: SSAP, STRUCTAL, DALI, LSQMAN, CE and SSM by aligning all 8,581,970 protein structure pairs in a test set of 2930 protein domains specially selected from CATH v.2.4 to ensure sequence diversity. We consider an alignment good if it matches many residues, and the two substructures are geometrically similar. Even with this definition, evaluating structural alignment methods is not straightforward. At first, we compared the rates of true and false positives using receiver operating characteristic (ROC) curves with the CATH classification taken as a gold standard. This proved unsatisfactory in that the quality of the alignments is not taken into account: sometimes a method that finds less good alignments scores better than a method that finds better alignments. We correct this intrinsic limitation by using four different geometric match measures (SI, MI, SAS, and GSAS) to evaluate the quality of each structural alignment. With this improved analysis we show that there is a wide variation in the performance of different methods; the main reason for this is that it can be difficult to find a good structural alignment between two proteins even when such an alignment exists. We find that STRUCTAL and SSM perform best, followed by LSQMAN and CE. Our focus on the intrinsic quality of each alignment allows us to propose a new method, called "Best-of-All" that combines the best results of all methods. Many commonly used methods miss 10-50% of the good Best-of-All alignments. By putting existing structural alignments into proper perspective, our study allows better comparison of protein structures. By highlighting limitations of existing methods, it will spur the further development of better structural alignment methods. This will have significant biological implications now that structural comparison has come to play a central role in the analysis of experimental work on protein structure, protein function and protein evolution.

Computational Biology↗

Relaxed specificity in aromatic prenyltransferases.

Prenylation represent a critical step in the biosynthesis of many natural products, A new study reveals how aromatic prenyltransferase enzymes tolerate diverse aromatic polyketides while still controlling the length of prenyl side chains.

Dimethylallyltranstransferase↗

A new lectin family with structure similarity to actinoporins revealed by the crystal structure of Xerocomus chrysenteron lectin XCL.

A newly defined family of fungal lectins displays no significant sequence similarity to any protein in the databases. These proteins, made of about 140 amino acid residues, have sequence identities ranging from 38% to 65% and share binding specificity to N-acetyl galactosamine. One member of this family, the lectin XCL from Xerocomus chrysenteron, induces drastic changes in the actin cytoskeleton after sugar binding at the cell surface and internalization, and has potent insecticidal activity. The crystal structure of XCL to 1.4 A resolution reveals the architecture of this new lectin family. The fold of the protein is not related to any of the several lectin folds documented so far. Unexpectedly, the structure similarity is significant with actinoporins, a family of pore-forming toxins. The specific structural features and sequence signatures in each protein family suggest a potential sugar binding site in XCL and a possible evolutionary relationship between these proteins. Finally, the tetrameric assembly of XCL reveals a complex network of protomer-protomer interfaces and generates a large, hydrated cavity of 1000 A3, which may become accessible to larger solutes after a small conformational change of the protein.

Amino Acid Sequence↗

The ASTRAL Compendium in 2004.

The ASTRAL Compendium provides several databases and tools to aid in the analysis of protein structures, particularly through the use of their sequences. Partially derived from the SCOP database of protein structure domains, it includes sequences for each domain and other resources useful for studying these sequences and domain structures. The current release of ASTRAL contains 54,745 domains, more than three times as many as the initial release 4 years ago. ASTRAL has undergone major transformations in the past 2 years. In addition to several complete updates each year, ASTRAL is now updated on a weekly basis with preliminary classifications of domains from newly released PDB structures. These classifications are available as a stand-alone database, as well as integrated into other ASTRAL databases such as representative subsets. To enhance the utility of ASTRAL to structural biologists, all SCOP domains are now made available as PDB-style coordinate files as well as sequences. In addition to sequences and representative subsets based on SCOP domains, sequences and subsets based on PDB chains are newly included in ASTRAL. Several search tools have been added to ASTRAL to facilitate retrieval of data by individual users and automated methods. ASTRAL may be accessed at http://astral.stanford. edu/.

Animals↗

The weighted-volume derivative of a space-filling diagram.

Computing the volume occupied by individual atoms in macromolecular structures has been the subject of research for several decades. This interest has grown in the recent years, because weighted volumes are widely used in implicit solvent models. Applications of the latter in molecular mechanics simulations require that the derivatives of these weighted volumes be known. In this article, we give a formula for the volume derivative of a molecule modeled as a space-filling diagram made up of balls in motion. The formula is given in terms of the weights, radii, and distances between the centers as well as the sizes of the facets of the power diagram restricted to the space-filling diagram. Special attention is given to the detection and treatment of singularities as well as discontinuities of the derivative.

Biophysical Phenomena↗

Sequence variations within protein families are linearly related to structural variations.

It is commonly believed that similarities between the sequences of two proteins infer similarities between their structures. Sequence alignments reliably recognize pairs of protein of similar structures provided that the percentage sequence identity between their two sequences is sufficiently high. This distinction, however, is statistically less reliable when the percentage sequence identity is lower than 30% and little is known then about the detailed relationship between the two measures of similarity. Here, we investigate the inverse correlation between structural similarity and sequence similarity on 12 protein structure families. We define the structure similarity between two proteins as the cRMS distance between their structures. The sequence similarity for a pair of proteins is measured as the mean distance between the sequences in the subsets of sequence space compatible with their structures. We obtain an approximation of the sequence space compatible with a protein by designing a collection of protein sequences both stable and specific to the structure of that protein. Using these measures of sequence and structure similarities, we find that structural changes within a protein family are linearly related to changes in sequence similarity.

Amino Acid Sequence↗

Small libraries of protein fragments model native protein structures accurately.

Prediction of protein structure depends on the accuracy and complexity of the models used. Here, we represent the polypeptide chain by a sequence of rigid fragments that are concatenated without any degrees of freedom. Fragments chosen from a library of representative fragments are fit to the native structure using a greedy build-up method. This gives a one-dimensional representation of native protein three-dimensional structure whose quality depends on the nature of the library. We use a novel clustering method to construct libraries that differ in the fragment length (four to seven residues) and number of representative fragments they contain (25-300). Each library is characterized by the quality of fit (accuracy) and the number of allowed states per residue (complexity). We find that the accuracy depends on the complexity and varies from 2.9A for a 2.7-state model on the basis of fragments of length 7-0.76A for a 15-state model on the basis of fragments of length 5. Our goal is to find representations that are both accurate and economical (low complexity). The models defined here are substantially better in this regard: with ten states per residue we approximate native protein structure to 1A compared to over 20 states per residue needed previously. For the same complexity, we find that longer fragments provide better fits. Unfortunately, libraries of longer fragments must be much larger (for ten states per residue, a seven-residue library is 100 times larger than a five-residue library). As the number of known protein native structures increases, it will be possible to construct larger libraries to better exploit this correlation between neighboring residues. Our fragment libraries, which offer a wide range of optimal fragments suited to different accuracies of fit, may prove to be useful for generating better decoy sets for ab initio protein folding and for generating accurate loop conformations in homology modeling.

Models, Molecular↗

Protein topology and stability define the space of allowed sequences.

We describe a new approach to explore and quantify the sequence space associated with a given protein structure. A set of sequences are optimized for a given target structure, using all-atom models and a physical energy function. Specificity of the sequence for its target is ensured by using the random energy model, which keeps the amino acid composition of the sequence constant. The designed sequences provide a multiple sequence alignment that describes the sequence space compatible with the structure of interest; here the size of this space is estimated by using an information entropy measure. In parallel, multiple alignments of naturally occurring sequences can be derived by using either sequence or structure alignments. We compared these 3 independent multiple sequence alignments for 10 different proteins, ranging in size from 56 to 310 residues. We observed that the subset of the sequence space derived by using our design procedure is similar in size to the sequence spaces observed in nature. These results suggest that the volume of sequence space compatible with a given protein fold is defined by the length of the protein as well as by the topology (i.e., geometry of the polypeptide chain) and the stability (i.e., free energy of denaturation) of the fold.

Amino Acid Sequence↗

Improved recognition of native-like protein structures using a family of designed sequences.

The goal of the inverse protein folding problem is to identify amino acid sequences that stabilize a given target protein conformation. Methods that attempt to solve this problem have proven useful for protein sequence design. Here we show that the same methods can provide valuable information for protein fold recognition and for ab initio protein structure prediction. We present a measure of the compatibility of a test sequence with a target model structure, based on computational protein design. The model structure is used as input to design a family of low free energy sequences, and these sequences are compared with the test sequence by using a metric in sequence space based on nearest-neighbor connectivity. We find that this measure is able to recognize the native fold of a myoglobin sequence among different globin folds. It is also powerful enough to recognize near-native protein structures among non-native models.

Amino Acid Sequence↗

ASTRAL compendium enhancements.

The ASTRAL compendium provides several databases and tools to aid in the analysis of protein structures, particularly through the use of their sequences. It is partially derived from the SCOP database of protein domains, and it includes sequences for each domain as well as other resources useful for studying these sequences and domain structures. Several major improvements have been made to the ASTRAL compendium since its initial release 2 years ago. The number of protein domain sequences included has doubled from 15 190 to 30 867, and additional databases have been added. The Rapid Access Format (RAF) database contains manually curated mappings linking the biological amino acid sequences described in the SEQRES records of PDB entries to the amino acid sequences structurally observed (provided in the ATOM records) in a format designed for rapid access by automated tools. This information is used to derive sequences for protein domains in the SCOP database. In cases where a SCOP domain spans several protein chains, all of which can be traced back to a single genetic source, a 'genetic domain' sequence is created by concatenating the sequences of each chain in the order found in the original gene sequence. Both the original-style library of SCOP sequences and a new library including genetic domain sequences are available. Selected representative subsets of each of these libraries, based on multiple criteria and degrees of similarity, are also included. ASTRAL may be accessed at http://astral.stanford.edu/.

Amino Acid Sequence↗