Search PubMed⌕ Search

Biomedical subjects

Gordon M Crippen

Publications and source records attributed to Gordon M Crippen.

13 recordsLinked to original sources

An iterative refinement algorithm for consistency based multiple structural alignment methods.

MOTIVATION: Multiple STructural Alignment (MSTA) provides valuable information for solving problems such as fold recognition. The consistency-based approach tries to find conflict-free subsets of alignments from a pre-computed all-to-all Pairwise Alignment Library (PAL). If large proportions of conflicts exist in the library, consistency can be hard to get. On the other hand, multiple structural superposition has been used in many MSTA methods to refine alignments. However, multiple structural superposition is dependent on alignments, and a superposition generated based on erroneous alignments is not guaranteed to be the optimal superposition. Correcting errors after making errors is not as good as avoiding errors from the beginning. Hence it is important to refine the pairwise library to reduce the number of conflicts before any consistency-based assembly. RESULTS: We present an algorithm, Iterative Refinement of Induced Structural alignment (IRIS), to refine the PAL. A new measurement for the consistency of a library is also proposed. Experiments show that our algorithm can greatly improve T-COFFEE performance for less consistent pairwise alignment libraries. The final multiple alignment outperforms most state-of-the-art MSTA algorithms at assembling 15 transglycosidases. Results on three other benchmarks showed that the algorithm consistently improves multiple alignment performance. AVAILABILITY: The C++ code of the algorithm is available upon request.

Algorithms↗

Fold recognition via a tree.

Recently, we developed a pairwise structural alignment algorithm using realistic structural and environmental information (SAUCE). In this paper, we at first present an automatic fold hierarchical classification based on SAUCE alignments. This classification enables us to build a fold tree containing different levels of multiple structural profiles. Then a tree-based fold search algorithm is described. We applied this method to a group of structures with sequence identity less than 35% and did a series of leave one out tests. These tests are approximately comparable to fold recognition tests on superfamily level. Results show that fold recognition via a fold tree can be faster and better at detecting distant homologues than classic fold recognition methods.

Algorithms↗

A novel approach to structural alignment using realistic structural and environmental information.

In the era of structural genomics, it is necessary to generate accurate structural alignments in order to build good templates for homology modeling. Although a great number of structural alignment algorithms have been developed, most of them ignore intermolecular interactions during the alignment procedure. Therefore, structures in different oligomeric states are barely distinguishable, and it is very challenging to find correct alignment in coil regions. Here we present a novel approach to structural alignment using a clique finding algorithm and environmental information (SAUCE). In this approach, we build the alignment based on not only structural coordinate information but also realistic environmental information extracted from biological unit files provided by the Protein Data Bank (PDB). At first, we eliminate all environmentally unfavorable pairings of residues. Then we identify alignments in core regions via a maximal clique finding algorithm. Two extreme value distribution (EVD) form statistics have been developed to evaluate core region alignments. With an optional extension step, global alignment can be derived based on environment-based dynamic programming linking. We show that our method is able to differentiate three-dimensional structures in different oligomeric states, and is able to find flexible alignments between multidomain structures without predetermined hinge regions. The overall performance is also evaluated on a large scale by comparisons to current structural classification databases as well as to other alignment methods.

Algorithms↗

Recognizing protein folds by cluster distance geometry.

Cluster distance geometry is a recent generalization of distance geometry whereby protein structures can be described at even lower levels of detail than one point per residue. With improvements in the clustering technique, protein conformations can be summarized in terms of alternative contact patterns between clusters, where each cluster contains four sequentially adjacent amino acid residues. A very simple potential function involving 210 adjustable parameters can be determined that favors the native contacts of 31 small, monomeric proteins over their respective sets of nonnative contacts. This potential then favors the native contacts for 174 small, monomeric proteins that have low sequence identity with any of the training set. A broader search finds 698 small protein chains from the Protein Data Bank where the native contacts are preferred over all alternatives, even though they have low sequence identity with the training set. This amounts to a highly predictive method for ab initio protein folding at low spatial resolution.

Algorithms↗

CASA: an efficient automated assignment of protein mainchain NMR data using an ordered tree search algorithm.

Rapid analysis of protein structure, interaction, and dynamics requires fast and automated assignments of 3D protein backbone triple-resonance NMR spectra. We introduce a new depth-first ordered tree search method of automated assignment, CASA, which uses hand-edited peak-pick lists of a flexible number of triple resonance experiments. The computer program was tested on 13 artificially simulated peak lists for proteins up to 723 residues, as well as on the experimental data for four proteins. Under reasonable tolerances, it generated assignments that correspond to the ones reported in the literature within a few minutes of CPU time. The program was also tested on the proteins analyzed by other methods, with both simulated and experimental peaklists, and it could generate good assignments in all relevant cases. The robustness was further tested under various situations.

Algorithms↗

Statistical mechanics of protein folding by cluster distance geometry.

This is our second type of model for protein folding where the configurational parameters and the effective potential energy function are chosen in such a way that all conformations are described and the canonical partition function can be evaluated analytically. Structure is described in terms of distances between pairs of sequentially contiguous blocks of eight residues, and all possible conformations are grouped into 71 subsets in terms of bounds on these distances. The energy is taken to be a sum of pairwise interactions between such blocks. The 210 energy parameters were adjusted so that the native folds of 32 small proteins are favored in free energy over the denatured state. We then found 146 proteins having negligible sequence similarity to any of the training proteins, yet the free energy of the respective correct native states were favored over the denatured state.

Algorithms↗

Cluster distance geometry of polypeptide chains.

Distance geometry has been a broadly useful tool for dealing with conformational calculations. Customarily each atom is represented as a point, constraints on the distances between some atoms are obtained from experimental or theoretical sources, and then a random sampling of conformations can be calculated that are consistent with the constraints. Although these methods can be applied to small proteins having on the order of 1000 atoms, for some purposes it is advantageous to view the problem at lower resolution. Here distance geometry is generalized to deal with distances between sets of points. In the end, much of the same techniques produce a sampling of different configurations of these sets of points subject to distance constraints, but now the radii of gyration of the different sets play an important role. A simple example is given of how the packing constraints for polypeptide chains combine with loose distance constraints to give good calculated protein conformers at a very low resolution.

Algorithms↗

Statistical mechanics of protein folding with separable energy functions.

We have initiated an entirely new approach to statistical mechanical models of strongly interacting systems where the configurational parameters and the potential energy function are both constructed so that the canonical partition function can be evaluated analytically. For a simplified model of proteins consisting of a single, fairly short polypeptide chain without cross-links, we can adjust the energy parameters to favor the experimentally determined native state of seven proteins having diverse types of folds. Then 497 test proteins are predicted to have stable native folds, even though they are also structurally diverse, and 480 of them have no significant sequence similarity to any of the training proteins.

Models, Theoretical↗

How to describe chirality and conformational flexibility.

Given atomic coordinates for a particular conformation of a molecule and some property value assigned to each atom, one can easily calculate a chirality function that distinguishes enantiomers, is zero for an achiral molecule, and is a continuous function of the coordinates and properties. This is useful as a quantitative measure of chirality for molecular modeling and structure-activity relations.

Molecular Conformation↗

A protein folding potential that places the native states of a large number of proteins near a local minimum.

BACKGROUND: We present a simple method to train a potential function for the protein folding problem which, even though trained using a small number of proteins, is able to place a significantly large number of native conformations near a local minimum. The training relies on generating decoys by energy minimization of the native conformations using the current potential and using a physically meaningful objective function (derivative of energy with respect to torsion angles at the native conformation) during the quadratic programming to place the native conformation near a local minimum. RESULTS: We also compare the performance of three different types of energy functions and find that while the pairwise energy function is trainable, a solvation energy function by itself is untrainable if decoys are generated by minimizing the current potential starting at the native conformation. The best results are obtained when a pairwise interaction energy function is used with solvation energy function. CONCLUSIONS: We are able to train a potential function using six proteins which places a total of 42 native conformations within approximately 4 A rmsd and 71 native conformations within approximately 6 A rmsd of a local minimum out of a total of 91 proteins. Furthermore, the threading test using the same 91 proteins ranks 89 native conformations to be first and the other two as second.

Models, Molecular↗

Three-dimensional molecular descriptors and a novel QSAR method.

A novel set of molecular descriptors suitable for use in quantitative structure-activity relationships and related methods is described. These descriptors are a smooth and interpretable representation of atomic physicochemical property values and intramolecular atom pair distances. Distance atomic physicochemical parameter energy relationships (DAPPER), a novel structure-activity relationship (QSAR) method using these descriptors, is validated on standard datasets.

Chemical Phenomena↗

Validation of DAPPER for 3D QSAR: conformational search and chirality metric.

Adequate conformational searching of small molecules and inclusion of a chirality identifier are necessary features of any current technique for quantitative structure-activity relationships (QSAR). However, implementation of these features can be difficult and computationally expensive, and some techniques can still lead to insufficient treatment of molecular conformation. We select the standard systematic conformational search as the default search method for our recent 3D QSAR program, DAPPER, and develop a novel chirality metric for use in QSAR. These techniques are implemented in DAPPER and validated on standard data sets.

Enzyme Inhibitors↗

Structure and specificity of a human valacyclovir activating enzyme: a homology model of BPHL.

Biphenyl hydrolase-like (BPHL) protein is a novel serine hydrolase which has been identified as human valacyclovirase (VACVase), catalyzing the hydrolytic activation of valine ester prodrugs of the antiviral drugs acyclovir and ganciclovir as well as other amino acid ester prodrugs of therapeutic nucleoside analogues. The broad specificity for nucleoside analogues as parent drugs suggests that BPHL may be particularly useful as a molecular target for prodrug activation. In order to develop an initial structural view of the specificity of BPHL, a homology model of BPHL based on the crystal structure of 2-hydroxy-6-oxo-7-methylocta-2,4-dienoate hydrolase was developed using the Molecular Operating Environment package (Chemical Computing Group, Montreal, Quebec), evaluated for its stereochemical quality and identification of free cysteines, and used in a molecular docking study. The BPHL model has residues S122, H255, and D227 comprising the putative catalytic triad in proximity and potential charge-charge interaction sites, M52 or D123 for the alpha-amino group. The model also suggested that the structural preference of BPHL for hydrophobic amino acyl promoieties and its limited activity for the secondary alcohol substrates may be attributed to the hydrophobic acyl-binding site formed by residues I158, G161, I162, and L229, and the spatial constraint around the catalytic site by a loop on one side, the active serine and histidine on the other side, and L53 and L179 on top. In addition, the broad specificity for nucleoside analogues may be due to the relatively less constrained nucleoside-binding site opening toward the entrance of the substrate-binding pocket. The homology model of BPHL provides a basis for further investigation of the catalytic and active site residues, can account for the observed structure activity profile of BPHL, and will be useful in the design of nucleoside prodrugs.

Acyclovir↗