Search PubMed⌕ Search

Biomedical subjects

Zaida Luthey-Schulten

Publications and source records attributed to Zaida Luthey-Schulten.

11 recordsLinked to original sources

MultiSeq: unifying sequence and structure data for evolutionary analysis.

BACKGROUND: Since the publication of the first draft of the human genome in 2000, bioinformatic data have been accumulating at an overwhelming pace. Currently, more than 3 million sequences and 35 thousand structures of proteins and nucleic acids are available in public databases. Finding correlations in and between these data to answer critical research questions is extremely challenging. This problem needs to be approached from several directions: information science to organize and search the data; information visualization to assist in recognizing correlations; mathematics to formulate statistical inferences; and biology to analyze chemical and physical properties in terms of sequence and structure changes. RESULTS: Here we present MultiSeq, a unified bioinformatics analysis environment that allows one to organize, display, align and analyze both sequence and structure data for proteins and nucleic acids. While special emphasis is placed on analyzing the data within the framework of evolutionary biology, the environment is also flexible enough to accommodate other usage patterns. The evolutionary approach is supported by the use of predefined metadata, adherence to standard ontological mappings, and the ability for the user to adjust these classifications using an electronic notebook. MultiSeq contains a new algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of a homologous group of distantly related proteins. The method, based on the multidimensional QR factorization of multiple sequence and structure alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. CONCLUSION: MultiSeq is a major extension of the Multiple Alignment tool that is provided as part of VMD, a structural visualization program for analyzing molecular dynamics simulations. Both are freely distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics and MultiSeq is included with VMD starting with version 1.8.5. The MultiSeq website has details on how to download and use the software: http://www.scs.uiuc.edu/~schulten/multiseq/

Algorithms↗

Visualizing the dual space of biological molecules.

An important part of protein structure characterization is the determination of excluded space such as fissures in contact interfaces, pores, inaccessible cavities, and catalytic pockets. We introduce a general tessellation method for visualizing the dual space around, within, and between biological molecules. Using Delaunay triangulation, a three-dimensional graph is constructed to provide a displayable discretization of the continuous volume. This graph structure is also used to compare the dual space of a system in two different states. Tessellator, a cross-platform implementation of the algorithm, is used to analyze the cavities within myoglobin, the protein-RNA docking interface between aspartyl-tRNA synthetase and tRNA(Asp), and the ammonia channel in the hisH-hisF complex of imidazole glycerol phosphate synthase.

Algorithms↗

Multiple Alignment of protein structures and sequences for VMD.

Multiple Alignment is a new interface for performing and analyzing multiple protein structure alignments. It enables viewing levels of sequence and structure similarity on the aligned structures and performing a variety of evolutionary and bioinformatic tasks, including the construction of structure-based phylogenetic trees and minimal basis sets of structures that best represent the topology of the phylogenetic tree. It is implemented as a plugin for VMD (Visual Molecular Dynamics), which is distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics at the University of Illinois.

Amino Acid Sequence↗

Evolutionary profiles from the QR factorization of multiple sequence alignments.

We present an algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of the homologous group. The method, based on the multidimensional QR factorization of numerically encoded multiple sequence alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. We observe a general trend that these smaller, more evolutionarily balanced profiles have comparable and, in many cases, better performance in database searches than conventional profiles containing hundreds of sequences, constructed in an iterative and computationally intensive procedure. For more diverse families or superfamilies, with sequence identity <30%, structural alignments, based purely on the geometry of the protein structures, provide better alignments than pure sequence-based methods. Merging the structure and sequence information allows the construction of accurate profiles for distantly related groups. These structure-based profiles outperformed other sequence-based methods for finding distant homologs and were used to identify a putative class II cysteinyl-tRNA synthetase (CysRS) in several archaea that eluded previous annotation studies. Phylogenetic analysis showed the putative class II CysRSs to be a monophyletic group and homology modeling revealed a constellation of active site residues similar to that in the known class I CysRS.

Algorithms↗

Evolutionary profiles derived from the QR factorization of multiple structural alignments gives an economy of information.

We present a new algorithm, based on the multidimensional QR factorization, to remove redundancy from a multiple structural alignment by choosing representative protein structures that best preserve the phylogenetic tree topology of the homologous group. The classical QR factorization with pivoting, developed as a fast numerical solution to eigenvalue and linear least-squares problems of the form Ax=b, was designed to re-order the columns of A by increasing linear dependence. Removing the most linear dependent columns from A leads to the formation of a minimal basis set which well spans the phase space of the problem at hand. By recasting the problem of redundancy in multiple structural alignments into this framework, in which the matrix A now describes the multiple alignment, we adapted the QR factorization to produce a minimal basis set of protein structures which best spans the evolutionary (phase) space. The non-redundant and representative profiles obtained from this procedure, termed evolutionary profiles, are shown in initial results to outperform well-tested profiles in homology detection searches over a large sequence database. A measure of structural similarity between homologous proteins, Q(H), is presented. By properly accounting for the effect and presence of gaps, a phylogenetic tree computed using this metric is shown to be congruent with the maximum-likelihood sequence-based phylogeny. The results indicate that evolutionary information is indeed recoverable from the comparative analysis of protein structure alone. Applications of the QR ordering and this structural similarity metric to analyze the evolution of structure among key, universally distributed proteins involved in translation, and to the selection of representatives from an ensemble of NMR structures are also discussed.

Algorithms↗

Water in protein structure prediction.

Proteins have evolved to use water to help guide folding. A physically motivated, nonpairwise-additive model of water-mediated interactions added to a protein structure prediction Hamiltonian yields marked improvement in the quality of structure prediction for larger proteins. Free energy profile analysis suggests that long-range water-mediated potentials guide folding and smooth the underlying folding funnel. Analyzing simulation trajectories gives direct evidence that water-mediated interactions facilitate native-like packing of supersecondary structural elements. Long-range pairing of hydrophilic groups is an integral part of protein architecture. Specific water-mediated interactions are a universal feature of biomolecular recognition landscapes in both folding and binding.

Computational Biology↗

Classical force field parameters for the heme prosthetic group of cytochrome c.

Accurate force fields are essential for describing biological systems in a molecular dynamics simulation. To analyze the docking of the small redox protein cytochrome c (cyt c) requires simulation parameters for the heme in both the reduced and oxidized states. This work presents parameters for the partial charges and geometries for the heme in both redox states with ligands appropriate to cyt c. The parameters are based on both protein X-ray structures and ab initio density functional theory (DFT) geometry optimizations at the B3LYP/6-31G* level. The simulations with the new parameter set reproduce the geometries of the X-ray structures and the interaction energies between water and heme prosthetic group obtained from B3LYP/6-31G* calculations. The parameter set developed here will provide new insights into docking processes of heme containing redox proteins.

Algorithms↗

Variations in the fast folding rates of the lambda-repressor: a hybrid molecular dynamics study.

The ability to predict the effects of mutations on protein folding rates and mechanisms would greatly facilitate folding studies. Using a realistic full atom potential coupled with a Gō-like potential biased to the native state structure, we have investigated the effects of point mutations on the folding rates of a small single domain protein. The hybrid potential provides a detailed level of description of the folding mechanism that we correlate to features of the folding energy landscapes of fast and slow mutants of an 80-residue-long fragment of the lambda-repressor. Our computational reconstruction of the folding events is compared to the recent experimental results of W. Y. Yang and M. Gruebele (see companion article) and T. G. Oas and co-workers on the lambda-repressor, and helps to clarify the differences observed in the folding mechanisms of the various mutants.

Algorithms↗

Developing an energy landscape for the novel function of a (beta/alpha)8 barrel: ammonia conduction through HisF.

HisH-hisF is a multidomain globular protein complex; hisH is a class I glutamine amidotransferase that hydrolyzes glutamine to form ammonia, and hisF is a (beta/alpha)8 barrel cyclase that completes the ring formation of imidizole glycerol phosphate synthase. Together, hisH and hisF form a glutamine amidotransferase that carries out the fifth step of the histidine biosynthetic pathway. Recently, it has been suggested that the (beta/alpha)8 barrel participates in a novel function: to channel ammonia from the active site of hisH to the active site of hisF. The present study presents a series of molecular dynamic simulations that investigate the channeling function of hisF. This article reconstructs potentials of mean force for the conduction of ammonia through the channel, and the entrance of ammonia through the strictly conserved channel gate, in both a closed and a hypothetical open conformation. The resulting energy landscape within the channel supports the idea that ammonia does indeed pass through the barrel, interacting with conserved hydrophilic residues along the way. The proposed open conformation, which involves an alternate rotamer state of one of the gate residues, presents only an approximately 2.5-kcal energy barrier to ammonia entry. Another alternate open-gate conformation, which may play a role in non-nitrogen-fixing organisms, is deduced through bioinformatics.

Algorithms↗

On the evolution of structure in aminoacyl-tRNA synthetases.

The aminoacyl-tRNA synthetases are one of the major protein components in the translation machinery. These essential proteins are found in all forms of life and are responsible for charging their cognate tRNAs with the correct amino acid. The evolution of the tRNA synthetases is of fundamental importance with respect to the nature of the biological cell and the transition from an RNA world to the modern world dominated by protein-enzymes. We present a structure-based phylogeny of the aminoacyl-tRNA synthetases. By using structural alignments of all of the aminoacyl-tRNA synthetases of known structure in combination with a new measure of structural homology, we have reconstructed the evolutionary history of these proteins. In order to derive unbiased statistics from the structural alignments, we introduce a multidimensional QR factorization which produces a nonredundant set of structures. Since protein structure is more highly conserved than protein sequence, this study has allowed us to glimpse the evolution of protein structure that predates the root of the universal phylogenetic tree. The extensive sequence-based phylogenetic analysis of the tRNA synthetases (Woese et al., Microbiol. Mol. Biol. Rev. 64:202-236, 2000) has further enabled us to reconstruct the complete evolutionary profile of these proteins and to make connections between major evolutionary events and the resulting changes in protein shape. We also discuss the effect of functional specificity on protein shape over the complex evolutionary course of the tRNA synthetases.

Amino Acid Sequence↗

Ab initio protein structure prediction.

Steady progress has been made in the field of ab initio protein folding. A variety of methods now allow the prediction of low-resolution structures of small proteins or protein fragments up to approximately 100 amino acid residues in length. Such low-resolution structures may be sufficient for the functional annotation of protein sequences on a genome-wide scale. Although no consistently reliable algorithm is currently available, the essential challenges to developing a general theory or approach to protein structure prediction are better understood. The energy landscapes resulting from the structure prediction algorithms are only partially funneled to the native state of the protein. This review focuses on two areas of recent advances in ab initio structure prediction-improvements in the energy functions and strategies to search the caldera region of the energy landscapes.

Chemistry, Physical↗