Search PubMed⌕ Search

Biomedical subjects

Eugene I Shakhnovich

Publications and source records attributed to Eugene I Shakhnovich.

At least 19 recordsLinked to original sources

A knowledge-based move set for protein folding.

The free energy landscape of protein folding is rugged, occasionally characterized by compact, intermediate states of low free energy. In computational folding, this landscape leads to trapped, compact states with incorrect secondary structure. We devised a residue-specific, protein backbone move set for efficient sampling of protein-like conformations in computational folding simulations. The move set is based on the selection of a small set of backbone dihedral angles, derived from clustering dihedral angles sampled from experimental structures. We show in both simulated annealing and replica exchange Monte Carlo (REMC) simulations that the knowledge-based move set, when compared with a conventional move set, shows statistically significant improved ability at overcoming kinetic barriers, reaching deeper energy minima, and achieving correspondingly lower RMSDs to native structures. The new move set is also more efficient, being able to reach low energy states considerably faster. Use of this move set in determining the energy minimum state and for calculating thermodynamic quantities is discussed.

Glycine↗

Description of atomic burials in compact globular proteins by Fermi-Dirac probability distributions.

We perform a statistical analysis of atomic distributions as a function of the distance R from the molecular geometrical center in a nonredundant set of compact globular proteins. The number of atoms increases quadratically for small R, indicating a constant average density inside the core, reaches a maximum at a size-dependent distance R(max), and falls rapidly for larger R. The empirical curves turn out to be consistent with the volume increase of spherical concentric solid shells and a Fermi-Dirac distribution in which the distance R plays the role of an effective atomic energy epsilon(R) = R. The effective chemical potential mu governing the distribution increases with the number of residues, reflecting the size of the protein globule, while the temperature parameter beta decreases. Interestingly, betamu is not as strongly dependent on protein size and appears to be tuned to maintain approximately half of the atoms in the high density interior and the other half in the exterior region of rapidly decreasing density. A normalized size-independent distribution was obtained for the atomic probability as a function of the reduced distance, r = R/R(g), where R(g) is the radius of gyration. The global normalized Fermi distribution, F(r), can be reasonably decomposed in Fermi-like subdistributions for different atomic types tau, F(tau)(r), with Sigma(tau)F(tau)(r) = F(r), which depend on two additional parameters mu(tau) and h(tau). The chemical potential mu(tau) affects a scaling prefactor and depends on the overall frequency of the corresponding atomic type, while the maximum position of the subdistribution is determined by h(tau), which appears in a type-dependent atomic effective energy, epsilon(tau)(r) = h(tau)r, and is strongly correlated to available hydrophobicity scales. Better adjustments are obtained when the effective energy is not assumed to be necessarily linear, or epsilon(tau)*(r) = h(tau)*r(alpha,), in which case a correlation with hydrophobicity scales is found for the product alpha(tau)h(tau)*. These results indicate that compact globular proteins are consistent with a thermodynamic system governed by hydrophobic-like energy functions, with reduced distances from the geometrical center, reflecting atomic burials, and provide a conceptual framework for the eventual prediction from sequence of a few parameters from which whole atomic probability distributions and potentials of mean force can be reconstructed.

Amino Acids↗

A structure-centric view of protein evolution, design, and adaptation.

Proteins, by virtue of their central role in most biological processes, represent one of the key subjects of the study of molecular evolution. Inherent in the indispensability of proteins for living cells is the fact that a given protein can adopt a specific three-dimensional shape that is specified solely by the protein's sequence of amino acids. Over the past several decades, structural biologists have demonstrated that the array of structures that proteins may adopt is quite astounding, and this has lead to a strong interest in understanding how protein structures change and evolve over time. In this review we consider a large body of recent work that attempts to illuminate this structure-centric picture of protein evolution. Much of this work has focused on the question of how completely new protein structures (i.e., new folds or topologies) are discovered by protein sequences as they evolve. Pursuant to this question of structural innovation has been a desire to describe and understand the observation that certain types of protein structures are far more abundant than others and how this uneven distribution of proteins implicates on the process through which new shapes are discovered. We consider a number of theoretical models that have been successful at explaining this heterogeneity in protein populations and discuss the increasing amount of evidence that indicates that the process of structural evolution involves the divergence of protein sequences and structures from one another. We also consider the topic of protein designability, which concerns itself with understanding how a protein's structure influences the number of sequences that can fold successfully into that structure. Understanding and quantifying the relationship between the physical feature of a structure and its designability has been a long-standing goal of the study of protein structure and evolution, and we discuss a number of recent advances that have yielded a promising answer to this question. Finally, we review the relatively new field of protein structural phylogeny, an area of study in which information about the distribution of protein structures among different organisms is used to reconstruct the evolutionary relationships between them. Taken together, the work that we review presents an increasingly coherent picture of how these unique polymers have evolved over the course of life on Earth.

Adaptation, Biological↗

All-atom ab initio folding of a diverse set of proteins.

Natural proteins fold to a unique, thermodynamically dominant state. Modeling of the folding process and prediction of the native fold of proteins are two major unsolved problems in biophysics. Here, we show successful all-atom ab initio folding of a representative diverse set of proteins by using a minimalist transferable-energy model that consists of two-body atom-atom interactions, hydrogen bonding, and a local sequence-energy term that models sequence-specific chain stiffness. Starting from a random coil, the native-like structure was observed during replica exchange Monte Carlo (REMC) simulation for most proteins regardless of their structural classes; the lowest energy structure was close to native-in the range of 2-6 A root-mean-square deviation (rmsd). Our results demonstrate that the successful folding of a protein chain to its native state is governed by only a few crucial energetic terms.

Models, Molecular↗

Protein and DNA sequence determinants of thermophilic adaptation.

There have been considerable attempts in the past to relate phenotypic trait--habitat temperature of organisms--to their genotypes, most importantly compositions of their genomes and proteomes. However, despite accumulation of anecdotal evidence, an exact and conclusive relationship between the former and the latter has been elusive. We present an exhaustive study of the relationship between amino acid composition of proteomes, nucleotide composition of DNA, and optimal growth temperature (OGT) of prokaryotes. Based on 204 complete proteomes of archaea and bacteria spanning the temperature range from -10 degrees C to 110 degrees C, we performed an exhaustive enumeration of all possible sets of amino acids and found a set of amino acids whose total fraction in a proteome is correlated, to a remarkable extent, with the OGT. The universal set is Ile, Val, Tyr, Trp, Arg, Glu, Leu (IVYWREL), and the correlation coefficient is as high as 0.93. We also found that the G + C content in 204 complete genomes does not exhibit a significant correlation with OGT (R = -0.10). On the other hand, the fraction of A + G in coding DNA is correlated with temperature, to a considerable extent, due to codon patterns of IVYWREL amino acids. Further, we found strong and independent correlation between OGT and the frequency with which pairs of A and G nucleotides appear as nearest neighbors in genome sequences. This adaptation is achieved via codon bias. These findings present a direct link between principles of proteins structure and stability and evolutionary mechanisms of thermophylic adaptation. On the nucleotide level, the analysis provides an example of how nature utilizes codon bias for evolutionary adaptation to extreme conditions. Together these results provide a complete picture of how compositions of proteomes and genomes in prokaryotes adjust to the extreme conditions of the environment.

Adaptation, Physiological↗

Understanding ensemble protein folding at atomic detail.

It has long been known that a protein's amino acid sequence dictates its native structure. However, despite significant recent advances, an ensemble description of how a protein achieves its native conformation from random coil under physiologically relevant conditions remains incomplete. Here we present a detailed all-atom model with a transferable potential that is capable of ab initio folding of entire protein domains using only sequence information. The computational efficiency of this model allows us to perform thousands of microsecond-time scale-folding simulations of the engrailed homeodomain and to observe thousands of complete independent folding events. We apply a graph-theoretic analysis to this massive data set to elucidate which intermediates and intermediary states are common to many trajectories and thus important for the folding process. This method provides an atomically detailed and complete picture of a folding pathway at the ensemble level. The approach that we describe is quite general and could be used to study the folding of proteins on time scales orders of magnitude longer than currently possible.

Algorithms↗

Divergent evolution of a structural proteome: phenomenological models.

We develop models of the divergent evolution of genomes; the elementary object of sequence dynamics is the protein structural domain. To identify patterns of organization that reflect mechanisms of evolution, we consider the individual genomes of many procaryote species, studying the arrangement of protein structural domains in the space of all polypeptide structures. We view the network of structural similarities as a graph, called the organismal Protein Domain Universe Graph (oPDUG); vertices represent types of structural domains and edges represent strong structural similarity. As observed before, each oPDUG is a highly nonrandom graph, as evidenced in the vertex degree distribution, which resembles a Pareto law (which has a power-law asymptotic). To explain this and other peculiar properties of the oPDUGs, we construct an evolving-graph model for the long-timescale evolutionary dynamics of oPDUGs, containing only divergent mechanisms of domain discovery. The model generates degree distributions (resembling Pareto laws) and clustering-coefficient distributions that are characteristic of the oPDUGs. In the infinite-graph limit, we analytically compute the exponent for specific biological parameters, as well as the complete phase diagram of the model, finding two distinct regimes of domain innovation dynamics. Thus, divergent evolutionary dynamics quantitatively explains the nonrandom organization of oPDUGs.

Bacterial Proteins↗

Common motifs and topological effects in the protein folding transition state.

Through extensive experiment, simulation, and analysis of protein S6 (1RIS), we find that variations in nucleation and folding pathway between circular permutations are determined principally by the restraints of topology and specific nucleation, and affected by changes in chain entropy. Simulations also relate topological features to experimentally measured stabilities. Despite many sizable changes in phi values and the structure of the transition state ensemble that result from permutation, we observe a common theme: the critical nucleus in each of the mutants share a subset of residues that can be mapped to the critical nucleus residues of the wild-type. Circular permutations create new N and C termini, which are the location of the largest disruption of the folding nucleus, leading to a decrease in both phi values and the role in nucleation. Mutant nuclei are built around the wild-type nucleus but are biased towards different parts of the S6 structure depending on the topological and entropic changes induced by the location of the new N and C termini.

Amino Acid Motifs↗

Semiconservative quasispecies equations for polysomic genomes: the haploid case.

This paper develops the semiconservative quasispecies equations for genomes consisting of an arbitrary number of chromosomes. We assume that the chromosomes are distinguishable, so that we are effectively considering haploid genomes. We derive the quasispecies equations under the assumption of arbitrary lesion repair efficiency, and consider the cases of both random and immortal strand chromosome segregation. We solve the model in the limit of infinite sequence length for the case of the static single fitness peak landscape, where the master genome has a first-order growth rate constant of k>1, and all other genomes have a first-order growth rate constant of 1. If we assume that each chromosome can tolerate an arbitrary number of lesions, so that only one master copy of the strands is necessary for a functional chromosome, then for random chromosome segregation we obtain an equilibrium mean fitness of [equation in text] below the error catastrophe, while for immortal strand co-segregation we obtain kappa (t=infinity)=k[e(-mu(1-lambda/2))+e(-mulambda/2)-1] (N denotes the number of chromosomes, lambda denotes the lesion repair efficiency, and mu is identical with epsilonL, where epsilon is the per base-pair mismatch probability, and L is the total genome length). It follows that immortal strand co-segregation leads to significantly better preservation of the master genome than random segregation when lesion repair is imperfect. Based on this result, we conjecture that certain classes of tumor cells exhibit immortal strand co-segregation.

Animals↗

Identification of the minimal protein-folding nucleus through loop-entropy perturbations.

To explore the plasticity and structural constraints of the protein-folding nucleus we have constructed through circular permutation four topological variants of the ribosomal protein S6. In effect, these topological variants represent entropy mutants with maintained spatial contacts. The proteins were characterized at two complementary levels of detail: by phi-value analysis estimating the extent of contact formation in the transition-state ensemble and by Hammond analysis measuring the site-specific growth of the folding nucleus. The results show that, although the loop-entropy alterations markedly influence the appearance and structural location of the folding nucleus, it retains a common motif of one helix docking against two strands. This nucleation motif is built around a shared subset of side chains in the center of the hydrophobic core but extends in different directions of the S6 structure following the permutant-specific differences in local loop entropies. The adjustment of the critical folding nucleus to alterations in loop entropies is reflected by a direct correlation between the phi-value change and the accompanying change in local sequence separation.

Biophysical Phenomena↗

Physical origins of protein superfamilies.

In this work, we discovered a fundamental connection between selection for protein stability and emergence of preferred structures of proteins. Using a standard exact three-dimensional lattice model we evolve sequences starting from random ones and determine the exact native structure after each mutation. Acceptance of mutations is biased to select for stable proteins. We found that certain structures, "wonderfolds", are independently discovered numerous times as native states of stable proteins in many unrelated runs of selection. The strong dependence of lattice fold usage on the structural determinant of designability quantitatively reproduces uneven fold usage in natural proteins. Diversity of sequences that fold into wonderfold structures gives rise to superfamilies, i.e. sets of dissimilar sequences that fold into the same or very similar structures. The present work establishes a model of pre-biotic structure selection, which identifies dominant structural patterns emerging upon optimization of proteins for survival in a hot environment. Convergently discovered pre-biotic initial superfamilies with wonderfold structures could have served as a seed for subsequent biological evolution involving gene duplications and divergence.

Amino Acid Sequence↗

PDB-UF: database of predicted enzymatic functions for unannotated protein structures from structural genomics.

BACKGROUND: The number of protein structures from structural genomics centers dramatically increases in the Protein Data Bank (PDB). Many of these structures are functionally unannotated because they have no sequence similarity to proteins of known function. However, it is possible to successfully infer function using only structural similarity. RESULTS: Here we present the PDB-UF database, a web-accessible collection of predictions of enzymatic properties using structure-function relationship. The assignments were conducted for three-dimensional protein structures of unknown function that come from structural genomics initiatives. We show that 4 hypothetical proteins (with PDB accession codes: 1VH0, 1NS5, 1O6D, and 1TO0), for which standard BLAST tools such as PSI-BLAST or RPS-BLAST failed to assign any function, are probably methyltransferase enzymes. CONCLUSION: We suggest that the structure-based prediction of an EC number should be conducted having the different similarity score cutoff for different protein folds. Moreover, performing the annotation using two different algorithms can reduce the rate of false positive assignments. We believe, that the presented web-based repository will help to decrease the number of protein structures that have functions marked as "unknown" in the PDB file. AVAILABILITY: http://paradox.harvard.edu/PDB-UF and http://bioinfo.pl/PDB-UF.

Chromosome Mapping↗

Genetic instability and the quasispecies model.

Genetic instability is a defining characteristic of cancers. Microsatellite instability (MIN) leads to by elevated point mutation rates, whereas chromosomal instability (CIN) refers to increased rates of losing or gaining whole chromosomes or parts of chromosomes during cell division. CIN and MIN are, in general, mutually exclusive. The quasispecies model is a very successful theoretical framework for the study of evolution at high mutation rates. It predicts the existence of an experimentally verified error catastrophe. This catastrophe occurs when the mutation rates exceed a threshold value, the error threshold, above which replicative infidelity is incompatible with cell survival. We analyse the semiconservative quasispecies model of both MIN and CIN tumors. We consider the role of post-methylation DNA repair in tumor cells and demonstrate that DNA repair is fundamental to the nature of the error catastrophe in both types of tumors. We find that CIN introduces a plateau in the maximum viable mutation rate for a repair-free model, which does not exist in the case of MIN. This provides a plausible explanation for the mutual exclusivity of CIN and MIN.

Chromosomal Instability↗

A simple physical model for scaling in protein-protein interaction networks.

It has recently been demonstrated that many biological networks exhibit a "scale-free" topology, for which the probability of observing a node with a certain number of edges (k) follows a power law: i.e., p(k) approximately k(-gamma). This observation has been reproduced by evolutionary models. Here we consider the network of protein-protein interactions (PPIs) and demonstrate that two published independent measurements of these interactions produce graphs that are only weakly correlated with one another despite their strikingly similar topology. We then propose a physical model based on the fundamental principle that (de)solvation is a major physical factor in PPIs. This model reproduces not only the scale-free nature of such graphs but also a number of higher-order correlations in these networks. A key support of the model is provided by the discovery of a significant correlation between the number of interactions made by a protein and the fraction of hydrophobic residues on its surface. The model presented in this paper represents a physical model for experimentally determined PPIs that comprehensively reproduces the topological features of interaction networks. These results have profound implications for understanding not only PPIs but also other types of scale-free networks.

Biophysical Phenomena↗

High-resolution protein folding with a transferable potential.

A generalized computational method for folding proteins with a fully transferable potential and geometrically realistic all-atom model is presented and tested on seven helix bundle proteins. The protocol, which includes graph-theoretical analysis of the ensemble of resulting folded conformations, was systematically applied and consistently produced structure predictions of approximately 3 A without any knowledge of the native state. To measure and understand the significance of the results, extensive control simulations were conducted. Graph theoretic analysis provides a means for systematically identifying the native fold and provides physical insight, conceptually linking the results to modern theoretical views of protein folding. In addition to presenting a method for prediction of structure and folding mechanism, our model suggests that an accurate all-atom amino acid representation coupled with a physically reasonable atomic interaction potential and hydrogen bonding are essential features for a realistic protein model.

Animals↗

Entropic stabilization of proteins and its proteomic consequences.

Evolutionary traces of thermophilic adaptation are manifest, on the whole-genome level, in compositional biases toward certain types of amino acids. However, it is sometimes difficult to discern their causes without a clear understanding of underlying physical mechanisms of thermal stabilization of proteins. For example, it is well-known that hyperthermophiles feature a greater proportion of charged residues, but, surprisingly, the excess of positively charged residues is almost entirely due to lysines but not arginines in the majority of hyperthermophilic genomes. All-atom simulations show that lysines have a much greater number of accessible rotamers than arginines of similar degree of burial in folded states of proteins. This finding suggests that lysines would preferentially entropically stabilize the native state. Indeed, we show in computational experiments that arginine-to-lysine amino acid substitutions result in noticeable stabilization of proteins. We then hypothesize that if evolution uses this physical mechanism as a complement to electrostatic stabilization in its strategies of thermophilic adaptation, then hyperthermostable organisms would have much greater content of lysines in their proteomes than comparably sized and similarly charged arginines. Consistent with that, high-throughput comparative analysis of complete proteomes shows extremely strong bias toward arginine-to-lysine replacement in hyperthermophilic organisms and overall much greater content of lysines than arginines in hyperthermophiles. This finding cannot be explained by genomic GC compositional biases or by the universal trend of amino acid gain and loss in protein evolution. We discovered here a novel entropic mechanism of protein thermostability due to residual dynamics of rotamer isomerization in native state and demonstrated its immediate proteomic implications. Our study provides an example of how analysis of a fundamental physical mechanism of thermostability helps to resolve a puzzle in comparative genomics as to why amino acid compositions of hyperthermophilic proteomes are significantly biased toward lysines but not similarly charged arginines.

Aminopeptidases↗

Physics and evolution of thermophilic adaptation.

Analysis of structures and sequences of several hyperthermostable proteins from various sources reveals two major physical mechanisms of their thermostabilization. The first mechanism is "structure-based," whereby some hyperthermostable proteins are significantly more compact than their mesophilic homologues, while no particular interaction type appears to cause stabilization; rather, a sheer number of interactions is responsible for thermostability. Other hyperthermostable proteins employ an alternative, "sequence-based" mechanism of their thermal stabilization. They do not show pronounced structural differences from mesophilic homologues. Rather, a small number of apparently strong interactions is responsible for high thermal stability of these proteins. High-throughput comparative analysis of structures and complete genomes of several hyperthermophilic archaea and bacteria revealed that organisms develop diverse strategies of thermophilic adaptation by using, to a varying degree, two fundamental physical mechanisms of thermostability. The choice of a particular strategy depends on the evolutionary history of an organism. Proteins from organisms that originated in an extreme environment, such as hyperthermophilic archaea (Pyrococcus furiosus), are significantly more compact and more hydrophobic than their mesophilic counterparts. Alternatively, organisms that evolved as mesophiles but later recolonized a hot environment (Thermotoga maritima) relied in their evolutionary strategy of thermophilic adaptation on "sequence-based" mechanism of thermostability. We propose an evolutionary explanation of these differences based on physical concepts of protein designability.

Acclimatization↗