Search PubMed⌕ Search

Biomedical subjects

Brian Kuhlman

Publications and source records attributed to Brian Kuhlman.

16 recordsLinked to original sources

Mis-translation of a computationally designed protein yields an exceptionally stable homodimer: implications for protein engineering and evolution.

We recently used computational protein design to create an extremely stable, globular protein, Top7, with a sequence and fold not observed previously in nature. Since Top7 was created in the absence of genetic selection, it provides a rare opportunity to investigate aspects of the cellular protein production and surveillance machinery that are subject to natural selection. Here we show that a portion of the Top7 protein corresponding to the final 49 C-terminal residues is efficiently mis-translated and accumulates at high levels in Escherichia coli. We used circular dichroism, size-exclusion chromatography, small-angle X-ray scattering, analytical ultra-centrifugation, and NMR spectroscopy to show that the resulting C-terminal fragment (CFr) protein adopts a compact, extremely stable, homo-dimeric structure. Based on the solution structure, we engineered an even more stable variant of CFr by disulfide-induced covalent circularisation that should be an excellent platform for design of novel functions. The accumulation of high levels of CFr exposes the high error rate of the protein translation machinery. The rarity of correspondingly stable fragments in natural proteins coupled with the observation that high quality ribosome binding sites are found to occur within E. coli protein-coding regions significantly less often than expected by random chance implies a stringent evolutionary pressure against protein sub-fragments that can independently fold into stable structures. The symmetric self-association between two identical mis-translated CFr sub-domains to generate an extremely stable structure parallels a mechanism for natural protein-fold evolution by modular recombination of protein sub-structures.

Amino Acid Sequence↗

RosettaDesign server for protein design.

The RosettaDesign server identifies low energy amino acid sequences for target protein structures (http://rosettadesign.med.unc.edu). The client provides the backbone coordinates of the target structure and specifies which residues to design. The server returns to the client the sequences, coordinates and energies of the designed proteins. The simulations are performed using the design module of the Rosetta program (RosettaDesign). RosettaDesign uses Monte Carlo optimization with simulated annealing to search for amino acids that pack well on the target structure and satisfy hydrogen bonding potential. RosettaDesign has been experimentally validated and has been used previously to stabilize naturally occurring proteins and design a novel protein structure.

Internet↗

Design of protein conformational switches.

Protein conformational switches are ubiquitous in nature and often regulate key biological processes. To design new proteins that can switch conformation, protein designers have focused on the two key components of protein switches: the amino acid sequence must be compatible with the multiple target states and there must be a mechanism for perturbing the relative stability of these states. Proteins have been designed that can switch between folded and disordered states, between distinct folded states and between different aggregation states. A variety of trigger mechanisms have been used, including pH shifts, post-translational modification and ligand binding. Recently, computational protein design methods have been applied to switch design. These include algorithms for designing novel ligand-binding sites and simultaneously optimizing a sequence for multiple target structures.

Protein Conformation↗

Protein design simulations suggest that side-chain conformational entropy is not a strong determinant of amino acid environmental preferences.

Loss of side-chain conformational entropy is an important force opposing protein folding and the relative preferences of the amino acids for being buried or solvent exposed may be partially determined by which amino acids lose more side-chain entropy when placed in the core of a protein. To investigate these preferences, we have incorporated explicit modeling of side-chain entropy into the protein design algorithm, RosettaDesign. In the standard version of the program, the energy of a particular sequence for a fixed backbone depends only on the lowest energy side-chain conformations that can be identified for that sequence. In the new model, the free energy of a single amino acid sequence is calculated by evaluating the average energy and entropy of an ensemble of structures generated by Monte Carlo sampling of amino acid side-chain conformations. To evaluate the impact of including explicit side-chain entropy, sequences were designed for 110 native protein backbones with and without the entropy model. In general, the differences between the two sets of sequences are modest, with the largest changes being observed for the longer amino acids: methionine and arginine. Overall, the identity between the designed sequences and the native sequences does not increase with the addition of entropy, unlike what is observed when other key terms are added to the model (hydrogen bonding, Lennard-Jones energies, and solvation energies). These results suggest that side-chain conformational entropy has a relatively small role in determining the preferred amino acid at each residue position in a protein.

Amino Acid Sequence↗

Computational design of a single amino acid sequence that can switch between two distinct protein folds.

The functions of many proteins are mediated by specific conformational changes, and therefore the ability to design primary sequences capable of secondary and tertiary changes is an important step toward the creation of novel functional proteins. To this end, we have developed an algorithm that can optimize a single amino acid sequence for multiple target structures. The algorithm consists of an outer loop, in which sequence space is sampled by a Monte Carlo search with simulated annealing, and an inner loop, in which the effect of a given mutation is evaluated on the various target structures by using the rotamer packing routine and composite energy function of the protein design software, RosettaDesign. We have experimentally tested the method by designing a peptide, Sw2, which can be switched from a 2Cys-2His zinc finger-like fold to a trimeric coiled-coil fold, depending upon the pH or the presence of transition metals. Physical characterization of Sw2 confirms that it is able to reversibly adopt each intended target fold.

Algorithms↗

Computer-based design of novel protein structures.

Over the past 10 years there has been tremendous success in the area of computational protein design. Protein design software has been used to stabilize proteins, solubilize membrane proteins, design intermolecular interactions, and design new protein structures. A key motivation for these studies is that they test our understanding of protein energetics and structure. De novo design of novel structures is a particularly rigorous test because the protein backbone must be designed in addition to the amino acid side chains. A priori it is not guaranteed that the target backbone is even designable. To address this issue, researchers have developed a variety of methods for generating protein-like scaffolds and for optimizing the protein backbone in conjunction with the amino acid sequence. These protocols have been used to design proteins from scratch and to explore sequence space for naturally occurring protein folds.

Algorithms↗

E2 conjugating enzymes must disengage from their E1 enzymes before E3-dependent ubiquitin and ubiquitin-like transfer.

During ubiquitin ligation, an E2 conjugating enzyme receives ubiquitin from an E1 enzyme and then interacts with an E3 ligase to modify substrates. Competitive binding experiments with three human E2-E3 protein pairs show that the binding of E1s and of E3s to E2s are mutually exclusive. These results imply that polyubiquitination requires recycling of E2 for addition of successive ubiquitins to substrate.

Binding, Competitive↗

A "solvated rotamer" approach to modeling water-mediated hydrogen bonds at protein-protein interfaces.

Water-mediated hydrogen bonds play critical roles at protein-protein and protein-nucleic acid interfaces, and the interactions formed by discrete water molecules cannot be captured using continuum solvent models. We describe a simple model for the energetics of water-mediated hydrogen bonds, and show that, together with knowledge of the positions of buried water molecules observed in X-ray crystal structures, the model improves the prediction of free-energy changes upon mutation at protein-protein interfaces, and the recovery of native amino acid sequences in protein interface design calculations. We then describe a "solvated rotamer" approach to efficiently predict the positions of water molecules, at protein-protein interfaces and in monomeric proteins, that is compatible with widely used rotamer-based side-chain packing and protein design algorithms. Finally, we examine the extent to which the predicted water molecules can be used to improve prediction of amino acid identities and protein-protein interface stability, and discuss avenues for overcoming current limitations of the approach.

Algorithms↗

An adaptive dynamic programming algorithm for the side chain placement problem.

Larger rotamer libraries, which provide a fine grained discretization of side chain conformation space by sampling near the canonical rotamers, allow protein designers to find better conformations, but slow down the algorithms that search for them. We present a dynamic programming solution to the side chain placement problem which treats rotamers at high or low resolution only as necessary. Dynamic programming is an exact technique; we turn it into an approximation, but can still analyze the error that can be introduced. We have used our algorithm to redesign the surface residues of ubiquitin's beta sheet.

Algorithms↗

Exploring folding free energy landscapes using computational protein design.

Recent advances in computational protein design have allowed exciting new insights into the sequence dependence of protein folding free energy landscapes. Whereas most previous studies have examined the sequence dependence of protein stability and folding kinetics by characterizing naturally occurring proteins and variants of these proteins that contain a small number of mutations, it is now possible to generate and characterize computationally designed proteins that differ significantly from naturally occurring proteins in sequence and/or structure. These computer-generated proteins provide insights into the determinants of protein structure, stability and folding, and make it possible to disentangle the properties of proteins that are the consequence of natural selection from those that reflect the fundamental physical chemistry of polypeptide chains.

Computational Biology↗

Design of a novel globular protein fold with atomic-level accuracy.

A major challenge of computational protein design is the creation of novel proteins with arbitrarily chosen three-dimensional structures. Here, we used a general computational strategy that iterates between sequence design and structure prediction to design a 93-residue alpha/beta protein called Top7 with a novel sequence and topology. Top7 was found experimentally to be folded and extremely stable, and the x-ray crystal structure of Top7 is similar (root mean square deviation equals 1.2 angstroms) to the design model. The ability to design a new protein fold makes possible the exploration of the large regions of the protein universe not yet observed in nature.

Algorithms↗

An improved protein decoy set for testing energy functions for protein structure prediction.

We have improved the original Rosetta centroid/backbone decoy set by increasing the number of proteins and frequency of near native models and by building on sidechains and minimizing clashes. The new set consists of 1,400 model structures for 78 different and diverse protein targets and provides a challenging set for the testing and evaluation of scoring functions. We evaluated the extent to which a variety of all-atom energy functions could identify the native and close-to-native structures in the new decoy sets. Of various implicit solvent models, we found that a solvent-accessible surface area-based solvation provided the best enrichment and discrimination of close-to-native decoys. The combination of this solvation treatment with Lennard Jones terms and the original Rosetta energy provided better enrichment and discrimination than any of the individual terms. The results also highlight the differences in accuracy of NMR and X-ray crystal structures: a large energy gap was observed between native and non-native conformations for X-ray structures but not for NMR structures.

Algorithms↗

A large scale test of computational protein design: folding and stability of nine completely redesigned globular proteins.

A previously developed computer program for protein design, RosettaDesign, was used to predict low free energy sequences for nine naturally occurring protein backbones. RosettaDesign had no knowledge of the naturally occurring sequences and on average 65% of the residues in the designed sequences differ from wild-type. Synthetic genes for ten completely redesigned proteins were generated, and the proteins were expressed, purified, and then characterized using circular dichroism, chemical and temperature denaturation and NMR experiments. Although high-resolution structures have not yet been determined, eight of these proteins appear to be folded and their circular dichroism spectra are similar to those of their wild-type counterparts. Six of the proteins have stabilities equal to or up to 7kcal/mol greater than their wild-type counterparts, and four of the proteins have NMR spectra consistent with a well-packed, rigid structure. These encouraging results indicate that the computational protein design methods can, with significant reliability, identify amino acid sequences compatible with a target protein backbone.

Amino Acid Sequence↗

Protein-protein docking with simultaneous optimization of rigid-body displacement and side-chain conformations.

Protein-protein docking algorithms provide a means to elucidate structural details for presently unknown complexes. Here, we present and evaluate a new method to predict protein-protein complexes from the coordinates of the unbound monomer components. The method employs a low-resolution, rigid-body, Monte Carlo search followed by simultaneous optimization of backbone displacement and side-chain conformations using Monte Carlo minimization. Up to 10(5) independent simulations are carried out, and the resulting "decoys" are ranked using an energy function dominated by van der Waals interactions, an implicit solvation model, and an orientation-dependent hydrogen bonding potential. Top-ranking decoys are clustered to select the final predictions. Small-perturbation studies reveal the formation of binding funnels in 42 of 54 cases using coordinates derived from the bound complexes and in 32 of 54 cases using independently determined coordinates of one or both monomers. Experimental binding affinities correlate with the calculated score function and explain the predictive success or failure of many targets. Global searches using one or both unbound components predict at least 25% of the native residue-residue contacts in 28 of the 32 cases where binding funnels exist. The results suggest that the method may soon be useful for generating models of biologically important complexes from the structures of the isolated components, but they also highlight the challenges that must be met to achieve consistent and accurate prediction of protein-protein interactions.

Algorithms↗

Accurate computer-based design of a new backbone conformation in the second turn of protein L.

The rational design of loops and turns is a key step towards creating proteins with new functions. We used a computational design procedure to create new backbone conformations in the second turn of protein L. The Protein Data Bank was searched for alternative turn conformations, and sequences optimal for these turns in the context of protein L were identified using a Monte Carlo search procedure and an energy function that favors close packing. Two variants containing 12 and 14 mutations were found to be as stable as wild-type protein L. The crystal structure of one of the variants has been solved at a resolution of 1.9 A, and the backbone conformation in the second turn is remarkably close to that of the in silico model (1.1 A RMSD) while it differs significantly from that of wild-type protein L (the turn residues are displaced by an average of 7.2 A). The folding rates of the redesigned proteins are greater than that of the wild-type protein and in contrast to wild-type protein L the second beta-turn appears to be formed at the rate limiting step in folding.

Amino Acid Sequence↗

Crystal structures and increased stabilization of the protein G variants with switched folding pathways NuG1 and NuG2.

We recently described two protein G variants (NuG1 and NuG2) with redesigned first hairpins that were almost twice as stable, folded 100-fold faster, and had a switched folding mechanism relative to the wild-type protein. To test the structural accuracy of our design algorithm and to provide insights to the dramatic changes in the kinetics and thermodynamics of folding, we have now determined the crystal structures of NuG1 and NuG2 to 1.8 A and 1.85 A, respectively. We find that they adopt hairpin structures that are closer to the computational models than to wild-type protein G; the RMSD of the NuG1 hairpin to the design model and the wild-type structure are 1.7 A and 5.1 A, respectively. The crystallographic B factor in the redesigned first hairpin of NuG1 is systematically higher than the second hairpin, suggesting that the redesigned region is somewhat less rigid. A second round of structure-based design yielded new variants of NuG1 and NuG2, which are further stabilized by 0.5 kcal/mole and 0.9 kcal/mole.

Crystallography, X-Ray↗