Protein structure prediction using Rosetta.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to David Baker.
Explore the source record for details and available documents.
Allogeneic bone marrow transplantation (BMT) without a total body irradiation (TBI) conditioning regimen was investigated in children with juvenile myelomonocytic leukemia (JMML). Eight consecutive patients with JMML (n = 6) or monosomy 7 (n = 2) underwent BMT at a median age of 20 months. Donor source included fully matched related (n = 3), mismatched related (n = 2), or fully matched unrelated (n = 3). The conditioning regimen included busulfan, cyclophosphamide, and etoposide (VP16) (melphalan was substituted for VP16 in one patient). The first patient in the series underwent TBI. Graft-versus-host disease prophylaxis was with cyclosporin and methotrexate and in vivo T-cell depletion (Campath 1 g) for mismatched and unrelated transplants. Seven and two patients, respectively, received chemotherapy and splenectomy before BMT. At a median follow-up of 48 months after BMT, five patients remained in remission. The overall survival rate was 63% at 5 years. All deaths occurred in patients with refractory disease at the time of BMT. Allogeneic BMT without TBI appears to be effective therapy for JMML and avoids some of the potential late sequelae of TBI in preschool children.
Understanding the sequence determinants of protein structure, stability and folding is critical for understanding how natural proteins have evolved and how proteins can be engineered to perform novel functions. The complexity of the protein folding problem requires the ability to search large volumes of sequence space for proteins with specific structural or functional characteristics. Here we describe our efforts to identify novel proteins using a phage-display selection strategy from a 'mini-exon' shuffling library generated from the yeast genome and from completely random sequence libraries, and compare the results to recent successes in generating novel proteins using in silico protein design.
Explore the source record for details and available documents.
Experimental structure determination by x-ray crystallography and NMR spectroscopy is slow and time-consuming compared with the rate at which new protein sequences are being identified. NMR spectroscopy has the advantage of rapidly providing the structurally relevant information in the form of unassigned chemical shifts (CSs), intensities of NOESY crosspeaks [nuclear Overhauser effects (NOEs)], and residual dipolar couplings (RDCs), but use of these data are limited by the time and effort needed to assign individual resonances to specific atoms. Here, we develop a method for generating low-resolution protein structures by using unassigned NMR data that relies on the de novo protein structure prediction algorithm, rosetta [Simons, K. T., Kooperberg, C., Huang, E. & Baker, D. (1997) J. Mol. Biol. 268, 209-225] and a Monte Carlo procedure that searches for the assignment of resonances to atoms that produces the best fit of the experimental NMR data to a candidate 3D structure. A large ensemble of models is generated from sequence information alone by using rosetta, an optimal assignment is identified for each model, and the models are then ranked based on their fit with the NMR data assuming the identified assignments. The method was tested on nine protein sequences between 56 and 140 amino acids and published CS, NOE, and RDC data. The procedure yielded models with rms deviations between 3 and 6 A, and, in four of the nine cases, the partial assignments obtained by the method could be used to refine the structures to high resolution (0.6-1.8 A) by repeated cycles of structure generation guided by the partial assignments, followed by reassignment using the newly generated models.
A major challenge of computational protein design is the creation of novel proteins with arbitrarily chosen three-dimensional structures. Here, we used a general computational strategy that iterates between sequence design and structure prediction to design a 93-residue alpha/beta protein called Top7 with a novel sequence and topology. Top7 was found experimentally to be folded and extremely stable, and the x-ray crystal structure of Top7 is similar (root mean square deviation equals 1.2 angstroms) to the design model. The ability to design a new protein fold makes possible the exploration of the large regions of the protein universe not yet observed in nature.
Angular potentials play an important role in the refinement of protein structures through angle-dependent restraints (e.g., those determined by cross-correlated relaxations, residual dipolar couplings, and hydrogen bonds). Analytic derivatives of such angular potentials with respect to the dihedral angles of proteins would be useful for optimizing such restraints and other types of angular potentials (i.e., such as we are now introducing into protein structure prediction) but have not been described. In this article, analytic derivatives are calculated for four types of angular potentials and integrated with the efficient recursive derivative calculation methods of Gō and coworkers. The formulas are implemented in publicly available software and illustrated by refining a low-resolution protein structure with idealized vector-angle, dipolar-coupling, and hydrogen-bond restraints. The method is now being used routinely to optimize hydrogen-bonding potentials in ROSETTA.
The strong coupling between secondary and tertiary structure formation in protein folding is neglected in most structure prediction methods. In this work we investigate the extent to which nonlocal interactions in predicted tertiary structures can be used to improve secondary structure prediction. The architecture of a neural network for secondary structure prediction that utilizes multiple sequence alignments was extended to accept low-resolution nonlocal tertiary structure information as an additional input. By using this modified network, together with tertiary structure information from native structures, the Q3-prediction accuracy is increased by 7-10% on average and by up to 35% in individual cases for independent test data. By using tertiary structure information from models generated with the ROSETTA de novo tertiary structure prediction method, the Q3-prediction accuracy is improved by 4-5% on average for small and medium-sized single-domain proteins. Analysis of proteins with particularly large improvements in secondary structure prediction using tertiary structure information provides insight into the feedback from tertiary to secondary structure.
We have improved the original Rosetta centroid/backbone decoy set by increasing the number of proteins and frequency of near native models and by building on sidechains and minimizing clashes. The new set consists of 1,400 model structures for 78 different and diverse protein targets and provides a challenging set for the testing and evaluation of scoring functions. We evaluated the extent to which a variety of all-atom energy functions could identify the native and close-to-native structures in the new decoy sets. Of various implicit solvent models, we found that a solvent-accessible surface area-based solvation provided the best enrichment and discrimination of close-to-native decoys. The combination of this solvation treatment with Lennard Jones terms and the original Rosetta energy provided better enrichment and discrimination than any of the individual terms. The results also highlight the differences in accuracy of NMR and X-ray crystal structures: a large energy gap was observed between native and non-native conformations for X-ray structures but not for NMR structures.
A previously developed computer program for protein design, RosettaDesign, was used to predict low free energy sequences for nine naturally occurring protein backbones. RosettaDesign had no knowledge of the naturally occurring sequences and on average 65% of the residues in the designed sequences differ from wild-type. Synthetic genes for ten completely redesigned proteins were generated, and the proteins were expressed, purified, and then characterized using circular dichroism, chemical and temperature denaturation and NMR experiments. Although high-resolution structures have not yet been determined, eight of these proteins appear to be folded and their circular dichroism spectra are similar to those of their wild-type counterparts. Six of the proteins have stabilities equal to or up to 7kcal/mol greater than their wild-type counterparts, and four of the proteins have NMR spectra consistent with a well-packed, rigid structure. These encouraging results indicate that the computational protein design methods can, with significant reliability, identify amino acid sequences compatible with a target protein backbone.
Sequence--and structure-based searching strategies have proven useful in the identification of remote homologs and have facilitated both structural and functional predictions of many uncharacterized protein families. We implement these strategies to predict the structure of and to classify a previously uncharacterized cluster of orthologs (COG3019) in the thioredoxin-like fold superfamily. The results of each searching method indicate that thioltransferases are the closest structural family to COG3019. We substantiate this conclusion using the ab initio structure prediction method rosetta, which generates a thioredoxin-like fold similar to that of the glutaredoxin-like thioltransferase (NrdH) for a COG3019 target sequence. This structural model contains the thiol-redox functional motif CYS-X-X-CYS in close proximity to other absolutely conserved COG3019 residues, defining a novel thioredoxin-like active site that potentially binds metal ions. Finally, the rosetta-derived model structure assists us in assembling a global multiple-sequence alignment of COG3019 with two other thioredoxin-like fold families, the thioltransferases and the bacterial arsenate reductases (ArsC).
Protein residues that are critical for structure and function are expected to be conserved throughout evolution. Here, we investigate the extent to which these conserved residues are clustered in three-dimensional protein structures. In 92% of the proteins in a data set of 79 proteins, the most conserved positions in multiple sequence alignments are significantly more clustered than randomly selected sets of positions. The comparison to random subsets is not necessarily appropriate, however, because the signal could be the result of differences in the amino acid composition of sets of conserved residues compared to random subsets (hydrophobic residues tend to be close together in the protein core), or differences in sequence separation of the residues in the different sets. In order to overcome these limits, we compare the degree of clustering of the conserved positions on the native structure and on alternative conformations generated by the de novo structure prediction method Rosetta. For 65% of the 79 proteins, the conserved residues are significantly more clustered in the native structure than in the alternative conformations, indicating that the clustering of conserved residues in protein structures goes beyond that expected purely from sequence locality and composition effects. The differences in the spatial distribution of conserved residues can be utilized in de novo protein structure prediction: We find that for 79% of the proteins, selection of the Rosetta generated conformations with the greatest clustering of the conserved residues significantly enriches the fraction of close-to-native structures.
Protein-protein docking algorithms provide a means to elucidate structural details for presently unknown complexes. Here, we present and evaluate a new method to predict protein-protein complexes from the coordinates of the unbound monomer components. The method employs a low-resolution, rigid-body, Monte Carlo search followed by simultaneous optimization of backbone displacement and side-chain conformations using Monte Carlo minimization. Up to 10(5) independent simulations are carried out, and the resulting "decoys" are ranked using an energy function dominated by van der Waals interactions, an implicit solvation model, and an orientation-dependent hydrogen bonding potential. Top-ranking decoys are clustered to select the final predictions. Small-perturbation studies reveal the formation of binding funnels in 42 of 54 cases using coordinates derived from the bound complexes and in 32 of 54 cases using independently determined coordinates of one or both monomers. Experimental binding affinities correlate with the calculated score function and explain the predictive success or failure of many targets. Global searches using one or both unbound components predict at least 25% of the native residue-residue contacts in 28 of the 32 cases where binding funnels exist. The results suggest that the method may soon be useful for generating models of biologically important complexes from the structures of the isolated components, but they also highlight the challenges that must be met to achieve consistent and accurate prediction of protein-protein interactions.
Multiple sclerosis is increasingly being recognized as a neurodegenerative disease that is triggered by inflammatory attack of the CNS. As yet there is no satisfactory treatment. Using experimental allergic encephalo myelitis (EAE), an animal model of multiple sclerosis, we demonstrate that the cannabinoid system is neuroprotective during EAE. Mice deficient in the cannabinoid receptor CB1 tolerate inflammatory and excitotoxic insults poorly and develop substantial neurodegeneration following immune attack in EAE. In addition, exogenous CB1 agonists can provide significant neuroprotection from the consequences of inflammatory CNS disease in an experimental allergic uveitis model. Therefore, in addition to symptom management, cannabis may also slow the neurodegenerative processes that ultimately lead to chronic disability in multiple sclerosis and probably other diseases.
We predicted structures for all seven targets in the CAPRI experiment using a new method in development at the time of the challenge. The technique includes a low-resolution rigid body Monte Carlo search followed by high-resolution refinement with side-chain conformational changes and rigid body minimization. Decoys (approximately 10(6) per target) were discriminated using a scoring function including van der Waals and solvation interactions, hydrogen bonding, residue-residue pair statistics, and rotamer probabilities. Decoys were ranked, clustered, manually inspected, and selected. The top ranked model for target 6 predicted the experimental structure to 1.5 A RMSD and included 48 of 65 correct residue-residue contacts. Target 7 was predicted at 5.3 A RMSD with 22 of 37 correct residue-residue contacts using a homology model from a known complex structure. Using a preliminary version of the protocol in round 1, target 1 was predicted within 8.8 A although few contacts were correct. For targets 2 and 3, the interface locations and a small fraction of the contacts were correctly identified.
Recently developed 2H spin relaxation experiments are applied to study the dynamics of methyl-containing side-chains in the B1 domain of protein L and in a pair of point mutants of the domain, F22L and A20V. X-ray and NMR studies of the three variants of protein L studied here establish that their structures are very similar, despite the fact that the F22L mutant is 3.2kcal/mol less stable. Measurements of methyl 2H spin relaxation rates, which probe dynamics on a picosecond-nanosecond time scale, and three-bond 3J(Cgamma-CO), 3J(Cgamma-N) and 3J(Calpha-Cdelta) scalar coupling constants, which are sensitive to motion spanning a wide range of time-scales, reveal changes in the magnitude of side-chain dynamics in response to mutation. Observed differences in the time-scale of motions between the variants have been related to changes in energetic barriers. Of interest, several of the residues with different motional properties across the variants are far from the site of mutation, suggesting the presence of long-range interactions within the protein that can be probed through studies of dynamics.
The cannabinoid system is a regulator of neurotransmission and is linked with hormonal control. We have found in experimental mouse studies that the progesterone receptor inhibitor mifepristone (RU38486, 80 mg/kg i.p.) or the 11beta-hydroxylase inhibitor metyrapone (100 mg/kg i.p.) when administered in combination with cannabinoids potentiates the transient-sedating cannabinoid receptor-1 effects of a high-dose Delta(9)-tetrahydrocannabinol (25 mg/kg i.p.), causing severe prolonged sedation associated with hypomotility, catalepsy and hypothermia. This observation has implications for human subjects taking these drugs and related compounds particularly because of the ubiquitous use of cannabis and the high potential for mifepristone and related compounds to become available on the 'black-market' as abortifacients.
The sensory neuron specific sodium channel Na(v)1.8 is normally detectable at only very low levels within cerebellar Purkinje cells. Annexin II light chain (p11) binds to the amino terminus of Na(v)1.8 and facilitates its functional expression within the cell membrane. We previously demonstrated that expression of Na(v)1.8 is up-regulated in cerebellar Purkinje cells in experimental allergic encephalomyelitis (EAE) and multiple sclerosis (MS). In this study we demonstrate that expression of p11 is significantly up-regulated in Purkinje cells in EAE (71 +/- 9.0% vs 21.3 +/- 4.9% in controls) and in MS(65.5 +/- 1.6% vs 21.8 +/- 6.2% in controls). We also demonstrate a high degree of co-expression of p11 and Na(v)1.8 (84.8 +/- 8.9%). Together with earlier results which show that experimental expression of Na(v)1.8 within Purkinje cells perturbs the temporal pattern of impulse generation in these cells, our results extend the evidence for an acquired channelopathy which interferes with cerebellar function in MS.