Search PubMed⌕ Search

Biomedical subjects

R L Jernigan

Publications and source records attributed to R L Jernigan.

At least 19 recordsLinked to original sources

Combining the GOR V algorithm with evolutionary information for protein secondary structure prediction from amino acid sequence.

We have modified and improved the GOR algorithm for the protein secondary structure prediction by using the evolutionary information provided by multiple sequence alignments, adding triplet statistics, and optimizing various parameters. We have expanded the database used to include the 513 non-redundant domains collected recently by Cuff and Barton (Proteins 1999;34:508-519; Proteins 2000;40:502-511). We have introduced a variable size window that allowed us to include sequences as short as 20-30 residues. A significant improvement over the previous versions of GOR algorithm was obtained by combining the PSI-BLAST multiple sequence alignments with the GOR method. The new algorithm will form the basis for the future GOR V release on an online prediction server. The average accuracy of the prediction of secondary structure with multiple sequence alignment and full jack-knife procedure was 73.5%. The accuracy of the prediction increases to 74.2% by limiting the prediction to 375 (of 513) sequences having at least 50 PSI-BLAST alignments. The average accuracy of the prediction of the new improved program without using multiple sequence alignments was 67.5%. This is approximately a 3% improvement over the preceding GOR IV algorithm (Garnier J, Gibrat JF, Robson B. Methods Enzymol 1996;266:540-553; Kloczkowski A, Ting K-L, Jernigan RL, Garnier J. Polymer 2002;43:441-449). We have discussed alternatives to the segment overlap (Sov) coefficient proposed by Zemla et al. (Proteins 1999;34:220-223).

Algorithms↗

Molecular mechanisms of chaperonin GroEL-GroES function.

The dynamics of the GroEL-GroES complex is investigated with a coarse-grained model. This model is one in which single-residue points are connected to other such points, which are nearby, by identical springs, forming a network of interactions. The nature of the most important (slowest) normal modes reveals a wide variety of motions uniquely dependent upon the central cavity of the structure, including opposed torsional rotation of the two GroEL rings accompanied by the alternating compression and expansion of the GroES cap binding region, bending, shear, opposed radial breathing of the cis and trans rings, and stretching and contraction along the protein assembly's long axis. The intermediate domains of the subunits are bifunctional due to the presence of two hinges, which are alternatively activated or frozen by an ATP-dependent mechanism. ATP binding stabilizes a relatively open conformation (with respect to the central cavity) and hinders the motion of the hinge site connecting the intermediate and equatorial domains, while enhancing the flexibility of the second hinge that sets in motion the apical domains. The relative flexibilities of the hinges are reversed in the nucleotide-free form. Cooperative cross-correlations between subunits provide information about the mechanism of action of the protein. The mechanical motions driven by the different modes provide variable binding surfaces and variable sized cavities in the interior to enable accommodation of a broad range of protein substrates. These modes of motion could be used to manipulate the substrate's conformations.

Adenosine Triphosphate↗

Sequence-dependent B<-->A transition in DNA evaluated with dimeric and trimeric scales.

Experimental data on the sequence-dependent B<-->A conformational transition in 24 oligo- and polymeric duplexes yield optimal dimeric and trimeric scales for this transition. The 10 sequence dimers and the 32 trimers of the DNA duplex were characterized by the free energy differences between the B and A forms in water solution. In general, the trimeric scale describes the sequence-dependent DNA conformational propensities more accurately than the dimeric scale, which is likely related to the trimeric model accounting for the two interfaces between adjacent base pairs on both sides (rather than only one interface in the dimeric model). The exceptional preference of the B form for the AA:TT dimers and AAN:N'TT trimers is consistent with the cooperative interactions in both grooves. In the minor groove, this is the hydration spine that stabilizes adenine runs in B form. In the major groove, these are hydrophobic interactions between the thymine methyls and the sugar methylene groups from the preceding nucleotides, occurring in B form. This interpretation is in accord with the key role played by hydration in the B<-->A transition in DNA. Importantly, our trimeric scale is consistent with the relative occurrences of the DNA trimers in A form in protein-DNA cocrystals. Thus, we suggest that the B/A scales developed here can be used for analyzing genome sequences in search for A-philic motifs, putatively operative in the protein-DNA recognition.

Carbohydrates↗

Anisotropy of fluctuation dynamics of proteins with an elastic network model.

Fluctuations about the native conformation of proteins have proven to be suitably reproduced with a simple elastic network model, which has shown excellent agreement with a number of different properties for a wide variety of proteins. This scalar model simply investigates the magnitudes of motion of individual residues in the structure. To use the elastic model approach further for developing the details of protein mechanisms, it becomes essential to expand this model to include the added details of the directions of individual residue fluctuations. In this paper a new tool is presented for this purpose and applied to the retinol-binding protein, which indicates enhanced flexibility in the region of entry to the ligand binding site and for the portion of the protein binding to its carrier protein.

Anisotropy↗

Proteins with similar architecture exhibit similar large-scale dynamic behavior.

We have investigated the similarities and differences in the computed dynamic fluctuations exhibited by six members of a protein fold family with a coarse-grained Gaussian network model. Specifically, we consider the cofactor binding fragment of CysB; the lysine/arginine/ornithine-binding protein (LAO); the enzyme porphobilinogen deaminase (PBGD); the ribose-binding protein (RBP); the N-terminal lobe of ovotransferrin in apo-form (apo-OVOT); and the leucine/isoleucine/valine-binding protein (LIVBP). All have domains that resemble a Rossmann fold, but there are also some significant differences. Results indicate that similar global dynamic behavior is preserved for the members of a fold family, and that differences usually occur in regions only where specific function is localized. The present work is a computational demonstration that the scaffold of a protein fold may be utilized for diverse purposes. LAO requires a bound ligand before it conforms to the large-scale fluctuation behavior of the three other members of the family, CysB, PBGD, and RBP, all of which contain a substrate (cofactor) at the active site cleft. The dynamics of the ligand-free enzymes LIVBP and apo-OVOT, on the other hand, concur with that of unliganded LAO. The present results suggest that it is possible to construct structure alignments based on dynamic fluctuation behavior.

Apoproteins↗

Identifying sequence-structure pairs undetected by sequence alignments.

We examine how effectively simple potential functions previously developed can identify compatibilities between sequences and structures of proteins for database searches. The potential function consists of pairwise contact energies, repulsive packing potentials of residues for overly dense arrangement and short-range potentials for secondary structures, all of which were estimated from statistical preferences observed in known protein structures. Each potential energy term was modified to represent compatibilities between sequences and structures for globular proteins. Pairwise contact interactions in a sequence-structure alignment are evaluated in a mean field approximation on the basis of probabilities of site pairs to be aligned. Gap penalties are assumed to be proportional to the number of contacts at each residue position, and as a result gaps will be more frequently placed on protein surfaces than in cores. In addition to minimum energy alignments, we use probability alignments made by successively aligning site pairs in order by pairwise alignment probabilities. The results show that the present energy function and alignment method can detect well both folds compatible with a given sequence and, inversely, sequences compatible with a given fold, and yield mostly similar alignments for these two types of sequence and structure pairs. Probability alignments consisting of most reliable site pairs only can yield extremely small root mean square deviations, and including less reliable pairs increases the deviations. Also, it is observed that secondary structure potentials are usefully complementary to yield improved alignments with this method. Remarkably, by this method some individual sequence-structure pairs are detected having only 5-20% sequence identity.

Algorithms↗

Characterization of anticancer agents by their growth inhibitory activity and relationships to mechanism of action and structure.

An analysis of the growth inhibitory potency of 122 anticancer agents available from the National Cancer Institute anticancer drug screen is presented. Methods of singular value decomposition (SVD) were applied to determine the matrix of distances between all compounds. These SVD-derived dissimilarity distances were used to cluster compounds that exhibit similar tumor growth inhibitory activity patterns against 60 human cancer cell lines. Cluster analysis divides the 122 standard agents into 25 statistically distinct groups. The first eight groups include structurally diverse compounds with reactive functionalities that act as DNA-damaging agents while the remaining 17 groups include compounds that inhibit nucleic acid biosynthesis and mitosis. Examination of the average activity patterns across the 60 tumor cell lines reveals unique 'fingerprints' associated with each group. A diverse set of structural features are observed for compounds within these groups, with frequent occurrences of strong within-group structural similarities. Clustering of cell types by their response to the 122 anticancer agents divides the 60 cell types into 21 groups. The strongest within-panel groupings were found for the renal, leukemia and ovarian cell panels. These results contribute to the basis for comparisons between log(GI(50)) screening patterns of the 122 anticancer agents and additional tested compounds.

Animals↗

Evaluation of short-range interactions as secondary structure energies for protein fold and sequence recognition.

Short-range interactions for secondary structures of proteins are evaluated as potentials of mean force from the observed frequencies of secondary structures in known protein structures which are assumed to have an equilibrium distribution with the Boltzmann factor of secondary structure energies. A secondary conformation at each residue position in a protein is described by a tripeptide, including one nearest neighbor on each side. The secondary structure potentials are approximated as additive contributions from neighboring residues along the sequence. These are part of an empirical potential to provide a crude estimate of protein conformational energy at a residue level. Unlike previous works, interactions are decoupled into intrinsic potentials of residues, potentials of backbone-backbone interactions, and of side chain-backbone interactions. Also interactions are decoupled into one-body, two-body, and higher order interactions between peptide backbone and side chain and between backbones. These decouplings are essential to correctly evaluate the total secondary structure energy of a protein structure without overcounting interactions. Each interaction potential is evaluated separately by taking account of the correlation in the amino acid order of protein sequences. Interactions among side chains are neglected, because of the relatively limited number of protein structures. Proteins 1999;36:347-356. Published 1999 Wiley-Liss, Inc.

Mathematics↗

An empirical energy potential with a reference state for protein fold and sequence recognition.

We consider modifications of an empirical energy potential for fold and sequence recognition to represent approximately the stabilities of proteins in various environments. A potential used here includes a secondary structure potential representing short-range interactions for secondary structures of proteins, and a tertiary structure potential consisting of a long-range, pairwise contact potential and a repulsive packing potential. This potential is devised to evaluate together the total conformational energy of a protein at the coarse grained residue level. It was previously estimated from the observed frequencies of secondary structures, from contact frequencies between residues, and from the distributions of the number of residues in contact in known protein structures by regarding those distributions as the equilibrium distributions with the Boltzmann factor of these interaction energies. The stability of native structures is assumed as a primary requirement for proteins to fold into their native structures. A collapse energy is subtracted from the contact energies to remove the protein size dependence and to represent protein stabilities for monomeric and multimeric states. The free energy of the whole ensemble of protein conformations that is subtracted from the conformational energy to represent protein stability is approximated as the average energy expected for a typical native structure with the same amino acid composition. This term may be constant in fold recognition but essentially varies in sequence recognition. A simple test of threading sequences into structures without gaps is employed to demonstrate the importance of the present modifications that permit the same potential to be utilized for both fold and sequence recognition. Proteins 1999;36:357-369. Published 1999 Wiley-Liss, Inc.

Models, Chemical↗

Cooperative fluctuations and subunit communication in tryptophan synthase.

Tryptophan synthase (TRPS), with linearly arrayed subunits alphabetabetaalpha, catalyzes the last two reactions in the biosynthesis of L-tryptophan. The two reactions take place in the respective alpha- and beta-subunits of the enzyme, and the intermediate product, indole, is transferred from the alpha- to the beta-site through a 25 A long hydrophobic tunnel. The occurrence of a unique ligand-mediated long-range cooperativity for substrate channeling, and a quest to understand the mechanism of allosteric control and coordination in metabolic cycles, have motivated many experimental studies on the structure and catalytic activity of the TRPS alpha2beta2 complex and its mutants. The dynamics of these complexes are analyzed here using a simple but rigorous theoretical approach, the Gaussian network model. Both wild-type and mutant structures, in the unliganded and various liganded forms, are considered. The substrate binding site in the beta-subunit is found to be closely coupled to a group of hinge residues (beta77-beta89 and beta376-beta379) near the beta-beta interface. These residues simultaneously control the anticorrelated motion of the two beta-subunits, and the opening or closing of the hydrophobic tunnel. The latter process is achieved by the large amplitude fluctuations of the so-called COMM domain in the same subunit. Intersubunit communications are strengthened in the presence of external aldimines bound to the beta-site. The motions of the COMM core residues are coordinated with those of the alpha-beta hinge residues beta174-beta179 on the interfacial helix betaH6 at the entrance of the hydrophobic tunnel. And the motions of betaH6 are coupled, via helix betaH1 and alphaL6, to those of the loop alphaL2 that includes the alpha-subunit catalytically active residue Asp60. Overall, our analysis sheds light on the molecular machinery underlying subunit communication, and identifies the residues playing a key role in the cooperative transmission of conformational motions across the two reaction sites.

Allosteric Regulation↗

p53-induced DNA bending and twisting: p53 tetramer binds on the outer side of a DNA loop and increases DNA twisting.

DNA binding activity of p53 is crucial for its tumor suppressor function. Our recent studies have shown that four molecules of the DNA binding domain of human p53 (p53DBD) bind the response elements with high cooperativity and bend the DNA. By using A-tract phasing experiments, we find significant differences between the bending and twisting of DNA by p53DBD and by full-length human wild-type (wt) p53. Our data show that four subunits of p53DBD bend the DNA by 32-36 degrees, whereas wt p53 bends it by 51-57 degrees. The directionality of bending is consistent with major groove bends at the two pentamer junctions in the consensus DNA response element. More sophisticated phasing analyses also demonstrate that p53DBD and wt p53 overtwist the DNA response element by approximately 35 degrees and approximately 70 degrees, respectively. These results are in accord with molecular modeling studies of the tetrameric complex. Within the constraints imposed by the protein subunits, the DNA can assume a range of conformations resulting from correlated changes in bend and twist angles such that the p53-DNA tetrameric complex is stabilized by DNA overtwisting and bending toward the major groove at the CATG tetramers. This bending is consistent with the inherent sequence-dependent anisotropy of the duplex. Overall, the four p53 moieties are placed laterally in a staggered array on the external side of the DNA loop and have numerous interprotein interactions that increase the stability and cooperativity of binding. The novel architecture of the p53 tetrameric complex has important functional implications including possible p53 interactions with chromatin.

Base Sequence↗

Collective motions in HIV-1 reverse transcriptase: examination of flexibility and enzyme function.

In order to study the inferences of structure for mechanism, the collective motions of the retroviral reverse transcriptase HIV-1 RT (RT) are examined using the Gaussian network model (GNM) of proteins. This model is particularly suitable for elucidating the global dynamic characteristics of large proteins such as the presently investigated heterodimeric RT comprising a total of 982 residues. Local packing density and coordination order of amino acid residues is inspected by the GNM to determine the type and range of motions, both at the residue level and on a global scale, such as the correlated movements of entire subdomains. Of the two subunits, p66 and p51, forming the RT, only p66 has a DNA-binding cleft and a functional polymerase active site. This difference in the structure of the two subunits is shown here to be reflected in their dynamic characteristics: only p66 has the potential to undergo large-scale cooperative motions in the heterodimer, while p51 is essentially rigid. Taken together, the global motion of the RT heterodimer is comprised of movements of the p66 thumb subdomain perpendicular to those of the p66 fingers, accompanied by anticorrelated fluctuations of the RNase H domain and p51 thumb, thus providing information about the details of one processivity mechanism. A few clusters of residues, generally distant in sequence but close in space, are identified in the p66 palm and connection subdomains, which form the hinge-bending regions that control the highly concerted motion of the subdomains. These regions include the catalytically active site and the non-nucleoside inhibitor binding pocket of p66 polymerase, as well as sites whose mutations have been shown to impair enzyme activity. It is easily conceivable that this hinge region, indicated by GNM analysis to play a critical role in modulating the global motion, is locked into an inactive conformation upon binding of an inhibitor. Comparative analysis of the dynamic characteristics of the unliganded and liganded dimers indicates severe repression of the mobility of the p66 thumb in RT's global mode, upon binding of non-nucleoside inhibitors.

Binding Sites↗

Self-consistent estimation of inter-residue protein contact energies based on an equilibrium mixture approximation of residues.

Pairwise contact energies for 20 types of residues are estimated self-consistently from the actual observed frequencies of contacts with regression coefficients that are obtained by comparing "input" and predicted values with the Bethe approximation for the equilibrium mixtures of residues interacting. This is premised on the fact that correlations between the "input" and the predicted values are sufficiently high although the regression coefficients themselves can depend to some extent on protein structures as well as interaction strengths. Residue coordination numbers are optimized to obtain the best correlation between "input" and predicted values for the partition energies. The contact energies self-consistently estimated this way indicate that the partition energies predicted with the Bethe approximation should be reduced by a factor of about 0.3 and the intrinsic pairwise energies by a factor of about 0.6. The observed distribution of contacts can be approximated with a small relative error of only about 0.08 as an equilibrium mixture of residues, if many proteins were employed to collect more than 20,000 contacts. Including repulsive packing interactions and secondary structure interactions further reduces the relative errors. These new contact energies are demonstrated by threading to have improved their ability to discriminate native structures from other non-native folds.

Computer Simulation↗

RNA bulge entropies in the unbound state correlate with peptide binding strengths for HIV-1 and BIV TAR RNA because of improved conformational access.

For the binding of peptides to wild-type HIV-1 and BIV TAR RNA and to mutants with bulges of various sizes, changes in the DeltaDelta G values of binding were determined from experimental K d values. The corresponding entropies of these bulges are estimated by enumerating all possible RNA bulge conformations on a lattice and then applying the Boltzmann relationship. Independent calculations of entropies from fluctuations are also carried out using the Gaussian network model (GNM) recently introduced for analyzing folded structures. Strong correlations are seen between the changes in free energy determined for binding and the two different unbound entropy calculations. The fact that the calculated entropy increase with larger bulge size is correlated with the enhanced experimental binding free energy is unusual. This system exhibits a dependence on the entropy of the unbound form that is opposite to usual binding models. Instead of a large initial entropy being unfavorable since it would be reduced upon binding, here the larger entropies actually favor binding. Several interpretations are possible: (i) the higher conformational freedom implies a higher competence for binding with a minimal strain, by suitable selection amongst the set of already accessible conformations; (ii) larger bulge entropies enhance the probability of the specific favorable conformation of the bound state; (iii) the increased freedom of the larger bulges contri-butes more to the bound state than to the unbound state; (iv) indirectly the large entropy of the bound state might have an unfavorable effect on the solvent structure. Nonetheless, this unusual effect is interesting.

Animals↗

Vibrational dynamics of transfer RNAs: comparison of the free and synthetase-bound forms.

The vibrational dynamics of transfer RNAs, both free, and complexed with the cognate synthetase, are analyzed using a model (Gaussian network model) which recently proved to satisfactorily describe the collective motions of folded proteins. The approach is similar to a normal mode analysis, with the major simplification that no residue specificity is taken into consideration, which permits us (i) to cast the problem into an analytical form applicable to biomolecular systems including about 10(3 )residues, and (ii) to acquire information on the essential dynamics of such large systems within computational times at least two orders of magnitude shorter than conventional simulations. On a local scale, the fluctuations calculated for yeast tRNAPhe and tRNAAsp in the free state, and for tRNAGln complexed with glutaminyl-tRNA synthetase (GlnRS) are in good agreement with the corresponding crystallographic B factors. On a global scale, a hinge-bending region comprising nucleotides U8 to C12 in the D arm, G20 to G22 in the D loop, and m7G46 to C48 in the variable loop (for tRNAPhe), is identified in the free tRNA, conforming with previous observations. The two regions subject to the largest amplitude anticorrelated fluctuations in the free form, i.e. the anticodon region and the acceptor arm are, at the same time, the regions that experience the most severe suppression in their flexibilities upon binding to synthetase, suggesting that their sampling of the conformational space facilitates their recognition by the synthetase. Likewise, examination of the global mode of motion of GlnRS in the complex indicates that residues 40 to 45, 260 to 270, 306 to 314, 320 to 327 and 478 to 485, all of which cluster near the ATP binding site, form a hinge-bending region controlling the cooperative motion, and thereby the catalytic function, of the enzyme. The distal beta-barrel and the tRNA acceptor binding domain, on the other hand, are distinguished by their high mobilities in the global modes of motion, a feature typical of recognition sites, also observed for other proteins. Most of the conserved bases and residues of tRNA and GlnRS are severely constrained in the global motions of the molecules, suggesting their having a role in stabilizing and modulating the global motion.

Amino Acyl-tRNA Synthetases↗

A role for CH...O interactions in protein-DNA recognition.

The concept of CH...O hydrogen bonds has recently gained much interest, with a number of reports indicating the significance of these non-classical hydrogen bonds in stabilizing nucleic acid and protein structures. Here, we analyze the CH...O interactions in the protein-DNA interface, based on 43 crystal structures of protein-DNA complexes. Surprisingly, we find that the number of close intermolecular CH...O contacts involving the thymine methyl group and position C5 of cytosine is comparable to the number of protein-DNA hydrogen bonds involving nitrogen and oxygen atoms as donors and acceptors. A comprehensive analysis of the geometries of these close contacts shows that they are similar to other CH...O interactions found in proteins and small molecules, as well as to classical NH...O hydrogen bonds. Thus, we suggest that C5 of cytosine and C5-Met of thymine form relatively weak CH...O hydrogen bonds with Asp, Asn, Glu, Gln, Ser, and Thr, contributing to the specificity of recognition. Including these interactions, in addition to the classical protein-DNA hydrogen bonds, enables the extraction of simple structural principles for amino acid-base recognition consistent with electrostatic considerations.

Base Composition↗

Correlation between native-state hydrogen exchange and cooperative residue fluctuations from a simple model.

Recently, we developed a simple analytical model based on local residue packing densities and the distribution of tertiary contacts for describing the conformational fluctuations of proteins in their folded state. This so-called Gaussian network model (GNM) is applied here to the interpretation of experimental hydrogen exchange (HX) behavior of proteins in their native state or under weakly denaturing conditions. Calculations are performed for five proteins: bovine pancreatic trypsin inhibitor, cytochrome c, plastocyanin, staphylococcal nuclease, and ribonuclease H. The results are significant in two respects. First, a good agreement is reached between calculated fluctuations and experimental measurements of HX despite the simplicity of the model and within computational times 2 or 3 orders of magnitude faster than earlier, more complex simulations. Second, the success of a theory, based on the coupled conformational fluctuations of residues near the native state, to satisfactorily describe the native-state HX behavior indicates the significant contribution of local, but cooperative, fluctuations to protein conformational dynamics. The correlation between the HX data and the unfolding kinetics of individual residues further suggests that local conformational susceptibilities as revealed by the GNM approach may have implications relevant to the global dynamics of proteins.

Aprotinin↗

Identification of kinetically hot residues in proteins.

A number of recent studies called attention to the presence of kinetically important residues underlying the formation and stabilization of folding nuclei in proteins, and to the possible existence of a correlation between conserved residues and those participating in the folding nuclei. Here, we use the Gaussian network model (GNM), which recently proved useful in describing the dynamic characteristics of proteins for identifying the kinetically hot residues in folded structures. These are the residues involved in the highest frequency fluctuations near the native state coordinates. Their high frequency is a manifestation of the steepness of the energy landscape near their native state positions. The theory is applied to a series of proteins whose kinetically important residues have been extensively explored: chymotrypsin inhibitor 2, cytochrome c, and related C2 proteins. Most of the residues previously pointed out to underlie the folding process of these proteins, and to be critically important for the stabilization of the tertiary fold, are correctly identified, indicating a correlation between the kinetic hot spots and the early forming structural elements in proteins. Additionally, a strong correlation between kinetically hot residues and loci of conserved residues is observed. Finally, residues that may be important for the stability of the tertiary structure of CheY are proposed.

Amino Acid Sequence↗