Search PubMed⌕ Search

Biomedical subjects

B Rost

Publications and source records attributed to B Rost.

At least 37 records · Page 2Linked to original sources

Twilight zone of protein sequence alignments.

Sequence alignments unambiguously distinguish between protein pairs of similar and non-similar structure when the pairwise sequence identity is high (>40% for long alignments). The signal gets blurred in the twilight zone of 20-35% sequence identity. Here, more than a million sequence alignments were analysed between protein pairs of known structures to re-define a line distinguishing between true and false positives for low levels of similarity. Four results stood out. (i) The transition from the safe zone of sequence alignment into the twilight zone is described by an explosion of false negatives. More than 95% of all pairs detected in the twilight zone had different structures. More precisely, above a cut-off roughly corresponding to 30% sequence identity, 90% of the pairs were homologous; below 25% less than 10% were. (ii) Whether or not sequence homology implied structural identity depended crucially on the alignment length. For example, if 10 residues were similar in an alignment of length 16 (>60%), structural similarity could not be inferred. (iii) The 'more similar than identical' rule (discarding all pairs for which percentage similarity was lower than percentage identity) reduced false positives significantly. (iv) Using intermediate sequences for finding links between more distant families was almost as successful: pairs were predicted to be homologous when the respective sequence families had proteins in common. All findings are applicable to automatic database searches.

Computer Simulation↗

MRS of the brain in patients with anorexia or bulimia nervosa.

Twenty patients with anorexia or bulimia nervosa were prospectively investigated by magnetic resonance spectroscopy (MRS) of the brain. Compared to healthy controls, MRS of those with eating disorders revealed metabolic changes, which seem to be a consequence of their nutritional deficiency.

Adolescent↗

Adaptation of protein surfaces to subcellular location.

In vivo, proteins occur in widely different physio-chemical environments, and, from in vitro studies, we know that protein structure can be very sensitive to environment. However, theoretical studies of protein structure have tended to ignore this complexity. In this paper, we have approached this problem by grouping proteins by their subcellular location and looking at structural properties that are characteristic to each location. We hypothesize that, throughout evolution, each subcellular location has maintained a characteristic physio-chemical environment, and that proteins in each location have adapted to these environments. If so, we would expect that protein structures from different locations will show characteristic differences, particularly at the surface, which is directly exposed to the environment. To test this hypothesis, we have examined all eukaryotic proteins with known three-dimensional structure and for which the subcellular location is known to be either nuclear, cytoplasmic, or extracellular. In agreement with previous studies, we find that the total amino acid composition carries a signal that identifies the subcellular location. This signal was due almost entirely to the surface residues. The surface residue signal was often strong enough to accurately predict subcellular location, given only a knowledge of which residues are at the protein surface. The results suggest how the accuracy of prediction of location from sequence can be improved. We concluded that protein surfaces show adaptation to their subcellular location. The nature of these adaptations suggests several principles that proteins may have used in adapting to particular physio-chemical environments; these principles may be useful for protein design.

Amino Acids↗

Role of transmembrane domains in the functions of B- and T-cell receptors.

The antigen receptors on the surface of B- and T-lymphocytes are complexes of several integral membrane proteins, essential for their proper expression and function. Recent studies demonstrated that transmembrane (TM) domains of the components of these receptors play a critical role in their association and function. It was specifically demonstrated that in many cases point mutations in the TM domains can partially or completely disrupt the receptor surface expression and function. Here we review studies of the TM domains of B- and T-cell receptors. Furthermore, we use a novel method, PHDtopology, to provide estimates of the exact locations and lengths of the TM domains of the subunit components of these receptors. Most previous studies used single residue hydrophobicity as a criterion for determining the position and length of the TM domains. In contrast, PHDtopology utilizes a system of neural networks and the evolutionary information contained in multiple alignments of related sequences to predict the location, length, and orientation of transmembrane helices. Present results significantly differ from most published estimates of the TM domains of the B- and T-cell receptor components, primarily in the length of the TM domains. These results may lead to modification of putative TM motifs and re-interpretation of the results of studies using mutated TM domains. The availability of PHDtopology on the Internet would make it a valuable tool in the future studies of the TM domains of integral membrane proteins.

Amino Acid Sequence↗

Protein fold recognition by prediction-based threading.

In fold recognition by threading one takes the amino acid sequence of a protein and evaluates how well it fits into one of the known three-dimensional (3D) protein structures. The quality of sequence-structure fit is typically evaluated using inter-residue potentials of mean force or other statistical parameters. Here, we present an alternative approach to evaluating sequence-structure fitness. Starting from the amino acid sequence we first predict secondary structure and solvent accessibility for each residue. We then thread the resulting one-dimensional (1D) profile of predicted structure assignments into each of the known 3D structures. The optimal threading for each sequence-structure pair is obtained using dynamic programming. The overall best sequence-structure pair constitutes the predicted 3D structure for the input sequence. The method is fine-tuned by adding information from direct sequence-sequence comparison and applying a series of empirical filters. Although the method relies on reduction of 3D information into 1D structure profiles, its accuracy is, surprisingly, not clearly inferior to methods based on evaluation of residue interactions in 3D. We therefore hypothesise that existing 1D-3D threading methods essentially do not capture more than the fitness of an amino acid sequence for a particular 1D succession of secondary structure segments and residue solvent accessibility. The prediction-based threading method on average finds any structurally homologous region at first rank in 29% of the cases (including sequence information). For the 22% first hits detected at highest scores, the expected accuracy rose to 75%. However, the task of detecting entire folds rather than homologous fragments was managed much better; 45 to 75% of the first hits correctly recognised the fold.

Algorithms↗

Better 1D predictions by experts with machines.

Accuracy of predicting protein secondary structure and solvent accessibility has been improved significantly by using evolutionary information contained in multiple sequence alignments. For the second Asilomar meeting, predictions were made automatically for all targets using the publicly available prediction service PredictProtein. Additionally, a semiautomatic procedure for generating more informative alignments was used in combination with the PHD prediction methods. Results confirmed the estimates for prediction accuracy. Furthermore, the more informative alignments yielded better predictions. The fairly accurate predictions of 1D structure were successfully used by various groups for the Asilomar meeting as first step toward predicting higher dimensions of protein structure.

Expert Systems↗

Protein structures sustain evolutionary drift.

A protein sequence folds into a unique three-dimensional protein structure. Different sequences, though, can fold into similar structures. How stable is a protein structure with respect to sequence changes? What percentage of the sequence is 'anchor' residues, that is, residues crucial for protein structure and function? Here, answers to these questions are pursued by analyzing large numbers of structurally homologous protein pairs. Most pairs of similar structures have sequence identity as low as expected from randomly related sequences (8-9%). On average, only 3-4% of all residues are 'anchor' residues. The symmetric shape of the distribution at low sequence identity suggests that for most structures, four billion years of evolution was sufficient to reach an equilibrium. The mean identities for convergent (different ancestor) and divergent (same ancestor) evolution of proteins to similar structures are quite close and hence, in most cases, it is difficult to distinguish between the two effects. In particular, low levels of sequence identity appear not to be indicative of convergent evolution.

Bias↗

Sisyphus and prediction of protein structure.

The problem of predicting protein structure from the sequence remains fundamentally unsolved despite more than three decades of intensive research effort. However, new and promising methods in three-dimensional (3D), 2D and 1D prediction have reopened the field. Mean-force-potentials derived from the protein databases can distinguish between correct and incorrect models (3D). Inter-residue contacts (2D) can be detected by analysis of correlated mutations, albeit with low accuracy. Secondary structure, solvent accessibility and transmembrane helices (1D) can be predicted with significantly improved accuracy using multiple sequence alignments. Some of these new prediction methods have proven accurate and reliable enough to be useful in genome analysis, and in experimental structure determination. Moreover, the new generation of theoretical methods is increasingly influencing experiments in molecular biology.

Computers↗

Carboxyl group protonation upon reduction of the Paracoccus denitrificans cytochrome c oxidase: direct evidence by FTIR spectroscopy.

The redox reactions of the cytochrome c oxidase from Paracoccus denitrificans were investigated in a thin-layer cell designed for the combination of electrochemistry under anaerobic conditions with UV/VIS and IR spectroscopy. Quantitative and reversible electrochemical reactions were obtained at a surface-modified electrode for all cofactors as indicated by the optical signals in the 400-700 nm range. Fourier transform infrared (FTIR) difference spectra of reduction and oxidation (reduced-minus-oxidized and oxidized-minus-reduced, respectively) obtained in the 1800-1000 cm(-1) range reveal highly structured band features with major contributions in the amide I (1620-1680 cm(-1)) and amide II (1580-1520 cm(-1)) range which indicate structural rearrangements in the cofactor vicinity. However, the small amplitude of the IR difference signals indicates that these conformational changes are small and affect only individual peptide groups. In the spectral region above 1700 cm(-1), a positive peak in the reduced state (1733 cm(-1)) and negative peak in the oxidized st ate (1745 cm(-1)) are characteristic for the formation and decay of a COOH mode upon reduction. The most obvious interpretation of this difference signal is proton uptake by one Asp or Glu side chain carboxyl group in the reduced state and deprotonation of another Asp or Glu residue. Moreover, both residues could well be coupled as a donor-acceptor pair in the proton transfer chain. An alternative interpretation is in terms of a protonated carboxyl group which shifts to a different environment in the reduced state. The relevance of this first direct observation of protein protonation changes in the cytochrome c oxidase for vectorial proton transfer and the catalytic reaction is discussed.

Electron Transport Complex IV↗

Topology prediction for helical transmembrane proteins at 86% accuracy.

Previously, we introduced a neural network system predicting locations of transmembrane helices (HTMs) based on evolutionary profiles (PHDhtm, Rost B, Casadio R, Fariselli P, Sander C, 1995, Protein Sci 4:521-533). Here, we describe an improvement and an extension of that system. The improvement is achieved by a dynamic programming-like algorithm that optimizes helices compatible with the neural network output. The extension is the prediction of topology (orientation of first loop region with respect to membrane) by applying to the refined prediction the observation that positively charged residues are more abundant in extra-cytoplasmic regions. Furthermore, we introduce a method to reduce the number of false positives, i.e., proteins falsely predicted with membrane helices. The evaluation of prediction accuracy is based on a cross-validation and a double-blind test set (in total 131 proteins). The final method appears to be more accurate than other methods published: (1) For almost 89% (+/-3%) of the test proteins, all HTMs are predicted correctly. (2) For more than 86% (+/-3%) of the proteins, topology is predicted correctly. (3) We define reliability indices that correlate with prediction accuracy: for one half of the proteins, segment accuracy raises to 98%; and for two-thirds, accuracy of topology prediction is 95%. (4) The rate of proteins for which HTMs are predicted falsely is below 2% (+/-1%). Finally, the method is applied to 1,616 sequences of Haemophilus influenzae. We predict 19% of the genome sequences to contain one or more HTMs. This appears to be lower than what we predicted previously for the yeast VIII chromosome (about 25%).

Algorithms↗

Bridging the protein sequence-structure gap by structure predictions.

The problem of accurately predicting protein three-dimensional structure from sequence has yet to be solved. Recently, several new and promising methods that work in one, two, or three dimensions have invigorated the field. Modeling by homology can yield fairly accurate three-dimensional structures for approximately 25% of the currently known protein sequences. Techniques for cooperatively fitting sequences into known three-dimensional folds, called threading methods, can increase this rate by detecting very remote homologies in favorable cases. Prediction of protein structure in two dimensions, i.e. prediction of interresidue contacts, is in its infancy. Prediction tools that work in one dimension are both mature and generally applicable; they predict secondary structure, residue solvent accessibility, and the location of transmembrane helices with reasonable accuracy. These and other prediction methods have gained immensely from the rapid increase of information in publicly accessible databases. Growing databases will lead to further improvements of prediction methods and, thus, to narrowing the gap between the number of known protein sequences and known protein structures.

Amino Acid Sequence↗

Development of a guinea-pig model for potency/immunogenicity evaluation of diphtheria, tetanus acellular pertussis (DTaP) and Haemophilus influenzae type b polysaccharide conjugate vaccines.

We have evaluated a guinea pig model for assessing the immunogenicity of Haemophilus influenzae type b (Hib) polysaccharide-protein conjugate vaccines, acellular pertussis vaccine and combination vaccines-consisting of tetanus toxoid (TT), diphtheria toxoid (DT), acellular pertussis vaccine and Hib-TT (Hib-T) conjugate vaccine. The model was based on the United States (US) potency test for TT and DT which requires injection of guinea pigs with a single dose of undiluted vaccine. Guinea pigs showed dose-dependent antibody responses to pertussis toxoid (PTxd) and filamentous haemagglutinin (FHA), two important components of acellular pertussis vaccine. Antibody response of guinea pigs to commercially available Hib conjugate vaccines qualitatively resembled those of human infants. Unconjugated polyribosylribitolphosphate (PRP) was not immunogenic; PRP-D conjugate produced a low antibody response, HbOC, PRP-T (Merieux) and Hib-T (MPHBL) produced a low response to the first dose and a strong anamnestic response to the booster dose. PRP-OMP uniquely produced a strong response after the first dose which was boosted by the second dose. In preliminary experiments, injection of guinea pigs with the combined vaccine formulations consisting of TT, DT, whole cell or acellular pertussis vaccine (Ptxd and FHA) and Hib-T conjugate showed that these vaccines were immunogenic when combined, with some effects on the antibody responses of certain components. This model for testing potency/immunogenicity of combined vaccines substantially reduces the number of animals needed to test each lot of vaccine. To reduce the use of animals in testing vaccines further, we propose the use of a Vero cell assay for titrating diphtheria antitoxin and ELISA for measuring IgG antibody to tetanus toxin. The guinea pig model may also be useful for evaluating combination vaccines.

Animal Testing Alternatives↗

Refining neural network predictions for helical transmembrane proteins by dynamic programming.

For transmembrane proteins experimental determination of three-dimensional structure is problematic. However, membrane proteins have important impact for molecular biology in general, and for drug design in particular. Thus, prediction method are needed. Here we introduce a method that started from the output of the profile-based neural network system PHDhtm (Rost, et al. 1995). Instead of choosing the neural network output unit with maximal value as prediction, we implemented a dynamic programming-like refinement procedure that aimed at producing the best model for all transmembrane helices compatible with the neural network output. The refined prediction was used successfully to predict transmembrane topology based on an empirical rule for the charge difference between extra- and intra-cytoplasmic regions (positive-inside rule). Preliminary results suggest that the refinement was clearly superior to the initial neural network system; and that the method predicted all transmembrane helices correctly for more proteins than a previously applied empirical filter. The resulting accuracy in predicting topology was better than 80%. Although a more thorough evaluation of the method on a larger data set will be required, the results compared favourably with alternative methods. The results reflected the strength of the refinement procedure which was the successful incorporation of global information: whereas the residue preferences output by the neural network were derived from stretches of 17 adjacent residues, the refinement procedure involved constraints on the level of the entire protein.

Algorithms↗

Transmembrane helices predicted at 95% accuracy.

We describe a neural network system that predicts the locations of transmembrane helices in integral membrane proteins. By using evolutionary information as input to the network system, the method significantly improved on a previously published neural network prediction method that had been based on single sequence information. The input data were derived from multiple alignments for each position in a window of 13 adjacent residues: amino acid frequency, conservation weights, number of insertions and deletions, and position of the window with respect to the ends of the protein chain. Additional input was the amino acid composition and length of the whole protein. A rigorous cross-validation test on 69 proteins with experimentally determined locations of transmembrane segments yielded an overall two-state per-residue accuracy of 95%. About 94% of all segments were predicted correctly. When applied to known globular proteins as a negative control, the network system incorrectly predicted fewer than 5% of globular proteins as having transmembrane helices. The method was applied to all 269 open reading frames from the complete yeast VIII chromosome. For 59 of these, at least two transmembrane helices were predicted. Thus, the prediction is that about one-fourth of all proteins from yeast VIII contain one transmembrane helix, and some 20%, more than one.

Amino Acid Sequence↗