Search PubMed⌕ Search

Biomedical subjects

H Margalit

Publications and source records attributed to H Margalit.

At least 19 recordsLinked to original sources

Examination of possible structural constraints of MHC-binding peptides by assessment of their native structure within their source proteins.

Antigenic peptides bind to major histocompatibility complex (MHC) molecules as a prerequisite for their presentation to T cells. In this study, we investigate possible structural preferences of MHC-binding peptides by examining the conformation space defined by the structures of these peptides within their native source proteins. Comparison of the conformation space of the native structures of MHC-binding nonamers and a corresponding conformation space defined by a random set of nonamers showed no significant difference. This suggests that the environment of the MHC binding groove has evolved to bind peptides with essentially any "structural background." A slight tendency for an extended beta-conformation at positions 8 and 9 was observed for the set of native structures. We suggest that such a preference may facilitate the binding of the C-terminal anchor position of processed peptides into the corresponding specificity pocket. MHC-binding peptides represent examples of short subsequences that are present in two different structural environments: within their native protein and within the MHC binding groove. Comparison of the native and of the bound structure of the peptides showed that peptides up to 14 residues long may adopt different conformations within different protein environments. This has direct implications for structure prediction algorithms.

Amino Acid Sequence↗

Correlated sequence-signatures as markers of protein-protein interaction.

As protein-protein interaction is intrinsic to most cellular processes, the ability to predict which proteins in the cell interact can aid significantly in identifying the function of newly discovered proteins, and in understanding the molecular networks they participate in. Here we demonstrate that characteristic pairs of sequence-signatures can be learned from a database of experimentally determined interacting proteins, where one protein contains the one sequence-signature and its interacting partner contains the other sequence-signature. The sequence-signatures that recur in concert in various pairs of interacting proteins are termed correlated sequence-signatures, and it is proposed that they can be used for predicting putative pairs of interacting partners in the cell. We demonstrate the potential of this approach on a comprehensive database of experimentally determined pairs of interacting proteins in the yeast Saccharomyces cerevisiae. The proteins in this database have been characterized by their sequence-signatures, as defined by the InterPro classification. A statistical analysis performed on all possible combinations of sequence-signature pairs has identified those pairs that are over-represented in the database of yeast interacting proteins. It is demonstrated how the use of the correlated sequence-signatures as identifiers of interacting proteins can reduce significantly the search space, and enable directed experimental interaction screens.

Computational Biology↗

Novel small RNA-encoding genes in the intergenic regions of Escherichia coli.

BACKGROUND: Small, untranslated RNA molecules were identified initially in bacteria, but examples can be found in all kingdoms of life. These RNAs carry out diverse functions, and many of them are regulators of gene expression. Genes encoding small, untranslated RNAs are difficult to detect experimentally or to predict by traditional sequence analysis approaches. Thus, in spite of the rising recognition that such RNAs may play key roles in bacterial physiology, many of the small RNAs known to date were discovered fortuitously. RESULTS: To search the Escherichia coli genome sequence for genes encoding small RNAs, we developed a computational strategy employing transcription signals and genomic features of the known small RNA-encoding genes. The search, for which we used rather restrictive criteria, has led to the prediction of 24 putative sRNA-encoding genes, of which 23 were tested experimentally. Here we report on the discovery of 14 genes encoding novel small RNAs in E. coli and their expression patterns under a variety of physiological conditions. Most of the newly discovered RNAs are abundant. Interestingly, the expression level of a significant number of these RNAs increases upon entry into stationary phase. CONCLUSIONS: Based on our results, we conclude that small RNAs are much more widespread than previously imagined and that these versatile molecules may play important roles in the fine-tuning of cell responses to changing environments.

Blotting, Northern↗

PromEC: An updated database of Escherichia coli mRNA promoters with experimentally identified transcriptional start sites.

PromEC is an updated compilation of Escherichia coli mRNA promoter sequences. It includes documentation on the location of experimentally identified mRNA transcriptional start sites on the E. coli chromosome, as well as the actual sequences in the promoter region. The database was updated as of July 2000 and includes 472 entries. PromEC is accessible at http://bioinfo.md.huji.ac. il/marg/promec

Chromosomes, Bacterial↗

Harnessing the cellular immune system to the gene-prediction cart.

Prediction of genes and verification of their bona fide expression in the cell are major challenges of the post-genomic era. Here, we demonstrate how information from the apparently unrelated field of cellular immunology can be recruited for these challenging tasks. The cellular immune system presents short peptides that are the degradation products of both foreign and self-proteins expressed in the cell. We carried out a comprehensive search comparing these peptides to all accumulated human sequence data. Our findings illustrate how these 'presented self-peptides' are informative for the identification of new genes, for hypothetical gene verification, for verifying gene expression at the protein level and for supporting splice junctions.

Antigen Presentation↗

Markovian domain fingerprinting: statistical segmentation of protein sequences.

MOTIVATION: Characterization of a protein family by its distinct sequence domains is crucial for functional annotation and correct classification of newly discovered proteins. Conventional Multiple Sequence Alignment (MSA) based methods find difficulties when faced with heterogeneous groups of proteins. However, even many families of proteins that do share a common domain contain instances of several other domains, without any common underlying linear ordering. Ignoring this modularity may lead to poor or even false classification results. An automated method that can analyze a group of proteins into the sequence domains it contains is therefore highly desirable. RESULTS: We apply a novel method to the problem of protein domain detection. The method takes as input an unaligned group of protein sequences. It segments them and clusters the segments into groups sharing the same underlying statistics. A Variable Memory Markov (VMM) model is built using a Prediction Suffix Tree (PST) data structure for each group of segments. Refinement is achieved by letting the PSTs compete over the segments, and a deterministic annealing framework infers the number of underlying PST models while avoiding many inferior solutions. We show that regions of similar statistics correlate well with protein sequence domains, by matching a unique signature to each domain. This is done in a fully automated manner, and does not require or attempt an MSA. Several representative cases are analyzed. We identify a protein fusion event, refine an HMM superfamily classification into the underlying families the HMM cannot separate, and detect all 12 instances of a short domain in a group of 396 sequences. CONTACT: jill@cs.huji.ac.il; tishby@cs.huji.ac.il.

Algorithms↗

A structure-based approach for prediction of protein binding sites in gene upstream regions.

The challenge of identifying DNA regulatory sequences based on sequence information only has been emphasized in view of the fast accumulation of new genes in the databases. While most predictive algorithms are based on multiple alignments of already known binding sites, here we examine the usefulness of a novel approach that is based on structural information of the protein-DNA complex. It has already been shown that specific recognition between a protein and its DNA target is achieved by stereo-chemical complementarity between the protein amino acids and the DNA bases. The proposed computational scheme uses crystallographic information to define the set of amino acid-base contacts between the proteins of a given DNA-binding protein family and their DNA targets. The compatibility of a given protein to bind to putative regulatory DNA sequences is then evaluated by knowledge-based parameters for amino acid-base interactions. By this procedure gene upstream regions may be screened for potential binding sites for regulatory proteins. Predictions are demonstrated for the E. coli cyclic AMP receptor protein (CRP) which recognizes the DNA via the helix-turn-helix motif, and for various Zif268-like proteins which belong to the Cys2His2 zinc finger family. The advantages and limitations of this approach are discussed.

Algorithms↗

Sequence signals for generation of antigenic peptides by the proteasome: implications for proteasomal cleavage mechanism.

Proteasomal cleavage of proteins is the first step in the processing of most antigenic peptides that are presented to cytotoxic T cells. Still, its specificity and mechanism are not fully understood. To identify preferred sequence signals that are used for generation of antigenic peptides by the proteasome, we performed a rigorous analysis of the residues at the termini and flanking regions of naturally processed peptides eluted from MHC class I molecules. Our results show that both the C terminus (position P1 of the cleavage site) and its immediate flanking position (P1') possess significant signals. The N termini of the peptides show these signals only weakly, consistent with previous findings that antigenic peptides may be cleaved by the proteasome with N-terminal extensions. Nevertheless, we succeed to demonstrate indirectly that the N-terminal cleavage sites contain the same preferred signals at position P1'. This reinforces previous findings regarding the role of the P1' position of a cleavage site in determining the cleavage specificity, in addition to the well-known contribution of position P1. Our results apply to the generation of antigenic peptides and bare direct implications for the mechanism of proteasomal cleavage. We propose a model for proteasomal cleavage mechanism by which both ends of cleaved fragments are determined by the same cleavage signals, involving preferred residues at both P1 and P1' positions of a cleavage site. The compatibility of this model with experimental data on protein degradation products and generation of antigenic peptides is demonstrated.

Animals↗

Evaluation of PSI-BLAST alignment accuracy in comparison to structural alignments.

The PSI-BLAST algorithm has been acknowledged as one of the most powerful tools for detecting remote evolutionary relationships by sequence considerations only. This has been demonstrated by its ability to recognize remote structural homologues and by the greatest coverage it enables in annotation of a complete genome. Although recognizing the correct fold of a sequence is of major importance, the accuracy of the alignment is crucial for the success of modeling one sequence by the structure of its remote homologue. Here we assess the accuracy of PSI-BLAST alignments on a stringent database of 123 structurally similar, sequence-dissimilar pairs of proteins, by comparing them to the alignments defined on a structural basis. Each protein sequence is compared to a nonredundant database of the protein sequences by PSI-BLAST. Whenever a pair member detects its pair-mate, the positions that are aligned both in the sequential and structural alignments are determined, and the alignment sensitivity is expressed as the percentage of these positions out of the structural alignment. Fifty-two sequences detected their pair-mates (for 16 pairs the success was bi-directional when either pair member was used as a query). The average percentage of correctly aligned residues per structural alignment was 43.5+/-2.2%. Other properties of the alignments were also examined, such as the sensitivity vs. specificity and the change in these parameters over consecutive iterations. Notably, there is an improvement in alignment sensitivity over consecutive iterations, reaching an average of 50.9+/-2.5% within the five iterations tested in the current study.

Algorithms↗

Structure-based prediction of binding peptides to MHC class I molecules: application to a broad range of MHC alleles.

Specific binding of antigenic peptides to major histocompatibility complex (MHC) class I molecules is a prerequisite for their recognition by cytotoxic T-cells. Prediction of MHC-binding peptides must therefore be incorporated in any predictive algorithm attempting to identify immunodominant T-cell epitopes, based on the amino acid sequence of the protein antigen. Development of predictive algorithms based on experimental binding data requires experimental testing of a very large number of peptides. A complementary approach relies on the structural conservation observed in crystallographically solved peptide-MHC complexes. By this approach, the peptide structure in the MHC groove is used as a template upon which peptide candidates are threaded, and their compatibility to bind is evaluated by statistical pairwise potentials. Our original algorithm based on this approach used the pairwise potential table of Miyazawa and Jernigan (Miyazawa S, Jernigan RL, 1996, J Mol Biol 256:623-644) and succeeded to correctly identify good binders only for MHC molecules with hydrophobic binding pockets, probably because of the high emphasis of hydrophobic interactions in this table. A recently developed pairwise potential table by Betancourt and Thirumalai (Betancourt MR, Thirumalai D, 1999, Protein Sci 8:361-369) that is based on the Miyazawa and Jernigan table describes the hydrophilic interactions more appropriately. In this paper, we demonstrate how the use of this table, together with a new definition of MHC contact residues by which only residues that contribute exclusively to sequence specific binding are included, allows the development of an improved algorithm that can be applied to a wide range of MHC class I alleles.

Alleles↗

Glimmers in the midnight zone: characterization of aligned identical residues in sequence-dissimilar proteins sharing a common fold.

Sequence comparison of proteins that adopt the same fold has revealed a large degree of sequence variation. There are many pairs of structurally similar proteins with only a very low percentage of identical residues at structurally aligned positions. It is not clear whether these few identical residues have been conserved just by coincidence, or due to their structural and/or functional role The current study focuses on characterization of STructurally Aligned Identical ResidueS (STAIRS) in a data set of protein pairs that are structurally similar but sequentially dissimilar. The conservation pattern of the residues at structurally aligned positions has been characterized within the protein families of the two pair members, and mutually highly and weakly conserved positions of STAIRS could be identified About 40% of the STAIRS are only moderately conserved, suggesting that their maintenance may have been coincidental. The mutually highly conserved STAIRS show distinct features that are associated with protein structure and function: a relatively high fraction of these STAIRS are buried within their protein structures. Glycine, cysteine, histidine, and tryptophan are significantly over-represented among the mutually conserved STAIRS. A detailed survey of these STAIRS reveals residue-specific roles in the determination of the protein's structure and function.

Animals↗

Molecular characterization of a common fragile site (FRA7H) on human chromosome 7 by the cloning of a simian virus 40 integration site.

Common fragile sites are chromosomal loci prone to breakage and rearrangement, hypothesized to provide targets for foreign DNA integration. We cloned a simian virus 40 integration site and showed by fluorescent in situ hybridization analysis that the integration event had occurred within a common aphidicolin-induced fragile site on human chromosome 7, FRA7H. A region of 161 kb spanning FRA7H was defined and sequenced. Several regions with a potential unusual DNA structure, including high-flexibility, low-stability, and non-B-DNA-forming sequences were identified in this region. We performed a similar analysis on the published FRA3B sequence and the putative partial FRA7G, which also revealed an impressive cluster of regions with high flexibility and low stability. Thus, these unusual DNA characteristics are possibly intrinsic properties of common fragile sites that may affect their replication and condensation as well as organization, and may lead to fragility.

Base Sequence↗

Quantitative parameters for amino acid-base interaction: implications for prediction of protein-DNA binding sites.

Inspection of the amino acid-base interactions in protein-DNA complexes is essential to the understanding of specific recognition of DNA target sites by regulatory proteins. The accumulation of information on protein-DNA co-crystals challenges the derivation of quantitative parameters for amino acid-base interaction based on these data. Here we use the coordinates of 53 solved protein-DNA complexes to extract all non-homologous pairs of amino acid-base that are in close contact, including hydrogen bonds and hydrophobic interactions. By comparing the frequency distribution of the different pairs to a theoretical distribution and calculating the log odds, a quantitative measure that expresses the likelihood of interaction for each pair of amino acid-base could be extracted. A score that reflects the compatibility between a protein and its DNA target can be calculated by summing up the individual measures of the pairs of amino acid-base involved in the complex, assuming additivity in their contributions to binding. This score enables ranking of different DNA binding sites given a protein binding site and vice versa and can be used in molecular design protocols. We demonstrate its validity by comparing the predictions using this score with experimental binding results of sequence variants of zif268 zinc fingers and their DNA binding sites.

Amino Acids↗

A role for CH...O interactions in protein-DNA recognition.

The concept of CH...O hydrogen bonds has recently gained much interest, with a number of reports indicating the significance of these non-classical hydrogen bonds in stabilizing nucleic acid and protein structures. Here, we analyze the CH...O interactions in the protein-DNA interface, based on 43 crystal structures of protein-DNA complexes. Surprisingly, we find that the number of close intermolecular CH...O contacts involving the thymine methyl group and position C5 of cytosine is comparable to the number of protein-DNA hydrogen bonds involving nitrogen and oxygen atoms as donors and acceptors. A comprehensive analysis of the geometries of these close contacts shows that they are similar to other CH...O interactions found in proteins and small molecules, as well as to classical NH...O hydrogen bonds. Thus, we suggest that C5 of cytosine and C5-Met of thymine form relatively weak CH...O hydrogen bonds with Asp, Asn, Glu, Gln, Ser, and Thr, contributing to the specificity of recognition. Including these interactions, in addition to the classical protein-DNA hydrogen bonds, enables the extraction of simple structural principles for amino acid-base recognition consistent with electrostatic considerations.

Base Composition↗

Knowledge-based structure prediction of MHC class I bound peptides: a study of 23 complexes.

BACKGROUND: The binding of T-cell antigenic peptides to MHC molecules is a prerequisite for their immunogenicity. The ability to identify binding peptides based on the protein sequence is of great importance to the rational design of peptide vaccines. As the requirements for peptide binding cannot be fully explained by the peptide sequence per se, structural considerations should be taken into account and are expected to improve predictive algorithms. The first step in such an algorithm requires accurate and fast modeling of the peptide structure in the MHC-binding groove. RESULTS: We have used 23 solved peptide-MHC class I complexes as a source of structural information in the development of a modeling algorithm. The peptide backbones and MHC structures were used as the templates for prediction. Sidechain conformations were built based on a rotamer library, using the 'dead end elimination' approach. A simple energy function selects the favorable combination of rotamers for a given sequence. It further selects the correct backbone structure from a limited library. The influence of different parameters on the prediction quality was assessed. With a specific rotamer library that incorporates information from the peptide sidechains in the solved complexes, the algorithm correctly identifies 85% (92%) of all (buried) sidechains and selects the correct backbones. Under cross-validation, 70% (78%) of all (buried) residues are correctly predicted and most of all backbones. The interaction between peptide sidechains has a negligible effect on the prediction quality. CONCLUSIONS: The structure of the peptide sidechains follows from the interactions with the MHC and the peptide backbone, as the prediction is hardly influenced by sidechain interactions. The proposed methodology was able to select the correct backbone from a limited set. The impairment in performance under cross-validation suggests that, currently, the specific rotamer library is not satisfactorily representative. The predictions might improve with an increase in the data.

Animals↗

A structure-based algorithm to predict potential binding peptides to MHC molecules with hydrophobic binding pockets.

Binding of peptides to MHC class I molecules is a prerequisite for their recognition by cytotoxic T cells. Consequently, identification of peptides that will bind to a given MHC molecule must constitute a central part of any algorithm for prediction of T-cell antigenic peptides based on the amino acid sequence of the protein. Binding motifs, defined by anchor positions only, have proven to be insufficient to ensure binding, suggesting that other positions along the peptide sequence also affect peptide-MHC interaction. The second phase of prediction schemes therefore take into account the effect of all positions along the peptide sequence, and are based on position-dependent-coefficients that are used in the calculation of a peptide score. These coefficients can be extracted from a large ensemble of binding sequences that were tested experimentally, or derived from structural considerations, as in the algorithm developed by us recently. This algorithm uses the coordinates of solved complexes to evaluate the interactions of peptide amino acids with MHC contact residues, and results in a peptide score that reflects its binding energy. Here we present our analysis for peptide binding to four MHC alleles (HLA-A2, HLA-A68, HLA-B27 and H-2Kb), and compare the predictions of the algorithm to experimental binding data. The algorithm performs successfully in predicting peptide binding to MHC molecules with hydrophobic binding pockets but not when MHC molecules with hydrophilic, charged pockets are considered. For MHC molecules with hydrophobic pockets it is demonstrated how the algorithm succeeds in distinguishing binding from non-binding peptides, and in high ranking of immunogenic peptides within all overlapping same-length peptides spanning their respective protein sequences. The latter property of the algorithm makes it a useful tool in the rational design of peptide vaccines aimed at T-cell immunity.

Algorithms↗

Microsatellite spreading in the human genome: evolutionary mechanisms and structural implications.

Microsatellites are tandem repeat sequences abundant in the genomes of higher eukaryotes and hitherto considered as "junk DNA." Analysis of a human genome representative data base (2.84 Mb) reveals a distinct juxtaposition of A-rich microsatellites and retroposons and suggests their coevolution. The analysis implies that most microsatellites were generated by a 3'-extension of retrotranscripts, similar to mRNA polyadenylylation, and that they serve in turn as "retroposition navigators," directing the retroposons via homology-driven integration into defined sites. Thus, they became instrumental in the preservation and extension of primordial genomic patterns. A role is assigned to these reiterating A-rich loci in the higher-order organization of the chromatin. The disease-associated triplet repeats are mostly found in coding regions and do not show an association with retroposons, constituting a unique set within the family of microsatellite sequences.

Biological Evolution↗