Search PubMed⌕ Search

Biomedical subjects

Hanah Margalit

Publications and source records attributed to Hanah Margalit.

10 recordsLinked to original sources

IntAct: an open source molecular interaction database.

IntAct provides an open source database and toolkit for the storage, presentation and analysis of protein interactions. The web interface provides both textual and graphical representations of protein interactions, and allows exploring interaction networks in the context of the GO annotations of the interacting proteins. A web service allows direct computational access to retrieve interaction networks in XML format. IntAct currently contains approximately 2200 binary and complex interactions imported from the literature and curated in collaboration with the Swiss-Prot team, making intensive use of controlled vocabularies to ensure data consistency. All IntAct software, data and controlled vocabularies are available at http://www.ebi.ac.uk/intact.

Animals↗

Detection of regulatory circuits by integrating the cellular networks of protein-protein interactions and transcription regulation.

The post-genomic era is marked by huge amounts of data generated by large-scale functional genomic and proteomic experiments. A major challenge is to integrate the various types of genome-scale information in order to reveal the intra- and inter- relationships between genes and proteins that constitute a living cell. Here we present a novel application of classical graph algorithms to integrate the cellular networks of protein-protein interactions and transcription regulation. We demonstrate how integration of these two networks enables the discovery of simple as well as complex regulatory circuits that involve both protein-protein and protein-DNA interactions. These circuits may serve for positive or negative feedback mechanisms. By applying our approach to data from the yeast Saccharomyces cerevisiae, we were able to identify known simple and complex regulatory circuits and to discover many putative circuits, whose biological relevance has been assessed using various types of experimental data. The newly identified relations provide new insight into the processes that take place in the cell, insight that could not be gained by analyzing each type of data independently. The computational scheme that we propose may be used to integrate additional functional genomic and proteomic data and to reveal other types of relations, in yeast as well as in higher organisms.

DNA, Fungal↗

How reliable are experimental protein-protein interaction data?

Data of protein-protein interactions provide valuable insight into the molecular networks underlying a living cell. However, their accuracy is often questioned, calling for a rigorous assessment of their reliability. The computation offered here provides an intelligible mean to assess directly the rate of true positives in a data set of experimentally determined interacting protein pairs. We show that the reliability of high-throughput yeast two-hybrid assays is about 50%, and that the size of the yeast interactome is estimated to be 10,000-16,600 interactions.

Protein Binding↗

Conserved sequence elements associated with exon skipping.

One of the major forms of alternative splicing, which generates multiple mRNA isoforms differing in the precise combinations of their exon sequences, is exon skipping. While in constitutive splicing all exons are included, in the skipped pattern(s) one or more exons are skipped. The regulation of this process is still not well understood; so far, cis- regulatory elements (such as exonic splicing enhancers) were identified in individual cases. We therefore set to investigate the possibility that exon skipping is controlled by sequences in the adjacent introns. We employed a computer analysis on 54 sequences documented as undergoing exon skipping, and identified two motifs both in the upstream and downstream introns of the skipped exons. One motif is highly enriched in pyrimidines (mostly C residues), and the other motif is highly enriched in purines (mostly G residues). The two motifs differ from the known cis-elements present at the 5' and 3' splice site. Interestingly, the two motifs are complementary, and their relative positional order is conserved in the flanking introns. These suggest that base pairing interactions can underlie a mechanism that involves secondary structure to regulate exon skipping. Remarkably, the two motifs are conserved in mouse orthologous genes that undergo exon skipping.

Alternative Splicing↗

A survey of small RNA-encoding genes in Escherichia coli.

Small RNA (sRNA) molecules have gained much interest lately, as recent genome-wide studies have shown that they are widespread in a variety of organisms. The relatively small family of 10 known sRNA-encoding genes in Escherichia coli has been significantly expanded during the past two years with the discovery of 45 novel genes. Most of these genes are still uncharacterized and their cellular roles are unknown. In this survey we examined the sequence and genomic features of the 55 currently known sRNA-encoding genes in E.coli, attempting to identify their common characteristics. Such characterization is important for both expanding our understanding of this unique gene family and for improving the methods to predict and identify sRNA-encoding genes based on genomic information.

Base Composition↗

Hierarchy of sequence-dependent features associated with prokaryotic translation.

Protein expression in the cell is affected by various sequence-dependent features. Several such sequence-dependent features have been individually studied,yet they have not been compared quantitatively in terms of their relative influence on protein expression,and a hierarchy of these elements has not been determined. Here we present a quantitative analysis examining sequence-dependent features involved in prokaryotic translation,namely,the base-pairing potential between the mRNA Shine-Dalgarno sequence and the ribosomal RNA,codon bias,and the identity of the stop codon. We analyzed these features both at intra- and intergenomic levels using the Escherichia coli and Haemophilus influenzae genomes. Within each genome,we examined the relationship between each feature and protein expression levels determined by 2D-gel analyses. At the intergenomic level,comparative genomic principles were applied to study the relative preservation of the different sequence-dependent properties between orthologs. From these analyses,we determined that biased codon usage is the property that is most highly associated with protein expression and that is most conserved. The identity of the stop codon and the base-pairing potential of the mRNA Shine-Dalgarno sequence and the rRNA seem to have less of an effect on protein expression.

3' Untranslated Regions↗

Molecular basis for expression of common and rare fragile sites.

Fragile sites are specific loci that form gaps, constrictions, and breaks on chromosomes exposed to partial replication stress and are rearranged in tumors. Fragile sites are classified as rare or common, depending on their induction and frequency within the population. The molecular basis of rare fragile sites is associated with expanded repeats capable of adopting unusual non-B DNA structures that can perturb DNA replication. The molecular basis of common fragile sites was unknown. Fragile sites from R-bands are enriched in flexible sequences relative to nonfragile regions from the same chromosomal bands. Here we cloned FRA7E, a common fragile site mapped to a G-band, and revealed a significant difference between its flexibility and that of nonfragile regions mapped to G-bands, similar to the pattern found in R-bands. Thus, in the entire genome, flexible sequences might play a role in the mechanism of fragility. The flexible sequences are composed of interrupted runs of AT-dinucleotides, which have the potential to form secondary structures and hence can affect replication. These sequences show similarity to the AT-rich minisatellite repeats that underlie the fragility of the rare fragile sites FRA16B and FRA10B. We further demonstrate that the normal alleles of FRA16B and FRA10B span the same genomic regions as the common fragile sites FRA16C and FRA10E. Our results suggest that a shared molecular basis, conferred by sequences with a potential to form secondary structures that can perturb replication, may underlie the fragility of rare fragile sites harboring AT-rich minisatellite repeats and aphidicolin-induced common fragile sites.

Alleles↗

Insights from MHC-bound peptides.

Cytotoxic T cells recognize short antigenic peptides, the processing products of protein antigens, when they are bound to major histocompatibility complex (MHC) class I molecules. Peptide binding to MHC molecules has been studied extensively in numerous laboratories, providing vast amounts of sequence and structure data that have been used as a rich source for bioinformatic research. MHC-bound peptides and their flanking sequences provide information about the sequence requirements of the different processing stages, in particular, the cleavage by the proteasome and the binding to MHC molecules. Elucidation of these sequence requirements sheds light on the evolutionary forces that have shaped and designed these peptides, and should lead to the development of an integrative predictive algorithm. Remarkably, the peptide sequence and structure data are also valuable for the study of biological questions that are apparently unrelated to cellular immunity, namely, sequence-structure relationship and genome annotation. Here we describe our computational analyses of MHC-bound peptides, applied to all these biological topics.

Allergy and Immunology↗

PeCoP: automatic determination of persistently conserved positions in protein families.

UNLABELLED: PeCoP is a WWW-based service which accepts a protein sequence, and reports positions that are conserved in close and distant sequence family members. The collation of family members is performed using iterative PSI-BLAST runs. Examining positional conservation in close and distant family members enables a better selection of positions that may play a role in determining the protein's structure and function. AVAILABILITY: http://bioinformatics.org/pecop CONTACT: idoerg@burnham.org SUPPLEMENTARY INFORMATION: http://bioinformatics.org/pecop/about_pecop

Amino Acid Sequence↗

Persistently conserved positions in structurally similar, sequence dissimilar proteins: roles in preserving protein fold and function.

Many protein pairs that share the same fold do not have any detectable sequence similarity, providing a valuable source of information for studying sequence-structure relationship. In this study, we use a stringent data set of structurally similar, sequence-dissimilar protein pairs to characterize residues that may play a role in the determination of protein structure and/or function. For each protein in the database, we identify amino-acid positions that show residue conservation within both close and distant family members. These positions are termed "persistently conserved". We then proceed to determine the "mutually" persistently conserved (MPC) positions: those structurally aligned positions in a protein pair that are persistently conserved in both pair mates. Because of their intra- and interfamily conservation, these positions are good candidates for determining protein fold and function. We find that 45% of the persistently conserved positions are mutually conserved. A significant fraction of them are located in critical positions for secondary structure determination, they are mostly buried, and many of them form spatial clusters within their protein structures. A substitution matrix based on the subset of MPC positions shows two distinct characteristics: (i) it is different from other available matrices, even those that are derived from structural alignments; (ii) its relative entropy is high, emphasizing the special residue restrictions imposed on these positions. Such a substitution matrix should be valuable for protein design experiments.

Amino Acid Motifs↗