Search PubMed⌕ Search

Biomedical subjects

István Simon

Publications and source records attributed to István Simon.

At least 19 recordsLinked to original sources

Metal-binding sites at the active site of restriction endonuclease BamHI can conform to a one-ion mechanism.

The number of metal ions required for phosphoryl transfer in restriction endonucleases is still an unresolved question in molecular biology. The two Ca(2+) and Mn(2+) ions observed in the pre- and post-reactive complexes of BamHI conform to the classical two-metal ion choreography. We probed the Mg(2+) cofactor positions at the active site of BamHI by molecular dynamics simulations with one and two metal ions present and identified several catalytically relevant sites. These can mark the pathway of a single ion during catalysis, suggesting its critical role, while a regulatory function is proposed for a possible second ion.

Binding Sites↗

Flexible segments modulate co-folding of dUTPase and nucleocapsid proteins.

The homotrimeric fusion protein nucleocapsid (NC)-dUTPase combines domains that participate in RNA/DNA folding, reverse transcription, and DNA repair in Mason-Pfizer monkey betaretrovirus infected cells. The structural organization of the fusion protein remained obscured by the N- and C-terminal flexible segments of dUTPase and the linker region connecting the two domains that are invisible in electron density maps. Small-angle X-ray scattering reveals that upon oligonucleotide binding the NC domains adopt the trimeric symmetry of dUTPase. High-resolution X-ray structures together with molecular modeling indicate that fusion with NC domains dramatically alters the conformation of the flexible C-terminus by perturbing the orientation of a critical beta-strand. Consequently, the C-terminal segment is capable of double backing upon the active site of its own monomer and stabilized by non-covalent interactions formed with the N-terminal segment. This co-folding of the dUTPase terminal segments, not observable in other homologous enzymes, is due to the presence of the fused NC domain. Structural and genomic advantages of fusing the NC domain to a shortened dUTPase in betaretroviruses and the possible physiological consequences are envisaged.

Amino Acid Sequence↗

The BiSearch web server.

BACKGROUND: A large number of PCR primer-design softwares are available online. However, only very few of them can be used for the design of primers to amplify bisulfite-treated DNA templates, necessary to determine genomic DNA methylation profiles. Indeed, the number of studies on bisulfite-treated templates exponentially increases as determining DNA methylation becomes more important in the diagnosis of cancers. Bisulfite-treated DNA is difficult to amplify since undesired PCR products are often amplified due to the increased sequence redundancy after the chemical conversion. In order to increase the efficiency of PCR primer-design, we have developed BiSearch web server, an online primer-design tool for both bisulfite-treated and native DNA templates. RESULTS: The web tool is composed of a primer-design and an electronic PCR (ePCR) algorithm. The completely reformulated ePCR module detects potential mispriming sites as well as undesired PCR products on both cDNA and native or bisulfite-treated genomic DNA libraries. Due to the new algorithm of the current version, the ePCR module became approximately hundred times faster than the previous one and gave the best performance when compared to other web based tools. This high-speed ePCR analysis made possible the development of the new option of high-throughput primer screening. BiSearch web server can be used for academic researchers at the http://bisearch.enzim.hu site. CONCLUSION: BiSearch web server is a useful tool for primer-design for any DNA template and especially for bisulfite-treated genomes. The ePCR tool for fast detection of mispriming sites and alternative PCR products in cDNA libraries and native or bisulfite-treated genomes are the unique features of the new version of BiSearch software.

DNA Primers↗

Phosphorylation-induced transient intrinsic structure in the kinase-inducible domain of CREB facilitates its recognition by the KIX domain of CBP.

Phosphorylation at Ser-133 of the kinase inducible domain of CREB (KID) triggers its binding to the KIX domain of CBP via a concomitant coil-to-helix transition. The exact role of this key event is still puzzling: it does not switch between disordered and ordered states, nor its direct interactions fully account for selectivity. Hence, we reasoned that phosphorylation may shift the conformational preferences of KID towards a binding-competent state. To this end we investigated the intrinsic structural properties of the unbound KID in phosphorylated and unphosphorylated forms by simulated annealing and molecular dynamics simulations. Although helical populations show subtle differences, phosphorylation reduces the flexibility of the turn segment connecting the two helices in the complexed structure and induces a transient structural element that corresponds to its bound conformation. It is stabilized by the pSer-133-Arg-131 interaction, which is absent from the unphosphorylated KID. Diminishing this coupling decreases the 3.1 kcal/mol contribution of pSer-133 to the binding free energy (DeltaGbind) of the phosphorylated KID to KIX by 1.1 kcal/mol, as computed in reference to Ser-133. In a binding competent form of the S133E KID mutant, the contribution of Glu-133 to DeltaGbind is by 1.5 kcal/mol smaller than that of pSer, suggesting that altered structural properties due to pSer --> Glu replacement impair the binding affinity. Thus, we propose that phoshorylation contributes to selectivity not merely by the direct interactions of the phosphate group with KIX, but also by promoting the formation of a transient structural element in the highly conserved turn segment.

Amino Acid Sequence↗

Disorder and sequence repeats in hub proteins and their implications for network evolution.

Protein interaction networks display approximate scale-free topology, in which hub proteins that interact with a large number of other proteins determine the overall organization of the network. In this study, we aim to determine whether hubs are distinguishable from other networked proteins by specific sequence features. Proteins of different connectednesses were compared in the interaction networks of Saccharomyces cerevisiae, Drosophila melanogaster, Caenorhabditis elegans, and Homo sapienswith respect to the distribution of predicted structural disorder, sequence repeats, low complexity regions, and chain length. Highly connected proteins ("hub proteins") contained significantly more of, and greater proportion of, these sequence features and tended to be longer overall as compared to less connected proteins. These sequence features provide two different functional means for realizing multiple interactions: (1) extended interaction surface and (2) flexibility and adaptability, providing a mechanism for the same region to bind distinct partners. Our view contradicts the prevailing view that scaling in protein interactomes arose from gene duplication and preferential attachment of equivalent proteins. We propose an alternative evolutionary network specialization process, in which certain components of the protein interactome improved their fitness for binding by becoming longer or accruing regions of disorder and/or internal repeats and have therefore become specialized in network organization.

Amino Acid Sequence↗

Membrane topology of human ABC proteins.

In this review, we summarize the currently available information on the membrane topology of some key members of the human ABC protein subfamilies, and present the predicted domain arrangements. In the lack of high-resolution structures for eukaryotic ABC transporters this topology is based only on prediction algorithms and biochemical data for the location of various segments of the polypeptide chain, relative to the membrane. We suggest that topology models generated by the available prediction methods should only be used as guidelines to provide a basis of experimental strategies for the elucidation of the membrane topology.

ATP-Binding Cassette Transporters↗

Flexibility of prolyl oligopeptidase: molecular dynamics and molecular framework analysis of the potential substrate pathways.

The flexibility of prolyl oligopeptidase has been investigated using molecular dynamics (MD) and molecular framework approaches to delineate the route of the substrate to the active site. The selectivity of the enzyme is mediated by a seven-bladed beta-propeller that in the crystal structure does not indicate the possible passage for the substrate to the catalytic center. Its open topology however, could allow the blades to move apart and let the substrate into the large central cavity. Flexibility analysis of prolyl oligopeptidase structure using the FIRST (Floppy Inclusion and Rigid Substructure Topology) approach and the atomic fluctuations derived from MD simulations demonstrated the rigidity of the propeller domain, which does not permit the substrate to approach the active site through this domain. Instead, a smaller tunnel at the inter-domain region comprising the highly flexible N-terminal segment of the peptidase domain and a facing hydrophilic loop from the propeller (residues 192-205) was identified by cross-correlation analysis and essential dynamics as the only potential pathway for the substrate. The functional importance of the flexible loop has been also verified by kinetic analysis of the enzyme with a split loop. Catalytic effect of engineered disulfide bridges was rationalized by characterizing the concerted motions of the two domains.

Animals↗

SRide: a server for identifying stabilizing residues in proteins.

Residues expected to play key roles in the stabilization of proteins [stabilizing residues (SRs)] are selected by combining several methods based mainly on the interactions of a given residue with its spatial, rather than its sequential neighborhood and by considering the evolutionary conservation of the residues. A residue is selected as a stabilizing residue if it has high surrounding hydrophobicity, high long-range order, high conservation score and if it belongs to a stabilization center. The definition of all these parameters and the thresholds used to identify the SRs are discussed in detail. The algorithm for identifying SRs was originally developed for TIM-barrel proteins [M. M. Gromiha, G. Pujadas, C. Magyar, S. Selvaraj, and I. Simon (2004), Proteins, 55, 316-329] and is now generalized for all proteins of known 3D structure. SRs could be applied in protein engineering and homology modeling and could also help to explain certain folds with significant stability. The SRide server is located at http://sride.enzim.hu.

Amino Acids↗

IUPred: web server for the prediction of intrinsically unstructured regions of proteins based on estimated energy content.

Intrinsically unstructured/disordered proteins and domains (IUPs) lack a well-defined three-dimensional structure under native conditions. The IUPred server presents a novel algorithm for predicting such regions from amino acid sequences by estimating their total pairwise interresidue interaction energy, based on the assumption that IUP sequences do not fold due to their inability to form sufficient stabilizing interresidue interactions. Optional to the prediction are built-in parameter sets optimized for predicting short or long disordered regions and structured domains.

Algorithms↗

Interfacial water as a "hydration fingerprint" in the noncognate complex of BamHI.

The molecular code of specific DNA recognition by proteins as a paradigm in molecular biology remains an unsolved puzzle primarily because of the subtle interplay between direct protein-DNA interaction and the indirect contribution from water and ions. Transformation of the nonspecific, low affinity complex to a specific, high affinity complex is accompanied by the release of interfacial water molecules. To provide insight into the conversion from the loose to the tight form, we characterized the structure and energetics of water at the protein-DNA interface of the BamHI complex with a noncognate sequence and in the specific complex. The fully hydrated models were produced with Grand Canonical Monte Carlo simulations. Proximity analysis shows that water distributions exhibit sequence dependent variations in both complexes and, in particular, in the noncognate complex they discriminate between the correct and the star site. Variations in water distributions control the number of water molecules released from a given sequence upon transformation from the loose to the tight complex as well as the local entropy contribution to the binding free energy. We propose that interfacial waters can serve as a "hydration fingerprint" of a given DNA sequence.

Base Sequence↗

The pairwise energy content estimated from amino acid composition discriminates between folded and intrinsically unstructured proteins.

The structural stability of a protein requires a large number of interresidue interactions. The energetic contribution of these can be approximated by low-resolution force fields extracted from known structures, based on observed amino acid pairing frequencies. The summation of such energies, however, cannot be carried out for proteins whose structure is not known or for intrinsically unstructured proteins. To overcome these limitations, we present a novel method for estimating the total pairwise interaction energy, based on a quadratic form in the amino acid composition of the protein. This approach is validated by the good correlation of the estimated and actual energies of proteins of known structure and by a clear separation of folded and disordered proteins in the energy space it defines. As the novel algorithm has not been trained on unstructured proteins, it substantiates the concept of protein disorder, i.e. that the inability to form a well-defined 3D structure is an intrinsic property of many proteins and protein domains. This property is encoded in their sequence, because their biased amino acid composition does not allow sufficient stabilizing interactions to form. By limiting the calculation to a predefined sequential neighborhood, the algorithm was turned into a position-specific scoring scheme that characterizes the tendency of a given amino acid to fall into an ordered or disordered region. This application we term IUPred and compare its performance with three generally accepted predictors, PONDR VL3H, DISOPRED2 and GlobPlot on a database of disordered proteins.

Amino Acids↗

BiSearch: primer-design and search tool for PCR on bisulfite-treated genomes.

Bisulfite genomic sequencing is the most widely used technique to analyze the 5-methylation of cytosines, the prevalent covalent DNA modification in mammals. The process is based on the selective transformation of unmethylated cytosines to uridines. Then, the investigated genomic regions are PCR amplified, subcloned and sequenced. During sequencing, the initially unmethylated cytosines are detected as thymines. The efficacy of bisulfite PCR is generally low; mispriming and non-specific amplification often occurs due to the T richness of the target sequences. In order to ameliorate the efficiency of PCR, we developed a new primer-design software called BiSearch, available on the World Wide Web. It has the unique property of analyzing the primer pairs for mispriming sites on the bisulfite-treated genome and determines potential non-specific amplification products with a new search algorithm. The options of primer-design and analysis for mispriming sites can be used sequentially or separately, both on bisulfite-treated and untreated sequences. In silico and in vitro tests of the software suggest that new PCR strategies may increase the efficiency of the amplification.

Algorithms↗

PDB_TM: selection and membrane localization of transmembrane proteins in the protein data bank.

PDB_TM is a database for transmembrane proteins with known structures. It aims to collect all transmembrane proteins that are deposited in the protein structure database (PDB) and to determine their membrane-spanning regions. These assignments are based on the TMDET algorithm, which uses only structural information to locate the most likely position of the lipid bilayer and to distinguish between transmembrane and globular proteins. This algorithm was applied to all PDB entries and the results were collected in the PDB_TM database. By using TMDET algorithm, the PDB_TM database can be automatically updated every week, keeping it synchronized with the latest PDB updates. The PDB_TM database is available at http://www.enzim.hu/PDB_TM.

Algorithms↗

TMDET: web server for detecting transmembrane regions of proteins by using their 3D coordinates.

The structure of integral membrane proteins is determined in the absence of the lipid bilayer; consequently the membrane localization of the protein is usually not specified in the corresponding PDB file. Recently, we have developed a new method called TMDET which determines the most possible localization of the membrane relative to the protein structure, and gives the annotation of the membrane embedded parts of the sequence. The entire Protein Data Bank has been scanned by the new TMDET algorithm resulting in the database of structurally determined transmembrane proteins (PDB_TM). Here we present the web interface of the TMDET algorithm to allow scientists to determine the membrane localization of structural data prior to deposition or to analyze model structures.

Algorithms↗

Functionally and structurally relevant residues of enzymes: are they segregated or overlapping?

There is a delicate balance between stability and flexibility needed for enzyme function. To avoid undesirable alteration of the functional properties during the evolutionary optimization of the structural stability under certain circumstances, and vice versa, to avoid unwanted changes of stability during the optimization of the functional properties of proteins, common sense would suggest that parts of the protein structure responsible for stability and parts responsible for function developed and evolved separately. This study shows that nature did not follow this anthropomorphic logic: the set of residues involved in function and those involved in structural stabilization of enzymes are rather overlapping than segregated.

Amino Acids↗

Transmembrane proteins in the Protein Data Bank: identification and classification.

MOTIVATION: Integral membrane proteins play important roles in living cells. Although these proteins are estimated to constitute 25% of proteins at a genomic scale, the Protein Data Bank (PDB) contains only a few hundred membrane proteins due to the difficulties with experimental techniques. The presence of transmembrane proteins in the structure data bank, however, is quite invisible, as the annotation of these entries is rather poor. Even if a protein is identified as a transmembrane one, the possible location of the lipid bilayer is not indicated in the PDB because these proteins are crystallized without their natural lipid bilayer, and currently no method is publicly available to detect the possible membrane plane using the atomic coordinates of membrane proteins. RESULTS: Here, we present a new geometrical approach to distinguish between transmembrane and globular proteins using structural information only and to locate the most likely position of the lipid bilayer. An automated algorithm (TMDET) is given to determine the membrane planes relative to the position of atomic coordinates, together with a discrimination function which is able to separate transmembrane and globular proteins even in cases of low resolution or incomplete structures such as fragments or parts of large multi chain complexes. This method can be used for the proper annotation of protein structures containing transmembrane segments and paves the way to an up-to-date database containing the structure of all known transmembrane proteins and fragments (PDB_TM) which can be automatically updated. The algorithm is equally important for the purpose of constructing databases purely of globular proteins.

Algorithms↗

Preformed structural elements feature in partner recognition by intrinsically unstructured proteins.

Intrinsically unstructured proteins (IUPs) are devoid of extensive structural order but often display signs of local and limited residual structure. To explain their effective functioning, we reasoned that such residual structure can be crucial in their interactions with their structured partner(s) in a way that preformed structural elements presage their final conformational state. To check this assumption, a database of 24 IUPs with known 3D structures in the bound state has been assembled and the distribution of secondary structure elements and backbone torsion angles have been analysed. The high proportion of residues in coil conformation and with phi, psi angles in the disallowed regions of the Ramachandran map compared to the reference set of globular proteins shows that IUPs are not fully ordered even in their bound form. To probe the effect of partner proteins on IUP folding, inherent conformational preferences of IUP sequences have been assessed by secondary structure predictions using the GOR, ALB and PROF algorithms. The accuracy of predicting secondary structure elements of IUPs is similar to that of their partner proteins and is significantly higher than the corresponding values for random sequences. We propose that strong conformational preferences mark regions in IUPs (mostly helices), which correspond to their final structural state, while regions with weak conformational preferences represent flexible linkers between them. In our interpretation, preformed elements could serve as initial contact points, the binding of which facilitates the reeling of the flexible regions onto the template. This finding implies that IUPs draw a functional advantage from preformed structural elements, as they enable their facile, kinetically and energetically less demanding, interaction with their physiological partner.

Databases, Protein↗

Servers for sequence-structure relationship analysis and prediction.

We describe several algorithms and public servers that were developed to analyze and predict various features of protein structures. These servers provide information about the covalent state of cysteine (CYSREDOX), as well as about residues involved in non-covalent cross links that play an important role in the structural stability of proteins (SCIDE and SCPRED). We also discuss methods and servers developed to identify helical transmembrane proteins from large databases and rough genomic data, including two of the most popular transmembrane prediction methods, DAS and HMMTOP. Several biologically interesting applications of these servers are also presented. The servers are available through http://www.enzim.hu/servers.html.

Algorithms↗