Search PubMed⌕ Search

Biomedical subjects

Sándor Pongor

Publications and source records attributed to Sándor Pongor.

At least 19 recordsLinked to original sources

A Protein Classification Benchmark collection for machine learning.

Protein classification by machine learning algorithms is now widely used in structural and functional annotation of proteins. The Protein Classification Benchmark collection (http://hydra.icgeb.trieste.it/benchmark) was created in order to provide standard datasets on which the performance of machine learning methods can be compared. It is primarily meant for method developers and users interested in comparing methods under standardized conditions. The collection contains datasets of sequences and structures, and each set is subdivided into positive/negative, training/test sets in several ways. There is a total of 6405 classification tasks, 3297 on protein sequences, 3095 on protein structures and 10 on protein coding regions in DNA. Typical tasks include the classification of structural domains in the SCOP and CATH databases based on their sequences or structures, as well as various functional and taxonomic classification problems. In the case of hierarchical classification schemes, the classification tasks can be defined at various levels of the hierarchy (such as classes, folds, superfamilies, etc.). For each dataset there are distance matrices available that contain all vs. all comparison of the data, based on various sequence or structure comparison methods, as well as a set of classification performance measures computed with various classifier algorithms.

Algorithms↗

Application of a simple likelihood ratio approximant to protein sequence classification.

MOTIVATION: Likelihood ratio approximants (LRA) have been widely used for model comparison in statistics. The present study was undertaken in order to explore their utility as a scoring (ranking) function in the classification of protein sequences. RESULTS: We used a simple LRA-based on the maximal similarity (or minimal distance) scores of the two top ranking sequence classes. The scoring methods (Smith-Waterman, BLAST, local alignment kernel and compression based distances) were compared on datasets designed to test sequence similarities between proteins distantly related in terms of structure or evolution. It was found that LRA-based scoring can significantly outperform simple scoring methods.

Algorithms↗

Gene synthesis, expression, purification, and characterization of human Jagged-1 intracellular region.

Notch signaling plays a key role in cell differentiation and is very well conserved from Drosophila to humans. Ligands of Notch receptors are type I, membrane spanning proteins composed of a large extracellular region and a 100-150 residue cytoplasmic tail. We report here, for the first time, the expression, purification, and characterization of the intracellular region of a Notch ligand. Starting from a set of synthetic oligonucleotides, we assembled a synthetic gene optimized for Escherichia coli codon usage and encoding the cytoplasmic region of human Jagged-1 (residues 1094-1218). The protein containing a N-terminal His(6)-tag was over-expressed in E. coli, and purified by affinity and reversed phase chromatography. After cleavage of the His(6)-tag by a dipeptidyl aminopeptidase, the protein was purified to homogeneity and characterized by spectroscopic techniques. Far-UV circular dichroism, fluorescence emission spectra, fluorescence anisotropy measurements, and (1)H nuclear magnetic resonance spectra, taken together, suggest that the cytoplasmic tail of human Jagged-1 behaves as an intrinsically unstructured domain in solution. This result was confirmed by the high susceptibility of the recombinant protein to proteolytic cleavage. The significance of this finding is discussed in relation to the recently proposed role of the intracellular region of Notch ligands in bi-directional signaling.

Calcium-Binding Proteins↗

Application of compression-based distance measures to protein sequence classification: a methodological study.

MOTIVATION: Distance measures built on the notion of text compression have been used for the comparison and classification of entire genomes and mitochondrial genomes. The present study was undertaken in order to explore their utility in the classification of protein sequences. RESULTS: We constructed compression-based distance measures (CBMs) using the Lempel-Zlv and the PPMZ compression algorithms and compared their performance with that of the Smith-Waterman algorithm and BLAST, using nearest neighbour or support vector machine classification schemes. The datasets included a subset of the SCOP protein structure database to test distant protein similarities, a 3-phosphoglycerate-kinase sequences selected from archaean, bacterial and eukaryotic species as well as low and high-complexity sequence segments of the human proteome, CBMs values show a dependence on the length and the complexity of the sequences compared. In classification tasks CBMs performed especially well on distantly related proteins where the performance of a combined measure, constructed from a CBM and a BLAST score, approached or even slightly exceeded that of the Smith-Waterman algorithm and two hidden Markov model-based algorithms.

Algorithms↗

The "first in-last out" hypothesis on protein folding revisited.

We calculated profiles for mean residue depth, contact order, and number of contacts in the native structure of a series of proteins for which folding has been studied extensively, the chymotrypsin inhibitor 2, the SH3 module from the src tyrosine kinase, the small ribonuclease barnase, the bacterial immunity protein Im7, and apomyoglobin. We compared these profiles with experimental data from equilibrium or pulse labeling hydrogen-deuterium exchange obtained from NMR and phi values obtained from the protein engineering approach. We find a good qualitative agreement between the hierarchy of formation of topological elements during the folding process and the ranking of secondary structure elements in terms of residue depth. Residues that are most deeply buried in the core of the native protein usually belong to stretches of secondary structure elements that are formed early in the folding pathway. Residue depth can thus provide a useful and simple tool for the design of folding experiments.

Algorithms↗

CX, DPX and PRIDE: WWW servers for the analysis and comparison of protein 3D structures.

The WWW servers at http://www.icgeb.org/protein/ are dedicated to the analysis of protein 3D structures submitted by the users as the Protein Data Bank (PDB) files. CX computes an atomic protrusion index that makes it possible to highlight the protruding atoms within a protein 3D structure. DPX calculates a depth index for the buried atoms and makes it possible to analyze the distribution of buried residues. CX and DPX return PDB files containing the calculated indices that can then be visualized using standard programs, such as Swiss-PDBviewer and Rasmol. PRIDE compares 3D structures using a fast algorithm based on the distribution of inter-atomic distances. The options include pairwise as well as multiple comparisons, and fold recognition based on searching the CATH fold database.

Algorithms↗

Multiple weak hits confuse complex systems: a transcriptional regulatory network as an example.

Robust systems, like the molecular networks of living cells, are often resistant to single hits such as those caused by high-specificity drugs. Here we show that partial weakening of the Escherichia coli and Saccharomyces cerevisiae transcriptional regulatory networks at a small number (3-5) of selected nodes can have a greater impact than the complete elimination of a single selected node. In both cases, the targeted nodes have the greatest possible impact; still, the results suggest that in some cases broader specificity compounds or multitarget drug therapies may be more effective than individual high-affinity, high-specificity ones. Multiple but partial attacks mimic well a number of in vivo scenarios and may be useful in the efficient modification of other complex systems.

Computer Simulation↗

Efficient recognition of folds in protein 3D structures by the improved PRIDE algorithm.

UNLABELLED: An improved version of the PRIDE (PRobaility of IDEntity) fold prediction algorithm has been developed, based on more solid statistical basis, fast search capabilities and efficient input structure processing. The new algorithm is effective in identifying protein structures at the 'H' level of the CATH hierarchy. AVAILABILITY: The new algorithm is integrated into the PRIDE2 web servers at http://pride.szbk.u-szeged.hu and http://www.icgeb.org/pride. SUPPLEMENTARY INFORMATION: Detailed documentation and performance evaluation is available in the description section of the PRIDE2 web server.

Algorithms↗

Graph-representation of oxidative folding pathways.

BACKGROUND: The process of oxidative folding combines the formation of native disulfide bond with conformational folding resulting in the native three-dimensional fold. Oxidative folding pathways can be described in terms of disulfide intermediate species (DIS) which can also be isolated and characterized. Each DIS corresponds to a family of folding states (conformations) that the given DIS can adopt in three dimensions. RESULTS: The oxidative folding space can be represented as a network of DIS states interconnected by disulfide interchange reactions that can either create/abolish or rearrange disulfide bridges. We propose a simple 3D representation wherein the states having the same number of disulfide bridges are placed on separate planes. In this representation, the shuffling transitions are within the planes, and the redox edges connect adjacent planes. In a number of experimentally studied cases (bovine pancreatic trypsin inhibitor, insulin-like growth factor and epidermal growth factor), the observed intermediates appear as part of contiguous oxidative folding pathways. CONCLUSIONS: Such networks can be used to visualize folding pathways in terms of the experimentally observed intermediates. A simple visualization template written for the Tulip package http://www.tulip-software.org/ can be obtained from V.A.

Animals↗

The SBASE domain sequence resource, release 12: prediction of protein domain-architecture using support vector machines.

SBASE (http://www.icgeb.trieste.it/sbase) is an online resource designed to facilitate the detection of domain homologies based on sequence database search. The present release of the SBASE A library of protein domain sequences contains 972,397 protein sequence segments annotated by structure, function, ligand-binding or cellular topology, clustered into 8547 domain groups. SBASE B contains 169,916 domain sequences clustered into 2526 less well-characterized groups. Domain prediction is based on an evaluation of database search results in comparison with a 'similarity network' of inter-sequence similarity scores, using support vector machines trained on similarity search results of known domains.

Artificial Intelligence↗

Efficient synthesis and comparative studies of the arginine and Nomega,Nomega-dimethylarginine forms of the human nucleolin glycine/arginine rich domain.

The Gly- and Arg-rich C-terminal region of human nucleolin is a 61-residue long domain involved in a number of protein-protein and protein-nucleic acid interactions. This domain contains 10 aDma residues in the form of aDma-GG repeats interspersed with Phe residues. The exact role of Arg dimethylation is not known, partly because of the lack of efficient synthetic methods. This work describes an effective synthetic strategy, generally applicable to long RGG peptides, based on side-chain protected aDma and backbone protected dipeptide Fmoc-Gly-(Dmob)Gly-OH. This strategy allowed us to synthesize both the unmodified (N61Arg) and the dimethylated (N61aDma) peptides with high yield ( approximately 26%) and purity. As detected by NMR spectroscopy, N61Arg does not possess any stable secondary or tertiary structure in solution and N(omega),N(omega)-dimethylation of the guanidino group does not alter the overall conformational propensity of this peptide. While both peptides bind single-stranded nucleic acids with similar affinities (K(d) = 1.5 x 10(-7) M), they exhibit a different behaviour in ssDNA affinity chromatography consistent with the difference in pK(a) values. It has been previously shown that N61Arg inhibits HIV infection at the stage of HIV attachment to cells. This study demonstrates that Arg-dimethylated C-terminal domain lacks any inhibition activity, raising the question of whether nucleolin expressed on the cell-surface is indeed dimethylated.

Adenosine Triphosphatases↗

The efficiency of multi-target drugs: the network approach might help drug design.

Despite considerable progress in genome- and proteome-based high-throughput screening methods and rational drug design, the number of successful single-target drugs did not increase appreciably during the past decade. Network models suggest that partial inhibition of a surprisingly small number of targets can be more efficient than the complete inhibition of a single target. This and the success stories of multi-target drugs and combinatorial therapies led us to suggest that systematic drug-design strategies should be directed against multiple targets. We propose that the final effect of partial, but multiple, drug actions might often surpass that of complete drug action at a single target. The future success of this novel drug-design paradigm will depend not only on a new generation of computer models to identify the correct multiple targets and their multi-fitting, low-affinity drug candidates but also on more-efficient in vivo testing.

Computational Biology↗

DNA-mediated assembly of weakly interacting DNA-binding protein subunits: in vitro recruitment of phage 434 repressor and yeast GCN4 DNA-binding domains.

The specificity of DNA-mediated protein assembly was studied in two in vitro systems, based on (i) the DNA-binding domain of bacteriophage 434 repressor cI (amino acid residues 1-69), or (ii) the DNA-binding domain of the yeast transcription factor GCN4, (amino acids 1-34) and their respective oligonucleotide cognates. In vivo, both of these peptides are part of larger protein molecules that also contain dimerization domains, and the resulting dimers recognize cognate palindromic DNA sequences that contain two half-sites of 4 bp each. The dimerization domains were not included in the peptides tested, so in solution-in the presence or absence of non-cognate DNA oligonucleotides-these molecules did not show appreciable dimerization, as determined by pyrene excimer fluorescence spectroscopy and oxidative cross-linking monitored by mass spectrometry. Oligonucleotides with only one 4 bp cognate half-site were able to initiate measurable dimerization, and two half-sites were able to select specific dimers even from a heterogeneous pool of molecules of closely related specificity (such as DNA-binding domains of the 434 repressor and their engineered mutants that mimic the binding helix of the related P22 phage repressor). The fluorescent technique allowed us to separately monitor the unspecific, ionic interaction of the peptides with DNA which produced a roughly similar signal in the case of both cognate and non-cognate oligonucleotides. But in the former case, a concomitant excimer fluorescence signal showed the formation of correctly positioned dimers. The results suggest that DNA acts as a highly specific template for the recruitment of weakly interacting protein molecules that can thus build up highly specific complexes.

Amino Acid Sequence↗

Exon 6 of human Jagged-1 encodes an autonomously folding unit.

Human Jagged-1 is predicted to contain 16 epidermal growth factor-like (EGF) repeats. The oxidative folding of EGF-2, despite the several conditions tested, systematically led to complex mixtures. A longer peptide spanning the C-terminal part of EGF-1 and the complete EGF-2 repeat, on the contrary, could be readily refolded. This peptide, which corresponds to the entire exon 6 of the Jagged-1 gene, thus represents an autonomously folding unit. We show that it is structured in solution, as suggested by circular dichroism and NMR spectroscopy, and displays an EGF-like disulfide bond topology, as determined by disulfide mapping.

Amino Acid Sequence↗

Vicinal disulfide bridge conformers by experimental methods and by ab initio and DFT molecular computations.

A systematic comparison is made between experimental and computational data gained on vicinal disulfide bridges in proteins and peptides. Structural and stability data of ab initio and density functional theory (DFT) calculations on the model compound 4,5-ditiaheptano-7-lactam and the model peptide HCO-ox-[Cys-Cys]-NH2 at RHF/3-21G*, B3LYP/6-31+G(d), and B3LYP/6-311++G(d,p) levels of theory are presented. The data on Xxx-Cys-Cys-Yyy type amino acid sequence units retrieved from PDB SELECT, along with data on sequence units that have vicinal disulfide bridge, taken from the Brookhaven Protein Data Bank, are conformationally characterized. Amino acid backbone conformations, cis-trans isomerism of the amide bond between the two cysteine residues, and ring puckering are studied. Ring puckers are characterized by their relation to the conformers of the parent 4,5-ditiaheptano-7-lactam. Computational precision and accuracy are proved by frequency calculation and solvent model optimization on selected conformers. It is found that the ox-[Cys-Cys] unit is able to accept types I, II, VIa, VIb, and VIII beta-turn structures.

Amino Acids↗

Oxidative folding of Amaranthus alpha-amylase inhibitor: disulfide bond formation and conformational folding.

Oxidative folding is the fusion of native disulfide bond formation with conformational folding. This complex process is guided by two types of interactions: first, covalent interactions between cysteine residues, which transform into native disulfide bridges, and second, non-covalent interactions giving rise to secondary and tertiary protein structure. The aim of this work is to understand both types of interactions in the oxidative folding of Amaranthus alpha-amylase inhibitor (AAI) by providing information both at the level of individual disulfide species and at the level of amino acid residue conformation. The cystine-knot disulfides of AAI protein are stabilized in an interdependent manner, and the oxidative folding is characterized by a high heterogeneity of one-, two-, and three-disulfide intermediates. The formation of the most abundant species, the main folding intermediate, is favored over other species even in the absence of non-covalent sequential preferences. Time-resolved NMR and photochemically induced dynamic nuclear polarization spectroscopies were used to follow the oxidative folding at the level of amino acid residue conformation. Because this is the first time that a complete oxidative folding process has been monitored with these two techniques, their results were compared with those obtained at the level of an individual disulfide species. The techniques proved to be valuable for the study of conformational developments and aromatic accessibility changes along oxidative folding pathways. A detailed picture of the oxidative folding of AAI provides a model study that combines different biochemical and biophysical techniques for a fuller understanding of a complex process.

Amaranthus↗

Fragmentation pathways of N(G)-methylated and unmodified arginine residues in peptides studied by ESI-MS/MS and MALDI-MS.

Protein methylation at arginine residues is a prevalent posttranslational modification in eukaryotic cells that has been implicated in processes from RNA-binding and transporting to protein sorting and transcription activation. Three main forms of methylarginine have been identified: N(G)-monomethylarginine (MMA), asymmetric N(G),N(G)-dimethylarginine (aDMA), and symmetric N(G),N'(G)-dimethylarginine (sDMA). To investigate gas-phase fragmentations and characteristic ions arising from methylated and unmodified arginine residues in detail, we subjected peptides containing these residues to electrospray triple-quadrupole tandem mass spectrometry. A variety of low mass ions including (methylated) ammonium, carbodiimidium, and guanidinium ions were observed. Fragment ions resulting from the loss of the corresponding neutral fragments (amines, carbodiimide, and guanidine) from intact molecular ions as well as from N- and C-terminal fragment ions were also identified. Furthermore, the peptides containing either methylated or unmodified arginines gave rise to abundant fragment ions at m/z 70, 112, and 115, for which cyclic ion structures are proposed. Electrospray ionization tandem mass spectra revealed that dimethylammonium (m/z 46) is a specific marker ion for aDMA. A precursor ion scanning method utilizing this fragment ion was developed, which allowed sensitive and specific detection of aDMA-containing peptides even in the presence of a five-fold excess of phosphorylase B digest. Interestingly, regular matrix-assisted laser desorption/ionization mass spectra recorded from aDMA- or sDMA-containing peptides showed metastable fragment ions resulting from cleavages of the arginine side chains. The neutral losses of mono- and dimethylamines permit the differentiation between aDMA and sDMA.

Arginine↗

Folding of epidermal growth factor-like repeats from human tenascin studied through a sequence frame-shift approach.

In order to investigate the factors that determine the correct folding of epidermal growth factor-like (EGF) repeats within a multidomain protein, we prepared a series of six peptides that, taken together, span the sequence of two EGF repeats of human tenascin, a large protein from the extracellular matrix. The peptides were selected by sliding a window of the average length of tenascin EGF repeats over the sequence of EGF repeats 13 and 14. We thus obtained six peptides, EGF-f1 to EGF-f6, that are 33 residues long, contain six cysteines each, and bear a partial overlap in the sequence. While EGF-f1 corresponds to the native EGF-14 repeat, the others are frame-shifted EGF repeats. We carried out the oxidative folding of these peptides in vitro, analyzed the reaction mixtures by acid trapping followed by LC-MS, and isolated some of the resulting products. The oxidative folding of the native EGF-14 peptide is fast, produces a single three-disulfide species with an EGF-like disulfide topology and a marked difference in the RP-HPLC retention time compared with the starting product. On the contrary, frame-shifted peptides fold more slowly and give mixtures of three-disulfide species displaying RP-HPLC retention times that are closer to those of the reduced peptides. In contrast to the native EGF-14, the three-disulfide products that could be isolated are mainly unstructured, as determined by CD and NMR spectroscopy. We conclude that both kinetics and thermodynamics drive the correct pairing of cysteines, and speculate about how cysteine mispairing could trigger disulfide reshuffling in vivo.

Amino Acid Sequence↗