Search PubMed⌕ Search

Biomedical subjects

Pier Luigi Martelli

Publications and source records attributed to Pier Luigi Martelli.

At least 19 recordsLinked to original sources

On the state of protein function prediction: a report on the fourth CAFA challenge.

BACKGROUND: The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). RESULTS: CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

Journal Article↗

Thinking the impossible: how to solve the protein folding problem with and without homologous structures and more.

Structure prediction of proteins is a difficult task as well as prediction of protein-protein interaction. When no homologous sequence with known structure is available for the target protein, search of distantly related proteins to the target may be done automatically (fold recognition/threading). However, there are difficult proteins for which still modeling on the basis of a putative scaffold is nearly impossible. In the following, we describe that for some specific examples, human expertise was able to derive alignments to proteins of similar function with the aid of machine learning-based methods specifically suited for predicting structural features. The manually curate search of putative templates was successful in generating low-resolution three-dimensional (3D) models in at least two cases: the human tissue transglutaminase and the alcohol dehydrogenase from Sulfolobus solfataricus. This is based on the structural comparison of the model with the 3D protein structure that became available after prediction. For protein-protein interaction, a knowledge-based method can give predictions of putative interaction patches on the protein surface; this feature may help in adding additional weight to specific nodes in nets of interacting proteins.

Alcohol Dehydrogenase↗

eSLDB: eukaryotic subcellular localization database.

Eukaryotic Subcellular Localization DataBase collects the annotations of subcellular localization of eukaryotic proteomes. So far five proteomes have been processed and stored: Homo sapiens, Mus musculus, Caenorhabditis elegans, Saccharomyces cerevisiae and Arabidopsis thaliana. For each sequence, the database lists localization obtained adopting three different approaches: (i) experimentally determined (when available); (ii) homology-based (when possible); and (iii) predicted. The latter is computed with a suite of machine learning based methods, developed in house. All the data are available at our website and can be searched by sequence, by protein code and/or by protein description. Furthermore, a more complex search can be performed combining different search fields and keys. All the data contained in the database can be freely downloaded in flat file format. The database is available at http://gpcr.biocomp.unibo.it/esldb/.

Animals↗

BaCelLo: a balanced subcellular localization predictor.

MOTIVATION: The knowledge of the subcellular localization of a protein is fundamental for elucidating its function. It is difficult to determine the subcellular location for eukaryotic cells with experimental high-throughput procedures. Computational procedures are then needed for annotating the subcellular location of proteins in large scale genomic projects. RESULTS: BaCelLo is a predictor for five classes of subcellular localization (secretory pathway, cytoplasm, nucleus, mitochondrion and chloroplast) and it is based on different SVMs organized in a decision tree. The system exploits the information derived from the residue sequence and from the evolutionary information contained in alignment profiles. It analyzes the whole sequence composition and the compositions of both the N- and C-termini. The training set is curated in order to avoid redundancy. For the first time a balancing procedure is introduced in order to mitigate the effect of biased training sets. Three kingdom-specific predictors are implemented: for animals, plants and fungi, respectively. When distributing the proteins from animals and fungi into four classes, accuracy of BaCelLo reach 74% and 76%, respectively; a score of 67% is obtained when proteins from plants are distributed into five classes. BaCelLo outperforms the other presently available methods for the same task and gives more balanced accuracy and coverage values for each class. We also predict the subcellular localization of five whole proteomes, Homo sapiens, Mus musculus, Caenorhabditis elegans, Saccharomyces cerevisiae and Arabidopsis thaliana, comparing the protein content in each different compartment. AVAILABILITY: BaCelLo can be accessed at http://www.biocomp.unibo.it/bacello/.

Algorithms↗

PONGO: a web server for multiple predictions of all-alpha transmembrane proteins.

The annotation efforts of the BIOSAPIENS European Network of Excellence have generated several distributed annotation systems (DAS) with the aim of integrating Bioinformatics resources and annotating metazoan genomes (http://www.biosapiens.info). In this context, the PONGO DAS server (http://pongo.biocomp.unibo.it) provides the annotation on predictive basis for the all-alpha membrane proteins in the human genome, not only through DAS queries, but also directly using a simple web interface. In order to produce a more comprehensive analysis of the sequence at hand, this annotation is carried out with four selected and high scoring predictors: TMHMM2.0, MEMSAT, PRODIV and ENSEMBLE1.0. The stored and pre-computed predictions for the human proteins can be searched and displayed in a graphical view. However the web service allows the prediction of the topology of any kind of putative membrane proteins, regardless of the organism and more importantly with the same sequence profile for a given sequence when required. Here we present a new web server that incorporates the state-of-the-art topology predictors in a single framework, so that putative users can interactively compare and evaluate four predictions simultaneously for a given sequence. Together with the predicted topology, the server also displays a signal peptide prediction determined with SPEP. The PONGO web server is available at http://pongo.biocomp.unibo.it/pongo.

Humans↗

Pressure and temperature as tools for investigating the role of individual non-covalent interactions in enzymatic reactions Sulfolobus solfataricus carboxypeptidase as a model enzyme.

Sulfolobus solfataricus carboxypeptidase, (CPSso), is a heat- and pressure-resistant zinc-metalloprotease. Thanks to its properties, it is an ideal tool for investigating the role of non-covalent interactions in substrate binding. It has a broad substrate specificity as it can cleave any N-blocked amino acid (except for N-blocked proline). Its catalytic and kinetic mechanisms are well understood, and the hydrolytic reaction is easily detectable spectrophotometrically. Here, we report investigations on the pressure- and temperature-dependence of the kinetic parameters (turnover number and Michaelis constant) of CPSso using several benzoyl- and 3-(2-furyl)acryloyl-amino acids as substrates. This approach enabled us to study these parameters in terms of individual rate constants and establish that the release of the free amino acid is the rate-limiting step, making it possible to dissect the individual non-covalent interactions participating in substrate binding. In keeping with molecular docking experiments performed on the 3D model of CPSso available to date, our results show that both hydrophobic and energetic interactions (i.e., stacking and van der Waals) are mainly involved, but their contribution varies strongly, probably due to changes in the conformational state of the enzyme.

Archaeal Proteins↗

A new decoding algorithm for hidden Markov models improves the prediction of the topology of all-beta membrane proteins.

BACKGROUND: Structure prediction of membrane proteins is still a challenging computational problem. Hidden Markov models (HMM) have been successfully applied to the problem of predicting membrane protein topology. In a predictive task, the HMM is endowed with a decoding algorithm in order to assign the most probable state path, and in turn the labels, to an unknown sequence. The Viterbi and the posterior decoding algorithms are the most common. The former is very efficient when one path dominates, while the latter, even though does not guarantee to preserve the HMM grammar, is more effective when several concurring paths have similar probabilities. A third good alternative is 1-best, which was shown to perform equal or better than Viterbi. RESULTS: In this paper we introduce the posterior-Viterbi (PV) a new decoding which combines the posterior and Viterbi algorithms. PV is a two step process: first the posterior probability of each state is computed and then the best posterior allowed path through the model is evaluated by a Viterbi algorithm. CONCLUSION: We show that PV decoding performs better than other algorithms when tested on the problem of the prediction of the topology of beta-barrel membrane proteins.

Algorithms↗

TRAMPLE: the transmembrane protein labelling environment.

TRAMPLE (http://gpcr.biocomp.unibo.it/biodec/) is a web application server dedicated to the detection and the annotation of transmembrane protein sequences. TRAMPLE includes different state-of-the-art algorithms for the prediction of signal peptides, transmembrane segments (both beta-strands and alpha-helices), secondary structure and fast fold recognition. TRAMPLE also includes a complete content management system to manage the results of the predictions. Each user of the server has his/her own workplace, where the data can be stored, organized, accessed and annotated with documents through a simple web-based interface. In this manner, TRAMPLE significantly improves usability with respect to other more traditional web servers.

Algorithms↗

The ectodomain of herpes simplex virus glycoprotein H contains a membrane alpha-helix with attributes of an internal fusion peptide, positionally conserved in the herpesviridae family.

Human herpesviruses enter cells by fusion with target membranes, a process that requires three conserved glycoproteins: gB, gH, and gL. How these glycoproteins execute fusion is unknown. Neural network bioinformatics predicted a membrane alpha-helix contained within the ectodomain of herpes simplex virus (HSV) gH, positionally conserved in the gH of all examined herpesviruses. Evidence that it has attributes of an internal fusion peptide rests on the following lines of evidence. (i) The predicted membrane alpha-helix has the attribute of a membrane segment, since it transformed a soluble form of gD into a membrane-bound gD. (ii) It represents a critical domain of gH. Its partial or entire deletion, or substitution of critical residues inhibited HSV infectivity and fusion in the cell-cell fusion assay. (iii) Its replacement with the fusion peptide from human immunodeficiency virus gp41 or from vesicular stomatitis virus G partially rescued HSV infectivity and cell-cell fusion. The corresponding antisense sequences did not. (iv) The predicted alpha-helix located in the varicella-zoster virus gH ectodomain can functionally substitute the native HSV gH membrane alpha-helix, suggesting a conserved function in the human herpesviruses. We conclude that HSV gH exhibits features typical of viral fusion glycoproteins and that this property is likely conserved in the Herpesviridae family.

Amino Acid Sequence↗

Prediction of disulfide-bonded cysteines in proteomes with a hidden neural network.

A hidden neural network-based method is used to predict the bonding state of cysteines starting from the residue sequence of the protein chain. The method scores as high as 89% and 86% per cysteine residue and per protein, respectively, and in this overcomes other predictors of the same category. We then explore the efficacy of our predictor in computing the disulfide content of the whole proteome of Escherichia coli (K12 and O157), Aeropirum pernix, Thermotoga maritima, and Homo sapiens. We find that the percentage of extracellular disulfide containing proteins is higher than that of intracellular one, and that the human proteome is by far the one with the highest content of sulfur-sulfur linkages in proteins.

Cysteine↗

MaxSubSeq: an algorithm for segment-length optimization. The case study of the transmembrane spanning segments.

MOTIVATION: A problem in predicting the topography of transmembrane proteins is the optimal localization of the transmembrane segments along the protein sequences, provided that each residue is associated with a propensity of being or not being included in the transmembrane protein region. From previous work it is known that post-processing of propensity signals with suited algorithms can greatly improve the quality and the accuracy of the predictions. In this paper we describe a general dynamic programming-like algorithm (MaxSubSeq, Maximal SubSequence) specifically designed to optimize the number and length of segments with constrained length in a given protein sequence. Previous application of our algorithm, has proved its effectiveness in the optimization task of both neural network and hidden Markov models output, and in this paper we present the detailed description of MaxSubSeq. RESULTS: We describe the application of MaxSubSeq to the location of both helical and beta strand transmembrane segments, optimizing the outputs derived with different predictive algorithms. For all-alpha transmembrane proteins we use both the standard Kyte-Doolittle (KD) hydropathy scale and the TMHMM predictor (http://www.cbs.dtu.dk/). Using a set of 188 well characterized membrane proteins, MaxSubSeq nearly doubles the correct location of transmembrane segments as compared to the standard KD hydrophobicity plot, reaching 51% accuracy. If MaxSubSeq is used to optimize the TMHMM method the accuracy increases from 68 to 72%. When used to regularize the prediction of beta transmembrane strands, obtained using both a neural network and a HMM based predictors, MaxSubSeq increases the accuracy per protein up to 72 and 73% respectively. AVAILABILITY: The program is available upon request to the authors, or it is accessible through our web server (http://gpcr.biocomp.unibo.it/predictors/)

Algorithms↗

3D structure of Sulfolobus solfataricus carboxypeptidase developed by molecular modeling is confirmed by site-directed mutagenesis and small angle X-ray scattering.

Sulfolobus solfataricus carboxypeptidase (CPSso) is a thermostable zinc-metalloenzyme with a M(r) of 43,000. Taking into account the experimentally determined zinc content of one ion per subunit, we developed two alternative 3D models, starting from the available structures of Thermoactinomyces vulgaris carboxypeptidase (Model A) and Pseudomonas carboxypeptidase G2 (Model B). The former enzyme is monomeric and has one metal ion in the active site, while the latter is dimeric and has two bound zinc ions. The two models were computed by exploiting the structural alignment of the one zinc- with the two zinc-containing active sites of the two templates, and with a threading procedure. Both computed structures resembled the respective template, with only one bound zinc with tetrahedric coordination in the active site. With these models, two different quaternary structures can be modeled: one using Model A with a hexameric symmetry, the other from Model B with a tetrameric symmetry. Mutagenesis experiments directed toward the residues putatively involved in metal chelation in either of the models disproved Model A and supported Model B, in which the metal-binding site comprises His(108), Asp(109), and His(168). We also identified Glu(142) as the acidic residue interacting with the water molecule occupying the fourth chelation site. Furthermore, the overall fold and the oligomeric structure of the molecule was validated by small angle x-ray scattering (SAXS). An ab initio original approach was used to reconstruct the shape of the CPSso in solution from the experimental curves. The results clearly support a tetrameric structure. The Monte Carlo method was then used to compare the crystallographic coordinates of the possible quaternary structures for CPSso with the SAXS profiles. The fitting procedure showed that only the model built using the Pseudomonas carboxypeptidase G2 structure as a template fitted the experimental data.

Amino Acid Sequence↗

In silico prediction of the structure of membrane proteins: is it feasible?

In the 'omic' era, hundreds of genomes are available for protein sequence analysis, and some 30 per cent of all sequences are of membrane proteins. Unlike globular proteins, a 3D model for membrane proteins can hardly be computed starting from the sequence. Why is this so? What can we really compute and with what reliability? These and other matters are outlined.

Computational Biology↗

An ENSEMBLE machine learning approach for the prediction of all-alpha membrane proteins.

MOTIVATION: All-alpha membrane proteins constitute a functionally relevant subset of the whole proteome. Their content ranges from about 10 to 30% of the cell proteins, based on sequence comparison and specific predictive methods. Due to the paucity of membrane proteins solved with atomic resolution, the training/testing sets of predictive methods for protein topography and topology routinely include very few well-solved structures mixed with a hundred proteins known with low resolution. Moreover, available predictors fail in predicting recently crystallised membrane proteins (Chen et al., 2002). Presently the number of well-solved membrane proteins comprises some 59 chains of low sequence homology. It is therefore possible to train/test predictors only with the set of proteins known with atomic resolution and evaluate more thoroughly the performance of different methods. RESULTS: We implement a cascade-neural network (NN), two different hidden Markov models (HMM), and their ensemble (ENSEMBLE) as a new method. We train and test in cross validation the three methods and ENSEMBLE on the 59 well resolved membrane proteins. ENSEMBLE scores with a per-protein accuracy of 90% for topography and 71% for topology, outperforming the best single method of 7 and 5 percentage points, respectively. When tested on a low resolution set of 151 proteins, with no homology with the 59 proteins, the per-protein accuracy of ENSEMBLE is 76% for topography and 68% for topology. Our results also indicate that the performance of ENSEMBLE is higher than that of the best predictors presently available on the Web.

Algorithms↗

Fishing new proteins in the twilight zone of genomes: the test case of outer membrane proteins in Escherichia coli K12, Escherichia coli O157:H7, and other Gram-negative bacteria.

We address the problem of clustering the whole protein content of genomes into three different categories-globular, all-alpha, and all-beta membrane proteins-with the aim of fishing new membrane proteins in the pool of nonannotated proteins (twilight zone). The focus is then mainly on outer membrane proteins. This is performed by using an integrated suite of programs (Hunter) specifically developed for predicting the occurrence of signal peptides in proteins of Gram-negative bacteria and the topography of all-alpha and all-beta membrane proteins. Hunter is tested on the well and partially annotated proteins (2160 and 760, respectively) of Escherichia coli K 12 scoring as high as 95.6% in the correct assignment of each chain to the category. Of the remaining 1253 nonannotated sequences, 1099 are predicted globular, 136 are all-alpha, and 18 are all-beta membrane proteins. In Escherichia coli 0157:H7 we filtered 1901 nonannotated proteins. Our analysis classifies 1564 globular chains, 327 inner membrane proteins, and 10 outer membrane proteins. With Hunter, new membrane proteins are added to the list of putative membrane proteins of Gram-negative bacteria. The content of outer membrane proteins per genome (nine are analyzed) ranges from 1.5% to 2.4%, and it is one order of magnitude lower than that of inner membrane proteins. The finding is particularly relevant when it is considered that this is the first large-scale analysis based on validated tools that can predict the content of outer membrane proteins in a genome and can allow cross-comparison of the same protein type between different species.

Bacterial Outer Membrane Proteins↗

Effect of molecular confinement on internal enzyme dynamics: frequency domain fluorometry and molecular dynamics simulation studies.

The tryptophanyl emission decay of the mesophilic beta-galactosidase from Aspergillus oryzae free in buffer and entrapped in agarose gel is investigated as a function of temperature and compared to that of the hyperthermophilic enzyme from Sulfolobus solfataricus. Both enzymes are tetrameric proteins with a large number of tryptophanyl residues, so the fluorescence emission can provide information on the conformational dynamics of the overall protein structure rather than that of the local environment. The tryptophanyl emission decays are best fitted by bimodal Lorentzian distributions. The long-lived component is ascribed to close, deeply buried tryptophanyl residues with reduced mobility; the short-lived one arises from tryptophanyl residues located in more flexible external regions of each subunit, some of which are involved in forming the catalytic site. The center of both lifetime distribution components at each temperature increases when going from the free in solution mesophilic enzyme to the gel-entrapped and hyperthermophilic enzyme, thus indicating that confinement of the mesophilic enzyme in the agarose gel limits the freedom of the polypeptide chain. A more complex dependence is observed for the distribution widths. Computer modeling techniques are used to recognize that the catalytic sites are similar for the mesophilic and hyperthermophilic beta-galactosidases. The effect due to gel entrapment is considered in dynamic simulations by imposing harmonic restraints to solvent-exposed atoms of the protein with the exclusion of those around the active site. The temperature dependence of the tryptophanyl fluorescence emission decay and the dynamic simulation confirm that more rigid structures, as in the case of the immobilized and/or hyperthermophilic enzyme, require higher temperatures to achieve the requisite conformational dynamics for an effective catalytic action and strongly suggest a link between conformational rigidity and enhanced thermal stability.

Aspergillus oryzae↗

A sequence-profile-based HMM for predicting and discriminating beta barrel membrane proteins.

MOTIVATION: Membrane proteins are an abundant and functionally relevant subset of proteins that putatively include from about 15 up to 30% of the proteome of organisms fully sequenced. These estimates are mainly computed on the basis of sequence comparison and membrane protein prediction. It is therefore urgent to develop methods capable of selecting membrane proteins especially in the case of outer membrane proteins, barely taken into consideration when proteome wide analysis is performed. This will also help protein annotation when no homologous sequence is found in the database. Outer membrane proteins solved so far at atomic resolution interact with the external membrane of bacteria with a characteristic beta barrel structure comprising different even numbers of beta strands (beta barrel membrane proteins). In this they differ from the membrane proteins of the cytoplasmic membrane endowed with alpha helix bundles (all alpha membrane proteins) and need specialised predictors. RESULTS: We develop a HMM model, which can predict the topology of beta barrel membrane proteins using, as input, evolutionary information. The model is cyclic with 6 types of states: two for the beta strand transmembrane core, one for the beta strand cap on either side of the membrane, one for the inner loop, one for the outer loop and one for the globular domain state in the middle of each loop. The development of a specific input for HMM based on multiple sequence alignment is novel. The accuracy per residue of the model is 83% when a jack knife procedure is adopted. With a model optimisation method using a dynamic programming algorithm seven topological models out of the twelve proteins included in the testing set are also correctly predicted. When used as a discriminator, the model is rather selective. At a fixed probability value, it retains 84% of a non-redundant set comprising 145 sequences of well-annotated outer membrane proteins. Concomitantly, it correctly rejects 90% of a set of globular proteins including about 1200 chains with low sequence identity (<30%) and 90% of a set of all alpha membrane proteins, including 188 chains.

Algorithms↗

Prediction of the disulfide bonding state of cysteines in proteins with hidden neural networks.

A hybrid system (hidden neural network) based on a hidden Markov model (HMM) and neural networks (NN) was trained to predict the bonding states of cysteines in proteins starting from the residue chains. Training was performed using 4136 cysteine-containing segments extracted from 969 non-homologous proteins of well-resolved 3D structure and without chain-breaks. After a 20-fold cross-validation procedure, the efficiency of the prediction scores as high as 80% using neural networks based on evolutionary information. When the whole protein is taken into account by means of an HMM, a hybrid system is generated, whose emission probabilities are computed using the NN output (hidden neural networks). In this case, the predictor accuracy increases up to 88%. Further, when tested on a protein basis, the hybrid system can correctly predict 84% of the chains in the data set, with a gain of at least 27% over the NN predictor.

Cysteine↗