Search PubMed⌕ Search

Biomedical subjects

Darren R Flower

Publications and source records attributed to Darren R Flower.

At least 19 recordsLinked to original sources

VaxiJen: a server for prediction of protective antigens, tumour antigens and subunit vaccines.

BACKGROUND: Vaccine development in the post-genomic era often begins with the in silico screening of genome information, with the most probable protective antigens being predicted rather than requiring causative microorganisms to be grown. Despite the obvious advantages of this approach--such as speed and cost efficiency--its success remains dependent on the accuracy of antigen prediction. Most approaches use sequence alignment to identify antigens. This is problematic for several reasons. Some proteins lack obvious sequence similarity, although they may share similar structures and biological properties. The antigenicity of a sequence may be encoded in a subtle and recondite manner not amendable to direct identification by sequence alignment. The discovery of truly novel antigens will be frustrated by their lack of similarity to antigens of known provenance. To overcome the limitations of alignment-dependent methods, we propose a new alignment-free approach for antigen prediction, which is based on auto cross covariance (ACC) transformation of protein sequences into uniform vectors of principal amino acid properties. RESULTS: Bacterial, viral and tumour protein datasets were used to derive models for prediction of whole protein antigenicity. Every set consisted of 100 known antigens and 100 non-antigens. The derived models were tested by internal leave-one-out cross-validation and external validation using test sets. An additional five training sets for each class of antigens were used to test the stability of the discrimination between antigens and non-antigens. The models performed well in both validations showing prediction accuracy of 70% to 89%. The models were implemented in a server, which we call VaxiJen. CONCLUSION: VaxiJen is the first server for alignment-independent prediction of protective antigens. It was developed to allow antigen classification solely based on the physicochemical properties of proteins without recourse to sequence alignment. The server can be used on its own or in combination with alignment-based prediction methods. It is freely-available online at the URL: http://www.jenner.ac.uk/VaxiJen.

Algorithms↗

Predicting Class II MHC-Peptide binding: a kernel based approach using similarity scores.

BACKGROUND: Modelling the interaction between potentially antigenic peptides and Major Histocompatibility Complex (MHC) molecules is a key step in identifying potential T-cell epitopes. For Class II MHC alleles, the binding groove is open at both ends, causing ambiguity in the positional alignment between the groove and peptide, as well as creating uncertainty as to what parts of the peptide interact with the MHC. Moreover, the antigenic peptides have variable lengths, making naive modelling methods difficult to apply. This paper introduces a kernel method that can handle variable length peptides effectively by quantifying similarities between peptide sequences and integrating these into the kernel. RESULTS: The kernel approach presented here shows increased prediction accuracy with a significantly higher number of true positives and negatives on multiple MHC class II alleles, when testing data sets from MHCPEP 1, MCHBN 2, and MHCBench 3. Evaluation by cross validation, when segregating binders and non-binders, produced an average of 0.824 AROC for the MHCBench data sets (up from 0.756), and an average of 0.96 AROC for multiple alleles of the MHCPEP database. CONCLUSION: The method improves performance over existing state-of-the-art methods of MHC class II peptide binding predictions by using a custom, knowledge-based representation of peptides. Similarity scores, in contrast to a fixed-length, pocket-specific representation of amino acids, provide a flexible and powerful way of modelling MHC binding, and can easily be applied to other dynamic sequence problems.

Binding Sites↗

SVRMHC prediction server for MHC-binding peptides.

BACKGROUND: The binding between antigenic peptides (epitopes) and the MHC molecule is a key step in the cellular immune response. Accurate in silico prediction of epitope-MHC binding affinity can greatly expedite epitope screening by reducing costs and experimental effort. RESULTS: Recently, we demonstrated the appealing performance of SVRMHC, an SVR-based quantitative modeling method for peptide-MHC interactions, when applied to three mouse class I MHC molecules. Subsequently, we have greatly extended the construction of SVRMHC models and have established such models for more than 40 class I and class II MHC molecules. Here we present the SVRMHC web server for predicting peptide-MHC binding affinities using these models. Benchmarked percentile scores are provided for all predictions. The larger number of SVRMHC models available allowed for an updated evaluation of the performance of the SVRMHC method compared to other well- known linear modeling methods. CONCLUSION: SVRMHC is an accurate and easy-to-use prediction server for epitope-MHC binding with significant coverage of MHC molecules. We believe it will prove to be a valuable resource for T cell epitope researchers.

Amino Acid Sequence↗

Identifying candidate subunit vaccines using an alignment-independent method based on principal amino acid properties.

Subunit vaccine discovery is an accepted clinical priority. The empirical approach is time- and labor-consuming and can often end in failure. Rational information-driven approaches can overcome these limitations in a fast and efficient manner. However, informatics solutions require reliable algorithms for antigen identification. All known algorithms use sequence similarity to identify antigens. However, antigenicity may be encoded subtly in a sequence and may not be directly identifiable by sequence alignment. We propose a new alignment-independent method for antigen recognition based on the principal chemical properties of protein amino acid sequences. The method is tested by cross-validation on a training set of bacterial antigens and external validation on a test set of known antigens. The prediction accuracy is 83% for the cross-validation and 80% for the external test set. Our approach is accurate and robust, and provides a potent tool for the in silico discovery of medically relevant subunit vaccines.

Algorithms↗

Benchmarking pK(a) prediction.

BACKGROUND: pKa values are a measure of the protonation of ionizable groups in proteins. Ionizable groups are involved in intra-protein, protein-solvent and protein-ligand interactions as well as solubility, protein folding and catalytic activity. The pKa shift of a group from its intrinsic value is determined by the perturbation of the residue by the environment and can be calculated from three-dimensional structural data. RESULTS: Here we use a large dataset of experimentally-determined pKas to analyse the performance of different prediction techniques. Our work provides a benchmark of available software implementations: MCCE, MEAD, PROPKA and UHBD. Combinatorial and regression analysis is also used in an attempt to find a consensus approach towards pKa prediction. The tendency of individual programs to over- or underpredict the pKa value is related to the underlying methodology of the individual programs. CONCLUSION: Overall, PROPKA is more accurate than the other three programs. Key to developing accurate predictive software will be a complete sampling of conformations accessible to protein structures.

Catalysis↗

Modeling the peptide-T cell receptor interaction by the comparative molecular similarity indices analysis-soft independent modeling of class analogy technique.

A set of 38 epitopes and 183 non-epitopes, which bind to alleles of the HLA-A3 supertype, was subjected to a combination of comparative molecular similarity indices analysis (CoMSIA) and soft independent modeling of class analogy (SIMCA). During the process of T cell recognition, T cell receptors (TCR) interact with the central section of the bound nonamer peptide; thus only positions 4-8 were considered in the study. The derived model distinguished 82% of the epitopes and 73% of the non-epitopes after cross-validation in five groups. The overall preference from the model is for polar amino acids with high electron density and the ability to form hydrogen bonds. These so-called "aggressive" amino acids are flanked by small-sized residues, which enable such residues to protrude from the binding cleft and take an active role in TCR-mediated T cell recognition. Combinations of "aggressive" and "passive" amino acids in the middle part of epitopes constitute a putative TCR binding motif.

Amino Acid Motifs↗

Quantitative prediction of mouse class I MHC peptide binding affinity using support vector machine regression (SVR) models.

BACKGROUND: The binding between peptide epitopes and major histocompatibility complex proteins (MHCs) is an important event in the cellular immune response. Accurate prediction of the binding between short peptides and the MHC molecules has long been a principal challenge for immunoinformatics. Recently, the modeling of MHC-peptide binding has come to emphasize quantitative predictions: instead of categorizing peptides as "binders" or "non-binders" or as "strong binders" and "weak binders", recent methods seek to make predictions about precise binding affinities. RESULTS: We developed a quantitative support vector machine regression (SVR) approach, called SVRMHC, to model peptide-MHC binding affinities. As a non-linear method, SVRMHC was able to generate models that out-performed existing linear models, such as the "additive method". By adopting a new "11-factor encoding" scheme, SVRMHC takes into account similarities in the physicochemical properties of the amino acids constituting the input peptides. When applied to MHC-peptide binding data for three mouse class I MHC alleles, the SVRMHC models produced more accurate predictions than those produced previously. Furthermore, comparisons based on Receiver Operating Characteristic (ROC) analysis indicated that SVRMHC was able to out-perform several prominent methods in identifying strongly binding peptides. CONCLUSION: As a method with demonstrated performance in the quantitative modeling of MHC-peptide binding and in identifying strong binders, SVRMHC is a promising immunoinformatics tool with not inconsiderable future potential.

Algorithms↗

Statistical deconvolution of enthalpic energetic contributions to MHC-peptide binding affinity.

BACKGROUND: MHC Class I molecules present antigenic peptides to cytotoxic T cells, which forms an integral part of the adaptive immune response. Peptides are bound within a groove formed by the MHC heavy chain. Previous approaches to MHC Class I-peptide binding prediction have largely concentrated on the peptide anchor residues located at the P2 and C-terminus positions. RESULTS: A large dataset comprising MHC-peptide structural complexes was created by re-modelling pre-determined x-ray crystallographic structures. Static energetic analysis, following energy minimisation, was performed on the dataset in order to characterise interactions between bound peptides and the MHC Class I molecule, partitioning the interactions within the groove into van der Waals, electrostatic and total non-bonded energy contributions. CONCLUSION: The QSAR techniques of Genetic Function Approximation (GFA) and Genetic Partial Least Squares (G/PLS) algorithms were used to identify key interactions between the two molecules by comparing the calculated energy values with experimentally-determined BL50 data. Although the peptide termini binding interactions help ensure the stability of the MHC Class I-peptide complex, the central region of the peptide is also important in defining the specificity of the interaction. As thermodynamic studies indicate that peptide association and dissociation may be driven entropically, it may be necessary to incorporate entropic contributions into future calculations.

Binding Sites↗

EpiJen: a server for multistep T cell epitope prediction.

BACKGROUND: The main processing pathway for MHC class I ligands involves degradation of proteins by the proteasome, followed by transport of products by the transporter associated with antigen processing (TAP) to the endoplasmic reticulum (ER), where peptides are bound by MHC class I molecules, and then presented on the cell surface by MHCs. The whole process is modeled here using an integrated approach, which we call EpiJen. EpiJen is based on quantitative matrices, derived by the additive method, and applied successively to select epitopes. EpiJen is available free online. RESULTS: To identify epitopes, a source protein is passed through four steps: proteasome cleavage, TAP transport, MHC binding and epitope selection. At each stage, different proportions of non-epitopes are eliminated. The final set of peptides represents no more than 5% of the whole protein sequence and will contain 85% of the true epitopes, as indicated by external validation. Compared to other integrated methods (NetCTL, WAPP and SMM), EpiJen performs best, predicting 61 of the 99 HIV epitopes used in this study. CONCLUSION: EpiJen is a reliable multi-step algorithm for T cell epitope prediction, which belongs to the next generation of in silico T cell epitope identification methods. These methods aim to reduce subsequent experimental work by improving the success rate of epitope prediction.

Computer Simulation↗

Class I T-cell epitope prediction: improvements using a combination of proteasome cleavage, TAP affinity, and MHC binding.

Cleavage by the proteasome is responsible for generating the C terminus of T-cell epitopes. Modeling the process of proteasome cleavage as part of a multi-step algorithm for T-cell epitope prediction will reduce the number of non-binders and increase the overall accuracy of the predictive algorithm. Quantitative matrix-based models for prediction of the proteasome cleavage sites in a protein were developed using a training set of 489 naturally processed T-cell epitopes (nonamer peptides) associated with HLA-A and HLA-B molecules. The models were validated using an external test set of 227 T-cell epitopes. The performance of the models was good, identifying 76% of the C-termini correctly. The best model of proteasome cleavage was incorporated as the first step in a three-step algorithm for T-cell epitope prediction, where subsequent steps predicted TAP affinity and MHC binding using previously derived models.

ATP-Binding Cassette Transporters↗

PPD v1.0--an integrated, web-accessible database of experimentally determined protein pKa values.

The Protein pK(a) Database (PPD) v1.0 provides a compendium of protein residue-specific ionization equilibria (pK(a) values), as collated from the primary literature, in the form of a web-accessible postgreSQL relational database. Ionizable residues play key roles in the molecular mechanisms that underlie many biological phenomena, including protein folding and enzyme catalysis. The PPD serves as a general protein pK(a) archive and as a source of data that allows for the development and improvement of pK(a) prediction systems. The database is accessed through an HTML interface, which offers two fast, efficient search methods: an amino acid-based query and a Basic Local Alignment Search Tool search. Entries also give details of experimental techniques and links to other key databases, such as National Center for Biotechnology Information and the Protein Data Bank, providing the user with considerable background information. The database can be found at the following URL: http://www.jenner.ac.uk/PPD.

Amino Acids↗

Receptor-binding sites: bioinformatic approaches.

It is increasingly clear that both transient and long-lasting interactions between biomacromolecules and their molecular partners are the most fundamental of all biological mechanisms and lie at the conceptual heart of protein function. In particular, the protein-binding site is the most fascinating and important mechanistic arbiter of protein function. In this review, I examine the nature of protein-binding sites found in both ligand-binding receptors and substrate-binding enzymes. I highlight two important concepts underlying the identification and analysis of binding sites. The first is based on knowledge: when one knows the location of a binding site in one protein, one can "inherit" the site from one protein to another. The second approach involves the a priori prediction of a binding site from a sequence or a structure. The full and complete analysis of binding sites will necessarily involve the full range of informatic techniques ranging from sequence-based bioinformatic analysis through structural bioinformatics to computational chemistry and molecular physics. Integration of both diverse experimental and diverse theoretical approaches is thus a mandatory requirement in the evaluation of binding sites and the binding events that occur within them.

Binding Sites↗

MHCPred 2.0: an updated quantitative T-cell epitope prediction server.

UNLABELLED: The accurate computational prediction of T-cell epitopes can greatly reduce the experimental overhead implicit in candidate epitope identification within genomic sequences. In this article we present MHCPred 2.0, an enhanced version of our online, quantitative T-cell epitope prediction server. The previous version of MHCPred included mostly alleles from the human leukocyte antigen A (HLA-A) locus. In MHCPred 2.0, mouse models are added and computational constraints removed. Currently the server includes 11 human HLA class I, three human HLA class II, and three mouse class I models. Additionally, a binding model for the human transporter associated with antigen processing (TAP) is incorporated into the new MHCPred. A tool for the design of heteroclitic peptides is also included within the server. To refine the veracity of binding affinities prediction, a confidence percentage is also now calculated for each peptide predicted. AVAILABILITY: As previously, MHCPred 2.0 is freely available at the URL http://www.jenner.ac.uk/MHCPred/ CONTACT: Darren R. Flower (darren.flower@jenner.ac.uk).

Algorithms↗

Receptor-ligand binding sites and virtual screening.

Within the pharmaceutical industry, the ultimate source of continuing profitability is the unremitting process of drug discovery. To be profitable, drugs must be marketable: legally novel, safe and relatively free of side effects, efficacious, and ideally inexpensive to produce. While drug discovery was once typified by a haphazard and empirical process, it is now increasingly driven by both knowledge of the receptor-mediated basis of disease and how drug molecules interact with receptors and the wider physiome. Medicinal chemistry postulates that to understand a congeneric ligand series, or set thereof, is to understand the nature and requirements of a ligand binding site. Likewise, structural molecular biology posits that to understand a binding site is to understand the nature of ligands bound therein. Reality sits somewhere between these extremes, yet subsumes them both. Complementary to rules of ligand design, arising through decades of medicinal chemistry, structural biology and computational chemistry are able to elucidate the nature of binding site-ligand interactions, facilitating, at both pragmatic and conceptual levels, the drug discovery process.

Binding Sites↗

DSD--an integrated, web-accessible database of Dehydrogenase Enzyme Stereospecificities.

BACKGROUND: Dehydrogenase enzymes belong to the oxidoreductase class and utilise the coenzymes NAD and NADP. Stereo-selectivity is focused on the C4 hydrogen atoms of the nicotinamide ring of NAD(P). Depending upon which hydrogen is transferred at the C4 location, the enzyme is designated as A or B stereospecific. DESCRIPTION: The Dehydrogenase Stereospecificity Database v1.0 (DSD) provides a compilation of enzyme stereochemical data, as sourced from the primary literature, in the form of a web-accessible database. There are two search engines, a menu driven search and a BLAST search. The entries are also linked to several external databases, including the NCBI and the Protein Data Bank, providing wide background information. The database is freely available online at: http://www.jenner.ac.uk/DSD/. CONCLUSION: DSD is a unique compilation available on-line for the first time which provides a key resource for the comparative analysis of reductase hydrogen transfer stereospecificity. As databases increasingly form the backbone of science, largely complete databases such as DSD, are a vital addition.

Computational Biology↗

Analysis of peptide-protein binding using amino acid descriptors: prediction and experimental verification for human histocompatibility complex HLA-A0201.

Amino acid descriptors are often used in quantitative structure-activity relationship (QSAR) analysis of proteins and peptides. In the present study, descriptors were used to characterize peptides binding to the human MHC allele HLA-A0201. Two sets of amino acid descriptors were chosen: 93 descriptors taken from the amino acid descriptor database AAindex and the z descriptors defined by Wold and Sandberg. Variable selection techniques (SIMCA, genetic algorithm, and GOLPE) were applied to remove redundant descriptors. Our results indicate that QSAR models generated using five z descriptors had the highest predictivity and explained variance (q2 between 0.6 and 0.7 and r2 between 0.6 and 0.9). Further to the QSAR analysis, 15 peptides were synthesized and tested using a T2 stabilization assay. All peptides bound to HLA-A0201 well, and four peptides were identified as high-affinity binders.

Algorithms↗

AntiJen: a quantitative immunology database integrating functional, thermodynamic, kinetic, biophysical, and cellular data.

AntiJen is a database system focused on the integration of kinetic, thermodynamic, functional, and cellular data within the context of immunology and vaccinology. Compared to its progenitor JenPep, the interface has been completely rewritten and redesigned and now offers a wider variety of search methods, including a nucleotide and a peptide BLAST search. In terms of data archived, AntiJen has a richer and more complete breadth, depth, and scope, and this has seen the database increase to over 31,000 entries. AntiJen provides the most complete and up-to-date dataset of its kind. While AntiJen v2.0 retains a focus on both T cell and B cell epitopes, its greatest novelty is the archiving of continuous quantitative data on a variety of immunological molecular interactions. This includes thermodynamic and kinetic measures of peptide binding to TAP and the Major Histocompatibility Complex (MHC), peptide-MHC complexes binding to T cell receptors, antibodies binding to protein antigens and general immunological protein-protein interactions. The database also contains quantitative specificity data from position-specific peptide libraries and biophysical data, in the form of diffusion co-efficients and cell surface copy numbers, on MHCs and other immunological molecules. The uses of AntiJen include the design of vaccines and diagnostics, such as tetramers, and other laboratory reagents, as well as helping parameterize the bioinformatic or mathematical in silico modeling of the immune system. The database is accessible from the URL: http://www.jenner.ac.uk/antijen.

Journal Article↗

Peptide recognition by the T cell receptor: comparison of binding free energies from thermodynamic integration, Poisson-Boltzmann and linear interaction energy approximations.

The binding to the T cell receptor of wild-type and variant HTLV-1 Tax peptide complexed to the major histocompatibility complex has been investigated by means of molecular dynamics simulations. The binding free energy difference is calculated using the molecular mechanics Poisson-Boltzmann surface area and linear interaction energy methods. These methods extract useful information on the binding energetics from simulations of the physical states of the ligands, which are more computationally expedient than the commonly used thermodynamic integration method. The successful reproduction of the relative binding free energies shows that these methods can be useful for free energy calculations and the rational design of drugs and vaccines.

Animals↗