Search PubMed⌕ Search

Biomedical subjects

James R Bradford

Publications and source records attributed to James R Bradford.

7 recordsLinked to original sources

Predicting the effect of missense mutations on protein function: analysis with Bayesian networks.

BACKGROUND: A number of methods that use both protein structural and evolutionary information are available to predict the functional consequences of missense mutations. However, many of these methods break down if either one of the two types of data are missing. Furthermore, there is a lack of rigorous assessment of how important the different factors are to prediction. RESULTS: Here we use Bayesian networks to predict whether or not a missense mutation will affect the function of the protein. Bayesian networks provide a concise representation for inferring models from data, and are known to generalise well to new data. More importantly, they can handle the noisy, incomplete and uncertain nature of biological data. Our Bayesian network achieved comparable performance with previous machine learning methods. The predictive performance of learned model structures was no better than a naïve Bayes classifier. However, analysis of the posterior distribution of model structures allows biologically meaningful interpretation of relationships between the input variables. CONCLUSION: The ability of the Bayesian network to make predictions when only structural or evolutionary data was observed allowed us to conclude that structural information is a significantly better predictor of the functional consequences of a missense mutation than evolutionary information, for the dataset used. Analysis of the posterior distribution of model structures revealed that the top three strongest connections with the class node all involved structural nodes. With this in mind, we derived a simplified Bayesian network that used just these three structural descriptors, with comparable performance to that of an all node network.

Algorithms↗

Insights into protein-protein interfaces using a Bayesian network prediction method.

Identifying the interface between two interacting proteins provides important clues to the function of a protein, and is becoming increasing relevant to drug discovery. Here, surface patch analysis was combined with a Bayesian network to predict protein-protein binding sites with a success rate of 82% on a benchmark dataset of 180 proteins, improving by 6% on previous work and well above the 36% that would be achieved by a random method. A comparable success rate was achieved even when evolutionary information was missing, a further improvement on our previous method which was unable to handle incomplete data automatically. In a case study of the Mog1p family, we showed that our Bayesian network method can aid the prediction of previously uncharacterised binding sites and provide important clues to protein function. On Mog1p itself a putative binding site involved in the SLN1-SKN7 signal transduction pathway was detected, as was a Ran binding site, previously characterized solely by conservation studies, even though our automated method operated without using homologous proteins. On the remaining members of the family (two structural genomics targets, and a protein involved in the photosystem II complex in higher plants) we identified novel binding sites with little correspondence to those on Mog1p. These results suggest that members of the Mog1p family bind to different proteins and probably have different functions despite sharing the same overall fold. We also demonstrated the applicability of our method to drug discovery efforts by successfully locating a number of binding sites involved in the protein-protein interaction network of papilloma virus infection. In a separate study, we attempted to distinguish between the two types of binding site, obligate and non-obligate, within our dataset using a second Bayesian network. This proved difficult although some separation was achieved on the basis of patch size, electrostatic potential and conservation. Such was the similarity between the two interacting patch types, we were able to use obligate binding site properties to predict the location of non-obligate binding sites and vice versa.

Animals↗

Arabidopsis Co-expression Tool (ACT): web server tools for microarray-based gene expression analysis.

The Arabidopsis Co-expression Tool, ACT, ranks the genes across a large microarray dataset according to how closely their expression follows the expression of a query gene. A database stores pre-calculated co-expression results for approximately 21,800 genes based on data from over 300 arrays. These results can be corroborated by calculation of co-expression results for user-defined sub-sets of arrays or experiments from the NASC/GARNet array dataset. Clique Finder (CF) identifies groups of genes which are consistently co-expressed with each other across a user-defined co-expression list. The parameters can be altered easily to adjust cluster size and the output examined for optimal inclusion of genes with known biological roles. Alternatively, a Scatter Plot tool displays the correlation coefficients for all genes against two user-selected queries on a scatter plot which can be useful for visual identification of clusters of genes with similar r-values. User-input groups of genes can be highlighted on the scatter plots. Inclusion of genes with known biology in sets of genes identified using CF and Scatter Plot tools allows inferences to be made about the roles of the other genes in the set and both tools can therefore be used to generate short lists of genes for further characterization. ACT is freely available at www.Arabidopsis.leeds.ac.uk/ACT.

Algorithms↗

Improved prediction of protein-protein binding sites using a support vector machines approach.

MOTIVATION: Structural genomics projects are beginning to produce protein structures with unknown function, therefore, accurate, automated predictors of protein function are required if all these structures are to be properly annotated in reasonable time. Identifying the interface between two interacting proteins provides important clues to the function of a protein and can reduce the search space required by docking algorithms to predict the structures of complexes. RESULTS: We have combined a support vector machine (SVM) approach with surface patch analysis to predict protein-protein binding sites. Using a leave-one-out cross-validation procedure, we were able to successfully predict the location of the binding site on 76% of our dataset made up of proteins with both transient and obligate interfaces. With heterogeneous cross-validation, where we trained the SVM on transient complexes to predict on obligate complexes (and vice versa), we still achieved comparable success rates to the leave-one-out cross-validation suggesting that sufficient properties are shared between transient and obligate interfaces. AVAILABILITY: A web application based on the method can be found at http://www.bioinformatics.leeds.ac.uk/ppi_pred. The dataset of 180 proteins used in this study is also available via the same web site. CONTACT: westhead@bmb.leeds.ac.uk SUPPLEMENTARY INFORMATION: http://www.bioinformatics.leeds.ac.uk/ppi-pred/supp-material.

Algorithms↗

Evaluation of lincomycin in drinking water for treatment of induced porcine proliferative enteropathy using a Swine challenge model.

A single-location, challenge-model study was conducted to evaluate the effectiveness of lincomycin against porcine proliferative enteropathy when administered through the drinking water at 125 and 250 mg/gallon. The primary variables of interest were pig removal rate, diarrhea scores, demeanor scores, and abdominal appearance scores. Ancillary performance variables examined included average daily feed intake, average daily gain, and feed per gain. After a 3-day acclimation period, pigs were challenged on 2 consecutive days with a mucosal homogenate containing a total dose of 1.4 x 10(9) cells of Lawsonia intracellularis. Five days later, when porcine proliferative enteropathy was well established, drinking water medicated with 125 mg (L125) or 250 mg (L250) lincomycin/gallon was provided to two groups of pigs for 10 days. Pigs were observed for 13 days following the treatment period. A third group of pigs served as controls and received unmedicated drinking water throughout the study. The L250 group experienced a significantly lower (P < .05) pig removal rate than the control group over the 23-day observation period. Additionally, for every primary variable, the L250 group experienced a significantly decreased (P < .01) number of abnormal days compared with the control group. The L125 group showed a significant reduction (P < .05) in abnormal demeanor and abnormal abdominal appearance scores compared with controls.

Administration, Oral↗

Asymmetric mutation rates at enzyme-inhibitor interfaces: implications for the protein-protein docking problem.

We have carried out a thorough and systematic sequence-structure study on how the pattern of conservation at the interface differs from the noninteracting surface in seven proteases and their inhibitors. As expected, the interface of a protease could be easily distinguished from the noninteracting surface by a concentrated area of conservation. In contrast, there was less distinction to be made between the interface and the noninteracting surface of inhibitors, and in five of the seven cases, a higher proportion of the interface area was variable compared to the rest of the surface. This is likely to cause a problem for binding-site prediction methods that assume the largest cluster of highly conserved residues on the surface of a protein corresponds to the interface. We conclude that such methods would succeed when applied to our protease test cases, but complications could arise with the inhibitors. These results also impact on methods to solve the protein-protein docking problem that use conservation at the interface to provide the location of the two protein binding sites prior to application of the docking algorithm.

Algorithms↗