Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

Sensor fusion by pseudo information measure: a mobile robot application.

In any autonomous mobile robot, one of the most important issues to be designed and implemented is environment perception. In this paper, a new approach is formulated in order to perform sensory data integration for generation of an occupancy grid map of the environment. This method is an extended version of the Bayesian fusion method for independent sources of information. The performance of the proposed method of fusion and its sensitivity are discussed. Map building simulation for a cylindrical robot with eight ultrasonic sensors and mapping implementation for a Khepera robot have been separately tried in simulation and experimental works. A new neural structure is introduced for conversion of proximity data that are given by Khepera IR sensors to occupancy probabilities. Path planning experiments have also been applied to the resulting maps. For each map, two factors are considered and calculated: the fitness and the augmented occupancy of the map with respect to the ideal map. The length and the least distance to obstacles were the other two factors that were calculated for the routes that are resulted by path planning experiments. Experimental and simulation results show that by using the new fusion formulas, more informative maps of the environment are obtained. By these maps more appropriate routes could be achieved. Actually, there is a tradeoff between the length of the resulting routes and their safety and by choosing the proper fusion function, this tradeoff is suitably tuned for different map building applications.

Algorithms↗

Hierarchical Bayesian estimation for MEG inverse problem.

Source current estimation from MEG measurement is an ill-posed problem that requires prior assumptions about brain activity and an efficient estimation algorithm. In this article, we propose a new hierarchical Bayesian method introducing a hierarchical prior that can effectively incorporate both structural and functional MRI data. In our method, the variance of the source current at each source location is considered an unknown parameter and estimated from the observed MEG data and prior information by using the Variational Bayesian method. The fMRI information can be imposed as prior information on the variance distribution rather than the variance itself so that it gives a soft constraint on the variance. A spatial smoothness constraint, that the neural activity within a few millimeter radius tends to be similar due to the neural connections, can also be implemented as a hierarchical prior. The proposed method provides a unified theory to deal with the following three situations: (1) MEG with no other data, (2) MEG with structural MRI data on cortical surfaces, and (3) MEG with both structural MRI and fMRI data. We investigated the performance of our method and conventional linear inverse methods under these three conditions. Simulation results indicate that our method has better accuracy and spatial resolution than the conventional linear inverse methods under all three conditions. It is also shown that accuracy of our method improves as MRI and fMRI information becomes available. Simulation results demonstrate that our method appropriately resolves the inverse problem even if fMRI data convey inaccurate information, while the Wiener filter method is seriously deteriorated by inaccurate fMRI information.

Algorithms↗

Predictive Bayesian neural network models of MHC class II peptide binding.

We used Bayesian regularized neural networks to model data on the MHC class II-binding affinity of peptides. Training data consisted of sequences and binding data for nonamer (nine amino acid) peptides. Independent test data consisted of sequences and binding data for peptides of length </=25. We assumed that MHC class II-binding activity of peptides depends only on the highest ranked embedded nonamer and that reverse sequences of active nonamers are inactive. We also internally validated the models by using 30% of the training data in an internal test set. We obtained robust models, with near identical statistics for multiple training runs. We determined how predictive our models were using statistical tests and area under the Receiver Operating Characteristic (ROC) graphs (A(ROC)). Most models gave training A(ROC) values close to 1.0 and test set A(ROC) values >0.8. We also used both amino acid indicator variables (bin20) and property-based descriptors to generate models for MHC class II-binding of peptides. The property-based descriptors were more parsimonious than the indicator variable descriptors, making them applicable to larger peptides, and their design makes them able to generalize to unknown peptides outside of the training space. None of the external test data sets contained any of the nonamer sequences in the training sets. Consequently, the models attempted to predict the activity of truly unknown peptides not encountered in the training sets. Our models were well able to tackle the difficult problem of correctly predicting the MHC class II-binding activities of a majority of the test set peptides. Exceptions to the assumption that nonamer motif activities were invariant to the peptide in which they were embedded, together with the limited coverage of the test data, and the fuzziness of the classification procedure, are likely explanations for some misclassifications.

Amino Acid Sequence↗

Comparison and testing of least-squares time domain inverse solutions in electrocardiography.

The use of several mathematical methods for estimating epicardial ECG potentials from arrays of body surface potentials has been reported in the literature; most of these methods are based on least-squares reconstruction principles and operate in the time-space domain. In this paper we introduce a general Bayesian maximum a posteriori (MAP) framework for time domain inverse solutions in the presence of noise. The two most popular previously applied least-squares methods, constrained (regularized) least-squares and low-rank approximation through the singular value decomposition, are placed in this framework, each of them requiring the a priori knowledge of a 'regularization parameter', which defines the degree of smoothing to be applied to the inversion. Results of simulations using these two methods are presented; they compare the ability of each method to reconstruct epicardial potentials. We used the geometric configuration of the torso and internal organs of an individual subject as reconstructed from CT scans. The accuracy of each method at each epicardial location was tested as a function of measurement noise, the size and shape of the subarray of torso sensors, and the regularization parameter. We paid particular attention to an assessment of the potential of these methods for clinical use by testing the effect of using compact, small-size subarrays of torso potentials while maintaining a high degree of resolution on the epicardium.

Algorithms↗

Ideal-observer performance under signal and background uncertainty.

We use the performance of the Bayesian ideal observer as a figure of merit for hardware optimization because this observer makes optimal use of signal-detection information. Due to the high dimensionality of certain integrals that need to be evaluated, it is difficult to compute the ideal observer test statistic, the likelihood ratio, when background variability is taken into account. Methods have been developed in our laboratory for performing this computation for fixed signals in random backgrounds. In this work, we extend these computational methods to compute the likelihood ratio in the case where both the backgrounds and the signals are random with known statistical properties. We are able to write the likelihood ratio as an integral over possible backgrounds and signals, and we have developed Markov-chain Monte Carlo (MCMC) techniques to estimate these high-dimensional integrals. We can use these results to quantify the degradation of the ideal-observer performance when signal uncertainties are present in addition to the randomness of the backgrounds. For background uncertainty, we use lumpy backgrounds. We present the performance of the ideal observer under various signal-uncertainty paradigms with different parameters of simulated parallel-hole collimator imaging systems. We are interested in any change in the rankings between different imaging systems under signal and background uncertainty compared to the background-uncertainty case. We also compare psychophysical studies to the performance of the ideal observer.

Algorithms↗

Comparative evaluation of a new effective population size estimator based on approximate bayesian computation.

We describe and evaluate a new estimator of the effective population size (N(e)), a critical parameter in evolutionary and conservation biology. This new "SummStat" N(e) estimator is based upon the use of summary statistics in an approximate Bayesian computation framework to infer N(e). Simulations of a Wright-Fisher population with known N(e) show that the SummStat estimator is useful across a realistic range of individuals and loci sampled, generations between samples, and N(e) values. We also address the paucity of information about the relative performance of N(e) estimators by comparing the SummStat estimator to two recently developed likelihood-based estimators and a traditional moment-based estimator. The SummStat estimator is the least biased of the four estimators compared. In 32 of 36 parameter combinations investigated using initial allele frequencies drawn from a Dirichlet distribution, it has the lowest bias. The relative mean square error (RMSE) of the SummStat estimator was generally intermediate to the others. All of the estimators had RMSE > 1 when small samples (n = 20, five loci) were collected a generation apart. In contrast, when samples were separated by three or more generations and N(e) < or = 50, the SummStat and likelihood-based estimators all had greatly reduced RMSE. Under the conditions simulated, SummStat confidence intervals were more conservative than the likelihood-based estimators and more likely to include true N(e). The greatest strength of the SummStat estimator is its flexible structure. This flexibility allows it to incorporate any potentially informative summary statistic from population genetic data.

Bayes Theorem↗

The diagnosis of polyarteritis nodosa. I. A literature-based decision analysis approach.

We investigated diagnostic testing in polyarteritis nodosa (PAN) by calculating, from published data, the sensitivity and specificity of visceral angiography and muscle, nerve, testicle, kidney, and liver biopsy. Test sequence strategies were constructed by Bayesian inference using a computer program written for this purpose. Test sequences were compared with an aggressive strategy consisting of repeated tests until there was a positive finding or until the available tests were exhausted, and a conservative strategy consisting of 1 biopsy procedure plus angiography. The Bayesian analysis agreed most closely with the conservative approach for most prior probabilities (degree of suspicion) that a patient had PAN. The aggressive strategy had an overall sensitivity of 90% and specificity of 91%, whereas the conservative strategy was 85% sensitive and 96% specific. Furthermore, the aggressive strategy was more costly ($2,986 versus $1,961) and had a higher rate of morbidity (3.8 versus 2.7 days of hospitalization per patient evaluated) than did the conservative strategy. The mortality rates of both strategies were equivalent (approximately 0.05 deaths per hundred patients evaluated). The per-case cost of diagnosis increased as prevalence decreased, and at 10% prevalence, the aggressive strategy cost more than $17,000 per case diagnosed. Sensitivity analysis revealed that the strategies were moderately affected by the test characteristics, within reasonable assumptions, but that the differences in conservative and aggressive approaches remained. Thus, our analysis based on available data and the assumption of test independence suggests that the preferred diagnostic evaluation of patients with symptoms suggestive of PAN consists, in most cases, of a single biopsy procedure, with angiographic evaluation if necessary.

Costs and Cost Analysis↗

Sparse Bayesian learning for efficient visual tracking.

This paper extends the use of statistical learning algorithms for object localization. It has been shown that object recognizers using kernel-SVMs can be elegantly adapted to localization by means of spatial perturbation of the SVM. While this SVM applies to each frame of a video independently of other frames, the benefits of temporal fusion of data are well-known. This is addressed here by using a fully probabilistic Relevance Vector Machine (RVM) to generate observations with Gaussian distributions that can be fused over time. Rather than adapting a recognizer, we build a displacement expert which directly estimates displacement from the target region. An object detector is used in tandem, for object verification, providing the capability for automatic initialization and recovery. This approach is demonstrated in real-time tracking systems where the sparsity of the RVM means that only a fraction of CPU time is required to track at frame rate. An experimental evaluation compares this approach to the state of the art showing it to be a viable method for long-term region tracking.

Algorithms↗

Statistical methods for gene map construction by fluorescence in situ hybridization.

Fluorescence in situ hybridization (FISH) provides an efficient and powerful technique for ordering loci both on metaphase chromosomes and in less condensed interphase chromatin. Two-color metaphase FISH can be used to order pairs of loci relative to the centromere; two- and three-color interphase FISH can be used to accurately order trios of loci spaced within 1 Mb relative to one another. Loci separated by a distance > 1-2 Mb exhibit chromatin loops that often give rise to a statistically significant but incorrect order. We derive Bayesian methods for selecting the best locus order based on microscopic evaluation for each of these types of FISH mapping data. We then describe how the results from several two- and three-locus analyses can be combined to evaluate the approximate posterior probability of a given multilocus order within the limits of the technology utilized. These methods directly address the question of interest: What is the probability that the inferred two-, three-, or multilocus order actually is correct? We illustrate our analysis methods by applying them to previously described FISH mapping data of 14 markers in the BRCA1 region on chromosome 17q12-q21. We also propose design strategies to order a group of closely spaced (< 1 Mb) loci, two and three loci at a time, using a bisection strategy for two-color FISH data and a trisection strategy for three-color FISH data. These strategies have the best worst-case performance for ordering a new locus relative to a group of ordered loci and are nearly optimal for ordering a group of loci of unknown order. These, in conjunction with physical mapping strategies, provide efficient and reliable methods for gene map construction by FISH.

BRCA1 Protein↗

The cost-effectiveness of basiliximab induction in "old-to-old" kidney transplant programs: Bayesian estimation, simulation, and uncertainty analysis.

INTRODUCTION: Markov models are employed in economic analyses to evaluate all possible expectations in a dilemna. The introduction of a new clinical protocol (Basiliximab induction with calcineurin-sparing protocols) for a group of kidney transplant recipients receiving organs from marginal donors was validated with a Markov simulation model, demonstrating the usefulness of combining simulation with Bayesian estimation methods for analysis of cost-effectiveness data collected alongside a clinical trial. We sought to determine whether calcineurin-sparing protocols using anti-interleukin-2/antibody induction (Simulect) would show a beneficial effect on initial kidney function and reduce transplantation costs upon admission, clinical incidences, graft function, and complications during the first month after transplant. PATIENTS AND METHODS: A Markov Chain Monte Carlo (MCMC) was used to estimate a system of generalized linear models relating costs and outcomes to a kidney transplant process affected by treatment under alternative therapies. The Markov simulation model was established following three chains: a calcineurin-free regimen with Basiliximab induction (chain A); a calcineurin-sparing protocol with Basiliximab induction (chain B); and a conventional immunosuppressive regimen (chain C). The MCMC draws were used as parameters in simulations that yielded inferences about the relative cost-effectiveness of the novel therapy under a variety of scenarios. After designing the Markov chain and cohorts, 31 patients from the "old-to-old" program were assigned; eight to chain A; eight to chain B; and 15 to chain C. A year after transplantation a cost-benefit study was performed guided by the three branches of the Markov model. RESULTS: The Markov model showed a benefit of induction therapies in elderly patients. A cost-benefit model showed that after a year, there was a clear benefit from calcineurin-free plus Basiliximab induction therapies, with a slight benefit from calcineurin-sparing protocols. CONCLUSIONS: Markov models are extremely useful when introducing new clinical therapies. The approach allows flexibility in assessing treatment using various premises and quantifies the global effect of parametric uncertainty on a decision maker's confidence to adopt one therapy over another. In our transplant program, a cost-effective analysis of outcomes in old patients using the Markov model showed a clear benefit of calcineurin-sparing protocols with Basixilimab induction.

Age Factors↗

Numbers of mutations to different types of colorectal cancer.

BACKGROUND: The numbers of oncogenic mutations required for transformation are uncertain but may be inferred from how cancer frequencies increase with aging. Cancers requiring more mutations will tend to appear later in life. This type of approach may be confounded by biologic heterogeneity because different cancer subtypes may require different numbers of mutations. For example, a sporadic cancer should require at least one more somatic mutation relative to its hereditary counterpart. METHODS: To better estimate numbers of mutations before transformation, 1,022 colorectal cancers were classified with respect to microsatellite instability (MSI) and germline DNA mismatch repair mutations characteristic of hereditary nonpolyposis colorectal cancer (HNPCC). MSI- cancers were also classified with respect to clinical stage. Ages at cancer and a Bayesian algorithm were used to estimate the numbers of oncogenic mutations required for transformation for each cancer subtype. RESULTS: Ages at MSI+ cancers were consistent with five or six oncogenic mutations for hereditary (HNPCC) cancers, and seven or eight mutations for its sporadic counterpart. Ages at cancer were consistent with seven mutations for sporadic MSI- cancers, and were similar (six to eight mutations) regardless of clinical cancer stage. CONCLUSION: Different biologic subtypes of colorectal cancer appear to require different numbers of oncogenic mutations before transformation. Sporadic MSI+ cancers may require more than a single additional somatic alteration compared to hereditary MSI+ cancers because the epigenetic inactivation of MLH1 commonly observed in sporadic MSI+ cancers may be a multistep process. Interestingly, estimated numbers of MSI- cancer mutations were similar (six to eight mutations) regardless of clinical cancer stage, suggesting a propensity to spread or metastasize does not require additional mutations after transformation. Estimates of oncogenic mutation numbers may help explain some of the biology underlying different cancer subtypes.

Adult↗

Identification of co-regulated genes through Bayesian clustering of predicted regulatory binding sites.

The identification of co-regulated genes and their transcription-factor binding sites (TFBS) are key steps toward understanding transcription regulation. In addition to effective laboratory assays, various computational approaches for the detection of TFBS in promoter regions of coexpressed genes have been developed. The availability of complete genome sequences combined with the likelihood that transcription factors and their cognate sites are often conserved during evolution has led to the development of phylogenetic footprinting. The modus operandi of this technique is to search for conserved motifs upstream of orthologous genes from closely related species. The method can identify hundreds of TFBS without prior knowledge of co-regulation or coexpression. Because many of these predicted sites are likely to be bound by the same transcription factor, motifs with similar patterns can be put into clusters so as to infer the sets of co-regulated genes, that is, the regulons. This strategy utilizes only genome sequence information and is complementary to and confirmative of gene expression data generated by microarray experiments. However, the limited data available to characterize individual binding patterns, the variation in motif alignment, motif width, and base conservation, and the lack of knowledge of the number and sizes of regulons make this inference problem difficult. We have developed a Gibbs sampling-based Bayesian motif clustering (BMC) algorithm to address these challenges. Tests on simulated data sets show that BMC produces many fewer errors than hierarchical and K-means clustering methods. The application of BMC to hundreds of predicted gamma-proteobacterial motifs correctly identified many experimentally reported regulons, inferred the existence of previously unreported members of these regulons, and suggested novel regulons.

Algorithms↗

Application of Bayesian regularized BP neural network model for analysis of aquatic ecological data-a case study of chlorophyll-a prediction in Nanzui water area of Dongting Lake.

Bayesian regularized BP neural network(BRBPNN) technique was applied in the chlorophyll-a prediction of Nanzui water area in Dongting Lake. Through BP network interpolation method, the input and output samples of the network were obtained. After the selection of input variables using stepwise/multiple linear regression method in SPSS 11.0 software, the BRBPNN model was established between chlorophyll-a and environmental parameters, biological parameters. The achieved optimal network structure was 3-11-1 with the correlation coefficients and the mean square errors for the training set and the test set as 0.999 and 0.00078426, 0.981 and 0.0216 respectively. The sum of square weights between each input neuron and the hidden layer of optimal BRBPNN models of different structures indicated that the effect of individual input parameter on chlorophyll-a declined in the order of alga amount > secchi disc depth (SD) > electrical conductivity (EC). Additionally, it also demonstrated that the contributions of these three factors were the maximal for the change of chlorophyll-a concentration, total phosphorus (TP) and total nitrogen (TN) were the minimal. All the results showed that BRBPNN model was capable of automated regularization parameter selection and thus it may ensure the excellent generation ability and robustness. Thus, this study laid the foundation for the application of BRBPNN model in the analysis of aquatic ecological data(chlorophyll-a prediction) and the explanation about the effective eutrophication treatment measures for Nanzui water area in Dongting Lake.

Bayes Theorem↗

Synergistic techniques for better understanding and classifying the environmental structure of landscapes.

The desire to capture natural regions in the landscape has been a goal of geographic and environmental classification and ecological land classification (ELC) for decades. Since the increased adoption of data-centric, multivariate, computational methods, the search for natural regions has become the search for the best classification that optimally trades off classification complexity for class homogeneity. In this study, three techniques are investigated for their ability to find the best classification of the physical environments of the Mt. Lofty Ranges in South Australia: AutoClass-C (a Bayesian classifier), a Kohonen Self-Organising Map neural network, and a k-means classifier with homogeneity analysis. AutoClass-C is specifically designed to find the classification that optimally trades off classification complexity for class homogeneity. However, AutoClass analysis was not found to be assumption-free because it was very sensitive to the user-specified level of relative error of input data. The AutoClass results suggest that there may be no way of finding the best classification without making critical assumptions as to the level of class heterogeneity acceptable in the classification when using continuous environmental data. Therefore, rather than relying on adjusting abstract parameters to arrive at a classification of suitable complexity, it is better to quantify and visualize the data structure and the relationship between classification complexity and class homogeneity. Individually and when integrated, the Self-Organizing Map and k-means classification with homogeneity analysis techniques also used in this study facilitate this and provide information upon which the decision of the scale of classification can be made. It is argued that instead of searching for the elusive classification of natural regions in the landscape, it is much better to understand and visualize the environmental structure of the landscape and to use this knowledge to select the best ELC at the required scale of analysis.

Bayes Theorem↗

Three globin lineages belonging to two structural classes in genomes from the three kingdoms of life.

Although most globins, including the N-terminal domains within chimeric proteins such as flavohemoglobins and globin-coupled sensors, exhibit a 3/3 helical sandwich structure, many bacterial, plant, and ciliate globins have a 2/2 helical sandwich structure. We carried out a comprehensive survey of globins in the genomes from the three kingdoms of life. Bayesian phylogenetic trees based on manually aligned sequences indicate the possibility of past horizontal globin gene transfers from bacteria to eukaryotes. blastp searches revealed the presence of 3/3 single-domain globins related to the globin domains of the bacterial and fungal flavohemoglobins in many bacteria, a red alga, and a diatom. Iterated psi-blast searches based on groups of globin sequences found that only the single-domain globins and flavohemoglobins recognize the eukaryote 3/3 globins, including vertebrate neuroglobins, alpha- and beta-globins, and cytoglobins. The 2/2 globins recognize the flavohemoglobins, as do the globin coupled sensors and the closely related single-domain protoglobins. However, the 2/2 globins and the globin-coupled sensors do not recognize each other. Thus, all globins appear to be distributed among three lineages: (i) the 3/3 plant and metazoan globins, single-domain globins, and flavohemoglobins; (ii) the bacterial 3/3 globin-coupled sensors and protoglobins; and (iii) the bacterial, plant, and ciliate 2/2 globins. The three lineages may have evolved from an ancestral 3/3 or 2/2 globin. Furthermore, it appears likely that the predominant functions of globins are enzymatic and that oxygen transport is a specialized development that accompanied the evolution of metazoans.

Bayes Theorem↗

Selective integration of multiple biological data for supervised network inference.

MOTIVATION: Inferring networks of proteins from biological data is a central issue of computational biology. Most network inference methods, including Bayesian networks, take unsupervised approaches in which the network is totally unknown in the beginning, and all the edges have to be predicted. A more realistic supervised framework, proposed recently, assumes that a substantial part of the network is known. We propose a new kernel-based method for supervised graph inference based on multiple types of biological datasets such as gene expression, phylogenetic profiles and amino acid sequences. Notably, our method assigns a weight to each type of dataset and thereby selects informative ones. Data selection is useful for reducing data collection costs. For example, when a similar network inference problem must be solved for other organisms, the dataset excluded by our algorithm need not be collected. RESULTS: First, we formulate supervised network inference as a kernel matrix completion problem, where the inference of edges boils down to estimation of missing entries of a kernel matrix. Then, an expectation-maximization algorithm is proposed to simultaneously infer the missing entries of the kernel matrix and the weights of multiple datasets. By introducing the weights, we can integrate multiple datasets selectively and thereby exclude irrelevant and noisy datasets. Our approach is favorably tested in two biological networks: a metabolic network and a protein interaction network. AVAILABILITY: Software is available on request.

Algorithms↗

Expert system support using Bayesian belief networks in the diagnosis of fine needle aspiration biopsy specimens of the breast.

AIM: To develop an expert system model for the diagnosis of fine needle aspiration cytology (FNAC) of the breast. METHODS: Knowledge and uncertainty were represented in the form of a Bayesian belief network which permitted the combination of diagnostic evidence in a cumulative manner and provided a final probability for the possible diagnostic outcomes. The network comprised 10 cytological features (evidence nodes), each independently linked to the diagnosis (decision node) by a conditional probability matrix. The system was designed to be interactive in that the cytopathologist entered evidence into the network in the form of likelihood ratios for the outcomes at each evidence node. RESULTS: The efficiency of the network was tested on a series of 40 breast FNAC specimens. The highest diagnostic probability provided by the network agreed with the cytopathologists' diagnosis in 100% of cases for the assessment of discrete, benign, and malignant aspirates. Atypical probably benign cases were given probabilities in favour of a benign diagnosis. Suspicious cases tended to have similar probabilities for both diagnostic outcomes and so, correctly, could not be assigned as benign or malignant. A closer examination of cumulative belief graphs for the diagnostic sequence of each case provided insight into the diagnostic process, and quantitative data which improved the identification of suspicious cases. CONCLUSION: The further development of such a system will have three important roles in breast cytodiagnosis: (1) to aid the cytologist in making a more consistent and objective diagnosis; (2) to provide a teaching tool on breast cytological diagnosis for the non-expert; and (3) it is the first stage in the development of a system capable of automated diagnosis through the use of expert system machine vision.

Bayes Theorem↗

Virtual screening to enrich a compound collection with CDK2 inhibitors using docking, scoring, and composite scoring models.

Docking programs can generate subsets of a compound collection with an increased percentage of actives against a target (enrichment) by predicting their binding mode (pose) and affinity (score), and retrieving those with the highest scores. Using the QXP and GOLD programs, we compared the ability of six single scoring functions (PLP, Ligscore, Ludi, Jain, ChemScore, PMF) and four composite scoring models (Mean Rank: MR, Rank-by-Vote: Vt, Bayesian Statistics: BS and PLS Discriminant Analysis: DA) to separate compounds that are active against CDK2 from inactives. We determined the enrichment for the entire set of actives (IC50 < 10 microM) and for three activity subsets. In all cases, the enrichment for each subset was lower than for the entire set of actives. QXP outperformed GOLD at pose prediction, but yielded only moderately better enrichments. Five to six scoring functions yielded good enrichments with GOLD poses, while typically only two worked well with QXP poses. For each program, two scoring functions generally performed better than the others (Ligscore2 and Ludi for GOLD; QXP and Jain for QXP). Composite scoring functions yielded better results than single scoring functions. The consensus approaches MR and Vt worked best when separating micromolar inhibitors from inactives. The statistical approaches BS and DA, which require training data, performed best when distinguishing between low and high nanomolar inhibitors. The key observation that all hit rate profiles for all four activity intervals for all scoring schemes for both programs are significantly better than random, is evidence that docking can be successfully applied to enrich compound collections.

Adenosine Triphosphate↗