Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Selection and combination of machine learning classifiers for prediction of linear B-cell epitopes on proteins.

Recently, new machine learning classifiers for the prediction of linear B-cell epitopes were presented. Here we show the application of Receiver Operator Characteristics (ROC) convex hulls to select optimal classifiers as well as possibilities to improve the post test probability (PTP) to meet real world requirements such as high throughput epitope screening of whole proteomes. The major finding is that ROC convex hulls present an easy to use way to rank classifiers based on their prediction conservativity as well as to select candidates for ensemble classifiers when validating against the antigenicity profile of 10 HIV-1 proteins. We also show that linear models are at least equally efficient to model the available data when compared to multi-layer feed-forward neural networks.

Algorithms↗

Optimal amnesic probabilistic automata or how to learn and classify proteins in linear time and space.

Statistical modeling of sequences is a central paradigm of machine learning that finds multiple uses in computational molecular biology and many other domains. The probabilistic automata typically built in these contexts are subtended by uniform, fixed-memory Markov models. In practice, such automata tend to be unnecessarily bulky and computationally imposing both during their synthesis and use. Recently, D. Ron, Y. Singer, and N. Tishby built much more compact, tree-shaped variants of probabilistic automata under the assumption of an underlying Markov process of variable memory length. These variants, called Probabilistic Suffix Trees (PSTs) were subsequently adapted by G. Bejerano and G. Yona and applied successfully to learning and prediction of protein families. The process of learning the automaton from a given training set S of sequences requires theta(Ln2) worst-case time, where n is the total length of the sequences in S and L is the length of a longest substring of S to be considered for a candidate state in the automaton. Once the automaton is built, predicting the likelihood of a query sequence of m characters may cost time theta(m2) in the worst case. The main contribution of this paper is to introduce automata equivalent to PSTs but having the following properties: Learning the automaton, for any L, takes O (n) time. Prediction of a string of m symbols by the automaton takes O (m) time. Along the way, the paper presents an evolving learning scheme and addresses notions of empirical probability and related efficient computation, which is a by-product possibly of more general interest.

Algorithms↗

Breast Cancer Recurrence Status Assessment in 5 Years Using Multimodal Integrated Learning: A Feasibility Study.

Despite advances in breast cancer detection and treatment, recurrence after curative therapy continues to impact long-term survival and quality of life. Therefore, early identification of high-risk patients is crucial to guide personalized treatment and follow-up strategies. Although genomic assays provide valuable prognostic insights, their high cost and limited accessibility hinder widespread adoption in clinical practice. Recent machine learning or deep learning approaches leveraging clinical, imaging, or multimodal data have shown promise but do not reflect real-world clinical scenarios. This study proposes a deep learning-based multimodal framework for predicting 5-year breast cancer recurrence using routinely collected clinical data. The framework consists of three main components. First, we adopted automated tumor segmentation with MedSAM to extract the tumor region from ultrasound images. The radiomics features are extracted from those tumor regions. Second, report features are extracted using a Med-Contrastive Pre-trained Transformers (MedCPT)-based approach incorporating predefined, clinically informed queries. Third, a multimodal integration model jointly processes image, radiomics, clinical features, and report features through modality-specific branches. The image branch employs the Ultrasound Foundation Model (USFM) as the backbone, while structured tabular data is processed using the FT-Transformer architecture. The features of all branches are fused using a mixture-of-experts (MoE)-based classifier, and the entire model is trained using a progressive fusion training strategy. Experimental results confirm the feasibility of using ultrasound images with tumor mask integration for recurrence prediction and demonstrate the additive value of integrating multiple data modalities through the proposed multimodal integration model. The final model for recurrence prediction achieved an AUC of 0.7540, accuracy of 74.61%, sensitivity of 70.41%, and specificity of 76.44%. This feasibility study's findings underscore the potential of the proposed multimodal deep learning framework to provide accessible, accurate, and generalizable recurrence risk prediction using routinely available clinical data, potentially supporting more informed treatment decisions and personalized post-treatment monitoring in real-world clinical practice.

Breast cancer recurrence↗

Real-time computing without stable states: a new framework for neural computation based on perturbations.

A key challenge for neural modeling is to explain how a continuous stream of multimodal input from a rapidly changing environment can be processed by stereotypical recurrent circuits of integrate-and-fire neurons in real time. We propose a new computational model for real-time computing on time-varying input that provides an alternative to paradigms based on Turing machines or attractor neural networks. It does not require a task-dependent construction of neural circuits. Instead, it is based on principles of high-dimensional dynamical systems in combination with statistical learning theory and can be implemented on generic evolved or found recurrent circuitry. It is shown that the inherent transient dynamics of the high-dimensional dynamical system formed by a sufficiently large and heterogeneous neural circuit may serve as universal analog fading memory. Readout neurons can learn to extract in real time from the current state of such recurrent neural circuit information about current and past inputs that may be needed for diverse tasks. Stable internal states are not required for giving a stable output, since transient internal states can be transformed by readout neurons into stable target outputs due to the high dimensionality of the dynamical system. Our approach is based on a rigorous computational model, the liquid state machine, that, unlike Turing machines, does not require sequential transitions between well-defined discrete internal states. It is supported, as the Turing machine is, by rigorous mathematical results that predict universal computational power under idealized conditions, but for the biologically more realistic scenario of real-time processing of time-varying inputs. Our approach provides new perspectives for the interpretation of neural coding, the design of experiments and data analysis in neurophysiology, and the solution of problems in robotics and neurotechnology.

Action Potentials↗

Using data mining to explore complex clinical decisions: A study of hospitalization after a suicide attempt.

BACKGROUND: Medical education is moving toward developing guidelines using the evidence-based approach; however, controlled data are missing for answering complex treatment decisions such as those made during suicide attempts. A new set of statistical techniques called data mining (or machine learning) is being used by different industries to explore complex databases and can be used to explore large clinical databases. METHOD: The study goal was to reanalyze, using data mining techniques, a published study of which variables predicted psychiatrists' decisions to hospitalize in 509 suicide attempters over the age of 18 years who were assessed in the emergency department. Patients were recruited for the study between 1996 and 1998. Traditional multivariate statistics were compared with data mining techniques to determine variables predicting hospitalization. RESULTS: Five analyses done by psychiatric researchers using traditional statistical techniques classified 72% to 88% of patients correctly. The model developed by researchers with no psychiatric knowledge and employing data mining techniques used 5 variables (drug consumption during the attempt, relief that the attempt was not effective, lack of family support, being a housewife, and family history of suicide attempts) and classified 99% of patients correctly (99% sensitivity and 100% specificity). CONCLUSIONS: This reanalysis of a published study fundamentally tries to make the point that these new multivariate techniques, called data mining, can be used to study large clinical databases in psychiatry. Data mining techniques may be used to explore important treatment questions and outcomes in large clinical databases and to help develop guidelines for problems where controlled data are difficult to obtain. New opportunities for good clinical research may be developed by using data mining analyses.

Adult↗

SNOSID, a proteomic method for identification of cysteine S-nitrosylation sites in complex protein mixtures.

Reversible addition of NO to Cys-sulfur in proteins, a modification termed S-nitrosylation, has emerged as a ubiquitous signaling mechanism for regulating diverse cellular processes. A key first-step toward elucidating the mechanism by which S-nitrosylation modulates a protein's function is specification of the targeted Cys (SNO-Cys) residue. To date, S-nitrosylation site specification has been laboriously tackled on a protein-by-protein basis. Here we describe a high-throughput proteomic approach that enables simultaneous identification of SNO-Cys sites and their cognate proteins in complex biological mixtures. The approach, termed SNOSID (SNO Site Identification), is a modification of the biotin-swap technique [Jaffrey, S. R., Erdjument-Bromage, H., Ferris, C. D., Tempst, P. & Snyder, S. H. (2001) Nat. Cell. Biol. 3, 193-197], comprising biotinylation of protein SNO-Cys residues, trypsinolysis, affinity purification of biotinylated-peptides, and amino acid sequencing by liquid chromatography tandem MS. With this approach, 68 SNO-Cys sites were specified on 56 distinct proteins in S-nitrosoglutathione-treated (2-10 microM) rat cerebellum lysates. In addition to enumerating these S-nitrosylation sites, the method revealed endogenous SNO-Cys modification sites on cerebellum proteins, including alpha-tubulin, beta-tubulin, GAPDH, and dihydropyrimidinase-related protein-2. Whereas these endogenous SNO proteins were previously recognized, we extend prior knowledge by specifying the SNO-Cys modification sites. Considering all 68 SNO-Cys sites identified, a machine learning approach failed to reveal a linear Cys-flanking motif that predicts stable transnitrosation by S-nitrosoglutathione under test conditions, suggesting that undefined 3D structural features determine S-nitrosylation specificity. SNOSID provides the first effective tool for unbiased elucidation of the SNO proteome, identifying Cys residues that undergo reversible S-nitrosylation.

Animals↗

Integration of single-cell transcriptomics and genomic mutation analysis identifies an immunotherapy-resistant tumor subcluster and validates ARNTL2 as a malignant driver in lung adenocarcinoma.

BACKGROUND: Immunotherapy resistance in lung adenocarcinoma (LUAD) remains a critical clinical challenge, and the mechanisms underlying resistance-associated intratumoral heterogeneity are poorly characterized. METHODS: We performed single-cell RNA sequencing of LUAD patients receiving neoadjuvant immunotherapy (responders vs. non-responders), integrating inferCNV, GSVA, and differential expression analyses. Cluster-specific genes were validated across seven independent cohorts (TCGA-LUAD, GSE13213, GSE26939, GSE29016, GSE30219, GSE31210, GSE42127). A multi-algorithm machine learning framework was used to construct a prognostic model, and the immune microenvironment was characterized using TCIA scoring, seven infiltration algorithms, and ESTIMATE. ARNTL2 function was assessed by CCK-8 and Transwell assays in A549 and H1299 cells. RESULTS: Non-responders showed significant enrichment of epithelial cells, depletion of cytotoxic T/NK cells, and elevated copy number variation burden versus responders (p < 0.0001). A resistance-enriched malignant subcluster (Cluster 2) exhibited hyperproliferative and metabolic reprogramming signatures with upregulated KRT17, S100A2, and CST6, which showed tumor-specific overexpression, adverse prognostic value, and genomic amplification across cohorts. CoxBoost combined with survivalSVM achieved optimal predictive performance (C-index = 0.686), yielding robust risk stratification (HR: 2.54-10.51, all p < 0.05). Low-risk patients showed greater immune infiltration and higher TCIA immunophenoscores. ARNTL2 was an independent prognostic factor (HR: 2.07-4.64) strongly correlated with risk score (r = 0.69), and its knockdown suppressed proliferation and invasion in both LUAD cell lines (all p < 0.05). CONCLUSION: This study identifies a resistance-associated malignant subcluster in LUAD, constructs a validated CoxBoost + survivalSVM prognostic model with robust immune stratification, and establishes ARNTL2 as a core oncogenic driver and therapeutic target.

ARNTL2↗

Semisupervised learning for molecular profiling.

Class prediction and feature selection are two learning tasks that are strictly paired in the search of molecular profiles from microarray data. Researchers have become aware how easy it is to incur a selection bias effect, and complex validation setups are required to avoid overly optimistic estimates of the predictive accuracy of the models and incorrect gene selections. This paper describes a semisupervised pattern discovery approach that uses the by-products of complete validation studies on experimental setups for gene profiling. In particular, we introduce the study of the patterns of single sample responses (sample-tracking profiles) to the gene selection process induced by typical supervised learning tasks in microarray studies. We originate sample-tracking profiles as the aggregated off-training evaluation of SVM models of increasing gene panel sizes. Genes are ranked by E-RFE, an entropy-based variant of the recursive feature elimination for support vector machines (RFE-SVM). A Dynamic Time Warping (DTW) algorithm is then applied to define a metric between sample-tracking profiles. An unsupervised clustering based on the DTW metric allows automating the discovery of outliers and of subtypes of different molecular profiles. Applications are described on synthetic data and in two gene expression studies.

Algorithms↗

Modeling of human cytochrome p450-mediated drug metabolism using unsupervised machine learning approach.

We developed a computational algorithm for evaluating the possibility of cytochrome P450-mediated metabolic transformations that xenobiotics molecules undergo in the human body. First, we compiled a database of known human cytochrome P-450 substrates, products, and nonsubstrates for 38 enzyme-specific groups (total of 2200 compounds). Second, we determined the cytochrome-mediated metabolic reactions most typical for each group and examined the substrates and products of these reactions. To assess the probability of P450 transformations of novel compounds, we built a nonlinear quantitative structure-metabolism relationships (QSMR) model based on Kohonen self-organizing maps (SOM). This neural network QSMR model incorporated a predefined set of physicochemical descriptors encoding the key molecular properties that define the metabolic fate of individual molecules. Isozyme-specific groups of substrate molecules were visualized, thus facilitating prediction of tissue-specific metabolism. The developed algorithm can be used in early stages of drug discovery as an efficient tool for the assessment of human metabolism and toxicity of novel compounds in designing discovery libraries and in lead optimization.

Algorithms↗

Gene networks inference using dynamic Bayesian networks.

This article deals with the identification of gene regulatory networks from experimental data using a statistical machine learning approach. A stochastic model of gene interactions capable of handling missing variables is proposed. It can be described as a dynamic Bayesian network particularly well suited to tackle the stochastic nature of gene regulation and gene expression measurement. Parameters of the model are learned through a penalized likelihood maximization implemented through an extended version of EM algorithm. Our approach is tested against experimental data relative to the S.O.S. DNA Repair network of the Escherichia coli bacterium. It appears to be able to extract the main regulations between the genes involved in this network. An added missing variable is found to model the main protein of the network. Good prediction abilities on unlearned data are observed. These first results are very promising: they show the power of the learning algorithm and the ability of the model to capture gene interactions.

Algorithms↗

Robust diagnosis of non-Hodgkin lymphoma phenotypes validated on gene expression data from different laboratories.

A major challenge in cancer diagnosis from microarray data is the need for robust, accurate, classification models which are independent of the analysis techniques used and can combine data from different laboratories. We propose such a classification scheme originally developed for phenotype identification from mass spectrometry data. The method uses a robust multivariate gene selection procedure and combines the results of several machine learning tools trained on raw and pattern data to produce an accurate meta-classifier. We illustrate and validate our method by applying it to gene expression datasets: the oligonucleotide HuGeneFL microarray dataset of Shipp et al. (www.genome.wi.mit.du/MPR/lymphoma) and the Hu95Av2 Affymetrix dataset (DallaFavera's laboratory, Columbia University). Our pattern-based meta-classification technique achieves higher predictive accuracies than each of the individual classifiers , is robust against data perturbations and provides subsets of related predictive genes. Our techniques predict that combinations of some genes in the p53 pathway are highly predictive of phenotype. In particular, we find that in 80% of DLBCL cases the mRNA level of at least one of the three genes p53, PLK1 and CDK2 is elevated, while in 80% of FL cases, the mRNA level of at most one of them is elevated.

Biomarkers, Tumor↗

Prediction of standard Gibbs energies of the transfer of peptide anions from aqueous solution to nitrobenzene based on support vector machine and the heuristic method.

Quantitative structure-property relationship (QSPR) method was performed for the prediction of the standard Gibbs energies (DeltaGtheta) of the transfer of peptide anions from aqueous solution to nitrobenzene. Descriptors calculated from the molecular structures alone were used to represent the characteristics of the peptides. The four molecular descriptors selected by the heuristic method (HM) in COmprehensive DEscriptors for Structural and Statistical Analysis (CODESSA) were used as inputs for support vector machine (SVM) and radial basis function neural networks (RNFNN). The results obtained by the novel machine learning technique, SVM, were compared with those obtained by HM and RBFNN. The root mean squared errors (RMS) of the training, predicted and overall data sets are 2.192, 2.541 and 2.267 unit (kJ/mol) for HM, 1.604, 2.478 and 1.817 unit (kJ/mol) for RBFNN and 1.5621, 2.364 and 1.756 unit (kJ/mol) for SVM, respectively. The prediction results were in agreement with the experimental values. This paper provided a potential method for predicting the physiochemical property (DeltaGtheta) of various small peptides.

Anions↗

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning↗

Off-road machine controls: investigating the risk of carpal tunnel syndrome.

Occupationally induced hand and wrist repetitive strain injuries (RSI) such as carpal tunnel syndrome (CTS) are a growing problem in North America. The purpose of this investigation was to apply a modification of the wrist flexion/ extension models of Armstrong and Chaffin (1978, 1979) to determine if joystick controller use in off-road machines could contribute to the development of CTS. A construction equipment cab in the laboratory was instrumented to allow force, displacement and angle measurements from 10 operators while they completed an approximately 30-min joystick motion protocol. The investigation revealed that both the external fingertip and predicted internal wrist forces resulting from the use of these joysticks were very low, indicating that the CTS risk associated with this factor was slight. However, the results also indicated that, particularly for the 'forward' and 'left' right side motions and for all left side motions, force was exerted by other portions of the fingers and hand, thereby under-predicting the tendon tension and internal wrist forces. Wrist angles observed were highest for motions that moved the joysticks to the sides rather than front to back. Thus, the 'right' and 'left' motions for both hands posed a higher risk for CTS development. When the right hand moved into the 'right' position and the left hand moved into the 'left' position, the wrist went into extension in both cases. Results indicate that neither learning nor fatigue affected the results.

Analysis of Variance↗

Ethanol sensitivity: a central role for CREB transcription regulation in the cerebellum.

BACKGROUND: Lowered sensitivity to the effects of ethanol increases the risk of developing alcoholism. Inbred mouse strains have been useful for the study of the genetic basis of various drug addiction-related phenotypes. Inbred Long-Sleep (ILS) and Inbred Short-Sleep (ISS) mice differentially express a number of genes thought to be implicated in sensitivity to the effects of ethanol. Concomitantly, there is evidence for a mediating role of cAMP/PKA/CREB signalling in aspects of alcoholism modelled in animals. In this report, the extent to which CREB signalling impacts the differential expression of genes in ILS and ISS mouse cerebella is examined. RESULTS: A training dataset for Machine Learning (ML) and Exploratory Data Analyses (EDA) was generated from promoter region sequences of a set of genes known to be targets of CREB transcription regulation and a set of genes whose transcription regulations are potentially CREB-independent. For each promoter sequence, a vector of size 132, with elements characterizing nucleotide composition features was generated. Genes whose expressions have been previously determined to be increased in ILS or ISS cerebella were identified, and their CREB regulation status predicted using the ML scheme C4.5. The C4.5 learning scheme was used because, of four ML schemes evaluated, it had the lowest predicted error rate. On an independent evaluation set of 21 genes of known CREB regulation status, C4.5 correctly classified 81% of instances with F-measures of 0.87 and 0.67 respectively for the CREB-regulated and CREB-independent classes. Additionally, six out of eight genes previously determined by two independent microarray platforms to be up-regulated in the ILS or ISS cerebellum were predicted by C4.5 to be transcriptionally regulated by CREB. Furthermore, 64% and 52% of a cross-section of other up-regulated cerebellar genes in ILS and ISS mice, respectively, were deemed to be CREB-regulated. CONCLUSION: These observations collectively suggest that ethanol sensitivity, as it relates to the cerebellum, may be associated with CREB transcription activity.

Alcoholism↗

Predicting risk of ischemic stroke: A transformer model using genomic data.

BACKGROUND AND OBJECTIVE: Ischemic stroke is a leading cause of mortality and long-term disability worldwide. Genetic factors contribute to IS susceptibility, yet conventional polygenic risk score approaches are primarily based on additive effects and may not fully capture non-linear relationships or positional context and interactions among genetic variants. This study aimed to develop and evaluate a transformer-based genomic model incorporating position-wise genotype embedding for IS risk prediction. METHODS: We conducted a genome-wide association study using the UK Biobank dataset to identify IS-associated loci. Gene prioritisation was subsequently performed using tissue-specific expression quantitative trait locus-based Mendelian randomisation and colocalization analyses in whole blood and brain cortex. We then developed a transformer-based model that encoded genotype and SNP-position information using a position-wise embedding layer. Model performance was evaluated across three UK Biobank control definitions and externally assessed in the independent All of Us cohort. Performance metrics included the area under the receiver operating characteristic curve (AUROC), precision, recall, and F1 score. RESULTS: Across the three UK Biobank control definitions, the proposed method achieved the numerically highest discrimination among the evaluated models, with AUROCs of 0.8109, 0.7843, and 0.7468 using MRF-negative, combined, and MRF-positive controls, respectively. In the external All of Us cohort, the proposed method achieved an AUROC of 0.7251 and retained the highest AUROC among the evaluated models. In a separate incident-stroke survival analysis, medium- and high-score groups had hazard ratios of 1.13 and 1.21, respectively, relative to the low-score group. A total of 18 IS-associated loci were identified. Among the tissue-specific MR results, EDEM2 in the brain cortex remained significant after Bonferroni correction, while DCHS2 showed a nominal association. CONCLUSIONS: The proposed transformer-based framework provides a genomic modelling approach that achieved the highest discrimination among the evaluated models in this study and retained comparative performance in an independent external cohort. In further applications, integrating this genomic framework with conventional clinical, lifestyle, and environmental risk factors may support more comprehensive and personalised IS risk assessment. Prospective, population-representative, and multi-ancestry validation will be important to establish its potential role in future prevention-oriented risk management.

Genomics and bioinformatics↗

CRNPRED: highly accurate prediction of one-dimensional protein structures by large-scale critical random networks.

BACKGROUND: One-dimensional protein structures such as secondary structures or contact numbers are useful for three-dimensional structure prediction and helpful for intuitive understanding of the sequence-structure relationship. Accurate prediction methods will serve as a basis for these and other purposes. RESULTS: We implemented a program CRNPRED which predicts secondary structures, contact numbers and residue-wise contact orders. This program is based on a novel machine learning scheme called critical random networks. Unlike most conventional one-dimensional structure prediction methods which are based on local windows of an amino acid sequence, CRNPRED takes into account the whole sequence. CRNPRED achieves, on average per chain, Q3 = 81% for secondary structure prediction, and correlation coefficients of 0.75 and 0.61 for contact number and residue-wise contact order predictions, respectively. CONCLUSION: CRNPRED will be a useful tool for computational as well as experimental biologists who need accurate one-dimensional protein structure predictions.

Algorithms↗

On the nature of cavities on protein surfaces: application to the identification of drug-binding sites.

In this article we introduce a new method for the identification and the accurate characterization of protein surface cavities. The method is encoded in the program SCREEN (Surface Cavity REcognition and EvaluatioN). As a first test of the utility of our approach we used SCREEN to locate and analyze the surface cavities of a nonredundant set of 99 proteins cocrystallized with drugs. We find that this set of proteins has on average about 14 distinct cavities per protein. In all cases, a drug is bound at one (and sometimes more than one) of these cavities. Using cavity size alone as a criterion for predicting drug-binding sites yields a high balanced error rate of 15.7%, with only 71.7% coverage. Here we characterize each surface cavity by computing a comprehensive set of 408 physicochemical, structural, and geometric attributes. By applying modern machine learning techniques (Random Forests) we were able to develop a classifier that can identify drug-binding cavities with a balanced error rate of 7.2% and coverage of 88.9%. Only 18 of the 408 cavity attributes had a statistically significant role in the prediction. Of these 18 important attributes, almost all involved size and shape rather than physicochemical properties of the surface cavity. The implications of these results are discussed. A SCREEN Web server is available at http://interface.bioc.columbia.edu/screen.

Binding Sites↗