Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Fairness-aware supervised hierarchical contrastive semantic learning for sexual dimorphism analysis.

MOTIVATION: Sexual dimorphism is a fundamental biological determinant driving systematic differences in disease susceptibility, progression, and clinical outcomes. However, current sex-combined AI-based genomic models often exhibit algorithmic bias and fail to capture these sex-specific mechanisms, creating a critical barrier to unbiased precision medicine. Ensuring fairness in the context of sexual dimorphism requires understanding and addressing the distinct biological mechanisms functioning in each sex, rather than focusing solely on equalizing predictive performance. RESULTS: We propose a fairness-aware supervised hierarchical contrastive learning approach, called FairHICON, to discover unbiased sex-common and sex-specific predictive features. Evaluations on cancer and asthma transcriptomic datasets demonstrate that FairHICON significantly outperforms state-of-the-art benchmarks, improving predictive performance by up to 9% while effectively reducing the performance gap between male and female sexes. Furthermore, prognostic validation confirms that the identified sex-specific pathways stratify patient survival significantly better within their corresponding sex groups. This validates FairHICON to elucidate the molecular heterogeneity of sexual dimorphism, advancing inclusive precision medicine. AVAILABILITY AND IMPLEMENTATION: The source code and data is available at https://github.com/datax-lab/FairHICON.

Sex Characteristics↗

Graph kernels for chemical informatics.

Increased availability of large repositories of chemical compounds is creating new challenges and opportunities for the application of machine learning methods to problems in computational chemistry and chemical informatics. Because chemical compounds are often represented by the graph of their covalent bonds, machine learning methods in this domain must be capable of processing graphical structures with variable size. Here, we first briefly review the literature on graph kernels and then introduce three new kernels (Tanimoto, MinMax, Hybrid) based on the idea of molecular fingerprints and counting labeled paths of depth up to d using depth-first search from each possible vertex. The kernels are applied to three classification problems to predict mutagenicity, toxicity, and anti-cancer activity on three publicly available data sets. The kernels achieve performances at least comparable, and most often superior, to those previously reported in the literature reaching accuracies of 91.5% on the Mutag dataset, 65-67% on the PTC (Predictive Toxicology Challenge) dataset, and 72% on the NCI (National Cancer Institute) dataset. Properties and tradeoffs of these kernels, as well as other proposed kernels that leverage 1D or 3D representations of molecules, are briefly discussed.

Anticarcinogenic Agents↗

Prediction of transporter family from protein sequence by support vector machine approach.

Transporters play key roles in cellular transport and metabolic processes, and in facilitating drug delivery and excretion. These proteins are classified into families based on the transporter classification (TC) system. Determination of the TC family of transporters facilitates the study of their cellular and pharmacological functions. Methods for predicting TC family without sequence alignments or clustering are particularly useful for studying novel transporters whose function cannot be determined by sequence similarity. This work explores the use of a machine learning method, support vector machines (SVMs), for predicting the family of transporters from their sequence without the use of sequence similarity. A total of 10,636 transporters in 13 TC subclasses, 1914 transporters in eight TC families, and 168,341 nontransporter proteins are used to train and test the SVM prediction system. Testing results by using a separate set of 4351 transporters and 83,151 nontransporter proteins show that the overall accuracy for predicting members of these TC subclasses and families is 83.4% and 88.0%, respectively, and that of nonmembers is 99.3% and 96.6%, respectively. The accuracies for predicting members and nonmembers of individual TC subclasses are in the range of 70.7-96.1% and 97.6-99.9%, respectively, and those of individual TC families are in the range of 60.6-97.1% and 91.5-99.4%, respectively. A further test by using 26,139 transmembrane proteins outside each of the 13 TC subclasses shows that 90.4-99.6% of these are correctly predicted. Our study suggests that the SVM is potentially useful for facilitating functional study of transporters irrespective of sequence similarity.

Amino Acid Sequence↗

Protein disorder prediction by condensed PSSM considering propensity for order or disorder.

BACKGROUND: More and more disordered regions have been discovered in protein sequences, and many of them are found to be functionally significant. Previous studies reveal that disordered regions of a protein can be predicted by its primary structure, the amino acid sequence. One observation that has been widely accepted is that ordered regions usually have compositional bias toward hydrophobic amino acids, and disordered regions are toward charged amino acids. Recent studies further show that employing evolutionary information such as position specific scoring matrices (PSSMs) improves the prediction accuracy of protein disorder. As more and more machine learning techniques have been introduced to protein disorder detection, extracting more useful features with biological insights attracts more attention. RESULTS: This paper first studies the effect of a condensed position specific scoring matrix with respect to physicochemical properties (PSSMP) on the prediction accuracy, where the PSSMP is derived by merging several amino acid columns of a PSSM belonging to a certain property into a single column. Next, we decompose each conventional physicochemical property of amino acids into two disjoint groups which have a propensity for order and disorder respectively, and show by experiments that some of the new properties perform better than their parent properties in predicting protein disorder. In order to get an effective and compact feature set on this problem, we propose a hybrid feature selection method that inherits the efficiency of uni-variant analysis and the effectiveness of the stepwise feature selection that explores combinations of multiple features. The experimental results show that the selected feature set improves the performance of a classifier built with Radial Basis Function Networks (RBFN) in comparison with the feature set constructed with PSSMs or PSSMPs that adopt simply the conventional physicochemical properties. CONCLUSION: Distinguishing disordered regions from ordered regions in protein sequences facilitates the exploration of protein structures and functions. Results based on independent testing data reveal that the proposed predicting model DisPSSMP performs the best among several of the existing packages doing similar tasks, without either under-predicting or over-predicting the disordered regions. Furthermore, the selected properties are demonstrated to be useful in finding discriminating patterns for order/disorder classification.

Amino Acid Sequence↗

Structural bioinformatics prediction of membrane-binding proteins.

Membrane-binding peripheral proteins play important roles in many biological processes, including cell signaling and membrane trafficking. Unlike integral membrane proteins, these proteins bind the membrane mostly in a reversible manner. Since peripheral proteins do not have canonical transmembrane segments, it is difficult to identify them from their amino acid sequences. As a first step toward genome-scale identification of membrane-binding peripheral proteins, we built a kernel-based machine learning protocol. Key features of known membrane-binding proteins, including electrostatic properties and amino acid composition, were calculated from their amino acid sequences and tertiary structures, which were then incorporated into the support vector machine to perform the classification. A data set of 40 membrane-binding proteins and 230 non-membrane-binding proteins was used to construct and validate the protocol. Cross-validation and holdout evaluation of the protocol showed that the accuracy of the prediction reached up to 93.7% and 91.6%, respectively. The protocol was applied to the prediction of membrane-binding properties of four C2 domains from novel protein kinases C. Although these C2 domains have 50% sequence identity, only one of them was predicted to bind the membrane, which was verified experimentally with surface plasmon resonance analysis. These results suggest that our protocol can be used for predicting membrane-binding properties of a wide variety of modular domains and may be further extended to genome-scale identification of membrane-binding peripheral proteins.

Amino Acid Sequence↗

The application of rule-based methods to class prediction problems in genomics.

We propose a method for constructing classifiers using logical combinations of elementary rules. The method is a form of rule-based classification, which has been widely discussed in the literature. In this work we focus specifically on issues that arise in the context of classifying cell samples based on RNA or protein expression measurements. The basic idea is to specify elementary rules that exhibit a locally strong pattern in favor of a single class. Strict admissibility criteria are imposed to produce a manageable universe of elementary rules. Then the elementary rules are combined using a set covering algorithm to form a composite rule that achieves a perfect fit to the training data. The user has explicit control over a parameter that determines the composite rule's level of redundancy and parsimony. This built-in control, along with the simplicity of interpreting the rules, makes the method particularly useful for classification problems in genomics. We demonstrate the new method using several microarray datasets and examine its generalization performance. We also draw comparisons to other machine-learning strategies such as CART, ID3, and C4.5.

Breast Neoplasms↗

Support vector machines for prediction of protein subcellular location.

Support Vector Machine (SVM), which is one kind of learning machines, was applied to predict the subcellular location of proteins from their amino acid composition. In this research, the proteins are classified into the following 12 groups: (1) chloroplast, (2) cytoplasm, (3) cytoskeleton, (4) endoplasmic reticulum, (5) extracall, (6) Golgi apparatus, (7) lysosome, (8) mitochondria, (9) nucleus, (10) peroxisome, (11) plasma membrane, and (12) vacuole, which have covered almost all the organelles and subcellular compartments in an animal or plant cell. The examination for the self-consistency and the jackknife test of the SVMs method was tested for the three sets: 2022 proteins, 2161 proteins, and 2319 proteins. As a result, the correct rate of self-consistency and jackknife test reaches 91 and 82% for 2022 proteins, 89 and 75% for 2161 proteins, and 85 and 73% for 2319 proteins, respectively. Furthermore, the predicting rate was tested by the three independent testing datasets containing 2240 proteins, 2513 proteins, and 2591 proteins. The correct prediction rates reach 82, 75, and 73% for 2240 proteins, 2513 proteins, and 2591 proteins, respectively.

Algorithms↗

Classification of a diverse set of Tetrahymena pyriformis toxicity chemical compounds from molecular descriptors by statistical learning methods.

Toxicity of various compounds has been measured in many studies by their toxic effects against Tetrahymena pyriformis. Efforts have also been made to use computational quantitative structure-activity relationship (QSAR) and statistical learning methods (SLMs) for predicting Tetrahymena pyriformis toxicity (TPT) at impressive accuracies. Because of the diversity of compounds and toxicity mechanisms, it is desirable to explore additional methods and to examine if these methods are applicable to more diverse sets of compounds. We tested several SLMs (logistic regression, C4.5 decision tree, k-nearest neighbor, probabilistic neural network, support vector machines) for their capability in predicting TPT by using 1129 compounds (841 TPT and 288 non-TPT agents) which are more diverse than those in other studies. A feature selection method was used for improving prediction performance and selecting molecular descriptors responsible for distinguishing TPT and non-TPT agents. The prediction accuracies are 86.9% approximately 94.2% for TPT and 71.2% approximately 87.5% for non-TPT agents based on 5-fold cross-validation studies, which are comparable to some of earlier studies despite the use of more diverse sets of compounds. The selected molecular descriptors are consistent with those used in other studies and experimental findings. These suggest that SLMs are useful for predicting TPT potential of diverse sets of compounds and for characterizing the molecular descriptors associated with TPT.

Animals↗

Prediction of ultrasound-mediated disruption of cell membranes using machine learning techniques and statistical analysis of acoustic spectra.

Although biological effects of ultrasound must be avoided for safe diagnostic applications, ultrasound's ability to disrupt cell membranes has attracted interest as a method to facilitate drug and gene delivery. This paper seeks to develop "prediction rules" for predicting the degree of cell membrane disruption based on specified ultrasound parameters and measured acoustic signals. Three techniques for generating prediction rules (regression analysis, classification trees and discriminant analysis) are applied to data obtained from a sequence of experiments on bovine red blood cells. For each experiment, the data consist of four ultrasound parameters, acoustic measurements at 400 frequencies, and a measure of cell membrane disruption. To avoid over-training, various combinations of the 404 predictor variables are used when applying the rule generation methods. The results indicate that the variable combination consisting of ultrasound exposure time and acoustic signals measured at the driving frequency and its higher harmonics yields the best rule for all three rule generation methods. The methods used for deriving the prediction rules are broadly applicable, and could be used to develop prediciton rules in other scenarios involving different cell types or tissues. These rules and the methods used to derive them could be used for real-time feedback about ultrasound's biological effects.

Animals↗

Classification of incidental carcinoma of the prostate using learning vector quantization and support vector machines.

The subclassification of incidental prostatic carcinoma into the categories T1a and T1b is of major prognostic and therapeutic relevance. In this paper an attempt was made to find out which properties mainly predispose to these two tumor categories, and whether it is possible to predict the category from a battery of clinical and histopathological variables using newer methods of multivariate data analysis. The incidental prostatic carcinomas of the decade 1990-99 diagnosed at our department were reexamined. Besides acquisition of routine clinical and pathological data, the tumours were scored by immunohistochemistry for proliferative activity and p53-overexpression. Tumour vascularization (angiogenesis) and epithelial texture were investigated by quantitative stereology. Learning vector quantization (LVQ) and support vector machines (SVM) were used for the purpose of prediction of tumour category from a set of 10 input variables (age, Gleason score, preoperative PSA value, immunohistochemical scores for proliferation and p53-overexpression, 3 stereological parameters of angiogenesis, 2 stereological parameters of epithelial texture). In a stepwise logistic regression analysis with the tumour categories T1a and T1b as dependent variables, only the Gleason score and the volume fraction of epithelial cells proved to be significant as independent predictor variables of the tumour category. Using LVQ and SVM with the information from all 10 input variables, more than 80 of the cases could be correctly predicted as T1a or T1b category with specificity, sensitivity, negative and positive predictive value from 74-92%. Using only the two significant input variables Gleason score and epithelial volume fraction, the accuracy of prediction was not worse. Thus, descriptive and quantitative texture parameters of tumour cells are of major importance for the extent of propagation in the prostate gland in incidental prostatic adenocarcinomas. Classical statistical tools and neuronal approaches led to consistent conclusions.

Adenocarcinoma↗

An in-silico method for prediction of polyadenylation signals in human sequences.

This paper presents a machine learning method to predict polyadenylation signals (PASes) in human DNA and mRNA sequences by analysing features around them. This method consists of three sequential steps of feature manipulation: generation, selection and integration of features. In the first step, new features are generated using k-gram nucleotide acid or amino acid patterns. In the second step, a number of important features are selected by an entropy-based algorithm. In the third step, support vector machines are employed to recognize true PASes from a large number of candidates. Our study shows that true PASes in DNA and mRNA sequences can be characterized by different features, and also shows that both upstream and downstream sequence elements are important for recognizing PASes from DNA sequences. We tested our method on several public data sets as well as our own extracted data sets. In most cases, we achieved better validation results than those reported previously on the same data sets. The important motifs observed are highly consistent with those reported in literature.

Base Sequence↗

Protein beta-turn prediction using nearest-neighbor method.

MOTIVATION: With the emerging success of protein secondary structure prediction through the applications of various statistical and machine learning techniques, similar techniques have been applied to protein beta-turn prediction. In this study, we perform protein beta-turn prediction using a k-nearest neighbor method, which is combined with a filter that uses predicted protein secondary structure information. Traditional beta-turn prediction from k-nearest neighbor method is modified to account for the unbalanced ratio of the natural occurrence of beta-turns and non-beta-turns. RESULTS: Our prediction scheme is tested on a set of 426 non-homologous protein sequences. The prediction scheme consists of two stages: k-nearest neighbor method stage and filtering stage. Variations of the k-nearest neighbor method were used to take property of beta-turns into consideration. Our filtering method uses beta-turn/non-beta-turn estimates from the k-nearest neighbor method stage and predicted protein secondary structure information from PSI-PRED in order to get new beta-turn/non-beta-turn estimate. Our result is compared with the previously best known beta-turn prediction method on the dataset of 426 non-homologous protein sequences and is shown to give slightly superior performance at significantly lower computational complexity. AVAILABILITY: Contact the author for information on the source code of the programs used.

Algorithms↗

The future of pediatric vesicoureteral reflux management.

BACKGROUND AND OBJECTIVE: Vesicoureteral reflux (VUR) is a common condition in pediatric urology, yet important uncertainties persist regarding risk stratification, imaging strategies, and prevention of long-term renal damage. Emerging technologies may help address these challenges. This review provides a forward-looking overview of recent advances in artificial intelligence (AI) and immunomodulation that may influence future management of pediatric VUR. METHODS: A forward-looking literature review was performed using the PubMed database (January 2000-March 2025), focusing on studies addressing AI, immunomodulation, or vaccination in the context of VUR and urinary tract infections. Criteria of inclusion were the relevance to pediatric VUR, the novelty of the proposed concept, the potential clinical implications and, for the AI literature, the existence of a clinical evaluation of the algorithm on a dataset from patients. KEY FINDINGS AND LIMITATIONS: AI-based models show promising performance in supporting clinical decision-making, including prediction of the need for voiding cystourethrography, automated grading of VUR, estimation of recurrent urinary tract infection risk and prediction of chemoprophylaxis. These tools may facilitate more individualized diagnostic and therapeutic strategies, although current evidence is largely retrospective and requires prospective validation. Immunization and immunomodulatory approaches aim to reduce infection burden and modulate inflammatory pathways associated with renal scarring. While early experimental and adult clinical data are encouraging, pediatric-specific evidence remains limited, and clinical applicability in children with VUR is not yet established. CONCLUSION: Artificial intelligence and immunologically targeted strategies represent complementary, emerging approaches that may contribute to more personalized management of pediatric VUR. At present, both should be regarded as exploratory tools whose clinical impact will depend on further validation and appropriately designed pediatric studies.

Humans↗

Learning interpretable SVMs for biological sequence classification.

BACKGROUND: Support Vector Machines (SVMs)--using a variety of string kernels--have been successfully applied to biological sequence classification problems. While SVMs achieve high classification accuracy they lack interpretability. In many applications, it does not suffice that an algorithm just detects a biological signal in the sequence, but it should also provide means to interpret its solution in order to gain biological insight. RESULTS: We propose novel and efficient algorithms for solving the so-called Support Vector Multiple Kernel Learning problem. The developed techniques can be used to understand the obtained support vector decision function in order to extract biologically relevant knowledge about the sequence analysis problem at hand. We apply the proposed methods to the task of acceptor splice site prediction and to the problem of recognizing alternatively spliced exons. Our algorithms compute sparse weightings of substring locations, highlighting which parts of the sequence are important for discrimination. CONCLUSION: The proposed method is able to deal with thousands of examples while combining hundreds of kernels within reasonable time, and reliably identifies a few statistically significant positions.

Algorithms↗

A new method to estimate ligand-receptor energetics.

In the discovery of new drugs, lead identification and optimization have assumed critical importance given the number of drug targets generated from genetic, genomics, and proteomic technologies. High-throughput experimental screening assays have been complemented recently by "virtual screening" approaches to identify and filter potential ligands when the characteristics of a target receptor structure of interest are known. Virtual screening mandates a reliable procedure for automatic ranking of structurally distinct ligands in compound library databases. Computing a rank score requires the accurate prediction of binding affinities between these ligands and the target. Many current scoring strategies require information about the target three-dimensional structure. In this study, a new method to estimate the free binding energy between a ligand and receptor is proposed. We extend a central idea previously reported (Bock, J. R., and Gough, D. A. (2001) Predicting protein-protein interactions from primary structure. Bioinformatics 17, 455-460; Bock, J. R., and Gough, D. A. (2002) Whole-proteome interaction mining. Bioinformatics, in press) that uses simple descriptors to represent biomolecules as input examples to train a support vector machine (Smola, A. J., and Schölkopf, B. (1998) A Tutorial on Support Vector Regression, Neuro-COLT Technical Report NC-TR-98-030, Royal Holloway College, University of London, UK) and the application of the trained system to previously unseen pairs, estimating their propensity for interaction. Here we seek to learn the function that maps features of a receptor-ligand pair onto their equilibrium free binding energy. These features do not comprise any direct information about the three-dimensional structures of ligand or target. In cross-validation experiments, it is demonstrated that objective measurements of prediction error rate and rank-ordering statistics are competitive with those of several other investigations, most of which depend on three-dimensional structural data. The size of the sample (n = 2,671) indicates that this approach is robust and may have widespread applicability beyond restricted families of receptor types. It is concluded that newly sequenced proteins, or those for which three-dimensional crystal structures are not easily obtained, can be rapidly analyzed for their binding potential against a library of ligands using this methodology.

Drug Design↗

Simultaneous quantification of carbamazepine crystal forms in ternary mixtures (I, III, and IV) by diffuse reflectance FTIR spectroscopy (DRIFTS) and multivariate calibration.

Diffuse reflectance FTIR spectroscopy (DRIFTS) coupled with modern multivariate calibration methods, namely artificial neural networks (ANNs) in two versions (ANN-raw and ANN-pca), support vector machines (SVMs), lazy learning (LL) and partial least squares (PLS) regression, is used in this study for the quantification of carbamazepine crystal forms in ternary powder mixtures (I, III, and IV). Two spectral regions (675-1180 and 3400-3600/cm) were selected and the data were partitioned into training and test subsets applying the Kennard-Stone design. It was found that all the selected algorithms perform better than the PLS regression (root mean squared error of prediction (RMSEP) from 3.0% to 8.2%). ANN-raw, trained on uncompressed spectral data, shows best predictive performance (RMSEP < 2.25%) but longest computation (up to 10 min). Principal component analysis (PCA) compression of the input spectral data accelerates significantly the computation (<16 s) at a relatively low cost in precision (RMSEP < 3.24%). The LL algorithm shows excellent performance in the 3400-3600/cm range (RMSEP < 1.6%), but in the 675-1180/cm range it shows strong dependence on data set structure (RMSEP between 1.6% and 8.9%). SVMs perform comparably well with ANNs (RMSEP < 3.1%), not showing the long computation time of ANNs (<1 s) and therefore may provide an attractive alternative to ANNs.

Algorithms↗

[Rule induction algorithm for brain glioma using support vector machine].

A new proposed data mining technique, support vector machine (SVM), is used to predict the degree of malignancy in brain glioma. Based on statistical learning theory, SVM realizes the principle of data dependent structure risk minimization, so it can depress the overfitting with better generalization performance, since the prediction in medical diagnosis often deals with a small sample. SVM based rule induction algorithm is implemented in comparison with other data mining techniques such as artificial neural networks, rule induction algorithm and fuzzy rule extraction algorithm based on fuzzy max-min neural networks (FRE-FMMNN) proposed recently. Computation results by 10 fold cross validation method show that SVM can get higher prediction accuracy than artificial neural networks and FRE-FMMNN, which implies SVM can get higher accuracy and more reliability. On the whole data sets, SVM gets one rule with the classification accuracy of 89.29%, while FRE-FMMNN gets two rules of 84. 64%, in which the rule got by SVM is of quantity relation and contains more information than the two rules by FRE-FMMNN. All the above show SVM is a potential algorithm for the medical diagnosis such as the prediction of the degree of malignancy in brain glioma.

Algorithms↗

Functional interaction of nitrogenous organic bases with cytochrome P450: a critical assessment and update of substrate features and predicted key active-site elements steering the access, binding, and orientation of amines.

The widespread use of nitrogenous organic bases as environmental chemicals, food additives, and clinically important drugs necessitates precise knowledge about the molecular principles governing biotransformation of this category of substrates. In this regard, analysis of the topological background of complex formation between amines and P450s, acting as major catalysts in C- and N-oxidative attack, is of paramount importance. Thus, progress in collaborative investigations, combining physico-chemical techniques with chemical-modification as well as genetic engineering experiments, enables substantiation of hypothetical work resulting from the design of pharmacophores or homology modelling of P450s. Based on a general, CYP2D6-related construct, the majority of prospective amine-docking residues was found to cluster near the distal heme face in the six known SRSs, made up by the highly variant helices B', F and G as well as the N-terminal portion of helix C and certain beta-structures. Most of the contact sites examined show a frequency of conservation < 20%, hinting at the requirement of some degree of conformational versatility, while a limited number of amino acids exhibiting a higher level of conservation reside close to the heme core. Some key determinants may have a dual role in amine binding and/or maintenance of protein integrity. Importantly, a series of non-SRS elements are likely to be operative via long-range effects. While hydrophobic mechanisms appear to dominate orientation of the nitrogenous compounds toward the iron-oxene species, polar residues seem to foster binding events through H-bonding or salt-bridge formation. Careful uncovering of structure-function relationships in amine-enzyme association together with recently developed unsupervised machine learning approaches will be helpful in both tailoring of novel amine-type drugs and early elimination of potentially toxic or mutagenic candidates. Also, chimeragenesis might serve in the construction of more efficient P450s for activation of amine drugs and/or bioremediation.

Amines↗