Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

A neural-network-based method for predicting protein stability changes upon single point mutations.

MOTIVATION: One important requirement for protein design is to be able to predict changes of protein stability upon mutation. Different methods addressing this task have been described and their performance tested considering global linear correlation between predicted and experimental data. Neither is direct statistical evaluation of their prediction performance available, nor is a direct comparison among different approaches possible. Recently, a significant database of thermodynamic data on protein stability changes upon single point mutation has been generated (ProTherm). This allows the application of machine learning techniques to predicting free energy stability changes upon mutation starting from the protein sequence. RESULTS: In this paper, we present a neural-network-based method to predict if a given mutation increases or decreases the protein thermodynamic stability with respect to the native structure. Using a dataset consisting of 1615 mutations, our predictor correctly classifies >80% of the mutations in the database. On the same task and using the same data, our predictor performs better than other methods available on the Web. Moreover, when our system is coupled with energy-based methods, the joint prediction accuracy increases up to 90%, suggesting that it can be used to increase also the performance of pre-existing methods, and generally to improve protein design strategies. AVAILABILITY: The server is under construction and will be available at http://www.biocomp.unibo.it

Algorithms↗

RASE: recognition of alternatively spliced exons in C.elegans.

MOTIVATION: Eukaryotic pre-mRNAs are spliced to form mature mRNA. Pre-mRNA alternative splicing greatly increases the complexity of gene expression. Estimates show that more than half of the human genes and at least one-third of the genes of less complex organisms, such as nematodes or flies, are alternatively spliced. In this work, we consider one major form of alternative splicing, namely the exclusion of exons from the transcript. It has been shown that alternatively spliced exons have certain properties that distinguish them from constitutively spliced exons. Although most recent computational studies on alternative splicing apply only to exons which are conserved among two species, our method only uses information that is available to the splicing machinery, i.e. the DNA sequence itself. We employ advanced machine learning techniques in order to answer the following two questions: (1) Is a certain exon alternatively spliced? (2) How can we identify yet unidentified exons within known introns? RESULTS: We designed a support vector machine (SVM) kernel well suited for the task of classifying sequences with motifs having positional preferences. In order to solve the task (1), we combine the kernel with additional local sequence information, such as lengths of the exon and the flanking introns. The resulting SVM-based classifier achieves a true positive rate of 48.5% at a false positive rate of 1%. By scanning over single EST confirmed exons we identified 215 potential alternatively spliced exons. For 10 randomly selected such exons we successfully performed biological verification experiments and confirmed three novel alternatively spliced exons. To answer question (2), we additionally used SVM-based predictions to recognize acceptor and donor splice sites. Combined with the above mentioned features we were able to identify 85.2% of skipped exons within known introns at a false positive rate of 1%. AVAILABILITY: Datasets, model selection results, our predictions and additional experimental results are available at http://www.fml.tuebingen.mpg.de/~raetsch/RASE SUPPLEMENTARY INFORMATION: http://www.fml.tuebingen.mpg.de/raetsch/RASE.

Algorithms↗

Descriptor-based protein remote homology identification.

Here, we report a novel protein sequence descriptor-based remote homology identification method, able to infer fold relationships without the explicit knowledge of structure. In a first phase, we have individually benchmarked 13 different descriptor types in fold identification experiments in a highly diverse set of protein sequences. The relevant descriptors were related to the fold class membership by using simple similarity measures in the descriptor spaces, such as the cosine angle. Our results revealed that the three best-performing sets of descriptors were the sequence-alignment-based descriptor using PSI-BLAST e-values, the descriptors based on the alignment of secondary structural elements (SSEA), and the descriptors based on the occurrence of PROSITE functional motifs. In a second phase, the three top-performing descriptors were combined to obtain a final method with improved performance, which we named DescFold. Class membership was predicted by Support Vector Machine (SVM) learning. In comparison with the individual PSI-BLAST-based descriptor, the rate of remote homology identification increased from 33.7% to 46.3%. We found out that the composite set of descriptors was able to identify the true remote homolog for nearly every sixth sequence at the 95% confidence level, or some 10% more than a single PSI-BLAST search. We have benchmarked the DescFold method against several other state-of-the-art fold recognition algorithms for the 172 LiveBench-8 targets, and we concluded that it was able to add value to the existing techniques by providing a confident hit for at least 10% of the sequences not identifiable by the previously known methods.

Algorithms↗

Machine learning in prognosis of the femoral neck fracture recovery.

We compare the performance of several machine learning algorithms in the problem of prognostics of the femoral neck fracture recovery: the K-nearest neighbours algorithm, the semi-naive Bayesian classifier, backpropagation with weight elimination learning of the multilayered neural networks, the LFC (lookahead feature construction) algorithm, and the Assistant-I and Assistant-R algorithms for top down induction of decision trees using information gain and RELIEFF as search heuristics, respectively. We compare the prognostic accuracy and the explanation ability of different classifiers. Among the different algorithms the semi-naive Bayesian classifier and Assistant-R seem to be the most appropriate. We analyze the combination of decisions of several classifiers for solving prediction problems and show that the combined classifier improves both performance and the explanation ability.

Algorithms↗

Use of an electronic nose to diagnose bacterial sinusitis.

BACKGROUND: Having previously established that an electronic nose (enose) can distinguish among bacteria samples, between cerebrospinal fluid leak and serum, and can identify patients with ventilator-associated pneumonia, we hypothesized that bacterial sinusitis could be diagnosed by sampling exhaled gas with an enose. METHODS: Using a nasal continuous positive airway pressure mask, we sampled gas exhaled through the nose of patients with sinusitis and compared them with controls. Data were first projected onto the principal components and then classified by support vector machine (SVM), a machine learning algorithm for pattern recognition. RESULTS: SVM analysis showed good discrimination using three approaches. First, 11 samples were used to create a training set that was used to predict whether individual samples from each set were a member of the control or infected sets. The enose was correct 98.4% of the time. Second, one-half of the samples from each of the same 11 control and infected groups were used to construct a training set, which was used to predict infection in the remaining samples. The enose was correct 82% of the time. Finally, 68 samples (34 positive and 34 controls) were analyzed using a leave-one-out scheme for creating training sets and testing sets. This method, designed to reflect the generalization property of the SVM classifier, scored a classification rate of 72%. CONCLUSION: Using the enose to sample nasal exhalation from patients with suspected sinusitis, we were able to predict correctly the diagnosis of sinusitis in at least 72% of the samples. The next step will be to do forward prediction using this model.

Algorithms↗

Predicting the toxicity of complex mixtures using artificial neural networks.

Industrial and municipal wastewaters constitute major sources of contamination of the aquatic compartment and represent a threat to aquatic life. Artificial neural networks based on three different learning paradigms were studied as a means of predicting acute toxicity to trout (5 days exposure to wastewaters) using input data from two simple microbiotests requiring only 5 or 15 min of incubation. These microbiotests were 1) the chemoluminescent peroxidase (Cl-Per) assay, which can detect radical scavengers and enzyme-inhibiting substances, and 2) the luminescent bacteria toxicity test (Microtox), in which reduction of light emission by bacteria during exposure is taken as a measure of toxicity. The responses obtained with the trout bioassay, the Cl-Per and the Microtox test were analyzed through statistical correlation (Pearson product-moment correlation), unsupervised learning by a self-organizing network, and assisted learning by the backpropagation and the Boltzmann machine (probabilistic) paradigms. No significant correlation (p < 0.05) was found between the responses obtained with either the Cl-Per assay (p = 0.121) or the Microtox (p = 0.061) microbiotest and those resulting from the trout bioassay. The self-organizing network was able to identify by itself a maximum of five classes that were more or less relevant for predicting toxicity to fish: class 1 contained 2 samples that were toxic to fish, class 2 contained 2/3 samples that were toxic, class 3 showed 6/8 samples that were non toxic, class 4 contained 5/6 samples that were non-toxic and class 5 comprised one sample that was toxic. Supervised learning with backpropagation analysis yielded two kinds of networks that hold potential. The first one was able to predict the actual toxic wastewater concentration with an overall performance of 65% when fed fresh data, while the second one, which was designed to differentiate between toxic and non-toxic effluents, exhibited a much better performance (90%). However, the probabilistic network also proved to be a very good predictive model for toxicity to fish, with an overall performance of 90%. Although more data are needed, the network based on the backpropagation paradigm seems to be a better predictor or classifier of trout toxicity when used with the Cl-Per and the Microtox microbiotests.

Algorithms↗

Symmetric and asymmetric multi-modality biclustering analysis for microarray data matrix.

Machine learning techniques offer a viable approach to cluster discovery from microarray data, which involves identifying and classifying biologically relevant groups in genes and conditions. It has been recognized that genes (whether or not they belong to the same gene group) may be co-expressed via a variety of pathways. Therefore, they can be adequately described by a diversity of coherence models. In fact, it is known that a gene may participate in multiple pathways that may or may not be co-active under all conditions. It is therefore biologically meaningful to simultaneously divide genes into functional groups and conditions into co-active categories--leading to the so-called biclustering analysis. For this, we have proposed a comprehensive set of coherence models to cope with various plausible regulation processes. Furthermore, a multivariate biclustering analysis based on fusion of different coherence models appears to be promising because the expression level of genes from the same group may follow more than one coherence models. The simulation studies further confirm that the proposed framework enjoys the advantage of high prediction performance.

Algorithms↗

Genetic algorithm learning as a robust approach to RNA editing site prediction.

BACKGROUND: RNA editing is one of several post-transcriptional modifications that may contribute to organismal complexity in the face of limited gene complement in a genome. One form, known as C --> U editing, appears to exist in a wide range of organisms, but most instances of this form of RNA editing have been discovered serendipitously. With the large amount of genomic and transcriptomic data now available, a computational analysis could provide a more rapid means of identifying novel sites of C --> U RNA editing. Previous efforts have had some success but also some limitations. We present a computational method for identifying C --> U RNA editing sites in genomic sequences that is both robust and generalizable. We evaluate its potential use on the best data set available for these purposes: C --> U editing sites in plant mitochondrial genomes. RESULTS: Our method is derived from a machine learning approach known as a genetic algorithm. REGAL (RNA Editing site prediction by Genetic Algorithm Learning) is 87% accurate when tested on three mitochondrial genomes, with an overall sensitivity of 82% and an overall specificity of 91%. REGAL's performance significantly improves on other ab initio approaches to predicting RNA editing sites in this data set. REGAL has a comparable sensitivity and higher specificity than approaches which rely on sequence homology, and it has the advantage that strong sequence conservation is not required for reliable prediction of edit sites. CONCLUSION: Our results suggest that ab initio methods can generate robust classifiers of putative edit sites, and we highlight the value of combinatorial approaches as embodied by genetic algorithms. We present REGAL as one approach with the potential to be generalized to other organisms exhibiting C --> U RNA editing.

Algorithms↗

A wiring of the human nucleolus.

Recent proteomic efforts have created an extensive inventory of the human nucleolar proteome. However, approximately 30% of the identified proteins lack functional annotation. We present an approach of assigning function to uncharacterized nucleolar proteins by data integration coupled to a machine-learning method. By assembling protein complexes, we present a first draft of the human ribosome biogenesis pathway encompassing 74 proteins and hereby assign function to 49 previously uncharacterized proteins. Moreover, the functional diversity of the nucleolus is underlined by the identification of a number of protein complexes with functions beyond ribosome biogenesis. Finally, we were able to obtain experimental evidence of nucleolar localization of 11 proteins, which were predicted by our platform to be associates of nucleolar complexes. We believe other biological organelles or systems could be "wired" in a similar fashion, integrating different types of data with high-throughput proteomics, followed by a detailed biological analysis and experimental validation.

Artificial Intelligence↗

Support vector machines for learning to identify the critical positions of a protein.

A method for identifying the positions in the amino acid sequence, which are critical for the catalytic activity of a protein using support vector machines (SVMs) is introduced and analysed. SVMs are supported by an efficient learning algorithm and can utilize some prior knowledge about the structure of the problem. The amino acid sequences of the variants of a protein, created by inducing mutations, along with their fitness are required as input data by the method to predict its critical positions. To investigate the performance of this algorithm, variants of the beta-lactamase enzyme were created in silico using simulations of both mutagenesis and recombination protocols. Results from literature on beta-lactamase were used to test the accuracy of this method. It was also compared with the results from a simple search algorithm. The algorithm was also shown to be able to predict critical positions that can tolerate two different amino acids and retain function.

Algorithms↗

Instance-based concept learning from multiclass DNA microarray data.

BACKGROUND: Various statistical and machine learning methods have been successfully applied to the classification of DNA microarray data. Simple instance-based classifiers such as nearest neighbor (NN) approaches perform remarkably well in comparison to more complex models, and are currently experiencing a renaissance in the analysis of data sets from biology and biotechnology. While binary classification of microarray data has been extensively investigated, studies involving multiclass data are rare. The question remains open whether there exists a significant difference in performance between NN approaches and more complex multiclass methods. Comparative studies in this field commonly assess different models based on their classification accuracy only; however, this approach lacks the rigor needed to draw reliable conclusions and is inadequate for testing the null hypothesis of equal performance. Comparing novel classification models to existing approaches requires focusing on the significance of differences in performance. RESULTS: We investigated the performance of instance-based classifiers, including a NN classifier able to assign a degree of class membership to each sample. This model alleviates a major problem of conventional instance-based learners, namely the lack of confidence values for predictions. The model translates the distances to the nearest neighbors into 'confidence scores'; the higher the confidence score, the closer is the considered instance to a pre-defined class. We applied the models to three real gene expression data sets and compared them with state-of-the-art methods for classifying microarray data of multiple classes, assessing performance using a statistical significance test that took into account the data resampling strategy. Simple NN classifiers performed as well as, or significantly better than, their more intricate competitors. CONCLUSION: Given its highly intuitive underlying principles--simplicity, ease-of-use, and robustness--the k-NN classifier complemented by a suitable distance-weighting regime constitutes an excellent alternative to more complex models for multiclass microarray data sets. Instance-based classifiers using weighted distances are not limited to microarray data sets, but are likely to perform competitively in classifications of high-dimensional biological data sets such as those generated by high-throughput mass spectrometry.

Algorithms↗

bioETH-PRS: confidential polygenic risk scoring with smart contracts on an FHE-enabled blockchain.

Polygenic risk scores (PRSs) aggregate genetic effect estimates to predict disease susceptibility, yet calculating one through an external service can require exposing raw genotype data. Homomorphic encryption hides those data during the calculation but, in prior work, still places a designated evaluator in a position of trust. We present bioETH-PRS, a protocol that replaces the evaluator with publicly auditable smart contracts on a blockchain supporting Fully Homomorphic Ethereum Virtual Machine (fhEVM). Using integer-exact encrypted arithmetic, bioETH-PRS computes the PRS dot product entirely in the encrypted domain, so genotype dosages and, at the model provider's discretion, the GWAS weights stay hidden from the parties performing the computation. A fixed-point encoding represents signed weights as nonnegative integers within a bound that rules out overflow, recovering the score to the precision of the published weights. A four-contract architecture separates data custody, model publication, computation, and output release, and supports both a classic path that stores encrypted inputs and an appreciably cheaper streaming path that discards them. A release oracle can return a randomized risk category instead of the raw score, limiting what a repeated querier learns. Prototype evaluation on real GWAS fixtures, including a run on a public testnet, shows cost growing linearly with variant count and suggests the approach may be practical where transaction fees are low. Trust is redistributed rather than removed: the system still depends on the contracts, the blockchain, and the fhEVM services. We evaluate additive models of moderate size, not genome-wide or clinical use.

Blockchain↗

Logistic-based patient grouping for multi-disciplinary treatment.

Present-day healthcare witnesses a growing demand for coordination of patient care. Coordination is needed especially in those cases in which hospitals have structured healthcare into specialty-oriented units, while a substantial portion of patient care is not limited to single units. From a logistic point of view, this multi-disciplinary patient care creates a tension between controlling the hospital's units, and the need for a control of the patient flow between units. A possible solution is the creation of new units in which different specialties work together for specific groups of patients. A first step in this solution is to identify the salient patient groups in need of multi-disciplinary care. Grouping techniques seem to offer a solution. However, most grouping approaches in medicine are driven by a search for pathophysiological homogeneity. In this paper, we present an alternative logistic-driven grouping approach. The starting point of our approach is a database with medical cases for 3,603 patients with peripheral arterial vascular (PAV) diseases. For these medical cases, six basic logistic variables (such as the number of visits to different specialist) are selected. Using these logistic variables, clustering techniques are used to group the medical cases in logistically homogeneous groups. In our approach, the quality of the resulting grouping is not measured by statistical significance, but by (i) the usefulness of the grouping for the creation of new multi-disciplinary units; (ii) how well patients can be selected for treatment in the new units. Given a priori knowledge of a patient (e.g. age, diagnosis), machine learning techniques are employed to induce rules that can be used for the selection of the patients eligible for treatment in the new units. In the paper, we describe the results of the above-proposed methodology for patients with PAV diseases. Two groupings and the accompanied classification rule sets are presented. One grouping is based on all the logistic variables, and another grouping is based on two latent factors found by applying factor analysis. On the basis of the experimental results, we can conclude that it is possible to search for medical logistic homogenous groups (i) that can be characterized by rules based on the aggregated logistic variables; (ii) for which we can formulate rules to predict to which cluster new patients belong.

Databases, Factual↗

Learning yeast gene functions from heterogeneous sources of data using hybrid weighted Bayesian networks.

We developed a machine learning system for determining gene functions from heterogeneous sources of data sets using a Weighted Naive Bayesian Network (WNB). The knowledge of gene functions is crucial for understanding many fundamental biological mechanisms such as regulatory pathways, cell cycles and diseases. Our major goal is to accurately infer functions of putative genes or ORFs (Open Reading Frames) from existing databases using computational methods. However, this task is intrinsically difficult since the underlying biological processes represent complex interactions of multiple entities. Therefore many functional links would be missing when only one or two source of data is used in the prediction. Our hypothesis is that integrating evidence from multiple and complementary sources could significantly improve the prediction accuracy. In this paper, our experimental results not only suggest that the above hypothesis is valid, but also provide guidelines for using the WNB system for data collection, training and predictions. The combined training data sets contain information from gene annotations, gene expressions, clustering outputs, keyword annotations and sequence homology from public databases. The current system is trained and tested on the genes of budding yeast Saccharomyces cerevisiae. Our WNB model can also be used to analyze the contribution of each source of information toward the prediction performance through the weight training process. The contribution analysis could potentially lead to significant scientific discovery by facilitating the interpretation and understanding of the complex relationships between biological entities.

Artificial Intelligence↗

EPIC: Event Prototyping via Information Constrained graph learning for personalized cancer driver gene prediction.

MOTIVATION: Precision oncology relies on accurately distinguishing patient-specific driver mutations from the vast background of passenger alterations. While graph-based computational methods have emerged as powerful tools for this task, they often struggle to preserve the distinct genomic context of individual mutations within complex biological networks. Consequently, subtle patient-specific driver signals are frequently obscured by dominant topological patterns, critically impeding the identification of individualized oncogenic events essential for personalized cancer therapy. RESULTS: To address this, we propose EPIC, a novel framework for Event Prototyping via Information Constrained Graph Learning. Unlike traditional node-centric approaches, EPIC redefines driver prediction as a metric learning task in an event embedding space. We introduce an information-constrained learning strategy that imposes explicit geometric constraints on feature variance, effectively preventing feature collapse and ensuring that low-frequency driver signals are distinctively preserved. Experiments on large-scale cancer cohorts demonstrate that EPIC significantly outperforms established baselines. Notably, the model prioritizes low-frequency driver variants typically overlooked by population-based methods, mapping them to critical oncogenic mechanisms associated with drug resistance and metastasis. Furthermore, clinical actionability analysis confirms that EPIC substantially expands the patient population eligible for targeted therapies. EPIC provides a robust and context-aware solution for personalized cancer driver discovery, bridging the gap between genomic data and actionable therapeutic insights. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/spcho-dev/EPIC.

Humans↗

Prediction of protein solvent accessibility using support vector machines.

A Support Vector Machine learning system has been trained to predict protein solvent accessibility from the primary structure. Different kernel functions and sliding window sizes have been explored to find how they affect the prediction performance. Using a cut-off threshold of 15% that splits the dataset evenly (an equal number of exposed and buried residues), this method was able to achieve a prediction accuracy of 70.1% for single sequence input and 73.9% for multiple alignment sequence input, respectively. The prediction of three and more states of solvent accessibility was also studied and compared with other methods. The prediction accuracies are better than, or comparable to, those obtained by other methods such as neural networks, Bayesian classification, multiple linear regression, and information theory. In addition, our results further suggest that this system may be combined with other prediction methods to achieve more reliable results, and that the Support Vector Machine method is a very useful tool for biological sequence analysis.

Bayes Theorem↗

Large-scale mapping and validation of Escherichia coli transcriptional regulation from a compendium of expression profiles.

Machine learning approaches offer the potential to systematically identify transcriptional regulatory interactions from a compendium of microarray expression profiles. However, experimental validation of the performance of these methods at the genome scale has remained elusive. Here we assess the global performance of four existing classes of inference algorithms using 445 Escherichia coli Affymetrix arrays and 3,216 known E. coli regulatory interactions from RegulonDB. We also developed and applied the context likelihood of relatedness (CLR) algorithm, a novel extension of the relevance networks class of algorithms. CLR demonstrates an average precision gain of 36% relative to the next-best performing algorithm. At a 60% true positive rate, CLR identifies 1,079 regulatory interactions, of which 338 were in the previously known network and 741 were novel predictions. We tested the predicted interactions for three transcription factors with chromatin immunoprecipitation, confirming 21 novel interactions and verifying our RegulonDB-based performance estimates. CLR also identified a regulatory link providing central metabolic control of iron transport, which we confirmed with real-time quantitative PCR. The compendium of expression data compiled in this study, coupled with RegulonDB, provides a valuable model system for further improvement of network inference algorithms using experimental data.

Algorithms↗

Protein solubility: sequence based prediction and experimental verification.

MOTIVATION: Obtaining soluble proteins in sufficient concentrations is a recurring limiting factor in various experimental studies. Solubility is an individual trait of proteins which, under a given set of experimental conditions, is determined by their amino acid sequence. Accurate theoretical prediction of solubility from sequence is instrumental for setting priorities on targets in large-scale proteomics projects. RESULTS: We present a machine-learning approach called PROSO to assess the chance of a protein to be soluble upon heterologous expression in Escherichia coli based on its amino acid composition. The classification algorithm is organized as a two-layered structure in which the output of primary support vector machine (SVM) classifiers serves as input for a secondary Naive Bayes classifier. Experimental progress information from the TargetDB database as well as previously published datasets were used as the source of training data. In comparison with previously published methods our classification algorithm possesses improved discriminatory capacity characterized by the Matthews Correlation Coefficient (MCC) of 0.434 between predicted and known solubility states and the overall prediction accuracy of 72% (75 and 68% for positive and negative class, respectively). We also provide experimental verification of our predictions using solubility measurements for 31 mutational variants of two different proteins.

Amino Acid Sequence↗