Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Molecular biology and integrated strategies for activating cryptic biosynthetic gene clusters toward next-generation antibiotic discovery.

Antimicrobial resistance (AMR) has been identified as one of the 21st century's severest global public health crises. AMR led to an estimated 4.95 million deaths in 2019 and will claim 10 million lives a year by 2050 in the absence of targeted interventions. During the same period, the number of novel antibiotics discovered has decreased drastically as many researchers are rediscovering known antibiotics, non-model microorganisms are poorly understood or difficult to culture and antibiotic research and development investment has declined drastically. However, high-throughput whole genome sequencing and the subsequent application of bioinformatics in bacterial and fungal genomes have shown that a numerous of cryptic or silent biosynthetic gene clusters (BGCs) remain latent at ambient laboratory conditions since their genes are transcriptionally inactive. Cryptic BGCs represent a vast source of unique secondary metabolites, many of which may yield novel antibacterial, antifungal, anti-cancer and other potentially valuable natural products. This review discusses the biological relevance of cryptic BGCs, the major limiting factors that restricts their activation and novel strategies that have been employed to activate them and exploit their potential to produce novel natural products. The review focuses on biological approaches including CRISPR-Cas mediation for the activation of cryptic BGCs, promoter engineering, pathway refactoring, and heterologous expression; biochemical strategies such as Osman, OsMAC, Precursor Feeding, Chemical Elicitation, Epigenetic Regulation and Co-cultivation and technology-based strategies such as Genome mining, Microfluidic Cultivation systems, High-Throughput Screening, Metabolomics, Molecular Networking and Artificial Intelligence and Machine Learning based prediction of BGCs and their metabolites. The use of multi-omics technologies combined with synthetic biology to achieve better discovery, characterization and large-scale production of novel natural products is also discussed herein. Finally, we will talk about the ecological significance and evolutionary advantage of cryptic BGCs' role in interactions between microorganisms, such as competition, communication, symbiosis and environmental adaptability, so as to provide a useful background for accelerating next-generation antibiotics.

CRISPR-Cas activation↗

Machine learning and multi-omics clustering to map cellular rewiring and immune evasion in ccRCC.

Immune checkpoint blockade (ICB) efficacy in clear cell renal cell carcinoma (ccRCC) is limited by tumor microenvironment (TME) heterogeneity. Because traditional bulk-derived models lack spatial resolution, we developed an integrated framework connecting macroscopic survival risks to microscopic TME structures. We applied ten algorithms to establish multi-omics subtypes and evaluated 101 machine-learning combinations across three independent cohorts to generate a Consensus Machine Learning-driven Signature (CMLS). The signature's spatial and cellular origins were decoded using spatial transcriptomics (ST) and a 140,000-cell scRNA-seq atlas. Expression of key genes was experimentally validated via RT-qPCR in 17 paired ccRCC clinical tissues. We identified two molecular subtypes with distinct clinical and epigenetic profiles. SuperPC optimization yielded a 24-gene CMLS serving as an independent prognostic factor. scRNA-seq and ST deconvolution revealed these signals predominantly originate from cancer-associated fibroblasts (CAFs) and malignant epithelial cells, which collaborate to drive spatial immune exclusion. RT-qPCR confirmed significant overexpression of five core CMLS genes in ccRCC versus adjacent normal tissues. Low CMLS scores correlated with enhanced ICB responsiveness, whereas high-CMLS tumors demonstrated specific vulnerability to dasatinib and dabrafenib. The CMLS translates spatial immune-exclusion dynamics into a quantifiable metric, outperforming tumor mutational burden in predicting ICB benefits, providing a robust tool for patient stratification in ccRCC.

Humans↗

Gene Specific Pathogenicity Predictor for Chromatin-Remodeling BAF Complex-Associated Neurodevelopmental Disorders.

Advancements in whole genome sequencing have increased the number of variants of uncertain significance (VUS) identified in patient genomes. This has created a diagnostic bottleneck for genetic counselors tasked with sifting through these variants and determining those most likely to be causative for a patient's clinical presentation. Machine learning (ML) tools can aid in identifying pathogenic variants from VUS, but there is a need for gene-specific algorithms that predict pathogenic variants with high accuracy. To address this need, we present a workflow for developing gene-specific, ensemble-learning ML tools, that leverage outputs from other algorithms, locations of variants within the gene, and evolutionary conservation data to make a prediction of pathogenicity. Variants in SMARCA2 and SMARCA4 that are associated with rare neurodevelopmental diseases were used to screen 15 ML algorithms. A random forest learner was tuned to yield a final accuracy of 0.93 on holdout data. Generalizing this predictor to other BAF complex proteins resulted in a sharp decline in performance. We trained a final predictor for all genes in the study to create a predictor that identifies pathogenic variants in these BAF subunits with an accuracy of 0.91 on holdout data. This predictor specific to BAF complex proteins performs with higher accuracy and AUROC than any other predictor. The decline in performance when generalized to other proteins emphasizes the need for the gene-specific calibration of predictors. Our workflow for the development of such models provides a quick, computationally inexpensive route for improving the ML tools available to genetic counselors.

Journal Article↗

A novel method of protein secondary structure prediction with high segment overlap measure: support vector machine approach.

We have introduced a new method of protein secondary structure prediction which is based on the theory of support vector machine (SVM). SVM represents a new approach to supervised pattern classification which has been successfully applied to a wide range of pattern recognition problems, including object recognition, speaker identification, gene function prediction with microarray expression profile, etc. In these cases, the performance of SVM either matches or is significantly better than that of traditional machine learning approaches, including neural networks.The first use of the SVM approach to predict protein secondary structure is described here. Unlike the previous studies, we first constructed several binary classifiers, then assembled a tertiary classifier for three secondary structure states (helix, sheet and coil) based on these binary classifiers. The SVM method achieved a good performance of segment overlap accuracy SOV=76.2 % through sevenfold cross validation on a database of 513 non-homologous protein chains with multiple sequence alignments, which out-performs existing methods. Meanwhile three-state overall per-residue accuracy Q(3) achieved 73.5 %, which is at least comparable to existing single prediction methods. Furthermore a useful "reliability index" for the predictions was developed. In addition, SVM has many attractive features, including effective avoidance of overfitting, the ability to handle large feature spaces, information condensing of the given data set, etc. The SVM method is conveniently applied to many other pattern classification tasks in biology.

Computer Simulation↗

Learning to lead at Toyota.

Many companies have tried to copy Toyota's famous production system--but without success. Why? Part of the reason, says the author, is that imitators fail to recognize the underlying principles of the Toyota Production System (TPS), focusing instead on specific tools and practices. This article tells the other part of the story. Building on a previous HBR article, "Decoding the DNA of the Toyota Production System," Spear explains how Toyota inculcates managers with TPS principles. He describes the training of a star recruit--a talented young American destined for a high-level position at one of Toyota's U.S. plants. Rich in detail, the story offers four basic lessons for any company wishing to train its managers to apply Toyota's system: There's no substitute for direct observation. Toyota employees are encouraged to observe failures as they occur--for example, by sitting next to a machine on the assembly line and waiting and watching for any problems. Proposed changes should always be structured as experiments. Employees embed explicit and testable assumptions in the analysis of their work. That allows them to examine the gaps between predicted and actual results. Workers and managers should experiment as frequently as possible. The company teaches employees at all levels to achieve continuous improvement through quick, simple experiments rather than through lengthy, complex ones. Managers should coach, not fix. Toyota managers act as enablers, directing employees but not telling them where to find opportunities for improvements. Rather than undergo a brief period of cursory walk-throughs, orientations, and introductions as incoming fast-track executives at most companies might, the executive in this story learned TPS the long, hard way--by practicing it, which is how Toyota trains any new employee, regardless of rank or function.

Administrative Personnel↗

Prediction of torsade-causing potential of drugs by support vector machine approach.

In an effort to facilitate drug discovery, computational methods for facilitating the prediction of various adverse drug reactions (ADRs) have been developed. So far, attention has not been sufficiently paid to the development of methods for the prediction of serious ADRs that occur less frequently. Some of these ADRs, such as torsade de pointes (TdP), are important issues in the approval of drugs for certain diseases. Thus there is a need to develop tools for facilitating the prediction of these ADRs. This work explores the use of a statistical learning method, support vector machine (SVM), for TdP prediction. TdP involves multiple mechanisms and SVM is a method suitable for such a problem. Our SVM classification system used a set of linear solvation energy relationship (LSER) descriptors and was optimized by leave-one-out cross validation procedure. Its prediction accuracy was evaluated by using an independent set of agents and by comparison with results obtained from other commonly used classification methods using the same dataset and optimization procedure. The accuracies for the SVM prediction of TdP-causing agents and non-TdP-causing agents are 97.4 and 84.6% respectively; one is substantially improved against and the other is comparable to the results obtained by other classification methods useful for multiple-mechanism prediction problems. This indicates the potential of SVM in facilitating the prediction of TdP-causing risk of small molecules and perhaps other ADRs that involve multiple mechanisms.

Algorithms↗

Exploiting the past and the future in protein secondary structure prediction.

MOTIVATION: Predicting the secondary structure of a protein (alpha-helix, beta-sheet, coil) is an important step towards elucidating its three-dimensional structure, as well as its function. Presently, the best predictors are based on machine learning approaches, in particular neural network architectures with a fixed, and relatively short, input window of amino acids, centered at the prediction site. Although a fixed small window avoids overfitting problems, it does not permit capturing variable long-rang information. RESULTS: We introduce a family of novel architectures which can learn to make predictions based on variable ranges of dependencies. These architectures extend recurrent neural networks, introducing non-causal bidirectional dynamics to capture both upstream and downstream information. The prediction algorithm is completed by the use of mixtures of estimators that leverage evolutionary information, expressed in terms of multiple alignments, both at the input and output levels. While our system currently achieves an overall performance close to 76% correct prediction--at least comparable to the best existing systems--the main emphasis here is on the development of new algorithmic ideas. AVAILABILITY: The executable program for predicting protein secondary structure is available from the authors free of charge. CONTACT: pfbaldi@ics.uci.edu, gpollast@ics.uci.edu, brunak@cbs.dtu.dk, paolo@dsi.unifi.it.

Algorithms↗

Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.

BACKGROUND: There has been an explosion in the number of single nucleotide polymorphisms (SNPs) within public databases. In this study we focused on non-synonymous protein coding single nucleotide polymorphisms (nsSNPs), some associated with disease and others which are thought to be neutral. We describe the distribution of both types of nsSNPs using structural and sequence based features and assess the relative value of these attributes as predictors of function using machine learning methods. We also address the common problem of balance within machine learning methods and show the effect of imbalance on nsSNP function prediction. We show that nsSNP function prediction can be significantly improved by 100% undersampling of the majority class. The learnt rules were then applied to make predictions of function on all nsSNPs within Ensembl. RESULTS: The measure of prediction success is greatly affected by the level of imbalance in the training dataset. We found the balanced dataset that included all attributes produced the best prediction. The performance as measured by the Matthews correlation coefficient (MCC) varied between 0.49 and 0.25 depending on the imbalance. As previously observed, the degree of sequence conservation at the nsSNP position is the single most useful attribute. In addition to conservation, structural predictions made using a balanced dataset can be of value. CONCLUSION: The predictions for all nsSNPs within Ensembl, based on a balanced dataset using all attributes, are available as a DAS annotation. Instructions for adding the track to Ensembl are at http://www.brightstudy.ac.uk/das_help.html.

Algorithms↗

Using automatically learnt verb selectional preferences for classification of biomedical terms.

In this paper, we present an approach to term classification based on verb selectional patterns (VSPs), where such a pattern is defined as a set of semantic classes that could be used in combination with a given domain-specific verb. VSPs have been automatically learnt based on the information found in a corpus and an ontology in the biomedical domain. Prior to the learning phase, the corpus is terminologically processed: term recognition is performed by both looking up the dictionary of terms listed in the ontology and applying the C/NC-value method for on-the-fly term extraction. Subsequently, domain-specific verbs are automatically identified in the corpus based on the frequency of occurrence and the frequency of their co-occurrence with terms. VSPs are then learnt automatically for these verbs. Two machine learning approaches are presented. The first approach has been implemented as an iterative generalisation procedure based on a partial order relation induced by the domain-specific ontology. The second approach exploits the idea of genetic algorithms. Once the VSPs are acquired, they can be used to classify newly recognised terms co-occurring with domain-specific verbs. Given a term, the most frequently co-occurring domain-specific verb is selected. Its VSP is used to constrain the search space by focusing on potential classes of the given term. A nearest-neighbour approach is then applied to select a class from the constrained space of candidate classes. The most similar candidate class is predicted for the given term. The similarity measure used for this purpose combines contextual, lexical, and syntactic properties of terms.

Abstracting and Indexing↗

Predicting functional family of novel enzymes irrespective of sequence similarity: a statistical learning approach.

The function of a protein that has no sequence homolog of known function is difficult to assign on the basis of sequence similarity. The same problem may arise for homologous proteins of different functions if one is newly discovered and the other is the only known protein of similar sequence. It is desirable to explore methods that are not based on sequence similarity. One approach is to assign functional family of a protein to provide useful hint about its function. Several groups have employed a statistical learning method, support vector machines (SVMs), for predicting protein functional family directly from sequence irrespective of sequence similarity. These studies showed that SVM prediction accuracy is at a level useful for functional family assignment. But its capability for assignment of distantly related proteins and homologous proteins of different functions has not been critically and adequately assessed. Here SVM is tested for functional family assignment of two groups of enzymes. One consists of 50 enzymes that have no homolog of known function from PSI-BLAST search of protein databases. The other contains eight pairs of homologous enzymes of different families. SVM correctly assigns 72% of the enzymes in the first group and 62% of the enzyme pairs in the second group, suggesting that it is potentially useful for facilitating functional study of novel proteins. A web version of our software, SVMProt, is accessible at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi.

Artificial Intelligence↗

Reinforcement learning-based dynamic ensemble for missense variant effect prediction and tiered prioritization of VUS.

BACKGROUND: Accurate classification of missense variants remains a challenging task despite major advances in genomics. Numerous computational models have been developed to assist in variant classification, but often require repeated integration and benchmarking efforts. Ensemble methods have been proposed to overcome the limitations of single predictors, but mostly rely on fixed, predefined weights that constrain their ability to capture interactions among predictive signals. METHODS: We present GenixRL, a dynamic ensemble framework that reformulates model fusion as a reinforcement learning optimization problem. GenixRL uses a Q-learning agent to learn a policy that dynamically weights the probabilistic outputs of complementary predictors, including BayesDel (addAF and noAF), ClinPred, and MetaRNN. Replacing static weighting with policy learning allows GenixRL to adaptively identify optimal weightings and substantially improve classification accuracy. RESULTS: In benchmark evaluation against 25 state-of-the-art predictors, GenixRL achieved an AUROC of 0.9644 on an independent ClinVar dataset. On saturation genome editing assays for BRCA1 and BRCA2, GenixRL achieved the best performance and ranked highest on 14 of 17 clinically significant genes in a zero-shot evaluation. Applied to uncertain and conflicting ClinVar variants, GenixRL enabled tiered, evidence-based prioritization of hundreds of thousands of variants as likely pathogenic or pathogenic with high confidence, supported by orthogonal population evidence from gnomAD. CONCLUSION: GenixRL advances pathogenicity prediction for missense variants and provides an adaptive ensemble that sorts variants of uncertain significance into tiered candidates for expert curation and functional validation.

Mutation, Missense↗

Improved protein secondary structure prediction using support vector machine with a new encoding scheme and an advanced tertiary classifier.

Prediction of protein secondary structures is an important problem in bioinformatics and has many applications. The recent trend of secondary structure prediction studies is mostly based on the neural network or the support vector machine (SVM). The SVM method is a comparatively new learning system which has mostly been used in pattern recognition problems. In this study, SVM is used as a machine learning tool for the prediction of secondary structure and several encoding schemes, including orthogonal matrix, hydrophobicity matrix, BLOSUM62 substitution matrix, and combined matrix of these, are applied and optimized to improve the prediction accuracy. Also, the optimal window length for six SVM binary classifiers is established by testing different window sizes and our new encoding scheme is tested based on this optimal window size via sevenfold cross validation tests. The results show 2% increase in the accuracy of the binary classifiers when compared with the instances in which the classical orthogonal matrix is used. Finally, to combine the results of the six SVM binary classifiers, a new tertiary classifier which combines the results of one-versus-one binary classifiers is introduced and the performance is compared with those of existing tertiary classifiers. According to the results, the Q3 prediction accuracy of new tertiary classifier reaches 78.8% and this is better than the best result reported in the literature.

Algorithms↗

A machine learning information retrieval approach to protein fold recognition.

MOTIVATION: Recognizing proteins that have similar tertiary structure is the key step of template-based protein structure prediction methods. Traditionally, a variety of alignment methods are used to identify similar folds, based on sequence similarity and sequence-structure compatibility. Although these methods are complementary, their integration has not been thoroughly exploited. Statistical machine learning methods provide tools for integrating multiple features, but so far these methods have been used primarily for protein and fold classification, rather than addressing the retrieval problem of fold recognition-finding a proper template for a given query protein. RESULTS: Here we present a two-stage machine learning, information retrieval, approach to fold recognition. First, we use alignment methods to derive pairwise similarity features for query-template protein pairs. We also use global profile-profile alignments in combination with predicted secondary structure, relative solvent accessibility, contact map and beta-strand pairing to extract pairwise structural compatibility features. Second, we apply support vector machines to these features to predict the structural relevance (i.e. in the same fold or not) of the query-template pairs. For each query, the continuous relevance scores are used to rank the templates. The FOLDpro approach is modular, scalable and effective. Compared with 11 other fold recognition methods, FOLDpro yields the best results in almost all standard categories on a comprehensive benchmark dataset. Using predictions of the top-ranked template, the sensitivity is approximately 85, 56, and 27% at the family, superfamily and fold levels respectively. Using the 5 top-ranked templates, the sensitivity increases to 90, 70, and 48%.

Algorithms↗

Computer prediction of allergen proteins from sequence-derived protein structural and physicochemical properties.

BACKGROUND: Computational methods have been developed for predicting allergen proteins from sequence segments that show identity, homology, or motif match to a known allergen. These methods achieve good prediction accuracies, but are less effective for novel proteins with no similarity to any known allergen. METHODS: This work tests the feasibility of using a statistical learning method, support vector machines, as such a method. The prediction system is trained and tested by using 1005 allergen proteins from the Allergome database and 22,469 non-allergen proteins from 7871 Pfam families. RESULTS: Testing results by an independent set of 229 allergen and 6717 non-allergen proteins from 7871 Pfam families show that 93.0% and 99.9% of these are correctly predicted, which are comparable to the best results of other methods. Of the 18 novel allergen proteins non-homologous to any other proteins in the Swissprot database, 88.9% is correctly predicted. A further screening of 168,128 proteins in the Swissprot database finds that 2.9% of the proteins are predicted as allergen proteins, which is consistent with the estimated numbers from motif-based methods. CONCLUSIONS: Our study suggests that SVM is a potentially useful method for predicting allergen proteins and it has certain capability for predicting novel allergen proteins. Our software can be accessed at .

Allergens↗

Proteomic profiling of bone for the estimation of post-mortem interval and post-mortem submersion interval: a systematic review.

Accurate estimation of the Post-Mortem Interval (PMI) and Post-Mortem Submersion Interval (PMSI) remains a persistent challenge in forensic science, especially when traditional morphological and entomological methods fail due to advanced decomposition or in aquatic environments. Proteomic profiling of bone tissues has recently emerged as a promising approach, leveraging the predictable degradation patterns of bone proteins to estimate time since death more reliably. This systematic review, conducted in accordance with PRISMA guidelines, analyzed 24 peer-reviewed studies focusing on the application of proteomic techniques to bone tissue for PMI and PMSI estimation. The included studies were evaluated based on sample type, analytical techniques used, identified biomarkers, environmental conditions assessed, and the overall reliability and reproducibility of the findings. The review found that specific bone proteins, particularly collagen, osteocalcin, fetuin-A, etc. exhibited consistent degradation patterns that correlated strongly with elapsed post-mortem time. Cortical bone was identified as a more stable and informative matrix compared to trabecular bone. Mass spectrometry, especially LC-MS/MS, emerged as the predominant analytical technique due to its high sensitivity and accuracy in detecting low-abundance proteins over extended PMIs and PMSIs. However, protein degradation rates were significantly influenced by environmental variables such as temperature, humidity, soil pH, and microbial activity. This review also emphasizes the transformative role of bone proteomics in advancing forensic science while identifying key gaps that must be addressed to achieve global standardization and practical implementation in diverse forensic contexts. The integration of proteomics with other emerging technologies, such as machine learning algorithms and computational modeling, may further enhance the precision of PMI and PMSI estimation in future applications.

Postmortem Changes↗

Decision tree based information integration for automated protein classification.

We propose a novel technique for automatically generating the SCOP classification of a protein structure with high accuracy. We achieve accurate classification by combining the decisions of multiple methods using the consensus of a committee (or an ensemble) classifier. Our technique, based on decision trees, is rooted in machine learning which shows that by judicially employing component classifiers, an ensemble classifier can be constructed to outperform its components. We use two sequence- and three structure-comparison tools as component classifiers. Given a protein structure and using the joint hypothesis, we first determine if the protein belongs to an existing category (family, superfamily, fold) in the SCOP hierarchy. For the proteins that are predicted as members of the existing categories, we compute their family-, superfamily-, and fold-level classifications using the consensus classifier. We show that we can significantly improve the classification accuracy compared to the individual component classifiers. In particular, we achieve error rates that are 3-12 times less than the individual classifiers' error rates at the family level, 1.5-4.5 times less at the superfamily level, and 1.1-2.4 times less at the fold level.

Algorithms↗

Understanding protein dispensability through machine-learning analysis of high-throughput data.

MOTIVATION: Protein dispensability is fundamental to the understanding of gene function and evolution. Recent advances in generating high-throughput data such as genomic sequence data, protein-protein interaction data, gene-expression data and growth-rate data of mutants allow us to investigate protein dispensability systematically at the genome scale. RESULTS: In our studies, protein dispensability is represented as a fitness score that is measured by the growth rate of gene-deletion mutants. By the analyses of high-throughput data in yeast Saccharomyces cerevisiae, we found that a protein's dispensability had significant correlations with its evolutionary rate and duplication rate, as well as its connectivity in protein-protein interaction network and gene-expression correlation network. Neural network and support vector machine were applied to predict protein dispensability through high-throughput data. Our studies shed some lights on global characteristics of protein dispensability and evolution. AVAILABILITY: The original datasets for protein dispensability analysis and prediction, together with related scripts, are available at http://digbio.missouri.edu/~ychen/ProDispen/ CONTACT: xudong@missouri.edu.

Artificial Intelligence↗

Organization of the state space of a simple recurrent network before and after training on recursive linguistic structures.

Recurrent neural networks are often employed in the cognitive science community to process symbol sequences that represent various natural language structures. The aim is to study possible neural mechanisms of language processing and aid in development of artificial language processing systems. We used data sets containing recursive linguistic structures and trained the Elman simple recurrent network (SRN) for the next-symbol prediction task. Concentrating on neuron activation clusters in the recurrent layer of SRN we investigate the network state space organization before and after training. Given a SRN and a training stream, we construct predictive models, called neural prediction machines, that directly employ the state space dynamics of the network. We demonstrate two important properties of representations of recursive symbol series in the SRN. First, the clusters of recurrent activations emerging before training are meaningful and correspond to Markov prediction contexts. We show that prediction states that naturally arise in the SRN initialized with small random weights approximately correspond to states of Variable Memory Length Markov Models (VLMM) based on individual symbols (i.e. words). Second, we demonstrate that during training, the SRN reorganizes its state space according to word categories and their grammatical subcategories, and the next-symbol prediction is again based on the VLMM strategy. However, after training, the prediction is based on word categories and their grammatical subcategories rather than individual words. Our conclusion holds for small depths of recursions that are comparable to human performances. The methods of SRN training and analysis of its state space introduced in this paper are of a general nature and can be used for investigation of processing of any other symbol time series by means of SRN.

Artificial Intelligence↗