Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

A coordinated interplay: proteins with multiple functions in DNA replication, DNA repair, cell cycle/checkpoint control, and transcription.

In eukaryotic cells, DNA transactions such as replication, repair, and transcription require a large set of proteins. In all of these events, complexes of more than 30 polypetides appear to function in highly organized and structurally well-defined machines. We have learned in the past few years that the three essential macromolecular events, replication, repair, and transcription, have common functional entities and are coordinated by complex regulatory mechanisms. This can be documented for replication and repair, for replication and checkpoint control, and for replication and cell cycle control, as well as for replication and transcription. In this review we cover the three different protein classes: DNA polymerases, DNA polymerase accessory proteins, and selected transcription factors. The "common enzyme-different pathway strategy" is fascinating from several points of view: first, it might guarantee that these events are coordinated; second, it can be viewed from an evolutionary angle; and third, this strategy might provide cells with backup mechanisms for essential physiological tasks.

Cell Cycle↗

A multi-scale fusion model based on multi-phase contrast-enhanced CT for predicting pancreatic cancer resectability.

Purpose.Develop a multi-scale fusion model (MSFM) based on multi-phase contrast-enhanced computed tomography (CECT) to predict pancreatic cancer (PC) resectability, thereby assisting expert decision-making.Methods.This retrospective study enrolled 280 patients with PC from four institutions, which were randomly divided into a training cohort (202 patients) and an independent test cohort (78 patients). Three-phase CECT images (arterial, venous, and delayed phases) were used for modeling. The MSFM comprises two sub-networks: (1) a multi-phase fusion network for extracting cross-phase shared fusion features, (2) a phase-specific branch network for capturing phase-specific features; and a post-fusion strategy to generate the final predictive score by integrating the shared fusion features and three groups of phase-specific features. Additionally, a human-machine fusion deep learning model (HMfDL) was constructed by fusing the predictive score of the MSFM with expert assessments.Results.In the independent test, the MSFM achieved an AUC (area under the receiver operating characteristic curve) of 0.8385 (95% CI: 0.7521-0.9249), accuracy of 84.62%, sensitivity of 72.00%, and specificity of 90.57%. This performance outperformed single-phase models (AUC range: 0.7638-0.7781), two-phase models (AUC range: 0.7826-0.7864), and ten states-of-the-art classifiers (AUC range: 0.7404-0.7796). The HMfDL further improved the performance, reaching an AUC of 0.8626 (95% CI: 0.7853-0.9400), accuracy of 91.03%, sensitivity of 80.00%, and specificity of 96.23%. Notably, the HMfDL corrected 58.82% of misdiagnosis made by experts.Conclusions. The MSFM effectively fuses multi-phase CECT to enable highly accurate predictions of PC resectability, and provides valuable support for expert decision-making through HMfDL.

Humans↗

Profiler: an open web platform for multi-omics analysis.

MOTIVATION: High-throughput multi-omics technologies produce increasingly large and heterogeneous datasets that are difficult to analyze without advanced computational expertise. Existing bioinformatics tools are often fragmented or limited to specific omics types, hindering reproducibility and accessibility. There is a critical need for an integrated, user-friendly, and scalable platform capable of supporting multi-omics analyses across different data modalities. RESULTS: We present Profiler, an open-source, modular platform that unifies data import, quality control, preprocessing, statistical testing, machine and deep learning, biomarker discovery, pathway and drug-target enrichment, and survival modeling within a single reproducible environment. Built in Python with Streamlit, Profiler is available as both a web-based platform deployed on high-performance computing and a desktop version for local execution, enabling flexible usage across computational infrastructures. Profiler supports diverse omics modalities, including proteomics, transcriptomics, lipidomics, and electroencephalogram data. Through applications to glioblastoma proteomic, pancancer, and multi-omics datasets, Profiler reproduced known molecular subtypes, revealed potential therapeutic targets, and generated fully traceable analysis reports within minutes. By integrating advanced analytics behind an intuitive interface, Profiler democratizes multi-omics analysis and provides a robust, scalable foundation for systems biology and precision medicine research. AVAILABILITY AND IMPLEMENTATION: Profiler is open-source and freely available via its web platform (https://prism-profiler.univ-lille.fr) and GitHub (web version: https://github.com/yanisZirem/Profiler_v1_requests_datatests, desktop version: https://github.com/yanisZirem/prism-profiler), and archived on Zenodo (DOI: https://doi.org/10.5281/zenodo.17478158).

Software↗

A sensitive, support-vector-machine method for the detection of horizontal gene transfers in viral, archaeal and bacterial genomes.

In earlier work, we introduced and discussed a generalized computational framework for identifying horizontal transfers. This framework relied on a gene's nucleotide composition, obviated the need for knowledge of codon boundaries and database searches, and was shown to perform very well across a wide range of archaeal and bacterial genomes when compared with previously published approaches, such as Codon Adaptation Index and C + G content. Nonetheless, two considerations remained outstanding: we wanted to further increase the sensitivity of detecting horizontal transfers and also to be able to apply the method to increasingly smaller genomes. In the discussion that follows, we present such a method, Wn-SVM, and show that it exhibits a very significant improvement in sensitivity compared with earlier approaches. Wn-SVM uses a one-class support-vector machine and can learn using rather small training sets. This property makes Wn-SVM particularly suitable for studying small-size genomes, similar to those of viruses, as well as the typically larger archaeal and bacterial genomes. We show experimentally that the new method results in a superior performance across a wide range of organisms and that it improves even upon our own earlier method by an average of 10% across all examined genomes. As a small-genome case study, we analyze the genome of the human cytomegalovirus and demonstrate that Wn-SVM correctly identifies regions that are known to be conserved and prototypical of all beta-herpesvirinae, regions that are known to have been acquired horizontally from the human host and, finally, regions that had not up to now been suspected to be horizontally transferred. Atypical region predictions for many eukaryotic viruses, including the alpha-, beta- and gamma-herpesvirinae, and 123 archaeal and bacterial genomes, have been made available online at http://cbcsrv.watson.ibm.com/HGT_SVM/.

Artificial Intelligence↗

Integrating Imaging-Derived Clinical Endotypes with Plasma Proteomics and External Polygenic Risk Scores Enhances Coronary Microvascular Disease Risk Prediction.

Coronary microvascular disease (CMVD) is an underdiagnosed but significant contributor to the burden of ischemic heart disease, characterized by angina and myocardial infarction. The development of risk prediction models such as polygenic risk scores (PRS) for CMVD has been limited by a lack of large-scale genome-wide association studies (GWAS). However, there is significant overlap between CMVD and enrollment criteria for coronary artery disease (CAD) GWAS. In this study, we developed CMVD PRS models by selecting variants identified in a CMVD GWAS and applying weights from an external CAD GWAS, using CMVD-associated loci as proxies for the genetic risk. We integrated plasma proteomics, clinical measures from perfusion PET imaging, and PRS to evaluate their contributions to CMVD risk prediction in comprehensive machine and deep learning models. We then developed a novel unsupervised endotyping framework for CMVD from perfusion PET-derived myocardial blood flow data, revealing distinct patient subgroups beyond traditional case-control definitions. This imaging-based stratification substantially improved classification performance alongside plasma proteomics and PRS, achieving AUROCs between 0.65 and 0.73 per class, significantly outperforming binary classifiers and existing clinical models, highlighting the potential of this stratification approach to enable more precise and personalized diagnosis by capturing the underlying heterogeneity of CMVD. This work represents the first application of imaging-based endotyping and the integration of genetic and proteomic data for CMVD risk prediction, establishing a framework for multimodal modeling in complex diseases.

Cardiovascular Disease↗

AI-Driven Precision Medicine in Alzheimer's Disease: Drug Repurposing, Digital Therapeutics and Clinical Decision Support.

Alzheimer's Disease (AD) is a neurodegenerative disease that causes significant clinical, social, and economic burden worldwide. Despite improvements in understanding its multifaceted pathogenesis, current treatments are mostly symptomatic and ineffective across varied patient populations. To overcome these constraints, AI-driven precision medicine allows tailored risk assessment, treatment selection, and disease monitoring. This review covers AI's role in AD precision medicine, focusing on drug repurposing, digital therapies and clinical decision support systems. Machine and deep learning models are used to predict medication response, integrate heterogeneous data sources such as genomics, transcriptomics, neuroimaging and electronic health records, and uncover pharmacogenomic treatment success factors. The paper covers AIenabled precision pharmacology, including tailored dosing algorithms, adaptive therapeutic monitoring, and adverse drug reaction prediction. Bioinformatics-based target identification, network pharmacology, graphbased AI models, virtual screening, and real-world and clinical data validation are emphasized in AI-driven medication repurposing. AI-powered digital treatments like personalized cognitive training platforms, wearable- derived digital biomarkers, virtual and mixed reality interventions, adherence monitoring, and digital twins for therapy optimization have been discussed. AI-based clinical decision support systems are also thoroughly assessed for clinical value, accuracy, and explainability in disease subtyping, trajectory prediction, and risk stratification in preclinical and prodromal AD. Despite these promises, data heterogeneity, algorithmic bias, legal barriers, and privacy concerns exist. Federated learning enables safe multi-center collaboration and hybrid AI-human approaches, and it represents the future. AI's ability to alter AD care opens the door to precision medicine paradigms that use repurposed medications, digital tools and intelligent decision-making to improve patient outcomes.

Alzheimer’s disease↗

Technology-based vs. traditional instruction. A comparison of two methods for teaching the skill of performing a 12-lead ECG.

The purpose of this study was to compare the effectiveness of an interactive, multimedia CD-ROM with traditional methods of teaching the skill of performing a 12-lead ECG. A randomized pre/posttest experimental design was used. Seventy-seven baccalaureate nursing students in a required, senior-level critical-care course at a large midwestern university were recruited for the study. Two teaching methods were compared. The traditional method included a self-study module, a brief lecture and demonstration by an instructor, and hands-on experience using a plastic manikin and a real 12-lead ECG machine in the learning laboratory. The second method covered the same content using an interactive, multimedia CD-ROM embedded with virtual reality and supplemented with a self-study module. There were no significant (p < .05) baseline differences in pretest scores between the two groups and no significant differences by group in cognitive gains, student satisfaction with their learning method, or perception of self-efficacy in performing the skill. Overall results indicated that both groups were satisfied with their instructional method and were similar in their ability to demonstrate the skill correctly on a live, simulated patient. This evaluation study is a beginning step to assess new and potentially more cost-effective teaching methods and their effects on student learning outcomes and behaviors, including the transfer of skill acquisition via a computer simulation to a real patient.

Adult↗

A support vector machine approach to the identification of phosphorylation sites.

We describe a bioinformatics tool that can be used to predict the position of phosphorylation sites in proteins based only on sequence information. The method uses the support vector machine (SVM) statistical learning theory. The statistical models for phosphorylation by various types of kinases are built using a dataset of short (9-amino acid long) sequence fragments. The sequence segments are dissected around post-translationally modified sites of proteins that are on the current release of the Swiss-Prot database, and that were experimentally confirmed to be phosphorylated by any kinase. We represent them as vectors in a multidimensional abstract space of short sequence fragments. The prediction method is as follows. First, a given query protein sequence is dissected into overlapping short segments. All the fragments are then projected into the multidimensional space of sequence fragments via a collection of different representations. Those points are classified with pre-built statistical models (the SVM method with linear, polynomial and radial kernel functions) either as phosphorylated or inactive ones. The resulting list of plausible sites for phosphorylation by various types of kinases in the query protein is returned to the user. The efficiency of the method for each type of phosphorylation is estimated using leave-one-out tests and presented here. The sensitivities of the models can reach over 70%, depending on the type of kinase. The additional information from profile representations of short sequence fragments helps in gaining a higher degree of accuracy in some phosphorylation types. The further development of an automatic phosphorylation site annotation predictor based on our algorithm should yield a significant improvement when using statistical algorithms in order to quantify the results.

Algorithms↗

A probabilistic active support vector learning algorithm.

The paper describes a probabilistic active learning strategy for support vector machine (SVM) design in large data applications. The learning strategy is motivated by the statistical query model. While most existing methods of active SVM learning query for points based on their proximity to the current separating hyperplane, the proposed method queries for a set of points according to a distribution as determined by the current separating hyperplane and a newly defined concept of an adaptive confidence factor. This enables the algorithm to have more robust and efficient learning capabilities. The confidence factor is estimated from local information using the k nearest neighbor principle. The effectiveness of the method is demonstrated on real-life data sets both in terms of generalization performance, query complexity, and training time.

Algorithms↗

Learning to generate combinatorial action sequences utilizing the initial sensitivity of deterministic dynamical systems.

This study shows how sensory-action sequences of imitating finite state machines (FSMs) can be learned by utilizing the deterministic dynamics of recurrent neural networks (RNNs). Our experiments indicated that each possible combinatorial sequence can be recalled by specifying its respective initial state value and also that fractal structures appear in this initial state mapping after the learning converges. We also observed that the sequences of mimicking FSMs are encoded utilizing the transient regions rather than the invariant sets of the evolved dynamical systems of the RNNs.

Computer Simulation↗

The learning curve for a colonoscopy simulator in the absence of any feedback: no feedback, no learning.

BACKGROUND: The hypothesis of this study is that working on the simulator without a structured feedback does not change performance; hence, any effects shown after structured feedback would amount to useful learning of the procedure. The aim was to investigate the learning curve for the HT Immersion Medical Colonoscopy Simulator without any structured feedback. This could then be potentially applied to validate the learning curve on the simulator when structured feedback is provided. There are no previous studies on this matter. METHODS: Candidates were asked to perform colonoscopy on the HT Immersion Medical Colonoscopy Simulator. Modules 3 and 4 were used at random. In total, each candidate was asked to perform five consecutive virtual colonoscopies on the same module. These five episodes were collectively referred to as one trial. A time result of 3,600 sec (1 h) was used to denote perforation. No guidance or feedback was given to candidates before, during, or after each procedure. A total of 26 postgraduate doctors were recruited, including nine research fellows, five preregistration house officers, six specialist registrars, and six consultants. Fourteen candidates recorded five attempts each (i.e., one trial each) on the same module of the colonoscopy simulator (14 trials over 70 episodes). Another 12 candidates recorded five attempts (i.e., one trial each) on two modules of the colonoscopy simulator (24 trials over 120 episodes). Hence, 190 episodes were recorded in total, representing 38 trials. RESULTS: There was no improvement in performance on the simulator from first attempt to the fifth in the absence of feedback. If there was any initial gain in any measurable outcome, this was lost in subsequent attempts indicating lack of learning. The outcomes measured included time taken to complete the test, percentage of the mucosa visualized, depth of the instrument inserted, and the path length used. The results were statistically significant for all outcomes. CONCLUSIONS: This study demonstrates that in the absence of feedback, it is not possible to improve performance on the HT Immersion Medical Colonoscopy Simulator. Thus, there is no learning curve for the machine. The information from this study is vital for using the simulators in training and assessment because any improvement in learning curves shown after training on simulators can be presumed to be due to learning the procedure and not the simulator.

Clinical Competence↗

Effect of training datasets on support vector machine prediction of protein-protein interactions.

Knowledge of protein-protein interaction is useful for elucidating protein function via the concept of 'guilt-by-association'. A statistical learning method, Support Vector Machine (SVM), has recently been explored for the prediction of protein-protein interactions using artificial shuffled sequences as hypothetical noninteracting proteins and it has shown promising results (Bock, J. R., Gough, D. A., Bioinformatics 2001, 17, 455-460). It remains unclear however, how the prediction accuracy is affected if real protein sequences are used to represent noninteracting proteins. In this work, this effect is assessed by comparison of the results derived from the use of real protein sequences with that derived from the use of shuffled sequences. The real protein sequences of hypothetical noninteracting proteins are generated from an exclusion analysis in combination with subcellular localization information of interacting proteins found in the Database of Interacting Proteins. Prediction accuracy using real protein sequences is 76.9% compared to 94.1% using artificial shuffled sequences. The discrepancy likely arises from the expected higher level of difficulty for separating two sets of real protein sequences than that for separating a set of real protein sequences from a set of artificial sequences. The use of real protein sequences for training a SVM classification system is expected to give better prediction results in practical cases. This is tested by using both SVM systems for predicting putative protein partners of a set of thioredoxin related proteins. The prediction results are consistent with observations, suggesting that real sequence is more practically useful in development of SVM classification system for facilitating protein-protein interaction prediction.

Algorithms↗

Qualified predictions for microarray and proteomics pattern diagnostics with confidence machines.

We focus on the problem of prediction with confidence and describe a recently developed learning algorithm called transductive confidence machine for making qualified region predictions. Its main advantage, in comparison with other classifiers, is that it is well-calibrated, with number of prediction errors strictly controlled by a given predefined confidence level. We apply the transductive confidence machine to the problems of acute leukaemia and ovarian cancer prediction using microarray and proteomics pattern diagnostics, respectively. We demonstrate that the algorithm performs well, yielding well-calibrated and informative predictions whilst maintaining a high level of accuracy.

Algorithms↗

Enzyme family classification by support vector machines.

One approach for facilitating protein function prediction is to classify proteins into functional families. Recent studies on the classification of G-protein coupled receptors and other proteins suggest that a statistical learning method, Support vector machines (SVM), may be potentially useful for protein classification into functional families. In this work, SVM is applied and tested on the classification of enzymes into functional families defined by the Enzyme Nomenclature Committee of IUBMB. SVM classification system for each family is trained from representative enzymes of that family and seed proteins of Pfam curated protein families. The classification accuracy for enzymes from 46 families and for non-enzymes is in the range of 50.0% to 95.7% and 79.0% to 100% respectively. The corresponding Matthews correlation coefficient is in the range of 54.1% to 96.1%. Moreover, 80.3% of the 8,291 correctly classified enzymes are uniquely classified into a specific enzyme family by using a scoring function, indicating that SVM may have certain level of unique prediction capability. Testing results also suggest that SVM in some cases is capable of classification of distantly related enzymes and homologous enzymes of different functions. Effort is being made to use a more comprehensive set of enzymes as training sets and to incorporate multi-class SVM classification systems to further enhance the unique prediction accuracy. Our results suggest the potential of SVM for enzyme family classification and for facilitating protein function prediction. Our software is accessible at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi.

Amino Acid Sequence↗

Support vector machines (SVMs) for monitoring network design.

In this paper we present a hydrologic application of a new statistical learning methodology called support vector machines (SVMs). SVMs are based on minimization of a bound on the generalized error (risk) model, rather than just the mean square error over a training set. Due to Mercer's conditions on the kernels, the corresponding optimization problems are convex and hence have no local minima. In this paper, SVMs are illustratively used to reproduce the behavior of Monte Carlo-based flow and transport models that are in turn used in the design of a ground water contamination detection monitoring system. The traditional approach, which is based on solving transient transport equations for each new configuration of a conductivity field, is too time consuming in practical applications. Thus, there is a need to capture the behavior of the transport phenomenon in random media in a relatively simple manner. The objective of the exercise is to maximize the probability of detecting contaminants that exceed some regulatory standard before they reach a compliance boundary, while minimizing cost (i.e., number of monitoring wells). Application of the method at a generic site showed a rather promising performance, which leads us to believe that SVMs could be successfully employed in other areas of hydrology. The SVM was trained using 510 monitoring configuration samples generated from 200 Monte Carlo flow and transport realizations. The best configurations of well networks selected by the SVM were identical with the ones obtained from the physical model, but the reliabilities provided by the respective networks differ slightly.

Artificial Intelligence↗

On the statistical assessment of classifiers using DNA microarray data.

BACKGROUND: In this paper we present a method for the statistical assessment of cancer predictors which make use of gene expression profiles. The methodology is applied to a new data set of microarray gene expression data collected in Casa Sollievo della Sofferenza Hospital, Foggia--Italy. The data set is made up of normal (22) and tumor (25) specimens extracted from 25 patients affected by colon cancer. We propose to give answers to some questions which are relevant for the automatic diagnosis of cancer such as: Is the size of the available data set sufficient to build accurate classifiers? What is the statistical significance of the associated error rates? In what ways can accuracy be considered dependant on the adopted classification scheme? How many genes are correlated with the pathology and how many are sufficient for an accurate colon cancer classification? The method we propose answers these questions whilst avoiding the potential pitfalls hidden in the analysis and interpretation of microarray data. RESULTS: We estimate the generalization error, evaluated through the Leave-K-Out Cross Validation error, for three different classification schemes by varying the number of training examples and the number of the genes used. The statistical significance of the error rate is measured by using a permutation test. We provide a statistical analysis in terms of the frequencies of the genes involved in the classification. Using the whole set of genes, we found that the Weighted Voting Algorithm (WVA) classifier learns the distinction between normal and tumor specimens with 25 training examples, providing e = 21% (p = 0.045) as an error rate. This remains constant even when the number of examples increases. Moreover, Regularized Least Squares (RLS) and Support Vector Machines (SVM) classifiers can learn with only 15 training examples, with an error rate of e = 19% (p = 0.035) and e = 18% (p = 0.037) respectively. Moreover, the error rate decreases as the training set size increases, reaching its best performances with 35 training examples. In this case, RLS and SVM have error rates of e = 14% (p = 0.027) and e = 11% (p = 0.019). Concerning the number of genes, we found about 6000 genes (p < 0.05) correlated with the pathology, resulting from the signal-to-noise statistic. Moreover the performances of RLS and SVM classifiers do not change when 74% of genes is used. They progressively reduce up to e = 16% (p < 0.05) when only 2 genes are employed. The biological relevance of a set of genes determined by our statistical analysis and the major roles they play in colorectal tumorigenesis is discussed. CONCLUSIONS: The method proposed provides statistically significant answers to precise questions relevant for the diagnosis and prognosis of cancer. We found that, with as few as 15 examples, it is possible to train statistically significant classifiers for colon cancer diagnosis. As for the definition of the number of genes sufficient for a reliable classification of colon cancer, our results suggest that it depends on the accuracy required.

Aged↗