Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

PharaCon: a new framework for identifying bacteriophages via conditional representation learning.

MOTIVATION: Identifying bacteriophages (phages) within metagenomic sequences is essential for understanding microbial community dynamics. Transformer-based foundation models have been successfully employed to address various biological challenges. However, these models are typically pre-trained with self-supervised tasks that do not consider label variance in the pre-training data. This presents a challenge for phage identification as pre-training on mixed bacterial and phage data may lead to information bias due to the imbalance between bacterial and phage samples. RESULTS: To overcome this limitation, we proposed a novel conditional BERT framework that incorporates label classes as special tokens during pre-training. Specifically, our conditional BERT model attaches labels directly during tokenization, introducing label constraints into the model's input. Additionally, we introduced a new fine-tuning scheme that enables the conditional BERT to be effectively utilized for classification tasks. This framework allows the BERT model to acquire label-specific contextual representations from mixed sequence data during pre-training and applies the conditional BERT as a classifier during fine-tuning, and we named the fine-tuned model as PharaCon. We evaluated PharaCon against several existing methods on both simulated sequence datasets and real metagenomic contig datasets. The results demonstrate PharaCon's effectiveness and efficiency in phage identification, highlighting the advantages of incorporating label information during both pre-training and fine-tuning. AVAILABILITY AND IMPLEMENTATION: The source code and associated data can be accessed at https://github.com/Celestial-Bai/PharaCon.

Bacteriophages↗

A bi-recursive neural network architecture for the prediction of protein coarse contact maps.

Prediction of contact maps may be seen as a strategic step towards the solution of fundamental open problems in structural genomics. In this paper we focus on coarse grained maps that describe the spatial neighborhood relation between secondary structure elements (helices, strands, and coils) of a protein. We introduce a new machine learning approach for scoring candidate contact maps. The method combines a specialized noncausal recursive connectionist architecture and a heuristic graph search algorithm. The network is trained using candidate graphs generated during search. We show how the process of selecting and generating training examples is important for tuning the precision of the predictor.

Algorithms↗

EPIC: Event Prototyping via Information Constrained graph learning for personalized cancer driver gene prediction.

MOTIVATION: Precision oncology relies on accurately distinguishing patient-specific driver mutations from the vast background of passenger alterations. While graph-based computational methods have emerged as powerful tools for this task, they often struggle to preserve the distinct genomic context of individual mutations within complex biological networks. Consequently, subtle patient-specific driver signals are frequently obscured by dominant topological patterns, critically impeding the identification of individualized oncogenic events essential for personalized cancer therapy. RESULTS: To address this, we propose EPIC, a novel framework for Event Prototyping via Information Constrained Graph Learning. Unlike traditional node-centric approaches, EPIC redefines driver prediction as a metric learning task in an event embedding space. We introduce an information-constrained learning strategy that imposes explicit geometric constraints on feature variance, effectively preventing feature collapse and ensuring that low-frequency driver signals are distinctively preserved. Experiments on large-scale cancer cohorts demonstrate that EPIC significantly outperforms established baselines. Notably, the model prioritizes low-frequency driver variants typically overlooked by population-based methods, mapping them to critical oncogenic mechanisms associated with drug resistance and metastasis. Furthermore, clinical actionability analysis confirms that EPIC substantially expands the patient population eligible for targeted therapies. EPIC provides a robust and context-aware solution for personalized cancer driver discovery, bridging the gap between genomic data and actionable therapeutic insights. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/spcho-dev/EPIC.

Humans↗

Parameter selection for and implementation of a web-based decision-support tool to predict extubation outcome in premature infants.

BACKGROUND: Approximately 30% of intubated preterm infants with respiratory distress syndrome (RDS) will fail attempted extubation, requiring reintubation and mechanical ventilation. Although ventilator technology and monitoring of premature infants have improved over time, optimal extubation remains challenging. Furthermore, extubation decisions for premature infants require complex informational processing, techniques implicitly learned through clinical practice. Computer-aided decision-support tools would benefit inexperienced clinicians, especially during peak neonatal intensive care unit (NICU) census. METHODS: A five-step procedure was developed to identify predictive variables. Clinical expert (CE) thought processes comprised one model. Variables from that model were used to develop two mathematical models for the decision-support tool: an artificial neural network (ANN) and a multivariate logistic regression model (MLR). The ranking of the variables in the three models was compared using the Wilcoxon Signed Rank Test. The best performing model was used in a web-based decision-support tool with a user interface implemented in Hypertext Markup Language (HTML) and the mathematical model employing the ANN. RESULTS: CEs identified 51 potentially predictive variables for extubation decisions for an infant on mechanical ventilation. Comparisons of the three models showed a significant difference between the ANN and the CE (p = 0.0006). Of the original 51 potentially predictive variables, the 13 most predictive variables were used to develop an ANN as a web-based decision-tool. The ANN processes user-provided data and returns the prediction 0-1 score and a novelty index. The user then selects the most appropriate threshold for categorizing the prediction as a success or failure. Furthermore, the novelty index, indicating the similarity of the test case to the training case, allows the user to assess the confidence level of the prediction with regard to how much the new data differ from the data originally used for the development of the prediction tool. CONCLUSION: State-of-the-art, machine-learning methods can be employed for the development of sophisticated tools to aid clinicians' decisions. We identified numerous variables considered relevant for extubation decisions for mechanically ventilated premature infants with RDS. We then developed a web-based decision-support tool for clinicians which can be made widely available and potentially improve patient care world wide.

Birth Weight↗

Building an ontology of adverse drug reactions for automated signal generation in pharmacovigilance.

Automated signal generation in pharmacovigilance implements unsupervised statistical machine learning techniques in order to discover unknown adverse drug reactions (ADR) in spontaneous reporting systems. The impact of the terminology used for coding ADRs has not been addressed previously. The Medical Dictionary for Regulatory Activities (MedDRA) used worldwide in pharmacovigilance cases does not provide formal definitions of terms. We have built an ontology of ADRs to describe semantics of MedDRA terms. Ontological subsumption and approximate matching inferences allow a better grouping of medically related conditions. Signal generation performances are significantly improved but time consumption related to modelization remains very important.

Adverse Drug Reaction Reporting Systems↗

The role of artificial intelligence in the diagnosis and prognosis of traumatic brain injury based on brain CT scans: a systematic review.

Traumatic brain injury (TBI) is a leading cause of emergency department visits and a major contributor to injury-related mortality and long-term neurological disability. Non-contrast computed tomography (CT) is the gold-standard imaging modality for the rapid diagnosis of TBI. Clinical outcomes depend strongly on early detection and prompt acute management. Artificial intelligence (AI)-based models may support faster automated identification of traumatic findings and early prediction of patient prognosis. A systematic literature search was conducted in PubMed/MEDLINE, Scopus, IEEE Xplore, ACM Digital Library, and the Cochrane Library in accordance with PRISMA 2020 guidelines to evaluate AI-based models for automated detection of TBI-related findings on CT and for prediction of clinical outcomes. Risk of bias and applicability were assessed using QUADAS-2 for diagnostic accuracy studies and PROBAST + AI for prediction model studies. Twenty-two studies were included. Sixteen studies evaluated diagnostic tasks and 10 evaluated prognostic outcomes, with four studies contributing to both categories. Diagnostic performance was generally high, with many studies reporting AUC values approaching or exceeding 0.90, particularly for larger lesion volumes.Prognostic performance was more variable, with moderate to high discrimination and substantial heterogeneity. Only 9 studies incorporated independent external validation, and performance was frequently lower in external cohorts. All prognostic model studies were judged to be at high overall risk of bias using PROBAST + AI, and most diagnostic accuracy studies also demonstrated high or unclear risk of bias in at least one QUADAS-2 domain, most frequently in patient selection. AI-based models applied to brain CT demonstrate strong technical performance for both diagnostic and prognostic tasks in TBI. However, most studies relied on retrospective designs and lacked independent external validation which limits models generalizability and raises concern for potential overfitting. Prospective, multicenter studies with standardized methodologies and rigorous external validation are required before widespread clinical implementation.

Humans↗

Prediction of transmembrane regions of beta-barrel proteins using ANN- and SVM-based methods.

This article describes a method developed for predicting transmembrane beta-barrel regions in membrane proteins using machine learning techniques: artificial neural network (ANN) and support vector machine (SVM). The ANN used in this study is a feed-forward neural network with a standard back-propagation training algorithm. The accuracy of the ANN-based method improved significantly, from 70.4% to 80.5%, when evolutionary information was added to a single sequence as a multiple sequence alignment obtained from PSI-BLAST. We have also developed an SVM-based method using a primary sequence as input and achieved an accuracy of 77.4%. The SVM model was modified by adding 36 physicochemical parameters to the amino acid sequence information. Finally, ANN- and SVM-based methods were combined to utilize the full potential of both techniques. The accuracy and Matthews correlation coefficient (MCC) value of SVM, ANN, and combined method are 78.5%, 80.5%, and 81.8%, and 0.55, 0.63, and 0.64, respectively. These methods were trained and tested on a nonredundant data set of 16 proteins, and performance was evaluated using "leave one out cross-validation" (LOOCV). Based on this study, we have developed a Web server, TBBPred, for predicting transmembrane beta-barrel regions in proteins (available at http://www.imtech.res.in/raghava/tbbpred).

Algorithms↗

Prediction of caspase cleavage sites using Bayesian bio-basis function neural networks.

MOTIVATION: Apoptosis has drawn the attention of researchers because of its importance in treating some diseases through finding a proper way to block or slow down the apoptosis process. Having understood that caspase cleavage is the key to apoptosis, we find novel methods or algorithms are essential for studying the specificity of caspase cleavage activity and this helps the effective drug design. As bio-basis function neural networks have proven to outperform some conventional neural learning algorithms, there is a motivation, in this study, to investigate the application of bio-basis function neural networks for the prediction of caspase cleavage sites. RESULTS: Thirteen protein sequences with experimentally determined caspase cleavage sites were downloaded from NCBI. Bayesian bio-basis function neural networks are investigated and the comparisons with single-layer perceptrons, multilayer perceptrons, the original bio-basis function neural networks and support vector machines are given. The impact of the sliding window size used to generate sub-sequences for modelling on prediction accuracy is studied. The results show that the Bayesian bio-basis function neural network with two Gaussian distributions for model parameters (weights) performed the best and the highest prediction accuracy is 97.15 +/- 1.13%. AVAILABILITY: The package of Bayesian bio-basis function neural network can be obtained by request to the author.

Algorithms↗

CSGL: chemical synthesis graph learning for molecule representation.

MOTIVATION: Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. RESULTS: Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-2023/CSGL.

Machine Learning↗

De novo Genes in Plants: Origins, Mechanisms, and Functional Implications.

De novo genes originate from previously non-coding genomic regions. They provide an important source of lineage-specific innovation. In plants, these genes may contribute to adaptation, trait diversity and crop evolution. This review summarizes recent progress in plant de novo gene research. It first discusses major routes of gene birth, including transcription-first, open reading frame (ORF)-first and concurrent models. It also examines how nascent loci acquire regulatory control and enter existing biological networks. The review then summarizes their evolutionary features, including weak early constraint, rapid molecular change, restricted expression and structural refinement. It further discusses plant de novo genes involved in stress responses, seed germination, kernel dehydration, subspecies divergence, reproductive isolation and floral scent diversification. Current methods for identifying de novo genes remain limited by rapid sequence evolution, genome annotation quality, polyploidy and transposable elements. Whole-genome synteny alignment, multi-omics evidence and machine-learning approaches can improve candidate discovery. However, each method has important limitations. Finally, this review highlights key future questions in functional validation, latent coding potential in long non-coding RNAs, epigenetic activation, regulatory-network integration and crop improvement. These perspectives clarify how de novo genes shape plant adaptation and how they may be used in precision breeding and synthetic biology.

adaptive evolution↗

Links between PPCA and subspace methods for complete Gaussian density estimation.

High-dimensional density estimation is a fundamental problem in pattern recognition and machine learning areas. In this letter, we show that, for complete high-dimensional Gaussian density estimation, two widely used methods, probabilistic principal component analysis and a typical subspace method using eigenspace decomposition, actually give the same results. Additionally, we present a unified view from the aspect of robust estimation of the covariance matrix.

Algorithms↗

Large-scale proteomics profiling of peripheral blood of DM1 patients identifies biomarkers for disease severity and functional capacity.

BackgroundMyotonic Dystrophy Type 1 (DM1), the most common genetic neuromuscular disorder in adults, poses significant challenges for drug development due to its multisystem nature and high clinical variability in symptoms and disease progression. With a growing number of therapies entering clinical trials, this study addresses the urgent need for biomarkers that can serve as surrogate endpoints.MethodsWe profiled 437 serum samples from adult DM1 patients collected at two timepoints of the OPTIMISTIC trial using bottom-up mass spectrometry with data-independent acquisition. Associations between protein expression, the disease-causing CTG-repeat and 25 clinical outcome measures were studied using linear mixed-effect models. All key study findings were validated in an independent cohort of 69 DM1 patients and 10 healthy controls.ResultsOf the 259 identified proteins, 161 showed significant associations with the CTG-repeat length (FDR&#x2009;<&#x2009;5%). Hypogammaglobulinemia was confirmed and shown to be worse in severely affected patients. A strong proteomic signature was associated with clinical measures of functional capacity, with the 6-Minute Walk Test showing the strongest signal (70 associations, FDR&#x2009;<&#x2009;5%). These novel associations reveal a compelling link between chronic inflammation and reduced functional capacity. A machine learning algorithm identified a minimal set of 13 proteins robustly reflecting both the underlying genetic defect and functional capacity.ConclusionsDM1 induces a broad disease fingerprint in the serum proteome, predominantly affecting proteins of the immune system. A carefully selected panel of proteins showed the greatest potential to meet the statistical criteria required for surrogate endpoints in clinical trials.

Humans↗

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33&#x202f;have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(&#x3b5;)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational↗

Handling missing values in support vector machine classifiers.

This paper discusses the task of learning a classifier from observed data containing missing values amongst the inputs which are missing completely at random. A non-parametric perspective is adopted by defining a modified risk taking into account the uncertainty of the predicted outputs when missing values are involved. It is shown that this approach generalizes the approach of mean imputation in the linear case and the resulting kernel machine reduces to the standard Support Vector Machine (SVM) when no input values are missing. Furthermore, the method is extended to the multivariate case of fitting additive models using componentwise kernel machines, and an efficient implementation is based on the Least Squares Support Vector Machine (LS-SVM) classifier formulation.

Algorithms↗

Using unsupervised learning with independent component analysis to identify patterns of glaucomatous visual field defects.

PURPOSE: Clustering by unsupervised learning with machine learning classifiers was shown to segment clusters of patterns in standard automated perimetry (SAP) for glaucoma in previous publications. In this study, unsupervised learning by independent component analysis decomposed SAP field patterns into axes, and the information represented by these axes was evaluated. METHODS: SAP fields were used that were obtained with the Humphrey Visual Field Analyzer (Carl Zeiss Meditec, Dublin, CA) from 189 normal eyes and 156 eyes with glaucomatous optic neuropathy (GON) determined by masked review with stereoscopic optic disc photographs. The variational Bayesian independent component analysis mixture model (vB-ICA-mm) partitioned the SAP fields into the most informative number of clusters. Simultaneously, the model learned an optimal number of maximally independent axes for each cluster. RESULTS: The most informative number of clusters in the SAP set was two. vB-ICA-mm placed 68.6% of the eyes with GON in a cluster labeled G and 98.4% of the eyes with normal optic discs in a cluster labeled N. Cluster G optimally contained six axes. Post hoc analysis of patterns generated at -1 SD and +2 SD from the cluster G mean on the six axes revealed defects similar to those identified by experts as indicative of glaucoma. SAP fields associated with an axis showed increasing severity, as they were located farther in the positive direction from the cluster G mean. CONCLUSIONS: vB-ICA-mm represented the SAP fields with patterns that were meaningful for glaucoma experts. This process also captured severity in the patterns uncovered. These findings should validate vB-ICA-mm as a data-mining technique for new and unfamiliar complex tests.

Artificial Intelligence↗

Classification and knowledge discovery in protein databases.

We consider the problem of classification in noisy, high-dimensional, and class-imbalanced protein datasets. In order to design a complete classification system, we use a three-stage machine learning framework consisting of a feature selection stage, a method addressing noise and class-imbalance, and a method for combining biologically related tasks through a prior-knowledge based clustering. In the first stage, we employ Fisher's permutation test as a feature selection filter. Comparisons with the alternative criteria show that it may be favorable for typical protein datasets. In the second stage, noise and class imbalance are addressed by using minority class over-sampling, majority class under-sampling, and ensemble learning. The performance of logistic regression models, decision trees, and neural networks is systematically evaluated. The experimental results show that in many cases ensembles of logistic regression classifiers may outperform more expressive models due to their robustness to noise and low sample density in a high-dimensional feature space. However, ensembles of neural networks may be the best solution for large datasets. In the third stage, we use prior knowledge to partition unlabeled data such that the class distributions among non-overlapping clusters significantly differ. In our experiments, training classifiers specialized to the class distributions of each cluster resulted in a further decrease in classification error.

Algorithms↗

Role of the cerebellum in the visual guidance of movement.

Mathematicians, control engineers and information technologists are beginning to take a greater interest in neuroscience. They are perhaps starting to realize that they may be able to learn a few tricks from nature with which to improve their machines. At the same time there is a good chance that neuroscientists will benefit from their input of fresh ideas and techniques with which to attack the problems of understanding neural processing. One area of the brain which seems particularly promising in these respects is the cerebellum.

Brain Mapping↗

Fedflow: cloud orchestration for federated learning with the FeatureCloud platform.

MOTIVATION: Federated learning (FL) enables collaborative model training on geographically distributed genomic and clinical datasets while complying with data privacy laws and regulatory constraints. FeatureCloud is an existing platform for FL that provides an accessible web-based interface and a large repository of implemented methods. However, due to its graphical interface, FeatureCloud requires manual interaction of all participants, limiting automation, iteration, and reproducibility. RESULTS: We introduce fedflow, a Python-based command-line tool for headless orchestration of FL tasks with FeatureCloud. This tool uses distributed computing resources such as virtual machines or cloud instances to automate such workflows. This allows for scalable federated computing either in local simulations or deployed in a trusted environment. Further, we demonstrate how fedflow can be used to integrate FeatureCloud in reproducible Snakemake workflows. For this, we reanalyse a metagenomic dataset with two federated algorithms and compare the results to the centralized approach with pooled data. Overall, fedflow enables automation of multi-client FL tasks, facilitates embedding of FeatureCloud in standard bioinformatics pipelines and thereby helps increase reproducibility. AVAILABILITY: Fedflow is open-source and available at https://github.com/W-L/fedflow.

Journal Article↗