Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Different classification techniques considering brain computer interface applications.

In this work the application of different machine learning techniques for classification of mental tasks from electroencephalograph (EEG) signals is investigated. The main application for this research is the improvement of brain computer interface (BCI) systems. For this purpose, Bayesian graphical network, neural network, Bayesian quadratic, Fisher linear and hidden Markov model classifiers are applied to two known EEG datasets in the BCI field. The Bayesian network classifier is used for the first time in this work for classification of EEG signals. The Bayesian network appeared to have a significant accuracy and more consistent classification compared to the other four methods. In addition to classical correct classification accuracy criteria, the mutual information is also used to compare the classification results with other BCI groups.

Algorithms↗

Target and biomarker exploration portal for drug discovery.

MOTIVATION: The discovery of novel drug targets and precision biomarkers remains a major challenge in drug development, with traditional differential expression analysis often overlooking key regulatory proteins. Here, we present a novel, web-based bioinformatics tool, the Target and Biomarker Exploration Portal (TBEP), designed to accelerate the drug discovery process by integrating large-scale biomedical data with network analysis techniques. RESULTS: TBEP harnesses machine-learning approaches to mine and combine multimodal datasets, including human genetics, functional genomics, and protein-protein interaction networks, to decode causal disease mechanisms and uncover novel therapeutic targets and precision biomarkers for specific phenotypes. A unique feature of the tool is its ability to process large-scale data in real-time, facilitated by an efficient cloud-based architecture. Additionally, the tool incorporates an integrated large language model (LLM), which assists researchers in exploring and interpreting complex biological relationships within the generated networks and multi-omics data using natural language (English). By offering an intuitive, interactive interface, the LLM enhances the exploration of biological insights, making it easier for scientists to derive actionable conclusions. This powerful integration of network analysis, multi-omics data, and LLM provides a robust framework for accelerating the identification of novel drug targets. AVAILABILITY AND IMPLEMENTATION: The tool is publicly available at https://tbep.missouri.edu. The source code, documentation and installation instructions are available at GitHub repository: https://github.com/mizzoudbl/tbep.

Drug Discovery↗

PCP: a program for supervised classification of gene expression profiles.

UNLABELLED: PCP (Pattern Classification Program) is an open-source machine learning program for supervised classification of patterns (vectors of measurements). The principal use of PCP in bioinformatics is design and evaluation of classifiers for use in clinical diagnostic tests based on measurements of gene expression. PCP implements leading pattern classification and gene selection algorithms and incorporates cross-validation estimation of classifier performance. Importantly, the implementation integrates gene selection and class prediction stages, which is vital for computing reliable performance estimates in small-sample scenarios. Additionally, the program includes automated and efficient model selection (optimization of parameters) for support vector machine (SVM) classifier. The distribution includes Linux and Windows/Cygwin binaries. The program can easily be ported to other platforms. AVAILABILITY: Free download at http://pcp.sourceforge.net

Algorithms↗

Prediction of estrogen receptor agonists and characterization of associated molecular descriptors by statistical learning methods.

Specific estrogen receptor (ER) agonists have been used for hormone replacement therapy, contraception, osteoporosis prevention, and prostate cancer treatment. Some ER agonists and partial-agonists induce cancer and endocrine function disruption. Methods for predicting ER agonists are useful for facilitating drug discovery and chemical safety evaluation. Structure-activity relationships and rule-based decision forest models have been derived for predicting ER binders at impressive accuracies of 87.1-97.6% for ER binders and 80.2-96.0% for ER non-binders. However, these are not designed for identifying ER agonists and they were developed from a subset of known ER binders. This work explored several statistical learning methods (support vector machines, k-nearest neighbor, probabilistic neural network and C4.5 decision tree) for predicting ER agonists from comprehensive set of known ER agonists and other compounds. The corresponding prediction systems were developed and tested by using 243 ER agonists and 463 ER non-agonists, respectively, which are significantly larger in number and structural diversity than those in previous studies. A feature selection method was used for selecting molecular descriptors responsible for distinguishing ER agonists from non-agonists, some of which are consistent with those used in other studies and the findings from X-ray crystallography data. The prediction accuracies of these methods are comparable to those of earlier studies despite the use of significantly more diverse range of compounds. SVM gives the best accuracy of 88.9% for ER agonists and 98.1% for non-agonists. Our study suggests that statistical learning methods such as SVM are potentially useful for facilitating the prediction of ER agonists and for characterizing the molecular descriptors associated with ER agonists.

Forecasting↗

Discovering hidden candidate plastic-degrading enzymes: Combined multi-omics and machine learning strategy.

Plastic pollution poses a major threat to the stability of natural ecosystems as well as human health. Microbial enzymes have long been considered a potential resource for targeted biodegradation but, except for a few successful cases, the discovery of efficient enzymes has proved challenging. Aiming to accelerate the process, we propose an approach combining metagenomics, metatranscriptomics and semi-supervised learning that selects promising plastic-degrading candidate enzymes from the proteome of relevant microorganisms. Tested on a dataset of over 10,000 microbial proteins, ranking models consistently prioritize known plastic-degrading enzymes, achieving an area under the cumulative distribution function curve above 0.96, with leave-one-family-out cross-validation indicating that performance is largely retained across protein families. As a case study, this work focuses on mixed microbial cultures exposed for extended periods to polyethylene, polyethylene terephthalate, and polyurethane substrates. The prevalent species after selective enrichment were functionally characterized, finding Rhodococcus aetherivorans as the most relevant species in two of the five cultures under investigation. Among the top-ranked proteins, several have high structural similarity with known enzymes despite not being identified by sequence similarity search. Moreover, according to metatranscriptomics results, several of these enzymes were found to be expressed at the same level or above that of annotated enzymes, suggesting that they may have functional relevance. Overall, this work highlights the potential of integrating multi-omics with data-driven methods for enzyme discovery and for accelerating the development of biotechnological solutions to plastic pollution.

Biodegradation, Environmental↗

PMGen: from peptide-MHC structure prediction to peptide generation.

MOTIVATION: Accurate structural modeling of peptide-major histocompatibility complex (pMHC) complexes is essential for structure-driven immunotherapy design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths, insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical insights from pMHC structures largely unexploited. RESULTS: We introduce peptide-MHC generator (PMGen), an integrated framework for structure prediction and structure-guided design of variable-length peptides across MHC Class I and II. PMGen enforces anchor constraints within AlphaFold2 through two complementary strategies, initial guess and template engineering, achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen outperforms all existing methods, yielding median peptide-core Cα RMSDs of 0.62 Å for MHC-I and 0.33 Å for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show that ProteinMPNN sampling on PMGen-predicted backbones yields higher affinity peptides while preserving the parental 3D conformation. Using PMGen to generate 63 817 high-confidence pMHC structures as training data, we further improve ProteinMPNN's peptide sequence recovery from 0.14 to 0.64 on a test set of 85 unseen MHC-I alleles, highlighting the value of accurate predicted structures for downstream machine learning tasks. AVAILABILITY AND IMPLEMENTATION: PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.

Peptides↗

FOSB is a key factor in the genetic link between inflammatory bowel disease and acute myocardial infarction: multiple bioinformatics analyses and validation.

BACKGROUND: Inflammatory Bowel Disease (IBD), which includes Crohn's disease and ulcerative colitis, is associated with an increased risk of Acute Myocardial Infarction (AMI). The genetic mechanisms underlying this link are not well understood. METHODS: We downloaded IBD and AMI-related microarray datasets from the NCBI Gene Expression Omnibus (GEO) database. Differentially expressed genes (DEGs) were identified and analyzed using enrichment analysis and Weighted Gene Co-expression Network Analysis (WGCNA). Machine learning techniques, including LASSO, random forest, and Boruta, were employed to screen for hub genes. These genes were validated through qRT-PCR and Western blotting. Single-cell sequencing was used to confirm findings. Additionally, potential therapeutic targets were identified using the Connectivity Map (CMap) database. RESULTS: Five key hub genes-THBD, FOSB, ADGPR3, IL1R2, and PLAUR-were identified as significantly involved in both IBD and AMI pathogenesis. A diagnostic model for AMI constructed using these hub genes demonstrated high predictive accuracy. Single-cell sequencing analysis and several potential drugs targeting these hub genes were identified, offering new therapeutic avenues. CONCLUSION: This study highlights the crucial role of FOSB and other hub genes in the comorbidity of IBD and AMI. The findings provide novel insights for early diagnosis and potential therapeutic strategies, emphasizing the importance of further investigation into these genetic links.

Humans↗

mmContext: an open framework for multimodal contrastive learning of omics and text data.

SUMMARY: Multimodal approaches are increasingly leveraged for integrating omics data with textual biological knowledge. Yet there is still no accessible, standardized framework that enables systematic comparison of omics representations with different text encoders within a unified workflow. We present mmContext, a lightweight and extensible multimodal embedding framework built on top of the open-source Sentence Transformers library. The software allows researchers to train or apply models that jointly embed omics and text data using any numeric representation stored in an AnnData.obsm layer and any text encoder available in Hugging Face. mmContext supports integration of diverse biological text sources and provides pipelines for training, evaluation, and data preparation. We train and evaluate models for a RNA-Seq and text integration task, and demonstrate their utility through zero-shot classification of cell types and diseases across four independent datasets. By releasing all models, datasets, and tutorials openly, mmContext enables reproducible and accessible multimodal learning for omics-text integration. AVAILABILITY AND IMPLEMENTATION: Pretrained checkpoints and full source code for our custom MMContextEncoder are available on Hugging Face huggingface.co/jo-mengr. The Python package github.com/mengerj/mmcontext provides the model implementation and training and evaluation scripts for custom training. The releases for the publication can be accessed via zenodo: adata_hf_datasets: doi.org/10.5281/zenodo.19185217 and mmContext: doi.org/10.5281/zenodo.19185493.

Computational Biology↗

Evolving rule-based systems in two medical domains using genetic programming.

OBJECTIVE: To demonstrate and compare the application of different genetic programming (GP) based intelligent methodologies for the construction of rule-based systems in two medical domains: the diagnosis of aphasia's subtypes and the classification of pap-smear examinations. MATERIAL: Past data representing (a) successful diagnosis of aphasia's subtypes from collaborating medical experts through a free interview per patient, and (b) correctly classified smears (images of cells) by cyto-technologists, previously stained using the Papanicolaou method. METHODS: Initially a hybrid approach is proposed, which combines standard genetic programming and heuristic hierarchical crisp rule-base construction. Then, genetic programming for the production of crisp rule based systems is attempted. Finally, another hybrid intelligent model is composed by a grammar driven genetic programming system for the generation of fuzzy rule-based systems. RESULTS: Results denote the effectiveness of the proposed systems, while they are also compared for their efficiency, accuracy and comprehensibility, to those of an inductive machine learning approach as well as to those of a standard genetic programming symbolic expression approach. CONCLUSION: The proposed GP-based intelligent methodologies are able to produce accurate and comprehensible results for medical experts performing competitive to other intelligent approaches. The aim of the authors was the production of accurate but also sensible decision rules that could potentially help medical doctors to extract conclusions, even at the expense of a higher classification score achievement.

Aphasia↗

Application of latent semantic analysis to protein remote homology detection.

MOTIVATION: Remote homology detection between protein sequences is a central problem in computational biology. The discriminative method such as the support vector machine (SVM) is one of the most effective methods. Many of the SVM-based methods focus on finding useful representations of protein sequence, using either explicit feature vector representations or kernel functions. Such representations may suffer from the peaking phenomenon in many machine-learning methods because the features are usually very large and noise data may be introduced. Based on these observations, this research focuses on feature extraction and efficient representation of protein vectors for SVM protein classification. RESULTS: In this study, a latent semantic analysis (LSA) model, which is an efficient feature extraction technique from natural language processing, has been introduced in protein remote homology detection. Several basic building blocks of protein sequences have been investigated as the 'words' of 'protein sequence language', including N-grams, patterns and motifs. Each protein sequence is taken as a 'document' that is composed of bags-of-word. The word-document matrix is constructed first. The LSA is performed on the matrix to produce the latent semantic representation vectors of protein sequences, leading to noise-removal and smart description of protein sequences. The latent semantic representation vectors are then evaluated by SVM. The method is tested on the SCOP 1.53 database. The results show that the LSA model significantly improves the performance of remote homology detection in comparison with the basic formalisms. Furthermore, the performance of this method is comparable with that of the complex kernel methods such as SVM-LA and better than that of other sequence-based methods such as PSI-BLAST and SVM-pairwise.

Algorithms↗

A fully automated method for lung nodule detection from postero-anterior chest radiographs.

In the past decades, a great deal of research work has been devoted to the development of systems that could improve radiologists' accuracy in detecting lung nodules. Despite the great efforts, the problem is still open. In this paper, we present a fully automated system processing digital postero-anterior (PA) chest radiographs, that starts by producing an accurate segmentation of the lung field area. The segmented lung area includes even those parts of the lungs hidden behind the heart, the spine, and the diaphragm, which are usually excluded from the methods presented in the literature. This decision is motivated by the fact that lung nodules may be found also in these areas. The segmented area is processed with a simple multiscale method that enhances the visibility of the nodules, and an extraction scheme is then applied to select potential nodules. To reduce the high number of false positives extracted, cost-sensitive support vector machines (SVMs) are trained to recognize the true nodules. Different learning experiments were performed on two different data sets, created by means of feature selection, and employing Gaussian and polynomial SVMs trained with different parameters; the results are reported and compared. With the best SVM models, we obtain about 1.5 false positives per image (fp/image) when sensitivity is approximately equal to 0.71; this number increases to about 2.5 and 4 fp/image when sensitivity is = 0.78 and = 0.85, respectively. For the highest sensitivity (= 0.92 and 1.0), we get 7 or 8 fp/image.

Algorithms↗

Best harmony, unified RPCL and automated model selection for unsupervised and supervised learning on Gaussian mixtures, three-layer nets and ME-RBF-SVM models.

After introducing the fundamentals of BYY system and harmony learning, which has been developed in past several years as a unified statistical framework for parameter learning, regularization and model selection, we systematically discuss this BYY harmony learning on systems with discrete inner-representations. First, we shown that one special case leads to unsupervised learning on Gaussian mixture. We show how harmony learning not only leads us to the EM algorithm for maximum likelihood (ML) learning and the corresponding extended KMEAN algorithms for Mahalanobis clustering with criteria for selecting the number of Gaussians or clusters, but also provides us two new regularization techniques and a unified scheme that includes the previous rival penalized competitive learning (RPCL) as well as its various variants and extensions that performs model selection automatically during parameter learning. Moreover, as a by-product, we also get a new approach for determining a set of 'supporting vectors' for Parzen window density estimation. Second, we shown that other special cases lead to three typical supervised learning models with several new results. On three layer net, we get (i) a new regularized ML learning, (ii) a new criterion for selecting the number of hidden units, and (iii) a family of EM-like algorithms that combines harmony learning with new techniques of regularization. On the original and alternative models of mixture-of-expert (ME) as well as radial basis function (RBF) nets, we get not only a new type of criteria for selecting the number of experts or basis functions but also a new type of the EM-like algorithms that combines regularization techniques and RPCL learning for parameter learning with either least complexity nature on the original ME model or automated model selection on the alternative ME model and RBF nets. Moreover, all the results for the alternative ME model are also applied to other two popular nonparametric statistical approaches, namely kernel regression and supporting vector machine. Particularly, not only we get an easily implemented approach for determining the smoothing parameter in kernel regression, but also we get an alternative approach for deciding the set of supporting vectors in supporting vector machine.

Algorithms↗

Efficient Detection and Characterization of Targets of Natural Selection Using Transfer Learning.

Natural selection leaves detectable patterns of altered spatial diversity within genomes, and identifying affected regions is crucial for understanding species evolution. Recently, machine learning approaches applied to raw population genomic data have been developed to uncover these adaptive signatures. Convolutional neural networks (CNNs) are particularly effective for this task, as they handle large data arrays while maintaining element correlations. However, shallow CNNs may miss complex patterns due to their limited capacity, while deep CNNs can capture these patterns but require extensive data and computational power. Transfer learning addresses these challenges by utilizing a deep CNN pretrained on a large dataset as a feature extraction tool for downstream classification and evolutionary parameter prediction. This approach reduces extensive training data generation requirements and computational needs while maintaining high performance. In this study, we developed TrIdent, a tool that uses transfer learning to enhance detection of adaptive genomic regions from image representations of multilocus variation. We evaluated TrIdent across various genetic, demographic, and adaptive settings, in addition to unphased data and other confounding factors. TrIdent demonstrated improved detection of adaptive regions compared to recent methods using similar data representations. We further explored model interpretability through class activation maps and adapted TrIdent to infer selection parameters for identified adaptive candidates. Using whole-genome haplotype data from European and African populations, TrIdent effectively recapitulated known sweep candidates and identified novel cancer, and other disease-associated genes as potential sweeps.

Selection, Genetic↗

[Ventilators for anesthesia. Models available in France. Criteria for choice].

This update article discusses the criteria for the choice of an anaesthetic machine and provides a short analysis of the main components of the models commercialized in France in 1994. The following items are considered: the design of the machine, the fresh gas delivery system, the anaesthesia breathing system(s), the ventilator and the waste gas scavenging system, the monitors associated with the machine and other criteria such as facility of learning to run the machine and of its daily use, ease of "in-house" maintenance and quality of after-sales service, cost of the machine and of its use (driving gas, disposable equipment).

Anesthesia, Inhalation↗

Modular DAG-RNN architectures for assembling coarse protein structures.

We develop and test machine learning methods for the prediction of coarse 3D protein structures, where a protein is represented by a set of rigid rods associated with its secondary structure elements (alpha-helices and beta-strands). First, we employ cascades of recursive neural networks derived from graphical models to predict the relative placements of segments. These are represented as discretized distance and angle maps, and the discretization levels are statistically inferred from a large and curated dataset. Coarse 3D folds of proteins are then assembled starting from topological information predicted in the first stage. Reconstruction is carried out by minimizing a cost function taking the form of a purely geometrical potential. We show that the proposed architecture outperforms simpler alternatives and can accurately predict binary and multiclass coarse maps. The reconstruction procedure proves to be fast and often leads to topologically correct coarse structures that could be exploited as a starting point for various protein modeling strategies. The fully integrated rod-shaped protein builder (predictor of contact maps + reconstruction algorithm) can be accessed at http://distill.ucd.ie/.

Algorithms↗

Semisupervised learning for molecular profiling.

Class prediction and feature selection are two learning tasks that are strictly paired in the search of molecular profiles from microarray data. Researchers have become aware how easy it is to incur a selection bias effect, and complex validation setups are required to avoid overly optimistic estimates of the predictive accuracy of the models and incorrect gene selections. This paper describes a semisupervised pattern discovery approach that uses the by-products of complete validation studies on experimental setups for gene profiling. In particular, we introduce the study of the patterns of single sample responses (sample-tracking profiles) to the gene selection process induced by typical supervised learning tasks in microarray studies. We originate sample-tracking profiles as the aggregated off-training evaluation of SVM models of increasing gene panel sizes. Genes are ranked by E-RFE, an entropy-based variant of the recursive feature elimination for support vector machines (RFE-SVM). A Dynamic Time Warping (DTW) algorithm is then applied to define a metric between sample-tracking profiles. An unsupervised clustering based on the DTW metric allows automating the discovery of outliers and of subtypes of different molecular profiles. Applications are described on synthetic data and in two gene expression studies.

Algorithms↗

Associative memory design using support vector machines.

The relation existing between support vector machines (SVMs) and recurrent associative memories is investigated. The design of associative memories based on the generalized brain-state-in-a-box (GBSB) neural model is formulated as a set of independent classification tasks which can be efficiently solved by standard software packages for SVM learning. Some properties of the networks designed in this way are evidenced, like the fact that surprisingly they follow a generalized Hebb's law. The performance of the SVM approach is compared to existing methods with nonsymmetric connections, by some design examples.

Algorithms↗

Onvergence and application of online active sampling using orthogonal pillar vectors.

The analysis of convergence and its application is shown for the Active Sampling-at-the-Boundary method applied to multidimensional space using orthogonal pillar vectors. Active learning method facilitates identifying an optimal decision boundary for pattern classification in machine learning. The result of this method is compared with the standard active learning method that uses random sampling on the decision boundary hyperplane. The comparison is done through simulation and application to the real-world data from the UCI benchmark data set. The boundary is modeled as a nonseparable linear decision hyperplane in multidimensional space with a stochastic oracle.

Algorithms↗