Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Towards the development of a conceptual distance metric for the UMLS.

The objective of this work is to investigate the feasibility of conceptual similarity metrics in the framework of the Unified Medical Language System (UMLS). We have investigated an approach based on the minimum number of parent links between concepts, and evaluated its performance relative to human expert estimates on three sets of concepts for three terminologies within the UMLS (i.e., MeSH, ICD9CM, and SNOMED). The resulting quantitative metric enables computer-based applications that use decision thresholds and approximate matching criteria. The proposed conceptual matching supports problem solving and inferencing (using high-level, generic concepts) based on readily available data (typically represented as low-level, specific concepts). Through the identification of semantically similar concepts, conceptual matching also enables reasoning in the absence of exact, or even approximate, lexical matching. Finally, conceptual matching is relevant for terminology development and maintenance, machine learning research, decision support system development, and data mining research in biomedical informatics and other fields.

Algorithms↗

Induction of comprehensible models for gene expression datasets by subgroup discovery methodology.

Finding disease markers (classifiers) from gene expression data by machine learning algorithms is characterized by a high risk of overfitting the data due the abundance of attributes (simultaneously measured gene expression values) and shortage of available examples (observations). To avoid this pitfall and achieve predictor robustness, state-of-the-art approaches construct complex classifiers that combine relatively weak contributions of up to thousands of genes (attributes) to classify a disease. The complexity of such classifiers limits their transparency and consequently the biological insights they can provide. The goal of this study is to apply to this domain the methodology of constructing simple yet robust logic-based classifiers amenable to direct expert interpretation. On two well-known, publicly available gene expression classification problems, the paper shows the feasibility of this approach, employing a recently developed subgroup discovery methodology. Some of the discovered classifiers allow for novel biological interpretations.

Algorithms↗

Classification and knowledge discovery in protein databases.

We consider the problem of classification in noisy, high-dimensional, and class-imbalanced protein datasets. In order to design a complete classification system, we use a three-stage machine learning framework consisting of a feature selection stage, a method addressing noise and class-imbalance, and a method for combining biologically related tasks through a prior-knowledge based clustering. In the first stage, we employ Fisher's permutation test as a feature selection filter. Comparisons with the alternative criteria show that it may be favorable for typical protein datasets. In the second stage, noise and class imbalance are addressed by using minority class over-sampling, majority class under-sampling, and ensemble learning. The performance of logistic regression models, decision trees, and neural networks is systematically evaluated. The experimental results show that in many cases ensembles of logistic regression classifiers may outperform more expressive models due to their robustness to noise and low sample density in a high-dimensional feature space. However, ensembles of neural networks may be the best solution for large datasets. In the third stage, we use prior knowledge to partition unlabeled data such that the class distributions among non-overlapping clusters significantly differ. In our experiments, training classifiers specialized to the class distributions of each cluster resulted in a further decrease in classification error.

Algorithms↗

Improving the performance of dictionary-based approaches in protein name recognition.

Dictionary-based protein name recognition is often a first step in extracting information from biomedical documents because it can provide ID information on recognized terms. However, dictionary-based approaches present two fundamental difficulties: (1) false recognition mainly caused by short names; (2) low recall due to spelling variations. In this paper, we tackle the former problem using machine learning to filter out false positives and present two alternative methods for alleviating the latter problem of spelling variations. The first is achieved by using approximate string searching, and the second by expanding the dictionary with a probabilistic variant generator, which we propose in this paper. Experimental results using the GENIA corpus revealed that filtering using a naive Bayes classifier greatly improved precision with only a slight loss of recall, resulting in 10.8% improvement in F-measure, and dictionary expansion with the variant generator gave further 1.6% improvement and achieved an F-measure of 66.6%.

Abstracting and Indexing↗

Prospective recruitment of patients with congestive heart failure using an ad-hoc binary classifier.

This paper addresses a very specific problem of identifying patients diagnosed with a specific condition for potential recruitment in a clinical trial or an epidemiological study. We present a simple machine learning method for identifying patients diagnosed with congestive heart failure and other related conditions by automatically classifying clinical notes dictated at Mayo Clinic. This method relies on an automatic classifier trained on comparable amounts of positive and negative samples of clinical notes previously categorized by human experts. The documents are represented as feature vectors, where features are a mix of demographic information as well as single words and concept mappings to MeSH and HICDA classification systems. We compare two simple and efficient classification algorithms (Naïve Bayes and Perceptron) and a baseline term spotting method with respect to their accuracy and recall on positive samples. Depending on the test set, we find that Naïve Bayes yields better recall on positive samples (95 vs. 86%) but worse accuracy than Perceptron (57 vs. 65%). Both algorithms perform better than the baseline with recall on positive samples of 71% and accuracy of 54%.

Artificial Intelligence↗

Automatic analysis of medical dialogue in the home hemodialysis domain: structure induction and summarization.

Spoken medical dialogue is a valuable source of information for patients and caregivers. This work presents a first step towards automatic analysis and summarization of spoken medical dialogue. We first abstract a dialogue into a sequence of semantic categories using linguistic and contextual features integrated in a supervised machine-learning framework. Our model has a classification accuracy of 73%, compared to 33% achieved by a majority baseline (p<0.01). We then describe and implement a summarizer that utilizes this automatically induced structure. Our evaluation results indicate that automatically generated summaries exhibit high resemblance to summaries written by humans. In addition, task-based evaluation shows that physicians can reasonably answer questions related to patient care by looking at the automatically generated summaries alone, in contrast to the physicians' performance when they were given summaries from a naïve summarizer (p<0.05). This work demonstrates the feasibility of automatically structuring and summarizing spoken medical dialogue.

Artificial Intelligence↗

Using MEDLINE as a knowledge source for disambiguating abbreviations and acronyms in full-text biomedical journal articles.

Biomedical abbreviations and acronyms are widely used in biomedical literature. Since many of them represent important content in biomedical literature, information retrieval and extraction benefits from identifying the meanings of those terms. On the other hand, many abbreviations and acronyms are ambiguous, it would be important to map them to their full forms, which ultimately represent the meanings of the abbreviations. In this study, we present a semi-supervised method that applies MEDLINE as a knowledge source for disambiguating abbreviations and acronyms in full-text biomedical journal articles. We first automatically generated from the MEDLINE abstracts a dictionary of abbreviation-full pairs based on a rule-based system that maps abbreviations to full forms when full forms are defined in the abstracts. We then trained on the MEDLINE abstracts and predicted the full forms of abbreviations in full-text journal articles by applying supervised machine-learning algorithms in a semi-supervised fashion. We report up to 92% prediction precision and up to 91% coverage.

Artificial Intelligence↗

Applying hybrid reasoning to mine for associative features in biological data.

We develop the means to mine for associative features in biological data. The hybrid reasoning schema for deterministic machine learning and its implementation via logic programming is presented. The methodology of mining for correlation between features is illustrated by the prediction tasks for protein secondary structure and phylogenetic profiles. The suggested methodology leads to a clearer approach to hierarchical classification of proteins and a novel way to represent evolutionary relationships. Comparative analysis of Jasmine and other statistical and deterministic systems (including Explanation-Based Learning and Inductive Logic Programming) are outlined. Advantages of using deterministic versus statistical data mining approaches for high-level exploration of correlation structure are analyzed.

Algorithms↗

Using support vector machines to optimally classify rotator cuff strength data and quantify post-operative strength in rotator cuff tear patients.

Shoulder strength data are important for post-operative assessment of shoulder function and have been used in diagnosis of rotator cuff pathology. Support vector machines (SVM) employ complex analysis techniques to solve classification and regression problems. A SVM, a machine learning technique, can be used for analysis and classification of shoulder strength data. The goals of this study were to determine the diagnostic competency of SVM based on shoulder strength data and to apply SVM analysis in efforts to derive a single representative shoulder strength score. Data were taken from fourteen isometric shoulder strength measurements of each shoulder (involved and uninvolved) in 45 rotator cuff tear patients. SVM diagnostic proficiency was found to be comparable to reported ultrasound values. Improvement of shoulder function was accurately represented by a single score in pairwise comparison of the pre-operative and the 12 month post-operative group (P < 0.004). Thus, the SVM-based score may be a promising metric for summarizing rotator cuff strength data.

Artificial Intelligence↗

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers↗

A review on the integration of artificial intelligence into coastal modeling.

With the development of computing technology, mechanistic models are often employed to simulate processes in coastal environments. However, these predictive tools are inevitably highly specialized, involving certain assumptions and/or limitations, and can be manipulated only by experienced engineers who have a thorough understanding of the underlying theories. This results in significant constraints on their manipulation as well as large gaps in understanding and expectations between the developers and practitioners of a model. The recent advancements in artificial intelligence (AI) technologies are making it possible to integrate machine learning capabilities into numerical modeling systems in order to bridge the gaps and lessen the demands on human experts. The objective of this paper is to review the state-of-the-art in the integration of different AI technologies into coastal modeling. The algorithms and methods studied include knowledge-based systems, genetic algorithms, artificial neural networks, and fuzzy inference systems. More focus is given to knowledge-based systems, which have apparent advantages over the others in allowing more transparent transfers of knowledge in the use of models and in furnishing the intelligent manipulation of calibration parameters. Of course, the other AI methods also have their individual contributions towards accurate and reliable predictions of coastal processes. The integrated model might be very powerful, since the advantages of each technique can be combined.

Artificial Intelligence↗

Integrative dual-track transcriptomics reveals stage-specific coordination, regulatory divergence, and HSP90AA1-associated remodeling in human folliculogenesis.

Human folliculogenesis depends on coordinated yet non-identical developmental remodeling in the oocyte and its surrounding granulosa cells. When these two compartments remain synchronized and when they diverge into lineage-specific regulatory states, however, remains incompletely resolved. Here we performed an integrative dual-track re-analysis of the human RNA-seq dataset GSE107746, modeling oocytes and granulosa cells as distinct but developmentally linked compartments across follicular progression. Analysis of 148 sequencing libraries showed that compartment identity was the dominant source of transcriptomic variation, supporting compartment-aware downstream interpretation. Within this framework, oocytes followed a relatively continuous developmental trajectory, with substantial transcriptional remodeling already evident across adjacent stages, whereas granulosa cells showed weaker early-stage contrasts but markedly stronger late-stage reorganization, particularly around the antral and preovulatory transitions. Functional enrichment indicated that oocyte maturation was associated with RNA-processing and broader genome-regulatory remodeling, whereas granulosa maturation was dominated by progressive mitochondrial and bioenergetic activation. Co-expression analysis showed that both compartments contained strong late-stage programmes together with inverse early-state modules, indicating a shared systems-level architecture of maturation, although the hub-gene composition and biological content of these programmes were largely compartment-specific. Machine-learning validation reinforced this asymmetry: oocyte stage classification was best recovered from a compact eigengene-based representation, whereas granulosa stage discrimination was better resolved by a broader differential-expression-derived feature set. At the gene level, HSP90AA1 emerged as a stage-associated marker with compartment-specific behavior, showing progressive attenuation across oocyte development, assignment to the selected oocyte blue module, and sharper transitional dynamics in granulosa cells. Together, these findings support a model in which human folliculogenesis proceeds through coordinated but non-equivalent transcriptomic remodeling, with shared developmental logic at the systems level but distinct molecular execution in germline and somatic compartments.

Co-expression networks↗

Automated interpretation of subcellular patterns from immunofluorescence microscopy.

Immunofluorescence microscopy is widely used to analyze the subcellular locations of proteins, but current approaches rely on visual interpretation of the resulting patterns. To facilitate more rapid, objective, and sensitive analysis, computer programs have been developed that can identify and compare protein subcellular locations from fluorescence microscope images. The basis of these programs is a set of features that numerically describe the characteristics of protein images. Supervised machine learning methods can be used to learn from the features of training images and make predictions of protein location for images not used for training. Using image databases covering all major organelles in HeLa cells, these programs can achieve over 92% accuracy for two-dimensional (2D) images and over 95% for three-dimensional images. Importantly, the programs can discriminate proteins that could not be distinguished by visual examination. In addition, the features can also be used to rigorously compare two sets of images (e.g., images of a protein in the presence and absence of a drug) and to automatically select the most typical image from a set. The programs described provide an important set of tools for those using fluorescence microscopy to study protein location.

Automation↗

Clinical applications of digital twin technology in In Vitro Fertilisation.

BACKGROUND: Digital twin technology, originating from aerospace and manufacturing industries, has emerged as a transformative tool in healthcare. In vitro fertilisation (IVF) faces persistent challenges including suboptimal embryo selection, unpredictable treatment outcomes, and limited personalisation of protocols. Despite advances in assisted reproductive technology, existing literature exhibits fragmentation: artificial intelligence applications in embryo selection, ovarian stimulation, and endometrial assessment have been developed independently without systematic integration into comprehensive treatment frameworks. Digital twin technology offers unprecedented opportunities to create virtual replicas of biological systems, enabling real-time monitoring, predictive modelling, and personalised treatment strategies. AIM: This narrative review aims to critically examine the current applications of digital twin technology in IVF, evaluate its potential benefits and limitations, synthesize existing evidence into an integrative conceptual model, and identify future directions for implementation in reproductive medicine. METHOD: A comprehensive narrative review was conducted using PubMed, Scopus, Web of Science, and IEEE Xplore databases. A narrative review approach was selected over systematic review to accommodate the heterogeneity of evidence types in this emerging field, including theoretical frameworks, simulation studies, and proof-of-concept implementations that would be excluded from systematic reviews. Search terms included "digital twin," "IVF," "in vitro fertilisation," "assisted reproductive technology," "embryo selection," and "predictive modelling." Studies published between 2015 and 2025 were included, focusing on original research articles, systematic reviews, and proof-of-concept studies describing digital twin applications in reproductive medicine. RESULTS: Digital twin technology in IVF demonstrates significant potential across multiple domains including embryo development simulation, ovarian response prediction, endometrial receptivity modelling, and personalised stimulation protocols. Current applications integrate artificial intelligence, machine learning algorithms, time-lapse imaging, and omics data to create comprehensive virtual models. Early evidence suggests improvements in embryo selection accuracy, ovarian response prediction, and treatment protocol optimization, though large-scale randomized controlled trials remain limited. Implementation challenges include data integration complexity, computational requirements, regulatory considerations, and validation requirements. CONCLUSION: Digital twin technology represents a paradigm shift in IVF practice, offering personalised, predictive, and precision medicine approaches. This review synthesizes existing evidence to propose an integrative conceptual model for digital twin implementation across the IVF treatment spectrum, identifies critical knowledge gaps, and establishes research priorities to advance clinical translation. Despite current limitations, continued advancement promises improved success rates and patient outcomes.

Humans↗

A flexible computational framework for detecting, characterizing, and interpreting statistical patterns of epistasis in genetic studies of human disease susceptibility.

Detecting, characterizing, and interpreting gene-gene interactions or epistasis in studies of human disease susceptibility is both a mathematical and a computational challenge. To address this problem, we have previously developed a multifactor dimensionality reduction (MDR) method for collapsing high-dimensional genetic data into a single dimension (i.e. constructive induction) thus permitting interactions to be detected in relatively small sample sizes. In this paper, we describe a comprehensive and flexible framework for detecting and interpreting gene-gene interactions that utilizes advances in information theory for selecting interesting single-nucleotide polymorphisms (SNPs), MDR for constructive induction, machine learning methods for classification, and finally graphical models for interpretation. We illustrate the usefulness of this strategy using artificial datasets simulated from several different two-locus and three-locus epistasis models. We show that the accuracy, sensitivity, specificity, and precision of a naïve Bayes classifier are significantly improved when SNPs are selected based on their information gain (i.e. class entropy removed) and reduced to a single attribute using MDR. We then apply this strategy to detecting, characterizing, and interpreting epistatic models in a genetic study (n = 500) of atrial fibrillation and show that both classification and model interpretation are significantly improved.

Atrial Fibrillation↗

Using pseudo-amino acid composition and support vector machine to predict protein structural class.

As a result of genome and other sequencing projects, the gap between the number of known protein sequences and the number of known protein structural classes is widening rapidly. In order to narrow this gap, it is vitally important to develop a computational prediction method for fast and accurately determining the protein structural class. In this paper, a novel predictor is developed for predicting protein structural class. It is featured by employing a support vector machine learning system and using a different pseudo-amino acid composition (PseAA), which was introduced to, to some extent, take into account the sequence-order effects to represent protein samples. As a demonstration, the jackknife cross-validation test was performed on a working dataset that contains 204 non-homologous proteins. The predicted results are very encouraging, indicating that the current predictor featured with the PseAA may play an important complementary role to the elegant covariant discriminant predictor and other existing algorithms.

Amino Acid Sequence↗

Estimating the association of antimicrobial resistance genes with minimum inhibitory concentration in Escherichia coli: an observational study.

BACKGROUND: Surveillance and prediction of antibiotic resistance in Escherichia coli relies on curated databases of genes and mutations. We aimed to quantify the effect of acquiring specific genetic elements on minimum inhibitory concentrations (MICs) for particular antibiotic-species combinations, addressing the current scarcity of such data in existing databases. METHODS: For this observational study, we evaluated a collection of E coli isolates with linked whole-genome sequencing and MIC data, originating from human urinary or bloodstream infections obtained from the Oxford University Hospitals National Health Service Foundation Trust in Oxfordshire, UK. We used multivariable interval regression models to estimate the change in MIC (with 95% CIs) for specific antibiotics associated with the acquisition of antibiotic resistance genes and associated mutations in the National Center for Biotechnology Information AMRFinder database, with and without an adjustment for population structure. We then tested the ability of these models to predict MIC and binary resistance or susceptibility using leave-one-out cross-validation. FINDINGS: We evaluated 2875 E coli isolates obtained during 2013-2018 and 2020. Although most ARGs and resistance mutations (89 [80%] of 111) were associated with an increased MIC, a much smaller number (27 [24%] of 111) was found to be putatively independently resistance-conferring (ie, associated with an MIC above the European Committee on Antimicrobial Susceptibility Testing breakpoint) when acquired in isolation. We found evidence of differential effects of acquired ARGs and resistance mutations between different generations of cephalosporin antibiotics and showed that sub-breakpoint variation in MIC can be linked to genetic mechanisms of resistance. 20&#x2009;697 (83&#xb7;3%; range 52&#xb7;9-97&#xb7;7 across all antibiotics) of 24&#x2009;858 MICs were correctly exactly predicted and 23&#x2009;677 (95&#xb7;2%; 87&#xb7;3-97&#xb7;7) of 24&#x2009;858 MICs were predicted to within one doubling dilution. INTERPRETATION: Quantitative estimates of the independent effect of the acquisition of ARGs on MIC add to the interpretability and utility of existing databases. Compared with approaches using machine learning models, the use of these estimates yields similar or better performance in the prediction of antibiotic resistance phenotype with more readily interpretable results. The methods outlined here could be readily applied to other antibiotic-pathogen combinations. FUNDING: The National Institute for Health and Care Research (NIHR) and the Medical Research Council (MRC).

Escherichia coli↗

Bioactive peptides for meat quality and preservation: Integrating peptidomics and computational screening.

Bioactive peptides generated from meat proteins, fermented meat products, and slaughter by-products have attracted increasing attention as functional molecules for improving meat quality and preservation. In meat systems, peptides can be produced through endogenous postmortem proteolysis, microbial fermentation, gastrointestinal digestion, or controlled enzymatic hydrolysis of underutilized animal by-products. These peptides are closely associated with key meat science endpoints, including postmortem tenderization, oxidative stability, color retention, flavor development, microbial inhibition, and the valorization of processing by-products. However, although high-resolution peptidomics has greatly expanded the identification of meat-derived peptide sequences, their translation into practical meat applications remains limited by matrix interactions, processing stability, sensory constraints, safety concerns, and insufficient validation in real meat systems. This review synthesizes recent advances in meat-related peptidomics and computational screening, including sequence-based prediction, machine learning, molecular docking, molecular dynamics, stability assessment, and safety-oriented filtering. Particular attention is given to how these approaches can prioritize peptides with antioxidant, antimicrobial, flavor-modulating, and preservation-related functions under meat-specific technological constraints. By integrating peptide generation pathways, mass spectrometry-based identification, in silico prioritization, and meat quality endpoints, this review proposes a stage-gated framework for translating meat-derived bioactive peptides from discovery to application. Future research should strengthen matrix-specific validation, standardized peptidomic reporting, and safety assessment to support the use of bioactive peptides in meat quality improvement, clean-label preservation, and circular utilization of meat industry by-products.

Animals↗