Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

QSAR and classification models of a novel series of COX-2 selective inhibitors: 1,5-diarylimidazoles based on support vector machines.

The support vector machine, which is a novel algorithm from the machine learning community, was used to develop quantitation and classification models which can be used as a potential screening mechanism for a novel series of COX-2 selective inhibitors. Each compound was represented by calculated structural descriptors that encode constitutional, topological, geometrical, electrostatic, and quantum-chemical features. The heuristic method was then used to search the descriptor space and select the descriptors responsible for activity. Quantitative modelling results in a nonlinear, seven-descriptor model based on SVMs with root mean-square errors of 0.107 and 0.136 for training and prediction sets, respectively. The best classification results are found using SVMs: the accuracy for training and test sets is 91.2% and 88.2%, respectively. This paper proposes a new and effective method for drug design and screening.

Algorithms↗

The prediction of human oral absorption for diffusion rate-limited drugs based on heuristic method and support vector machine.

Support vector machine (SVM), as a novel machine learning technique, was used for the prediction of the human oral absorption for a large and diverse data set using the five descriptors calculated from the molecular structure alone. The molecular descriptors were selected by heuristic method (HM) implemented in CODESSA. At the same time, in order to show the influence of different molecular descriptors on absorption and to well understand the absorption mechanism, HM was used to build several multivariable linear models using different numbers of molecular descriptors. Both the linear and non-linear model can give satisfactory prediction results: the square of correlation coefficient R(2) was 0.78 and 0.86 for the training set, and 0.70 and 0.73 for the test set respectively. In addition, this paper provides a new and effective method for predicting the absorption of the drugs from their structures and gives some insight into structural features related to the absorption of the drugs.

Administration, Oral↗

Prediction of the tissue/blood partition coefficients of organic compounds based on the molecular structure using least-squares support vector machines.

The accurate nonlinear model for predicting the tissue/blood partition coefficients (PC) of organic compounds in different tissues was firstly developed based on least-squares support vector machines (LS-SVM), as a novel machine learning technique, by using the compounds' molecular descriptors calculated from the structure alone and the composition features of tissues. The heuristic method (HM) was used to select the appropriate molecular descriptors and build the linear model. The prediction result of the LS-SVM model is much better than that obtained by HM method and the prediction values of tissue/blood partition coefficients based on the LS-SVM model are in good agreement with the experimental values, which proved that nonlinear model can simulate the relationship between the structural descriptors, the tissue composition and the tissue/blood partition coefficients more accurately as well as LS-SVM was a powerful and promising tool in the prediction of the tissue/blood partition behaviour of compounds. Furthermore, this paper provided a new and effective method for predicting the tissue/blood partition behaviour of the compounds in the different tissues from their structures and gave some insight into structural features related to the partition process of the organic compounds in different tissues.

Least-Squares Analysis↗

Prediction of standard Gibbs energies of the transfer of peptide anions from aqueous solution to nitrobenzene based on support vector machine and the heuristic method.

Quantitative structure-property relationship (QSPR) method was performed for the prediction of the standard Gibbs energies (DeltaGtheta) of the transfer of peptide anions from aqueous solution to nitrobenzene. Descriptors calculated from the molecular structures alone were used to represent the characteristics of the peptides. The four molecular descriptors selected by the heuristic method (HM) in COmprehensive DEscriptors for Structural and Statistical Analysis (CODESSA) were used as inputs for support vector machine (SVM) and radial basis function neural networks (RNFNN). The results obtained by the novel machine learning technique, SVM, were compared with those obtained by HM and RBFNN. The root mean squared errors (RMS) of the training, predicted and overall data sets are 2.192, 2.541 and 2.267 unit (kJ/mol) for HM, 1.604, 2.478 and 1.817 unit (kJ/mol) for RBFNN and 1.5621, 2.364 and 1.756 unit (kJ/mol) for SVM, respectively. The prediction results were in agreement with the experimental values. This paper provided a potential method for predicting the physiochemical property (DeltaGtheta) of various small peptides.

Anions↗

Feature-map vectors: a new class of informative descriptors for computational drug discovery.

In order to develop robust machine-learning or statistical models for predicting biological activity, descriptors that capture the essence of the protein-ligand interaction are required. In the absence of structural information from X-ray or NMR experiments, deriving informative descriptors can be difficult. We have developed feature-map vectors (FMVs), a new class of descriptors based on chemical features, to address this challenge. FMVs, which are derived from the conformational models of a few actives, are low dimensional, problem specific, and highly interpretable. By using shape-based alignments and scoring with chemical features, FMVs can combine information about a molecule's shape and the pharmacophores it can match. In five validation studies, bag classifiers built using FMVs have shown high enrichments for identifying actives for five diverse targets: CDK2, 5-HT(3), DHFR, thrombin, and ACE. The interpretability of these descriptors has been demonstrated for CDK2 and 5-HT(3), where the method automatically discovers the standard literature pharmacophore.

Algorithms↗

DNA methylation biomarkers for early detection of ovarian cancer.

Ovarian cancer (OC) remains difficult to detect at an early stage, and current screening approaches using CA125 and transvaginal ultrasonography have not demonstrated sufficient benefit for population screening. DNA methylation is a promising biomarker class because epigenetic alterations may arise early in tumourigenesis, can be detected in circulating cell-free DNA (cfDNA), and may provide tissue-of-origin information. This review critically evaluates recent evidence on DNA methylation biomarkers for early OC detection. PubMed/MEDLINE, Web of Science, and Scopus were searched for studies published between January 2020 and September 2025, supplemented by selected earlier studies of biological or methodological relevance. Evidence was synthesised across single-gene biomarkers, multi-locus panels, genome-wide signatures, assay platforms, and machine-learning classifiers, with emphasis on early-stage performance, histological representation, comparator populations, analytical methodology, and validation design. Single-gene markers such as BRCA1, RASSF1A, OPCML, HOXA9, and HIC1 show variable performance, while multi-gene and classifier-based approaches generally provide stronger discrimination. However, many studies remain limited by retrospective case-control designs, small FIGO stage I-II subsets, predominance of serous disease, and insufficient prospective validation. Integration with CA125 may improve sensitivity but can reduce specificity, which is critical in low-prevalence screening. Clinical translation will therefore require minimal and reproducible methylation signatures, standardised low-input cfDNA workflows, rigorous external validation, and prospective longitudinal evaluation in intended-use populations.

Humans↗

Collateral sensitivity-harnessing microbial vulnerabilities as a solution to antimicrobial resistance.

Bacteria exhibit an evolutionary trade-off through their development of collateral sensitivity (CS) which allows them to resist one antibiotic while becoming more vulnerable to another. This vulnerability offers a compelling therapeutic opportunity by selecting against resistant isolates. Laboratory evolution studies, genome sequencing, deep mutagenesis and use of artificial intelligence and machine learning can design the bespoke strategy against multi-drug-resistant bacteria. This review discusses about recent studies that are rationally designed to harness this evolutionary trade-off for the development of alternative antimicrobial strategies. The translational barriers to the clinical implementation of CS are addressed and evidence-based design principles for optimization of CS-guided therapy are discussed.

Bacteria↗

DNA microarrays in prostate cancer.

DNA microarray technology provides a means to examine large numbers of molecular changes related to a biological process in a high throughput manner. This review discusses plausible utilities of this technology in prostate cancer research, including definition of prostate cancer predisposition, global profiling of gene expression patterns associated with cancer initiation and progression, identification of new diagnostic and prognostic markers, and discovery of novel patient classification schemes. The technology, at present, has only been explored in a limited fashion in prostate cancer research. Some hurdles to be overcome are the high cost of the technology, insufficient sample size and repeated experiments, and the inadequate use of bioinformatics. With the completion of the Human Genome Project and the advance of several highly complementary technologies, such as laser capture microdissection, unbiased RNA amplification, customized functional arrays (eg, single-nucleotide polymorphism chips), and amenable bioinformatics software, this technology will become widely used by investigators in the field. The large amount of novel, unbiased hypotheses and insights generated by this technology is expected to have a significant impact on the diagnosis, treatment, and prevention of prostate cancer. Finally, this review emphasizes existing, but currently underutilized, data-mining tools, such as multivariate statistical analyses, neural networking, and machine learning techniques, to stimulate wider usage.

Algorithms↗

Oncogenic EME1 promotes tumor progression and immune modulation in human cancers with therapeutic targeting potential.

BACKGROUND: EME1, a critical DNA repair endonuclease, has emerged as a potential oncogene implicated in genome instability and cancer progression. However, its pan-cancer roles, prognostic significance, immune interactions, and therapeutic targeting remain underexplored. METHODS: We conducted a comprehensive pan-cancer analysis integrating multi-omics data from public databases, including TIMER2.0, GEPIA2, TISIDB, and cBioPortal, to evaluate EME1 expression, genetic alterations, and their association with clinical outcomes, immune infiltration, and molecular pathways. Virtual screening of 3180 FDA-approved drugs and molecular dynamics (MD) simulations were employed to identify and validate potential EME1 inhibitors. RESULTS: EME1 was significantly overexpressed in various human cancers and positively associated with advanced tumor grade and stage. High EME1 expression and mutations were linked to poor overall and disease-free survival. Immunogenomic profiling revealed strong positive correlations between EME1 and myeloid-derived suppressor cells (MDSCs), alongside a negative association with endothelial cell function, suggesting immunosuppressive roles. Machine learning models based on EME1-associated genes demonstrated high predictive accuracy for liver hepatocellular carcinoma (AUC > 0.90). Virtual screening identified eight promising drug candidates, including Everolimus and Dioscin, with strong binding affinities. MD simulations confirmed the stability of these interactions, particularly for Dioscin. CONCLUSION: This study reveals the multifaceted oncogenic roles of EME1 in tumor progression, immune evasion, and prognosis. It proposes EME1 as a promising biomarker and therapeutic target across multiple cancer types. The identified drug candidates warrant further in vitro and in vivo validation for potential repurposing in EME1-targeted cancer therapy.

EME1↗

International Federation of Clinical Chemistry. Use of artificial intelligence in analytical systems for the clinical laboratory. IFCC Committee on Analytical Systems.

The incorporation of information-processing technology into analytical systems in the form of standard computing software has recently been advanced by the introduction of artificial intelligence (AI) both as expert systems and as neural networks. This paper considers the role of software in system operation, control and automation and attempts to define intelligence. AI is characterized by its ability to deal with incomplete and imprecise information and to accumulate knowledge. Expert systems, building on standard computing techniques, depend heavily on the domain experts and knowledge engineers that have programmed them to represent the real world. Neural networks are intended to emulate the pattern-recognition and parallel-processing capabilities of the human brain and are taught rather than programmed. The future may lie in a combination of the recognition ability of the neural network and the rationalization capability of the expert system. In the second part of this paper, examples are given of applications of AI in stand-alone systems for knowledge engineering and medical diagnosis and in embedded systems for failure detection, image analysis, user interfacing, natural language processing, robotics and machine learning, as related to clinical laboratories. It is concluded that AI constitutes a collective form of intellectual property and that there is a need for better documentation, evaluation and regulation of the systems already being used widely in clinical laboratories.

Artificial Intelligence↗

Use of artificial intelligence in analytical systems for the clinical laboratory.

OBJECTIVE: To consider the role of software in system operation, control and automation, and attempts to define intelligence. METHODS AND RESULTS: Artificial intelligence (Al) is characterized by its ability to deal with incomplete and imprecise information and to accumulate knowledge. Expert systems, building on standard computing techniques, depend heavily on the domain experts and knowledge engineers that have programmed them to represent the real world. Neural networks are intended to emulate the pattern-recognition and parallel processing capabilities of the human brain and are taught rather than programmed. The future may lie in a combination of the recognition ability of the neural network and the rationalization capability of the expert system. In the second part of this paper, examples are given of applications of Al in stand-alone systems for knowledge engineering and medical diagnosis and in embedded systems for failure detection, image analysis, user interfacing, natural language processing, robotics and machine learning, as related to clinical laboratories. CONCLUSION: Al constitutes a collective form of intellectual property, and that there is a need for better documentation, evaluation and regulation of the systems already being used widely in clinical laboratories.

Artificial Intelligence↗

Representation and semiautomatic acquisition of medical knowledge in CADIAG-1 and CADIAG-2.

CADIAG-1 and CADIAG-2 (Computer-Assisted DIAGnosis) are medical expert systems especially designed for ill-defined areas such as internal medicine. Both systems are being tested in the setting of a medical information system. With respect to their knowledge representation, CADIAG-1 has obvious advantages in totally ill-defined areas such as syndromes in internal medicine, whereas CADIAG-2 seems more suited for domains with basic laboratory programs, e.g., hepatology or gall bladder and bile duct diseases. The formalization of relationships between medical entities led to first-order predicate calculus formulas in the case of CADIAG-1 and to a model based on fuzzy set theory in the case of CADIAG-2. In both systems two kinds of relationships between medical entities are considered: (1) necessity of occurrence and (2) sufficiency of occurrence. Statistical interpretations using the 2 X 2 table paradigm yield a way to calculate these relationships automatically from samples of patient data. Results obtained by exploiting 3530 patient records from a rheumatological hospital are presented. The described application is a machine-learning program that allows inductive learning from examples under statistical uncertainty.

Artificial Intelligence↗

Knowledge discovery in biomedical databases: a machine induction approach.

The increase in the number and size of available databases by far exceeds the growth of the corresponding knowledge. Furthermore, many databases contain information which is not possessed by an existing human expert. This creates both a need and an opportunity for extracting knowledge from databases. An unsolved problem in molecular biology is the problem of predicting a protein's secondary structure from its primary structure. Inductive machine learning is a search for a plausible general description which can explain the given input data, and is useful for predicting new data. In this paper we present a statistical inductive algorithm which can be used to produce new rules for predicting multiple protein secondary structures from protein primary structure databases.

Algorithms↗

Likelihood linkage analysis (LLA) classification method: an example treated by hand.

This paper describes a very general method of data analysis using a hierarchical classification. The data can be provided by observation, experiment or knowledge; their nature can be numerical, qualitative or logical. First, the classical view of the context of data representation, in which the algorithm of hierarchical ascendant construction of the classification tree is set, is treated in a synthetic manner. The main notion in our method is one of 'similarity'. This must be elaborated in the best way, taking into account the mathematical nature of the objects to be compared. Here we adopt a set of theoretical and combinatorial representation of the descriptive attributes, which are interpreted in terms of relations. Then we introduce a probability scale for similarity measurement by using a likelihood concept. The largest part of the paper concerns an illustrating example, moderately sized, detailing very minutely the different steps and the different calculations assumed by the method. The data structure handled with this example is the simplest possible. Then, general aspects and methodological extensions are evoked. We end by indicating the interest of the described approach in future works, in which we are involved, concerning typological organization of genetic sequences. We emphasize the 'explanation' aspect of the obtained results, with respect to a given description. For this purpose, classifications (on the object set and on the attribute set) on the one hand and machine learning techniques on the other, intervene efficiently.

Algorithms↗

A computer aided system for systematic production and revision of sequence patterns.

We used two complementary fields, object-oriented databases and machine learning, to produce and revise a set of protein sequence patterns. In a first stage, we show that object-oriented query languages are well suited for the production of patterns as well as for the interpretation of the biological function of new (uncharacterized) sequences. In a second stage, a classification is built from the set of sequences according to the pattern matches. This classification may be criticized by a specific analysis method, which yields back to revise sequences and patterns. In our application, we have used concept lattices as a classification method and sequence multiple alignment for criticism.

Algorithms↗

New approaches in molecular structure prediction.

In the past years, much effort has been put on the development of new methodologies and algorithms for the prediction of protein secondary and tertiary structures from (sequence) data; this is reviewed in detail. New approaches for these predictions such as neural network methods, genetic algorithms, machine learning, and graph theoretical methods are discussed. Secondary structure prediction algorithms were improved mostly by considering families of related proteins; however, for the reliable tertiary structure modeling of proteins, knowledge-based techniques are still preferred. Methods and examples with more or less successful results are described. Also, programs and parameterizations for energy minimisations, molecular dynamics, and electrostatic interactions have been improved, especially with respect to their former limits of applicability. Other topics discussed in this review include the use of traditional and on-line databases, the docking problem and surface properties of biomolecules, packing of protein cores, de novo design and protein engineering, prediction of membrane protein structures, the verification and reliability of model structures, and progress made with currently available software and computer hardware. In summary, the prediction of the structure, function, and other properties of a protein is still possible only within limits, but these limits continue to be moved.

Chemical Phenomena↗

Protein secondary structure prediction.

The past year has seen a consolidation of protein secondary structure prediction methods. The advantages of prediction from an aligned family of proteins have been highlighted by several accurate predictions made 'blind', before any X-ray or NMR structure was known for the family. New techniques that apply machine learning and discriminant analysis show promise as alternatives to neural networks.

Animals↗

Mining metagenomes from extremophiles as a resource for novel glycoside hydrolases for industrial applications.

The exploration of metagenomes from extremophiles has emerged as a promising approach for discovering novel glycoside hydrolases (GHs) with potential industrial applications. Extremophiles, which thrive in harsh conditions such as high salinity, extreme temperatures, and acidic or alkaline environments, produce enzymes naturally adapted to function under these conditions. This unique adaptability makes them highly desirable for industrial processes requiring robust and efficient biocatalysts. These biocatalysts reduce reliance on harsh chemicals and energy-intensive processes, contributing to greener industrial operations. This review underscores the power of metagenomics in bypassing the need to culture large libraries of extremophiles in the lab. High-throughput sequencing and bioinformatics enable the identification of novel GH-encoding genes directly from environmental DNA. While metagenomic mining has yielded promising results, challenges such as the expression of extremophile-derived genes in mesophilic hosts, low activity yields, and scalability remain. Advances in synthetic biology and protein engineering could address these bottlenecks, enabling more efficient utilization of GHs. Additionally, integrating machine learning for predictive functional annotation may accelerate the identification of high-value candidates.

Glycoside Hydrolases↗