Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

ORBIT: Oncogenic Representation Learning via Bi-Prototype Contrastive Learning in Hyperbolic Space for cancer driver gene identification.

Accurate identification of cancer driver genes is crucial for precision oncology but remains challenging due to the complexity of integrating heterogeneous data and modeling dynamic biological systems. To address these limitations, we propose ORBIT (Oncogenic Representation Learning via Bi-Prototype Contrastive Learning in Hyperbolic Space). Our framework synergistically fuses multi-omics profiles with functional network data using a context-adaptive graph reweighting mechanism to capture cancer-specific dynamics. The model employs a bi-prototype contrastive learning strategy within hyperbolic space, which aligns gene representations around distinct driver and non-driver semantic anchors while preserving the intrinsic hierarchy of biological networks. Comprehensive evaluations demonstrate that ORBIT achieves highly competitive stability in pan-cancer analysis while consistently outperforming state-of-the-art methods in cancer-specific predictions. Furthermore, functional enrichment analysis confirms that the model effectively segregates core cancer pathways, and drug sensitivity profiling validates the clinical relevance of the identified drivers. By integrating hyperbolic geometry with context-adaptive learning, ORBIT offers a robust and interpretable paradigm for precision medicine. The source codes and datasets are publicly accessible at https://github.com/spcho-dev/ORBIT.

Humans↗

Unsupervised learning with independent component analysis can identify patterns of glaucomatous visual field defects.

PURPOSE: We previously reported the use of clustering by unsupervised learning with machine learning classifiers to segment clusters of patterns in standard automated perimetry (SAP) for glaucoma. In this study, the process of unsupervised learning by independent component analysis decomposed SAP field patterns into axes, and the information represented by these axes was evaluated. METHODS: SAP fields were obtained with the Humphrey Visual Field Analyzer on 189 normal eyes and 156 eyes with glaucomatous optic neuropathy (GON) determined by masked review with stereoscopic optic disc photos. The variational Bayesian independent component analysis mixture model (vB-ICA-mm) partitioned the SAP fields into the most informative number of clusters. Simultaneously, it learned an optimal number of maximally independent axes for each cluster. RESULTS: The most informative number of clusters was two. vB-ICA-mm placed 68.6% of the SAP fields from eyes with GON in a cluster labeled G and 98.4% of the fields from eyes with normal optic discs in a cluster labeled N. Cluster G optimally contained six axes. Post hoc analysis of patterns generated at -1 SD and +2 SD from the cluster G mean on the six axes revealed defects similar to those identified by experts as indicative of glaucoma. SAP fields associated with an axis showed increasing severity as they were located farther in the positive direction from the cluster G mean. CONCLUSIONS: vB-ICA-mm represented the SAP fields with patterns that were meaningful for glaucoma experts. This process also captured severity in the patterns uncovered. These findings should validate vB-ICA-mm as a data mining technique for new and unfamiliar complex tests.

Artificial Intelligence↗

Prediction of 'drug-likeness'.

Recent developments in combinatorial chemistry and high-throughput screening have dramatically increased the scale on which drug discovery programs are carried out. Along with these advances has come a need for automated methods of determining which compounds from a library should be synthesized and screened. These methods range from simple counting schemes to sophisticated machine learning techniques such as neural networks. While many of these methods have performed well in validation studies, the field is still in its formative stage. This paper reviews a number of computational techniques for identifying drug-like molecules and examines challenges facing the field.

Artificial Intelligence↗

Unsupervised pattern recognition: an introduction to the whys and wherefores of clustering microarray data.

Clustering has become an integral part of microarray data analysis and interpretation. The algorithmic basis of clustering -- the application of unsupervised machine-learning techniques to identify the patterns inherent in a data set -- is well established. This review discusses the biological motivations for and applications of these techniques to integrating gene expression data with other biological information, such as functional annotation, promoter data and proteomic data.

Algorithms↗

An approach to biological computation: unicellular core-memory creatures evolved using genetic algorithms.

A novel machine language genetic programming system that uses one-dimensional core memories is proposed and simulated. The core is compared to a biochemical reaction space, and in imitation of biological molecules, four types of data words (Membrane, Pure data, Operator, and Instruction) are prepared in the core. A program is represented by a sequence of Instructions. During execution of the core, Instructions are transcribed into corresponding Operators, and Operators modify, create, or transfer Pure data. The core is hierarchically partitioned into sections by the Membrane data, and the data transfer between sections by special channel Operators constitutes a tree data-flow structure among sections in the core. In the experiment, genetic algorithms are used to modify program information. A simple machine learning problem is prepared for the environment data set of the creatures (programs), and the fitness value of a creature is calculated from the Pure data excreted by the creature. Breeding of programs that can output the predefined answer is successfully carried out. Several future plans to extend this system are also discussed.

Algorithms↗

[Fish-based assessment methods for the ecological status of aquatic systems].

A short overview of fish-based assessment methods for aquatic systems is presented. Multimetric indices, as, e.g., the index of biotic integrity (IBI), firstly developed in USA and later adapted for European river basins and other countries, are shortly described. Non-multimetric indices are also discussed, e.g. the ichthyological index (II) and the index of ecological status of fish communities (ISECI), both proposed for monitoring Italian rivers. Moreover, statistical and predicting methods based on machine learning techniques are described. Finally, a new approach for developing standardised fish-based methods useful to assess ecological status of Italian rivers is proposed. Although rather complex, the use of bony fish in biomonitoring is promising and requires a multidisciplinary approach to be adopted.

Animals↗

Constructing biological knowledge bases by extracting information from text sources.

Recently, there has been much effort in making databases for molecular biology more accessible and interoperable. However, information in text form, such as MEDLINE records, remains a greatly underutilized source of biological information. We have begun a research effort aimed at automatically mapping information from text sources into structured representations, such as knowledge bases. Our approach to this task is to use machine-learning methods to induce routines for extracting facts from text. We describe two learning methods that we have applied to this task--a statistical text classification method, and a relational learning method--and our initial experiments in learning such information-extraction routines. We also present an approach to decreasing the cost of learning information-extraction routines by learning from "weakly" labeled training data.

Artificial Intelligence↗

Activity theory as a framework for analyzing and redesigning work.

Cultural-historical activity theory is a new framework aimed at transcending the dichotomies of micro- and macro-, mental and material, observation and intervention in analysis and redesign of work. The approach distinguishes between short-lived goal-directed actions and durable, object-oriented activity systems. A historically evolving collective activity system, seen in its network relations to other activity systems, is taken as the prime unit of analysis against which scripted strings of goal-directed actions and automatic operations are interpreted. Activity systems are driven by communal motives that are often difficult to articulate for individual participants. Activity systems are in constant movement and internally contradictory. Their systemic contradictions, manifested in disturbances and mundane innovations, offer possibilities for expansive developmental transformations. Such transformations proceed through stepwise cycles of expansive learning which begin with actions of questioning the existing standard practice, then proceed to actions of analyzing its contradictions and modelling a vision for its zone of proximal development, then to actions of examining and implementing the new model in practice. New forms of work organization increasingly require negotiated 'knotworking' across boundaries. Correspondingly, expansive learning increasingly involves horizontal widening of collective expertise by means of debating, negotiating and hybridizing different perspectives and conceptualizations. Findings from a longitudinal intervention study of children's medical care illuminate the theoretical arguments.

Adult↗

Discrimination of modes of action of antifungal substances by use of metabolic footprinting.

Diploid cells of Saccharomyces cerevisiae were grown under controlled conditions with a Bioscreen instrument, which permitted the essentially continuous registration of their growth via optical density measurements. Some cultures were exposed to concentrations of a number of antifungal substances with different targets or modes of action (sterol biosynthesis, respiratory chain, amino acid synthesis, and the uncoupler). Culture supernatants were taken and analyzed for their "metabolic footprints" by using direct-injection mass spectrometry. Discriminant function analysis and hierarchical cluster analysis allowed these antifungal compounds to be distinguished and classified according to their modes of action. Genetic programming, a rule-evolving machine learning strategy, allowed respiratory inhibitors to be discriminated from others by using just two masses. Metabolic footprinting thus represents a rapid, convenient, and information-rich method for classifying the modes of action of antifungal substances.

Antifungal Agents↗

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n = 549) and a validation set (n = 236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60 mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60 mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans↗

The use of misclassification costs to learn rule-based decision support models for cost-effective hospital admission strategies.

Cost-effective health care is at the forefront of today's important health-related issues. A research team at the University of Pittsburgh has been interested in lowering the cost of medical care by attempting to define a subset of patients with community-acquire pneumonia for whom outpatient therapy is appropriate and safe. Sensitivity and specificity requirements for this domain make it difficult to use rule-based learning algorithms with standard measures of performance based on accuracy. This paper describes the use of misclassification costs to assist a rule-based machine-learning program in deriving a decision-support aid for choosing outpatient therapy for patients with community-acquired pneumonia.

Algorithms↗

Rigorous proof of termination of SMO algorithm for support vector machines.

Sequential minimal optimization (SMO) algorithm is one of the simplest decomposition methods for learning of support vector machines (SVMs). Keerthi and Gilbert have recently studied the convergence property of SMO algorithm and given a proof that SMO algorithm always stops within a finite number of iterations. In this letter, we point out the incompleteness of their proof and give a more rigorous proof.

Algorithms↗

Learning bounds for kernel regression using effective data dimensionality.

Kernel methods can embed finite-dimensional data into infinite-dimensional feature spaces. In spite of the large underlying feature dimensionality, kernel methods can achieve good generalization ability. This observation is often wrongly interpreted, and it has been used to argue that kernel learning can magically avoid the "curse-of-dimensionality" phenomenon encountered in statistical estimation problems. This letter shows that although using kernel representation, one can embed data into an infinite-dimensional feature space; the effective dimensionality of this embedding, which determines the learning complexity of the underlying kernel machine, is usually small. In particular, we introduce an algebraic definition of a scale-sensitive effective dimension associated with a kernel representation. Based on this quantity, we derive upper bounds on the generalization performance of some kernel regression methods. Moreover, we show that the resulting convergent rates are optimal under various circumstances.

Artificial Intelligence↗

Computer-derived nuclear "grade" and breast cancer prognosis.

Visual assessments of nuclear grade are subjective yet still prognostically important. Now, computer-based analytical techniques can objectively and accurately measure size, shape and texture features, which constitute nuclear grade. The cell samples used in this study were obtained by fine needle aspiration (FNA) during the diagnosis of 187 consecutive patients with invasive breast cancer. Regions of FNA preparations to be analyzed were digitized and displayed on a computer monitor. Nuclei to be analyzed were roughly outlined by an operator using a mouse. Next, the computer generated a "snake" that precisely enclosed each designated nucleus. Ten nuclear features were then calculated for each nucleus based on these snakes. These results were analyzed statistically and by an inductive machine learning technique that we developed and call "recurrence surface approximation" (RSA). Both the statistical and RSA machine learning analyses demonstrated that computer-derived nuclear features are prognostically more important than are the classic prognostic features, tumor size and lymph node status.

Adult↗

Integrated unit performance testing of powered, air-purifying particulate respirators using a DOP challenge aerosol.

Although workplace protection factor (WPF) and simulated workplace protection factor (SWPF) studies provide useful information regarding the performance capabilities of powered air-purifying respirators (PAPRs) under certain workplace or simulated workplace conditions, some fail to address the issue of total PAPR unit performance over extended time. PAPR unit performance over time is of paramount importance in protecting worker health over the course of a work shift or at least for the recommended service lifetime of the PAPR battery pack, whichever is shorter. The need for PAPR unit performance testing has become even more important with the inception of 42 CFR 84 and the recent introduction of electrostatic respirator filter media into the PAPR market. This study was conducted to learn how current PAPRs certified by the National Institute for Occupational Safety and Health would perform under an 8-hour unit performance test similar to the dioctyl phthalate (DOP) loading test described in 42 CFR 84 for R- and P-series filters for nonpowered, air-purifying particulate respirators. In this study, entire PAPR units, four with mechanical filters and one with an electrostatic filter, were tested using a TSI Model 8122 Automated Respirator Tester, with and without the built-in breathing machine. The two, tight-fitting PAPRs, both with mechanical filters, showed little effect on performance resulting from the breathing machine. The two loose-fitting helmet PAPRs indicate that unit performance testing without the breathing machine is a more stringent test than testing with the breathing machine under the conditions used. The PAPR with a loose-fitting hood gave inconclusive results as to which testing condition is more stringent. The PAPR unit equipped with electrostatic filters gave the highest maximum penetration values during unit performance testing.

Aerosols↗

Evolving mobile robots in simulated and real environments.

The problem of the validity of simulation is particularly relevant for methodologies that use machine learning techniques to develop control systems for autonomous robots, as, for instance, the artificial life approach known as evolutionary robotics. In fact, although it has been demonstrated that training or evolving robots in real environments is possible, the number of trials needed to test the system discourages the use of physical robots during the training period. By evolving neural controllers for a Khepera robot in computer simulations and then transferring the agents obtained to the real environment we show that (a) an accurate model of a particular robot-environment dynamics can be built by sampling the real world through the sensors and the actuators of the robot; (b) the performance gap between the obtained behaviors in simulated and real environments may be significantly reduced by introducing a "conservative" form of noise; (c) if a decrease in performance is observed when the system is transferred to a real environment, successful and robust results can be obtained by continuing the evolutionary process in the real environment for a few generations.

Algorithms↗

Bayesian Gaussian process classification with the EM-EP algorithm.

Gaussian process classifiers (GPCs) are Bayesian probabilistic kernel classifiers. In GPCs, the probability of belonging to a certain class at an input location is monotonically related to the value of some latent function at that location. Starting from a Gaussian process prior over this latent function, data are used to infer both the posterior over the latent function and the values of hyperparameters to determine various aspects of the function. Recently, the expectation propagation (EP) approach has been proposed to infer the posterior over the latent function. Based on this work, we present an approximate EM algorithm, the EM-EP algorithm, to learn both the latent function and the hyperparameters. This algorithm is found to converge in practice and provides an efficient Bayesian framework for learning hyperparameters of the kernel. A multiclass extension of the EM-EP algorithm for GPCs is also derived. In the experimental results, the EM-EP algorithms are as good or better than other methods for GPCs or Support Vector Machines (SVMs) with cross-validation.

Algorithms↗

Implementation of automated signal generation in pharmacovigilance using a knowledge-based approach.

Automated signal generation is a growing field in pharmacovigilance that relies on data mining of huge spontaneous reporting systems for detecting unknown adverse drug reactions (ADR). Previous implementations of quantitative techniques did not take into account issues related to the medical dictionary for regulatory activities (MedDRA) terminology used for coding ADRs. MedDRA is a first generation terminology lacking formal definitions; grouping of similar medical conditions is not accurate due to taxonomic limitations. Our objective was to build a data-mining tool that improves signal detection algorithms by performing terminological reasoning on MedDRA codes described with the DAML+OIL description logic. We propose the PharmaMiner tool that implements quantitative techniques based on underlying statistical and bayesian models. It is a JAVA application displaying results in tabular format and performing terminological reasoning with the Racer inference engine. The mean frequency of drug-adverse effect associations in the French database was 2.66. Subsumption reasoning based on MedDRA taxonomical hierarchy produced a mean number of occurrence of 2.92 versus 3.63 (p < 0.001) obtained with a combined technique using subsumption and approximate matching reasoning based on the ontological structure. Semantic integration of terminological systems with data mining methods is a promising technique for improving machine learning in medical databases.

Adverse Drug Reaction Reporting Systems↗