Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

A support vector machine using the lazy learning approach for multi-class classification.

Support vector machines can be used in a new machine learning technique based on statistical learning. In this paper, we develop least squares support vector machines (LS-SVMs) using the lazy learning approach to classify data in unclassifiable regions in the case of multi-class classification. LS-SVMs use a set of linear equations while SVMs use a quadratic programming problem. The lazy learning approach is a local and memory-based technique. Therefore, it is an alternative technique to fuzzy inference systems. Our studies show that LS-SVMs with the lazy learning approach can give comparable results to fuzzy LS-SVMs for multi-class classification.

Algorithms↗

[Closed system anaesthesia with continuous inspiratory oxygen measurement (author's transl)].

Relatively simple and inexpensive electrochemical oxygen-sensors have led to clinical evaluation of oxygen and nitrous oxide in completely closed circuits. Thus, a timely solution relating to possibilities of theatre-pollution and the economy this method allows, was obtained. "Fading" is a characteristic of electrochemical sensors; the principal factor involved seems to be the condensation of water vapor on the sensor, attenuating its sensitivity with the passage of time. However, a new model, that of Dräger, showed uniformly less than 1% fading after several hours of closed system anaesthesia, which is clinically negligible. Other failures were unmistakable and immediately recognizable. The techniques for neuroleptanalgesia and halothane anaesthesia with standard anaesthetic machines are described. The dosage of halothane with Vapor vaporizer, the reliability of absorption of CO2, the conditions which may be associated with accumulation of nitrogen up to 10 Vol.-%, as well as the relative merits of different ventilatores were explored. Studies performed with aid of oxygen sensors during clinical closed system anaesthetics are shown in relation to N2O-uptake and N2-elimination. In comparison with conventional methods the understanding of anaesthetic-uptake, the learning process and personal interest are enhanced, the patient's safety being ensured at the same time. The method of using a closed circuit here described was used in fifteen hundred patients ranging from 5-97 years without complications and is now routine at our clinic. Other advantages of the technique for the patient are retention of adequate humidification and heat of inspired gas and an almost complete dissociation of alveolar ventilation and uptake of anaesthetic.

Absorption↗

Learning in higher order Boltzmann machines using linear response.

We introduce an efficient method for learning and inference in higher order Boltzmann machines. The method is based on mean field theory with the linear response correction. We compute the correlations using the exact and the approximated method for a fully connected third order network of ten neurons. In addition, we compare the results of the exact and approximate learning algorithm. Finally we use the presented method to solve the shifter problem. We conclude that the linear response approximation gives good results as long as the couplings are not too large.

Artificial Intelligence↗

The linear separability problem: some testing methods.

The notion of linear separability is used widely in machine learning research. Learning algorithms that use this concept to learn include neural networks (single layer perceptron and recursive deterministic perceptron), and kernel machines (support vector machines). This paper presents an overview of several of the methods for testing linear separability between two classes. The methods are divided into four groups: Those based on linear programming, those based on computational geometry, one based on neural networks, and one based on quadratic programming. The Fisher linear discriminant method is also presented. A section on the quantification of the complexity of classification problems is included.

Algorithms↗

Human-AI Interaction With AI-Assisted Tumor Overlays in Pediatric Whole-Body Magnetic Resonance Imaging: Exploratory Reader Study.

BACKGROUND: AI tools have the potential to enhance personalized clinical care, particularly in radiology. However, their integration into clinical workflows remains complex, especially in pediatric oncology, where early cancer detection is critical. Children with Li-Fraumeni syndrome (LFS), a rare cancer predisposition disorder, undergo regular surveillance whole-body magnetic resonance imaging (wbMRI), which presents an opportunity for AI-assisted tumor detection. OBJECTIVE: We evaluated the feasibility of an AI-assisted overlay for highlighting tumor-like regions in pediatric surveillance wbMRI and explored how access to the overlay influenced radiologist workflow, candidate-lesion marking behavior, follow-up recommendations, and perceived workload. METHODS: We developed a patch-based AI segmentation model trained on augmented 2D slices from 675 surveillance wbMRI volumes of pediatric patients with LFS. The model was designed to highlight regions with high tumor probability. A reader study was conducted with 2 radiologists who independently reviewed wbMRI cases both with and without AI assistance. We measured evaluation time, number and location of reader-marked candidate lesions, type of follow-up recommendation, and subjective feedback using structured questionnaires. RESULTS: AI assistance altered interpretation workflows for both radiologists, with mixed effects. On average, the time required to evaluate each case increased when using the AI tool for both radiologists. However, one radiologist had an increase in the number of candidate lesion locations selected with the tool, and one had a decrease in the number of candidate lesion locations selected with the tool. Subjective feedback indicated that one of the radiologists reported lower mental demand with the AI tool, while both radiologists reported lower stress with the AI tool. Interrater variability was evident, underscoring the need for personalized calibration of AI tools. CONCLUSIONS: AI-assisted wbMRI interpretation can improve tumor detection in pediatric cancer surveillance by reducing false negatives. However, its influence on workflow efficiency and interradiologist variability highlights the importance of careful implementation. Successful integration requires addressing challenges such as improving the predictive precision of AI models, offering intuitive end-user designs and instructions, and building trust in AI outputs. AI outputs can influence workflow and behavior in reader-specific ways. Clinical translation will require larger, randomized, multireader studies and model refinement to reduce false positives and quantify lesion-level reader performance. This can help ensure better patient outcomes in addition to reduced clinician burnout.

Humans↗

Analysis and prediction of leucine-rich nuclear export signals.

We present a thorough analysis of nuclear export signals and a prediction server, which we have made publicly available. The machine learning prediction method is a significant improvement over the generally used consensus patterns. Nuclear export signals (NESs) are extremely important regulators of the subcellular location of proteins. This regulation has an impact on transcription and other nuclear processes, which are fundamental to the viability of the cell. NESs are studied in relation to cancer, the cell cycle, cell differentiation and other important aspects of molecular biology. Our conclusion from this analysis is that the most important properties of NESs are accessibility and flexibility allowing relevant proteins to interact with the signal. Furthermore, we show that not only the known hydrophobic residues are important in defining a nuclear export signals. We employ both neural networks and hidden Markov models in the prediction algorithm and verify the method on the most recently discovered NESs. The NES predictor (NetNES) is made available for general use at http://www.cbs.dtu.dk/.

Active Transport, Cell Nucleus↗

Individuality of handwriting.

Motivated by several rulings in United States courts concerning expert testimony in general, and handwriting testimony in particular, we undertook a study to objectively validate the hypothesis that handwriting is individual. Handwriting samples of 1,500 individuals, representative of the U.S. population with respect to gender, age, ethnic groups, etc., were obtained. Analyzing differences in handwriting was done by using computer algorithms for extracting features from scanned images of handwriting. Attributes characteristic of the handwriting were obtained, e.g., line separation, slant, character shapes, etc. These attributes, which are a subset of attributes used by forensic document examiners (FDEs), were used to quantitatively establish individuality by using machine learning approaches. Using global attributes of handwriting and very few characters in the writing, the ability to determine the writer with a high degree of confidence was established. The work is a step towards providing scientific support for admitting handwriting evidence in court. The mathematical approach and the resulting software also have the promise of aiding the FDE.

Adolescent↗

Predicting thermal displacements in modular tool systems.

In the last decade, there has been an increasing interest in compensating thermally induced errors to improve the manufacturing accuracy of modular tool systems. These modular tool systems are interfaces between spindle and workpiece and consist of several complicatedly formed parts. Their thermal behavior is dominated by nonlinearities, delay and hysteresis effects even in tools with simpler geometry and it is difficult to describe it theoretically. Due to the dominant nonlinear nature of this behavior the so far used linear regression between the temperatures and the displacements is insufficient. Therefore, in this study we test the hypothesis whether we can reliably predict such thermal displacements via nonlinear temperature-displacement regression functions. These functions are estimated first from learning measurements using the alternating conditional expectation (ACE) algorithm and then tested on independent data sets. First, we analyze data that were generated by a finite element spindle model. We find that our approach is a powerful tool to describe the relation between temperatures and displacements for simulated data. Next, we analyze the temperature-displacement relationship in a silent real experimental setup, where the tool system is thermally forced. Again, the ACE algorithm is powerful to estimate the deformation with high precision. The corresponding errors obtained by using the nonlinear regression approach are 10-fold lower in comparison to multiple linear regression analysis. Finally, we investigate the thermal behavior of a modular tool system in a working milling machine and again get promising results. The thermally induced errors can be estimated with 1-2 microm accuracy using this nonlinear regression analysis. Therefore, this approach seems to be very useful for the development of new modular tool systems.

Computer Simulation↗

On the nature of cavities on protein surfaces: application to the identification of drug-binding sites.

In this article we introduce a new method for the identification and the accurate characterization of protein surface cavities. The method is encoded in the program SCREEN (Surface Cavity REcognition and EvaluatioN). As a first test of the utility of our approach we used SCREEN to locate and analyze the surface cavities of a nonredundant set of 99 proteins cocrystallized with drugs. We find that this set of proteins has on average about 14 distinct cavities per protein. In all cases, a drug is bound at one (and sometimes more than one) of these cavities. Using cavity size alone as a criterion for predicting drug-binding sites yields a high balanced error rate of 15.7%, with only 71.7% coverage. Here we characterize each surface cavity by computing a comprehensive set of 408 physicochemical, structural, and geometric attributes. By applying modern machine learning techniques (Random Forests) we were able to develop a classifier that can identify drug-binding cavities with a balanced error rate of 7.2% and coverage of 88.9%. Only 18 of the 408 cavity attributes had a statistically significant role in the prediction. Of these 18 important attributes, almost all involved size and shape rather than physicochemical properties of the surface cavity. The implications of these results are discussed. A SCREEN Web server is available at http://interface.bioc.columbia.edu/screen.

Binding Sites↗

Firemaster 550 differentially alters gene expression underlying synaptic function in amygdala of prairie voles after gestational or lactational exposure.

Neurodevelopmental disorders often share similar behavioral diagnostic criteria including socioemotional and cognitive deficits. The prairie vole is a uniquely suitable model to study these deficits because they demonstrate strong social affiliation, bi-parental care, and partner attachment. Previously, we have shown that developmental exposure to the flame-retardant mixture Firemaster 550 (FM 550) impairs socioemotional behavior in the prairie vole and alters underlying neuroanatomy and function. However, the mechanisms for impaired pair bonding in males and increased anxiety in females remain unknown, along with the specific critical window(s) of vulnerability. Herein, we exposed prairie vole dams to FM 550 during gestation or lactation, and performed bulk RNA-seq on the amygdala, a hub of socioemotional processing, in their adult offspring. Two mathematically orthogonal methods were utilized for analysis, a linear statistical method and an ensemble machine learning method, incorporating sex as a biological variable. Gene ontology (GO) pathway analysis was performed following both and results compared to identify potential mechanisms of toxicity. GO results indicated consistent expression changes in the Synapse cellular component in all conditions, and implicated glutamatergic signaling specifically. Additionally, gestational exposure (GE) altered genes underlying modulation of synaptic transmission and neural development, while lactational exposure (LE) impacted genes underlying synaptic plasticity, axon guidance, and mitophagy. Machine learning identified disruption of endocrine system development, regulation of biosynthetic processes in GE animals, and suppression of various neuroinflammatory genes across multiple groups. Finally, we performed RNA expression analysis using Nanostring and demonstrated stronger correlation with the differentially expressed genes (DEG) of interest in females than males. Overall, this study demonstrates both the intersecting and distinct impacts of FM 550 exposure on amygdalar gene expression depending on sex and timing of exposure.

Animals↗

bioETH-PRS: confidential polygenic risk scoring with smart contracts on an FHE-enabled blockchain.

Polygenic risk scores (PRSs) aggregate genetic effect estimates to predict disease susceptibility, yet calculating one through an external service can require exposing raw genotype data. Homomorphic encryption hides those data during the calculation but, in prior work, still places a designated evaluator in a position of trust. We present bioETH-PRS, a protocol that replaces the evaluator with publicly auditable smart contracts on a blockchain supporting Fully Homomorphic Ethereum Virtual Machine (fhEVM). Using integer-exact encrypted arithmetic, bioETH-PRS computes the PRS dot product entirely in the encrypted domain, so genotype dosages and, at the model provider's discretion, the GWAS weights stay hidden from the parties performing the computation. A fixed-point encoding represents signed weights as nonnegative integers within a bound that rules out overflow, recovering the score to the precision of the published weights. A four-contract architecture separates data custody, model publication, computation, and output release, and supports both a classic path that stores encrypted inputs and an appreciably cheaper streaming path that discards them. A release oracle can return a randomized risk category instead of the raw score, limiting what a repeated querier learns. Prototype evaluation on real GWAS fixtures, including a run on a public testnet, shows cost growing linearly with variant count and suggests the approach may be practical where transaction fees are low. Trust is redistributed rather than removed: the system still depends on the contracts, the blockchain, and the fhEVM services. We evaluate additive models of moderate size, not genome-wide or clinical use.

Blockchain↗

Neurofuzzy adaptive controlling of selective stimulation for FES: a case study.

A controller was designed for the selective stimulation of the sciatic nerve with a multiple contact cuff electrode to generate a desired torque in the ankle joint of cat. The design integrates three approaches, artificial neural network (ANN) modeling, fuzzy logical adaptation, and geometrical mapping. The geometrical mapping refers to the vector transformation from the joint coordinates to the virtual muscle coordinates which have been conceptually developed to represent the major recruitment features of contact-based functional units in the physical plant. This method reduces the complexity of generating a data set for training the neural network in the feedforward path and implementing the on-line learning algorithm embedded in the feedback loop. The controller was evaluated by computer simulation with the experimental data obtained from the torque generation in five acute cats. The results show that the ANN-based feedforward is capable of predicting 65% of a given desired isometric torque, and the fuzzy logical machine is able to provide suitable gains for feedback modulation to reduce the error from 35 to 8.5% and produce a robust control.

Animals↗

Providing QoS through machine-learning-driven adaptive multimedia applications.

We investigate the optimization of the quality of service (QoS) offered by real-time multimedia adaptive applications through machine learning algorithms. These applications are able to adapt in real time their internal settings (i.e., video sizes, audio and video codecs, among others) to the unpredictably changing capacity of the network. Traditional adaptive applications just select a set of settings to consume less than the available bandwidth. We propose a novel approach in which the selected set of settings is the one which offers a better user-perceived QoS among all those combinations which satisfy the bandwidth restrictions. We use a genetic algorithm to decide when to trigger the adaptation process depending on the network conditions (i.e., loss-rate, jitter, etc.). Additionally, the selection of the new set of settings is done according to a set of rules which model the user-perceived QoS. These rules are learned using the SLIPPER rule induction algorithm over a set of examples extracted from scores provided by real users. We will demonstrate that the proposed approach guarantees a good user-perceived QoS even when the network conditions are constantly changing.

Algorithms↗

Evaluation of clustering algorithms for gene expression data.

BACKGROUND: Cluster analysis is an integral part of high dimensional data analysis. In the context of large scale gene expression data, a filtered set of genes are grouped together according to their expression profiles using one of numerous clustering algorithms that exist in the statistics and machine learning literature. A closely related problem is that of selecting a clustering algorithm that is "optimal" in some sense from a rather impressive list of clustering algorithms that currently exist. RESULTS: In this paper, we propose two validation measures each with two parts: one measuring the statistical consistency (stability) of the clusters produced and the other representing their biological functional congruence. Smaller values of these indices indicate better performance for a clustering algorithm. We illustrate this approach using two case studies with publicly available gene expression data sets: one involving a SAGE data of breast cancer patients and the other involving a time course cDNA microarray data on yeast. Six well known clustering algorithms UPGMA, K-Means, Diana, Fanny, Model-Based and SOM were evaluated. CONCLUSION: No single clustering algorithm may be best suited for clustering genes into functional groups via expression profiles for all data sets. The validation measures introduced in this paper can aid in the selection of an optimal algorithm, for a given data set, from a collection of available clustering algorithms.

Algorithms↗

Dimension reduction-based penalized logistic regression for cancer classification using microarray data.

The use of penalized logistic regression for cancer classification using microarray expression data is presented. Two dimension reduction methods are respectively combined with the penalized logistic regression so that both the classification accuracy and computational speed are enhanced. Two other machine-learning methods, support vector machines and least-squares regression, have been chosen for comparison. It is shown that our methods have achieved at least equal or better results. They also have the advantage that the output probability can be explicitly given and the regression coefficients are easier to interpret. Several other aspects, such as the selection of penalty parameters and components, pertinent to the application of our methods for cancer classification are also discussed.

Algorithms↗

The generalized LASSO.

In the last few years, the support vector machine (SVM) method has motivated new interest in kernel regression techniques. Although the SVM has been shown to exhibit excellent generalization properties in many experiments, it suffers from several drawbacks, both of a theoretical and a technical nature: the absence of probabilistic outputs, the restriction to Mercer kernels, and the steep growth of the number of support vectors with increasing size of the training set. In this paper, we present a different class of kernel regressors that effectively overcome the above problems. We call this approach generalized LASSO regression. It has a clear probabilistic interpretation, can handle learning sets that are corrupted by outliers, produces extremely sparse solutions, and is capable of dealing with large-scale problems. For regression functionals which can be modeled as iteratively reweighted least-squares (IRLS) problems, we present a highly efficient algorithm with guaranteed global convergence. This defies a unique framework for sparse regression models in the very rich class of IRLS models, including various types of robust regression models and logistic regression. Performance studies for many standard benchmark datasets effectively demonstrate the advantages of this model over related approaches.

Algorithms↗

Symmetry breaking and training from incomplete data with Radial Basis Boltzmann Machines.

A Radial Basis Boltzmann Machine (RBBM) is a specialized Boltzmann Machine architecture that combines feed-forward mapping with probability estimation in the input space, and for which very efficient learning rules exist. The hidden representation of the network displays symmetry breaking as a function of the noise in the dynamics. Thus, generalization can be studied as a function of the noise in the neuron dynamics instead of as a function of the number of hidden units. We show that the RBBM can be seen as an elegant alternative of k-nearest neighbor, leading to comparable performance without the need to store all data. We show that the RBBM has good classification performance compared to the MLP. The main advantage of the RBBM is that simultaneously with the input-output mapping, a model of the input space is obtained which can be used for learning with missing values. We derive learning rules for the case of incomplete data, and show that they perform better on incomplete data than the traditional learning rules on a 'repaired' data set.

Computer Simulation↗

Data-driven approaches in green microbiology: strategies for plant growth-promoting bacteria.

Plant growth-promoting bacteria (PGPB) are gaining attention as scalable biological solutions to enhance crop productivity and resilience. However, accurately identifying and characterizing PGPB remains challenging, particularly under variable environmental conditions where microbial functions are context-dependent and shaped by complex plant-microbe interactions. Advances in high-throughput sequencing have shifted the field from culture-dependent approaches to genome-informed strategies, enabling large-scale taxonomic and functional profiling. Although trait-based databases support the prediction of plant-beneficial genes, they capture only a fraction of the underlying biological complexity and often require labor-intensive analyses. Machine learning (ML) and deep learning (DL) have emerged as powerful tools to integrate genomic, physiological, and ecological data, enabling the prioritization of candidate strains with plant growth-promoting potential. To evaluate advances in the field, we conducted a systematic review of studies integrating ML and DL with PGPB characterization, assessing algorithm selection, performance, and target plant systems. Across 248 observations, only 6.0% of studies directly addressed PGPB screening, whereas the majority (77.4%) focused on plant disease detection, revealing a substantial gap in the application of AI to beneficial microorganisms for plant growth. Convolutional neural networks (CNNs) were the most frequently applied algorithms, largely driven by image-based phenotyping tasks. Overall, the field is constrained by limited datasets, high computational demands, and challenges in modeling multispecies and host-associated interactions. We highlight the need for integrative and interpretable ML and DL frameworks that bridge genomic data and functional validation. Such approaches represent a promising path toward scalable, data-driven discovery and deployment of bioinoculants in sustainable agriculture.

Agriculture↗