Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

A knowledge model for the interpretation and visualization of NLP-parsed discharged summaries.

At our institution, a Natural Language Processing (NLP) tool called MedLEE is used on a daily basis to parse medical texts including complete discharge summaries. MedLEE transforms written text into a generic structured format, which preserves the richness of the underlying natural language expressions by the use of concept modifiers (like change, certainty, degree and status). As a tradeoff, extraction of application-specific medical information is difficult without a clear understanding of how these modifiers combine. We report on a knowledge model for MedLEE modifiers that is helpful for a high level interpretation of NLP data and is used for the generation of two distinct views on NLP-parsed discharge summaries: A physician view offering a condensed overview of the severity of patient problems and a data mining view featuring binary problem states useful for machine learning.

Artificial Intelligence↗

DNA splice site detection: a comparison of specific and general methods.

In an era when whole organism genomes are being routinely sequenced, the problem of gene finding has become a key issue on the road to understanding. For eukaryotic organisms a large part of locating the genes is accomplished by predicting the likely location of splice sites on a DNA strand. This problem of splice site location has been ap- proached using a number of machine learning or statistical methods tailored more or less specifically to the nature of the problem. Recently large margin classifiers and boosting methods have been found to give improvements over more traditional methods in a number of areas. Here we compare large margin classifiers (SVM and CMLS) and boosted decision trees with the three most common models used for splice site detection (WMM, WAM, and MDT). We find that the newer methods compare favorably in all cases and can yield significant improvement in some cases.

Algorithms↗

Machine learning approaches to lung cancer prediction from mass spectra.

We addressed the problem of discriminating between 24 diseased and 17 healthy specimens on the basis of protein mass spectra. To prepare the data, we performed mass to charge ratio (m/z) normalization, baseline elimination, and conversion of absolute peak height measures to height ratios. After preprocessing, the major difficulty encountered was the extremely large number of variables (1676 m/z values) versus the number of examples (41). Dimensionality reduction was treated as an integral part of the classification process; variable selection was coupled with model construction in a single ten-fold cross-validation loop. We explored different experimental setups involving two peak height representations, two variable selection methods, and six induction algorithms, all on both the original 1676-mass data set and on a prescreened 124-mass data set. Highest predictive accuracies (1-2 off-sample misclassifications) were achieved by a multilayer perceptron and Naïve Bayes, with the latter displaying more consistent performance (hence greater reliability) over varying experimental conditions. We attempted to identify the most discriminant peaks (proteins) on the basis of scores assigned by the two variable selection methods and by neural network based sensitivity analysis. These three scoring schemes consistently ranked four peaks as the most relevant discriminators: 11683, 1403, 17350 and 66107.

Algorithms↗

GRUMB: a genome-resolved metagenomic framework for monitoring urban microbiomes and diagnosing pathogen risk.

SUMMARY: Urban infrastructure hosts dynamic microbial communities that complicate biosurveillance and AMR monitoring. Existing tools rarely combine genome-resolved reconstruction with ecological modeling and batch-aware analytics tailored to infrastructure-scale studies. We present GRUMB (Genome-Resolved Urban Microbiome Biosurveillance), an open-source, SLURM-compatible pipeline that reconstructs high-quality metagenome-assembled genomes (MAGs) from shotgun sequencing reads and integrates taxonomic/functional annotation (CARD, VFDB), batch-aware normalization, ecological diagnostics and machine learning classification of environment types with uncertainty and risk scoring. GRUMB accepts either SRA project accessions or paired-end FASTQ files with metadata, and produces assemblies, MAGs, taxonomic and functional profiles, ecological outputs and risk-informed classification. Its modular design enables reproducible, infrastructure-scale biosurveillance across diverse environments. AVAILABILITY AND IMPLEMENTATION: GRUMB is freely available under the MIT License at: https://github.com/SuleimanAminu/genome-resolved-urban-microbiome-biosurveillance; Zenodo DOI: https://doi.org/10.5281/zenodo.15505402. Requirements: Linux (Ubuntu 20.04+), Python 3.11, R 4.2+, SLURM. Issues and feature requests are tracked on GitHub.

Microbiota↗

Optical coherence tomography machine learning classifiers for glaucoma detection: a preliminary study.

PURPOSE: Machine-learning classifiers are trained computerized systems with the ability to detect the relationship between multiple input parameters and a diagnosis. The present study investigated whether the use of machine-learning classifiers improves optical coherence tomography (OCT) glaucoma detection. METHODS: Forty-seven patients with glaucoma (47 eyes) and 42 healthy subjects (42 eyes) were included in this cross-sectional study. Of the glaucoma patients, 27 had early disease (visual field mean deviation [MD] > or = -6 dB) and 20 had advanced glaucoma (MD < -6 dB). Machine-learning classifiers were trained to discriminate between glaucomatous and healthy eyes using parameters derived from OCT output. The classifiers were trained with all 38 parameters as well as with only 8 parameters that correlated best with the visual field MD. Five classifiers were tested: linear discriminant analysis, support vector machine, recursive partitioning and regression tree, generalized linear model, and generalized additive model. For the last two classifiers, a backward feature selection was used to find the minimal number of parameters that resulted in the best and most simple prediction. The cross-validated receiver operating characteristic (ROC) curve and accuracies were calculated. RESULTS: The largest area under the ROC curve (AROC) for glaucoma detection was achieved with the support vector machine using eight parameters (0.981). The sensitivity at 80% and 95% specificity was 97.9% and 92.5%, respectively. This classifier also performed best when judged by cross-validated accuracy (0.966). The best classification between early glaucoma and advanced glaucoma was obtained with the generalized additive model using only three parameters (AROC = 0.854). CONCLUSIONS: Automated machine classifiers of OCT data might be useful for enhancing the utility of this technology for detecting glaucomatous abnormality.

Adult↗

Posterior probability support vector machines for unbalanced data.

This paper proposes a complete framework of posterior probability support vector machines (PPSVMs) for weighted training samples using modified concepts of risks, linear separability, margin, and optimal hyperplane. Within this framework, a new optimization problem for unbalanced classification problems is formulated and a new concept of support vectors established. Furthermore, a soft PPSVM with an interpretable parameter v is obtained which is similar to the v-SVM developed by Schölkopf et al., and an empirical method for determining the posterior probability is proposed as a new approach to determine v. The main advantage of an PPSVM classifier lies in that fact that it is closer to the Bayes optimal without knowing the distributions. To validate the proposed method, two synthetic classification examples are used to illustrate the logical correctness of PPSVMs and their relationship to regular SVMs and Bayesian methods. Several other classification experiments are conducted to demonstrate that the performance of PPSVMs is better than regular SVMs in some cases. Compared with fuzzy support vector machines (FSVMs), the proposed PPSVM is a natural and an analytical extension of regular SVMs based on the statistical learning theory.

Algorithms↗

Functional discrimination of gene expression patterns in terms of the gene ontology.

The ever-growing amount of experimental data in molecular biology and genetics requires its automated analysis, by employing sophisticated knowledge discovery tools. We use an Inductive Logic Programming (ILP) learner to induce functional discrimination rules between genes studied using microarrays and found to be differentially expressed in three recently discovered subtypes of adenocarcinoma of the lung. The discrimination rules involve functional annotations from the Proteome HumanPSD database in terms of the Gene Ontology, whose hierarchical structure is essential for this task. While most of the lower levels of gene expression data (pre)processing have been automated, our work can be seen as a step toward automating the higher level functional analysis of the data. We view our application not just as a prototypical example of applying more sophisticated machine learning techniques to the functional analysis of genes, but also as an incentive for developing increasingly more sophisticated functional annotations and ontologies, that can be automatically processed by such learning algorithms.

Adenocarcinoma↗

Mismatch string kernels for discriminative protein classification.

MOTIVATION: Classification of proteins sequences into functional and structural families based on sequence homology is a central problem in computational biology. Discriminative supervised machine learning approaches provide good performance, but simplicity and computational efficiency of training and prediction are also important concerns. RESULTS: We introduce a class of string kernels, called mismatch kernels, for use with support vector machines (SVMs) in a discriminative approach to the problem of protein classification and remote homology detection. These kernels measure sequence similarity based on shared occurrences of fixed-length patterns in the data, allowing for mutations between patterns. Thus, the kernels provide a biologically well-motivated way to compare protein sequences without relying on family-based generative models such as hidden Markov models. We compute the kernels efficiently using a mismatch tree data structure, allowing us to calculate the contributions of all patterns occurring in the data in one pass while traversing the tree. When used with an SVM, the kernels enable fast prediction on test sequences. We report experiments on two benchmark SCOP datasets, where we show that the mismatch kernel used with an SVM classifier performs competitively with state-of-the-art methods for homology detection, particularly when very few training examples are available. Examination of the highest-weighted patterns learned by the SVM classifier recovers biologically important motifs in protein families and superfamilies.

Algorithms↗

Neural network mosaic model for pupillary responses to spatial stimuli.

A neural network mosaic model was developed to investigate the spatial-temporal properties of the human pupillary control system. It was based on the double-layer neural network model developed by Cannon and Robinson and the pupillary dual-path model developed by Sun and Stark. The neural network portion of the model received its input from a sensor array and consisted of a retina-like two-dimensional neuronal layer. The dual-path portion of the model was composed of interconnections of the neurons that formed a mosaic of AC transient and DC sustained paths. The spatial aggregates of the AC and DC signals were input to the AC and DC summing neurons, respectively. Finally, the weighted sum of the aggregate AC and DC signals provided the output for driving the pupillary response. An important property of the model was that it could adaptively learn from training samples by adjustment of the weights. The neural network mosaic model showed excellent performance in simulating both the traditional pupillary phenomena and the new spatial stimulation findings such as responses to change in stimulus pattern and shift of light spot. Moreover, the model could also be used for the diagnosis of clinical deficits and image processing in machine vision.

Brain Stem↗

Automated expert multiexponential biomodeling interactively over the Internet.

DIMSUM, an acronym for DIMension of a SUM of exponentials, is a highly automated expert system for fitting multiexponential models of increasing dimension to time series data. Up to now, a researcher has needed an individual copy of DIMSUM on his or her own computer as well as support to learn how to use it. W3DIMSUM, a new implementation of DIMSUM, is web-based, new territory for interactive biomodeling, allowing interactive multiexponential model building and model discrimination over the Internet. The algorithms used are numerically intensive, so we have implemented a distributed system, with numerical processing done on our server. Only the user interface is run on the client machine, but users can load and save data and results on their machines, facilitated by our use of Java WebStart.

Algorithms↗

A teaching and research simulator for therapeutic embolization.

A teaching machine that simulates intravascular conditions found during a human therapeutic embolization has been constructed. A submersible pump drives fluid through a circuit of tubing. One limb of a Y (a vascular bifurcation) located in this circuit leads to a model arteriovenous malformation. Catheters placed in this limb may introduce embolic materials by various techniques, and those techniques may be learned and practiced under safe and stress-free conditions. Loss of an embolus into the other limb which supposedly leads to normal tissues is caught and displayed by a sieve, providing immediate feedback that a technique error has occurred.

Embolization, Therapeutic↗

High-volume hemofiltration in septic shock.

In the past decade we have learned a lot about the pathophysiology of septic shock. A lot of experimental research has been performed in vitro and in vivo, showing that hemofiltration can improve hemodynamics and survival. With modern machines, hemofiltration is becoming a sepsis treatment in patients.

Animals↗

Probabilistic finite-state machines--part I.

Probabilistic finite-state machines are used today in a variety of areas in pattern recognition, or in fields to which pattern recognition is linked: computational linguistics, machine learning, time series analysis, circuit testing, computational biology, speech recognition, and machine translation are some of them. In Part I of this paper, we survey these generative objects and study their definitions and properties. In Part II, we will study the relation of probabilistic finite-state automata with other well-known devices that generate strings as hidden Markov models and n-grams and provide theorems, algorithms, and properties that represent a current state of the art of these objects.

Algorithms↗

Learning machines applied to potential forest distribution.

The clearing of forests to obtain land for pasture and agriculture and the replacement of autochthonous species by other faster-growing varieties of trees for timber have both led to the loss of vast areas of forest worldwide. At present, many developed countries are attempting to reverse these effects, establishing policies for the restoration of older woodland systems. Reforestation is a complex matter, planned and carried out by experts who need objective information regarding the type of forest that can be sustained in each area. This information is obtained by drawing up feasibility models constructed using statistical methods that make use of the information provided by morphological and environmental variables (height, gradient, rainfall, etc.) that partially condition the presence or absence of a specific kind of forestation in an area. The aim of this work is to construct a set of feasibility models for woodland located in the basin of the River Liébana (NW Spain), to serve as a support tool for the experts entrusted with carrying out the reforestation project. The techniques used are multilayer perceptron neural networks and support vector machines. Their results will be compared to the results obtained by traditional techniques (such as discriminant analysis and logistic regression) by measuring the degree of fit between each model and the existing distribution of woodlands. The interpretation and problems of the feasibility models are commented on in the Discussion section.

Artificial Intelligence↗

Long-term depression as a memory process in the cerebellum.

When details of neuronal network structures of the cerebellum were uncovered in the 1960's, a hope emerged that functions of the cerebellum would eventually be explained in terms of operation of the cerebellar neuronal network. While various network models were proposed, involvement of synaptic plasticity in the cerebellar neuronal network as a memory process became a focus of discussion. The characteristic dual inputs to Purkinje cells, one from parallel fibers (axons of granule cells) and the other from climbing fibers, were suggested to represent such synaptic plasticity, and under this assumption, the cerebellar cortex was envisaged as a learning machine for pattern recognition. Despite these theoretical suggestions, earlier efforts to reveal the postulated synaptic plasticity in the cerebellar cortex were unsuccessful. It had then to wait for a decade before long-term depression (LTD) was finally found as its possible substrate. LTD is a long-lasting depression of parallel fiber-to-Purkinje cell transmission that occurs following conjunctive activation of parallel fibers and a climbing fiber both converging onto one and the same Purkinje cell. LTD has now been established by means of various testing methods, and recent efforts have been directed toward its molecular mechanisms. Efforts have also been devoted to demonstrate roles of LTD in motor learning through studies of adaptation of the vestibulo-ocular reflex, adaptive adjustment of hand movement, and more recently eyelid blink conditioned reflex. This article reviews recent efforts to characterize the LTD as a memory process, presumably the major, in the cerebellum.

Animals↗

Health. Care. Anywhere. Today.

What if clinical quality medical equipment were available to every consumer in a form factor that was inexpensive, accurate, and easy to use? What if this equipment provided information that previously was un-measurable or very difficult to measure? What if the physiological state of individuals, at resolutions measured in thousandths of a second instead of in visits per year, could be measured easily, making it possible to ascertain caloric intake and expenditure, patterns of sleep, contextual activities such as working-out and driving, even parameters of mental state and health. What aspect of healthcare would not change? We present a system that is available today that enables this vision. This award-winning multi-channel wearable physiological monitor has enabled the collection of more than 90 million minutes of data in natural settings from thousands of subjects engaged in diverse activities. Data modeling efforts are resulting in applications that present meaningful and actionable information in real-time to users and their designated collaborators (physicians, family members, counselors, coaches, etc.) We describe the SenseWear system, its design, and a summary of validation studies, current commercial applications, and ongoing research. This discussion will show how the convergence of design for wearability, advances in machine learning, and improvements in wireless technology will manifest the future of health care as personal, ubiquitous, and collaborative.

Clothing↗

Medical diagnostic system using Fuzzy Coloured Petri Nets under uncertainty.

We propose a medical diagnostic system using Fuzzy Coloured Petri Nets (FCPN) in this paper. For complex real-world knowledge Fuzzy Petri Net (FPN) models have been proposed to perform fuzzy reasoning automatically. However, in the Petri Net we have to represent all kinds of processes by separate subnets even though the process has the same behavior of other one. Real-world knowledge often contains many parts which are similar, but not identical. This means that the total PTN becomes very large. The kind of problems may be annoying for a small system, and it may be catastrophic for the description of large-scale system. To avoid this kind of problems we propose a learning and reasoning method using FCPNs under uncertainty. On the other hand to correct the rules of knowledge-based system hand-built classifier and empirical learning method both based on domain theory have been proposed as machine learning methods, where there is a significant gap between the knowledge-intensive approach in the former and the virtually knowledge-free approach in the later. To resolve such problems simultaneously we propose a hybrid learning method which is built on the top of knowledge-based FCPN and Genetic Algorithms (GA). To verify the validity and the effectiveness of the proposed system, we have successfully applied it to the diagnosis of intervertebral diseases.

Algorithms↗

CAKR: commutative algebra k-mer representations for genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer representations as a nonlinear algebraic framework for analyzing genomic sequences. This representation bridges commutative algebra, algebraic topology, combinatorics, and machine learning to establish a mathematical framework for comparative genomic analysis. We evaluate its effectiveness on three tasks including genetic variant classification, phylogenetic tree reconstruction, and viral classification, typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. In this work, we show that commutative algebra k-mer representations outperform five state-of-the-art sequence analysis methods across twelve primary datasets, with two additional supplementary fragment-placement benchmarks, especially in viral classification, and maintain relatively stable predictive accuracy as dataset size increases, underscoring scalability and robustness.

Genomics↗