Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Bio-medical entity extraction using support vector machines.

OBJECTIVE: Support vector machines (SVMs) have achieved state-of-the-art performance in several classification tasks. In this article we apply them to the identification and semantic annotation of scientific and technical terminology in the domain of molecular biology. This illustrates the extensibility of the traditional named entity task to special domains with large-scale terminologies such as those in medicine and related disciplines. METHODS AND MATERIALS: The foundation for the model is a sample of text annotated by a domain expert according to an ontology of concepts, properties and relations. The model then learns to annotate unseen terms in new texts and contexts. The results can be used for a variety of intelligent language processing applications. We illustrate SVMs capabilities using a sample of 100 journal abstracts texts taken from the {human, blood cell, transcription factor} domain of MEDLINE. RESULTS: Approximately 3400 terms are annotated and the model performs at about 74% F-score on cross-validation tests. A detailed analysis based on empirical evidence shows the contribution of various feature sets to performance. CONCLUSION: Our experiments indicate a relationship between feature window size and the amount of training data and that a combination of surface words, orthographic features and head noun features achieve the best performance among the feature sets tested.

Algorithms↗

Grounding words in perception and action: computational insights.

We use words to communicate about things and kinds of things, their properties, relations and actions. Researchers are now creating robotic and simulated systems that ground language in machine perception and action, mirroring human abilities. A new kind of computational model is emerging from this work that bridges the symbolic realm of language with the physical realm of real-world referents. It explains aspects of context-dependent shifts of word meaning that cannot easily be explained by purely symbolic models. An exciting implication for cognitive modeling is the use of grounded systems to 'step into the shoes' of humans by directly processing first-person-perspective sensory data, providing a new methodology for testing various hypotheses of situated communication and learning.

Cognition↗

Design and construction of a long-term continuous video-EEG monitoring unit for simultaneous recording of multiple small animals.

In recent years several new rat models of human limbic/mesial temporal lobe epilepsy have been described [1,2,4-7,11,15-17]. Unlike earlier models such as kindling in which the seizures are induced by an exogenous stimulus, these new models are characterized by seizures that occur spontaneously at random intervals. Although the spontaneity of the seizures makes these models more like human epilepsy, documentation of these seizures by direct observation is highly inefficient, and sub-behavioral electrographic seizures could be missed. Continuous paper EEG and video recording have been used [5-7,15], but these techniques are resource intensive. The slow paper speed required by long-term paper recordings limits the ability to differentiate between true seizure activity and electrical artifact. Subtle behavioral seizures are likely to be missed during rapid review of video recordings alone [16]. Ambulatory cassette EEG recordings have been used [3], but the systems require expensive proprietary hardware, and the systems have limited channels for recording (8-16). To improve the utility of the models, we developed a long-term EEG/video monitoring system to detect the electrographic seizures and document their behavioral accompaniment. The system is based on commercially available components, including a computerized EEG seizure detection system that was initially developed for human seizure monitoring [8,9,13]. Seizures are reliably detected and the data are reduced so that 24 h of recording can be reviewed in 30-90 min. Although the computer program is accurate, special care must be taken in system design and construction to reduce sources of electrical artifact that can cause false detections when multiple animals are recorded simultaneously on a single EEG machine. During data review it is necessary to differentiate between electrical artifact induced by animal activity from true seizure activity by key EEG patterns. Certain seizure patterns (less than 3 hz. low amplitude) will not be detected by the seizure detection program, but the system is highly effective for typical limbic seizures and may be useful for the animal models of absence epilepsy [12,14]. It can also be used as a continuous or intermittent EEG/physiological recording device for experiments that examine animals' spontaneous behavior and the EEG correlate (e.g. sleep/wake cycles, learning and memory tasks).

Animals↗

Melanie II--a third-generation software package for analysis of two-dimensional electrophoresis images: II. Algorithms.

After two generations of software systems for the analysis of two-dimensional electrophoresis (2-DE) images, a third generation of such software packages has recently emerged that combines state-of-the-art graphical user interfaces with comprehensive spot data analysis capabilities. A key characteristic common to most of these software packages is that many of their tools are implementations of algorithms that resulted from research areas such as image processing, vision, artificial intelligence or machine learning. This article presents the main algorithms implemented in the Melanie II 2-D PAGE software package. The applications of these algorithms, embodied as the feature of the program, are explained in an accompanying article (R. D. Appel et al.; Electrophoresis 1997, 18, 2724-2734).

Algorithms↗

An ENSEMBLE machine learning approach for the prediction of all-alpha membrane proteins.

MOTIVATION: All-alpha membrane proteins constitute a functionally relevant subset of the whole proteome. Their content ranges from about 10 to 30% of the cell proteins, based on sequence comparison and specific predictive methods. Due to the paucity of membrane proteins solved with atomic resolution, the training/testing sets of predictive methods for protein topography and topology routinely include very few well-solved structures mixed with a hundred proteins known with low resolution. Moreover, available predictors fail in predicting recently crystallised membrane proteins (Chen et al., 2002). Presently the number of well-solved membrane proteins comprises some 59 chains of low sequence homology. It is therefore possible to train/test predictors only with the set of proteins known with atomic resolution and evaluate more thoroughly the performance of different methods. RESULTS: We implement a cascade-neural network (NN), two different hidden Markov models (HMM), and their ensemble (ENSEMBLE) as a new method. We train and test in cross validation the three methods and ENSEMBLE on the 59 well resolved membrane proteins. ENSEMBLE scores with a per-protein accuracy of 90% for topography and 71% for topology, outperforming the best single method of 7 and 5 percentage points, respectively. When tested on a low resolution set of 151 proteins, with no homology with the 59 proteins, the per-protein accuracy of ENSEMBLE is 76% for topography and 68% for topology. Our results also indicate that the performance of ENSEMBLE is higher than that of the best predictors presently available on the Web.

Algorithms↗

Support vector analysis of color-Doppler images: a new approach for estimating indices of left ventricular function.

Reliable noninvasive estimators of global left ventricular (LV) chamber function remain unavailable. We have previously demonstrated a potential relationship between color-Doppler M-mode (CDMM) images and two basic indices of LV function: peak-systolic elastance (Emax) and the time-constant of LV relaxation (tau). Thus, we hypothesized that these two indices could be estimated noninvasively by adequate postprocessing of CDMM recordings. A semiparametric regression (SR) version of support vector machine (SVM) is here proposed for building a blind model, capable of analyzing CDMM images automatically, as well as complementary clinical information. Simultaneous invasive and Doppler tracings were obtained in nine mini-pigs in a high-fidelity experimental setup. The model was developed using a test and validation leave-one-out design. Reasonably acceptable prediction accuracy was obtained for both Emax (intraclass correlation coefficient Ric, = 0.81) and tau (Ric, = 0.61). For the first time, a quantitative, noninvasive estimation of cardiovascular indices is addressed by processing Doppler-echocardiography recordings using a learning-from-samples method.

Algorithms↗

The challenge of computer mathematics.

Progress in the foundations of mathematics has made it possible to formulate all thinkable mathematical concepts, algorithms and proofs in one language and in an impeccable way. This is not in spite of, but partially based on the famous results of Gödel and Turing. In this way statements are about mathematical objects and algorithms, proofs show the correctness of statements and computations, and computations are dealing with objects and proofs. Interactive computer systems for a full integration of defining, computing and proving are based on this. The human defines concepts, constructs algorithms and provides proofs, while the machine checks that the definitions are well formed and the proofs and computations are correct. Results formalized so far demonstrate the feasibility of this 'computer mathematics'. Also there are very good applications. The challenge is to make the systems more mathematician-friendly, by building libraries and tools. The eventual goal is to help humans to learn, develop, communicate, referee and apply mathematics.

Algorithms↗

Predicting essential genes in fungal genomes.

Essential genes are required for an organism's viability, and the ability to identify these genes in pathogens is crucial to directed drug development. Predicting essential genes through computational methods is appealing because it circumvents expensive and difficult experimental screens. Most such prediction is based on homology mapping to experimentally verified essential genes in model organisms. We present here a different approach, one that relies exclusively on sequence features of a gene to estimate essentiality and offers a promising way to identify essential genes in unstudied or uncultured organisms. We identified 14 characteristic sequence features potentially associated with essentiality, such as localization signals, codon adaptation, GC content, and overall hydrophobicity. Using the well-characterized baker's yeast Saccharomyces cerevisiae, we employed a simple Bayesian framework to measure the correlation of each of these features with essentiality. We then employed the 14 features to learn the parameters of a machine learning classifier capable of predicting essential genes. We trained our classifier on known essential genes in S. cerevisiae and applied it to the closely related and relatively unstudied yeast Saccharomyces mikatae. We assessed predictive success in two ways: First, we compared all of our predictions with those generated by homology mapping between these two species. Second, we verified a subset of our predictions with eight in vivo knockouts in S. mikatae, and we present here the first experimentally confirmed essential genes in this species.

Computational Biology↗

Biomarkers that discriminate multiple myeloma patients with or without skeletal involvement detected using SELDI-TOF mass spectrometry and statistical and machine learning tools.

Multiple Myeloma (MM) is a severely debilitating neoplastic disease of B cell origin, with the primary source of morbidity and mortality associated with unrestrained bone destruction. Surface enhanced laser desorption/ionization time-of-flight mass spectrometry (SELDI-TOF MS) was used to screen for potential biomarkers indicative of skeletal involvement in patients with MM. Serum samples from 48 MM patients, 24 with more than three bone lesions and 24 with no evidence of bone lesions were fractionated and analyzed in duplicate using copper ion loaded immobilized metal affinity SELDI chip arrays. The spectra obtained were compiled, normalized, and mass peaks with mass-to-charge ratios (m/z) between 2000 and 20,000 Da identified. Peak information from all fractions was combined together and analyzed using univariate statistics, as well as a linear, partial least squares discriminant analysis (PLS-DA), and a non-linear, random forest (RF), classification algorithm. The PLS-DA model resulted in prediction accuracy between 96-100%, while the RF model was able to achieve a specificity and sensitivity of 87.5% each. Both models as well as multiple comparison adjusted univariate analysis identified a set of four peaks that were the most discriminating between the two groups of patients and hold promise as potential biomarkers for future diagnostic and/or therapeutic purposes.

Adult↗

Approaches to the automatic discovery of patterns in biosequences.

This paper surveys approaches to the discovery of patterns in biosequences and places these approaches within a formal framework that systematises the types of patterns and the discovery algorithms. Patterns with expressive power in the class of regular languages are considered, and a classification of pattern languages in this class is developed, covering the patterns that are the most frequently used in molecular bioinformatics. A formulation is given of the problem of the automatic discovery of such patterns from a set of sequences, and an analysis is presented of the ways in which an assessment can be made of the significance of the discovered patterns. It is shown that the problem is related to problems studied in the field of machine learning. The major part of this paper comprises a review of a number of existing methods developed to solve the problem and how these relate to each other, focusing on the algorithms underlying the approaches. A comparison is given of the algorithms, and examples are given of patterns that have been discovered using the different methods.

Algorithms↗

Induction of medical expert system rules based on rough sets and resampling methods.

Automated knowledge acquisition is an important research issue in improving the efficiency of medical expert systems. Rules for medical expert systems consists of two parts: one is a proposition part, which represent a if-then rule, and the other is probabilistic measures, which represents reliability of that rule. Therefore, acquisition of both knowledge is very important for application of machine learning methods to medical domains. Extending concepts of rough set theory to probabilistic domain, we introduce a new approach to knowledge acquisition, which induces probabilistic rules based on rough set theory (PRIMEROSE) and develop a program that extracts rules for an expert system from clinical database, using this method. The results show that the derived rules almost correspond to those of medical experts.

Artificial Intelligence↗

Metric learning for text documents.

Many algorithms in machine learning rely on being given a good distance metric over the input space. Rather than using a default metric such as the Euclidean metric, it is desirable to obtain a metric based on the provided data. We consider the problem of learning a Riemannian metric associated with a given differentiable manifold and a set of points. Our approach to the problem involves choosing a metric from a parametric family that is based on maximizing the inverse volume of a given data set of points. From a statistical perspective, it is related to maximum likelihood under a model that assigns probabilities inversely proportional to the Riemannian volume element. We discuss in detail learning a metric on the multinomial simplex where the metric candidates are pull-back metrics of the Fisher information under a Lie group of transformations. When applied to text document classification the resulting geodesic distance resemble, but outperform, the tfidf cosine similarity measure.

Algorithms↗

Using feature generation and feature selection for accurate prediction of translation initiation sites.

Correct prediction of the translation initiation site (TIS) is an important issue in genomic research. We show that feature generation together with correlation based feature selection can be used with a variety of machine learning algorithms to give highly accurate translation initiation site prediction. Only very few features are needed and the results achieve comparable accuracy to the best existing approaches. Our approach has the advantage that it does not require one to devise a special prediction method; rather standard machine learning classifiers are shown to give very good performance on the selected features. The raw and generated features which we have found to be important are the following: positions -3 and -1 in the sequence; upstream k-grams for k=3, 4, and 5; stop-codon frequency; downstream in-frame 3-gram; and the distance of ATG to the beginning of the sequence. The best result, with an overall accuracy of 90%, is obtained by selecting only seven features from this set. The same features retrained with the use of a scanning model achieves an overall accuracy of 94% on this dataset.

Codon, Initiator↗

Automated Deep Learning-Based Detection of Early Atherosclerotic Plaques in Carotid Ultrasound Imaging.

BACKGROUND: Carotid plaque presence is associated with cardiovascular risk, even among asymptomatic individuals. While deep learning has shown promise for carotid plaque phenotyping in patients with advanced atherosclerosis, its application in population-based settings of asymptomatic individuals remains unexplored. METHODS: We developed a YOLOv8-based model for plaque detection using carotid ultrasound images from 19,499 participants of the population-based UK Biobank (UKB) and fine-tuned it for external validation in the BiDirect study (N = 2,105). Cox regression was used to estimate the impact of plaque presence and count on major cardiovascular events. To explore the genetic architecture of carotid atherosclerosis, we conducted a genome-wide association study (GWAS) meta-analysis of the UKB and CHARGE cohorts. Mendelian randomization (MR) assessed the effect of genetic predisposition to vascular risk factors on carotid atherosclerosis. RESULTS: Our model demonstrated high performance with accuracy, sensitivity, and specificity exceeding 85%, enabling identification of carotid plaques in 45% of the UKB population (aged 47-83 years). In the external BiDirect cohort, a fine-tuned model achieved 86% accuracy, 78% sensitivity, and 90% specificity. Plaque presence and count were associated with risk of major adverse cardiovascular events (MACE) over a follow-up of up to seven years, improving risk reclassification beyond the Pooled Cohort Equations. A GWAS meta-analysis of carotid plaques uncovered two novel genomic loci, with downstream analyses implicating targets of investigational drugs in advanced clinical development. Observational and MR analyses showed associations between smoking, LDL cholesterol, hypertension, and odds of carotid atherosclerosis. CONCLUSIONS: Our model offers a scalable solution for early carotid plaque detection, potentially enabling automated screening in asymptomatic individuals and improving plaque phenotyping in population-based cohorts. This approach could advance large-scale atherosclerosis research.

atherosclerosis↗

Mining mass spectra for diagnosis and biomarker discovery of cerebral accidents.

In this paper we try to identify potential biomarkers for early stroke diagnosis using surface-enhanced laser desorption/ionization mass spectrometry coupled with analysis tools from machine learning and data mining. Data consist of 42 specimen samples, i.e., mass spectra divided in two big categories, stroke and control specimens. Among the stroke specimens two further categories exist that correspond to ischemic and hemorrhagic stroke; in this paper we limit our data analysis to discriminating between control and stroke specimens. We performed two suites of experiments. In the first one we simply applied a number of different machine learning algorithms; in the second one we have chosen the best performing algorithm as it was determined from the first phase and coupled it with a number of different feature selection methods. The reason for this was 2-fold, first to establish whether feature selection can indeed improve performance, which in our case it did not seem to confirm, but more importantly to acquire a small list of potentially interesting biomarkers. Of the different methods explored the most promising one was support vector machines which gave us high levels of sensitivity and specificity. Finally, by analyzing the models constructed by support vector machines we produced a small set of 13 features that could be used as potential biomarkers, and which exhibited good performance both in terms of sensitivity, specificity and model stability.

Adult↗

Learning systems in biosignal analysis.

In biosignal analysis, the utility of artificial neural networks (ANN) in classifying electromyographic (EMG) data trained with the momentum back propagation algorithm has recently been demonstrated. In the current study, the self-organizing feature map algorithm, the genetics-based machine learning (GBML) paradigm, and the K-means nearest neighbour clustering algorithm are applied on the same set of data. The aim of this exercise is to show how these three paradigms can be used in practice, given that their diagnostic performance is problem- and parameter-dependent. A total of 720 macro EMG recordings were carried out from four groups, from seven normal, nine motor neuron disease, 14 Becker's muscular dystrophy, and six spinal muscular atrophy subjects, respectively. Twenty-three of the subjects were used for training and 13 for evaluating the various models. For each subject, the mean and the standard deviation of the parameters (i) amplitude, (ii) area, (iii) average power and (iv) duration were extracted. The feature vector was structured in two different ways for input to the models: an eight-input feature vector that consisted of both the mean and the standard deviation of the four parameters measured, and a four-input feature vector that included only the mean of the parameters. Also, due to the heterogenous nature of the spinal muscular atrophy group, three class models that excluded this group were investigated. In general, self-organizing feature map and GBML models resulted in comparable diagnostic performance of the order of 80-90% correct classifications (CCs) score for the evaluation set, whereas the K-means nearest neighbour algorithm models gave lower percentage CCs. Furthermore, for all three learning paradigms: better diagnostic performance was obtained for the three class models compared with the four class models; similar diagnostic performance was obtained for both the eight- and four-input feature vectors. Finally, it is claimed that the proposed methodology followed in this work can be applied for the development of diagnostic systems in the analysis of biosignals.

Algorithms↗

Confidence-based active learning.

This paper proposes a new active learning approach, confidence-based active learning, for training a wide range of classifiers. This approach is based on identifying and annotating uncertain samples. The uncertainty value of each sample is measured by its conditional error. The approach takes advantage of current classifiers' probability preserving and ordering properties. It calibrates the output scores of classifiers to conditional error. Thus, it can estimate the uncertainty value for each input sample according to its output score from a classifier and select only samples with uncertainty value above a user-defined threshold. Even though we cannot guarantee the optimality of the proposed approach, we find it to provide good performance. Compared with existing methods, this approach is robust without additional computational effort. A new active learning method for support vector machines (SVMs) is implemented following this approach. A dynamic bin width allocation method is proposed to accurately estimate sample conditional error and this method adapts to the underlying probabilities. The effectiveness of the proposed approach is demonstrated using synthetic and real data sets and its performance is compared with the widely used least certain active learning method.

Algorithms↗

Support vector machine based training of multilayer feedforward neural networks as optimized by particle swarm algorithm: application in QSAR studies of bioactivity of organic compounds.

Multilayer feedforward neural networks (MLFNNs) are important modeling techniques widely used in QSAR studies for their ability to represent nonlinear relationships between descriptors and activity. However, the problems of overfitting and premature convergence to local optima still pose great challenges in the practice of MLFNNs. To circumvent these problems, a support vector machine (SVM) based training algorithm for MLFNNs has been developed with the incorporation of particle swarm optimization (PSO). The introduction of the SVM based training mechanism imparts the developed algorithm with inherent capacity for combating the overfitting problem. Moreover, with the implementation of PSO for searching the optimal network weights, the SVM based learning algorithm shows relatively high efficiency in converging to the optima. The proposed algorithm has been evaluated using the Hansch data set. Application to QSAR studies of the activity of COX-2 inhibitors is also demonstrated. The results reveal that this technique provides superior performance to backpropagation (BP) and PSO training neural networks.

Algorithms↗