Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Population size and quality in genetics-based rule learning from medical data.

Population size and quality are parameters which control the performance of genetic algorithms. We researched these parameters in a genetic-based machine learning system Galactica which was used to discover the differential diagnostic rules for female urinary incontinence from case data. The performance of the system was measured with on-line and off-line criteria. Surprisingly, randomly generated small populations (30 and 70 rules) did not promote the best on-line performance as earlier results suggested. Probable explanation is the lack of diversity in initial populations. The seeding of population with positive learning examples was used to obtain more divergent populations. As expected, the seeding increased the on-line performance of small populations. The results are mainly in accord with the earlier results indicating that large randomly generated populations (150 rules) lead to the better off-line performance. Again, the seeding of small populations was successful producing even the better off-line performance than a large population. In conclusion, the seeding allowed the small populations to converge to good rules in relatively short period of time.

Algorithms↗

A neural network architecture for data classification.

This article aims at showing an architecture of neural networks designed for the classification of data distributed among a high number of classes. A significant gain in the global classification rate can be obtained by using our architecture. This latter is based on a set of several little neural networks, each one discriminating only two classes. The specialization of each neural network simplifies their structure and improves the classification. Moreover, the learning step automatically determines the number of hidden neurons. The discussion is illustrated by tests on databases from the UCI machine learning database repository. The experimental results show that this architecture can achieve a faster learning, simpler neural networks and an improved performance in classification.

Computer Systems↗

On classification capability of neural networks: a case study with otoneurological data.

We investigated the capability of multilayer perceptron neural networks and Kohonen neural networks to recognize difficult otoneurological diseases from each other. We found that they are efficient methods, but the distribution of a learning set should be rather uniform. Also it is important that the number of learning cases is sufficient. If the two mentioned conditions are satisfied, these neural networks are similarly efficient as some other machine learning methods. The conditions are known in the theory of neural networks [1,2], but not often taken seriously in practice. Both networks functioned as well, excluding the case with several input variables, where the Kohonen neural networks surpassed the perceptron.

Algorithms↗

Computer simulation of FES standing up in paraplegia: a self-adaptive fuzzy controller with reinforcement learning.

Using computer simulation, the theoretical feasibility of functional electrical stimulation (FES) assisted standing up is demonstrated using a closed-loop self-adaptive fuzzy logic controller based on reinforcement machine learning (FLC-RL). The control goal was to minimize upper limb forces and the terminal velocity of the knee joint. The reinforcement learning (RL) technique was extended to multicontroller problems in continuous state and action spaces. The validated algorithms were used to synthesize FES controllers for the knee and hip joints in simulated paraplegic standing up. The FLC-RL controller was able to achieve the maneuver with only 22% of the upper limb force required to stand-up without FES and to simultaneously reduce the terminal velocity of the knee joint close to zero. The FLC-RL controller demonstrated, as expected, the closed loop fuzzy logic control and on-line self-adaptation capability of the RL was able to accommodate for simulated disturbances due to voluntary arm forces, FES induced muscle fatigue and anthropometric differences between individuals. A method of incorporating a priori heuristic rule based knowledge is described that could reduce the number of the learning trials required to establish a usable control strategy. We also discuss how such heuristics may also be incorporated into the initial FLC-RL controller to ensure safe operation from the onset.

Arm↗

Learning Petri net models of non-linear gene interactions.

Understanding how an individual's genetic make-up influences their risk of disease is a problem of paramount importance. Although machine-learning techniques are able to uncover the relationships between genotype and disease, the problem of automatically building the best biochemical model or "explanation" of the relationship has received less attention. In this paper, I describe a method based on random hill climbing that automatically builds Petri net models of non-linear (or multi-factorial) disease-causing gene-gene interactions. Petri nets are a suitable formalism for this problem, because they are used to model concurrent, dynamic processes analogous to biochemical reaction networks. I show that this method is routinely able to identify perfect Petri net models for three disease-causing gene-gene interactions recently reported in the literature.

Algorithms↗

Learning optimized features for hierarchical models of invariant object recognition.

There is an ongoing debate over the capabilities of hierarchical neural feedforward architectures for performing real-world invariant object recognition. Although a variety of hierarchical models exists, appropriate supervised and unsupervised learning methods are still an issue of intense research. We propose a feedforward model for recognition that shares components like weight sharing, pooling stages, and competitive nonlinearities with earlier approaches but focuses on new methods for learning optimal feature-detecting cells in intermediate stages of the hierarchical network. We show that principles of sparse coding, which were previously mostly applied to the initial feature detection stages, can also be employed to obtain optimized intermediate complex features. We suggest a new approach to optimize the learning of sparse features under the constraints of a weight-sharing or convolutional architecture that uses pooling operations to achieve gradual invariance in the feature hierarchy. The approach explicitly enforces symmetry constraints like translation invariance on the feature set. This leads to a dimension reduction in the search space of optimal features and allows determining more efficiently the basis representatives, which achieve a sparse decomposition of the input. We analyze the quality of the learned feature representation by investigating the recognition performance of the resulting hierarchical network on object and face databases. We show that a hierarchy with features learned on a single object data set can also be applied to face recognition without parameter changes and is competitive with other recent machine learning recognition approaches. To investigate the effect of the interplay between sparse coding and processing nonlinearities, we also consider alternative feedforward pooling nonlinearities such as presynaptic maximum selection and sum-of-squares integration. The comparison shows that a combination of strong competitive nonlinearities with sparse coding offers the best recognition performance in the difficult scenario of segmentation-free recognition in cluttered surround. We demonstrate that for both learning and recognition, a precise segmentation of the objects is not necessary.

Learning↗

Boolean matrix logic programming for active learning of gene functions in genome-scale metabolic network models.

Reasoning about hypotheses and updating knowledge through empirical observations are central to scientific discovery. In this work, we applied logic-based machine learning methods to drive biological discovery by guiding experimentation. Genome-scale metabolic network models (GEMs) - comprehensive representations of metabolic genes and reactions - are widely used to evaluate genetic engineering of biological systems. However, GEMs often fail to accurately predict the behaviour of genetically engineered cells, primarily due to incomplete annotations of gene interactions. The task of learning the intricate genetic interactions within GEMs presents computational and empirical challenges. To efficiently predict using GEM, we describe a novel approach called Boolean Matrix Logic Programming (BMLP) by leveraging Boolean matrices to evaluate large logic programs. We developed a new system, [Formula: see text], which guides cost-effective experimentation and uses interpretable logic programs to encode a state-of-the-art GEM of a model bacterial organism. Notably, [Formula: see text] successfully learned the interaction between a gene pair with fewer training examples than random experimentation, overcoming the increase in experimental design space. [Formula: see text] enables rapid optimisation of metabolic models to reliably engineer biological systems for producing useful compounds. It offers a realistic approach to creating a self-driving lab for biological discovery, which would then facilitate microbial engineering for practical applications.

Active learning↗

Deep learning-based multimodal pathogenomics integration for precision cancer prognosis.

BACKGROUND: Recent studies have revealed valuable prognostic insights in haematoxylin and eosin (H&E)-stained histological sections and transcriptomic profiles, suggesting potential applications in machine learning. However, existing methods lack sufficient intra- and inter-modal interactions, and face challenges in clinical validation due to incomplete multimodal data. METHODS: We proposed PathoGems (PathoGenomics-based integrative survival prediction), a weakly-supervised, interpretable multimodal learning framework that integrates histology and genomic profiles for precise cancer prognosis prediction. To evaluate the robustness of PathoGems, we initially curated a dataset of 1965 cases across four cohorts from The Cancer Genome Atlas (TCGA), including breast, colorectal, glioblastoma, and esophageal cancers. For external validation, PathoGems was further evaluated on four independent cohorts, consisting of 76 breast cancer and 41 esophageal squamous cell carcinoma cases from Zhejiang Cancer Hospital, as well as 102 colorectal cancer and 58 glioblastoma cases from the Clinical Proteomic Tumor Analysis Consortium (CPTAC). RESULTS: PathoGems effectively stratified patients into favorable and unfavorable risk groups, revealing significant differences in histological patterns, genomic features, and overall survival (log-rank test, p&#x2009;<&#x2009;0.05). Moreover, the model&#x2019;s predictions are further supported by visualization and transcriptomic analysis, enhancing interpretability and reliability. CONCLUSIONS: By fusing histological and clinicogenomic multimodal models, PathoGems will provide a solid foundation for developing an innovative tool that aids clinicians in making informed decisions and selection personalized treatment strategies for cancer patients.

Humans↗

Spectral imaging perspective on cytomics.

BACKGROUND: Cytomics involves the analysis of cellular morphology and molecular phenotypes, with reference to tissue architecture and to additional metadata. To this end, a variety of imaging and nonimaging technologies need to be integrated. Spectral imaging is proposed as a tool that can simplify and enrich the extraction of morphological and molecular information. Simple-to-use instrumentation is available that mounts on standard microscopes and can generate spectral image datasets with excellent spatial and spectral resolution; these can be exploited by sophisticated analysis tools. METHODS: This report focuses on brightfield microscopy-based approaches. Cytological and histological samples were stained using nonspecific standard stains (Giemsa; hematoxylin and eosin (H&E)) or immunohistochemical (IHC) techniques employing three chromogens plus a hematoxylin counterstain. The samples were imaged using the Nuance system, a commercially available, liquid-crystal tunable-filter-based multispectral imaging platform. The resulting data sets were analyzed using spectral unmixing algorithms and/or learn-by-example classification tools. RESULTS: Spectral unmixing of Giemsa-stained guinea-pig blood films readily classified the major blood elements. Machine-learning classifiers were also successful at the same task, as well in distinguishing normal from malignant regions in a colon-cancer example, and in delineating regions of inflammation in an H&E-stained kidney sample. In an example of a multiplexed ICH sample, brown, red, and blue chromogens were isolated into separate images without crosstalk or interference from the (also blue) hematoxylin counterstain. CONCLUSION: Cytomics requires both accurate architectural segmentation as well as multiplexed molecular imaging to associate molecular phenotypes with relevant cellular and tissue compartments. Multispectral imaging can assist in both these tasks, and conveys new utility to brightfield-based microscopy approaches.

Animals↗

Learning deterministic finite automata with a smart state labeling evolutionary algorithm.

Learning a Deterministic Finite Automaton (DFA) from a training set of labeled strings is a hard task that has been much studied within the machine learning community. It is equivalent to learning a regular language by example and has applications in language modeling. In this paper, we describe a novel evolutionary method for learning DFA that evolves only the transition matrix and uses a simple deterministic procedure to optimally assign state labels. We compare its performance with the Evidence Driven State Merging (EDSM) algorithm, one of the most powerful known DFA learning algorithms. We present results on random DFA induction problems of varying target size and training set density. We also studythe effects of noisy training data on the evolutionary approach and on EDSM. On noise-free data, we find that our evolutionary method outperforms EDSM on small sparse data sets. In the case of noisy training data, we find that our evolutionary method consistently outperforms EDSM, as well as other significant methods submitted to two recent competitions.

Algorithms↗

The use of misclassification costs to learn rule-based decision support models for cost-effective hospital admission strategies.

Cost-effective health care is at the forefront of today's important health-related issues. A research team at the University of Pittsburgh has been interested in lowering the cost of medical care by attempting to define a subset of patients with community-acquire pneumonia for whom outpatient therapy is appropriate and safe. Sensitivity and specificity requirements for this domain make it difficult to use rule-based learning algorithms with standard measures of performance based on accuracy. This paper describes the use of misclassification costs to assist a rule-based machine-learning program in deriving a decision-support aid for choosing outpatient therapy for patients with community-acquired pneumonia.

Algorithms↗

Protein names precisely peeled off free text.

MOTIVATION: Automatically identifying protein names from the scientific literature is a pre-requisite for the increasing demand in data-mining this wealth of information. Existing approaches are based on dictionaries, rules and machine-learning. Here, we introduced a novel system that combines a pre-processing dictionary- and rule-based filtering step with several separately trained support vector machines (SVMs) to identify protein names in the MEDLINE abstracts. RESULTS: Our new tagging-system NLProt is capable of extracting protein names with a precision (accuracy) of 75% at a recall (coverage) of 76% after training on a corpus, which was used before by other groups and contains 200 annotated abstracts. For our estimate of sustained performance, we considered partially identified names as false positives. One important issue frequently ignored in the literature is the redundancy in evaluation sets. We suggested some guidelines for removing overly inadequate overlaps between training and testing sets. Applying these new guidelines, our program appeared to significantly out-perform other methods tagging protein names. NLProt was so successful due to the SVM-building blocks that succeeded in utilizing the local context of protein names in the scientific literature. We challenge that our system may constitute the most general and precise method for tagging protein names. AVAILABILITY: http://cubic.bioc.columbia.edu/services/nlprot/

Abstracting and Indexing↗

Learning about protein hydrogen bonding by minimizing contrastive divergence.

Defining the strength and geometry of hydrogen bonds in protein structures has been a challenging task since early days of structural biology. In this article, we apply a novel statistical machine learning technique, known as contrastive divergence, to efficiently estimate both the hydrogen bond strength and the geometric characteristics of strong interpeptide backbone hydrogen bonds, from a dataset of structures representing a variety of different protein folds. Despite the simplifying assumptions of the interatomic energy terms used, we determine the strength of these hydrogen bonds to be between 1.1 and 1.5 kcal/mol, in good agreement with earlier experimental estimates. The geometry of these strong backbone hydrogen bonds features an almost linear arrangement of all four atoms involved in hydrogen bond formation. We estimate that about a quarter of all hydrogen bond donors and acceptors participate in these strong interpeptide hydrogen bonds.

Amino Acids↗

LSAT: learning about alternative transcripts in MEDLINE.

MOTIVATION: Generation of alternative transcripts from the same gene is an important biological event due to their contribution in creating functional diversity in eukaryotes. In this work, we choose the task of extracting information around this complex topic using a two-step procedure involving machine learning and information extraction. RESULTS: In the first step, we trained a classifier that inductively learns to identify sentences about physiological transcript diversity from the MEDLINE abstracts. Using a large hand-built corpus, we compared the sentence classification performance of various text categorization methods. Support vector machines (SVMs) followed by the maximum entropy classifier outperformed other methods for the sentence classification task. The SVM with the radial basis function kernel and optimized parameters achieved Fbeta-measure of 91% during the 4-fold cross validation and of 74% when applied to all sentences in more than 12 million abstracts of MEDLINE. In the second step, we identified eight frequently present semantic categories in the sentences and performed a limited amount of semantic role labeling. The role labeling step also achieved very high Fbeta-measure for all eight categories. AVAILABILITY: The results of our two-step procedure are summarized in the LSAT database of alternative transcripts. LSAT is available at http://www.bork.embl.de/LSAT CONTACT: shah@embl.de SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Benchmarking↗

Systematic identification of statistically significant network measures.

We present a graph embedding space (i.e., a set of measures on graphs) for performing statistical analyses of networks. Key improvements over existing approaches include discovery of "motif hubs" (multiple overlapping significant subgraphs), computational efficiency relative to subgraph census, and flexibility (the method is easily generalizable to weighted and signed graphs). The embedding space is based on scalars, functionals of the adjacency matrix representing the network. Scalars are global, involving all nodes; although they can be related to subgraph enumeration, there is not a one-to-one mapping between scalars and subgraphs. Improvements in network randomization and significance testing--we learn the distribution rather than assuming Gaussianity--are also presented. The resulting algorithm establishes a systematic approach to the identification of the most significant scalars and suggests machine-learning techniques for network classification.

Algorithms↗

A hybrid SOM-SVM approach for the zebrafish gene expression analysis.

Microarray technology can be employed to quantitatively measure the expression of thousands of genes in a single experiment. It has become one of the main tools for global gene expression analysis in molecular biology research in recent years. The large amount of expression data generated by this technology makes the study of certain complex biological problems possible, and machine learning methods are expected to play a crucial role in the analysis process. In this paper, we present our results from integrating the self-organizing map (SOM) and the support vector machine (SVM) for the analysis of the various functions of zebrafish genes based on their expression. The most distinctive characteristic of our zebrafish gene expression is that the number of samples of different classes is imbalanced. We discuss how SOM can be used as a data-filtering tool to improve the classification performance of the SVM on this data set.

Animals↗

Predicting the insurgence of human genetic diseases associated to single point protein mutations with support vector machines and evolutionary information.

MOTIVATION: Human single nucleotide polymorphisms (SNPs) are the most frequent type of genetic variation in human population. One of the most important goals of SNP projects is to understand which human genotype variations are related to Mendelian and complex diseases. Great interest is focused on non-synonymous coding SNPs (nsSNPs) that are responsible of protein single point mutation. nsSNPs can be neutral or disease associated. It is known that the mutation of only one residue in a protein sequence can be related to a number of pathological conditions of dramatic social impact such as Alzheimer's, Parkinson's and Creutzfeldt-Jakob's diseases. The quality and completeness of presently available SNPs databases allows the application of machine learning techniques to predict the insurgence of human diseases due to single point protein mutation starting from the protein sequence. RESULTS: In this paper, we develop a method based on support vector machines (SVMs) that starting from the protein sequence information can predict whether a new phenotype derived from a nsSNP can be related to a genetic disease in humans. Using a dataset of 21 185 single point mutations, 61% of which are disease-related, out of 3587 proteins, we show that our predictor can reach more than 74% accuracy in the specific task of predicting whether a single point mutation can be disease related or not. Our method, although based on less information, outperforms other web-available predictors implementing different approaches. AVAILABILITY: A beta version of the web tool is available at http://gpcr.biocomp.unibo.it/cgi/predictors/PhD-SNP/PhD-SNP.cgi

Algorithms↗

Learning systems in biosignal analysis.

In biosignal analysis, the utility of artificial neural networks (ANN) in classifying electromyographic (EMG) data trained with the momentum back propagation algorithm has recently been demonstrated. In the current study, the self-organizing feature map algorithm, the genetics-based machine learning (GBML) paradigm, and the K-means nearest neighbour clustering algorithm are applied on the same set of data. The aim of this exercise is to show how these three paradigms can be used in practice, given that their diagnostic performance is problem- and parameter-dependent. A total of 720 macro EMG recordings were carried out from four groups, from seven normal, nine motor neuron disease, 14 Becker's muscular dystrophy, and six spinal muscular atrophy subjects, respectively. Twenty-three of the subjects were used for training and 13 for evaluating the various models. For each subject, the mean and the standard deviation of the parameters (i) amplitude, (ii) area, (iii) average power and (iv) duration were extracted. The feature vector was structured in two different ways for input to the models: an eight-input feature vector that consisted of both the mean and the standard deviation of the four parameters measured, and a four-input feature vector that included only the mean of the parameters. Also, due to the heterogenous nature of the spinal muscular atrophy group, three class models that excluded this group were investigated. In general, self-organizing feature map and GBML models resulted in comparable diagnostic performance of the order of 80-90% correct classifications (CCs) score for the evaluation set, whereas the K-means nearest neighbour algorithm models gave lower percentage CCs. Furthermore, for all three learning paradigms: better diagnostic performance was obtained for the three class models compared with the four class models; similar diagnostic performance was obtained for both the eight- and four-input feature vectors. Finally, it is claimed that the proposed methodology followed in this work can be applied for the development of diagnostic systems in the analysis of biosignals.

Algorithms↗