Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Ensemble learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

[Vector coding and neuronal maps].

The model of vector coding is proposed. The excitation vector generated in a neuronal ensemble simultaneously acts on a map of selective detectors (selectors) producing a local excitation maximum that represents the input stimulus. Vector coding is suggested also to explain associative learning and memory. The output responses in the model are specified by excitation vectors triggered by the command neurons in the ensembles of premotor neurons.

Animals↗

The Gene Set Builder: collation, curation, and distribution of sets of genes.

BACKGROUND: In bioinformatics and genomics, there are many applications designed to investigate the common properties for a set of genes. Often, these multi-gene analysis tools attempt to reveal sequential, functional, and expressional ties. However, while tremendous effort has been invested in developing tools that can analyze a set of genes, minimal effort has been invested in developing tools that can help researchers compile, store, and annotate gene sets in the first place. As a result, the process of making or accessing a set often involves tedious and time consuming steps such as finding identifiers for each individual gene. These steps are often repeated extensively to shift from one identifier type to another; or to recreate a published set. In this paper, we present a simple online tool which - with the help of the gene catalogs Ensembl and GeneLynx - can help researchers build and annotate sets of genes quickly and easily. DESCRIPTION: The Gene Set Builder is a database-driven, web-based tool designed to help researchers compile, store, export, and share sets of genes. This application supports the 17 eukaryotic genomes found in version 32 of the Ensembl database, which includes species from yeast to human. User-created information such as sets and customized annotations are stored to facilitate easy access. Gene sets stored in the system can be "exported" in a variety of output formats - as lists of identifiers, in tables, or as sequences. In addition, gene sets can be "shared" with specific users to facilitate collaborations or fully released to provide access to published results. The application also features a Perl API (Application Programming Interface) for direct connectivity to custom analysis tools. A downloadable Quick Reference guide and an online tutorial are available to help new users learn its functionalities. CONCLUSION: The Gene Set Builder is an Ensembl-facilitated online tool designed to help researchers compile and manage sets of genes in a user-friendly environment. The application can be accessed via http://www.cisreg.ca/gsb/.

Computational Biology↗

Brain systems and long-term memory.

This paper focuses mainly on those findings derived from lesion studies on the rat which help to identify ensembles of neural structures concerned with the expression of previously learned responses. At the outset, the use of the lesion method in the search for those neurological circuits underlying memory is defended. This is followed by an evaluation of neocortical and subcortical systems in long-term memory. Subsequently, a modest list of tentative functional neural "complexes" involved in the maintenance of certain classes of learned responses is given, based largely upon the author's own research. It is concluded that the key to the understanding of the neurological substrates of long-term memory lies in the identification of those subcortical sites which interact with neocortical sites in the performance of complex learned tasks. The most likely subcortical sites involved in this interaction appear to inhabit the regions of the basal ganglia, limbic midbrain area, and ventral portions of the brainstem reticular formation.

Animals↗

Bagging to improve the accuracy of a clustering procedure.

MOTIVATION: The microarray technology is increasingly being applied in biological and medical research to address a wide range of problems such as the classification of tumors. An important statistical question associated with tumor classification is the identification of new tumor classes using gene expression profiles. Essential aspects of this clustering problem include identifying accurate partitions of the tumor samples into clusters and assessing the confidence of cluster assignments for individual samples. RESULTS: Two new resampling methods, inspired from bagging in prediction, are proposed to improve and assess the accuracy of a given clustering procedure. In these ensemble methods, a partitioning clustering procedure is applied to bootstrap learning sets and the resulting multiple partitions are combined by voting or the creation of a new dissimilarity matrix. As in prediction, the motivation behind bagging is to reduce variability in the partitioning results via averaging. The performances of the new and existing methods were compared using simulated data and gene expression data from two recently published cancer microarray studies. The bagged clustering procedures were in general at least as accurate and often substantially more accurate than a single application of the partitioning clustering procedure. A valuable by-product of bagged clustering are the cluster votes which can be used to assess the confidence of cluster assignments for individual observations. SUPPLEMENTARY INFORMATION: For supplementary information on datasets, analyses, and software, consult http://www.stat.berkeley.edu/~sandrine and http://www.bioconductor.org.

Algorithms↗

Exploring bias in the Protein Data Bank using contrast classifiers.

In this study we analyzed the bias existing in the Protein Data Bank (PDB) using the novel contrast classifier approach. We trained an ensemble of neural network classifiers, called a contrast classifier, to learn the distributional differences between non-redundant sequence subsets of PDB and SWISS-PROT. Assuming that SWISS-PROT is a representative of the sequence diversity in nature while the PDB is a biased sample, output of the contrast classifier can be used to measure whether the properties of a given sequence or its region are underrepresented in PDB. We applied the contrast classifier to SWISS-PROT sequences to analyze the bias in PDB towards different functional protein properties. The results showed that transmembrane, signal, disordered, and low complexity regions are significantly underrepresented in PDB, while disulfide bonds, metal binding sites, and sites involved in enzyme activity are overrepresented. Additionally, hydroxylation and phosphorylation posttranslational modification sites were found to be underrepresented while acetylation sites were significantly overrepresented. These results suggest the potential usefulness of contrast classifiers in the selection of target proteins for structural characterization experiments.

Bias↗

Sequential-context-dependent hippocampal activity is not necessary to learn sequences with repeated elements.

Learning sequences of events (e.g., a-b-c) is conceptually a simple problem that can be solved using asymmetrically linked cell assemblies [e.g., "phase sequences" (Hebb, 1949)], provided that the elements of the sequence are unique. When elements repeat within the sequence, however (e.g., a-b-c-d-b-e), the same element belongs to two separate "contexts," and a more complex sequence encoding mechanism is required to differentiate between the two contexts. Some neural structure must form sequential-context-dependent, or "differential," representations of the two contexts (i.e., b as an element of "a-b-c" as opposed to "d-b-e") to allow the correct choice to be made after the repeated element. To investigate the possible role of hippocampus in complex sequence encoding, rats were trained to remember repeated-location sequences under three conditions: (1) reward was given at each location; (2) during training, moveable barriers were placed at the entry and exit of the repeated segment to direct the rat and were removed once the sequence was learned; and (3) reward was withheld at the entry and exit of the repeated segment. In the first condition, hippocampal ensemble activity did not differentiate the sequential context of the repeated segment, indicating that complex sequences with repeated segments can be learned without differential encoding within the hippocampus. Differential hippocampal encoding was observed, however, under the latter two conditions, suggesting that long-term memory for discriminative cues present only during training, working memory of the most recently visited reinforcement sites, or anticipation of the subsequent reinforcement site can separate hippocampal activity patterns at the same location.

Acoustic Stimulation↗

Hippocampal replay contributes to within session learning in a temporal difference reinforcement learning model.

Temporal difference reinforcement learning (TDRL) algorithms, hypothesized to partially explain basal ganglia functionality, learn more slowly than real animals. Modified TDRL algorithms (e.g. the Dyna-Q family) learn faster than standard TDRL by practicing experienced sequences offline. We suggest that the replay phenomenon, in which ensembles of hippocampal neurons replay previously experienced firing sequences during subsequent rest and sleep, may provide practice sequences to improve the speed of TDRL learning, even within a single session. We test the plausibility of this hypothesis in a computational model of a multiple-T choice-task. Rats show two learning rates on this task: a fast decrease in errors and a slow development of a stereotyped path. Adding developing replay to the model accelerates learning the correct path, but slows down the stereotyping of that path. These models provide testable predictions relating the effects of hippocampal inactivation as well as hippocampal replay on this task.

Algorithms↗

ESMpHLA: Evolutionary Scale Model-Based Deep Learning Prediction of HLA Class I Binding Peptides.

The recognition of endogenous peptides by HLA class I plays a crucial role in CD8+ T cell immune responses and human adaptive cell immune. Thus, the prediction of HLA class I-peptide binding affinities is always the core issue for the research of immune recognition and vaccine development. In this study, an evolutionary scale model (ESM) combined with parallel CNN blocks and a cross attention mechanism was used to construct a novel ESMpHLA model for predicting HLA class I binding peptides. Based on the 91,560 binding peptides of 41 HLA-A alleles, 56,731 of 50 HLA-B alleles and 2444 of 10 HLA-C alleles, the ESMpHLA model was successfully established and achieved satisfying prediction performances with the overall accuracy and AUC values of 0.874 and 0.938 for the test dataset. The results indicate that the ESMpHLA model performs well in dealing with different HLA class I 2-field alleles as well as the peptides with different lengths. Then, the generalisation ability of the ESMpHLA model was validated by an independent test dataset compiled from recent IEDB weekly benchmark datasets. The results showed that the ESMpHLA model achieved the highest ROC-AUC and PR-AUC values when compared with the latest BVMHC, CapsNet-MHC, STMHCpan and BVLSTM models. In addition, two ensemble models were also established by integrating the above 5 deep learning models using soft-voting and hard-voting strategies.

Humans↗

A role for protein kinase C in associative learning.

Recent work suggests that protein kinase C (PKC), an enzyme that has a critical role in the regulation of cell growth and differentiation, also participates in the sequence of molecular events that underlie learning and memory. By means of electrophysiological, biochemical, and neuro-imaging methods it has been demonstrated that, in the brain, the distribution of PKC changes as a result of memory storage. The changes in distribution occur within the same ensembles of nerve cells that are necessary for the acquisition and performance of various learning tasks in several species. Here we review the data pertaining to a model that has been proposed to account for the participation of PKC as a molecular signal for cotemporal synaptic input during associative learning.

Alzheimer Disease↗

Neuronal networks and synaptic plasticity: understanding complex system dynamics by interfacing neurons with silicon technologies.

Information processing in the central nervous system is primarily mediated through synaptic connections between neurons. This connectivity in turn defines how large ensembles of neurons may coordinate network output to execute complex sensory and motor functions including learning and memory. The synaptic connectivity between any given pair of neurons is not hard-wired; rather it exhibits a high degree of plasticity, which in turn forms the basis for learning and memory. While there has been extensive research to define the cellular and molecular basis of synaptic plasticity, at the level of either pairs of neurons or smaller networks, analysis of larger neuronal ensembles has proved technically challenging. The ability to monitor the activities of larger neuronal networks simultaneously and non-invasively is a necessary prerequisite to understanding how neuronal networks function at the systems level. Here we describe recent breakthroughs in the area of various bionic hybrids whereby neuronal networks have been successfully interfaced with silicon devices to monitor the output of synaptically connected neurons. These technologies hold tremendous potential for future research not only in the area of synaptic plasticity but also for the development of strategies that will enable implantation of electronic devices in live animals during various memory tasks.

Animals↗

Cognitive navigation based on nonuniform Gabor space sampling, unsupervised growing networks, and reinforcement learning.

We study spatial learning and navigation for autonomous agents. A state space representation is constructed by unsupervised Hebbian learning during exploration. As a result of learning, a representation of the continuous two-dimensional (2-D) manifold in the high-dimensional input space is found. The representation consists of a population of localized overlapping place fields covering the 2-D space densely and uniformly. This space coding is comparable to the representation provided by hippocampal place cells in rats. Place fields are learned by extracting spatio-temporal properties of the environment from sensory inputs. The visual scene is modeled using the responses of modified Gabor filters placed at the nodes of a sparse Log-polar graph. Visual sensory aliasing is eliminated by taking into account self-motion signals via path integration. This solves the hidden state problem and provides a suitable representation for applying reinforcement learning in continuous space for action selection. A temporal-difference prediction scheme is used to learn sensorimotor mappings to perform goal-oriented navigation. Population vector coding is employed to interpret ensemble neural activity. The model is validated on a mobile Khepera miniature robot.

Cognition↗

A study on several machine-learning methods for classification of malignant and benign clustered microcalcifications.

In this paper, we investigate several state-of-the-art machine-learning methods for automated classification of clustered microcalcifications (MCs). The classifier is part of a computer-aided diagnosis (CADx) scheme that is aimed to assisting radiologists in making more accurate diagnoses of breast cancer on mammograms. The methods we considered were: support vector machine (SVM), kernel Fisher discriminant (KFD), relevance vector machine (RVM), and committee machines (ensemble averaging and AdaBoost), of which most have been developed recently in statistical learning theory. We formulated differentiation of malignant from benign MCs as a supervised learning problem, and applied these learning methods to develop the classification algorithm. As input, these methods used image features automatically extracted from clustered MCs. We tested these methods using a database of 697 clinical mammograms from 386 cases, which included a wide spectrum of difficult-to-classify cases. We analyzed the distribution of the cases in this database using the multidimensional scaling technique, which reveals that in the feature space the malignant cases are not trivially separable from the benign ones. We used receiver operating characteristic (ROC) analysis to evaluate and to compare classification performance by the different methods. In addition, we also investigated how to combine information from multiple-view mammograms of the same case so that the best decision can be made by a classifier. In our experiments, the kernel-based methods (i.e., SVM, KFD, and RVM) yielded the best performance (Az = 0.85, SVM), significantly outperforming a well-established, clinically-proven CADx approach that is based on neural network (Az = 0.80).

Algorithms↗

Learning generative models of natural images.

This work proposes an unsupervised learning process for analysis of natural images. The derivation is based on a generative model, a stochastic coin-flip process directly operating on many disjoint multivariate Gaussian distributions. Following the maximal likelihood principle and using the Potts encoding, the goodness-of-fit of the generative model to tremendous patches randomly sampled from natural images is quantitatively expressed by an objective function subject to a set of constraints. By further combination of the objective function and the minimal wiring criterion, we achieve a mixed integer and linear programming. A hybrid of the mean field annealing and the gradient descent method is applied to the mathematical framework and produces three sets of interactive dynamics for the learning process. Numerical simulations show that the learning process is effective for extraction of orientation, localization and bandpass features and the generative model can make an ensemble of a sparse code for natural images.

Learning↗

Neural representations of location outside the hippocampus.

Place cells of the rat hippocampus are a dominant model system for understanding the role of the hippocampus in learning and memory at the level of single-unit and neural ensemble responses. A complete understanding of the information processing and computations performed by the hippocampus requires detailed knowledge about the properties of the representations that are present in hippocampal afferents and efferents in order to decipher the transformations that occur to these representations in the hippocampal circuitry. Neural recordings in behaving rats have revealed a number of brain areas that contain place-related firing properties in the parahippocampal regions and in other brain regions that are thought to interact with the hippocampus in certain behavioral tasks. Although investigators have just begun to scratch the surface in terms of understanding these properties, differences in the precise nature of the spatial firing between the hippocampus and these other regions promise to reveal important clues regarding the exact role of the hippocampus in learning and memory and the nature of its interactions with other brain systems to support adaptive behavior.

Animals↗

Improved classification of medical data using abductive network committees trained on different feature subsets.

This paper demonstrates the use of abductive network classifier committees trained on different features for improving classification accuracy in medical diagnosis. In an earlier publication, committee members were trained on different subsets of the training set to ensure enough diversity for improved committee performance. In situations characterized by high data dimensionality, i.e. a large number of features and a relatively few training examples, it may be more advantageous to split the feature set rather than the training set. We describe a novel approach for tentatively ranking the features and forming subsets of uniform predictive quality for training individual members. The abductive network training algorithm is used to select optimum predictors from the feature set at various levels of model complexity specified by the user. Using the resulting tentative ranking, the features are grouped into mutually exclusive subsets of approximately equal predictive power for training the members. The approach is demonstrated on three standard medical diagnosis datasets (breast cancer, heart disease, and diabetes). Three-member committees trained on different feature subsets and using simple output combination methods reduce classification errors by up to 20% compared to the best single model developed with the full feature set. Results are compared with those reported previously with members trained through splitting the training set. Training abductive committee members on feature subsets of approximately equal predictive power achieves both diversity and quality for improved committee performance. Ensemble feature subset selection can be performed using GMDH-based learning algorithms. The approach should be advantageous in situations characterized by high data dimensionality.

Algorithms↗

Application of an ensemble technique based on singular spectrum analysis to daily rainfall forecasting.

In previous work, we have proposed a constructive methodology for temporal data learning supported by results and prescriptions related to the embedding theorem, and using the singular spectrum analysis both in order to reduce the effects of the possible discontinuity of the signal and to implement an efficient ensemble method. In this paper we present new results concerning the application of this approach to the forecasting of the individual rain-fall intensities series collected by 135 stations distributed in the Tiber basin. The average RMS error of the obtained forecasting is less than 3mm of rain.

Forecasting↗

Cortical ensemble adaptation to represent velocity of an artificial actuator controlled by a brain-machine interface.

Monkeys can learn to directly control the movements of an artificial actuator by using a brain-machine interface (BMI) driven by the activity of a sample of cortical neurons. Eventually, they can do so without moving their limbs. Neuronal adaptations underlying the transition from control of the limb to control of the actuator are poorly understood. Here, we show that rapid modifications in neuronal representation of velocity of the hand and actuator occur in multiple cortical areas during the operation of a BMI. Initially, monkeys controlled the actuator by moving a hand-held pole. During this period, the BMI was trained to predict the actuator velocity. As the monkeys started using their cortical activity to control the actuator, the activity of individual neurons and neuronal populations became less representative of the animal's hand movements while representing the movements of the actuator. As a result of this adaptation, the animals could eventually stop moving their hands yet continue to control the actuator. These results show that, during BMI control, cortical ensembles represent behaviorally significant motor parameters, even if these are not associated with movements of the animal's own limb.

Adaptation, Physiological↗

Clustering ensembles of neural network models.

We show that large ensembles of (neural network) models, obtained e.g. in bootstrapping or sampling from (Bayesian) probability distributions, can be effectively summarized by a relatively small number of representative models. In some cases this summary may even yield better function estimates. We present a method to find representative models through clustering based on the models' outputs on a data set. We apply the method on an ensemble of neural network models obtained from bootstrapping on the Boston housing data, and use the results to discuss bootstrapping in terms of bias and variance. A parallel application is the prediction of newspaper sales, where we learn a series of parallel tasks. The results indicate that it is not necessary to store all samples in the ensembles: a small number of representative models generally matches, or even surpasses, the performance of the full ensemble. The clustered representation of the ensemble obtained thus is much better suitable for qualitative analysis, and will be shown to yield new insights into the data.

Algorithms↗