Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Probability Learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

A comparison of algorithms for inference and learning in probabilistic graphical models.

Research into methods for reasoning under uncertainty is currently one of the most exciting areas of artificial intelligence, largely because it has recently become possible to record, store, and process large amounts of data. While impressive achievements have been made in pattern classification problems such as handwritten character recognition, face detection, speaker identification, and prediction of gene function, it is even more exciting that researchers are on the verge of introducing systems that can perform large-scale combinatorial analyses of data, decomposing the data into interacting components. For example, computational methods for automatic scene analysis are now emerging in the computer vision community. These methods decompose an input image into its constituent objects, lighting conditions, motion patterns, etc. Two of the main challenges are finding effective representations and models in specific applications and finding efficient algorithms for inference and learning in these models. In this paper, we advocate the use of graph-based probability models and their associated inference and learning algorithms. We review exact techniques and various approximate, computationally efficient techniques, including iterated conditional modes, the expectation maximization (EM) algorithm, Gibbs sampling, the mean field method, variational techniques, structured variational techniques and the sum-product algorithm ("loopy" belief propagation). We describe how each technique can be applied in a vision model of multiple, occluding objects and contrast the behaviors and performances of the techniques using a unifying cost function, free energy.

Algorithms↗

Executive amnesia in a patient with pre-frontal damage due to a gunshot wound.

This paper reports the case of a young patient with extensive pre-frontal damage in whom we tested the hypothesis that intensive training improves executive performance as assessed by the Wisconsin Card Sorting Test (WCST). As long as her declarative memory, complex perceptual abilities and global cognitive status were spared, we surmised that any deficit in executive learning would have occurred in relative isolation. We showed that her abnormal performance on the WCST, both on the standard as well as on the post-instruction condition, was due to an impairment of shifting attention across perceptual dimensions (extra-dimensional). In contrast, her ability to shift attention within perceptual categories (intra-dimensional) was spared, as were her declarative memory, object and visuospatial perception, oral language comprehension and praxis (ideomotor, tool use and constructional). This case supports the hypothesis that executive amnesia is a type of amnesic disorder distinct from the classic amnesic syndrome due to mamillo-temporomedial damage. As such, it is probably closely related to procedural learning and may depend on the same fronto-subcortical loops that mediate the actual execution of behaviour.

Adolescent↗

Nonparametric supervised learning by linear interpolation with maximum entropy.

Nonparametric neighborhood methods for learning entail estimation of class conditional probabilities based on relative frequencies of samples that are "near-neighbors" of a test point. We propose and explore the behavior of a learning algorithm that uses linear interpolation and the principle of maximum entropy (LIME). We consider some theoretical properties of the LIME algorithm: LIME weights have exponential form; the estimates are consistent; and the estimates are robust to additive noise. In relation to bias reduction, we show that near-neighbors contain a test point in their convex hull asymptotically. The common linear interpolation solution used for regression on grids or look-up-tables is shown to solve a related maximum entropy problem. LIME simulation results support use of the method, and performance on a pipeline integrity classification problem demonstrates that the proposed algorithm has practical value.

Algorithms↗

Bayesian analysis of interleaved learning and response bias in behavioral experiments.

Accurate characterizations of behavior during learning experiments are essential for understanding the neural bases of learning. Whereas learning experiments often give subjects multiple tasks to learn simultaneously, most analyze subject performance separately on each individual task. This analysis strategy ignores the true interleaved presentation order of the tasks and cannot distinguish learning behavior from response preferences that may represent a subject's biases or strategies. We present a Bayesian analysis of a state-space model for characterizing simultaneous learning of multiple tasks and for assessing behavioral biases in learning experiments with interleaved task presentations. Under the Bayesian analysis the posterior probability densities of the model parameters and the learning state are computed using Monte Carlo Markov Chain methods. Measures of learning, including the learning curve, the ideal observer curve, and the learning trial translate directly from our previous likelihood-based state-space model analyses. We compare the Bayesian and current likelihood-based approaches in the analysis of a simulated conditioned T-maze task and of an actual object-place association task. Modeling the interleaved learning feature of the experiments along with the animal's response sequences allows us to disambiguate actual learning from response biases. The implementation of the Bayesian analysis using the WinBUGS software provides an efficient way to test different models without developing a new algorithm for each model. The new state-space model and the Bayesian estimation procedure suggest an improved, computationally efficient approach for accurately characterizing learning in behavioral experiments.

Bayes Theorem↗

On the relationship between recognition speed and accuracy for words rehearsed via rote versus elaborative rehearsal.

Tacit within both lay and cognitive conceptualizations of learning is the notion that those conditions of learning that foster "good" retention do so by increasing both the probability and the speed of access to the relevant information. In 3 experiments, time pressure during recognition is shown to decrease accessibility more for words learned via elaborative rehearsal than for words learned via rote rehearsal, despite the fact that elaborative rehearsal is a more efficacious learning strategy as measured by the probability of access. In Experiment 1, participants learned each word using both types of rehearsal, and the results show that access to the products of elaborative rehearsal is more compromised by time pressure than is access to the products of rote rehearsal. The results of Experiment 2, in which each word was learned via either pure rote or pure elaborative rehearsal, exhibit the same pattern. Experiment 3, in which the authors used the response-signal procedure, provides evidence that this difference in accessibility owes not to differences in the rate of access to the 2 types of traces, but rather to the higher asymptotic level of stored information for words learned via elaborative rehearsal.

Adult↗

Judgments of proportions.

This study investigated the processes that underlie estimates of relative frequency. Ss performed 4 tasks using the same stimuli (squares containing black and white dots); they judged "percentages" of white dots, "percentages" of black dots, "ratios" of black dots to white dots, and "differences" between the number of black and white dots. Results were consistent with the theory that Ss used the instructed operations with the same scale values in all tasks. Despite the use of the correct operation, Ss consistently overestimated small proportions and underestimated large proportions. Variations in the distributions of actual proportions affected the extent to which Ss overestimated small proportions and underestimated large proportions in the direction predicted by range-frequency theory. Results suggest that proportion judgments, and by analogy probability judgments, should not be taken at face value.

Adult↗

Probability matching, accuracy maximization, and a test of the optimal classifier's independence assumption in perceptual categorization.

Observers completed perceptual categorization tasks that included 25 base-rate/payoff conditions constructed from the factorial combination of five base-rate ratios (1:3, 1:2, 1:1, 2:1, and 3:1) with five payoff ratios (1:3, 1:2, 1:1, 2:1, and 3:1). This large database allowed an initial comparison of the competition between reward and accuracy maximization (COBRA) hypothesis with a competition between reward maximization and probability matching (COBRM) hypothesis, and an extensive and critical comparison of the flat-maxima hypothesis with the independence assumption of the optimal classifier. Model-based instantiations of the COBRA and COBRM hypotheses provided good accounts of the data, but there was a consistent advantage for the COBRM instantiation early in learning and for the COBRA instantiation later in learning. This pattern held in the present study and in a reanalysis of Bohil and Maddox (2003). Strong support was obtained for the flat-maxima hypothesis over the independence assumption, especially as the observers gained experience with the task. Model parameters indicated that observers' reward-maximizing decision criterion rapidly approaches the optimal value and that more weight is placed on accuracy maximization in separate base-rate/payoff conditions than in simultaneous base-rate/payoff conditions. The superiority of the flat-maxima hypothesis suggests that violations of the independence assumption are to be expected, and are well captured by the flat-maxima hypothesis, with no need for any additional assumptions.

Decision Making↗

Inflation of conditional predictions.

The authors report 7 experiments indicating that conditional predictions--the assessed probability that a certain outcome will occur given a certain condition--tend to be markedly inflated. The results suggest that this inflation derives in part from backward activation in which the target outcome highlights aspects of the condition that are consistent with that outcome, thus supporting the plausibility of that outcome. One consequence of this process is that alternative outcomes are not conceived to compete as fully as they should. Another consequence is that prediction inflation is resistant to manipulations that induce participants to consider alternative outcomes to the target outcome.

Adolescent↗

Probability judgment and subadditivity: the role of working memory capacity and constraining retrieval.

In this research, we examined the role that individual differences in working memory (WM) capacity, the strength of alternatives, and time constraints play in probability judgment and subadditivity. With a laboratory-based learning task, Experiment 1 revealed that the degree to which participants' probability judgments were subadditive was negatively correlated with a measure of WM capacity, even when variance due to short-term memory capacity was removed. In addition, participants were more subadditive when the viable alternatives were all rather weak. Experiment 2 extended the WM-capacity-subadditivity correlation to a population judgment task and revealed that subadditivity increases when the judgment task is performed under time constraints. Results support a model that assumes that people make probability judgments by comparing the focal hypothesis with relevant alternatives retrieved from long-term memory and that people high in WM span include more alternatives in the comparison process. Time constraints are assumed to truncate the alternative generation process, leading to fewer alternatives being recalled from long-term memory.

Humans↗

Plasma glucose and insulin levels in monkeys anticipating feeding.

To see whether plasma glucose or insulin changed in anticipation of feeding, we provided seven rhesus monkeys with four-hour access to food every other day. Blood was sampled before and during a 30-minute signal which ended with food availability and before and during a 30-minute signal which was not closely and reliably linked with food availability. Plasma insulin showed no evidence of conditioning. Plasma glucose was higher during the signal than prior to the signal in both experiments. This probably reflects the arousing nature of the signal rather than appetitive-associated learning. However, the differences, while statistically significant, were probably biologically trivial because they fall within the normal fluctuations of meal-fed monkeys. Under the conditions of this experiments, it appears that conditional changes in glucose and insulin do not reliably occur in monkeys anticipating access to food.

Animals↗

Instructional and probability manipulations of bias in multiletter matching.

Ratcliff (1985) performed fits of his diffusion model to the results of multiletter-matching experiments conducted by Ratcliff and Hacker (1981) and Proctor, Rao, and Hurst (1984), in which bias to respond "same" or "different" was manipulated by instructions and probabilities, respectively. The fits showed that both bias manipulations affected settings of a goodness-of-match criterion, whereas instructions also affected sensitivity. Evaluations of the experimental procedures and of Ratcliff's model-fitting procedures were performed in the present study. Three experiments showed that instructions and probabilities had similar effects, regardless of whether the different pairs were blocked or randomized according to the number of mismatching positions. The most salient feature of the results--that "same" reaction times were traded off more than were "different" reaction times, with no corresponding asymmetry in the error rates--was evident in all situations. The evaluation of Ratcliff's model-fitting procedures indicated that the apparent influence of instructions on sensitivity likely is an artifact of unequal variance for the sets of same and different pairs. Moreover, the effects of bias can be explained in terms of settings of response criteria, rather than of the goodness-of-match criterion, as in Ratcliff's fits.

Adult↗

Brain mechanisms of selective learning: event-related potentials provide evidence for error-driven learning in humans.

Selective learning has been observed in Pavlovian conditioning in animals and in judgements of event contingencies in humans. This analogy led to the suggestion that the formation of associations underlies both types of learning. An alternative theory proposes that both tasks involve the computation of event contingencies as prescribed by probability theory. Error-driven models of learning incorporate trial-by-trial error-correction mechanisms during training whereas probabilistic models view learning merely as the storage of frequency information for later use during judgement of event contingencies. Competitive interaction between cues was observed in a contingency judgement task. Event-related brain potentials (ERPs) provided evidence for brain events related to the discrepancy between actual and expected outcomes during training thus supporting error-driven accounts of selective learning.

Adult↗

Disparity tuning as simulated by a neural net.

Previous research has suggested that the processing of binocular disparity in complex cells may be described with an energy formalism. The energy formalism allows for a representation of disparity by differences in the position or in the phase of monocular receptive subfields of binocular cells, or by combination of these two types. We studied the coding of disparities with an approach complementary to previous algorithmic investigations. Since realization of these representations is probably not genetically determined but learned during ontogeny, we used backpropagation networks to study which of these three possibilities were realized within neural nets. Three types of networks were trained with noise patterns in analogy to the three types of energy models. The networks learned the task and generalized to untrained correlated noise pattern input. Outputs were broadly tuned to spatial frequency and did not respond to anti-correlated noise patterns. Although the energy model was not explicitly implemented, we could analyze the outputs of the networks using predictions of the energy formalism. After learning was completed, the model neurons preferred position shifts over phase shifts in representing disparity. We discuss the general meaning of these findings and the correspondences and deviations between the energy model, V1 neurons, and our networks.

Computer Simulation↗

[Nocebo effect: the other side of placebo].

Administration of drugs is often followed by beneficial (placebo effects) and harmful (nocebo effects) effects that are not always related to their mechanism of action. Nocebo effects are rather unknown even when may be the source of many adverse reactions which could be erroneously attributed to drug therapy. Some mechanisms have been postulated which might be associated with the development of nocebo effects. Expectancy, learning and classical conditioning are probably important in the psychological domain. The neuropharmacological substrate is much less known yet an opioid peptide-cholecystokinin interaction has been suggested. At the clinical setting, a nocebo effect should be suspected in those patients who present common unspecific symptoms after drug administration and have a tendency to somatize. An early detection of these patients may contribute to the prevention of the nocebo effect.

Clinical Trials as Topic↗

An Extension of the Back-Propagation Algorithm to Complex Numbers.

This paper presents a complex-valued version of the back-propagation algorithm (called 'Complex-BP'), which can be applied to multi-layered neural networks whose weights, threshold values, input and output signals are all complex numbers. Some inherent properties of this new algorithm are studied. The results may be summarized as follows. The updating rule of the Complex-BP is such that the probability for a "standstill in learning" is reduced. The average convergence speed is superior to that of the real-valued back-propagation, whereas the generalization performance remains unchanged. In addition, the number of weights and thresholds needed is only about the half of real-valued back-propagation, where a complex-valued parameter z=x+iy (where i=-1) is counted as two because it consists of a real part x and an imaginary part y. The Complex-BP can transform geometric figures, e.g. rotation, similarity transformation and parallel displacement of straight lines, circles, etc., whereas the real-valued back-propagation cannot. Mathematical analysis indicates that a Complex-BP network which has learned a transformation, has the ability to generalize that transformation with an error which is represented by the sine. It is interesting that the above characteristics appear only by extending neural networks to complex numbers.

Journal Article↗

Multilayer neural networks and Bayes decision theory.

There are many applications of multilayer neural networks to pattern classification problems in the engineering field. Recently, it has been shown that Bayes a posteriori probability can be estimated by feedforward neural networks through computer simulation. In this paper, Bayes decision theory is combined with the approximation theory on three-layer neural networks, and the two-category n-dimensional Gaussian classification problem is studied. First, we prove theoretically that three-layer neural networks with at least 2n hidden units have the capability of approximating the a posteriori probability in the two-category classification problem with arbitrary accuracy. Second, we prove that the input-output function of neural networks with at least 2n hidden units tends to the a posteriori probability as Back-Propagation learning proceeds ideally. These results provide a theoretical basis for the study of pattern classification by computer simulation.

Journal Article↗