Search PubMed⌕ Search

Biomedical subjects

P Read Montague

Publications and source records attributed to P Read Montague.

15 recordsLinked to original sources

Policy adjustment in a dynamic economic game.

Making sequential decisions to harvest rewards is a notoriously difficult problem. One difficulty is that the real world is not stationary and the reward expected from a contemplated action may depend in complex ways on the history of an animal's choices. Previous functional neuroimaging work combined with principled models has detected brain responses that correlate with computations thought to guide simple learning and action choice. Those works generally employed instrumental conditioning tasks with fixed action-reward contingencies. For real-world learning problems, the history of reward-harvesting choices can change the likelihood of rewards collected by the same choices in the near-term future. We used functional MRI to probe brain and behavioral responses in a continuous decision-making task where reward contingency is a function of both a subject's immediate choice and his choice history. In these more complex tasks, we demonstrated that a simple actor-critic model can account for both the subjects' behavioral and brain responses, and identified a reward prediction error signal in ventral striatal structures active during these non-stationary decision tasks. However, a sudden introduction of new reward structures engages more complex control circuitry in the prefrontal cortex (inferior frontal gyrus and anterior insula) and is not captured by a simple actor-critic model. Taken together, these results extend our knowledge of reward-learning signals into more complex, history-dependent choice tasks. They also highlight the important interplay between striatum and prefrontal cortex as decision-makers respond to the strategic demands imposed by non-stationary reward environments more reminiscent of real-world tasks.

Adult↗

Motor-sensory recalibration leads to an illusory reversal of action and sensation.

To judge causality, organisms must determine the temporal order of their actions and sensations. However, this judgment may be confounded by changing delays in sensory pathways, suggesting the need for dynamic temporal recalibration. To test for such a mechanism, we artificially injected a fixed delay between participants' actions (keypresses) and subsequent sensations (flashes). After participants adapted to this delay, flashes at unexpectedly short delays after the keypress were often perceived as occurring before the keypress, demonstrating a recalibration of motor-sensory temporal order judgments. When participants experienced illusory reversals, fMRI BOLD signals increased in anterior cingulate cortex/medial frontal cortex (ACC/MFC), a brain region previously implicated in conflict monitoring. This illusion-specific activation suggests that the brain maintains not only a recalibrated representation of timing, but also a less-plastic representation against which to compare it.

Brain↗

Agent-specific responses in the cingulate cortex during economic exchanges.

Interactions with other responsive agents lie at the core of all social exchange. During a social exchange with a partner, one fundamental variable that must be computed correctly is who gets credit for a shared outcome; this assignment is crucial for deciding on an optimal level of cooperation that avoids simple exploitation. We carried out an iterated, two-person economic exchange and made simultaneous hemodynamic measurements from each player's brain. These joint measurements revealed agent-specific responses in the social domain ("me" and "not me") arranged in a systematic spatial pattern along the cingulate cortex. This systematic response pattern did not depend on metrical aspects of the exchange, and it disappeared completely in the absence of a responding partner.

Brain Mapping↗

Imaging valuation models in human choice.

To make a decision, a system must assign value to each of its available choices. In the human brain, one approach to studying valuation has used rewarding stimuli to map out brain responses by varying the dimension or importance of the rewards. However, theoretical models have taught us that value computations are complex, and so reward probes alone can give only partial information about neural responses related to valuation. In recent years, computationally principled models of value learning have been used in conjunction with noninvasive neuroimaging to tease out neural valuation responses related to reward-learning and decision-making. We restrict our review to the role of these models in a new generation of experiments that seeks to build on a now-large body of diverse reward-related brain responses. We show that the models and the measurements based on them point the way forward in two important directions: the valuation of time and the valuation of fictive experience.

Animals↗

When things are better or worse than expected: the medial frontal cortex and the allocation of processing resources.

Access to limited-capacity neural systems of cognitive control must be restricted to the most relevant information. How the brain identifies and selects items for preferential processing is not fully understood. Anatomical models often place the selection mechanism in the medial frontal cortex (MFC), and one computational model proposes that the mesotelencephalic dopamine (DA) system, via its reward prediction properties, provides a "gate" through which information gains access to limited-capacity systems. There is a medial frontal event-related potential (ERP) index of attention selection, the anterior positivity (P2a), associated with DA reward system input to the MFC for the identification of task-relevant perceptual representations. The P2a has a similar spatio-temporal distribution as the medial frontal negativity (MFN), elicited to error responses or choices resulting in monetary loss. The MFN has also been linked to DA projections to the MFC but for action monitoring rather than attention selection. This study proposes that the P2a and the MFN reflect the same MFC evaluation function and use a passive reward prediction design containing neither instructed attention nor response to demonstrate that the ERP over medial frontal leads at the P2a/MFN latency is consistent with activity of midbrain DA neurons, positive to unpredicted rewards and negative when a predicted reward is withheld. This result suggests that MFC activity is regulated by DA reward system input and may function to identify items or actions that exceed or fail to meet motivational prediction.

Electroencephalography↗

Getting to know you: reputation and trust in a two-person economic exchange.

Using a multiround version of an economic exchange (trust game), we report that reciprocity expressed by one player strongly predicts future trust expressed by their partner-a behavioral finding mirrored by neural responses in the dorsal striatum. Here, analyses within and between brains revealed two signals-one encoded by response magnitude, and the other by response timing. Response magnitude correlated with the "intention to trust" on the next play of the game, and the peak of these "intention to trust" responses shifted its time of occurrence by 14 seconds as player reputations developed. This temporal transfer resembles a similar shift of reward prediction errors common to reinforcement learning models, but in the context of a social exchange. These data extend previous model-based functional magnetic resonance imaging studies into the social domain and broaden our view of the spectrum of functions implemented by the dorsal striatum.

Caudate Nucleus↗

Neural correlates of behavioral preference for culturally familiar drinks.

Coca-Cola (Coke) and Pepsi are nearly identical in chemical composition, yet humans routinely display strong subjective preferences for one or the other. This simple observation raises the important question of how cultural messages combine with content to shape our perceptions; even to the point of modifying behavioral preferences for a primary reward like a sugared drink. We delivered Coke and Pepsi to human subjects in behavioral taste tests and also in passive experiments carried out during functional magnetic resonance imaging (fMRI). Two conditions were examined: (1) anonymous delivery of Coke and Pepsi and (2) brand-cued delivery of Coke and Pepsi. For the anonymous task, we report a consistent neural response in the ventromedial prefrontal cortex that correlated with subjects' behavioral preferences for these beverages. In the brand-cued experiment, brand knowledge for one of the drinks had a dramatic influence on expressed behavioral preferences and on the measured brain responses.

Adult↗

Computational roles for dopamine in behavioural control.

Neuromodulators such as dopamine have a central role in cognitive disorders. In the past decade, biological findings on dopamine function have been infused with concepts taken from computational theories of reinforcement learning. These more abstract approaches have now been applied to describe the biological algorithms at play in our brains when we form value judgements and make choices. The application of such quantitative models has opened up new fields, ripe for attack by young synthesizers and theoreticians.

Animals↗

Dynamic gain control of dopamine delivery in freely moving animals.

Activity changes in a large subset of midbrain dopamine neurons fulfill numerous assumptions of learning theory by encoding a prediction error between actual and predicted reward. This computational interpretation of dopaminergic spike activity invites the important question of how changes in spike rate are translated into changes in dopamine delivery at target neural structures. Using electrochemical detection of rapid dopamine release in the striatum of freely moving rats, we established that a single dynamic model can capture all the measured fluctuations in dopamine delivery. This model revealed three independent short-term adaptive processes acting to control dopamine release. These short-term components generalized well across animals and stimulation patterns and were preserved under anesthesia. The model has implications for the dynamic filtering interposed between changes in spike production and forebrain dopamine release.

Adaptation, Physiological↗

The neural substrates of reward processing in humans: the modern role of FMRI.

Experimental work in animals has identified numerous neural structures involved in reward processing and reward-dependent learning. Until recently, this work provided the primary basis for speculations about the neural substrates of human reward processing. The widespread use of neuroimaging technology has changed this situation dramatically over the past decade through the use of PET and fMRI. Here, the authors focus on the role played by fMRI studies, where recent work has replicated the animal results in human subjects and has extended the view of putative reward-processing neural structures. In particular, fMRI work has identified a set of reward-related brain structures including the orbitofrontal cortex, amygdala, ventral striatum, and medial prefrontal cortex. Moreover, the human experiments have probed the dependence of human reward responses on learned expectations, context, timing, and the reward dimension. Current experiments aim to assess the function of human reward-processing structures to determine how they allow us to predict, assess, and act in response to rewards. The authors review current accomplishments in the study of human reward processing and focus their discussion on explanations directed particularly at the role played by the ventral striatum. They discuss how these findings may contribute to a better understanding of deficits associated with Parkinson's disease.

Algorithms↗

Temporal prediction errors in a passive learning task activate human striatum.

Functional MRI experiments in human subjects strongly suggest that the striatum participates in processing information about the predictability of rewarding stimuli. However, stimuli can be unpredictable in character (what stimulus arrives next), unpredictable in time (when the stimulus arrives), and unpredictable in amount (how much arrives). These variables have not been dissociated in previous imaging work in humans, thus conflating possible interpretations of the kinds of expectation errors driving the measured brain responses. Using a passive conditioning task and fMRI in human subjects, we show that positive and negative prediction errors in reward delivery time correlate with BOLD changes in human striatum, with the strongest activation lateralized to the left putamen. For the negative prediction error, the brain response was elicited by expectations only and not by stimuli presented directly; that is, we measured the brain response to nothing delivered (juice expected but not delivered) contrasted with nothing delivered (nothing expected).

Adult↗

A computational substrate for incentive salience.

Theories of dopamine function are at a crossroads. Computational models derived from single-unit recordings capture changes in dopaminergic neuron firing rate as a prediction error signal. These models employ the prediction error signal in two roles: learning to predict future rewarding events and biasing action choice. Conversely, pharmacological inhibition or lesion of dopaminergic neuron function diminishes the ability of an animal to motivate behaviors directed at acquiring rewards. These lesion experiments have raised the possibility that dopamine release encodes a measure of the incentive value of a contemplated behavioral act. The most complete psychological idea that captures this notion frames the dopamine signal as carrying 'incentive salience'. On the surface, these two competing accounts of dopamine function seem incommensurate. To the contrary, we demonstrate that both of these functions can be captured in a single computational model of the involvement of dopamine in reward prediction for the purpose of reward seeking.

Algorithms↗

Neural economics and the biological substrates of valuation.

A recent flurry of neuroimaging and decision-making experiments in humans, when combined with single-unit data from orbitofrontal cortex, suggests major additions to current models of reward processing. We review these data and models and use them to develop a specific computational relationship between the value of a predictor and the future rewards or punishments that it promises. The resulting computational model, the predictor-valuation model (PVM), is shown to anticipate a class of single-unit neural responses in orbitofrontal and striatal neurons. The model also suggests how neural responses in the orbitofrontal-striatal circuit may support the conversion of disparate types of future rewards into a kind of internal currency, that is, a common scale used to compare the valuation of future behavioral acts or stimuli.

Animals↗

Hyperscanning: simultaneous fMRI during linked social interactions.

"Plain question and plain answer make the shortest road out of most perplexities." Mark Twain-Life on the Mississippi. A new methodology for the measurement of the neural substrates of human social interaction is described. This technology, termed "Hyperscan," embodies both the hardware and the software necessary to link magnetic resonance scanners through the internet. Hyperscanning allows for the performance of human behavioral experiments in which participants can interact with each other while functional MRI is acquired in synchrony with the behavioral interactions. Data are presented from a simple game of deception between pairs of subjects. Because people may interact both asymmetrically and asynchronously, both the design and the analysis must accommodate this added complexity. Several potential approaches are described.

Brain↗

Activity in human ventral striatum locked to errors of reward prediction.

The mesolimbic dopaminergic system has long been known to be involved in the processing of rewarding stimuli, although recent evidence from animal research has suggested a more specific role of signaling errors in the prediction of rewards. We tested this hypothesis in humans, using functional magnetic resonance imaging (fMRI) and an operant conditioning paradigm for the discrete delivery of small quantities of fruit juice, along with a control experiment in which juice was substituted with a neutral visual stimulus. A local estimation of the activity in the ventral striatum showed a significant differentiation when the juice was withheld at the expected time of delivery; this finding was not replicated in the case of visual stimulation, providing evidence for time-locked processing of reward prediction errors in human ventral striatum.

Adult↗