Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “REINFORCEMENT”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Effects of drugs on response duration differentiation. III. Acute variation of reinforced duration.

Rats trained to hold a lever down for at least 1.0 s but less than 1.3 s could differentiate the reinforced response duration on about 50% of the trials. The response duration frequency distribution was a normal distribution with a peak near the minimum reinforced response duration. Dose-effect curves were determined for the effects of phencyclidine (PCP) and methamphetamine. Subsequently, rats continued to be trained for 3 days a week with responses between 1.0 and 1.3 s reinforced, but on days when injections were given either the maximum reinforced duration was increased to 2.3 s, or the minimum reinforced duration was lowered to 0.5. When the maximum duration was increased to 2.3 s, the percentage of reinforced responses increased to 60% and when the minimum reinforced duration was decreased to 0.5 s, the percentage of reinforced responses increased to 89%. Despite the increased percentage of reinforced responses when the time window was widened, the effects of PCP and methamphetamine were not changed. These data suggest that the effects of drugs on response duration differentiation are not greatly influenced by transient changes in reinforcement frequency.

Animals↗

Effects of intraaccumbens injections of dopamine agonists and antagonists on sucrose and sucrose-ethanol reinforced responding.

The present experiment tested the effects of intraaccumbens injections of dopamine (DA) agonists and antagonists on operant responding reinforced by sucrose and sucrose/ethanol solutions. The mixed DA agonist d-amphetamine (20.0 micrograms/microliters) significantly reduced responding reinforced by a low concentration sucrose solution (2% w/v) by 48% and 38% compared to no injection and sham control values, respectively. The addition of ethanol (10%) to a low concentration sucrose solution (3%) presented as the reinforcer changed the response pattern from a continuous moderate response rate, over a 30 min session, to an initial high response rate that terminated after approximately 10 min. With sucrose/ethanol reinforcement, d-amphetamine slowed the initial high response rate but extended responding throughout the 30 min sessions. However, no significant changes were observed in number of responses per session. When 75% sucrose (w/v) was presented as the reinforcer, d-amphetamine did not change the total number of responses/session, but response patterns were again altered from high initial rates with early offset to slow steady rates that continued for the duration of sessions. The D2 DA antagonist raclopride (0.1-5.0 micrograms/microliters) resulted in a dose-dependent decrease in responding reinforced by 75% sucrose. The baseline patterns, response totals, and effects of the DA antagonists resemble our previously reported findings with 10% ethanol (v/v) reinforcement. These data support the conclusion that mesolimbic DA activity may be a common mechanism in ethanol reinforced behavior and behavior reinforced by other substances, but suggest that the nature of behavioral change may depend upon the reinforcer.

Animals↗

A neural network model with dopamine-like reinforcement signal that learns a spatial delayed response task.

This study investigated how the simulated response of dopamine neurons to reward-related stimuli could be used as reinforcement signal for learning a spatial delayed response task. Spatial delayed response tasks assess the functions of frontal cortex and basal ganglia in short-term memory, movement preparation and expectation of environmental events. In these tasks, a stimulus appears for a short period at a particular location, and after a delay the subject moves to the location indicated. Dopamine neurons are activated by unpredicted rewards and reward-predicting stimuli, are not influenced by fully predicted rewards, and are depressed by omitted rewards. Thus, they appear to report an error in the prediction of reward, which is the crucial reinforcement term in formal learning theories. Theoretical studies on reinforcement learning have shown that signals similar to dopamine responses can be used as effective teaching signals for learning. A neural network model implementing the temporal difference algorithm was trained to perform a simulated spatial delayed response task. The reinforcement signal was modeled according to the basic characteristics of dopamine responses to novel stimuli, primary rewards and reward-predicting stimuli. A Critic component analogous to dopamine neurons computed a temporal error in the prediction of reinforcement and emitted this signal to an Actor component which mediated the behavioral output. The spatial delayed response task was learned via two subtasks introducing spatial choices and temporal delays, in the same manner as monkeys in the laboratory. In all three tasks, the reinforcement signal of the Critic developed in a similar manner to the responses of natural dopamine neurons in comparable learning situations, and the learning curves of the Actor replicated the progress of learning observed in the animals. Several manipulations demonstrated further the efficacy of the particular characteristics of the dopamine-like reinforcement signal. Omission of reward induced a phasic reduction of the reinforcement signal at the time of the reward and led to extinction of learned actions. A reinforcement signal without prediction error resulted in impaired learning because of perseverative errors. Loss of learned behavior was seen with sustained reductions of the reinforcement signal, a situation in general comparable to the loss of dopamine innervation in Parkinsonian patients and experimentally lesioned animals. The striking similarities in teaching signals and learning behavior between the computational and biological results suggest that dopamine-like reward responses may serve as effective teaching signals for learning behavioral tasks that are typical for primate cognitive behavior, such as spatial delayed responding.

Animals↗

Noncontingent reinforcement: effects of satiation versus choice responding.

Recent research findings suggest that the initial reductive effects of noncontingent reinforcement (NCR) schedules on destructive behavior result from the establishing effects of an antecedent stimulus (i.e., the availability of "free" reinforcement) rather than extinction. A number of authors have suggested that these antecedent effects result primarily from reinforcer satiation, but an alternative hypothesis is that the individual attempts to access contingent reinforcement primarily when noncontingent reinforcement is unavailable, but chooses not to access contingent reinforcement when noncontingent reinforcement is available. If the satiation hypothesis is more accurate, then the reductive effects of NCR should increase over the course of a session, especially for denser schedules of NCR, and should occur during both NCR delivery and the NCR inter-reinforcement interval (NCR IRI). If the choice hypothesis is more accurate, then the reductive effects of NCR should be relatively constant over the course of a session for both denser and leaner schedules of NCR and should occur almost exclusively during the NCR interval (rather than the NCR IRI). To evaluate these hypotheses, we examined within-session trends of destructive behavior with denser and leaner schedules of NCR (without extinction), and also measured responding in the NCR interval separate from responding in the NCR IRI. Reductions in destructive behavior were mostly due to the participants choosing not to access contingent reinforcement when NCR was being delivered and only minimally due to reinforcer satiation.

Adolescent↗

Ethanol-maintained responding of rats is more resistant to change in a context with added non-drug reinforcement.

Alternative non-drug reinforcers reliably decrease drug-maintained responding in self-administration procedures. Studies of the resistance to change of food-maintained behavior, however, have found that responding in the presence of a stimulus associated with an alternative reinforcer is more resistant to disruption. This increase in persistence occurs despite lower response rates when the alternative reinforcer is present. The present experiment examined if, in addition to decreasing response rates, an alternative non-drug reinforcer also increases the persistence of drug-maintained responding. Rats self-administered oral ethanol in a multiple schedule of reinforcement in which responding was reinforced in two components signaled by different stimuli. In one component, response-independent food was delivered in addition to the earned ethanol. The effects of the alternative food reinforcer on response rates and resistance to extinction in the two components were examined. As in previous experiments on the resistance to change of food-maintained operant behavior, response rates were lower, but more resistant to extinction in the presence of the stimulus associated with the alternative reinforcer. These findings suggest that all the reinforcers obtained in a context in which drugs are consumed may contribute to the persistence of drug seeking in that context. This increase in persistence may occur even if the alternative reinforcers interfere with drug seeking.

Administration, Oral↗

Revealed preference between reinforcers used to examine hypotheses about behavioral consistencies.

New techniques for measuring preference between reinforcers have emerged in a field known as behavioral economics. Preference is assessed from the relative shapes of reinforcer demand functions, shown in graphs in which rate of reinforcement is plotted against schedule requirement. In economic terminology, a schedule requirement sets the price of a reinforcer as it sets the numbers of responses needed to obtain a reinforcer. Relative shapes of demand functions for alternative reinforcers are interpreted using the principle of revealed preference, as the shape of a demand function reflects the numbers of responses emitted to obtain reinforcers at each schedule requirement. Individual preferences between reinforcers are measured from differences in shapes of demand functions. Demand functions from two single subject experiments are examined to assess the hypothesis that individuals may generate differently shaped demand functions for the same reinforcers. It is hypothesized that individual differences in reinforcer preference may be related to consistent differences in behavior such as those observed in personality traits.

Female↗

The relative motivational properties of sensory and edible reinforcers in teaching autistic children.

We compared the effects of sensory and edible reinforcers on resistance to satiation in three autistic children while learning visual discrimination tasks. Within-subject designs were used to compare a single sensory reinforcer with a single edible reinforcer and to compare multiple sensory reinforcers with multiple edibles. Results indicated that multiple sensory reinforcers maintained responding over more trials than did multiple edible reinforcers; however, the use of single sensory reinforcers and single edibles resulted in about equal numbers of trials to satiation. Both multiple and single sensory reinforcers produced higher percentages of correct responses than edible reinforcers. The findings are discussed in terms of the advantages of sensory reinforcers in teaching autistic children.

Autistic Disorder↗

Establishing discriminative control of responding using functional and alternative reinforcers during functional communication training.

Functional communication training (FCT) is a popular treatment for problem behaviors, but its effectiveness may be compromised when the client emits the target communication response and reinforcement is either delayed or denied. In the current investigation, we trained 2 individuals to emit different communication responses to request (a) the reinforcer for destructive behavior in a given situation (e.g., contingent attention in the attention condition of a functional analysis) and (b) an alternative reinforcer (e.g., toys in the attention condition of a functional analysis). Next, we taught the participants to request each reinforcer in the presence of a different discriminative stimulus (SD). Then, we evaluated the effects of differential reinforcement of communication (DRC) using the functional and alternative reinforcers and correlated SDs, with and without extinction of destructive behavior. During all applications, DRC (in combination with SDs that signaled available reinforcers) rapidly reduced destructive behavior to low levels regardless of whether the functional reinforcer or an alternative reinforcer was available or whether reinforcement for destructive behavior was discontinued (i.e., extinction).

Adolescent↗

Reinforcement schedule thinning following treatment with functional communication training.

We evaluated four methods for increasing the practicality of functional communication training (FCT) by decreasing the frequency of reinforcement for alternative behavior. Three participants whose problem behaviors were maintained by positive reinforcement were treated successfully with FCT in which reinforcement for alternative behavior was initially delivered on fixed-ratio (FR) 1 schedules. One participant was then exposed to increasing delays to reinforcement under FR 1, a graduated fixed-interval (FI) schedule, and a graduated multiple-schedule arrangement in which signaled periods of reinforcement and extinction were alternated. Results showed that (a) increasing delays resulted in extinction of the alternative behavior, (b) the FI schedule produced undesirably high rates of the alternative behavior, and (c) the multiple schedule resulted in moderate and stable levels of the alternative behavior as the duration of the extinction component was increased. The other 2 participants were exposed to graduated mixed-schedule (unsignaled alternation between reinforcement and extinction components) and multiple-schedule (signaled alternation between reinforcement and extinction components) arrangements in which the durations of the reinforcement and extinction components were modified. Results obtained for these 2 participants indicated that the use of discriminative stimuli in the multiple schedule facilitated reinforcement schedule thinning. Upon completion of treatment, problem behavior remained low (or at zero), whereas alternative behavior was maintained as well as differentiated during a multiple-schedule arrangement consisting of a 4-min extinction period followed by a 1-min reinforcement period.

Adult↗

The roles of stimulus control and reinforcement frequency in modulating the behavioral effects of d-amphetamine in the rat.

The behavioral effects of d-amphetamine have been shown to be modulated by stimulus control, with less impairment of performance occurring when control is great. When the fixed-consecutive-number schedule is used (on which at least a specified consecutive number of responses must be made on one operandum before a single response on another will produce a reinforcer), response rate tends to be invariant but reinforcement frequency is not. This study asks whether the differences in reinforcement frequency that usually accompany changes in stimulus control could themselves be responsible for the performance differences. Two versions of the fixed-consecutive-number schedule of reinforcement were combined into a multiple schedule within which stimulus control was varied but differences in reinforcement frequency were minimized by omitting some reinforcer deliveries during the component that usually had the higher reinforcement frequency. In one component, a compound discriminative stimulus was added with the eighth consecutive response on the first lever; a single response on the second lever was then reinforced. In the other component, no such stimulus was presented. With no added stimulus, large decreases occurred in the number of runs satisfying the minimum requirement for reinforcement at doses of drug that produced only minimal changes when an added stimulus controlled behavior. Thus, increased stimulus control diminishes the behavioral changes produced by d-amphetamine even when the possible contribution by baseline reinforcement rate is minimized.

Animals↗

Probability and delay of reinforcement as factors in discrete-trial choice.

Pigeons chose between two alternatives that differed in the probability of reinforcement and the delay to reinforcement. A peck on the red key always produced a delay of 5 s and then a possible reinforcer. The probability of reinforcement for responding on this key varied from .05 to 1.0 in different conditions. A response on the green key produced a delay of adjustable duration and then a possible reinforcer, with the probability of reinforcement ranging from .25 to 1.0 in different conditions. The green-key delay was increased or decreased many times per session, depending on a subject's previous choices. The purpose of these adjustments was to estimate an indifference point, or a delay that resulted in a subject's choosing each alternative about equally often. In conditions where the probability of reinforcement was five times higher on the green key, the green-key delay averaged about 12 s at the indifference point. In conditions where the probability of reinforcement was twice as high on the green key, the green-key delay at the indifference point was about 8 s with high probabilities and about 6 s with low probabilities. An analysis based on these results and those from studies on delay of reinforcement suggests that pigeons' choices are relatively insensitive to variations in the probability of reinforcement between .2 and 1.0, but quite sensitive to variations in probability between .2 and 0.

Animals↗

Theories of probabilistic reinforcement.

In three experiments, pigeons chose between two alternatives that differed in the probability of reinforcement and the delay to reinforcement. A peck at a red key led to a delay of 5 s and then a possible reinforcer. A peck at a green key led to an adjusting delay and then a certain reinforcer. This delay was adjusted over trials so as to estimate an indifference point, or a duration at which the two alternatives were chosen about equally often. In Experiments 1 and 2, the intertrial interval was varied across conditions, and these variations had no systematic effects on choice. In Experiment 3, the stimuli that followed a choice of the red key differed across conditions. In some conditions, a red houselight was presented for 5 s after each choice of the red key. In other conditions, the red houselight was present on reinforced trials but not on nonreinforced trials. Subjects exhibited greater preference for the red key in the latter case. The results were used to evaluate four different theories of probabilistic reinforcement. The results were most consistent with the view that the value or effectiveness of a probabilistic reinforcer is determined by the total time per reinforcer spent in the presence of stimuli associated with the probabilistic alternative. According to this view, probabilistic reinforcers are analogous to reinforcers that are delivered after variable delays.

Animals↗

Choice behavior in transition: development of preference for the higher probability of reinforcement.

Ten acquisition curves were obtained from each of 4 pigeons in a two-choice discrete-trial procedure. In each of these 10 conditions, the two response keys initially had equal probabilities of reinforcement, and subjects' choice responses were about equally divided between the two keys. Then the reinforcement probabilities were changed so that one key had a higher probability of reinforcement (the left key in half of the conditions and the right key in the other half), and in nearly every case the subjects developed a preference for this key. The rate of acquisition of preference for this key was faster when the ratio of the two reinforcement probabilities was higher. For instance, acquisition of preference was faster in conditions with reinforcement probabilities of .12 and .02 than in conditions with reinforcement probabilities of .40 and .30, even though the pairs of probabilities differed by .10 in both cases. These results were used to evaluate the predictions of some theories of transitional behavior in choice situations. A trial-by-trial analysis of individual responses and reinforcers suggested that reinforcement had both short-term and long-term effects on choice. The short-term effect was an increased probability of returning to the same key on the one or two trials following a reinforcer. The long-term effect was a gradual increase in the proportion of responses on the key with the higher probability of reinforcement, an increase that usually continued for several hundred trials.

Animals↗

Concurrent schedules: reinforcer magnitude effects.

Five pigeons were trained on pairs of concurrent variable-interval schedules in a switching-key procedure. The arranged overall rate of reinforcement was constant in all conditions, and the reinforcer-magnitude ratios obtained from the two alternatives were varied over five levels. Each condition remained in effect for 65 sessions and the last 50 sessions of data from each condition were analyzed. At a molar level of analysis, preference was described well by a version of the generalized matching law, consistent with previous reports. More local analyses showed that recently obtained reinforcers had small measurable effects on current preference, with the most recently obtained reinforcer having a substantially larger effect. Larger reinforcers resulted in larger and longer preference pulses, and a small preference was maintained for the larger-magnitude alternative even after long inter-reinforcer intervals. These results are consistent with the notion that the variables controlling choice have both short- and long-term effects. Moreover, they suggest that control by reinforcer magnitude is exerted in a manner similar to control by reinforcer frequency. Lower sensitivities when reinforcer magnitude is varied are likely to be due to equal frequencies of different sized preference pulses, whereas higher sensitivities when reinforcer rates are varied might result from changes in the frequencies of different sized preference pulses.

Animals↗

Response rates and choices of schizophrenics under fixed-ratio contingencies of reinforcement.

Schizophrenics (n = 12) were conditioned under different multiple fixed-ratio (mult FR FR) schedules of monetary reinforcement. The two FR components of these schedules differed in terms of ratio requirements (reinforcement frequency) or amounts of reinforcement per occurrence of reinforcement. Relatively low rates of responding were emitted by the schizophrenics under these schedules. Further, their response rates were positively correlated with the frequency and amount of FR reinforcement. In previous studies under comparable conditions, normal subjects tended to maximize reinforcement by responding at higher rates and to maintain these rates irrespective of the frequency or amount of FR reinforcement. When given the opportunity to select from among the two components of the mult FR FR schedules, the schizophrenics in the present study tended to respond like normal subjects in previous studies in that they chose to work predominantly under that FR component which provided the highest frequency or amount of reinforcement. It was concluded that schizophrenics resemble normals more and act more rationally in terms of maximizing reinforcement when reinforcement is less dependent upon rates of responding.

Choice Behavior↗

The effect of rate of reinforcement and time in session on preference for variability.

Pigeons pecked keys on concurrent-chains schedules that provided a variable interval 30-sec schedule in the initial link. One terminal link provided reinforcers in a fixed manner; the other provided reinforcers in a variable manner with the same arithmetic mean as the fixed alternative. In Experiment 1, the terminal links provided fixed and variable interval schedules. In Experiment 2, the terminal links provided reinforcers after a fixed or a variable delay following the response that produced them. In Experiment 3, the terminal links provided reinforcers that were fixed or variable in size. Rate of reinforcement was varied by changing the scheduled interreinforcer interval in the terminal link from 5 to 225 sec. The subjects usually preferred the variable option in Experiments 1 and 2 but differed in preference in Experiment 3. The preference for variability was usually stronger for lower (longer terminal links) than for higher (shorter terminal links) rates of reinforcement. Preference did not change systematically with time in the session. Some aspects of these results are inconsistent with explanations for the preference for variability in terms of scaling factors, scalar expectancy theory, risk-sensitive models of optimal foraging theory, and habituation to the reinforcer. Initial-link response rates also changed within sessions when the schedules provided high, but not low, rates of reinforcement. Within-session changes in responding were similar for the two initial links. These similarities imply that habituation to the reinforcer is represented differently in theories of choice than are other variables related to reinforcement.

Animals↗

Ethanol as a reinforcer in the newborn's first suckling experience.

BACKGROUND: Recent evidence suggests that human infants prefer alcohol-flavored milk when fed through a bottle. Animal models also indicate a surprising predisposition for neonatal and infant rats to voluntarily and willingly ingest ethanol. These findings suggest high susceptibility to the reinforcing properties of ethanol early in ontogeny. METHODS: A surrogate nipple technique-a highly effective tool for investigation of the reinforcing properties of different fluids-was applied in the present study. Tests of ethanol reinforcement were accomplished in terms of two basic paradigms of Pavlovian conditioning. In one paradigm, the conditioned stimulus (CS) was the surrogate nipple, and in the other, the CS was a novel odor. RESULTS: Newborn rats showed sustained attachment to the nipple providing 5% ethanol, and later reproduced this behavioral pattern toward the empty nipple (CS alone). Ingestion of ethanol yielding appetitive reinforcement was accompanied by detectable blood alcohol concentrations, with most in the range of 20-30 mg/dl. The reinforcing efficacy of ethanol was also confirmed in the classical olfactory conditioning paradigm: following pairing with intraoral ethanol infusions, the odor (CS) alone elicited sustained attachment to an empty nipple. Females showed better olfactory conditioning with low concentrations of ethanol, whereas males were effectively more conditioned to high concentrations. Although there were no reinforcing consequences of intraperitoneally injected ethanol [as an unconditioned stimulus (US)] when a neutral odor was the CS, when paired with ingestion of water from a nipple, the injection of ethanol had a reinforcing effect. CONCLUSIONS: The present series of experiments revealed ethanol reinforcement in the newborn rat. Two varieties of Pavlovian conditioning established that ethanol can serve as an effective US, and hence reinforcer, in such a way as to increase the approach and responsiveness toward stimuli paired with that US, indicating appetitive reinforcement.

Animals↗

Lesions of the orbitofrontal but not medial prefrontal cortex disrupt conditioned reinforcement in primates.

The ventromedial prefrontal cortex (PFC) is implicated in affective and motivated behaviors. Damage to this region, which includes the orbitofrontal cortex as well as ventral sectors of medial PFC, causes profound changes in emotional and social behavior, including impairments in certain aspects of decision making. One reinforcement mechanism that may well contribute to these behaviors is conditioned reinforcement, whereby previously neutral stimuli in the environment, by virtue of their association with primary rewards, take on reinforcing value and come to support instrumental action. Conditioned reinforcers are powerful determinants of behavior and can maintain responding over protracted periods of time in the absence of and potentially in conflict with primary reinforcers. It has already been shown that conditioned reinforcement is dependent on the amygdala, and because the amygdala projects to both the orbitofrontal cortex and the medial PFC, the present study determined whether conditioned reinforcement was also dependent on one or the other of these prefrontal regions. Comparison of the behavioral effects of selective excitotoxic lesions of the PFC in the common marmoset revealed that orbitofrontal but not medial PFC lesions disrupted two distinct measures of conditioned reinforcement: (1) acquisition of a new response and (2) sensitivity to conditioned stimulus omission on a second-order schedule. In contrast, the orbitofrontal lesion did not affect sensitivity to primary reinforcement as measured by responding on a progressive-ratio schedule and a home cage consumption test. Together, these findings demonstrate the critical and specific involvement of the orbitofrontal cortex but not the medial PFC in conditioned reinforcement.

Acoustic Stimulation↗