Search PubMedSearch

Biomedical subjects

A G Barto

Publications and source records attributed to A G Barto.

9 recordsLinked to original sources

Reinforcement learning control.

Reinforcement learning refers to improving performance through trial-and-error. Despite recent progress in developing artificial learning systems, including new learning methods for artificial neural networks, most of these systems learn under the tutelage of a knowledgeable 'teacher' able to tell them how to respond to a set of training stimuli. Learning under these conditions is not adequate, however, when it is costly, or even impossible, to obtain this kind of training information. Reinforcement learning is attracting increasing attention in computer science and engineering because it can be used by autonomous systems to learn from their experiences instead of from knowledgeable teachers, and it is attracting attention in computational neuroscience because it is consonant with biological principles. Recent research has improved the efficiency of reinforcement learning and has provided some striking examples of its capabilities.

Animals

Distributed motor commands in the limb premotor network.

Neuroanatomical studies have demonstrated extensive interconnections between the motor cortex, red nucleus and cerebellum, forming a premotor network for controlling limb movement. Single-unit studies indicate that command signals for limb movements are distributed broadly throughout this network. Cellular studies have demonstrated multiple recurrent loops in this network, and the presence of excitatory and inhibitory amino acid neurotransmitters. A recent model suggests that movement commands are initiated by sensory inputs to these loops, and that positive feedback, regulated by inhibition from cerebellar Purkinje cells, distributes commands throughout the limb premotor network. This model offers a new framework for exploring relationships between basic neural mechanisms and concepts of motor performance that derive from experimental psychology.

Animals

Linear systems analysis of the relationship between firing of deep cerebellar neurons and the classically conditioned nictitating membrane response in rabbits.

The correlation of the activity of neurons in the interposed and dentate nuclei of the cerebellum with conditioned movements of the nictitating membrane was investigated using linear systems analysis. The activity of single deep cerebellar nuclear cells was assumed to be the input to a linear system that produced nictitating membrane movement. Data were initially analyzed with a causal model to assess the degree to which past neural activity predicted the conditioned response. 55 of 165 cells had correlation coefficients of 0.50 or greater between the model's moment-to-moment output and the actual output, with two interpositus cells having correlation coefficients of greater than 0.90. Double-sided impulse responses indicated that afference from the face and efference copy probably affect deep cerebellar neural activity. Nonlinearities were also found in the relationship between neuronal activity and conditioned movement. It was concluded that cerebellar deep nuclear firing is highly correlated with future nictitating membrane movements but that the firing-movement relationship contains noncausal and nonlinear components.

Animals

Simulation of the classically conditioned nictitating membrane response by a neuron-like adaptive element: response topography, neuronal firing, and interstimulus intervals.

A neuron-like adaptive element with computational features suitable for classical conditioning, the Sutton-Barto (S-B) model, was extended to simulate real-time aspects of the conditioned nictitating membrane (NM) response. The aspects of concern were response topography, CR-related neuronal firing, and interstimulus interval (ISI) effects for forward-delay and trace conditioning paradigms. The topography of the NM CR has the following features: response latency after CS onset decreases over trials; response amplitude increases gradually within the ISI and attains its maximum coincidentally with the UR. A similar pattern characterizes the firing of some (but not all) neurons in brain regions demonstrated experimentally to be important for NM conditioning. The variant of the S-B model described in this paper consists of a set of parameters and implementation rules based on 10-ms computational time steps. It differs from the original S-B model in a number of ways. The main difference is the assumption that CS inputs to the adaptive element are not instantaneous but are instead shaped by unspecified coding processes so as to produce outputs that conform with the real-time properties of NM conditioning. The model successfully simulates the aforementioned features of NM response topography. It is also capable of simulating appropriate ISI functions, i.e. with maximum conditioning strength with ISIs of 250 ms, for forward-delay and trace paradigms. The original model's successful treatment of multiple-CS phenomena, such as blocking, conditioned inhibition, and higher-order conditioning, are retained by the present model.

Adaptation, Physiological

Learning by statistical cooperation of self-interested neuron-like computing elements.

Since the usual approaches to cooperative computation in networks of neuron-like computating elements do not assume that network components have any "preferences", they do not make substantive contact with game theoretic concepts, despite their use of some of the same terminology. In the approach presented here, however, each network component, or adaptive element, is a self-interested agent that prefers some inputs over others and "works" toward obtaining the most highly preferred inputs. Here we describe an adaptive element that is robust enough to learn to cooperate with other elements like itself in order to further its self-interests. It is argued that some of the longstanding problems concerning adaptation and learning by networks might be solvable by this form of cooperativity, and computer simulation experiments are described that show how networks of self-interested components that are sufficiently robust can solve rather difficult learning problems. We then place the approach in its proper historical and theoretical perspective through comparison with a number of related algorithms. A secondary aim of this article is to suggest that beyond what is explicitly illustrated here, there is a wealth of ideas from game theory and allied disciplines such as mathematical economics that can be of use in thinking about cooperative computation in both nervous systems and man-made systems.

Adaptation, Physiological

Synthesis of nonlinear control surfaces by a layered associative search network.

An approach to solving nonlinear control problems is illustrated by means of a layered associative network composed of adaptive elements capable of reinforcement learning. The first layer adaptively develops a representation in terms of which the second layer can solve the problem linearly. The adaptive elements comprising the network employ a novel type of learning rule whose properties, we argue, are essential to the adaptive behavior of the layered network. The behavior of the network is illustrated by means of a spatial learning problem that requires the formation of nonlinear associations. We argue that this approach to nonlinearity can be extended to a large class of nonlinear control problems.

Animals

Simulation of anticipatory responses in classical conditioning by a neuron-like adaptive element.

A neuron-like adaptive element is described that produces an important feature of the anticipatory nature of classical conditioning. The response that occurs after training (conditioned response) usually begins earlier than the reinforcing stimulus (unconditioned stimulus). The conditioned response therefore usually anticipates the unconditioned stimulus. This aspect of classical conditioning has been largely neglected by hypotheses that neurons provide single unit analogs of conditioning. This paper briefly presents the model and extends earlier results by computer simulation of conditioned inhibition and chaining of associations.

Animals

Landmark learning: an illustration of associative search.

In a previous paper we defined the associative search problem and presented a system capable of solving it under certain conditions. In this paper we interpret a spatial learning problem as an associative search task and describe the behavior of an adaptive network capable of solving it. This example shows how naturally the associative search problem can arise and permits the search, association, and generalization properties of the adaptive network to be clearly illustrated.

Association Learning