Search PubMed⌕ Search

PubMed · 14692633

Inter-module credit assignment in modular reinforcement learning.

Abstract

Critical issues in modular or hierarchical reinforcement learning (RL) are (i) how to decompose a task into sub-tasks, (ii) how to achieve independence of learning of sub-tasks, and (iii) how to assure optimality of the composite policy for the entire task. The second and last requirements are often under trade-off. We propose a method for propagating the reward for the entire task achievement between modules. This is done in the form of a 'modular reward', which is calculated from the temporal difference of the module gating signal and the value of the succeeding module. We implement modular reward for a multiple model-based reinforcement learning (MMRL) architecture and show its effectiveness in simulations of a pursuit task with hidden states and a continuous-time non-linear control task.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kazuyuki Samejima, Kenji Doya, Mitsuo Kawato. 2003. Inter-module credit assignment in modular reinforcement learning.. https://doi.org/10.1016/s0893-6080(02)00235-6

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

[Working memory in basic learning processes].

INTRODUCTION AND DEVELOPMENT: Working or operative memory is considered to be a distinctive element of executive functioning. Nowadays, thanks to neuroimaging studies, it is known that the dorsolateral prefrontal cortex plays a crucial role in working memory. It has been observed that during the intervals when information is being retained, intense and persistent activity is going on in the region, as shown by the delayed response times. Working memory is fundamental for the analysis and synthesis of information, the retention of data needed to perform a particular mental process, carrying out priming (impression in memory of something that has been experienced, such as words, objects or events, for example), carrying out pre-functional tutoring activities and post-functional monitoring. CONCLUSIONS: Disorders affecting the fundamental mechanisms of working memory will give rise to a dysfunction that will exert an influence on innumerable formal academic learning processes such as difficulty in focusing attention, difficulty in inhibiting irrelevant stimuli, difficulty in recognising priority patterns, inability to recognise hierarchies and the meaning of stimuli (analysis and synthesis), problems in establishing an intention, and difficulty in recognising and selecting the goals that are best suited to solving a problem. It will also involve the impossibility to establish a plan to achieve goals, inability to analyse the activities required to accomplish an objective and difficulties in carrying out a plan, since it becomes impossible to monitor or modify the task to fit the original plans.

Learning↗

Stability criteria for unsupervised temporal association networks.

A biologically realizable, unsupervised learning rule is described for the online extraction of object features, suitable for solving a range of object recognition tasks. Alterations to the basic learning rule are proposed which allow the rule to better suit the parameters of a given input space. One negative consequence of such modifications is the potential for learning instability. The criteria for such instability are modeled using digital filtering techniques and predicted regions of stability and instability tested. The result is a family of learning rules which can be tailored to the specific environment, improving both convergence times and accuracy over the standard learning rule, while simultaneously insuring learning stability.

Learning↗

Self-organizing learning array.

A new machine learning concept--self-organizing learning array (SOLAR)--is presented. It is a sparsely connected, information theory-based learning machine, with a multilayer structure. It has reconfigurable processing units (neurons) and an evolvable system structure, which makes it an adaptive classification system for a variety of machine learning problems. Its multilayer structure can handle complex problems. Based on the entropy estimation, information theory-based learning is performed locally at each neuron. Neural parameters and connections that correspond to minimum entropy are adaptively set for each neuron. By choosing connections for each neuron, the system sets up its wiring and completes its self-organization. SOLAR classifies input data based on the weighted statistical information from all the neurons. The system classification ability has been simulated and experiments were conducted using test-bench data. Results show a very good performance compared to other classification methods. An important advantage of this structure is its scalability to a large system and ease of hardware implementation on regular arrays of cells.

Learning↗