Search PubMed⌕ Search

Biomedical subjects

K Doya

Publications and source records attributed to K Doya.

9 recordsLinked to original sources

Parallel cortico-basal ganglia mechanisms for acquisition and execution of visuomotor sequences - a computational approach.

Experimental studies have suggested that many brain areas, including the basal ganglia (BG), contribute to procedural learning. Focusing on the basal ganglia-thalamocortical (BG-TC) system, we propose a computational model to explain how different brain areas work together in procedural learning. The BG-TC system is composed of multiple separate loop circuits. According to our model, two separate BG-TC loops learn a visuomotor sequence concurrently but using different coordinates, one visual, and the other motor. The visual loop includes the dorsolateral prefrontal (DLPF) cortex and the anterior part of the BG, while the motor loop includes the supplementary motor area (SMA) and the posterior BG. The concurrent learning in these loops is based on reinforcement signals carried by dopaminergic (DA) neurons that project divergently to the anterior ("visual") and posterior ("motor") parts of the striatum. It is expected, however, that the visual loop learns a sequence faster than the motor loop due to their different coordinates. The difference in learning speed may lead to inconsistent outputs from the visual and motor loops, and this problem is solved by a mechanism called a "coordinator," which adjusts the contribution of the visual and motor loops to a final motor output. The coordinator is assumed to be in the presupplementary motor area (pre-SMA). We hypothesize that the visual and motor loops, with the help of the coordinator, achieve both the quick acquisition of novel sequences and the robust execution of well-learned sequences. A computational model based on the hypothesis is examined in a series of computer simulations, referring to the results of the 2 x 5 task experiments that have been used on both monkeys and humans. We found that the dual mechanism with the coordinator was superior to the single (visual or motor) mechanism. The model replicated the following essential features of the experimental results: (1) the time course of learning, (2) the effect of opposite hand use, (3) the effect of sequence reversal, and (4) the effects of localized brain inactivations. Our model may account for a common feature of procedural learning: A spatial sequence of discrete actions (subserved by the visual loop) is gradually replaced by a robust motor skill (subserved by the motor loop).

Basal Ganglia↗

Statistical characteristics of climbing fiber spikes necessary for efficient cerebellar learning.

Mean firing rates (MFRs), with analogue values, have thus far been used as information carriers of neurons in most brain theories of learning. However, the neurons transmit the signal by spikes, which are discrete events. The climbing fibers (CFs), which are known to be essential for cerebellar motor learning, fire at the ultra-low firing rates (around 1 Hz), and it is not yet understood theoretically how high-frequency information can be conveyed and how learning of smooth and fast movements can be achieved. Here we address whether cerebellar learning can be achieved by CF spikes instead of conventional MFR in an eye movement task, such as the ocular following response (OFR), and an arm movement task. There are two major afferents into cerebellar Purkinje cells: parallel fiber (PF) and CF, and the synaptic weights between PFs and Purkinje cells have been shown to be modulated by the stimulation of both types of fiber. The modulation of the synaptic weights is regulated by the cerebellar synaptic plasticity. In this study we simulated cerebellar learning using CF signals as spikes instead of conventional MFR. To generate the spikes we used the following four spike generation models: (1) a Poisson model in which the spike interval probability follows a Poisson distribution, (2) a gamma model in which the spike interval probability follows the gamma distribution, (3) a max model in which a spike is generated when a synaptic input reaches maximum, and (4) a threshold model in which a spike is generated when the input crosses a certain small threshold. We found that, in an OFR task with a constant visual velocity, learning was successful with stochastic models, such as Poisson and gamma models, but not in the deterministic models, such as max and threshold models. In an OFR with a stepwise velocity change and an arm movement task, learning could be achieved only in the Poisson model. In addition, for efficient cerebellar learning, the distribution of CF spike-occurrence time after stimulus onset must capture at least the first, second and third moments of the temporal distribution of error signals.

Action Potentials↗

Unsupervised learning of granule cell sparse codes enhances cerebellar adaptive control.

Marr [J. Physiol. (1969) 202, 437-470] and Albus [Math. Biosci. (1971) 10, 25-61] hypothesized that cerebellar learning is facilitated by a granule cell sparse code, i.e. a neural code in which the fraction of active neurons is low at any one time. In this paper, we re-examine this hypothesis in light of recent experimental and theoretical findings. We argue that cerebellar motor learning is enhanced by a sparse code that simultaneously maximizes information transfer between mossy fibers and granule cells, minimizes redundancies between granule cell discharges, and re-codes the mossy fiber inputs with an adaptive resolution such that inputs corresponding to large errors are finely encoded. We then propose that a set of biologically plausible unsupervised learning rules can produce such a code. To maintain a low mean firing rate compatible with a sparse code, an activity-dependent homeostatic mechanism sets the cells' thresholds. Then, to maximize information transfer, the mossy fiber--granule cell synapses are adjusted by a Hebbian rule. Furthermore, to minimize redundancies between granule cell discharges, the inhibitory Golgi cell--granule cell synapses are tuned by an anti-Hebbian rule. Finally, to allow adaptive resolution, a performance-based neuromodulator-like signal gates these three plastic processes. We integrate these gated learning rules into a simplified model of the cerebellum for arm movement control, and show that unsupervised learning of granule cell sparse codes greatly improves cerebellar adaptive motor control in comparison to a "fixed" Marr--Albus-type model. Until recently, activity-dependent cerebellar plasticity was thought to be largely confined to the granule cell--Purkinje cell synapses. This static view of the cerebellum is, however, quickly being replaced by an extremely dynamic view in which plasticity is omnipresent. The present theoretical study shows how several forms of plasticity in the granular layer of the cerebellum can produce fast, accurate and stable cerebellar learning.

Algorithms↗

Evidence for effector independent and dependent representations and their differential time course of acquisition during motor sequence learning.

To investigate the representation of motor sequence, we tested transfer effects in a motor sequence learning paradigm. We hypothesize that there are two sequence representations, effector independent and dependent. Further, we postulate that the effector independent representation is in visual/spatial coordinates, that the effector dependent representation is in motor coordinates, and that their time courses of acquisition during learning are different. Twelve subjects were tested in a modified 2x10 task. Subjects learned to press two keys (called a set) successively on a keypad in response to two lighted squares on a 3x3 display. The complete sequence to be learned was composed of ten such sets, called a hyperset. Training was given in the normal condition and sequence recall was assessed in the early, intermediate, and late stages in three conditions, normal, visual, and motor. In the visual condition, finger-keypad mapping was rotated 90 degrees while the keypad-display mapping was kept identical to normal. In the motor condition, the keypad-display mapping was also rotated 90 degrees, resulting in an identical finger-display mapping as in normal. Subjects formed two groups with each group using a different normal condition. One group learned the sequence in a standard keypad-hand setting and subsequently recalled the sequence using a rotated keypad-hand setting in the test conditions. The second group learned the sequence with a rotated keypad-hand setting and subsequently recalled the sequence with a standard keypad-hand setting in the test conditions. Response time (RT) and sequencing errors during recall were recorded. Although subjects committed more sequencing errors in both testing conditions, visual and motor, as compared to the normal condition, the errors were below chance level. Sequencing errors did not differ significantly between visual and motor conditions. Further, the sequence recall accuracy was over 70% even by the early stage when the subjects performed the sequence for the first time with the altered conditions, visual and motor. There were parallel improvements thereafter in all the conditions. These results of positive transfer of sequence knowledge across conditions that use dissimilar finger movements point to an effector independent sequence representation, possibly in visual/spatial coordinates. Initially the RTs were similar in the visual and the motor conditions, but with training RTs in the motor condition became significantly shorter than in the visual condition, as revealed by significant interaction for the testing stage and condition term in the repeated measures ANOVA. Moreover, using RTs for single key pressing in the three conditions as baseline indices, it was again observed that RTs in the visual and motor conditions were not significantly different in the early stage, but motor RTs became significantly shorter by the late testing stage. These results support the hypothesis that the motor condition benefits more than the visual because it uses identical effector movements to the normal condition. Further, these results argue for the existence of effector dependent sequence representation, in motor coordinates, which is acquired relatively slowly. The difference in the time course of learning of these two representations may account for the differential involvement of brain areas in early and late learning phases found in lesion and imaging studies.

Adult↗

Complementary roles of basal ganglia and cerebellum in learning and motor control.

The classical notion that the basal ganglia and the cerebellum are dedicated to motor control has been challenged by the accumulation of evidence revealing their involvement in non-motor, cognitive functions. From a computational viewpoint, it has been suggested that the cerebellum, the basal ganglia, and the cerebral cortex are specialized for different types of learning: namely, supervised learning, reinforcement learning and unsupervised learning, respectively. This idea of learning-oriented specialization is helpful in understanding the complementary roles of the basal ganglia and the cerebellum in motor control and cognitive functions.

Animals↗

Reinforcement learning in continuous time and space.

This article presents a reinforcement learning framework for continuous-time dynamical systems without a priori discretization of time, state, and action. Based on the Hamilton-Jacobi-Bellman (HJB) equation for infinite-horizon, discounted reward problems, we derive algorithms for estimating value functions and improving policies with the use of function approximators. The process of value function estimation is formulated as the minimization of a continuous-time form of the temporal difference (TD) error. Update methods based on backward Euler approximation and exponential eligibility traces are derived, and their correspondences with the conventional residual gradient, TD(0), and TD(lambda) algorithms are shown. For policy improvement, two methods-a continuous actor-critic method and a value-gradient-based greedy policy-are formulated. As a special case of the latter, a nonlinear feedback control law using the value gradient and the model of the input gain is derived. The advantage updating, a model-free algorithm derived previously, is also formulated in the HJB-based framework. The performance of the proposed algorithms is first tested in a nonlinear control task of swinging a pendulum up with limited torque. It is shown in the simulations that (1) the task is accomplished by the continuous actor-critic method in a number of trials several times fewer than by the conventional discrete actor-critic method; (2) among the continuous policy update methods, the value-gradient-based policy with a known or learned dynamic model performs several times better than the actor-critic method; and (3) a value function update using exponential eligibility traces is more efficient and stable than that based on Euler approximation. The algorithms are then tested in a higher-dimensional task: cart-pole swing-up. This task is accomplished in several hundred trials using the value-gradient-based policy with a learned dynamic model.

Algorithms↗

Parallel neural networks for learning sequential procedures.

Recent studies have shown that multiple brain areas contribute to different stages and aspects of procedural learning. On the basis of a series of studies using a sequence-learning task with trial-and-error, we propose a hypothetical scheme in which a sequential procedure is acquired independently by two cortical systems, one using spatial coordinates and the other using motor coordinates. They are active preferentially in the early and late stages of learning, respectively. Both of the two systems are supported by loop circuits formed with the basal ganglia and the cerebellum, the former for reward-based evaluation and the latter for processing of timing. The proposed neural architecture would operate in a flexible manner to acquire and execute multiple sequential procedures.

Animals↗

Electrophysiological properties of inferior olive neurons: A compartmental model.

As a step in exploring the functions of the inferior olive, we constructed a biophysical model of the olivary neurons to examine their unique electrophysiological properties. The model consists of two compartments to represent the known distribution of ionic currents across the cell membrane, as well as the dendritic location of the gap junctions and synaptic inputs. The somatic compartment includes a low-threshold calcium current (I(Ca_l)), an anomalous inward rectifier current (I(h)), a sodium current (I(Na)), and a delayed rectifier potassium current (I(K_dr)). The dendritic compartment contains a high-threshold calcium current (I(Ca_h)), a calcium-dependent potassium current (I(K_Ca)), and a current flowing into other cells through electrical coupling (I(c)). First, kinetic parameters for these currents were set according to previously reported experimental data. Next, the remaining free parameters were determined to account for both static and spiking properties of single olivary neurons in vitro. We then performed a series of simulated pharmacological experiments using bifurcation analysis and extensive two-parameter searches. Consistent with previous studies, we quantitatively demonstrated the major role of I(Ca_l) in spiking excitability. In addition, I(h) had an important modulatory role in the spike generation and period of oscillations, as previously suggested by Bal and McCormick. Finally, we investigated the role of electrical coupling in two coupled spiking cells. Depending on the coupling strength, the hyperpolarization level, and the I(Ca_l) and I(h) modulation, the coupled cells had four different synchronization modes: the cells could be in-phase, phase-shifted, or anti-phase or could exhibit a complex desynchronized spiking mode. Hence these simulation results support the counterintuitive hypothesis that electrical coupling can desynchronize coupled inferior olive cells.

Calcium Channels↗

Near-saddle-node bifurcation behavior as dynamics in working memory for goal-directed behavior.

In consideration of working memory as a means for goal-directed behavior in nonstationary environments, we argue that the dynamics of working memory should satisfy two opposing demands: long-term maintenance and quick transition. These two characteristics are contradictory within the linear domain. We propose the near-saddle-node bifurcation behavior of a sigmoidal unit with a self-connection as a candidate of the dynamical mechanism that satisfies both of these demands. It is shown in evolutionary programming experiments that the near-saddle-node bifurcation behavior can be found in recurrent networks optimized for a task that requires efficient use of working memory. The results suggests that the near-saddle-node bifurcation behavior may be a functional necessity for survival in nonstationary environments.

Animals↗