Search PubMed⌕ Search

PubMed · 11506674

Subspace information criterion for model selection.

Abstract

The problem of model selection is considerably important for acquiring higher levels of generalization capability in supervised learning. In this article, we propose a new criterion for model selection, the subspace information criterion (SIC), which is a generalization of Mallows's C(L). It is assumed that the learning target function belongs to a specified functional Hilbert space and the generalization error is defined as the Hilbert space squared norm of the difference between the learning result function and target function. SIC gives an unbiased estimate of the generalization error so defined. SIC assumes the availability of an unbiased estimate of the target function and the noise covariance matrix, which are generally unknown. A practical calculation method of SIC for least-mean-squares learning is provided under the assumption that the dimension of the Hilbert space is less than the number of training examples. Finally, computer simulations in two examples show that SIC works well even when the number of training examples is small.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

M Sugiyama, H Ogawa. 2001. Subspace information criterion for model selection.. https://doi.org/10.1162/08997660152469387

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A statistical property of multiagent learning based on Markov decision process.

We exhibit an important property called the asymptotic equipartition property (AEP) on empirical sequences in an ergodic multiagent Markov decision process (MDP). Using the AEP which facilitates the analysis of multiagent learning, we give a statistical property of multiagent learning, such as reinforcement learning (RL), near the end of the learning process. We examine the effect of the conditions among the agents on the achievement of a cooperative policy in three different cases: blind, visible, and communicable. Also, we derive a bound on the speed with which the empirical sequence converges to the best sequence in probability, so that the multiagent learning yields the best cooperative result.

Learning↗

Second order neurons and learning in Cohen-Grossberg networks.

The well known Cohen-Grossberg network is modified to include second order neural interconnections and also to have a learning component. Sufficient conditions are obtained for the existence of a globally exponentially stable equilibrium. The model provides a two-fold generalization of the Cohen-Grossberg network in the sense if one removes the learning component, then one gets a network with second order synaptic interactions; if both the learning component and the second order interactions are removed, then the model reduces to the standard Cohen-Grossberg network.

Learning↗