Search PubMed⌕ Search

PubMed · 11674845

Predictability, complexity, and learning.

Abstract

We define predictive information I(pred)(T) as the mutual information between the past and the future of a time series. Three qualitatively different behaviors are found in the limit of large observation times T:I(pred)(T) can remain finite, grow logarithmically, or grow as a fractional power law. If the time series allows us to learn a model with a finite number of parameters, then I(pred)(T) grows logarithmically with a coefficient that counts the dimensionality of the model space. In contrast, power-law growth is associated, for example, with the learning of infinite parameter (or nonparametric) models such as continuous functions with smoothness constraints. There are connections between the predictive information and measures of complexity that have been defined both in learning theory and the analysis of physical systems through statistical mechanics and dynamical systems theory. Furthermore, in the same way that entropy provides the unique measure of available information consistent with some simple and plausible conditions, we argue that the divergent part of I(pred)(T) provides the unique measure for the complexity of dynamics underlying a time series. Finally, we discuss how these ideas may be useful in problems in physics, statistics, and biology.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

W Bialek, I Nemenman, N Tishby. 2001. Predictability, complexity, and learning.. https://doi.org/10.1162/089976601753195969

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

[Quality management in health care: error prevention and managing errors in medicine].

BACKGROUND: About 3.7% of in-house-treated patients in Switzerland, the USA and Australia are victims of treatment-related health problems which probably are related to avoidable "adverse events" in more than 50% of the occurrences. Reasons are primarily systematic incidents, e.g., organizational deficiencies in the health system and only secondarly individual mistakes. As there are no systematic studies available, it is not proven if those figures can be transferred to the German Health Care System. Here, experts anticipate up to 12.000 proven treatment errors per year. PREVENTION OF AND DEALING WITH ADVERSE EVENTS: Dedicated programs for identification and prevention of adverse events should be implemented--besides systematic quality improvement--to improve professional handling and prevention of adverse events. This consists of a) assessment of the existing problem using existing data bases and/or implementation of mandatory documentation and information routines as well as reporting systems, b) development of sanction-free reporting routines within the legal framework, c) dissemination of behavior-oriented training systems for recognition and prevention of adverse events as well as incentives for the participation in such training systems, d) implementation of automatic routines for prevention of adverse events (e.g., computer-based monitoring of ADE or computer-based reminder systems based on clinical guidelines).

Forecasting↗

Statistical approaches to estimating mean water quality concentrations with detection limits.

We review statistical methodology for estimating mean concentrations of potentially toxic pollutants in water, for small samples that are not normally distributed and often contain substantial numbers of nondetects, i.e. samples that are only known to be below some set of fixed thresholds. Maximum likelihood estimation (MLE) and regression on order statistics (ROS) are two main approaches that dominate the literature, with transformation bias under non-normality that increases with the severity of censoring being the main problem. We consider exact maximum likelihood estimators in conjunction with the Box-Cox transformation and propose the Quenouille-Tukey Jackknife as a method for bias reduction and variance estimation. Exact maximum likelihood estimators resulting from the expectation-maximization (EM) algorithm are exhibited in a simple heuristic form that also provides estimated values for the nondetects as subsidiary outputs. We show in simulationsthatthetwo main approaches perform well for the log-normal and gamma distributions as long as the jackknife is employed to reduce bias. Bias corrections to MLE used in the literature are shown to correct in the wrong direction under severe censoring. The jackknife is also used for estimating the variance of the both the MLE and ROS estimators. Robustness is improved by searching a class of power transformations (Box-Cox) for the best approximating normal distribution. We conclude that both the exact MLE and ROS procedures can be useful under varying experimental conditions. Limited simulations indicate that the ROS procedure is unbiased and has a smaller variance than the MLE under the log-normal distribution and is robust. The MLE performed better in simulations involving the gamma as the underlying distribution. We also compare the estimators for the mean and variance that one obtains from typical sets of water quality data, analyzing for copper, alumnium, arsenic, chromium, nickel, and lead.

Forecasting↗