Forgetting is Everywhere
Abstract
A fundamental challenge in developing general learning algorithms is their tendency to forget past knowledge as they adapt to new data.
Addressing this problem requires a principled understanding of forgetting.
Yet, despite decades of study, no unified definition has emerged that offers insight into the underlying dynamics of learning.
We propose an algorithm- and task-agnostic theory that characterises forgetting as a lack of self-consistency in a learner's predictive distribution, manifesting as a loss of predictive information.
Our theory naturally yields a general measure of an algorithm's propensity to forget, proves that exact Bayesian inference allows for adaptation without forgetting, and provides a tautological explanation for why generative models forget when trained on their own synthetic outputs.
To validate these claims, we design a comprehensive set of experiments that span classification, regression, generative modelling, and reinforcement learning.
We demonstrate that forgetting is present across all deep learning settings and plays a significant role in determining learning efficiency.
Together, these results establish a principled understanding of forgetting and lay the foundation for analysing and improving the information retention capabilities of general learning algorithms.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요