Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control
Abstract
We derive statistical limit theorems for sample-based approximations of infinite-horizon discounted stochastic optimal control problems in discrete time.
Our first result is a functional central limit theorem for the sample-based value function under a uniqueness-type condition on population optimal policies.
The limiting law is a mean-zero Gaussian process characterized by a linear fixed point equation that resembles a dynamic programming principle.
We compare these asymptotics with those obtained from sample-based policy optimization and illustrate that their limiting variances can be different.
We also derive a limit theorem for models with nonunique optimal policies, where the limiting law may be non-Gaussian.
Applications to inventory control and renewable harvesting illustrate the theory.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요