Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies
Abstract
We study continuous-time reinforcement learning for $N$-player noncooperative stochastic differential games.
Each player adopts an entropy-regularized exploratory policy; given the others' actions, the optimal response is a Gibbs distribution, and a Nash equilibrium requires these $N$ conditional distributions to be jointly compatible.
We prove that the natural equilibrium concept -- simultaneous Hamiltonian maximization -- is equivalent to this compatibility, and establish a necessary and sufficient condition expressed as a computable cross-partial criterion on the optimal $q$-functions.
Nash equilibria exist unconditionally for decoupled and symmetric games.
When compatibility fails, a coordinate path integral construction yields an approximate correlated equilibrium with explicit quadratic KL-divergence bounds that vanish locally uniformly as the exploration weight $\gamma\to\infty$.
A $q$-function framework for the $N$-player game extends the single-agent $q$-learning theory of [21], with weak martingale characterizations motivating model-free on-policy and off-policy algorithms.
The framework extends to the ergodic (infinite-horizon) setting with the same locally uniform $O_R(1/\gamma)$ asymptotic rates.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요