Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning
Abstract
We study a structured class of Markov Decision Processes, known as Exo-MDPs, in which the state space is partitioned into exogenous and endogenous components.
Exogenous states evolve stochastically, independent of the agent's actions, while endogenous states evolve deterministically based on both state components and actions.
Exo-MDPs capture many operations research settings, including inventory control, resource management, and ride-sharing.
Our first contribution is structural: we establish a representational equivalence between discrete MDPs, Exo-MDPs, and discrete linear mixture MDPs.
Our second contribution is statistical.
We characterize the minimax regret of learning in Exo-MDPs when the effective dimension r is small relative to the endogenous state and action spaces.
When the exogenous states are unobserved, we prove matching upper and lower regret bounds of order $\Theta(Hr \sqrt{K})$ over $K$ episodes of horizon $H$, where $r$ is the effective dimension of the Exo-MDP.
When exogenous states are observed, the minimax regret improves to $\Theta(H\sqrt{ r K})$, revealing a $\Theta(\sqrt{r})$ statistical gap due to observation of the exogenous states.
These results show that Exo-MDPs decouple sample complexity from action space and endogenous state space.
We validate these insights with experiments on inventory control and resource allocation.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요