What structures make model-free RL possible? an elliptic theory for controlled Markov diffusions
Wenlong Mou, University of Toronto
Can offline reinforcement learning with function approximation ever be as easy as supervised learning? In general, the answer is no — the Bellman operator contracts only in the sup-norm, not in the L^2-norm induced by the data distribution. This geometric mismatch makes model-free value learning with function approximation provably harder than regression. However, real-world problems often come with additional structures that may facilitate reinforcement learning.In this talk, I will discuss recent advances in understanding the structures that enable model-free offline RL. Focusing on controlled Markov diffusions—a widely used class of dynamical systems—I will provide an affirmative answer to the question above. Specifically, I will identify ellipticity as a key structure that makes model-free RL with function approximation tractable with offline data. Leveraging ellipticity, I will demonstrate desirable geometric properties of Bellman operators in an appropriate Sobolev space. Based on these insights, I will introduce a new class of algorithms for model-free RL with function approximation that achieve near-optimal oracle inequalities efficiently. Finally, I will discuss an application to fine-tuning diffusion-based generative models, where the ellipticity structure is exploited to design a PDE-based algorithm that attains fast convergence rates.