This was part of Reinforcement Learning from Offline Data and Human Feedback

What structures make model-free RL possible? an elliptic theory for controlled Markov diffusions

Wenlong Mou, University of Toronto

Thursday, April 23, 2026



Slides
Abstract:

Can offline reinforcement learning with function approximation ever be as easy as supervised learning? In general, the answer is no — the Bellman operator contracts only in the sup-norm, not in the L^2-norm induced by the data distribution. This geometric mismatch makes model-free value learning with function approximation provably harder than regression. However, real-world problems often come with additional structures that may facilitate reinforcement learning.In this talk, I will discuss recent advances in understanding the structures that enable model-free offline RL. Focusing on controlled Markov diffusions—a widely used class of dynamical systems—I will provide an affirmative answer to the question above. Specifically, I will identify ellipticity as a key structure that makes model-free RL with function approximation tractable with offline data. Leveraging ellipticity, I will demonstrate desirable geometric properties of Bellman operators in an appropriate Sobolev space. Based on these insights, I will introduce a new class of algorithms for model-free RL with function approximation that achieve near-optimal oracle inequalities efficiently. Finally, I will discuss an application to fine-tuning diffusion-based generative models, where the ellipticity structure is exploited to design a PDE-based algorithm that attains fast convergence rates.