This was part of Foundations of Multi-Agent and Mean Field Reinforcement Learning

Policy Gradient for Continuous-Time Mean-Field Control from Discrete-Time Data

Yuhua Zhu, University of California, Los Angeles (UCLA)

Thursday, May 21, 2026



Slides
Abstract: We consider mean-field optimal control problems where the dynamics are governed by an unknown McKean–Vlasov equation and only discrete-time observations are available. Existing RL algorithms applied to such data do not exploit the smooth structure of the underlying dynamics, which can limit their effectiveness in physical-world settings. We first develop a new policy gradient algorithm for the case of known dynamics. We then consider the more challenging setting with unknown dynamics and discrete-time data, and propose a new framework that incorporates the smooth structure of the underlying system into discrete-time information, and establish its first-order accuracy. Moreover, in the linear–quadratic mean-field control setting, we show that this error decreases as the running cost gets less discounted, a regime typically regarded as more challenging. Finally, we derive a model-free policy gradient algorithm and validate both the theoretical results and the proposed method through numerical experiments.