This was part of Foundations of Multi-Agent and Mean Field Reinforcement Learning

Reinforcement Learning for Quantal Mean-Field Leader-Follower Games

Sebastian Jaimungal, University of Toronto

Tuesday, May 19, 2026



Slides
Abstract: We study discrete-time episodic leader–follower mean-field games with a single far-sighted leader and an infinite population of boundedly rational, myopic followers who use quantal response. The leader seeks to maximize long-run reward while learning both the environment and the followers’ response behavior under information asymmetry. We formulate learning the resulting equilibrium as an online reinforcement learning problem and propose optimistic value-iteration algorithms using linear and RKHS function approximation. Our method integrates optimistic value-based planning with a dynamically shrinking confidence set over followers’ response parameters, constructed from observed aggregate behavior. We establish high-probability sublinear regret bounds that scale with intrinsic function class complexity. The results provide sample-efficient guarantees for learning quantal equilibria in leader-follower games.