This was part of Foundations of Multi-Agent and Mean Field Reinforcement Learning

Efficiently searching for good agent state based policies in Dec-POMDPs

Aditya Mahajan, McGill University

Thursday, May 21, 2026



Slides
Abstract:

Decentralized partially observable Markov decision processes (Dec-POMDPs) are becoming increasingly popular in various applications ranging from decentralized control of fleet of autonomous vehicles to that of smart grids. Optimally solving Dec-POMDPs is notoriously hard as is illustrated by the non-stationary problem and the search complexity of finding best history based policies (which is NEXP complete). Agent-state based policies have emerged as a popular paradigm to address some of these challenges.

In this talk, we review the existing solution approaches to find optimal agent state base policies and present a novel policy search algorithm which has monotonic improvement guarantee and converges to a locally optimal solution. The approach has an interesting similarity with forward-backward algorithms in mean-field games! We conclude by presenting experimental results that show that that the proposed algorithm identifies close to optimal policies in various POMDP and Dec-POMDP benchmarks.

Joint work with Amit Sinha and Matthieu Geist.