Efficiently searching for good agent state based policies in Dec-POMDPs
Aditya Mahajan, McGill University
Decentralized partially observable Markov decision processes (Dec-POMDPs) are becoming increasingly popular in various applications ranging from decentralized control of fleet of autonomous vehicles to that of smart grids. Optimally solving Dec-POMDPs is notoriously hard as is illustrated by the non-stationary problem and the search complexity of finding best history based policies (which is NEXP complete). Agent-state based policies have emerged as a popular paradigm to address some of these challenges.
In this talk, we review the existing solution approaches to find optimal agent state base policies and present a novel policy search algorithm which has monotonic improvement guarantee and converges to a locally optimal solution. The approach has an interesting similarity with forward-backward algorithms in mean-field games! We conclude by presenting experimental results that show that that the proposed algorithm identifies close to optimal policies in various POMDP and Dec-POMDP benchmarks.
Joint work with Amit Sinha and Matthieu Geist.