This was part of Reinforcement Learning from Offline Data and Human Feedback

Model simulation using offline observations with low-rank factor model

Devavrat Shah, Massachusetts Institute of Technology (MIT)

Wednesday, April 22, 2026



Slides
Abstract: We will discuss the role of low-rank factor models in developing model simulation using offline observations that are likely biased and coming from potentially heterogenous settings. We do so by positing that the transition dynamics can be represented as a latent function of latent factors associated with agents, states, and actions. Such naturally leads to approximate low-rank decomposition of separable agent, state, and action latent functions. This enables effective learning of the transition dynamics per agent, even with limited, offline data. This naturally extends the literature on causal inference rooted in panel data setting in Econometrics. I will discuss application of this approach in developing CausalSim, simulation platform for communication network protocols. Time permitting, I will discuss some of the ongoing theoretical inquiries suggested by the empirical success of such an approach.