This was part of Reinforcement Learning from Offline Data and Human Feedback

Off-policy Evaluation via Particle Filtering and Moment Matching

Nan Jiang, University of Illinois at Urbana-Champaign

Wednesday, April 22, 2026



Slides
Abstract:

I will present a new algorithmic framework and analysis for off-policy evaluation (OPE) in finite-horizon MDPs. The algorithm learns a scalar weight for each data point by a moment matching objective against a discriminator class F that realizes QÏ€. Notably, the theoretical guarantee of the algorithm is dimension-free, that the finite-sample error does NOT depend on the statistical complexity of the function class F (e.g., no log|F| dependence), and generalizes the standard error bound for linear regression with a fixed design. The algorithm is also closely connected to several existing methods, such as linear FQE, (sequential) importance sampling, and trajectory stitching, providing connections and novel perspectives to the foundational task of OPE.