Off-policy Evaluation via Particle Filtering and Moment Matching
Nan Jiang, University of Illinois at Urbana-Champaign
I will present a new algorithmic framework and analysis for off-policy evaluation (OPE) in finite-horizon MDPs. The algorithm learns a scalar weight for each data point by a moment matching objective against a discriminator class F that realizes QÏ€. Notably, the theoretical guarantee of the algorithm is dimension-free, that the finite-sample error does NOT depend on the statistical complexity of the function class F (e.g., no log|F| dependence), and generalizes the standard error bound for linear regression with a fixed design. The algorithm is also closely connected to several existing methods, such as linear FQE, (sequential) importance sampling, and trajectory stitching, providing connections and novel perspectives to the foundational task of OPE.