This was part of Reinforcement Learning from Offline Data and Human Feedback

Optimal offline policy learning under unknown confounding factors

Zhimei Ren, University of Pennsylvania

Thursday, April 23, 2026



Slides
Abstract:

We investigate the problem of offline policy learning in the presence of unobserved confounders, which may arise in both observational studies and adaptive experiments (e.g., self-selection and noncompliance in sequential medical settings). In particular, we study this problem under the f -sensitivity model, which characterizes the confounding effect by its “average” strength. Under the f -sensitivity model, we characterize the distribution shift from the observable to the counterfactual and design a distributionally robust policy learning algorithm, f -SR(ad)L, which maximizes the expected outcome within a given policy class Π . We show that the sub-optimality gap of f -SR(ad)L learned from a sequential (i.i.d. or adaptively collected) dataset is of the order O(κ(Π)n), where κ(Π) is the entropy integral of Π under the Hamming distance and n is the sample size. A matching lower bound of is provided to show the optimality of the rate. Finally, we assess our method on synthetic and a real-world data on lung cancer treatments to demonstrate its advantage over existing benchmarks.