This was part of
Reinforcement Learning from Offline Data and Human Feedback
Statistical Inference under Adaptive Sampling with LinUCB
Yuting Wei, University of Pennsylvania
Wednesday, April 22, 2026
Abstract: Adaptively collected data has become ubiquitous in modern practice. Yet even seemingly benign adaptive sampling schemes can introduce severe biases, rendering traditional statistical inference tools inapplicable. Focusing on the linear bandit problem, a fundamental and influential framework in reinforcement learning and the bandit literature, we characterize the performance of LinUCB, a canonical upper-confidence-bound algorithm that balances exploration and exploitation, and derive inferential procedures that remain valid despite the challenges posed by adaptive data collection. A central difficulty is to understand the behavior of the eigenvalues and eigenvectors of the random feature covariance matrix generated by LinUCB without imposing the stability assumptions that prior work relied upon. Our analysis provides this characterization and, in turn, enables us to establish a central limit theorem for LinUCB: the estimation error converges in distribution at a $T^{-1/4}$ rate and is asymptotically normal. The resulting Wald-type confidence sets and hypothesis tests do not depend on the feature covariance matrix and are asymptotically tighter than existing nonasymptotic confidence sets. Numerical simulations corroborate our theoretical findings.