This was part of
Frontiers in Online Reinforcement Learning
Toward a Statistical Perspective on LLM Post-training: Preference Sampling and Gradient Reweighting
Yaqi Duan, New York University
Tuesday, March 31, 2026
Abstract: Post-training—the process of fine-tuning large language models (LLMs) with human preference or verifiable feedback—has rapidly evolved into a central problem in LLM development, yet its principles remain largely heuristic. This talk explores how statistical thinking can provide structure and rigor to this stage of learning. I will present two studies as steps toward formulating a statistical perspective on post-training. The first, PILAF, views human preference collection as an optimal experimental design problem, deriving sampling strategies that maximize reward-model information under budget constraints. The second, LENS, reinterprets reinforcement learning with verifiable rewards (RLVR) as a likelihood-based estimation problem, showing how confidence-weighted corrections on negative responses recover gradients otherwise lost in standard policy optimization. Together, these results illustrate how classical statistical reasoning can strengthen the foundations of post-training data collection and policy-update procedures, advancing efficiency, stability, and theoretical clarity.