This was part of Frontiers in Online Reinforcement Learning

Toward a Statistical Perspective on LLM Post-training: Preference Sampling and Gradient Reweighting

Yaqi Duan, New York University

Tuesday, March 31, 2026



Abstract: Post-training—the process of fine-tuning large language models (LLMs) with human preference or verifiable feedback—has rapidly evolved into a central problem in LLM development, yet its principles remain largely heuristic. This talk explores how statistical thinking can provide structure and rigor to this stage of learning. I will present two studies as steps toward formulating a statistical perspective on post-training. The first, PILAF, views human preference collection as an optimal experimental design problem, deriving sampling strategies that maximize reward-model information under budget constraints. The second, LENS, reinterprets reinforcement learning with verifiable rewards (RLVR) as a likelihood-based estimation problem, showing how confidence-weighted corrections on negative responses recover gradients otherwise lost in standard policy optimization. Together, these results illustrate how classical statistical reasoning can strengthen the foundations of post-training data collection and policy-update procedures, advancing efficiency, stability, and theoretical clarity.