This was part of Reinforcement Learning from Offline Data and Human Feedback

From Reward Learning to Leaderboards: Uncertainty Quantification for LLMs under Heterogeneous Human Feedback

Will Wei Sun, Purdue University

Tuesday, April 21, 2026



Slides
Abstract:

Pairwise human feedback is now widely used in both LLM alignment and LLM evaluation, from reward modeling in RLHF to public leaderboards based on head-to-head comparisons. However, these data are noisy, heterogeneous, and highly non-uniform, making uncertainty quantification a central statistical challenge. In this talk, I will present two recent works on this topic. The first studies reward learning under heterogeneous human feedback, jointly modeling latent rewards and annotator rationality, with asymptotic guarantees that enable valid reward comparison and uncertainty-aware best-of-N sampling. The second studies LLM evaluation as inference on a low-rank latent score tensor observed through pairwise comparisons, leading to efficient debiased inference and a score-whitening method for handling anisotropic information under non-uniform sampling. Together, these works illustrate how statistical inference can provide principled uncertainty quantification for both alignment and evaluation of large language models.