This was part of Reinforcement Learning from Offline Data and Human Feedback

Sampler Stochasticity in Training Diffusion Models for RLHF

Wenpin Tang, Columbia University

Monday, April 20, 2026



Slides
Abstract: In this talk, I will talk about the reward gap problem, which sees a tradeoff between RL training and diffusion inference. This provides some insights in choosing the level of stochasticity in diffusion generation.