This was part of
Reinforcement Learning from Offline Data and Human Feedback
Sampler Stochasticity in Training Diffusion Models for RLHF
Wenpin Tang, Columbia University
Monday, April 20, 2026
Abstract: In this talk, I will talk about the reward gap problem, which sees a tradeoff between RL training and diffusion inference. This provides some insights in choosing the level of stochasticity in diffusion generation.