This was part of Frontiers in Online Reinforcement Learning

Reward-Guided Generation in Diffusion Models

Masatoshi Uehara, Evolutionary Scale

Thursday, April 2, 2026



Abstract: Diffusion models are celebrated for their strong generative capabilities. However, practical applications often demand sample generation that not only produces realistic outputs but also optimizes specific objectives (e.g., human preference scores in computer vision, binding affinity in proteins). To address this, diffusion models can be adapted to explicitly maximize desired reward metrics. While many methods have been developed for domains like computer vision, applying reward-guided generation to biological design poses unique challenges: (1) reward functions are often non-differentiable, and (2) biological data frequently involves discrete data. In this talk, I will present our recent advances in test-time controlled generation methods that address these challenges. I will also discuss how these techniques enable real-world applications across molecular design tasks, including protein, DNA, RNA, and small molecule generation.