This was part of
Reinforcement Learning from Offline Data and Human Feedback
Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach
Renyuan Xu, Stanford University
Monday, April 20, 2026
Abstract:
We study how to steer diffusion models under hard constraints, so that generated samples satisfy prescribed events almost surely. This problem arises naturally in safety-critical generation, constrained decision-making, and rare-event simulation, where one seeks to condition a pretrained model using only offline trajectories while guaranteeing exact constraint satisfaction.
Our approach builds on Doob’s h-transform and introduces a novel martingale-based loss to learn an additive guidance term, without retraining the full score network. We propose two off-policy objectives for estimating this guidance term from pretrained trajectories, establish non-asymptotic guarantees for the resulting sampler, and demonstrate strong performance on stress testing for financial assets and queueing networks.
This is based on joint work with Wenpin Tang and Zhengyi Guo (Columbia University).