This was part of Reinforcement Learning from Offline Data and Human Feedback

Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach

Renyuan Xu, Stanford University

Monday, April 20, 2026



Slides
Abstract:
We study how to steer diffusion models under hard constraints, so that generated samples satisfy prescribed events almost surely. This problem arises naturally in safety-critical generation, constrained decision-making, and rare-event simulation, where one seeks to condition a pretrained model using only offline trajectories while guaranteeing exact constraint satisfaction.
Our approach builds on Doob’s h-transform and introduces a novel martingale-based loss to learn an additive guidance term, without retraining the full score network. We propose two off-policy objectives for estimating this guidance term from pretrained trajectories, establish non-asymptotic guarantees for the resulting sampler, and demonstrate strong performance on stress testing for financial assets and queueing networks.
This is based on joint work with Wenpin Tang and Zhengyi Guo (Columbia University).