Statistical Foundations of Generative Modeling

Theory, Evaluation, Applications

Due to the high interest in this workshop, this event will be held in two locations: David Rubenstein Forum at 1201 E 60th Street, Chicago, IL 60637 and an overflow space at the IMSI Building at 1155 E 60th Street Chicago, IL 60637

Description

Back to top

Generative models have rapidly become a central tool in modern AI and data science. At a high level, a generative model learns an underlying probability distribution from data and can sample from it to create synthetic yet realistic outputs: text, images, financial scenarios, molecules, patient records, climate simulations, and more. Recent advances such as variational autoencoders, generative adversarial networks, normalizing flows, diffusion models, and flow matching have delivered striking empirical performance. At the same time, many of the most pressing questions remain fundamentally statistical: What distribution is being learned, and under what assumptions is it identifiable? Which distributional features are easy or hard to capture (e.g., modes with complex geometry, rare events, and tail behavior)? How can we quantify uncertainty, control bias, and ensure calibration, especially in high-stakes settings where downstream decisions depend on faithful modeling of extremes?

These challenges become even sharper in domain-specific contexts. Open-ended text generation lacks a single “correct” output, making objective evaluation of quality, coherence, diversity, and fluency a critical open problem. In finance and other dependent-data regimes, correlations and selection effects in training data raise questions about generalization and downstream validity. Across application areas, statistics plays a key role in designing reliable metrics to evaluate and compare generative models, and in understanding the properties of common fine-tuning and alignment procedures. More broadly, determining when synthetic data is “good enough” for inference, prediction, or decision-making remains an open question, as do opportunities to use generative models for tasks such as anomaly and changepoint detection.

This workshop brings together statisticians, machine learning researchers, and practitioners from domains including language modeling, finance, biomedicine, and the natural sciences to develop a shared language and research agenda. The goal is to connect modern generative modeling techniques to classical statistical principles, while advancing theory, methodology, and practices that enable reliable deployment in real-world scientific and societal applications.

Some of the funding for this workshop is provided by the Stevanovich Center.

Poster Session

This workshop will include a poster session for early career researchers (including graduate students). In order to propose a poster, you must first register for the workshop, and then submit a proposal using the form that will become available on this page after you register. The registration form should not be used to propose a poster. The organizers may offer the opportunity to give a short lightning talk to a subset of accepted poster proposals.

The deadline for proposing is Sunday, August 9, 2026. If your proposal is accepted, you should plan to attend the event in-person.

In-Person Registration

Seats are limited at the venue, which means that in-person registration may be capped prior to the workshop start date. If capacity is reached, a waitlist will be imposed, which the registration form will reflect. Early registration is strongly encouraged.

All in-person registrants must wait to receive an invitation to attend in-person from IMSI before traveling, which generally begin to be sent out 4-6 weeks in advance.

All registrants (online and in-person) will receive zoom links and are welcome to attend online.

Registration Fee

A non-refundable registration fee will be payable by credit card or debit card for any participants invited to attend this workshop in-person. In-person participants agree to pay the non-refundable fee by the deadline given by IMSI. Failure to pay the fee by the deadline may mean that the invitation to attend in-person is revoked.

Current fees:

  • $25 for students
  • $50 for non-students

Organizers

Back to top
F B
Florentina Bunea Cornell University
L M
Li Ma University of Chicago
R W
Rebecca Willett University of Chicago
A Z
Anru Zhang Duke University

Speakers

Back to top
B A
Bryon Aragam University of Chicago, Booth School of Business
R B
Randall Balestriero Brown University
R B
Ricardo Baptista University of Toronto
P B
Peter Bartlett University of California, Berkeley and Google
X ( B
Xin (Mike) Bing University of Toronto
J ( C
Junhui (Jeff) Cai Notre Dame University
A D
Arnak Dalalyan ENSAE Paris
F K
Fred Koehler University of Chicago
J L
Jessica Li Fred Hutchinson Cancer Center, Biostatistics
T L
Tengyuan Liang University of Chicago
X L
Xihong Lin Harvard University, Biostatistics
Q L
Qiao Liu Yale University, Biostatistics
A R
Alessandro Rinaldo University of Texas at Austin
V R
Veronika Rockova University of Chicago, Econometrics and Statistics
J S
Jeremias Sulam Johns Hopkins University
A T
Alex Tong AITHYRA Research Institute for Biomedical Artificial Intelligence of the Austrian Academy of Sciences
B T
Brian Trippe Stanford University, Statistics
K W
Kaizheng Wang Columbia University
M W
Mengdi Wang Princeton University, Electrical & Computer Engineering
Y W
Yuexi Wang University of Illinois Urbana-Champaign, Statistics
X W
Xiao Wang Purdue University
Z W
Zhaoran Wang Northwestern University
H Z
Hongtu Zhu University of North Carolina-Chapel Hill, Biostatistics

Schedule

Monday, October 5, 2026
8:30-9:00 CDT
Breakfast/Welcome
9:00-10:00 CDT
Implicit versus explicit complexity regularization

Speaker: Peter Bartlett (University of California, Berkeley and Google)

10:00-10:05 CDT
Tech Break
10:05-11:05 CDT
What is a world model?

Speaker: Bryon Aragam (University of Chicago, Booth School of Business)

11:05-11:35 CDT
Coffee Break
11:35-12:35 CDT
A mixture of expert model for quantizing embedded text

Speaker: Xin (Mike) Bing (University of Toronto)

12:35-13:35 CDT
Lunch Break – Rubenstein Dining Room
13:35-14:35 CDT
Learning to augment statistical inference with generative models

Speaker: Kaizheng Wang (Columbia University)

14:35-15:00 CDT
Break
15:00-16:00 CDT
Poster Session/Social Hour at the IMSI Building | 1155 E 60th Street Chicago, IL 60637
Tuesday, October 6, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-9:55 CDT
Coreset-Induced Conditional Velocity Flow Matching

Speaker: Xiao Wang (Purdue University)

9:55-10:00 CDT
Tech Break
10:00-11:00 CDT
Optimal Self-Distillation in Ridge Regression

Speaker: Alessandro Rinaldo (University of Texas at Austin)

11:00-11:15 CDT
Coffee Break
11:15-12:15 CDT
Calibration methods for generative models with application to protein structure modeling

Speaker: Brian Trippe (University of Texas at Austin)

12:15-13:15 CDT
Lunch Break – Rubenstein Dining Room
13:15-14:10 CDT
Synthetic Data: From Generative AI to Valid and Powerful Statistical Inference

Speaker: Xihong Lin (Harvard University, Biostatistics)

14:10-14:15 CDT
Tech Break
14:15-15:15 CDT
Generative AI and Entrepreneurship

Speaker: Junhui (Jeff) Cai (Notre Dame University)

15:15-15:30 CDT
Coffee Break
15:30-16:30 CDT
Training World Models with Statistical Guarantees

Speaker: Randall Balestriero (Brown University)

Wednesday, October 7, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-10:00 CDT
Denoiser-Based Diffusion Models for Discrete Ordinal Data

Speaker: Ricardo Baptista (University of Toronto)

10:00-10:05 CDT
Tech Break
10:05-11:05 CDT
Distributional Shrinkage

Speaker: Tengyuan Liang (University of Chicago)

11:05-11:35 CDT
Coffee Break
11:35-12:35 CDT
AI-powered Bayesian Generative Modeling Approaches for Statistical Inference

Speaker: Qiao Liu (Yale University)

12:35-13:35 CDT
Lunch Break – Rubenstein Dining Room
13:35-14:35 CDT
FAST-Brain: A Flow-Aligned Spatio-Temporal Surrogate Brain Model

Speaker: Hongtu Zhu (University of North Carolina at Chapel Hill)

14:35-15:00 CDT
Coffee Break
15:00-16:00 CDT
Simulation-based inference via structured score matching

Speaker: Yuexi Wang (University of Illinois at Urbana-Champaign)

Thursday, October 8, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-10:00 CDT
TBA

Speaker: Veronika Rockova (University of Chicago, Econometrics and Statistics)

10:00-10:05 CDT
Tech Break
10:05-11:05 CDT
Beyond Scores: Proximal Diffusion Models

Speaker: Jeremias Sulam (Johns Hopkins University)

11:05-11:35 CDT
Coffee Break
11:35-12:35 CDT
Improved Guarantees for Langevin Monte Carlo with Average Smoothness

Speaker: Arnak Dalalyan (ENSAE Paris)

12:35-13:35 CDT
Lunch Break – 504 Foyer Dining
13:35-14:35 CDT
Generation Without Creation: What Can Generative Models Actually Add to Statistical Inference?

Speaker: Jessica Li (Fred Hutchinson Cancer Center)

14:35-15:00 CDT
Coffee Break
15:00-16:00 CDT
Scalable Conformation Sampling via Generative Modeling

Speaker: Alex Tong (Aithyra)

Friday, October 9, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-10:00 CDT
A Hierarchical Language Model with Predictable Scaling Laws

Speaker: Frederic Koehler (University of Chicago)

10:00-10:05 CDT
Tech Break
10:05-11:05 CDT
TBA

Speaker: Mengdi Wang (Princeton University, Electrical & Computer Engineering)

11:05-11:35 CDT
Coffee Break
11:35-12:35 CDT
TBA

Speaker: Zhaoran Wang (Northwestern University)

12:35-12:45 CDT
Workshop Survey and Close

Registration

IMSI is committed to making all of our programs and events inclusive and accessible. Contact [email protected] to request disability-related accommodations.

In order to register for this workshop, you must have an IMSI account and be logged in.