Due to the high interest in this workshop, this event will be held in two locations: David Rubenstein Forum at 1201 E 60th Street, Chicago, IL 60637 and an overflow space at the IMSI Building at 1155 E 60th Street Chicago, IL 60637
Generative models have rapidly become a central tool in modern AI and data science. At a high level, a generative model learns an underlying probability distribution from data and can sample from it to create synthetic yet realistic outputs: text, images, financial scenarios, molecules, patient records, climate simulations, and more. Recent advances such as variational autoencoders, generative adversarial networks, normalizing flows, diffusion models, and flow matching have delivered striking empirical performance. At the same time, many of the most pressing questions remain fundamentally statistical: What distribution is being learned, and under what assumptions is it identifiable? Which distributional features are easy or hard to capture (e.g., modes with complex geometry, rare events, and tail behavior)? How can we quantify uncertainty, control bias, and ensure calibration, especially in high-stakes settings where downstream decisions depend on faithful modeling of extremes?
These challenges become even sharper in domain-specific contexts. Open-ended text generation lacks a single “correct” output, making objective evaluation of quality, coherence, diversity, and fluency a critical open problem. In finance and other dependent-data regimes, correlations and selection effects in training data raise questions about generalization and downstream validity. Across application areas, statistics plays a key role in designing reliable metrics to evaluate and compare generative models, and in understanding the properties of common fine-tuning and alignment procedures. More broadly, determining when synthetic data is “good enough” for inference, prediction, or decision-making remains an open question, as do opportunities to use generative models for tasks such as anomaly and changepoint detection.
This workshop brings together statisticians, machine learning researchers, and practitioners from domains including language modeling, finance, biomedicine, and the natural sciences to develop a shared language and research agenda. The goal is to connect modern generative modeling techniques to classical statistical principles, while advancing theory, methodology, and practices that enable reliable deployment in real-world scientific and societal applications.
Some of the funding for this workshop is provided by the Stevanovich Center.
Poster Session
This workshop will include a poster session for early career researchers (including graduate students). In order to propose a poster, you must first register for the workshop, and then submit a proposal using the form that will become available on this page after you register. The registration form should not be used to propose a poster. The organizers may offer the opportunity to give a short lightning talk to a subset of accepted poster proposals.
The deadline for proposing is Sunday, August 9, 2026. If your proposal is accepted, you should plan to attend the event in-person.
In-Person Registration
Seats are limited at the venue, which means that in-person registration may be capped prior to the workshop start date. If capacity is reached, a waitlist will be imposed, which the registration form will reflect. Early registration is strongly encouraged.
All in-person registrants must wait to receive an invitation to attend in-person from IMSI before traveling, which generally begin to be sent out 4-6 weeks in advance.
All registrants (online and in-person) will receive zoom links and are welcome to attend online.
Registration Fee
A non-refundable registration fee will be payable by credit card or debit card for any participants invited to attend this workshop in-person. In-person participants agree to pay the non-refundable fee by the deadline given by IMSI. Failure to pay the fee by the deadline may mean that the invitation to attend in-person is revoked.
Bryon Aragam
University of Chicago, Booth School of Business
R
B
Randall Balestriero
Brown University
R
B
Ricardo Baptista
University of Toronto
P
B
Peter Bartlett
University of California, Berkeley and Google
X
(
B
Xin (Mike) Bing
University of Toronto
J
(
C
Junhui (Jeff) Cai
Notre Dame University
A
D
Arnak Dalalyan
ENSAE Paris
F
K
Fred Koehler
University of Chicago
J
L
Jessica Li
Fred Hutchinson Cancer Center, Biostatistics
T
L
Tengyuan Liang
University of Chicago
X
L
Xihong Lin
Harvard University, Biostatistics
Q
L
Qiao Liu
Yale University, Biostatistics
A
R
Alessandro Rinaldo
University of Texas at Austin
V
R
Veronika Rockova
University of Chicago, Econometrics and Statistics
J
S
Jeremias Sulam
Johns Hopkins University
A
T
Alex Tong
AITHYRA Research Institute for Biomedical Artificial Intelligence of the Austrian Academy of Sciences
B
T
Brian Trippe
Stanford University, Statistics
K
W
Kaizheng Wang
Columbia University
M
W
Mengdi Wang
Princeton University, Electrical & Computer Engineering
Y
W
Yuexi Wang
University of Illinois Urbana-Champaign, Statistics
X
W
Xiao Wang
Purdue University
Z
W
Zhaoran Wang
Northwestern University
H
Z
Hongtu Zhu
University of North Carolina-Chapel Hill, Biostatistics
Schedule
Monday, October 5, 2026
8:30-9:00 CDT
Breakfast/Welcome
9:00-10:00 CDT
Implicit versus explicit complexity regularization
Speaker: Peter Bartlett (University of California, Berkeley and Google)
The impressive performance of generative models trained using modern machine learning methods seems to arise through different mechanisms from those of classical statistical learning theory and mathematical statistics. Simple gradient methods, without any explicit effort to control model complexity, lead to excellent empirical performance. This talk will consider the statistical properties of solutions that are favored by gradient optimization methods, and in particular describe recent results comparing the finite-sample performance of early-stopped gradient methods with more classical regularization approaches.
10:00-10:05 CDT
Tech Break
10:05-11:05 CDT
What is a world model?
Speaker: Bryon Aragam (University of Chicago, Booth School of Business)
TBA
11:05-11:35 CDT
Coffee Break
11:35-12:35 CDT
A mixture of expert model for quantizing embedded text
Speaker: Xin (Mike) Bing (University of Toronto)
Token‑level contextual embeddings for text corpora are now readily available for both natural language and genetic sequences. We take a top‑down perspective and introduce CODL, a continuous, doubly-latent, version of a topic model that leverages these embeddings to uncover the latent domains represented in a corpus. The model is a very large $p$-mixture of high‑dimensional Gaussian distributions, where the mixture weights themselves are mixtures of $K$ softmax probability vectors supported on the latent $p$ Gaussian means. The softmax mixture atoms are the main directions explored in the embedding space, and provide a $K$-level quantization of the data. We develop a fitting strategy designed for direct theoretical analysis and interpretation. By first learning the many latent Gaussian means from the tokenized corpus, we construct a corpus‑specific, contextualy embedded, vocabulary. This vocabulary forms the support of a discrete $K$-Softmax mixture ensemble whose atoms encode the corpus’ latent domains in the embedding space. These learned latent atoms can then be mapped back to natural language—via the corpus‑specific vocabulary—enabling interpretable topic discovery. A central component of our procedure is a new method for learning the number of latent domains $K$ from data, prior to atom estimation. We show that this reduces to detecting whether a Softmax mixture is primarily concentrated on a single leading component. To achieve this, we introduce a detection criterion that directly estimates functionals of a Softmax mixing measure, bypassing the challenging task of estimating the mixing measure itself in high dimensions. Using the leading Softmax atom estimates as powerful warm starts, we provide end‑to‑end theoretical guarantees for the final EM‑based atom estimates, supported by empirical demonstrations.
12:35-13:35 CDT
Lunch Break – Rubenstein Dining Room
13:35-14:35 CDT
Learning to augment statistical inference with generative models
Speaker: Kaizheng Wang (Columbia University)
Generative models can produce seemingly unlimited synthetic data, yet discrepancies between synthetic and real populations can introduce bias and undermine statistical conclusions. For a new task with scarce or no real data, how much synthetic data can be safely used? This talk develops a general framework that leverages historical tasks to calibrate synthetic-data augmentation and quantify the resulting uncertainty. When no real observations are available, the framework adaptively selects the synthetic sample size to achieve nominal average-case coverage. In the application to LLM-generated survey responses, this calibrated size can be interpreted as the number of human respondents the LLM is effectively worth, providing a measure of its simulation fidelity. More generally, when real observations are available, the augmentation scheme is characterized jointly by the number of synthetic observations and the weight assigned to each. We learn a size–weight frontier that identifies the largest synthetic contribution compatible with reliable uncertainty quantification. Together, these results show how generative models can strengthen statistical analysis when real data are limited.
14:35-15:00 CDT
Break
15:00-16:00 CDT
Poster Session/Social Hour at the IMSI Building | 1155 E 60th Street Chicago, IL 60637
Flow matching drives many of today's generative models. The standard recipe is simple: learn the average velocity that carries noise to data, then integrate it. But averaging discards structure. When a point could plausibly move toward several destinations, the mean velocity points somewhere between them, and the multimodal geometry of the transport is lost. This talk asks what happens if we keep that structure. Rather than regressing the conditional mean velocity, CCVFM models the full conditional law of the velocity $V = X_1 - X_0$. What makes this tractable is a surrogate: replace the target by an entropy-regularized coreset, a Gaussian mixture with far fewer components than there are data points, and the induced conditional velocity law becomes available in closed form. A Sinkhorn-anchored coupling then pairs source and data velocities so that the surrogate marginal is preserved exactly and pairs concentrate on matched mixture components. The construction also closes a gap between theory and practice. Known rates for flow matching leave slack in the exponent, and the proofs rely on a diffusion-style path with the time interval partitioned into subintervals, each handled by a separately trained network. Deployed flow matching does none of this. We obtain the exact minimax exponent using a single global network on a linear path, with no time partition. Experiments on MNIST, CIFAR-10, ImageNet-32, and CelebA-HQ confirm the predicted regime.
9:55-10:00 CDT
Tech Break
10:00-11:00 CDT
Optimal Self-Distillation in Ridge Regression
Speaker: Alessandro Rinaldo (University of Texas at Austin)
Self-distillation (SD) is the process of retraining a student model on a mixture of ground-truth labels and a teacher model’s predictions, using the same architecture and training data. While SD has been empirically shown to improve generalization in regression tasks, a rigorous theoretical understanding of its mechanics and properties remains elusive. We study SD for ridge regression in the general unconstrained setting where the mixing weight is allowed to lie outside the unit interval. Without any distributional assumptions, we prove that the squared prediction risk — including the out-of-distribution risk — of the optimally mixed student model strictly improves upon the teacher model for every value of regularization at which the teacher's risk is non-stationary. We express the optimal mixing weight in terms of the teacher's risk derivative, thereby characterizing the somewhat surprising scenarios in which a negative weight is optimal. To quantify the magnitude of these improvements, we derive the limiting SD risk in the proportional asymptotic regime under general anisotropic covariances and deterministic signals. Building on this theory, we propose a novel one-shot tuning procedure to consistently estimate the optimal mixing weight without retraining, sample splitting, or grid search. We further investigate the use of SD in prediction-only settings in which training data are not available, and one has access only to the trained predictor and fresh unlabeled covariates. We show that the limiting optimally mixed student prediction risk is strictly smaller than the teacher risk for almost every pair of teacher and student regularization levels, including the case of out-of-distribution fresh covariates. Joint work with Hien Dang and Pratik Patil, from UT Austin.
11:00-11:15 CDT
Coffee Break
11:15-12:15 CDT
Calibration methods for generative models with application to protein structure modeling
Speaker: Brian Trippe (University of Texas at Austin)
Generative models frequently suffer miscalibration, where the statistics of their generations deviate from desired values. Protein structure diffusion models, for example, produce highly realistic samples but often underestimate the probabilities of important modes. These deviations are not addressed by existing fine-tuning methods for image and text generation. We frame calibration of generative models as a constrained optimization problem that we seek to solve by fine-tuning. Because the natural objective is intractable, we introduce two surrogate objectives for which we can compute low-variance gradient estimates amenable to stochastic optimization. The resulting procedures reduce the majority of calibration error across hundreds of simultaneous constraints and models with up to nine billion parameters. Lastly, we describe an application for assimilating thermodynamic measurements to calibrate a generative model of protein structure Boltzmann ensembles.
12:15-13:15 CDT
Lunch Break – Rubenstein Dining Room
13:15-14:10 CDT
Synthetic Data: From Generative AI to Valid and Powerful Statistical Inference
Speaker: Xihong Lin (Harvard University, Biostatistics)
Integration of statistics and generative AI plays a pivotal role for accelerating trustworthy cross-domain scientific discovery. Recent advances in generative models have dramatically increased the availability and use of synthetic data across scientific domains. While these developments create exciting opportunities for empowering data analysis, they also raise fundamental statistical challenges regarding how synthetic data can be used in a valid, reliable, and principled manner. In this talk, we first discuss the current landscape of high-fidelity high-dimensional synthetic data generation using generative AI models such as transformer, diffusion, and flow matching based models. More importantly, we present a principled framework for incorporating synthetic data in downstream statistical analysis that ensures valid statistical inference even when generative AI models are misspecified. We show that the proposed synthetic data assisted and augmented methods for continuous and discreate outcomes that integrate observed and synthetic data are robust to misspecified black-box generative models and can improve statistical inferential power when the generative AI models are informative. We demonstrate the utility of these synthetic data assisted methods to the analysis of the UK biobank data, by performing genome-wide association studies (GWAS) of proteomic data and whole-genome sequencing (WGS) analyses of brain imaging phenotypes, and binary mental health survey data both characterized by substantial missingness (about 70-90%).
14:10-14:15 CDT
Tech Break
14:15-15:15 CDT
Generative AI and Entrepreneurship
Speaker: Junhui (Jeff) Cai (Notre Dame University)
TBA
15:15-15:30 CDT
Coffee Break
15:30-16:30 CDT
Training World Models with Statistical Guarantees
Speaker: Randall Balestriero (Brown University)
This talk presents a sequence of recent results on Joint-Embedding Predictive Architectures (JEPAs), spanning self-supervised representation learning, action-conditioned world models, and counterfactual reasoning. We begin with LeJEPA, which addresses representation collapse through a distributional principle: embeddings are regularized toward an isotropic Gaussian using Sketched Isotropic Gaussian Regularization, producing a simple predictive objective with theoretical guarantees and without teacher–student networks or stop-gradient heuristics. We then introduce LeWorldModel, which extends this construction to sequential data by learning an encoder and action-conditioned latent transition model, (z_{t+1} approx f_theta(z_t,a_t)), end-to-end from raw observations. By combining next-embedding prediction with Gaussian latent regularization, the model supports planning and control directly in representation space, without reconstructing future pixels. Finally, we discuss new directions in counterfactual JEPAs, where object-level masking and latent interventions create structured partial-observability problems that force the predictor to model interaction-dependent dynamics rather than exploit local correlations. Throughout, we emphasize the geometric and statistical principles connecting these methods, and ask when latent prediction is sufficient for learning causal, controllable abstractions of a dynamical world.
Wednesday, October 7, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-10:00 CDT
Denoiser-Based Diffusion Models for Discrete Ordinal Data
Speaker: Ricardo Baptista (University of Toronto)
Discrete diffusion models provide a probabilistic framework for representing and sampling discrete data such as text sequences. For continuous-valued data, score-based diffusion models have rapidly become state of the art for tasks such as generating images and video. A key reason for their success is that estimating the score function, the gradient of the log-density of a perturbed data distribution, can be linked to a denoising problem through Tweedie's formula, which enables the use of supervised methods for learning the score. To date, diffusion models for discrete data have not matched this success and often require specialized loss functions and architectures. In this work, we introduce a new family of diffusion models for ordinal data that parallels continuous score-based generative modeling. Specifically, we propose processes for ordinal data that only require learning a denoiser and allow for bi-directional (up and down) perturbations of the data. The simplicity of the training objective, compared with approaches that train discrete scores, rates, or distributions, yields desirable theoretical properties and leads to competitive results across different data modalities.
10:00-10:05 CDT
Tech Break
10:05-11:05 CDT
Distributional Shrinkage
Speaker: Tengyuan Liang (University of Chicago)
This talk focuses on distributional shrinkage, a new phenomenon and statistical framework at the intersection of empirical Bayes, deconvolution, and generative modeling. It is based on a two-part series: “Distributional Shrinkage I: Universal Denoiser Beyond Tweedie’s Formula.” arXiv:2511.09500. “Distributional Shrinkage II: Higher-Order Scores Encode Brenier Map.” arXiv:2512.09295.
11:05-11:35 CDT
Coffee Break
11:35-12:35 CDT
AI-powered Bayesian Generative Modeling Approaches for Statistical Inference
Speaker: Qiao Liu (Yale University)
Flow matching drives many of today's generative models. The standard recipe is simple: learn the average velocity that carries noise to data, then integrate it. But averaging discards structure. When a point could plausibly move toward several destinations, the mean velocity points somewhere between them, and the multimodal geometry of the transport is lost. This talk asks what happens if we keep that structure. Rather than regressing the conditional mean velocity, CCVFM models the full conditional law of the velocity 𝑉 =𝑋1 −𝑋0. What makes this tractable is a surrogate: replace the target by an entropy-regularized coreset, a Gaussian mixture with far fewer components than there are data points, and the induced conditional velocity law becomes available in closed form. A Sinkhorn-anchored coupling then pairs source and data velocities so that the surrogate marginal is preserved exactly and pairs concentrate on matched mixture components. The construction also closes a gap between theory and practice. Known rates for flow matching leave slack in the exponent, and the proofs rely on a diffusion-style path with the time interval partitioned into subintervals, each handled by a separately trained network. Deployed flow matching does none of this. We obtain the exact minimax exponent using a single global network on a linear path, with no time partition. Experiments on MNIST, CIFAR-10, ImageNet-32, and CelebA-HQ confirm the predicted regime.
12:35-13:35 CDT
Lunch Break – Rubenstein Dining Room
13:35-14:35 CDT
FAST-Brain: A Flow-Aligned Spatio-Temporal Surrogate Brain Model
Speaker: Hongtu Zhu (University of North Carolina at Chapel Hill)
Modeling resting-state functional magnetic resonance imaging (rs-fMRI) data is crucial for understanding brain-wide neural activity. However, traditional methods struggle to capture complex temporal dynamics over long horizons, to account for the brain's anatomical spatial structure, and to model high-dimensional ambient signals that lie on a low-dimensional intrinsic subspace. We propose FAST-Brain, a unified flow-aligned spatio-temporal surrogate brain model that addresses all three challenges. At its core is a flow-aligned generative framework that directly predicts the clean blood-oxygen-level-dependent (BOLD) signal, paired with a graph convolutional network that captures spatial structural constraints and a Transformer that models long-range temporal dependencies. Theoretically, we show that under a low-dimensional subspace assumption, the approximation error of our model scales with the intrinsic dimension rather than the ambient dimension, which justifies our direct modeling of the BOLD signal. Extensive experiments on synthetic and Human Connectome Project datasets demonstrate that FAST-Brain achieves state-of-the-art performance in recovering functional connectivity, effective connectivity, and the implicit low-dimensional signal subspace.
14:35-15:00 CDT
Coffee Break
15:00-16:00 CDT
Simulation-based inference via structured score matching
Speaker: Yuexi Wang (University of Illinois at Urbana-Champaign)
Simulation-based inference (SBI) provides a powerful framework for statistical analysis when the likelihood function is intractable but simulations are available. I will present a unified framework that leverages structured score matching to enable both maximum likelihood estimation and Bayesian posterior sampling in such likelihood-free settings. The key idea is to approximate the intractable likelihood score function with neural networks that embed the statistical structures of likelihood score through architectural regularization. These structural constraints enhance estimation accuracy, scalability, and uncertainty quantification. We establish theoretical guarantees for the proposed methods and demonstrate their practical advantages on benchmark tasks and challenging high-dimensional problems, where they perform favorably against existing approaches.
Thursday, October 8, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-10:00 CDT
TBA
Speaker: Veronika Rockova (University of Chicago, Econometrics and Statistics)
Diffusion models have quickly become some of the most popular and powerful generative models for high-dimensional data. The key insight that enabled their development was the realization that access to the score (the gradient of the log-density at different noise levels) allows for sampling from data distributions by solving a reverse-time stochastic differential equation (SDE) via forward discretization, and that popular denoisers allow for unbiased estimators of this score. In this talk I will demonstrate that an alternative, backward discretization of these SDEs, using proximal maps in place of the score, leads to theoretical and practical benefits. We will leverage recent results in proximal matching to learn proximal operators of the log-density and, with them, develop Proximal Diffusion Models (ProxDM). Theoretically, and assuming oracle access to distributional quantities, these models provide faster convergence to target distributions. Empirically, I will show that ProxDM achieves significantly faster convergence within just a few sampling steps compared to conventional score-matching methods for unconditional, conditional, and latent diffusion models.
11:05-11:35 CDT
Coffee Break
11:35-12:35 CDT
Improved Guarantees for Langevin Monte Carlo with Average Smoothness
Speaker: Arnak Dalalyan (ENSAE Paris)
We establish improved nonasymptotic bounds for Langevin Monte Carlo in the strongly log-concave setting, when the error is measured by the Wasserstein distance. The main result shows that the discretization error is governed by an average coordinate-wise smoothness constant, rather than by the usual global smoothness constant. The proof is short and probabilistic, and relies on a refined use of the synchronous coupling. We further show that the same ideas lead to improved bounds for variable step sizes, for potentials whose Laplacian is Lipschitz-continuous, and for finite-sum problems sampled by stochastic-gradient Langevin dynamics with fixed point control variates. In the Laplacian-smooth case, the usual Hessian-Lipschitz contribution is replaced by a weaker trace-type third-order smoothness quantity. In the finite-sum setting, the resulting SGLD bound improves the dependence on the root mean square smoothness of the component functions. Applications to generalized linear models with Gaussian design show that these refinements can yield substantial, dimension-dependent improvements over previously known bounds, especially for correlated covariates. The talk is based on a joint work with Avetik Karagulyan. https://arxiv.org/abs/2605.31413
12:35-13:35 CDT
Lunch Break – 504 Foyer Dining
13:35-14:35 CDT
Generation Without Creation: What Can Generative Models Actually Add to Statistical Inference?
Speaker: Jessica Li (Fred Hutchinson Cancer Center)
Generative models are typically evaluated by the realism of their synthetic data. In biomedical research, however, the data themselves are rarely the ultimate object of interest. What matters is whether a generative model improves inference about effects of interest when data are scarce or difficult to collect, as in rare-disease trials, small patient subgroups, or rare cell states. This talk examines a fundamental limitation: any information a generative model contributes to an inference problem must come from its training data or from prior knowledge and assumptions encoded in the model, its architecture, or its training objective. Successful uses of generative models for inference include synthetic control arms in rare-disease trials and synthetic null distributions for false-discovery-rate control in differential-expression analysis. In these examples, synthetic data serve as flexible devices for encoding informative priors or statistical null hypotheses, respectively. When synthetic data are added to real data to improve power, the potential benefit follows a classical bias–variance trade-off: variance may decrease, but bias can increase when the encoded prior is misspecified. I relate this problem to prediction-powered inference, where a debiasing rectifier comparing predicted and observed labels is needed to obtain validity guarantees. Comparable guarantees are less straightforward for what one might call generation-powered inference. For example, single-cell foundation models trained on atlas data dominated by normal conditions face fundamental limitations in generating cells from previously unseen biological conditions and may underperform simple baselines. I contrast generation-powered inference with generation-calibrated inference, in which synthetic data are generated to reflect a null hypothesis under test. Whereas generation-powered inference requires scrutiny of the encoded prior and tools to quantify or control its induced bias, generation-calibrated inference is often more tractable: its goal is valid calibration under a specified null. Our work on scDesign3, ClusterDE, and Nullstrap illustrates this latter approach. I argue that, beyond realism, generative models should be judged by criteria appropriate to their inferential role, including calibration, coverage, bias–variance trade-offs, and robustness to discrepancies between the generative model and the true data-generating process.
14:35-15:00 CDT
Coffee Break
15:00-16:00 CDT
Scalable Conformation Sampling via Generative Modeling
Speaker: Alex Tong (Aithyra)
Sampling molecular conformations at thermodynamic equilibrium is a fundamental challenge in computational chemistry and statistical physics. While molecular dynamics and Monte Carlo methods provide principled solutions, their computational cost can make it difficult to explore high-dimensional, rugged energy landscapes, particularly for flexible peptides and larger molecular systems. In this talk, I will discuss recent advances in generative modeling for scalable molecular sampling. The central idea is to learn amortized samplers that can rapidly propose diverse, approximately independent molecular conformations, while retaining principled mechanisms for correcting model bias and recovering the target Boltzmann distribution. I will describe a progression of approaches based on normalizing flows, diffusion models, and sequential Monte Carlo, including methods for efficient likelihood evaluation, stable simulation-free training, inference-time annealing, and transfer across related molecular systems. These techniques lead to practical Boltzmann generators that operate directly in Cartesian coordinates and scale beyond the small systems traditionally used to benchmark learned samplers. In particular, sequential transport and importance-sampling corrections allow generative proposals to be refined toward equilibrium without sacrificing the speed advantages of amortization. I will also discuss transferable models that learn reusable sampling algorithms across peptide sequences and molecular sizes. Together, these results suggest a new paradigm for molecular simulation: rather than paying the full exploration cost independently for every system, we can train generative models to amortize conformation sampling across families of related energy landscapes. This combination of learned generation, physical correction, and inference-time scaling offers a path toward faster and more scalable equilibrium sampling for molecular discovery and design.
Friday, October 9, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-10:00 CDT
A Hierarchical Language Model with Predictable Scaling Laws
Speaker: Frederic Koehler (University of Chicago)
I will discuss recent work studying the behavior of standard transformer models when they are trained on a mathematically natural hierarchical language arising from the generalized broadcast model on trees. Based on joint work with Jason Gaitonde, Elchanan Mossel, Joonhyung Shin, and Allan Sly.
10:00-10:05 CDT
Tech Break
10:05-11:05 CDT
TBA
Speaker: Mengdi Wang (Princeton University, Electrical & Computer Engineering)
TBA
11:05-11:35 CDT
Coffee Break
11:35-12:35 CDT
TBA
Speaker: Zhaoran Wang (Northwestern University)
TBA
12:35-12:45 CDT
Workshop Survey and Close
Registration
IMSI is committed to making all of our programs and events inclusive and accessible.
Contact [email protected] to request
disability-related accommodations.
In order to register for this workshop, you must have an IMSI account and be logged in.