The goal of this workshop is to provide an introduction to the mathematical, statistical, and computational frameworks for the analysis of samples of data residing in non-Euclidean metric spaces. Motivating applications span various scientific disciplines, with an abundance of such examples to be found in brain connectomics. These include networks and graphs, shapes, and probability distributions quantifying functional connectivity.
Topics of interest include foundational tools for exploratory data analysis, modeling, and statistical inference for non-Euclidean data. For instance, the Fréchet mean and variance typically serve as quantifications of the center and spread of data distributions, both at the population and sample level. Theory and practice will each be emphasized throughout the workshop to justify the methodologies. Theoretical properties of interest include finite-sample concentration bounds and asymptotic properties for the evaluation of estimators. Computational algorithms that implement the methodologies will also be demonstrated using complex, real-world data sets.
Lightning Talks and Poster Session
This workshop will include lightning talks and a poster session for early career researchers (including graduate students). In order to propose a lightning talk and poster, you must first register for the workshop, and then submit a proposal using the form that will become available on this page after you register. The registration form should not be used to propose a lightning talk or poster.
The deadline for proposingis Thursday September 10, 2026.. If your proposal is accepted, you should plan to attend the event in-person.
In-Person Registration
Seats are limited at the venue, which means that in-person registration may be capped prior to the workshop start date. If capacity is reached, a waitlist will be imposed, which the registration form will reflect. Early registration is strongly encouraged.
All in-person registrants must wait to receive an invitation to attend in-person from IMSI before traveling, which generally begin to be sent out 4-6 weeks in advance.
All registrants (online and in-person) will receive zoom links and are welcome to attend online.
Registration Fee
A non-refundable registration fee will be payable by credit card or debit card for any participants invited to attend this workshop in-person. In-person participants agree to pay the non-refundable fee by the deadline given by IMSI. Failure to pay the fee by the deadline may mean that the invitation to attend in-person is revoked.
Bharath Sriperumbudur
Pennsylvania State University
W
Z
Wenjun Zhao
Wake Forest University
Y
Z
Yize Zhao
Yale University
R
Z
Ralf Zimmermann
University of Southern Denmark
Schedule
Monday, October 19, 2026
8:30-8:55 CDT
Check-in/Breakfast
8:55-9:00 CDT
Welcome Remarks
9:00-9:45 CDT
Power Frechet Means in Hadamard Spaces
Speaker: Christof Schötz (Technical University Munich (TUM))
The 2-Frechet mean generalizes the expectation to metric spaces by minimizing the expected squared distance to a random object, while the 1-Frechet mean generalizes the median. $alpha$-Frechet means with $1<alpha<2$ interpolate between these two notions and inherit corresponding intermediate robustness properties. For Hadamard spaces (complete metric spaces of non-positive curvature), finite-sample bounds in expectation are established for empirical $alpha$-Frechet means under the sole assumption of a finite moment of order $alpha$. The bounds are governed by the moment of order $2alpha-2$, which is lower than order $alpha$, making power Frechet means particularly attractive for heavy-tailed distributions. The analysis extends beyond power functions to a broad class of non-decreasing convex transformations with concave derivative. The results apply in infinite-dimensional settings and are new even in Euclidean spaces.
9:45-10:00 CDT
Q&A
10:00-10:05 CDT
Tech Break
10:05-10:50 CDT
Flag spaces for Geometric statistics
Speaker: Xavier Pennec (Institut National de Recherche en Informatique et Automatique (INRIA))
Flags are sequences of properly embedded linear subspaces. They appear in multiscale dimension reduction methods such as principal or independent component analysis. Non-linear flags also appear to be the right geometric objects to work with for generalizations of PCA to manifolds such as principal nested spheres or barycentric subspace analysis. One can actually show that extracting order principal components actually optimizes a criterion on the flag space and not on Grassmannians as usually thought. Flag spaces are Riemannian homogeneous spaces that generalize Grassmmann and Steifel manifolds. However, they are usually not symmetric. In this talk, I will first present the extension of PCA to manifolds with flags of barycentric subspaces. In the second part, I will extend classical PCA in Euclidean spaces with non-complete flags of linear subspaces. The resulting Principal Subspace Analysis (PSA) method turns out to be much more stable that the classical PCA decomposition into unidimensional modes. This part is joint work with Tom Szwagier.
10:50-11:05 CDT
Q&A
11:05-11:35 CDT
Coffee Break
11:35-12:20 CDT
On retraction maps onto Grassmann, Stiefel, and symplectic Stiefel, and the question of to what extent one can trust one’s local coordinates.
Speaker: Ralf Zimmermann (University of Southern Denmark)
In Riemannian computing applications, it is crucial to map manifold data to a Euclidean domain, where vector space arithmetic is available, and back. Classical manifold theory guarantees the existence of such mappings, called charts and parameterizations, or, collectively, local coordinates. When computational efficiency is of the essence, practitioners usually resort to retraction maps to define local coordinates. Retractions yield first-order approximations of the Riemannian normal coordinates. This talk we discuss known and new retraction maps on the Grassmann manifold, the Stiefel manifold and the symplectic Stiefel manifold. We expose numerical properties, in particular condition numbers and error propagation between the coordinate domains and the manifold. Next, we will discuss two specific retraction maps, namely, the only known ones that exhibit second-order accuracy under certain metrics while still allowing for an explicit formula for their inverses. As an application, we will consider the interpolation of manifold-valued functions.
12:20-12:35 CDT
Q&A
12:35-13:35 CDT
Lunch
13:35-14:20 CDT
Estimation of barycenters in CAT spaces
Speaker: Victor-Emmanuel Brunel (ENSAE)
Barycenters provide a canonical extension of classical expectations for random variables with values in non-Euclidean metric spaces. In this talk, we will present new results on the estimation of barycenters, given iid random variables, when the ambient space satisfies a global curvature upper bound. When the data generating distribution is supported on a small enough ball, its barycenter is the solution to a geodesically strongly convex optimization problem. This fact allows one to obtain high probability estimation error bounds in a fairly simple manner. By resorting to more involved tools, we are also able to cover the case where the distribution is supported on larger balls (e.g., on a whole open hemisphere of a Euclidean ball). Statistical guarantees are obtained for both empirical barycenters and stochastic proximal algorithms (yielding iterated barycenters), which are more tractable and allow for streaming data in online setups. The talk is based on join works with Yassine Boukhateb (PhD student, ENSAE), Jordan Serres (Sorbonne) and Austin Stromme (ENSAE).
14:20-14:35 CDT
Q&A
14:35-14:40 CDT
Tech Break
14:40-15:00 CDT
Lightning Talks
15:00-16:00 CDT
Social Hour & Poster Session
Tuesday, October 20, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-9:45 CDT
Metric Statistics: Distribution-valued Processes
Speaker: Hans-Georg Müller (University of California, Davis (UC Davis))
Statistical methods for data situated in general metric spaces are not yet well developed. Such data have been referred to as Random Objects and the field dealing with them as Metric Statistics. Relevant basic tools include measures of location and spread as well as regression models. Following a review of some of these basic tools and the challenges arising due to the absence of algebraic operations, the focus is a scenario where one aims to analyze samples of realizations of a distribution-valued stochastic processes. Specific examples of longitudinally observed one-dimensional distributions include age-at-death distributions in demography. We consider an intrinsic optimal transport model. In an initial centering step the observed distributions are converted to optimal transports from the barycenter at each sampling time t. Utilizing a transport algebra, each realization of the process is decomposed into a time-independent transport and a real-valued random trajectory. A connection with real-valued Gaussian processes and functional principal component analysis makes it possible to consider not only continuously sampled distributions but also distributions obtained under sparse/longitudinal sampling designs, where due to the sparse and irregular sampling the underlying distributional process remains latent and cannot be directly observed. Based on joint work with Hang Zhou (University of North Carolina).
9:45-10:00 CDT
Q&A
10:00-10:05 CDT
Tech Break
10:05-10:50 CDT
DiPMInd: Distance Profile based Mutual Independence testing for random objects
Speaker: Yaqing Chen (Rutgers University)
We develop a novel unified framework for testing mutual independence among random objects residing in possibly different metric spaces. The framework generalizes existing methodologies and introduces new measures of mutual independence, and proposes associated tests that achieve minimax rate optimality and exhibit strong empirical power. The foundation of the proposed tests is the new concept of joint distance profiles, which uniquely characterize the joint law of random objects under a mild condition on either the joint law or the metric spaces. Our test statistics quantify the difference of the joint distance profiles of each data point with respect to the joint law and the product of marginal laws of the vector of random objects. To enhance power, we consider integrating this difference with respect to different measures and incorporate flexible data-adaptive weight profiles in the test statistics. We derive the limiting distribution of the test statistics under the null hypothesis of mutual independence and show that the proposed tests with certain weight profiles are asymptotically distribution-free if the marginal distance profiles are continuous. Furthermore, we establish the consistency of the tests under sequences of alternative hypotheses converging to the null. For practical implementations, we employ a permutation scheme to approximate the $p$-values and provide theoretical guarantees that the permutation-based tests maintain type I error control under the null and achieve consistency under alternatives. We demonstrate the power of the proposed tests across various types of data objects through simulations and real data applications, where the new tests exhibit better performance compared with popular existing approaches.
10:50-11:05 CDT
Q&A
11:05-11:35 CDT
Coffee Break
11:35-12:20 CDT
Hypothesis tests in metric spaces
Speaker: Helle Sørensen (University of Copenhagen)
Hypothesis testing is one of the core tools in statistical inference and is also relevant for so-called random objects, i.e., data that reside in general (non-Euclidian) metric spaces. Computations must be based on the distance function only, not on algebraic operations applied to the data objects themselves. I am going to talk about hypothesis tests in two specific situations. In the first situation we consider an auto-regressive model for time series of random objects and test the hypothesis about independence over time. The test can be carried out as a permutation test. In the second situation we consider a sample of iid random objects and the hypothesis that the common Frechet mean is identical to a hypothesized value. The test is carried out as a randomization test and requires symmetry structures on the metric space. The talk is based on joint work with Matthieu Bulte.
12:20-12:35 CDT
Q&A
12:35-13:35 CDT
Lunch
13:35-14:20 CDT
DFNN: A Deep Fréchet Neural Network Framework for Learning Metric-Space-Valued Responses
Speaker: Paromita Dubey (University of Southern California (USC))
Regression with non-Euclidean responses—e.g., probability distributions, networks, symmetric positive-definite matrices, and compositions—has become increasingly important in modern applications. In this paper, we propose deep Fréchet neural networks (DFNNs), an end-to-end deep learning framework for predicting non-Euclidean responses—which are considered as random objects in a metric space—from Euclidean predictors. Our method utilizes the representation-learning power of deep neural networks (DNNs) to the task of approximating conditional Fréchet means of the response given the predictors, the metric-space analogue of conditional expectations, by minimizing a Fréchet risk. The framework is highly flexible, accommodating diverse metrics and high-dimensional predictors. We establish a universal approximation theorem for DFNNs, advancing the state-of-the-art of neural network approximation theory to general metric-space-valued responses,
without making model assumptions or relying on local smoothing. We further establish rigorous generalization guarantees for DFNNs and derive corresponding risk bounds, providing, to the best of our knowledge, the first such theoretical results for deep learning regression with metric-space-valued responses. Empirical studies on synthetic distributional and network-valued
responses, as well as real-world applications to predicting compositional responses in an Aitchison simplex and spherical responses, demonstrate that DFNNs consistently outperform all existing methods.
14:20-14:35 CDT
Q&A
14:35-15:00 CDT
Coffee Break
15:00-15:45 CDT
Nonparametric Riemannian Empirical Bayes, and Denoising Measurements on Manifolds
Speaker: Adam Quinn Jaffe (Columbia University)
We initiate the study of nonparametric empirical Bayes denoising methods in the setting where both the latent variables and their measurements lie on a compact Riemannian manifold, and where the likelihood is a Riemannian Gaussian distribution. Our starting point is a novel Tweedie-Eddington formula for Riemannian Gaussian mixture models which identifies a certain surrogate oracle denoiser in terms of the marginal distribution of the measurements; it avoids the explicit computation of the posterior Fréchet mean (as required by the Bayes denoiser) via a first-order approximation, hence we refer to it as the "tangential" Bayes denoiser. We show that this surrogate oracle achieves nearly the Bayes risk in a low-noise regime, we construct a fully data-driven approximation of it using the spectral theory of the Laplace-Beltrami operator, and we establish finite-sample rates of convergence for the distance between the the surrogate oracle and its approximation. Contrasting the nearly-parametric rates from the Euclidean setting, the rates in the Riemannian setting are slower due to the singularities of the Riemannian Gaussian density at the cut locus of its Fréchet mean; in the special case of the circle we establish matching lower bounds which show that our proposed denoiser is minimax-optimal, and that the denoising problem exhibits a genuinely nonparametric rate of convergence. Lastly, we implement our methodology in two scientific applications: in astronomy, the sphere-valued problem of denoising the locations of gamma ray bursts; in structural biology, the torus-valued problem of denoising pairs of torsion angles of adjacent amino acids in a protein (i.e., the Ramachandran plot).
15:45-16:00 CDT
Q&A
Wednesday, October 21, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-9:45 CDT
Data Integration Via Analysis of Manifolds (DIVAM)
Speaker: J. S. (Steve) Marron (University of North Carolina Chapel Hill)
A major challenge in the age of Big Data is the integration of disparate data types into a single data analysis. That was tackled by Data Integration Via Analysis of Subspaces (DIVAS) in the context of data blocks measured on a common set of experimental cases. Joint variation was defined in terms of modes of variation having identical scores across data blocks. DIVAS allowed mathematically rigorous formulation of individual variation within each data block in terms of individual modes. The goal of DIVAM is to intrinsically extend the DIVAS approach to data objects lying in manifolds, such as shape data.
9:45-10:00 CDT
Q&A
10:00-10:05 CDT
Tech Break
10:05-10:50 CDT
Frechet Mean Estimation: Beyond Sample Frechet Means
Speaker: Andrew McCormack (University of Alberta)
10:50-11:05 CDT
Q&A
11:05-11:35 CDT
Coffee Break
11:35-12:20 CDT
Beyond Euclidean summaries: distributional methods for wearable device data
Speaker: Irina Gaynanova (University of Michigan)
Wearable devices such as continuous glucose monitors and accelerometers record thousands of measurements per subject, yet analyses typically reduce these recordings to a handful of scalar summaries, discarding most of the available information. An appealing alternative is to treat each subject's measurements as a sample from a subject-specific distribution and to take that distribution itself as the unit of analysis. Doing so moves the problem out of Euclidean space: distributions equipped with a transport-based metric form a metric space rather than a vector space, so standard tools for regression, dimension reduction, and interpretation do not directly apply. In this talk I will discuss statistical and computational challenges that arise in this setting, including efficient estimation and variable selection with distribution-valued responses, construction of interpretable low-dimensional summaries that respect the underlying geometry, and extensions beyond the univariate case where optimal transport becomes computationally prohibitive. The methods will be illustrated on data from continuous glucose monitoring studies.
12:20-12:35 CDT
Q&A
12:35-13:35 CDT
Lunch
13:35-14:20 CDT
TBA
Speaker: Olivier Bisson (Institut National de Recherche en Informatique et Automatique (INRIA))
14:20-14:35 CDT
Q&A
14:35-15:00 CDT
Coffee Break
15:00-15:45 CDT
TBA
Speaker: Ian L. Dryden (University of Nottingham)
15:45-16:00 CDT
Q&A
Thursday, October 22, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-9:45 CDT
Duality, Statistics and Geometry of Gromov-Wasserstein Distance
Speaker: Bharath Sriperumbudur (Pennsylvania State University)
The Gromov-Wasserstein (GW) distance, rooted in optimal transport (OT) theory, quantifies dissimilarity between metric measure spaces and provides a natural framework for aligning them. As such, GW distance enables applications including object matching, single-cell genomics, and matching language models. While computational aspects of the GW distance have been studied heuristically, most of the mathematical theories about GW duality, Brenier maps, geometry, etc., remained elusive, despite the rapid progress these aspects have seen under the classical OT paradigm in recent decades. This talk will cover recent progress on closing these gaps for the GW. We present (i) sharp statistical estimation rates through duality, (ii) a thorough investigation of the Jordan-Kinderlehrer-Otto (JKO) scheme for the gradient flow of inner product GW (IGW) distance, and (iii) a dynamical formulation of IGW, which generalizes the Benamou-Brenier formula for the Wasserstein distance. Central to (ii) and (iii) is a Riemannian structure on the space of probability distributions, based on which we also propose novel numerical schemes for measure evolution and deformation. [Joint work with Zhengxin Zhang (Cornell), Ziv Goldfeld (Cornell), Kristjan Greenewald (IBM Research), and Youssef Mroueh (IBM Research)]
9:45-10:00 CDT
Q&A
10:00-10:05 CDT
Tech Break
10:05-10:50 CDT
TBA
Speaker: Karthik Bharath (University of Nottingham)
10:50-11:05 CDT
Q&A
11:05-11:35 CDT
Coffee Break
11:35-12:20 CDT
Some Statistical Challenges and Approaches in Non-Euclidean and Distribution-Valued Omics Data
Speaker: Wenjun Zhao (Wake Forest University)
Modern omics technologies increasingly produce data that are not naturally represented as vectors. Examples include distributions of cells, geometric shapes, spatially resolved molecular measurements, and multimodal observations in which geometry is coupled with high-dimensional features. These settings raise statistical questions about how to compare, summarize, and align structured observations while respecting their non-Euclidean geometry.
In this talk, I will discuss several approaches motivated by these settings. I will consider problems ranging from comparing large collections of biological shapes, to analyzing shapes equipped with features supported on their geometry, to summarizing spatial transcriptomic tissue sections in which molecular measurements are represented as distributions over geometric domains. Most of the approaches I will discuss are based on optimal transport. Along the way, I will introduce the relevant biological settings and highlight practical challenges that arise in these applications, including heterogeneity, incomplete observations, and geometric distortions such as slicing bias. I hope the talk will also provide a starting point for discussion with researchers working on related non-Euclidean data problems, and I would be very happy to explore potential connections and collaborations.
12:20-12:35 CDT
Q&A
12:35-13:35 CDT
Lunch
13:35-14:20 CDT
TBA
Speaker: Lizhen Lin (University of Maryland)
14:20-14:35 CDT
Q&A
14:35-15:00 CDT
Coffee Break
15:00-15:45 CDT
TBA
Speaker: Luís Pereira (University of California, Santa Barbara (UCSB))
15:45-16:00 CDT
Q&A
Friday, October 23, 2026
8:30-9:00 CDT
Check-in/Breakfast
9:00-9:45 CDT
Harmonic Map Regression: Rate-Optimal Nonparametric Estimation on Manifolds with Topological Recovery
Speaker: Xiaoyu Chen (University at Buffalo)
We study harmonic map regression, a nonparametric estimator for manifold-valued responses, that penalizes the empirical Fréchet risk by the Dirichlet energy. By connecting penalized regression to the theory of harmonic maps, the estimator acquires a structural theory that parallels the classical Euclidean smoothing spline. The Euler-Lagrange equation characterizes the solution as a piecewise-geodesic spline, an equivalent kernel controls pointwise risk at the rate n^?2/3, and the infinite-dimensional variational problem reduces exactly to a finite-dimensional optimization. Such newly established connection reveals a topological phenomenon that has no analogue in Euclidean nonparametric regression and, to our knowledge, has not been studied in the manifold regression literature. On manifolds whose regression curves can wrap around in topologically distinct ways, maps in distinct homotopy classes are separated by energy barriers intrinsic to the geometry of the target, and the Dirichlet penalty makes the estimator sensitive to this structure, recovering the correct topological class with probability tending to one, a phase transition we call topological recovery. A curvature-dependent oracle inequality yields the minimax rate n^?2s/(2s+1) for Sobolev order s, matching the Euclidean constant on non-positively curved targets. Simulations on five Riemannian manifolds corroborate the theory, and an application to wind-direction data on S1 illustrates practical advantages.
9:45-10:00 CDT
Q&A
10:00-10:30 CDT
Coffee Break
10:30-11:15 CDT
A novel Bayesian framework uncovering brain connectivity-to-shape relationship in preclinical Alzheimer’s disease
Speaker: Yize Zhao (Yale University)
Alzheimer’s disease (AD) is a progressive neurodegenerative disorder characterized by amyloid-beta plaques and tau tangles, with significant pathological changes occurring in subcortical brain regions. While previous research has focused primarily on volumetric reductions in areas such as the hippocampus, thalamus, and caudate, emerging evidence suggests that their fine-grained shape deformations may offer greater sensitivity to early disease pathology. Moreover, understanding how these shape alterations influence brain functional connectivity (FC) networks could provide critical insights into the neurobiological mechanisms underlying the progression of AD. In this context, we propose a novel statistical approach, the Connectivity-on-Shape Regression (COSR) model, designed to investigate the spatially varying impact of brain subcortical shape on FC, accounting for the intrinsic modularity of functional networks. Under a Bayesian framework, COSR employs a relaxed-thresholded Gaussian process prior model to promote feature selection and integrates a stochastic block model to capture the unknown modular organization of FC. To facilitate the practical application of COSR with vertex-level shape measurements, we develop a computationally efficient variational inference approach to achieve posterior inference. Extensive simulations demonstrate the superiority of COSR over existing alternatives in accurately uncovering connectivity-to-shape associations and identifying neurobiological signals. Applying COSR to data from the Anti-Amyloid Treatment in Asymptomatic Alzheimer’s study, we discover meaningful neural structural-functional relationships in amyloid-positive individuals, highlighting the potentially complex interplay between structural and functional brain alterations during this crucial preclinical stage of AD.
11:15-11:30 CDT
Q&A
11:30-11:45 CDT
Workshop Survey and Closing Remarks
Registration
IMSI is committed to making all of our programs and events inclusive and accessible.
Contact [email protected] to request
disability-related accommodations.
In order to register for this workshop, you must have an IMSI account and be logged in.