Learning Pipelines for Adaptive Control
Florian Dorfler, University of Pennsylvania
The adjacent fields of reinforcement learning (RL) and adaptive control share the same objectives, yet they are separated by a wide cultural gap. In this presentation, I attempt to bridge this gap for the linear quadratic regulator (LQR) problem, which serves as a cornerstone and the benchmark for both fields. I begin by discussing different learning pipelines, including direct and indirect (model-based) approaches, as well as episodic and online (adaptive) approaches. Despite the extensive literature spanning several decades, numerous problems remain unsolved. For instance, RL methods are seldom concerned with closed-loop stability certificates or efficient implementations, while the adaptive control community has dedicated minimal effort to optimality. We address the data-driven LQR problem in an adaptive setting, which entails online recursive algorithms and closed-loop data, and we seek both algorithmic as well as closed-loop certificates. Our approach encompasses different variations of policy gradient methods and employs a novel covariance parameterization of the LQR problem. Finally, all our theoretical results are validated through simulations and experiments in diverse domains, demonstrating the computational and sample efficiency of our method.