Baroncini, Simone
(2026)
Control-theoretic methods for reinforcement learning, [Dissertation thesis], Alma Mater Studiorum Università di Bologna.
Dottorato di ricerca in
Ingegneria biomedica, elettrica e dei sistemi, 38 Ciclo.
Documenti full-text disponibili:
![baroncini_simone_tesi.pdf [thumbnail of baroncini_simone_tesi.pdf]](https://amsdottorato.unibo.it/style/images/fileicons/application_pdf.png) |
Documento PDF (English)
- Richiede un lettore di PDF come Xpdf o Adobe Acrobat Reader
Disponibile con Licenza: Salvo eventuali più ampie autorizzazioni dell'autore, la tesi può essere liberamente consultata e può essere effettuato il salvataggio e la stampa di una copia per fini strettamente personali di studio, di ricerca e di insegnamento, con espresso divieto di qualunque utilizzo direttamente o indirettamente commerciale. Ogni altro diritto sul materiale è riservato.
Download (11MB)
|
Abstract
Reinforcement Learning (RL) has become a cornerstone of modern artificial intelligence and control theory, enabling complex decision-making tasks in a wide variety of domains such as robotics, autonomous driving and strategic games. While deep neural networks allow RL methods to address large-scale problems, exploiting the mathematical structures underlying these problems remains essential for designing efficient and theoretically grounded algorithms. This thesis bridges two complementary perspectives, mathematical analysis of reinforcement learning problems and control-theoretic principles, to develop a rigorous foundation for the design and the analysis of novel reinforcement learning algorithms. The first part of this work investigates a class of learning algorithms that generalize the classical RL paradigm of concurrent policy evaluation and policy improvement. Extending the classical two-timescale framework, a third concurrent process is introduced to incorporate system measurements. Using tools based on singular perturbations theory and LaSalle's invariance principle, sufficient conditions for the convergence of this class of methods are derived. Two novel actor-critic variants are then proposed to address two scenarios with different available measurements from the system. The second part focuses on reward-balancing methods, a recently proposed family of algorithms that approach RL from a dual viewpoint. Unlike standard methods that iteratively improve the policy while keeping the reward fixed, these methods fix the policy to be greedy and iteratively adjust the reward function to make that policy optimal, without altering the underlying problem. This thesis provides a detailed characterization of the transformation group underlying this process and introduces a control-theoretic reformulation of the approach. The resulting framework enables principled inclusion of constraints and penalty terms to improve robustness under model uncertainty. Overall, by exploiting the inherent generality of systems theory, this thesis takes a step toward a unifying framework for the design and the analysis of RL algorithms that leverage their underlying structural properties.
Abstract
Reinforcement Learning (RL) has become a cornerstone of modern artificial intelligence and control theory, enabling complex decision-making tasks in a wide variety of domains such as robotics, autonomous driving and strategic games. While deep neural networks allow RL methods to address large-scale problems, exploiting the mathematical structures underlying these problems remains essential for designing efficient and theoretically grounded algorithms. This thesis bridges two complementary perspectives, mathematical analysis of reinforcement learning problems and control-theoretic principles, to develop a rigorous foundation for the design and the analysis of novel reinforcement learning algorithms. The first part of this work investigates a class of learning algorithms that generalize the classical RL paradigm of concurrent policy evaluation and policy improvement. Extending the classical two-timescale framework, a third concurrent process is introduced to incorporate system measurements. Using tools based on singular perturbations theory and LaSalle's invariance principle, sufficient conditions for the convergence of this class of methods are derived. Two novel actor-critic variants are then proposed to address two scenarios with different available measurements from the system. The second part focuses on reward-balancing methods, a recently proposed family of algorithms that approach RL from a dual viewpoint. Unlike standard methods that iteratively improve the policy while keeping the reward fixed, these methods fix the policy to be greedy and iteratively adjust the reward function to make that policy optimal, without altering the underlying problem. This thesis provides a detailed characterization of the transformation group underlying this process and introduces a control-theoretic reformulation of the approach. The resulting framework enables principled inclusion of constraints and penalty terms to improve robustness under model uncertainty. Overall, by exploiting the inherent generality of systems theory, this thesis takes a step toward a unifying framework for the design and the analysis of RL algorithms that leverage their underlying structural properties.
Tipologia del documento
Tesi di dottorato
Autore
Baroncini, Simone
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Reinforcement Learning, Control Theory, Optimal Control, Actor Critic, Timescale Separation, Reward-Balancing Methods
Data di discussione
9 Aprile 2026
URI
Altri metadati
Tipologia del documento
Tesi di dottorato
Autore
Baroncini, Simone
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Reinforcement Learning, Control Theory, Optimal Control, Actor Critic, Timescale Separation, Reward-Balancing Methods
Data di discussione
9 Aprile 2026
URI
Statistica sui download
Gestione del documento: