MD

Moritz Diehl

info

Please Note

8 records found

Actor-Critic Reinforcement Learning for Guiding Model Predictive Control

Journal article (2026) - Rudolf Reiter, Andrea Ghezzi, Katrin Baumgartner, Jasper Hoffmann, Robert D. McAllister, Moritz Diehl
Nonlinear model predictive control (MPC) and reinforcement learning (RL) are two powerful control strategies with complementary advantages. This work shows how actor-critic RL techniques can be leveraged to improve the performance of MPC. The RL critic is used as an approximation of the optimal value function, and an actor rollout provides an initial guess for the primal variables of the MPC. A parallel control architecture is proposed where each MPC instance is solved twice for different initial guesses. Besides the actor rollout initialization, a shifted initialization from the previous solution is used. The control actions from the lowest-cost trajectory are applied to the system at each time step. We provide some theoretical justification of the proposed algorithm by establishing that the discounted closed-loop cost is upper-bounded by the discounted closed-loop cost of the original RL actor plus an error term that depends on the (sub)optimality of the RL actor and the accuracy of the critic. These results do not require globally optimal solutions and indicate that larger horizons mitigate the effect of errors in the critic approximation. The proposed algorithm is intended for applications where standard methods to construct terminal costs or constraints for MPC are impractical. The approach is demonstrated in an illustrative toy example and an autonomous driving overtaking scenario. ...
Review (2021) - Chris Vermillion, Mitchell Cobb, Lorenzo Fagiano, Rachel Leuthold, Moritz Diehl, Roy S. Smith, Tony A. Wood, Sebastian Rapp, Roland Schmehl, More Authors...
Airborne wind energy systems convert wind energy into electricity using tethered flying devices, typically flexible kites or aircraft. Replacing the tower and foundation of conventional wind turbines can substantially reduce the material use and, consequently, the cost of energy, while providing access to wind at higher altitudes. Because the flight operation of tethered devices can be adjusted to a varying wind resource, the energy availability increases in comparison to conventional wind turbines. Ultimately, this represents a rich topic for the study of real-time optimal control strategies that must function robustly in a spatiotemporally varying environment. With all of the opportunities that airborne wind energy systems bring, however, there are also a host of challenges, particularly those relating to robustness in extreme operating conditions and launching/landing the system (especially in the absence of wind). Thus, airborne wind energy systems can be viewed as a control system designer’s paradise or nightmare, depending on one’s perspective. This survey article explores insights from the development and experimental deployment of control systems for airborne wind energy platforms over approximately the past two decades, highlighting both the optimal control approaches that have been used to extract the maximal amount of power from tethered systems and the robust modal control approaches that have been used to achieve reliable launch, landing, and extreme wind operation. This survey will detail several of the many prototypes that have been deployed over the last decade and will discuss future directions of airborne wind energy technology as well as its nascent adoption in other domains, such as ocean energy. ...
Conference paper (2020) - Fabian Girrbach, Manon Kok, Raymond Zandbergen, Tijmen Hageman, Moritz Diehl
Robust and accurate pose estimation of moving systems is a challenging task that is often tackled by combining information from different sensor subsystems in a multi-sensor fusion setup. To obtain robust and accurate estimates, it is crucial to respect the exact time of each measurement. Data fusion is additionally challenged when the sensors are running at different rates and the information is subject to processing- and transmission delays. In this paper, we present an optimization-based moving horizon estimator which allows to estimate and compensate for time-varying measurement delays without the need for any synchronization signals between the sensors. By adopting a direct collocation approach, we find a continuous-time solution for the navigation states which allows us to incorporate the discrete-time sensor measurements in an optimal way despite the presence of unknown time delays. The presented sensor fusion algorithm is applied to the problem of pose estimation by fusing data of a high-rate inertial measurement unit and a low-rate centimeter-accurate global navigation satellite system receiver using simulated and real-data experiments. ...
Conference paper (2019) - Fabian Girrbach, Raymond Zandbergen, Manon Kok, Tijmen Hageman, Giovanni Bellusci, Moritz DIehl
Inertial sensors are used in an increasing number of autonomous applications. Integrating such sensors into dynamic systems, the problem of their calibration arises naturally. Existing methods often require the sensor to be accurately placed in certain poses, which can be infeasible in practice. In this paper, we present an optimization-based estimator for in-field identification of inertial biases and scale factors. Instead of predefined poses, we use measurements of an accurate global navigation satellite system receiver in the calibration algorithm. By adopting a moving horizon scheme, the resulting estimator has the potential to run on embedded hardware allowing for online calibration without sacrificing robustness. We also present an approach for the simulation of realistic sensor data. The resulting datasets are used to analyze the performance of the optimization-based estimator. The evaluated statistics clearly show that moving horizon estimation improves the robustness and accuracy of the presented calibration approach in the presence of uncertain initial conditions and outperforms traditional recursive filters. ...
Journal article (2018) - Jesus Lago, Michael Erhard, Moritz Diehl
Fast online generation of feasible and optimal reference trajectories is crucial in tracking model predictive control, especially for stability and optimality in presence of a time varying parameter. In this paper, in order to circumvent the operational efforts of handling a discrete set of precomputed trajectories and switching between them, time warping of a single trajectory is proposed as an alternative concept. In particular, the conceptual ideas of warping theory are presented and illustrated based on the example of a tethered kite system for airborne wind energy. In detail, for warpable systems, feasibility and optimality of trajectories are discussed. Subsequently, the full algorithm of a nonlinear model predictive control implementation based on warping a single precomputed reference is presented. Finally, the warping algorithm is applied to the airborne wind energy system. Simulation results in presence of real world perturbations are evaluated and compared. ...
Journal article (2017) - Jesus Lago Garcia, Michael Erhard, Moritz Diehl
Generation of feasible and optimal reference trajectories is crucial in tracking Nonlinear Model Predictive Control. Especially, for stability and optimality in presence of a time varying parameter, adaptation of the tracking trajectory has to be implemented. General approaches are real-time generation of trajectories or switching between a discrete set of precomputed trajectories. In order to circumvent the operational efforts of these methods for a special type of dynamical systems, we propose time warping as an alternative approach. This algorithm implements online generation of tracking trajectories by warping a single precomputed reference. In detail, warpable systems, feasibility and optimality of trajectories and the controller implementation are discussed. Finally, as an application example, simulation results of a tethered kite system for airborne wind energy generation are presented. ...
Conference paper (2017) - Minh Dang Doan, Moritz Diehl, Tamas Keviczky, Bart De Schutter
In this paper we introduce an iterative distributed Jacobi algorithm for solving convex optimization problems, which is motivated by distributed model predictive control (MPC) for linear time-invariant systems. Starting from a given feasible initial guess, the algorithm iteratively improves the value of the cost function with guaranteed feasible solutions at every iteration step, and is thus suitable for MPC applications in which hard constraints are important. The proposed iterative approach involves solving local optimization problems consisting of only few subsystems, depending on the flexible choice of decomposition and the sparsity structure of the couplings. This makes our approach more applicable to situations where the number of subsystems is large, the coupling is sparse, and local communication is available. We also provide a method for checking a posteriori centralized optimality of the converging solution, using comparison between Lagrange multipliers of the local problems. Furthermore, a theoretical result on convergence to optimality for a particular distributed setting is also provided. ...
Book chapter (2013) - Uwe Ahrens, Moritz Diehl, Roland Schmehl