K. He
Please Note
7 records found
1
State-Action Control Barrier Functions
Imposing Safety on Learning-Based Control With Low Online Computational Costs
Learning-based control with safety guarantees usually requires real-time safety certification and modifications of possibly unsafe learning-based policies. The control barrier function (CBF) method uses a safety filter (SF) containing a constrained optimization problem to produce safe policies. However, finding a valid CBF for a general nonlinear system requires a complex function parameterization, which in general makes the policy optimization problem difficult to solve in real time. For nonlinear systems with nonlinear state constraints, this paper proposes the novel concept of state-action CBFs (SACBFs), which do not only characterize the safety at each state but also evaluate the control inputs taken at each state. SACBFs, in contrast to CBFs, enable a flexible parameterization, resulting in a SF that involves a convex quadratic optimization problem, which significantly alleviates the online computational burden. We propose a learning-based approach to synthesize SACBFs. The effect of learning errors on the effectiveness of SACBFs is addressed by constraint tightening and introducing a new concept called contractive-set CBFs. This ensures formal safety guarantees for the learned CBFs and control policies. Simulation results on an inverted pendulum with elastic walls validate the proposed CBFs in terms of constraint satisfaction and CPU time.
Infinite-horizon optimal control of constrained piecewise affine (PWA) systems has been approximately addressed by hybrid model predictive control (MPC), which, however, has computational limitations, both in offline design and online implementation. In this article, we consider an alternative approach based on approximate dynamic programming (ADP), an important class of methods in reinforcement learning. We accommodate nonconvex union-of-polyhedra state constraints and linear input constraints into ADP by designing PWA penalty functions. PWA function approximation is used, which allows for a mixed-integer encoding to implement ADP. The main advantage of the proposed ADP method is its online computational efficiency. Particularly, we propose two control policies, which lead to solving a smaller-scale mixed-integer linear program than conventional hybrid MPC, or a single convex quadratic program, depending on whether the policy is implicitly determined online or explicitly computed offline. We characterize the stability and safety properties of the closed-loop systems, as well as the suboptimality of the proposed policies, by quantifying the approximation errors of value functions and policies. We also develop an offline mixed-integer-linear-programming-based method to certify the reliability of the proposed method. Simulation results on an inverted pendulum with elastic walls and on an adaptive cruise control problem validate the control performance in terms of constraint satisfaction and CPU time.
Integrating Learning-Based and MPC-Based Control for PWA Systems
Challenges and Opportunities
Learning-based control, in particularReinforcement Learning (RL) reinforcementReinforcement learning, and optimization-based control, in particular model predictive control, each have their advantages and disadvantages for online, real-timeOptimal control optimal controlOptimal control of systems with complex dynamicsDynamic. However, both approaches are highly complementary and therefore there is an increased interest in combining their advantages in an integrated approach. In this chapter, we provide an overview of recent results, challenges, and opportunities on an integrated learning-based and optimization-based control approach. We focus in particular on piecewise affine systems as they are an extension of linear systemsLinear systems that can model or approximate hybridHybrid or nonlinearNonlinearbehaviorBehavior and as they still allow for effective numerical solutionSolution approaches.
Approximate dynamic programming for constrained linear systems
A piecewise quadratic approximation approach
Approximate dynamic programming (ADP) faces challenges in dealing with constraints in control problems. Model predictive control (MPC) is, in comparison, well-known for its accommodation of constraints and stability guarantees, although its computation is sometimes prohibitive. This paper introduces an approach combining the two methodologies to overcome their individual limitations. The predictive control law for constrained linear quadratic regulation (CLQR) problems has been proven to be piecewise affine (PWA) while the value function is piecewise quadratic. We exploit these formal results from MPC to design an ADP method for CLQR problems with a known model. A novel convex and piecewise quadratic neural network with a local–global architecture is proposed to provide an accurate approximation of the value function, which is used as the cost-to-go function in the online dynamic programming problem. An efficient decomposition algorithm is developed to generate the control policy and speed up the online computation. Rigorous stability analysis of the closed-loop system is conducted for the proposed control scheme under the condition that a good approximation of the value function is achieved. Comparative simulations are carried out to demonstrate the potential of the proposed method in terms of online computation and optimality.
This paper deals with the problem of active disturbance rejection control (ADRC) design for a class of uncertain nonlinear systems with sporadic measurements. A novel extended state observer (ESO) is designed in a cascade form consisting of a continuous time estimator, a continuous observation error predictor, and a reset compensator. The proposed ESO estimates not only the system state but also the total uncertainty, which may include the effects of the external perturbation, the parametric uncertainty, and the unknown nonlinear dynamics. Such a reset compensator, whose state is reset to zero whenever a new measurement arrives, is used to calibrate the predictor. Due to the cascade structure, the resulting error dynamics system is presented in a non-hybrid form, and accordingly, analyzed in a general sampled-data system framework. Based on the output of the ESO, a continuous ADRC law is then developed. The convergence of the resulting closed-loop system is proved under given conditions. Two numerical simulations demonstrate the effectiveness of the proposed control method.