|
|
|
题名
|
作者
|
年代
|
出处
|
被引量
|
| 1 | Discrete-time dynamic graphical games:model-free reinforcement learning solution显示文摘This paper introduces a model-free reinforcement learning technique that is used to solve a class of dynamic games known as dynamic graphical games. The graphical game results from multi-agent dynamical systems, where pinning control is used to make all the agents synchronize to the state of a command generator or a leader agent. Novel coupled Bellman equations and Hamiltonian functions are developed for the dynamic graphical games. The Hamiltonian mechanics are used to derive the necessary conditions for optimality. The solution for the dynamic graphical game is given in terms of the solution to a set of coupled Hamilton-Jacobi-Bellman equations developed herein. Nash equilibrium solution for the graphical game is given in terms of the solution to the underlying coupled Hamilton-Jacobi-Bellman equations. An online model-free policy iteration algorithm is developed to learn the Nash solution for the dynamic graphical game. This algorithm does not require any knowledge of the agents' dynamics. A proof of convergence for this multi-agent learning algorithm is given under mild assumption about the inter-connectivity properties of the graph. A gradient descent technique with critic network structures is used to implement the policy iteration algorithm to solve the graphical game online in real-time. | Mohammed I.ABOUHEAF Frank L.LEWIS Magdi S.MAHMOUD Dariusz G.MIKULSKI | 2015 | Control Theory and Technology2015,13,1: | 6 |
| 2 | Stochastic DoS Attack Allocation Against Collaborative Estimation in Sensor Networks显示文摘In this paper,denial of service(DoS)attack management for destroying the collaborative estimation in sensor networks and minimizing attack energy from the attacker perspective is studied.In the communication channels between sensors and a remote estimator,the attacker chooses some channels to randomly jam DoS attacks to make their packets randomly dropped.A stochastic power allocation approach composed of three steps is proposed.Firstly,the minimum number of channels and the channel set to be attacked are given.Secondly,a necessary condition and a sufficient condition on the packet loss probabilities of the channels in the attack set are provided for general and special systems,respectively.Finally,by converting the original coupling nonlinear programming problem to a linear programming problem,a method of searching attack probabilities and power to minimize the attack energy is proposed.The effectiveness of the proposed scheme is verified by simulation examples. | Ya Zhang Lishuang Du Frank L.Lewis | 2020 | IEEE/CAA Journal of Automatica Sinica2020,7,5: | 2 |
| 3 | Cost-effective distributed FTFC for uncertain nonholonomic mobile robot fleet with collision avoidance and connectivity preservation显示文摘In this paper,the fault-tolerant formation control(FTFC)problem is investigated for a group of uncertain nonholonomic mobile robots with limited communication ranges and unpredicted actuator faults,where the communication between the robots is in a directed one-to-one way.In order to guarantee the connectivity preservation and collision avoidance among the robots,some properly chosen performance functions are incorporated into the controller to per-assign the asymmetrical bounds for relative distance and bearing angle between each pair of adjacent mobile robots.Particularly,the resultant control scheme remains at a costeffective level because its design does not use any velocity information from neighbors,any prior knowledge of system nonlinearities or any nonlinear approximator to account for them despite the presence of modeling uncertainties,unknown external disturbances,and unexpected actuator faults.Meanwhile,each follower is derived to track the leader with the tracking errors regarding relative distance and bearing angle subject to prescribed transient and steady-state performance guarantees,respectively.Moreover,all the closed-loop signals are ensured to be ultimately uniformly bounded.Finally,a numerical example is simulated to verify the effectiveness of this methodology. | Xiucai Huang Zhengguo Li Frank L.Lewis | 2023 | Journal of Automation and Intelligence2023,2,1: | 0 |
| 4 | Guest Editorial for Special Issue on Extensions of Reinforcement Learning and Adaptive Control显示文摘ADAPTIVE control is a proven method for learning feedback controllers for systems with unknown dynamic models,exogenous disturbances,nonzero setpoints,and unmodeled nonlinearities.Adaptive control has been applied for years in process control,industry,aerospace systems。 | Frank L.Lewis Warren Dixon Zhongsheng Hou Tansel Yucelen | 2014 | IEEE/CAA Journal of Automatica Sinica2014,1,3: | 0 |
| 5 | Heterogeneous multi-player imitation learning显示文摘This paper studies imitation learning in nonlinear multi-player game systems with heterogeneous control input dynamics.We propose a model-free data-driven inverse reinforcement learning(RL)algorithm for a leaner to find the cost functions of a N-player Nash expert system given the expert's states and control inputs.This allows us to address the imitation learning problem without prior knowledge of the expert's system dynamics.To achieve this,we provide a basic model-based algorithm that is built upon RL and inverse optimal control.This serves as the foundation for our final model-free inverse RL algorithm which is implemented via neural network-based value function approximators.Theoretical analysis and simulation examples verify the methods. | Bosen Lian Wenqian Xue Frank L.Lewis | 2023 | Control Theory and Technology2023,21,3: | 0 |
| 6 | Adaptive Uniform Performance Control of Strict-Feedback Nonlinear Systems With Time-Varying Control Gain显示文摘In this paper,we present a novel adaptive performance control approach for strict-feedback nonparametric systems with unknown time-varying control coefficients,which mainly includes the following steps.Firstly,by introducing several key transformation functions and selecting the initial value of the time-varying scaling function,the symmetric prescribed performance with global and semi-global properties can be handled uniformly,without the need for control re-design.Secondly,to handle the problem of unknown time-varying control coefficient with an unknown sign,we propose an enhanced Nussbaum function(ENF)bearing some unique properties and characteristics,with which the complex stability analysis based on specific Nussbaum functions as commonly used is no longer required.Thirdly,by utilizing the core-function information technique,the nonparametric uncertainties in the system are gracefully handled so that no approximator is required.Furthermore,simulation results verify the effectiveness and benefits of the approach. | Kai Zhao Changyun Wen Yongduan Song Frank L.Lewis | 2023 | IEEE/CAA Journal of Automatica Sinica2023,10,2: | 0 |
| 7 | Practical prescribed-time tracking control for uncertain strict-feedback systems with guaranteed performance under unknown control directions显示文摘In this paper,we consider the practical prescribed-time performance guaranteed tracking control problem for a class of uncertain strict-feedback systems subject to unknown control direction.Due to the existence of unknown nonlinearities and uncertainties,it is challenging to design a controller that can ensure the stability of closed-loop system within a predetermined finite time while maintaining the specified transient performance.The underlying problem becomes further complex as the control directions are unknown.To deal with the above problems,a special translation function as well as Nussbaum type function are introduced in the prescribed performance control(PPC)framework.Finally,a PPC as well as preset finite time tracking control scheme is designed,and its effectiveness is confirmed by both theoretical analysis and numerical simulation. | Zhou Yang Yujuan Wang Frank L.Lewis | 2023 | Journal of Automation and Intelligence2023,2,2: | 0 |
| 8 | Recent Progress in Reinforcement Learning and Adaptive Dynamic Programming for Advanced Control Applications显示文摘Reinforcement learning(RL) has roots in dynamic programming and it is called adaptive/approximate dynamic programming(ADP) within the control community. This paper reviews recent developments in ADP along with RL and its applications to various advanced control fields. First, the background of the development of ADP is described, emphasizing the significance of regulation and tracking control problems. Some effective offline and online algorithms for ADP/adaptive critic control are displayed, where the main results towards discrete-time systems and continuous-time systems are surveyed, respectively.Then, the research progress on adaptive critic control based on the event-triggered framework and under uncertain environment is discussed, respectively, where event-based design, robust stabilization, and game design are reviewed. Moreover, the extensions of ADP for addressing control problems under complex environment attract enormous attention. The ADP architecture is revisited under the perspective of data-driven and RL frameworks,showing how they promote ADP formulation significantly.Finally, several typical control applications with respect to RL and ADP are summarized, particularly in the fields of wastewater treatment processes and power systems, followed by some general prospects for future research. Overall, the comprehensive survey on ADP and RL for advanced control applications has d emonstrated its remarkable potential within the artificial intelligence era. In addition, it also plays a vital role in promoting environmental protection and industrial intelligence. | Ding Wang Ning Gao Derong Liu Jinna Li Frank L.Lewis | 2024 | IEEE/CAA Journal of Automatica Sinica2024,11,1: | 0 |
| 9 | Neural network solution for finite-horizon H-infinity constrained optimal control of nonlinear systems显示文摘In this paper,neural networks are used to approximately solve the finite-horizon constrained input H-infinity state feedback control problem.The method is based on solving a related Hamilton-Jacobi-Isaacs equation of the corresponding finite-horizon zero-sum game.The game value function is approximated by a neural network with time-varying weights.It is shown that the neural network approximation converges uniformly to the game-value function and the resulting almost optimal constrained feedback controller provides closed-loop stability and bounded L2 gain.The result is an almost optimal H-infinity feedback controller with time-varying coefficients that is solved a priori off-line.The effectiveness of the method is shown on the Rotational/Translational Actuator benchmark nonlinear control problem. | Frank L.LEWIS | 2007 | 控制理论与应用(英文版)2007,5,1: | 0 |