详细信息
Data-Driven H∞ Output Consensus for Heterogeneous Multiagent Systems Under Switching Topology via Reinforcement Learning ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Data-Driven H∞ Output Consensus for Heterogeneous Multiagent Systems Under Switching Topology via Reinforcement Learning
作者:Liu, Qiwei[1];Yan, Huaicheng[1,2];Zhang, Hao[3];Wang, Meng[1];Tian, Yongxiao[4]
机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China;[2]Hubei Univ Technol, Sch Elect & Elect Engn, Minist Educ, Wuhan 430068, Peoples R China;[3]Tongji Univ, Dept Control Sci & Engn, Shanghai 200092, Peoples R China;[4]Shanghai Univ, Sch Future Technol, Shanghai 200444, Peoples R China
年份:2024
卷号:54
期号:12
起止页码:7865
外文期刊名:IEEE TRANSACTIONS ON CYBERNETICS
收录:;EI(收录号:20244917469336);WOS:【SCI-EXPANDED(收录号:WOS:001288382200001)】;
基金:This work was supported in part by the National Natural Science Foundation of China under Grant 62333005, Grant 62073143, and Grant 62373152; and in part by the Innovation Program of Shanghai Municipal Education Commission under Grant 2021-01-07-00-02-E00105.
语种:英文
外文关键词:Data-driven control; multiagent systems (MASs); multiagent systems (MASs); policy gradient; H-infinity control; reinforcement learning (RL); reinforcement learning (RL); zero-sum dynamic game; zero-sum dynamic game; reinforcement learning (RL); zero-sum dynamic game
摘要:In this article, a novel model-free policy gradient reinforcement learning algorithm is proposed to solve the H-infinity tracking problem for discrete-time heterogeneous multiagent systems with external disturbances over switching topology. The dynamics of the followers and the leader are unknown, and the leader's information is missing for each agent due to the switching topology. Therefore, a distributed adaptive observer is introduced to learn the leader's dynamic model and estimate its state for each agent. For the H-infinity tracking problem, an exponential discount value function is established and the related discrete-time game algebraic Riccati equation (DTGARE) is derived, which is the key to obtaining the control strategy. Furthermore, a data-based policy gradient algorithm is proposed to approximate the solution of the GAREs online and the utilization of agents' accurate knowledge is avoided. To improve the efficiency of data utilization, an offline dataset and the experience replay scheme are used. In addition, the lower bound of the exponential discount value is explored to ensure the stability of the systems. In the end, a simulation is provided to show the validity of the proposed method.
参考文献:
正在载入数据...
