详细信息
Data-Driven Optimal Bipartite Consensus Control for Second-Order Multiagent Systems via Policy Gradient Reinforcement Learning ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Data-Driven Optimal Bipartite Consensus Control for Second-Order Multiagent Systems via Policy Gradient Reinforcement Learning
作者:Liu, Qiwei[1];Yan, Huaicheng[1,2];Wang, Meng[1];Li, Zhichen[1];Liu, Shuai[3]
机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China;[2]Chengdu Univ, Sch Informat Sci & Engn, Chengdu 610106, Peoples R China;[3]Shandong Univ, Sch Control Sci & Engn, Jinan 250061, Peoples R China
年份:2024
卷号:54
期号:6
起止页码:3468
外文期刊名:IEEE TRANSACTIONS ON CYBERNETICS
收录:;EI(收录号:20232614289521);WOS:【SCI-EXPANDED(收录号:WOS:001012447000001)】;
基金:This work was supported in part by the National Natural Science Foundation of China under Grant 62073143, Grant 62273255,and Grant 62003139; in part by the Shanghai International Science and Technology Cooperation Project under Grant 18510711100; and in part bythe Innovation Program of Shanghai Municipal Education Commission underGrant 2021-01-07-00-02-E00107.
语种:英文
外文关键词:Data-driven; optimal bipartite consensus control (OBCC); policy gradient; reinforcement learning (RL); second-order multiagent systems (MASs)
摘要:This article investigates the optimal bipartite consensus control (OBCC) problem for unknown second-order discrete-time multiagent systems (MASs). First, the coopetition network is constructed to describe the cooperative and competitive relationships between agents, and the OBCC problem is proposed by the tracking error and related performance index function. Based on the distributed policy gradient reinforcement learning (RL) theory, a data-driven distributed optimal control strategy is obtained to guarantee the bipartite consensus of all agents' position and velocity states. In addition, the offline data sets ensure the learning efficiency of the system. These data sets are generated by running the system in real time. Besides, the designed algorithm is an asynchronous version, which is essential to solve the challenge caused by the computational ability difference between nodes in MASs. Then, by means of the functional analysis and Lyapunov theory, the stability of the proposed MASs and the convergence of the learning process are analyzed. Furthermore, an actor-critic structure containing two neural networks is used to implement the proposed methods. Finally, a numerical simulation shows the effectiveness and validity of the results.
参考文献:
正在载入数据...
