详细信息

Data-Driven Optimal Bipartite Consensus Control for Second-Order Multiagent Systems via Policy Gradient Reinforcement Learning  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Data-Driven Optimal Bipartite Consensus Control for Second-Order Multiagent Systems via Policy Gradient Reinforcement Learning

作者:Liu, Qiwei[1];Yan, Huaicheng[1,2];Wang, Meng[1];Li, Zhichen[1];Liu, Shuai[3]

机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China;[2]Chengdu Univ, Sch Informat Sci & Engn, Chengdu 610106, Peoples R China;[3]Shandong Univ, Sch Control Sci & Engn, Jinan 250061, Peoples R China

年份:2024

卷号:54

期号:6

起止页码:3468

外文期刊名:IEEE TRANSACTIONS ON CYBERNETICS

收录:;EI(收录号:20232614289521);WOS:【SCI-EXPANDED(收录号:WOS:001012447000001)】;

基金:This work was supported in part by the National Natural Science Foundation of China under Grant 62073143, Grant 62273255,and Grant 62003139; in part by the Shanghai International Science and Technology Cooperation Project under Grant 18510711100; and in part bythe Innovation Program of Shanghai Municipal Education Commission underGrant 2021-01-07-00-02-E00107.

语种:英文

外文关键词:Data-driven; optimal bipartite consensus control (OBCC); policy gradient; reinforcement learning (RL); second-order multiagent systems (MASs)

摘要:This article investigates the optimal bipartite consensus control (OBCC) problem for unknown second-order discrete-time multiagent systems (MASs). First, the coopetition network is constructed to describe the cooperative and competitive relationships between agents, and the OBCC problem is proposed by the tracking error and related performance index function. Based on the distributed policy gradient reinforcement learning (RL) theory, a data-driven distributed optimal control strategy is obtained to guarantee the bipartite consensus of all agents' position and velocity states. In addition, the offline data sets ensure the learning efficiency of the system. These data sets are generated by running the system in real time. Besides, the designed algorithm is an asynchronous version, which is essential to solve the challenge caused by the computational ability difference between nodes in MASs. Then, by means of the functional analysis and Lyapunov theory, the stability of the proposed MASs and the convergence of the learning process are analyzed. Furthermore, an actor-critic structure containing two neural networks is used to implement the proposed methods. Finally, a numerical simulation shows the effectiveness and validity of the results.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心