详细信息

Decentralized Natural Policy Gradient with Variance Reduction for Collaborative Multi-Agent Reinforcement Learning  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Decentralized Natural Policy Gradient with Variance Reduction for Collaborative Multi-Agent Reinforcement Learning

作者:Chen, Jinchi[1,2];Feng, Jie[1];Gao, Weiguo[1,3];Wei, Ke[1]

机构:[1]Fudan Univ, Sch Data Sci, Shanghai, Peoples R China;[2]East China Univ Sci & Technol, Sch Math, Shanghai, Peoples R China;[3]Fudan Univ, Sch Math Sci, Shanghai, Peoples R China

年份:2024

卷号:25

外文期刊名:JOURNAL OF MACHINE LEARNING RESEARCH

收录:;EI(收录号:20254219343859);WOS:【SCI-EXPANDED(收录号:WOS:001263168800001)】;

基金:Acknowledgments This work was partially supported by the National Key R&D Program of China (Grant No. 2021YFA1003300) , Natural Science Foundation of Shanghai (Grant No. 23ZR1406400) , and Science and Technology Commission of Shanghai Municipality (No. 23JC1401000) .

语种:英文

外文关键词:multi-agent reinforcement learning; natural policy gradient; decentralized optimization; variance reduction

摘要:This paper studies a policy optimization problem arising from collaborative multi-agent reinforcement learning in a decentralized setting where agents communicate with their neighbors over an undirected graph to maximize the sum of their cumulative rewards. A novel decentralized natural policy gradient method, dubbed Momentum-based Decentralized Natural Policy Gradient (MDNPG), is proposed, which incorporates natural gradient, momentum-based variance reduction, and gradient tracking into the decentralized stochastic gradient ascent framework. The O( n - 1 f - 3 ) sample complexity for MDNPG to converge to an epsilon-stationary point has been established under standard assumptions, where n is the number of agents. It indicates that MDNPG can achieve the optimal convergence rate for decentralized policy gradient methods and possesses a linear speedup in contrast to centralized optimization methods. Moreover, superior empirical performance of MDNPG over other state -of -the -art algorithms has been demonstrated by extensive numerical experiments.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心