详细信息
A Cooperative Multi-Agent Reinforcement Learning Method Based On Reward Decomposition In Increasing Number Of Agents Environment ( EI收录)
文献类型:期刊文献
英文题名:A Cooperative Multi-Agent Reinforcement Learning Method Based On Reward Decomposition In Increasing Number Of Agents Environment
作者:Wei, Suhang[1]; Zhao, Zonghang[1]; Feng, Xiang[1]; Yu, Huiqun[1]
机构:[1] East China University of Science and Technology, Department of Computer Science and Engineering, Shanghai, 200237, China
年份:2025
外文期刊名:IEEE Transactions on Cognitive and Developmental Systems
收录:EI(收录号:20254319391815)
语种:英文
外文关键词:Decomposition - Intelligent agents - Intelligent systems - Knowledge management - Knowledge transfer - Learning systems - Scalability - Starting - Students - Teaching - Transfer learning
摘要:Cooperative multi-agent reinforcement learning (MARL) is a promising approach for complex collaborative tasks. However, practical deployment remains challenging due to ambiguous credit assignment, inefficient exploration, and the cold-start problem, particularly in systems where the number of agents grows dynamically. Inspired by human cognitive mechanisms for task decomposition and experience-based knowledge transfer, we propose Multi-Agent Reward Decomposition and Knowledge Transfer (MARDKT), a unified method that jointly addresses the challenges of credit assignment, exploration inefficiency, and cold-start in dynamic multi-agent settings. We introduce a four-channel reward decomposition mechanism that separates reward signals along local/global and extrinsic/intrinsic dimensions: local rewards drive individual exploration, global rewards foster cooperation, and intrinsic curiosity at both individual and team levels promotes discovery of novel states. To enable rapid integration of new agents, we further design a teacher-student framework where students inherit knowledge from trained teachers via policy imitation and value function distillation. We prove that MARDKT ensures monotonic policy improvement from a local perspective and demonstrate its effectiveness in multi-vehicle on-ramp merging and cooperative boxpushing tasks. Furthermore, MARDKT achieves rapid convergence as the number of agents increases dynamically, showcasing strong scalability in dynamic environments. ? 2016 IEEE.
参考文献:
正在载入数据...
