详细信息
A Brain-Inspired Incremental Multitask Reinforcement Learning Approach ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:A Brain-Inspired Incremental Multitask Reinforcement Learning Approach
作者:Jin, Chenying[1];Feng, Xiang[1];Yu, Huiqun[1]
机构:[1]East China Univ Sci & Technol, Shanghai 200237, Peoples R China
年份:2024
卷号:16
期号:3
起止页码:1147
外文期刊名:IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS
收录:;EI(收录号:20235115237196);WOS:【SCI-EXPANDED(收录号:WOS:001247154200005)】;
基金:No Statement Available
语种:英文
外文关键词:Task analysis; Training; Reinforcement learning; Convergence; Scalability; Multitasking; Heuristic algorithms; Brain-inspired reinforcement learning; conscious and subconscious mode; distributed and incremental architecture; multitask reinforcement learning; off-policy correction
摘要:Recently, there have been growing interests in multitask reinforcement learning (MTRL), which is viewed as a promising framework for training agents to execute multiple tasks simultaneously. However, limitations in scalability and convergence remain key obstacles for scaling these MTRL algorithms to dynamic and complex tasks. To address these, we propose a method called brain-inspired incremental multitask reinforcement learning (BIMTRL) that aims to improve parallelism and scalability of multiple tasks. Inspired by learning processes in human brain, we integrate conscious and subconscious modes into the agents' exploration of environments. Our two-step strategy of policy loosening and importance tradeoff enables an effective switch between these modes. Additionally, in order to overcome the convergence dilemma, we adopt the V-trace method as a stable and robust off-policy correction technique for our actor-critic agents. Experimental evaluations on various tasks in OpenAI Gym, Atari, and PyBullet have demonstrated that BIMTRL achieves a 61.4% greater average return and 57.6% higher speed than specific multitask baselines. Furthermore, its distributed and incremental architecture endows the BIMTRL approach with a desired scalability in both discrete and continuous environments, ultimately leading to larger rewards, higher speed, and better convergence.
参考文献:
正在载入数据...
