详细信息
Enhancing robot reinforcement learning with experience-acquirer-aided training method ( EI收录)
文献类型:期刊文献
英文题名:Enhancing robot reinforcement learning with experience-acquirer-aided training method
作者:Ji, Xiancheng[1]; Yi, Jianjun[1]; Su, Lin[1]
机构:[1] School of Mechanical and Power Engineering, East China University of Science and Technology, Shanghai, China
年份:2025
起止页码:1
外文期刊名:Robotic Intelligence and Automation
收录:EI(收录号:20260319912321)
语种:英文
外文关键词:Complex networks - Computer aided instruction - Deep learning - Deep reinforcement learning - Learning algorithms - Machine design - Parallel architectures - Robot learning - Robotic assembly - Transfer learning
摘要:Purpose – The purpose of this paper is to study a reinforcement learning training method with action decoupling and experience transfer capabilities, aiming to overcome the difficulty of robots learning composite actions caused by sparse rewards. When learning complex, compound actions under sparse reward environments, robots require an enormous number of training steps to trigger an effective reward by chance, which can indefinitely prolong the training cycle or even lead to learning failure. Therefore, this paper proposed a method to alleviate the sparsity issue through action decomposition and experience aid. Design/methodology/approach – A reinforcement learning method, termed experience-acquirer-aided training (EAAT), is proposed to enhance the reinforcement learning capability of robots when learning complex actions under sparse rewards. EAAT integrates two identical actor-critic reinforcement learning algorithms, which focus on local and global compound actions respectively with differential observations. Through mutual fusion via time-variant Q-functions and their respective critic networks – where local critics act as experience acquirers and continuous reward generators – this method enables efficient learning in sparse reward scenarios. Moreover, a fuzzy reward function is designed based on Dempster–Shafer fusion theory, which reduces the complexity of designing traditional continuous reward functions and provides further efficiency support for EAAT. Findings – Core experimental validation on a robotic autonomous precision assembly platform demonstrates that EAAT achieves high-quality convergence with 50% fewer required time steps. In addition, the approach effectively generalizes to inverse kinematics learning, thereby broadening its application scope. Originality/value – This work introduces a novel hybrid training method EAAT, a non-hierarchical dual-agent parallel architecture that can automate reward generation via a pretrained critic network and integrate fuzzy evaluation to simplify reward design. In addition, for specific peg-hole assembly tasks, an innovative hole-searching strategy and improved insertion control methods are added. ? 2025 Emerald Publishing Limited
参考文献:
正在载入数据...
