详细信息
Enhancing robot reinforcement learning with experience-acquirer-aided training method ( SCI-EXPANDED收录)
文献类型:期刊文献
英文题名:Enhancing robot reinforcement learning with experience-acquirer-aided training method
作者:Ji, Xiancheng[1];Yi, Jianjun[1];Su, Lin[1]
机构:[1]East China Univ Sci & Technol, Sch Mech & Power Engn, Shanghai, Peoples R China
年份:2026
卷号:46
期号:1
起止页码:49
外文期刊名:ROBOTIC INTELLIGENCE AND AUTOMATION
收录:;WOS:【SCI-EXPANDED(收录号:WOS:001652216100001)】;
基金:This paper was supported by Shanghai Science and Technology Action Plan under Grant no. 21JM0010300, Shanghai Aerospace Science and Technology Innovation Foundation (SAST) under Grant no. 2021-037 and Special Fund Technology Innovation Support Project of Shanghai under Grant no. HCXBCY-2023-046, National Defense Basic Scientific Research Program of China (Grant no. JCKY2021606B002). During the preparation of this work, the author used AI-powered language polishing services provided by ChatGPT and Claude to enhance the language and readability of the manuscript. After using this tool, the authors thoroughly reviewed and edited the content as necessary and take full responsibility for the content of the publication.
语种:英文
外文关键词:Deep reinforcement learning; Sparse reward; Compound actions; Fuzzy reward; DS theory; Peg-hole assembly
摘要:PurposeThe purpose of this paper is to study a reinforcement learning training method with action decoupling and experience transfer capabilities, aiming to overcome the difficulty of robots learning composite actions caused by sparse rewards. When learning complex, compound actions under sparse reward environments, robots require an enormous number of training steps to trigger an effective reward by chance, which can indefinitely prolong the training cycle or even lead to learning failure. Therefore, this paper proposed a method to alleviate the sparsity issue through action decomposition and experience aid.Design/methodology/approachA reinforcement learning method, termed experience-acquirer-aided training (EAAT), is proposed to enhance the reinforcement learning capability of robots when learning complex actions under sparse rewards. EAAT integrates two identical actor-critic reinforcement learning algorithms, which focus on local and global compound actions respectively with differential observations. Through mutual fusion via time-variant Q-functions and their respective critic networks - where local critics act as experience acquirers and continuous reward generators - this method enables efficient learning in sparse reward scenarios. Moreover, a fuzzy reward function is designed based on Dempster-Shafer fusion theory, which reduces the complexity of designing traditional continuous reward functions and provides further efficiency support for EAAT.FindingsCore experimental validation on a robotic autonomous precision assembly platform demonstrates that EAAT achieves high-quality convergence with 50% fewer required time steps. In addition, the approach effectively generalizes to inverse kinematics learning, thereby broadening its application scope.Originality/valueThis work introduces a novel hybrid training method EAAT, a non-hierarchical dual-agent parallel architecture that can automate reward generation via a pretrained critic network and integrate fuzzy evaluation to simplify reward design. In addition, for specific peg-hole assembly tasks, an innovative hole-searching strategy and improved insertion control methods are added.
参考文献:
正在载入数据...
