详细信息

Two-Time Scale Tracking Control of Flexible Robots With Primal-Dual Inverse Reinforcement Learning  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Two-Time Scale Tracking Control of Flexible Robots With Primal-Dual Inverse Reinforcement Learning

作者:Que, Xuejie[1,2];Wang, Zhenlei[1,2];Zhang, Yanqi[1,2];Su, Guanghao[1,2]

机构:[1]East China Univ Sci & Technol, State Key Lab Ind Control Technol, Minist Educ, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China

年份:2025

卷号:36

期号:4

起止页码:6383

外文期刊名:IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS

收录:;EI(收录号:20242216162162);WOS:【SCI-EXPANDED(收录号:WOS:001230775700001)】;

基金:No Statement Available

语种:英文

外文关键词:Vibrations; Tracking; Task analysis; Trajectory; Cost function; Space exploration; Convergence; Flexible robots (FRs); incomplete reference signals; inverse reinforcement learning (IRL); primal-dual (PD); two-time scale tracking control

摘要:Flexible robots (FRs) are generally designed to be lightweight to achieve rapid motion. However, accompanying vibrations and modeling errors influence tracking control, especially in situations involving reference signal loss. This article develops a two-time scale primal-dual inverse reinforcement learning (PD-IRL) framework for FRs to perform tracking tasks with incomplete reference signals. First, consider the admissible policy as a nonconvex input constraint to guarantee the stable operation of the equipment. Then, FRs imitate the demonstration behaviors of an expert, including both rigid and flexible motions, to achieve a balance in tracking speed and vibration suppression. During the imitation process, nonconvex optimization problems of FRs are transformed into corresponding dual problems to obtain the global optimal policy. Moreover, employing multiple linearly independent paths to explore the state space simultaneously can improve convergence speed. Convergence and stability are studied rigorously. Finally, simulations and comparisons show the effectiveness and superiority of the proposed method.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心