详细信息

Monotonic Quantile Network for Worst-Case Offline Reinforcement Learning  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Monotonic Quantile Network for Worst-Case Offline Reinforcement Learning

作者:Bai, Chenjia[1];Xiao, Ting[2];Zhu, Zhoufan[3];Wang, Lingxiao[4,5];Zhou, Fan[3];Garg, Animesh[6];He, Bin[7];Liu, Peng[8];Wang, Zhaoran[4,5]

机构:[1]Shanghai AI Lab, Shanghai 200232, Peoples R China;[2]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[3]Shanghai Univ Finance & Econ, Sch Stat & Management, Shanghai 200433, Peoples R China;[4]Northwestern Univ, Dept Ind Engn, Evanston, IL 60208 USA;[5]Northwestern Univ, Dept Management Sci, Evanston, IL 60208 USA;[6]Univ Toronto, Vector Inst, Toronto, ON M5S 0A5, Canada;[7]Tongji Univ, Sch Elect & Informat Engn, Shanghai Res Inst Intelligent Autonomous Syst, Shanghai 201210, Peoples R China;[8]Harbin Inst Technol, Fac Comp, Harbin 150001, Peoples R China

年份:2024

卷号:35

期号:7

起止页码:8954

外文期刊名:IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS

收录:;EI(收录号:20224613111890);WOS:【SCI-EXPANDED(收录号:WOS:000881967100001)】;

基金:This work was supported by the Shanghai AI Laboratory, Shanghai, China.

语种:英文

外文关键词:Behavioral sciences; Task analysis; Standards; Estimation; Training; Service robots; Learning systems; Monotonic quantile network (MQN); offline reinforcement learning (RL); quantile regression; risk-sensitive learning

摘要:A key challenge in offline reinforcement learning (RL) is how to ensure the learned offline policy is safe, especially in safety-critical domains. In this article, we focus on learning a distributional value function in offline RL and optimizing a worst-case criterion of returns. However, optimizing a distributional value function in offline RL can be hard, since the crossing quantile issue is serious, and the distribution shift problem needs to be addressed. To this end, we propose monotonic quantile network (MQN) with conservative quantile regression (CQR) for risk-averse policy learning. First, we propose an MQN to learn the distribution over returns with non-crossing guarantees of the quantiles. Then, we perform CQR by penalizing the quantile estimation for out-of-distribution (OOD) actions to address the distribution shift in offline RL. Finally, we learn a worst-case policy by optimizing the conditional value-at-risk (CVaR) of the distributional value function. Furthermore, we provide theoretical analysis of the fixed-point convergence in our method. We conduct experiments in both risk-neutral and risk-sensitive offline settings, and the results show that our method obtains safe and conservative behaviors in robotic locomotion tasks.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心