详细信息

基于KL散度的策略优化    

KL-divergence-based Policy Optimization

文献类型:期刊文献

中文题名:基于KL散度的策略优化

英文题名:KL-divergence-based Policy Optimization

作者:李建国[1];赵海涛[1];孙韶媛[2]

机构:[1]华东理工大学信息科学与工程学院,上海200237;[2]东华大学信息科学与技术学院,上海201620

年份:2019

卷号:46

期号:6

起止页码:212

中文期刊名:计算机科学

外文期刊名:Computer Science

收录:CSTPCD;;北大核心:【北大核心2017】;CSCD:【CSCD_E2019_2020】;

基金:国家自然科学基金(61375007);上海市科委基础研究项目(15JC1400600)资助

语种:中文

中文关键词:强化学习;KL散度;策略优化;连续动作空间

外文关键词:Reinforcement learning;KL-divergence;Policy optimization;Continuous action space

摘要:强化学习(Reinforcement Learning,RL)在复杂的优化和控制问题中具有广泛的应用前景。针对传统的策略梯度方法在处理高维的连续动作空间环境时无法有效学习复杂策略,导致收敛速度慢甚至无法收敛的问题,提出了一种在线学习的基于KL散度的策略优化算法(KL-divergence-based Policy Optimization,KLPO)。在Actor-Critic方法的基础上,通过引入KL散度构造惩罚项,将“新”“旧”策略间的散度结合到损失函数中,以对Actor部分的策略更新进行优化;并进一步利用KL散度控制算法更新学习步长,以确保策略每次在由KL散度定义的合理范围内以最大学习步长进行更新。分别在经典的倒立摆仿真环境和公开的连续动作空间的机器人运动环境中对所提算法进行了测试。实验结果表明,KLPO算法能够更好地学习复杂的策略,收敛速度快,并且可获取更高的回报。
Reinforcement learning has wide application prospects in dealing with the problem of complex optimization and control.Since traditional policy gradient method cannot learn the complex policy effectively in addressing with the environment with high-dimensional and continuous action space,that causes slow convergence rate or even non-convergence,this paper proposed an online KL-divergence-based policy optimization algorithm to solve this issue.Based on Actor-Critic algorithm,the KL-divergence is introduced to construct a penalty which adds the distance between“new”and“old”into policy loss function to optimization the policy update of Actor.Furthermore,the learning step is controlled by KL-divergence to ensure the policy update with maximum learning step within security region.On the experiment of Pendulum and Humanoid,simulation results show that KLPO algorithm can learn complex strategies better,converge faster and get higher returns.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心