详细信息
文献类型:期刊文献
中文题名:平均报酬模型强化学习理论、算法及应用
英文题名:Average Reward Reinforcement Learning Theory Algorithms and Its Application
作者:黄炳强[1];曹广益[1];李建华[2]
机构:[1]上海交通大学自动化系,上海200030;[2]华东理工大学计算机系,上海200237
年份:2007
卷号:33
期号:18
起止页码:18
中文期刊名:计算机工程
外文期刊名:Computer Engineering
收录:CSTPCD;;Scopus;北大核心:【北大核心2004】;CSCD:【CSCD2011_2012】;
语种:中文
中文关键词:平均报酬强化学习;R学习;H学习
外文关键词:average reward reinforcement learning;R-learning;H-learning
摘要:折扣报酬模型强化学习是目前强化学习研究的主流,但折扣因子的选取使得近期期望报酬的影响大于远期期望报酬的影响,而有时候较大远期期望报酬的策略有可能是最优的,因此比较合理的方法是采用平均报酬模型强化学习。该文介绍了平均报酬模型强化学习的两个主要算法以及主要应用。
Discounted reward reinforcement learning is the mainstream of reinforcement learning research and its short-term reward is more important than a long-term reward owing to the discount factor.However,sometimes the long-term reward is optimal and it is reasonable to use the average reward reinforcement learning method.This paper presents average reward reinforcement learning including R-learning and H-learning.The application is proposed.
参考文献:
正在载入数据...
