详细信息

Global Convergence of Natural Policy Gradient with Hessian-Aided Momentum Variance Reduction  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Global Convergence of Natural Policy Gradient with Hessian-Aided Momentum Variance Reduction

作者:Feng, Jie[1];Wei, Ke[1];Chen, Jinchi[2]

机构:[1]Fudan Univ, Sch Data Sci, Shanghai, Peoples R China;[2]East China Univ Sci & Technol, Sch Math, Shanghai, Peoples R China

年份:2024

卷号:101

期号:2

外文期刊名:JOURNAL OF SCIENTIFIC COMPUTING

收录:;EI(收录号:20240023090);WOS:【SCI-EXPANDED(收录号:WOS:001325828900001)】;

基金:Ke Wei was partially supported by National Natural Science Foundation of China (Grant No. 92370105).

语种:英文

外文关键词:Natural policy gradient; Reinforcement learning; Sample complexity; Variance reduction

摘要:Natural policy gradient (NPG) and its variants are widely-used policy search methods in reinforcement learning. Inspired by prior work, a new NPG variant coined NPG-HM is developed in this paper, which utilizes the Hessian-aided momentum technique for variance reduction, while the sub-problem is solved via the stochastic gradient descent method. It is shown that NPG-HM can achieve the global last iterate epsilon\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varepsilon $$\end{document}-optimality with a sample complexity of O(epsilon-2)\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal{O}(\varepsilon <^>{-2})$$\end{document}, which is the best known result for natural policy gradient type methods under the generic Fisher non-degenerate policy parameterizations. The convergence analysis is built upon a relaxed weak gradient dominance property tailored for NPG under the compatible function approximation framework, as well as a neat way to decompose the error when handling the sub-problem. Moreover, numerical experiments on Mujoco-based environments demonstrate the superior performance of NPG-HM over other state-of-the-art policy gradient methods.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心