详细信息
Global Convergence of Natural Policy Gradient with Hessian-Aided Momentum Variance Reduction ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Global Convergence of Natural Policy Gradient with Hessian-Aided Momentum Variance Reduction
作者:Feng, Jie[1];Wei, Ke[1];Chen, Jinchi[2]
机构:[1]Fudan Univ, Sch Data Sci, Shanghai, Peoples R China;[2]East China Univ Sci & Technol, Sch Math, Shanghai, Peoples R China
年份:2024
卷号:101
期号:2
外文期刊名:JOURNAL OF SCIENTIFIC COMPUTING
收录:;EI(收录号:20240023090);WOS:【SCI-EXPANDED(收录号:WOS:001325828900001)】;
基金:Ke Wei was partially supported by National Natural Science Foundation of China (Grant No. 92370105).
语种:英文
外文关键词:Natural policy gradient; Reinforcement learning; Sample complexity; Variance reduction
摘要:Natural policy gradient (NPG) and its variants are widely-used policy search methods in reinforcement learning. Inspired by prior work, a new NPG variant coined NPG-HM is developed in this paper, which utilizes the Hessian-aided momentum technique for variance reduction, while the sub-problem is solved via the stochastic gradient descent method. It is shown that NPG-HM can achieve the global last iterate epsilon\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varepsilon $$\end{document}-optimality with a sample complexity of O(epsilon-2)\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathcal{O}(\varepsilon <^>{-2})$$\end{document}, which is the best known result for natural policy gradient type methods under the generic Fisher non-degenerate policy parameterizations. The convergence analysis is built upon a relaxed weak gradient dominance property tailored for NPG under the compatible function approximation framework, as well as a neat way to decompose the error when handling the sub-problem. Moreover, numerical experiments on Mujoco-based environments demonstrate the superior performance of NPG-HM over other state-of-the-art policy gradient methods.
参考文献:
正在载入数据...
