详细信息

基于改进A2C目标驱动的室内无地图导航方法    

An Indoor Map-free Navigation Method Based on Improved A2C Target-driven Model

文献类型:期刊文献

中文题名:基于改进A2C目标驱动的室内无地图导航方法

英文题名:An Indoor Map-free Navigation Method Based on Improved A2C Target-driven Model

作者:王彦臻[1,2];胡晗[1,2];李文倩[1,2];袁士博[1,2];和望利[1,2]

机构:[1]华东理工大学信息科学与工程学院,上海200237;[2]华东理工大学能源化工过程智能制造教育部重点实验室,上海200237

年份:2022

卷号:29

期号:3

起止页码:474

中文期刊名:控制工程

外文期刊名:Control Engineering of China

收录:CSTPCD;;北大核心:【北大核心2020】;CSCD:【CSCD_E2021_2022】;

基金:国家优秀青年科学基金资助项目(61922030)。

语种:中文

中文关键词:深度强化学习;室内导航;视觉目标驱动模型;A2C模型

外文关键词:Deep reinforcement learning;indoor navigation;a visual-target-driven model;A2C model

摘要:室内无先验地图场景下的目标驱动式导航是机器人领域的公认难题,近年来兴起的深度强化学习方法为该问题的求解提供了新思路,同时也产生了诸如模型泛化能力不足、难以收敛的新问题。为解决上述问题,提出了一种基于深度强化学习的视觉目标驱动式室内无地图导航方法,设计了一种新的稠密奖励机制,同时引入目标驱动模型并嵌入深度残差网络进行场景特征提取,通过Actor-Critic强化学习算法进行模型训练。以室内导航模拟器Ai2thor为仿真环境,通过对比实验验证了算法具有更快的训练收敛速率及良好的泛化性能。
Target-driven navigation in indoor environment without prior maps is a recognized problem in the field of robotics.In recent years,the deep reinforcement learning method has provided a new idea for solving the problem.At the same time,new problems have emerged,such as insufficient generalization ability and difficult convergence.In order to solve the above problems,a visual-target-driven indoor map-free navigation method based on deep reinforcement learning is proposed in this paper.A new dense reward mechanism is designed,a target-driven model is introduced and a deep residual network is embedded to extract scene features.The model is trained through Actor-Critic reinforcement learning algorithm.The indoor navigation simulator Ai2thor is used as the simulation environment for comparative experiments.The experimental results show that the algorithm has faster training convergence rate and good generalization performance.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心