详细信息
A Meta Reinforcement Learning-Based Task Offloading Strategy for IoT Devices in an Edge Cloud Computing Environment ( SCI-EXPANDED收录)
文献类型:期刊文献
英文题名:A Meta Reinforcement Learning-Based Task Offloading Strategy for IoT Devices in an Edge Cloud Computing Environment
作者:Yang, He[1];Ding, Weichao[2];Min, Qi[2];Dai, Zhiming[2,3];Jiang, Qingchao[2];Gu, Chunhua[2]
机构:[1]SINOPEC Res Inst Petr Proc Co Ltd, Beijing 100083, Peoples R China;[2]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[3]Shanghai Jian Qiao Univ, Sch Informat Technol, Shanghai 201306, Peoples R China
年份:2023
卷号:13
期号:9
外文期刊名:APPLIED SCIENCES-BASEL
收录:;WOS:【SCI-EXPANDED(收录号:WOS:000987248300001)】;
基金:This work was sponsored by the Shanghai Sailing Program (No. 20YF1410900), the Shanghai Natural Science Foundation (23ZR1414900), the National Natural Science Foundation (No. 61472139), the Shanghai Automobile Industry Science and Technology Development Foundation (No. 1915), and the Shanghai Science and Technology Innovation Action Plan (No. 20dz1201400). Any opinions, findings, and conclusions are those of the authors, and do not necessarily reflect the views of the above agencies.
语种:英文
外文关键词:task offloading; mobile edge computing; meta reinforcement learning; IoT devices
摘要:Developing an effective task offloading strategy has been a focus of research to improve the task processing speed of IoT devices in recent years. Some of the reinforcement learning-based policies can improve the dependence of heuristic algorithms on models through continuous interactive exploration of the edge environment; however, when the environment changes, such reinforcement learning algorithms cannot adapt to the environment and need to spend time on retraining. This paper proposes an adaptive task offloading strategy based on meta reinforcement learning with task latency and device energy consumption as optimization targets to overcome this challenge. An edge system model with a wireless charging module is developed to improve the ability of IoT devices to provide service constantly. A Seq2Seq-based neural network is built as a task strategy network to solve the problem of difficult network training due to different dimensions of task sequences. A first-order approximation method is proposed to accelerate the calculation of the Seq2Seq network meta-strategy training, which involves quadratic gradients. The experimental results show that, compared with existing methods, the algorithm in this paper has better performance in different tasks and network environments, can effectively reduce the task processing delay and device energy consumption, and can quickly adapt to new environments.
参考文献:
正在载入数据...
