详细信息

Temporal-Aware DDPG for Delay and Stability Optimization in Vehicular Edge Computing  ( EI收录)  

文献类型:期刊文献

英文题名:Temporal-Aware DDPG for Delay and Stability Optimization in Vehicular Edge Computing

作者:Fan, Guisheng[1,2,3]; Wang, Jiatong[1]; Zhang, Hengrun[1,4]; Yu, Huiqun[1]

机构:[1] East China University of Science and Technology, Department of Computer Science and Engineering, Shanghai, 200237, China; [2] Shanghai Key Laboratory of Computer Software Evaluating and Testing, Shanghai, 201112, China; [3] Shanghai Engineering Research Center of Smart Energy, Shanghai, 201103, China; [4] Nanjing University, State Key Laboratory for Novel Software Technology, Nanjing, 210023, China

年份:2026

外文期刊名:IEEE Transactions on Vehicular Technology

收录:EI(收录号:20262921133132);Scopus(收录号:2-s2.0-105045110757)

语种:英文

外文关键词:Computation offloading - Computation theory - Lyapunov methods - Memory architecture - Network architecture - Optimization - Queueing theory

摘要:Roadside Units (RSUs) deployed along roads provide proximate computation services in Vehicular Edge Computing (VEC). However, heterogeneous node capabilities and diverse task sources make offloading decisions highly challenging. To address this, we propose an improved Deep Deterministic Policy Gradient (DDPG) method, called Temporal Lyapunov-enhanced DDPG (T-LyDDPG), which integrates Lyapunov optimization theory into a DDPG framework. In T-LyDDPG, a Long Short- Term Memory (LSTM) module is incorporated into both the Actor and Critic networks to capture temporal dependencies and remember past state information, while an attention mechanism is used to highlight and extract the most relevant features of the high-dimensional state. These designs allow the agent to smooth its decisions over time and focus on key state elements, thus improving decision stability and accuracy under dynamic conditions. We also adopt a double-Critic architecture, which mitigates Q-value overestimation and lowers target variance by taking the minimum across two Critics, thereby improving policy stability under time-varying arrivals and channels. In addition, we employ Prioritized Experience Replay (PER) to improve sample efficiency by focusing updates on rare but consequential transitions with high Temporal Difference (TD) error, while using importance sampling to control bias. Through Lyapunov drift-plus-penalty design, the long-term optimization of average task delay and queue stability is decomposed into perslot decisions. The simulation results show that our method has great advantages in metrics such as average reward, task delay, queue backlog, and completion ratio. ? 1967-2012 IEEE.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心