详细信息

Dynamic replenishment policy for perishable goods using change point detection-based soft actor-critic reinforcement learning  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Dynamic replenishment policy for perishable goods using change point detection-based soft actor-critic reinforcement learning

作者:Kou, Aiqing[1];Cheng, Yan[1];Huang, Xiangyu[1];Jin, Jing[1]

机构:[1]East China Univ Sci & Technol, Sch Business, Shanghai 200237, Peoples R China

年份:2025

卷号:270

外文期刊名:EXPERT SYSTEMS WITH APPLICATIONS

收录:;EI(收录号:20250417756879);WOS:【SCI-EXPANDED(收录号:WOS:001410083300001)】;

语种:英文

外文关键词:Perishable goods; Change point detection; Reinforcement learning; Soft actor-critic; Replenishment policy

摘要:This paper examines the problem of establishing a dynamic replenishment policy that minimizes the costs associated with selling perishable goods. The perishable inventory is highly desired to match the realized demand. However, the demand exhibits significant non-stationarity, which is characterized by the dynamic change of stochastic demand distribution patterns. In this paper, the replenishment problem is modeled as a non- stationary Markov decision process (NSMDP) with unknown transition probabilities, and a deep reinforcement learning (DRL)-based solution framework is proposed for the NSMDP model. In this framework, the feature- enhanced long short-term memory (LSTM) is employed to detect change points in real time. On this basis, the paper develops a change point detection-based soft actor-critic (CPD-SAC) algorithm that dynamically adjusts replenishment decisions to adapt to different states across various stochastic demand distribution patterns. The numerical experiments first analyze the effect of sliding window selection on the accuracy of change point detection (CPD). Furthermore, the proposed approach is compared against several benchmark DRL algorithms and the static base stock policy. Finally, a sensitivity analysis is conducted on key parameters, including lead time, lifetime, and unit shortage cost for perishable goods. The results confirm the effectiveness of the proposed approach and demonstrate the applicability scenarios for the dynamic replenishment policy.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心