详细信息
Dynamic replenishment policy for perishable goods using change point detection-based soft actor-critic reinforcement learning ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Dynamic replenishment policy for perishable goods using change point detection-based soft actor-critic reinforcement learning
作者:Kou, Aiqing[1];Cheng, Yan[1];Huang, Xiangyu[1];Jin, Jing[1]
机构:[1]East China Univ Sci & Technol, Sch Business, Shanghai 200237, Peoples R China
年份:2025
卷号:270
外文期刊名:EXPERT SYSTEMS WITH APPLICATIONS
收录:;EI(收录号:20250417756879);WOS:【SCI-EXPANDED(收录号:WOS:001410083300001)】;
语种:英文
外文关键词:Perishable goods; Change point detection; Reinforcement learning; Soft actor-critic; Replenishment policy
摘要:This paper examines the problem of establishing a dynamic replenishment policy that minimizes the costs associated with selling perishable goods. The perishable inventory is highly desired to match the realized demand. However, the demand exhibits significant non-stationarity, which is characterized by the dynamic change of stochastic demand distribution patterns. In this paper, the replenishment problem is modeled as a non- stationary Markov decision process (NSMDP) with unknown transition probabilities, and a deep reinforcement learning (DRL)-based solution framework is proposed for the NSMDP model. In this framework, the feature- enhanced long short-term memory (LSTM) is employed to detect change points in real time. On this basis, the paper develops a change point detection-based soft actor-critic (CPD-SAC) algorithm that dynamically adjusts replenishment decisions to adapt to different states across various stochastic demand distribution patterns. The numerical experiments first analyze the effect of sliding window selection on the accuracy of change point detection (CPD). Furthermore, the proposed approach is compared against several benchmark DRL algorithms and the static base stock policy. Finally, a sensitivity analysis is conducted on key parameters, including lead time, lifetime, and unit shortage cost for perishable goods. The results confirm the effectiveness of the proposed approach and demonstrate the applicability scenarios for the dynamic replenishment policy.
参考文献:
正在载入数据...
