详细信息

Research on Dynamic Subsidy Based on Deep Reinforcement Learning for Non-Stationary Stochastic Demand in Ride-Hailing  ( SCI-EXPANDED收录)  

文献类型:期刊文献

英文题名:Research on Dynamic Subsidy Based on Deep Reinforcement Learning for Non-Stationary Stochastic Demand in Ride-Hailing

作者:Huang, Xiangyu[1];Cheng, Yan[1];Jin, Jing[1];Kou, Aiqing[1]

机构:[1]East China Univ Sci & Technol, Sch Business, Shanghai 200237, Peoples R China

年份:2024

卷号:16

期号:15

外文期刊名:SUSTAINABILITY

收录:;WOS:【SSCI(收录号:WOS:001287191900001),SCI-EXPANDED(收录号:WOS:001287191900001)】;

基金:This work was finished when Xiangyu Huang is studying at East China University of Science and Technology. The support provided by the East China University of Science and Technology during Xiangyu Huang's pursuit of a master's degree.

语种:英文

外文关键词:ride-hailing; nonstationary stochastic demand; change point detection; non-stationary Markov decision; deep reinforcement learning

摘要:The ride-hailing market often experiences significant fluctuations in traffic demand, resulting in supply-demand imbalances. In this regard, the dynamic subsidy strategy is frequently employed by ride-hailing platforms to incentivize drivers to relocate to zones with high demand. However, determining the appropriate amount of subsidy at the appropriate time remains challenging. First, traffic demand exhibits high non-stationarity, characterized by multi-context patterns with time-varying statistical features. Second, high-dimensional state/action spaces contain multiple spatiotemporal dimensions and context patterns. Third, decision-making should satisfy real-time requirements. To address the above challenges, we first construct a Non-Stationary Markov Decision Process (NSMDP) based on the assumption of ride-hailing service systems dynamics. Then, we develop a solution framework for the NSMDP. A change point detection method based on feature-enhanced LSTM within the framework can identify the changepoints and time-varying context patterns of stochastic demand. Moreover, the framework also includes a deterministic policy deep reinforcement learning algorithm to optimize. Finally, through simulated experiments with real-world historical data, we demonstrate the effectiveness of the proposed approach. It performs well in improving the platform's profits and alleviating supply-demand imbalances under the dynamic subsidy strategy. The results also prove that a scientific dynamic subsidy strategy is particularly effective in the high-demand context pattern with more drastic fluctuations. Additionally, the profitability of dynamic subsidy strategy will increase with the increase of the non-stationary level.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心