详细信息
Reinforcement learning-guided two-stage optimization framework for multi-product batch scheduling ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Reinforcement learning-guided two-stage optimization framework for multi-product batch scheduling
作者:Zhu, Jiawen[1,2];Du, Wenli[1,2,3];Fan, Chen[1,2];Huang, Muyi[1,2];Wang, Chuan[1,2];Zhang, Furong[4]
机构:[1]East China Univ Sci & Technol, State Key Lab Ind Control Technol, Shanghai, Peoples R China;[2]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai, Peoples R China;[3]Huzhou Inst Ind Control Technol, Huzhou 313099, Peoples R China;[4]Sinopec Zhenhai Refining & Chem Co, Ningbo 315207, Peoples R China
年份:2026
卷号:204
外文期刊名:COMPUTERS & CHEMICAL ENGINEERING
收录:;EI(收录号:20253819192342);WOS:【SCI-EXPANDED(收录号:WOS:001576767200001)】;
基金:The work was supported by National Key Research and Development Program of China (2022YFB3305900) , National Natural Science Foundation of China (62394343, 62303186, 62394345) , Fundamental Research Funds for the Central Universities (222202517006) .
语种:英文
外文关键词:Multi-product batch; Production scheduling; Reinforcement learning; Mathematical programming; Two-stage optimization
摘要:With the increasing demand for high-end and fine manufacturing, multi-product batch scheduling has become essential in process industries. Its inherent complexity stems from hybrid decision variables and tightly coupled constraints. To address these challenges, this study proposes a two-stage optimization framework that integrates reinforcement learning (RL) and mathematical programming (MP). The RL layer determines batch allocations and production sequences, which are then transmitted as time windows within which the MP layer optimizes continuous variables to ensure feasibility. To handle hybrid action spaces, a mapping mechanism is introduced to unify discrete and continuous decisions. In addition, dynamic short-term targets based on reformulated constraints are designed to address the sparsity of rewards caused by long-horizon objectives. Experiments on polyolefin production scheduling demonstrate that the proposed method outperforms MP and standalone RL in terms of economic profit, production stability, and computational performance.
参考文献:
正在载入数据...
