详细信息
Multi-agent reinforcement learning for multi-objective optimization of crude distillation units ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Multi-agent reinforcement learning for multi-objective optimization of crude distillation units
作者:Xue, Dong[1];Tu, Jicheng[1];Guo, Yuan[1];Long, Jian[1];Zhao, Liang[1]
机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China
年份:2026
卷号:336
外文期刊名:CHEMICAL ENGINEERING SCIENCE
收录:;EI(收录号:20262621019004);WOS:【SCI-EXPANDED(收录号:WOS:001814402400001)】;
基金:This work was supported by the National Science and Technology Major Project (No. 2025ZD1601700) and the National Natural Science Foundation of China (No. 62373155, 62373154, and 62173147) .
语种:英文
外文关键词:Multi-agent reinforcement learning; Multi-objective optimization; Hybrid deep learning; Crude distillation unit; Dynamic operating conditions
摘要:Crude distillation units (CDUs) are vital to refinery operations, but optimizing them involves multiple conflicting objectives as well as challenges arising from process complexity, high-dimensional data, and dynamic environments, all of which limit the efficiency and adaptability of traditional methods. In this article, a novel multi-objective optimization framework based on hybrid data-driven modeling techniques and multi-agent reinforcement learning (MARL) methods is proposed for optimally operating CDUs. First, a hybrid model that integrates convolutional neural networks (CNNs), bidirectional long short-term memory networks (BiLSTMs), and attention mechanisms is developed to characterize the dynamic production process of CDUs. In particular, this developed data-driven model captures dependencies and dynamic correlations in high-dimensional process data, providing accurate predictions of production processes. Furthermore, a multi-agent deep deterministic policy gradient (MADDPG) method is employed to tackle the multi-objective real-time optimization of CDUs. More specifically, the optimization problems are modeled as partially observable Markov games, and the developed hybrid model serves as the environment, where multiple agents collaboratively learn policies to achieve Pareto-optimal solutions. Experimental evaluations demonstrate that the Pareto solutions obtained by the proposed framework are comparable or superior to those generated by traditional NSGA-II algorithms in terms of solution quality under dynamic operating conditions, effectively balancing the trade-off between product yields and energy costs. Meanwhile, the real-time performance of the proposed method ensures its applicability to the online optimization of CDUs. Robustness evaluation of the proposed framework is further conducted from multiple perspectives, including model scale, regularization, reward function sensitivity analysis and strategy selection.
参考文献:
正在载入数据...
