详细信息

An offline-to-online reinforcement learning framework with trajectory-guided exploration for industrial process control  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:An offline-to-online reinforcement learning framework with trajectory-guided exploration for industrial process control

作者:Chen, Jiyang[1];Luo, Na[2]

机构:[1]Minist Educ, Key Lab Smart Mfg Energy Chem Proc, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai, 200237, Peoples R China

年份:2025

卷号:154

外文期刊名:JOURNAL OF PROCESS CONTROL

收录:;EI(收录号:20253519089283);WOS:【SCI-EXPANDED(收录号:WOS:001564582900001)】;

语种:英文

外文关键词:Offline reinforcement learning; Process optimization; Expert knowledge integration; Interactive imitation learning

摘要:Reinforcement learning (RL) in industrial process control faces critical challenges, including limited data availability, unsafe exploration, and the high cost of high-fidelity simulators. These issues limit the practical adoption of RL in process control systems. To address these limitations, this paper presents a comprehensive framework that combines offline pre-training with online finetuning. Specifically, the framework first employs offline RL method to learn conservative policies from historical data, preventing overestimation of unseen actions. It then transitions to fine-tuning using online RL method with a mixed replay buffer that gradually shifts from offline to online data. To further enhance safety during online exploration, this work introduces a trajectory-guided strategy that leverages timestamped sub-optimal expert demonstrations. Rather than replacing agent actions entirely, the proposed method computes a weighted combination of agent and expert actions based on a decaying intervention rate. Both components are designed as modular additions that can be integrated into existing actor-critic algorithms without structural modifications. Case studies on penicillin fermentation and simulated moving bed (SMB) processes demonstrate that the proposed framework outperforms baseline algorithms in terms of learning efficiency, stability, computation costs, and operational safety.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心