详细信息
Safe offline-to-online reinforcement learning via execution-time risk-aware selection for industrial process control ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Safe offline-to-online reinforcement learning via execution-time risk-aware selection for industrial process control
作者:Chen, Jiyang;Luo, Na[2]
机构:[1]Minist Educ, Key Lab Smart Mfg Energy Chem Proc, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China
年份:2026
卷号:214
外文期刊名:COMPUTERS & CHEMICAL ENGINEERING
收录:;EI(收录号:20262721039184);Scopus(收录号:2-s2.0-105043630703);WOS:【SCI-EXPANDED(收录号:WOS:001819526000001)】;
语种:英文
外文关键词:Safe reinforcement learning; Constrained Markov decision process; Optimal operation; Advanced process control
摘要:Industrial process control often needs online adaptation, but direct online exploration is costly and can trigger safety violations. We study a safe offline-to-online reinforcement learning (RL) method for low-fault-tolerance processes that adds execution-time risk control to our earlier backbone. Online fine-tuning is cast as an episodic constrained Markov decision process (CMDP) with an episode-level chance constraint and absorbing violation states. At execution time, Risk-Aware Selection (RAS) samples candidate actions from the actor and executes the one with the largest Lagrangian utility, computed from the reward critic and a conservative estimate of violation probability, while a single Lagrange multiplier is updated by a primal-dual rule. Experiments on penicillin fermentation and simulated moving-bed (SMB) separation show a better reward-safety trade-off during online fine-tuning without changing the underlying training backbone. The gain is larger in penicillin. In SMB, the same-backbone comparison also yields higher return and more stable rollouts. These results suggest that execution-time risk-aware selection can reduce online violation risk in industrial offline-to-online RL.
参考文献:
正在载入数据...
