详细信息

Scene text image super-resolution algorithm based on directional feature modeling  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Scene text image super-resolution algorithm based on directional feature modeling

作者:Guo, Qin[1];Wang, Nan[1];Hong, Yubo[1];Liu, Yanbin[1];Zhu, Yu[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Meilong Rd, Shanghai 200237, Peoples R China

年份:2026

卷号:32

期号:4

外文期刊名:MULTIMEDIA SYSTEMS

收录:;EI(收录号:20262020730824);WOS:【SCI-EXPANDED(收录号:WOS:001758553200002)】;

基金:This work was supported by the National Natural Science Foundation of China under Grant 62476088, and the Science and Technology Commission of Shanghai Municipality under Grant 20DZ2254400.

语种:英文

外文关键词:Scene text image super-resolution; Direction feature modeling; State space model; Feature fusion; Gate control mechanism

摘要:Scene Text Image Super-Resolution (STISR) is a critical method for enhancing the readability of low-resolution text images and making text recognition tasks more accurate. Its core challenge is how to effectively model the unique textual feature of text images. Although current methods that are based on Convolutional Neural Networks (CNNs) and Transformers have achieved notable progress, their inherent modeling mechanisms fail to precisely capture the distinctive directional features present in text images. To address this issue, this paper extends the sequential modeling approach of state space models to the modeling of pixel arrangement directions in text images, thereby capturing vertical stroke structures and horizontal semantic correlations between characters. Building on this idea, we propose a novel Scene Text Image Super-Resolution framework based on directional feature modeling (DirTextSR). Within this framework, the Fusion State Space Model conducts integrated modeling of image features and text priors along the vertical, horizontal, and their reverse directions of the feature maps, while incorporating a multi-branch structure to capture feature information at different levels. The Gated State Space Model also has a special gating mechanism that controls the flow of character semantic relationships and stroke structure features separately within the model, with the final output controlled via a Gated Recurrent Unit (GRU). Additionally, Directional Gradient Loss is constructed to reinforce error propagation for direction-dependent features. Experiments demonstrate that DirTextSR outperforms other state-of-the-art methods by achieving an average improvement of 1.3% in text recognition accuracy on the TextZoom dataset, and that the proposed framework can effectively collaborate with existing text prior extraction methods. Ablation studies further validate the rationality and effectiveness of the DirTextSR.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心