详细信息
Hybrid CNN-RWKV with high-frequency enhancement for real-world chinese-english scene text image super-resolution ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Hybrid CNN-RWKV with high-frequency enhancement for real-world chinese-english scene text image super-resolution
作者:Liu, Yanbin[1];Zhu, Yu[1,2];Li, Hangyu[1];Ling, Xiaofeng[1]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Shanghai Engn Res Ctr Internet Things Resp Med, Shanghai 200032, Peoples R China
年份:2025
卷号:55
期号:13
外文期刊名:APPLIED INTELLIGENCE
收录:;EI(收录号:20253619108290);WOS:【SCI-EXPANDED(收录号:WOS:001563530300007)】;
基金:This work was supported in part by the National Natural Science Foundation of China under Grant 62476088, and the Shanghai Automotive Industry Science and Technology Development Foundation under Grant 2304.
语种:英文
外文关键词:Text image super-resolution; Recurrent bi-directional WKV attention; Multi-scale large kernel convolution; High-frequency enhancement; Multi-frequency channel attention
摘要:Existing scene text image super-resolution (STISR) methods primarily focus on the restoration of fixed-size English text images. Compared to English characters, Chinese characters present a greater variety of categories and more intricate stroke structures. In recent years, Transformer-based methods have achieved significant progress in image super-resolution task, but face the dilemma between global modeling and efficient computation. The emerging Receptance Weighted Key Value (RWKV) model can serve as a promising alternative to Transformer, enabling long-distance modeling with linear computational complexity. In this paper, we propose a Hybrid CNN-RWKV with High-Frequency Enhancement (HCR-HFE) model for STISR task. First, we design a recurrent bidirectional WKV (Re-Bi-WKV) attention which integrates bidirectional WKV (Bi-WKV) attention with a recurrent mechanism. Bi-WKV achieves global receptive field with linear complexity, while the recurrent mechanism establishes 2D image dependencies from different scanning directions. Additionally, a computationally efficient high-frequency enhancement module (HFEM) is incorporated to enhance high-frequency details, such as character edge information. Furthermore, we design a multi-scale large kernel convolutional (MLKC) block which integrates large kernel decomposition, gated aggregation and multi-scale mechanism to capture various-range dependencies with reduced computational cost. Finally, we introduce a multi-frequency channel attention (MFCA) which extends channel attention to the frequency domain, enabling the model to focus on critical features. Extensive experiments on real-world Chinese-English (Real-CE) dataset demonstrate that HCR-HFE outperforms previous methods in both quantitative metrics and visual results. Furthermore, HCR-HFE achieves excellent performance on natural image datasets, demonstrating its broad applicability.
参考文献:
正在载入数据...
