详细信息

Perceiving Multiple Representations for scene text image super-resolution guided by text recognizer  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Perceiving Multiple Representations for scene text image super-resolution guided by text recognizer

作者:Shi, Qin[1,4];Zhu, Yu[1];Liu, Yatong[1];Ye, Jiongyao[1];Yang, Dawei[2,3,4]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Fudan Univ, Zhongshan Hosp, Dept Pulm & Crit Care Med, Shanghai 200032, Peoples R China;[3]Fudan Univ, Zhongshan Hosp Xiamen, Dept Pulm & Crit Care Med, Shanghai 361015, Peoples R China;[4]Shanghai Engn Res Ctr Internet Things Resp Med, Shanghai 200032, Peoples R China

年份:2023

卷号:124

外文期刊名:ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE

收录:;EI(收录号:20232414209914);WOS:【SCI-EXPANDED(收录号:WOS:001019764000001)】;

基金:This work was supported in part by the National Natural Science Foundation of China under Grant 82170110, and the Science and Technology Commission of Shanghai Municipality, China under Grant 20DZ22544000, 21DZ2200600, 20DZ2261200, ZD2021CY001. Fujian Province Department of Science and Technology, China (2022D014)

语种:英文

外文关键词:Scene text image super-resolution; Scene text recognition; Contextual information; Visual features; Frequency domain learning

摘要:Single image super-resolution (SISR) aims to recover clear high-resolution images from low-resolution images, which has made great progress with the development of deep learning these years. Scene text image super -resolution (STISR) is a subfield of SISR with the goal of increasing the resolution of a low-resolution text image and enhancing the readability of characters in the image. Despite significant improvements in recent approaches, STISR remains a challenging task due to the diversity of background, text appearances and layouts, etc. This paper presents a Perceiving Multiple Representations (PerMR) method for better super -resolution performances in scene text images. PerMR is a unified network that combines super-resolution with text recognition and exploits the recognizer's feedback to facilitate super-resolution. Specifically, contextual information from the text decoder is extracted to provide sequence-specific guidance and enable the super -resolution model to pay more attention to the text region. Meanwhile, low-level and high-level visual features from the vision backbone of the recognition network are integrated to further improve visual quality. Additionally, we incorporate a frequency branch into the vanilla convolution unit, which efficiently enhances global and local feature representations. Experiments on the STISR benchmark dataset TextZoom validate that PerMR can not only generate more distinguishable images, but also outperforms the current state-of-the-art methods. PerMR boosts the average recognition accuracy by 5.9% using ASTER, 5.8% using MORAN and 10.6% using CRNN compared to the baseline model TSRN. PerMR outperforms the advanced method TPGSR-3 by 1.4% on ASTER, 0.1% on MORAN, 0.2% on CRNN and boosts TATT by 0.6% on ASTER and 1.1% on MORAN respectively. Furthermore, PerMR demonstrates good robustness and generalization when tackling low-quality text images in multiple scene text recognition datasets. The experiment results verify the capabilities of PerMR to boost text recognition performance.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心