详细信息

TSRGAN: Real-world text image super-resolution based on adversarial learning and triplet attention  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:TSRGAN: Real-world text image super-resolution based on adversarial learning and triplet attention

作者:Fang, Chuantao[1];Zhu, Yu[1];Liao, Lei[1];Ling, Xiaofeng[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China

年份:2021

卷号:455

起止页码:88

外文期刊名:NEUROCOMPUTING

收录:;EI(收录号:20212410501670);WOS:【SCI-EXPANDED(收录号:WOS:000672811000009)】;

基金:The authors greatly appreciate the financial supports of Shanghai Association for Science and Technology under Grant 17DZ1100808, Natural Science Foundation of Shanghai under Grant 19ZR1413400.

语种:英文

外文关键词:Text image super-resolution; Adversarial learning; Triplet attention; Wavelet loss; Scene text recognition

摘要:The text in a low-resolution (LR) image is usually hard to read. Super-resolution (SR) is an intuitive solution to this issue. Existing single image super-resolution (SISR) models are mainly trained on synthetic datasets whose LR images are obtained by performing bicubic interpolation or gaussian blur on high-resolution (HR) images. However, these models can hardly generalize to practical scenarios because real-world LR images are more difficult to super-resolve. The newly proposed TextZoom dataset is the first dataset for real-world text image super-resolution. We propose a new model termed TSRGAN trained on this dataset. First, a discriminator is designed to prevent the SR network from generating over-smoothed images. Second, we introduce triplet attention into the SR network for better representational ability. Moreover, besides L-2 loss and adversarial loss, wavelet loss is incorporated to help reconstruct sharper character edges. Since TextZoom provides text labels, the recognition accuracy of scene text recognition (STR) model can be used to evaluate the quality of SR images. It can reflect the performance of text image SR models better than traditional SR evaluation metrics such as PSNR and SSIM. Comprehensive experiments show the superiority of our TSRGAN. Compared with the state-of-the-art method, the proposed TSRGAN improves the average recognition accuracy of ASTER, MORAN and CRNN by 0.8%, 1.5% and 3.2% on TextZoom respectively. (C) 2021 Elsevier B.V. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心