详细信息
RAOD: refined oriented detector with augmented feature in remote sensing images object detection ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:RAOD: refined oriented detector with augmented feature in remote sensing images object detection
作者:Shi, Qin[1];Zhu, Yu[1];Fang, Chuantao[1];Wang, Nan[1];Lin, Jiajun[1]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China
年份:2022
卷号:52
期号:13
起止页码:15278
外文期刊名:APPLIED INTELLIGENCE
收录:;EI(收录号:20221111786251);WOS:【SCI-EXPANDED(收录号:WOS:000767910000001)】;
基金:The authors greatly appreciate the financial supports of the Shanghai Association for Science and Technology under Grant 17DZ1100808.
语种:英文
外文关键词:Remote sensing image; Oriented object detection; Augmented feature pyramid; Deformable RoI pooling; Rotated RoI align
摘要:Object detection is a challenging task in remote sensing. Aerial images are distinguished by complex backgrounds, arbitrary orientations, and dense distributions. Considering those difficulties, this paper proposes a two-stage refined oriented detector with augmented features named RAOD. First, a novel Augmented Feature Pyramid Network (A-FPN) is built to enhance fusion both in spatial and channel dimensions. Specifically, it mainly consists of three modules: Scale Transfer Module (STM), Feature Aggregate Module (FAM) and Feature Refinement Module (FRM). STM reduces information loss when fusing features in the top-down pathway. FAM aggregates features from different scales. FRM aims to refine the integrated features using a lightweight attention module. Then, we adopt a two-step processing, which consists of a coarse stage and a refinement stage. In the coarse stage, deformable RoI pooling is adopted to improve the network's ability of modeling spatial transformations and then horizontal proposals are transformed into oriented ones. In the refinement stage, Rotated RoI align (RRoI align) is used to extract rotation-invariant features from rotated RoIs and further optimize the localization. To enhance stability and robustness during training, smooth Ln is chosen as regression loss as it has better ability in terms of robustness and stability than smooth L-1 loss. Extensive experiments on several rotation detection datasets demonstrate the effectiveness of our method. Results show that our method is able to achieve 79.78%, 74.7% and 94.82% on DOTA-v1.0, DOTA-v1.5 and HRSC2016, respectively.
参考文献:
正在载入数据...
