详细信息
Dual-Guided Frequency Prototype Network for Few-Shot Semantic Segmentation ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Dual-Guided Frequency Prototype Network for Few-Shot Semantic Segmentation
作者:Wen, Chunlin[1];Huang, Hui[2];Ma, Yan[1];Yuan, Feiniu[1,3,4];Zhu, Hongqing[5]
机构:[1]Shanghai Normal Univ, Coll Informat Mech & Elect Engn, Shanghai 201418, Peoples R China;[2]Shanghai Normal Univ, Coll Informat Mech & Elect Engn, Shanghai Engn Res Ctr Intelligent Educ & Bigdata, Shanghai 201418, Peoples R China;[3]Res Base Online Educ Shanghai Middle & Primary Sch, Shanghai, Peoples R China;[4]Shanghai Engn Res Ctr Intelligent Educ & Big Data, Shanghai 201418, Peoples R China;[5]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China
年份:2024
卷号:26
起止页码:8874
外文期刊名:IEEE TRANSACTIONS ON MULTIMEDIA
收录:;EI(收录号:20241515873570);WOS:【SCI-EXPANDED(收录号:WOS:001297535300009)】;
基金:This work was supported by the National Nature Science Foundation of China under Grant 61872143.
语种:英文
外文关键词:Prototypes; Feature extraction; Frequency-domain analysis; Semantic segmentation; Task analysis; Predictive models; Training; Few-shot segmentation; few-shot learning; prototype learning; frequency domain learning; dual-guidance
摘要:Few-shot semantic segmentation is a challenging task that aims to segment novel classes in the query images given only a few annotated support samples. Most existing prototype-based approaches extract global or local prototypes by global average pooling (GAP) or clustering to represent all object information. Subsequently, the prototype information is employed as guidance for query image segmentation. However, these frameworks fail to fully mine the object details and ignore information from query images. Consequently, we propose a Dual-Guided Frequency Prototype Network (DGFPNet) to solve these issues. Specifically, to mine the global and local object information, a Frequency Prototype Generation Module (FPGM) is first proposed to extract more comprehensive frequency prototypes by multi-frequency pooling (MFP) in the DCT domain. Then, with the guidance of support and query information, a Dual-Guided Selection Module (DGSM) is presented to produce the query attention mask and select more effective prototypes. Based on the query attention mask and support information, the generalized object information is integrated into the feature with the proposed Feature Generalization Module (FGM). Finally, we propose a Multi-Dimension Feature Enrichment Decoder Module (MDFEDM) to capture multi-dimension object information and tackle hard pixels for refining the final segmentation results. Extensive experiments on PASCAL- 5(i) and COCO- 20(i) show that our model achieves new state-of-the-art performances.
参考文献:
正在载入数据...
