详细信息
文献类型:期刊文献
中文题名:基于多模态特征融合的小样本图像分类方法
英文题名:Few-shot image classification method based on multimodal feature fusion
作者:李琛琛[1];王喆[1];肖婷[1];马学铭[1];杨孟平[1]
机构:[1]华东理工大学信息科学与工程学院,上海200237
年份:2025
卷号:46
期号:9
起止页码:2441
中文期刊名:计算机工程与设计
外文期刊名:Computer Engineering and Design
收录:;北大核心:【北大核心2023】;
基金:国家自然科学基金项目(62076094);国防科技领域基金项目(2021-JCJQ-JJ-0041)。
语种:中文
中文关键词:小样本图像分类;多模态;注意力机制;特征筛选;语义融合;视觉语言模型;大语言模型
外文关键词:few-shot image classification;multimodal;attention mechanism;feature selection;semantic fusion;visual language model;large language models
摘要:为提升小样本图像分类任务中的模型性能,提出了一种基于多模态特征融合的小样本图像分类方法(AMTAdapter)。该方法通过引入注意力机制模块增强模型对图像的关注能力,利用类内相似度与类间相似度对特征进行筛选,并采用MiniGPT4生成描述,结合CLIP模型获取语义信息,实现图像与文本信息的融合。在9个图像分类数据集上的实验结果表明,该方法在小样本场景下有效提升了分类准确率,展现了良好的泛化能力和适用性。研究还通过消融实验验证了各模块对性能提升的贡献,验证了所提方法的有效性。
To improve the model performance in few-shot image classification tasks,a few-shot image classification method based on multimodal feature fusion(AMT-Adapter)was proposed.The model’s attention to images was enhanced by incorporating an attention mechanism module,while features were filtered based on intra-class and inter-class similarity metrics.Descriptions were generated via MiniGPT4,and semantic information was extracted using the CLIP model,achieving the integration of visual and textual information.Experimental results on nine image classification datasets demonstrate that the proposed method effectively improves classification accuracy under few-shot scenarios,exhibiting strong generalization capability and applicability.Ablation studies further validate the contribution of each module to performance gains,confirming the effectiveness of the proposed method.
参考文献:
正在载入数据...
