详细信息
EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning in Vision Transformers ( EI收录)
文献类型:期刊文献
英文题名:EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning in Vision Transformers
作者:Liao, Wenwen[1]; Ruan, Hang[1]; Yu, Jianbo[2]; Song, Bing[3]; Wang, Yuansong[4]; Yang, Xiaofeng[2]
机构:[1] College of Intelligent Robotics and Advance Manufacturing, Fudan University, Shanghai, China; [2] School of Microelectronics, Fudan University, Shanghai, China; [3] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China; [4] Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China
年份:2026
外文期刊名:arXiv
收录:EI(收录号:20260057930)
语种:英文
外文关键词:Classification (of information) - Machine vision - Program processors - Query processing - Structured Query Language - Tuning
摘要:Large models such as Vision Transformers (ViTs) have demonstrated remarkable superiority over smaller architectures like ResNet in few-shot classification, owing to their powerful representational capacity. However, fine-tuning such large models demands extensive GPU memory and prolonged training time, making them impractical for many real-world low-resource scenarios. To bridge this gap, we propose EfficientFSL, a query-only fine-tuning framework tailored specifically for few-shot classification with ViT, which achieves competitive performance while significantly reducing computational overhead. EfficientFSL fully leverages the knowledge embedded in the pre-trained model and its strong comprehension ability, achieving high classification accuracy with an extremely small number of tunable parameters. Specifically, we introduce a lightweight trainable Forward Block to synthesize task-specific queries that extract informative features from the intermediate representations of the pre-trained model in a query-only manner. We further propose a Combine Block to fuse multi-layer outputs, enhancing the depth and robustness of feature representations. Finally, a Support-Query Attention Block mitigates distribution shift by adjusting prototypes to align with the query set distribution. With minimal trainable parameters, EfficientFSL achieves state-of-the-art performance on four in-domain few-shot datasets and six cross-domain datasets, demonstrating its effectiveness in real-world applications. Copyright ? 2026, The Authors. All rights reserved.
参考文献:
正在载入数据...
