详细信息
Slimmable neural architecture design based on cross architecture and token distillation ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Slimmable neural architecture design based on cross architecture and token distillation
作者:Qiu, Guhao[1];Chen, Zhihua[1];Dai, Lei[1];Li, Ping[2,3];Sheng, Bin[4]
机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[2]Hong Kong Polytech Univ, Dept Comp, Hong Kong, Peoples R China;[3]Hong Kong Polytech Univ, Sch Design, Hong Kong, Peoples R China;[4]Shanghai Jiao Tong Univ, Dept Comp Sci & Engn, Shanghai 200240, Peoples R China
年份:2026
卷号:166
外文期刊名:ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE
收录:;EI(收录号:20255319834236);WOS:【SCI-EXPANDED(收录号:WOS:001655506400001)】;
基金:Acknowledgments This work was supported by the National Natural Science Foundation of China (Grant Number: 62272164, Grant Number: 62572188, and Grant Number: 62306113) .
语种:英文
外文关键词:Slimmable neural network; Self-supervised learning; Masked autoencoder
摘要:The manually designed neural networks have the drawbacks of requiring a large amount of training data and high computational costs. In this paper, we propose the masked autoencoder based lightweight network search algorithm which leverages the efficient channel search algorithm and specific distillation strategy to obtain the optimal architecture. During SuperNet training process, we design the cross-token distillation and cross-architecture strategy. Token distillation strategy is used to enforce the similar representation obtained from different masks in one image. Architecture distillation strategy is used to fully utilize representation from the sampled subnetwork and use the feature maps from one image but the same token. In the subnetwork searching process, we further pretrain the selected network considering the component dependency. Comprehensive experiments verify that our proposed method is efficient and flexible than baseline self-supervised learning algorithm and structured pruning algorithms. For example, our method obtains 4.4% improvement in TOP-1 metrics compared with the classic Masked Autoencoder algorithm designed lightweight Transformer architecture with less than 10M parameter.
参考文献:
正在载入数据...
