详细信息

Gradient amplification for gradient matching based dataset distillation  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Gradient amplification for gradient matching based dataset distillation

作者:Zhang, Jingxuan[1];Chen, Zhihua[1];Dai, Lei[1];Li, Ping[2,3];Sheng, Bin[4]

机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[2]Hong Kong Polytech Univ, Dept Comp, Hong Kong, Peoples R China;[3]Hong Kong Polytech Univ, Sch Design, Hong Kong, Peoples R China;[4]Shanghai Jiao Tong Univ, Dept Comp Sci & Engn, Shanghai 200240, Peoples R China

年份:2025

卷号:191

外文期刊名:NEURAL NETWORKS

收录:;EI(收录号:20252818734856);WOS:【SCI-EXPANDED(收录号:WOS:001529939200008)】;

基金:This work was supported by the National Natural Science Foundation of China (Grant No. 62272164 and No. 62306113) .

语种:英文

外文关键词:Dataset distillation; Gradient matching; Image classification; Dataset ensembling

摘要:Dataset distillation (DD) aims to construct a smaller dataset compared to the original cumbersome one. Models trained on both datasets are expected to achieve almost the same accuracy on the test set. Previous work using a gradient matching (GM) framework achieved suboptimal performance because it only matched the gradient information for correct labels, neglecting to account for the model's surprise regarding incorrect answers. In this paper, we aim to produce more informative gradient information during the matching process and present a novel framework for DD by leveraging label cycle shifting. Specifically, it involves using pre-trained neural networks to process mismatched image-label pairs, resulting in the generation of diverse and substantial gradients during the backpropagation of the cross-entropy loss. Furthermore, GM with larger gradients tends to converge more rapidly compared to conventional GM approaches, prompting us to propose an early exit mechanism. To enhance the performance further, we employ an ensemble approach by applying an exponential moving average to the distilled dataset and introduce distribution matching to the total matching function. We demonstrate that the model implicitly considers gradient experiences from past rounds, and we have delved into mechanisms where gradient matching and distribution matching mutually enhance each other. Our design can outperform most previous DD methods with fewer training iterations. Experiments on the benchmark datasets (CIFAR10, CIFAR100, TinyImageNet, and a subset of ImageNet) present the effectiveness of our method.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心