详细信息
Dataset Distillation via a Noise-Unconstrained Generative Model ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Dataset Distillation via a Noise-Unconstrained Generative Model
作者:Zhang, Jingxuan[1,2];Dai, Lei[1,2];Ye, Fei[3];Chen, Zhihua[1,2];Li, Ping[4];Yang, Xiaokang[5];Sheng, Bin[5]
机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, State Key Lab Ind Control Technol, Shanghai 200237, Peoples R China;[3]Univ Elect Sci & Technol China, Sch Informat & Software Engn, Chengdu 610054, Peoples R China;[4]Hong Kong Polytech Univ, Dept Comp, Hong Kong, Peoples R China;[5]Shanghai Jiao Tong Univ, Sch Comp Sci, Shanghai 200240, Peoples R China
年份:2026
卷号:48
期号:9
起止页码:10809
外文期刊名:IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
收录:;EI(收录号:20262020701540);WOS:【SCI-EXPANDED(收录号:WOS:001843530100009)】;
基金:This work was supported by the National Natural Science Foundation of China under Grant T2525004, Grant 62272164, Grant 62572188, and Grant 62306113, in part by Shenzhen Medical Research Fund under Grant E250200613, and in part by the State Key Laboratory of Industrial Control Technology, China under Grant ICT2026A25.
语种:英文
外文关键词:Circuits and systems; Pixel; Radio access networks; Regional area networks; Electronic mail; Network architecture; Massive machine type communications; Videos; Digital images; Video equipment; Dataset distillation; generative adversarial network; diffusion model; image classification
摘要:Dataset distillation (DD) aims to synthesize a more compact dataset than the original one and models trained on it are expected to have the same generalization capabilities as on the original dataset. Previous work via a generative model (GM) faces several limitations. First, GM struggles to generate representative samples due to a lack of constraints. Second, it overlooks the relationships between generated samples, limiting its effectiveness. In this paper, a new noise-unconstrained GM-based DD framework is proposed. In the distillation stage, an adaptive matching coefficient is introduced to align generated images with representative class elements and the MiniMax loss function is extended to reduce the optimization difficulty. In the deployment stage, features among each generative image are ensembled by gradient-matching based DD. Theoretical analysis based on McDiarmid's inequality demonstrates that the proposed components can reduce the generalization error of the original baseline method. We also provide insights into the potential of generated images as an effective proxy dataset for DD. For example, on the ImageWoof dataset with 50 distilled images per class using a 6-layer ConvNet for evaluation, generated images outperform 25%, 50%, and 75% original images by 8.4%, 6.3%, and 8.3% in distillation performance. Our method effectively handles both low- and high-resolution datasets, with experiments on 11 benchmarks demonstrating its efficacy.
参考文献:
正在载入数据...
