ProtoVAR: Efficient Dataset Distillation via Prototype-Guided Visual Autoregressive Modeling
Mingyu Wang, Wei Jiang
摘要
Recent advances in generative distillation have shown strong potential in constructing high quality surrogate datasets within a fraction of the time required by optimization-based approaches. However, most existing generative solutions rely on diffusion models, which suffer from two limitations. (i) Indirect matching objectives. Their sequential denoising process makes it difficult to directly match representative prototypes. (ii) Target-agnostic generation. The generation process is often decoupled from the target task, causing the synthesized samples to drift from the desired distribution. Building on this insight, We propose ProtoVAR, a prototype-guided visual autoregressive framework. Instead of relying on latent space, ProtoVAR uses the coarse-to-fine next-scale prediction of Visual AutoRegressive (VAR) modeling to maintain semantic consistency during generation. By injecting multi-scale class prototypes, ProtoVAR enforces clear representativeness constraints while preserving diversity. A pool-based selector further distills the prototype-guided outputs into a compact, task-aligned surrogate dataset. Extensive experiments show that ProtoVAR achieves state-of-the-art performance with comparable or lower computational cost than diffusion-based distillation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun 等ICML 2022 · 被引用 234 次
- Scaling Up Dataset Distillation to ImageNet-1K with Constant MemoryJustin Cui, Ruochen Wang, Si Si, Cho-Jui HsiehICML 2023 · 被引用 223 次
相关 Paper
- CaO2: Rectifying Inconsistencies in Diffusion-Based Dataset DistillationHaoxuan Wang, Zhenghao Zhao, Junyi Wu, Yuzhang Shang 等ICCV 2025 · 被引用 1 次
- An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and DiversitySunbeom Jeong, Sehwan Kim, Hyeonggeun Han, Hyungjun Joo 等AAAI 2026
- Fine-Tuning Visual Autoregressive Models for Subject-Driven GenerationJiwoo Chung, Sangeek Hyun, Hyunjun Kim, Eunseo Koh 等ICCV 2025 · 被引用 1 次
- Efficient Dataset Distillation via Minimax DiffusionJianyang Gu, Saeed Vahidian, Vyacheslav Kungurtsev, Haonan Wang 等CVPR 2024
- ProtoVAE: A Trustworthy Self-Explainable Prototypical Variational ModelSrishti Gautam, Ahcène Boubekki, Stine Hansen, Suaiba Amina Salahuddin 等NeurIPS 2022 · 被引用 55 次
