OPTICAL: Leveraging Optimal Transport for Contribution Allocation in Dataset Distillation
Xiao Cui, Yulei Qin, Wengang Zhou, Hongsheng Li, Houqiang Li
摘要
The demands for increasingly large-scale datasets pose substantial storage and computation challenges to building deep learning models. Dataset distillation methods, especially those via sample generation techniques, rise in response to condensing large original datasets into small synthetic ones while preserving critical information. Existing subset synthesis methods simply minimize the homogeneous distance where uniform contributions from all real instances are allocated to shaping each synthetic sample. We demonstrate that such equal allocation fails to consider the instance-level relationship between each real-synthetic pair and gives rise to insufficient modeling of geometric structural nuances between the distilled and original sets. In this paper, we propose a novel framework named OP-TICAL to reformulate the homogeneous distance minimization into a bi-level optimization problem via matching-andapproximating. In the matching step, we leverage optimal transport matrix to dynamically allocate contributions from real instances. Subsequently, we polish the generated samples in accordance with the established allocation scheme for approximating the real ones. Such a strategy better measures intricate geometric characteristics and handles intraclass variations for high fidelity of data distillation. Extensive experiments across seven datasets and three model architectures demonstrate our method's versatility and effectiveness. Its plug-and-play characteristic makes it compatible with a wide range of distillation frameworks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Efficient Multimodal Dataset Distillation via Generative ModelsZhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang 等NeurIPS 2025 · 被引用 7 次
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset DistillationXiao Cui, Yulei Qin, Wengang Zhou, Hongsheng Li 等NeurIPS 2025 · 被引用 5 次
- DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion ModelsQichao Wang, Yunhong Lu, Hengyuan Cao, Junyi Zhang 等CVPR 2026 · 被引用 4 次
- Multimodal Distribution Matching for Vision-Language Dataset DistillationJongoh Jeong, Hoyong Kwon, Minseok Kim, Kuk-Jin YoonCVPR 2026 · 被引用 3 次
- GeoDM: Geometry-aware Distribution Matching for Dataset DistillationXuhui Li, Zhengquan Luo, Zihui Cui, Kai Zhao 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 被引用 313 次
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 被引用 307 次
相关 Paper
- Dataset Distillation via the Wasserstein MetricHaoyang Liu, Yijiang Li, Tiancheng Xing, Peiran Wang 等ICCV 2025 · 被引用 39 次
- DREAM: Efficient Dataset Distillation by Representative MatchingYanqing Liu, Jianyang Gu, Kai Wang, Zheng Zhu 等ICCV 2023 · 被引用 114 次
- Dataset Distillation of 3D Point Clouds via Distribution MatchingJae-Young Yim, Dongwook Kim, Jae-Young SimNeurIPS 2025
- Hyperbolic Dataset DistillationWenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa 等NeurIPS 2025 · 被引用 17 次
- Diversified Semantic Distribution Matching for Dataset DistillationHongcheng Li, Yucan Zhou, Xiaoyan Gu, Bo Li 等ACM MM 2024 · 被引用 10 次
