Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation
Muquan Li, Hang Gou, Yingyi Ma, Rongzheng Wang, Ke Qin, Tao He
Abstract
Decoupled dataset distillation (DD) compresses large corpora into a few synthetic images by matching a frozen teacher's statistics. However, current residual-matching pipelines rely on static real patches, creating a fit-complexity gap and a pull-to-anchor effect that reduce intra-class diversity and hurt generalization. To address these issues, we introduce RETA -- a Retrieval and Topology Alignment framework for decoupled DD. First, Dynamic Retrieval Connection (DRC) selects a real patch from a prebuilt pool by minimizing a fit-complexity score in teacher feature space; the chosen patch is injected via a residual connection to tighten feature fit while controlling injected complexity. Second, Persistent Topology Alignment (PTA) regularizes synthesis with persistent homology: we build a mutual k-NN feature graph, compute persistence images of components and loops, and penalize topology discrepancies between real and synthetic sets, mitigating pull-to-anchor effect. Across CIFAR-100, Tiny-ImageNet, ImageNet-1K, and multiple ImageNet subsets, RETA consistently outperforms various baselines under comparable time and memory, especially reaching 64.3% top-1 accuracy on ImageNet-1K with ResNet-18 at 50 images per class, +3.1% over the best prior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d3a1df7-7c84-4f6c-beea-90134e20badfCited by top-tier papers4
- Mind Your Margin and Boundary: Are Your Distilled Datasets Truly Robust?Muquan Li, Yingyi Ma, Yihong Huang, Hang Gou et al.ICML 2026
- Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow MatchingXin Hu, Ke Qin, Wen Yin, Yuan-Fang Li et al.CVPR 2026
- TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D TilingHongyaoxing Gu, Xinzhe Chen, LIJUAN HU, Liu fangfangICML 2026
- PADA-Coder: Improving Plan-Following Code Generation via Perturbation-Verified Attention Distillation and Dynamic AlignmentYihong Huang, KE QIN, Rongzheng Wang, Muquan Li et al.ICML 2026
Builds on26
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun et al.ICML 2022 · 234 citations
- Dataset Distillation by Matching Training TrajectoriesGeorge Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros et al.CVPR 2022 · 198 citations
- Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New PerspectiveZeyuan Yin, Eric P. Xing, Zhiqiang ShenNeurIPS 2023 · 180 citations
Related papers
- Topology-aware Knowledge Preservation for Class-Incremental LearningHan Zang, Yongfeng Dong, Linhao Li, Liang Yang et al.AAAI 2026
- Do Topological Characteristics Help in Knowledge Distillation?Jungeun Kim, Junwon You, Dongjin Lee, Ha Young Kim et al.ICML 2024 · 11 citations
- TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionFengli Ran, Xiao Pu, Bo Liu, Xiuli Bi et al.AAAI 2026
- Diversity-Enhanced Distribution Alignment for Dataset DistillationHongcheng Li, Yucan Zhou, Xiaoyan Gu, Bo Li et al.ICCV 2025 · 1 citation
- DELT: A Simple Diversity-driven EarlyLate Training for Dataset DistillationZhiqiang Shen, Ammar Sherif, Zeyuan Yin, Shitong ShaoCVPR 2025
