DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledge
Xin Jiang, Hao Tang, Meiqi Cao, Junyao Gao, Fei Shen, Zechao Li
摘要
Open-set fine-grained retrieval (OSFR) is a challenging task where models must generalize to unseen subcategories. Existing methods often fail this, as they embed categoryspecific semantics from closed-set training labels. Recently, diffusion transformers (DiT) have shown promise by encoding attribute-centric, generative curriculum knowledge that is agnostic to these labels. However, the vanilla DiT is not optimized for fine-grained visual discrepancies and its massive size makes deployment infeasible. To solve this, we propose DiT-Distill, a framework to first refine and then distill this knowledge. We introduce a conditional discrepancy refinement strategy to fine-tune the DiT, forcing it to focus on discrepancy-aware, attribute-centric details rather than holistic context. Subsequently, a generative curriculum distillation mechanism transfers this refined, hierarchical knowledge from multiple diffusion timesteps of the DiT into a lightweight backbone using a generative infusion module and a curriculum alignment loss. This process results in an efficient retrieval model that enables DiT-free inference. Extensive experiments show DiT-Distill achieves state-ofthe-art performance on open-set fine-grained datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- DiffusionDet: Diffusion Model for Object DetectionShoufa Chen, Peize Sun, Yibing Song, Ping LuoICCV 2023 · 被引用 715 次
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera 等NeurIPS 2023 · 被引用 371 次
- Unleashing Text-to-Image Diffusion Models for Visual PerceptionWenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu 等ICCV 2023 · 被引用 327 次
相关 Paper
- Adversarial Reconstruction Feedback for Robust Fine-Grained GeneralizationShijie Wang, Jian Shi, Haojie LiICCV 2025 · 被引用 2 次
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang 等CVPR 2023
- DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval GuidelinesXin Jiang, Hao Tang, Rui Yan, Jinhui Tang 等ACM MM 2024 · 被引用 18 次
- DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any ArchitectureQianlong Xiang, Miao Zhang, Yuzhang Shang, Jianlong Wu 等CVPR 2025
- Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World ModelXiaodong Wang, Zhirong Wu, Peixi PengAAAI 2026
