HieRD: Hierarchical Relational Distillation for Vision-Language Embedding Models
Vinh Le, Nguyen Dang, Tu Vu, Linh Van, Duc Nguyen, Trung Le
Abstract
Knowledge distillation is crucial for compressing large Vision–Language Models (VLMs) into efficient architectures. While prior VLM research has primarily focused on reasoning tasks like visual question answering, multimodal embedding learning, a key component for large-scale retrieval, has received comparatively less attention. Existing distillation methods typically align static global representations, overlooking hierarchical feature structure and fine-grained cross-modal interactions. This leads to a structural gap where student models fail to inherit object-level semantics and spatial relationships from teachers. To address this limitation, we propose HieRD , a Hierarchical Representation Distillation framework that preserves hierarchical structure within and across modalities throughout the distillation process by leveraging clustered visual tokens and multi-granular alignment with phrase-level text. Experimental results on multimodal embedding and downstream tasks show that HieRD consistently outperforms strong baselines, reflecting the effectiveness of its fine-grained semantic and spatial modeling, while enabling compact and efficient embedding models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0aa758d9-2f87-44ff-a354-673566dc0e39Builds on12
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat et al.NeurIPS 2021 · 863 citations
- Knowledge Fusion of Large Language ModelsFanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan et al.ICLR 2024 · 113 citations
- Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch MiningRaghuveer Thirukovalluru, Rui Meng, Ye Liu, Karthikeyan K et al.NeurIPS 2025 · 30 citations
- Token-Level Self-Play with Importance-Aware Guidance for Large Language ModelsTue Le, Hoang Tran Vuong, Quyen Tran, Linh Van Ngo et al.NeurIPS 2025 · 5 citations
Related papers
- KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsDa Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao et al.ACM MM 2025 · 10 citations
- EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision TokensZe Feng, Sen Yang, Boqiang Duan, Wankou Yang et al.AAAI 2026
- Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Large Model EnhancementQianhan Feng, Wenshuo Li, Tong Lin, Xinghao ChenCVPR 2025
- Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token InteractionsLin Chen, zhaoxiaoke, Kun Ding, Weiwei Feng et al.ICML 2026 · 4 citations
- Masking Teacher and Reinforcing Student for Distilling Vision-Language ModelsByung-Kwan Lee, Yu-Chiang Frank Wang, Ryo HachiumaCVPR 2026 · 7 citations
