Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression
Hyoungseob Park, Lipeng Ke, Pritish Mohapatra, Huajun Ying, sankar venkataraman, Alex Wong
摘要
We propose a method for audio-visual knowledge distillation. Existing methods typically distill a student model from the latent embeddings or outputs of a teacher. The former requires matching feature dimensions, if not the same architecture, between teacher and student models while the latter supports any teacher-student pairing, but tends to be less performant. Unlike them, we do not explicitly distill from latent embeddings or outputs, but the pairwise relationships between embeddings across samples for each modality; this is realized as a kernel, which is the crux of our method, "Kernelized Token Distillation (KTD)". Specifically, we tokenize and embed the input for a given modality, and compute the Gram matrix across tokens, from which we distill. As audio and visual modalities afford different information for a task, we adaptively modulate distillation by measuring the entropy of each modality, leading to an Entropy-Monitored Kernelized Token Distillation (EM-KTD) scheme. Our method allows for flexibility in complexity of kernel function to model relationships across tokens, which are selectively distilled to ensure high-fidelity supervision for the student. We evaluate EM-KTD on VGGSound and AVS-Bench, where we use 94% fewer parameters than the teacher while preserving 96.9% in performance for audio-visual event recognition and 96.5% on audio-visual segmentation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
相关 Paper
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep 等EMNLP 2025
- C2KD: Bridging the Modality Gap for Cross-Modal Knowledge DistillationFushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang 等CVPR 2024
- XKD: Cross-Modal Knowledge Distillation with Domain Alignment for Video Representation LearningPritam Sarkar, Ali EtemadAAAI 2024 · 被引用 45 次
- Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and AlgorithmT. K Tran, Duc Chu Anh, Quang Hung Pham, Phi Le Nguyen 等ICML 2026
- Video Event Extraction with Multi-View Interaction Knowledge DistillationKaiwen Wei, Runyan Du, Li Jin, Jian Liu 等AAAI 2024 · 被引用 5 次
