ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation
Qizhen Lan, Qing Tian
摘要
Dense visual prediction tasks, such as detection and segmentation, are crucial for time-critical applications (e.g., autonomous driving and video surveillance). While deep models achieve strong performance, their efficiency remains a challenge. Knowledge distillation (KD) is an effective model compression technique, but existing feature-based KD methods rely on static, teacher-driven feature selection, failing to adapt to the student's evolving learning state or leverage dynamic student-teacher interactions. To address these limitations, we propose Adaptive student-teacher Cooperative Attention Masking for Knowledge Distillation (ACAM-KD), which introduces two key components: (1) Student-Teacher Cross-Attention Feature Fusion (STCA ), which adaptively integrates features from both models for a more interactive distillation process, and (2) Adaptive Spatial-Channel Masking (ASCM), which dynamically generates importance masks to enhance both spatial and channel-wise feature selection. Unlike conventional KD methods, ACAM-KD adapts to the student's evolving needs throughout the entire distillation process. Extensive experiments on multiple benchmarks validate its effectiveness. For instance, on COCO2017, ACAM-KD improves object detection performance by up to 1.4 mAP over the state-of-the-art when distilling a ResNet-50 student from a ResNet101 teacher. For semantic segmentation on Cityscapes, it boosts mIoU by 3.09 over the baseline with DeepLabV3MobileNetV2 as the student model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video UnderstandingYuhao Su, Anwesa Choudhuri, Zhongpai Gao, Benjamin Planche 等CVPR 2026 · 被引用 10 次
- Momentum Memory for Knowledge Distillation in Computational Pathologyyongxin guo, Hao Lu, Onur C., Zhengjie Zhu 等CVPR 2026 · 被引用 5 次
- SepPrune: Structured Pruning for Efficient Deep Speech SeparationYuqi Li, Kai Li, Xin Yin, Zhifei Yang 等AAAI 2026 · 被引用 4 次
- Pansharpening for Thin-Cloud Contaminated Remote Sensing Images: A Unified Framework and Benchmark DatasetSongcheng Du, Yang Zou, Jiaxin Li, Mingxuan Liu 等AAAI 2026 · 被引用 3 次
- FlowTime: Towards Continuous Generative Watch Time Prediction via Flow-based Personalized PriorsHongxu Ma, Han Zhou, Chenghou Jin, Jie Zhang 等KDD 2026 · 被引用 1 次
它引用的顶会 Paper11
- RepPoints: Point Set Representation for Object DetectionZe Yang, Shaohui Liu, Han Hu, Liwei Wang 等ICCV 2019 · 被引用 1,056 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
- Channel-wise Knowledge Distillation for Dense Prediction*Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan 等ICCV 2021 · 被引用 432 次
- Focal and Global Knowledge Distillation for DetectorsZhendong Yang, Zhe Li, Xiaohu Jiang, Yuan Gong 等CVPR 2022 · 被引用 325 次
- Improve Object Detection with Feature-based Knowledge Distillation: Towards Accurate and Efficient DetectorsLinfeng Zhang, Kaisheng MaICLR 2021 · 被引用 251 次
相关 Paper
- DCSF-KD: Dynamic Channel-wise Spatial Feature Knowledge Distillation for Object DetectionTao Dai, Yang Lin, Hang Guo, Jinbao Wang 等AAAI 2025 · 被引用 7 次
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li 等CVPR 2024 · 被引用 93 次
- Masked Autoencoders Are Stronger Knowledge DistillersShanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu 等ICCV 2023 · 被引用 11 次
- Multi-Knowledge Aggregation and Transfer for Semantic SegmentationYuang Liu, Wei Zhang, Jun WangAAAI 2022 · 被引用 11 次
- Pixel-Wise Contrastive DistillationJunqiang Huang, Zichao GuoICCV 2023 · 被引用 8 次
