Shortcut-Resistant CAM Distillation for Long-Tailed Recognition
Wenhai Wan, Teng Zhang, Shao-Yuan Li, Xinrui Wang, Qiang-Sheng Hua, Songcan Chen
Abstract
Real-world datasets often follow a long-tailed distribution, making generalization to tail classes difficult. We revisit this problem through the lens of shortcut learning, where models prefer the easiest predictive cues (e.g., background or textures) over object-centric semantics, especially under scarce and biased supervision. We find that this tendency is amplified for tail classes: limited examples often share similar contexts, making non-semantic signals highly correlated and thus tempting shortcuts, whereas head classes with diverse appearances and environments encourage more stable object-focused representations. Motivated by this observation, we propose Shortcut-Resistant CAM Distillation (SRCD), a plug-and-play framework that transfers object-focused explanations from head to tail classes. SRCD operates in the Class Activation Map (CAM) space, where a CAM provides a class-specific spatial evidence map for a prediction. SRCD aggregates CAMs from a small set of head-class candidates into a shortcut-resistant teacher using an energy-model weighting based on coherence and concentration, and distills it to the tail-class CAM. We provide a theoretical analysis that quantifies shortcut reliance as shortcut-region evidence mass in CAM space and shows that SRCD suppresses tail shortcuts. Extensive experiments on long-tailed benchmarks consistently improve strong baselines. The code is available at https://github.com/Haifeng3/SRCD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 928de1c8-ff2f-406d-9526-11680c209398Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
Related papers
- SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object DetectionHao Vo, Khoa Vo, Thinh Phan, Ngo Xuan Cuong et al.CVPR 2026 · 1 citation
- Distilling Balanced Knowledge from a Biased TeacherSeonghak KimCVPR 2026 · 1 citation
- Distilling Virtual Examples for Long-tailed RecognitionYin-Yin He, Jianxin Wu, Xiu-Shen WeiICCV 2021 · 129 citations
- Decoupled Contrastive Learning for Long-Tailed RecognitionShiyu Xuan, Shiliang ZhangAAAI 2024 · 29 citations
- SFC: Shared Feature Calibration in Weakly Supervised Semantic SegmentationXinqiao Zhao, Feilong Tang, Xiaoyang Wang, Jimin XiaoAAAI 2024 · 66 citations
