Knowledge Consistency between Neural Networks and Beyond
Ruofan Liang, Tianlin Li, Longfei Li, Jing Wang, Quanshi Zhang
摘要
This paper aims to analyze knowledge consistency between pre-trained deep neural networks. We propose a generic definition for knowledge consistency between neural networks at different fuzziness levels. A task-agnostic method is designed to disentangle feature components, which represent the consistent knowledge, from raw intermediate-layer features of each neural network. As a generic tool, our method can be broadly used for different applications. In preliminary experiments, we have used knowledge consistency as a tool to diagnose knowledge representations of neural networks. Knowledge consistency provides new insights to explain the success of existing deep-learning techniques, such as knowledge distillation and network compression. More crucially, knowledge consistency can also be used to refine pre-trained networks and boost performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff PerspectiveHelong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou 等ICLR 2021 · 被引用 209 次
- PMET: Precise Model Editing in a TransformerXiaopeng Li, Shasha Li, Shezheng Song, Jing Yang 等AAAI 2024 · 被引用 208 次
- Grounding Representation Similarity Through Statistical TestingFrances Ding, Jean-Stanislas Denain, Jacob SteinhardtNeurIPS 2021 · 被引用 88 次
- Interpreting Multivariate Shapley Interactions in DNNsHao Zhang, Yichen Xie, Longjie Zheng, Die Zhang 等AAAI 2021 · 被引用 70 次
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li 等NeurIPS 2022 · 被引用 46 次
相关 Paper
- Interpreting and Disentangling Feature Components of Various Complexity from DNNsJie Ren, Mingjie Li, Zexu Liu, Quanshi ZhangICML 2021 · 被引用 20 次
- Explaining Knowledge Distillation by Quantifying the KnowledgeXu Cheng, Zhefan Rao, Yilan Chen, Quanshi ZhangCVPR 2020
- Show, Attend and Distill: Knowledge Distillation via Attention-based Feature MatchingMingi Ji, Byeongho Heo, Sungrae ParkAAAI 2021 · 被引用 194 次
- DEPARA: Deep Attribution Graph for Deep Knowledge TransferabilityJie Song, Yixin Chen, Jingwen Ye, Xinchao Wang 等CVPR 2020
- Lipschitz Continuity Guided Knowledge DistillationYuzhang Shang, Bin Duan, Ziliang Zong, Liqiang Nie 等ICCV 2021 · 被引用 31 次
