Sharpness-Aware Minimization Activates the Interactive Teaching's Understanding and Optimization
Mingwei Xu, Xiaofeng Cao, Ivor W. Tsang
摘要
Teaching is a potentially effective approach for understanding interactions among multiple intelligences. Previous explorations have convincingly shown that teaching presents additional opportunities for observation and demonstration within the learning model, such as data distillation and selection. However, the underlying optimization principles and convergence of interactive teaching lack theoretical analysis, and in this regard co-teaching serves as a notable prototype. In this paper, we discuss its role as a reduction of the larger loss landscape derived from Sharpness-Aware Minimization (SAM). Then, we classify it as an iterative parameter estimation process using Expectation-Maximization. The convergence of this typical interactive teaching is achieved by continuously optimizing a variational lower bound on the log marginal likelihood. This lower bound represents the expected value of the log posterior distribution of the latent variables under a scaled, factorized variational distribution. To further enhance interactive teaching’s performance, we incorporate SAM’s strong generalization information into interactive teaching, referred as Sharpness Reduction Interactive Teaching (SRIT). This integration can be viewed as a novel sequential optimization process. Finally, we validate the performance of our approach through multiple experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Part-dependent Label Noise: Towards Instance-dependent Label NoiseXiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang 等NeurIPS 2020 · 被引用 329 次
相关 Paper
- Mitigating Parameter Interference in Model Merging via Sharpness-Aware Fine-TuningYeoreum Lee, Jinwook Jung, Sungyong BaikICLR 2025
- Nonparametric Teaching for Graph Property LearnersChen Zhang, Weixin Bu, Zeyi Ren, Zhengwu Liu 等ICML 2025
- SAMO: A Lightweight Sharpness-Aware Approach for Multi-Task Optimization with Joint Global-Local PerturbationHao Ban, Gokul Ram Subramani, Kaiyi JiICCV 2025 · 被引用 3 次
- Nonparametric Teaching for Multiple LearnersChen Zhang, Xiaofeng Cao, Weiyang Liu, Ivor W. Tsang 等NeurIPS 2023 · 被引用 8 次
- Revisiting Sharpness-Aware Minimization: A More Faithful and Effective ImplementationJianlong Chen, Zhiming ZhouICLR 2026 · 被引用 1 次
