Adaptive Mixing of Auxiliary Losses in Supervised Learning
Durga Sivasubramanian, Ayush Maheshwari, Prathosh AP, Pradeep Shenoy, Ganesh Ramakrishnan
摘要
In many supervised learning scenarios, auxiliary losses are used in order to introduce additional information or constraints into the supervised learning objective. For instance, knowledge distillation aims to mimic outputs of a powerful teacher model; similarly, in rule-based approaches, weak labeling information is provided by labeling functions which may be noisy rule-based approximations to true labels. We tackle the problem of learning to combine these losses in a principled manner. Our proposal, AMAL, uses a bi-level optimization criterion on validation data to learn optimal mixing weights, at an instance-level, over the training data. We describe a meta-learning approach towards solving this bi-level objective, and show how it can be applied to different scenarios in supervised learning. Experiments in a number of knowledge distillation and rule denoising domains show that AMAL provides noticeable gains over competitive baselines in those domains. We empirically analyze our method and share insights into the mechanisms through which it provides performance gains. The code for AMAL is at: https://github.com/durgas16/AMAL.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Instance-Conditional Timescales of Decay for Non-Stationary LearningNishant Jain, Pradeep ShenoyAAAI 2024 · 被引用 8 次
- Learning model uncertainty as variance-minimizing instance weightsNishant Jain, Karthikeyan Shanmugam, Pradeep ShenoyICLR 2024 · 被引用 7 次
- Auxiliary Gene Learning: Spatial Gene Expression Estimation by Auxiliary Gene SelectionKaito Shiku, Kazuya Nishimura, Shinnosuke Matsuo, Yasuhiro Kojima 等AAAI 2026
它引用的顶会 Paper11
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Densely Guided Knowledge Distillation using Multiple Teacher AssistantsWonchul Son, Jaemin Na, Junyong Choi, Wonjun HwangICCV 2021 · 被引用 158 次
- A statistical perspective on distillationAditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim 等ICML 2021 · 被引用 97 次
相关 Paper
- Module-Aware Optimization for Auxiliary LearningHong Chen, Xin Wang, Yue Liu, Yuwei Zhou 等NeurIPS 2022 · 被引用 11 次
- Auxiliary Learning by Implicit DifferentiationAviv Navon, Idan Achituve, Haggai Maron, Gal Chechik 等ICLR 2021 · 被引用 72 次
- MetaASSIST: Robust Dialogue State Tracking with Meta LearningFanghua Ye, Xi Wang, Jie Huang, Shenghui Li 等EMNLP 2022 · 被引用 10 次
- A Nested Bi-level Optimization Framework for Robust Few Shot LearningKrishnaTeja Killamsetty, Changbin Li, Chen Zhao, Feng Chen 等AAAI 2022 · 被引用 12 次
- Partial Multi-Label Learning with Meta DisambiguationMing-Kun Xie, Feng Sun, Sheng-Jun HuangKDD 2021 · 被引用 25 次
