PolyLoss: A Polynomial Expansion Perspective of Classification Loss Functions
Zhaoqi Leng, Mingxing Tan, Chenxi Liu, Ekin Dogus Cubuk, Jay Shi, Shuyang Cheng, Dragomir Anguelov
摘要
Cross-entropy loss and focal loss are the most common choices when training deep neural networks for classification problems. Generally speaking, however, a good loss function can take on much more flexible forms, and should be tailored for different tasks and datasets. Motivated by how functions can be approximated via Taylor expansion, we propose a simple framework, named PolyLoss, to view and design loss functions as a linear combination of polynomial functions. Our PolyLoss allows the importance of different polynomial bases to be easily adjusted depending on the targeting tasks and datasets, while naturally subsuming the aforementioned cross-entropy loss and focal loss as special cases. Extensive experimental results show that the optimal choice within the PolyLoss is indeed dependent on the task and dataset. Simply by introducing one extra hyperparameter and adding one line of code, our Poly-1 formulation outperforms the cross-entropy loss and focal loss on 2D image classification, instance segmentation, object detection, and 3D object detection tasks, sometimes by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai 等NeurIPS 2022 · 被引用 1,270 次
- Test Time Adaptation via Conjugate Pseudo-labelsSachin Goyal, Mingjie Sun, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 被引用 152 次
- Learning Sample Difficulty from Pre-trained Models for Reliable PredictionPeng Cui, Dan Zhang, Zhijie Deng, Yinpeng Dong 等NeurIPS 2023 · 被引用 21 次
- GradTree: Learning Axis-Aligned Decision Trees with Gradient DescentSascha Marton, Stefan Lüdtke, Christian Bartelt, Heiner StuckenschmidtAAAI 2024 · 被引用 17 次
- Temporal Label Smoothing for Early Event PredictionHugo Yèche, Alizée Pace, Gunnar Rätsch, Rita KuznetsovaICML 2023 · 被引用 16 次
它引用的顶会 Paper13
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui 等NeurIPS 2020 · 被引用 755 次
- Revisiting ResNets: Improved Training and Scaling StrategiesIrwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk 等NeurIPS 2021 · 被引用 378 次
- Can gradient clipping mitigate label noise?Aditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Sanjiv KumarICLR 2020 · 被引用 163 次
相关 Paper
- Are All Losses Created Equal: A Neural Collapse PerspectiveJinxin Zhou, Chong You, Xiao Li, Kangning Liu 等NeurIPS 2022 · 被引用 93 次
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz 等NeurIPS 2020 · 被引用 674 次
- Loss Function Discovery for Object Detection via Convergence-Simulation Driven SearchPeidong Liu, Gengwei Zhang, Bochao Wang, Hang Xu 等ICLR 2021 · 被引用 30 次
- Searching for Robustness: Loss Learning for Noisy Classification TasksBoyan Gao, Henry Gouk, Timothy M. HospedalesICCV 2021 · 被引用 22 次
- Anchor Loss: Modulating Loss Scale Based on Prediction DifficultySerim Ryou, Seong-Gyun Jeong, Pietro PeronaICCV 2019 · 被引用 46 次
