Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies
Itai Gat, Idan Schwartz, Alexander G. Schwing, Tamir Hazan
摘要
Many recent datasets contain a variety of different data modalities, for instance, image, question, and answer data in visual question answering (VQA). When training deep net classifiers on those multi-modal datasets, the modalities get exploited at different scales, i.e., some modalities can more easily contribute to the classification results than others. This is suboptimal because the classifier is inherently biased towards a subset of the modalities. To alleviate this shortcoming, we propose a novel regularization term based on the functional entropy. Intuitively, this term encourages to balance the contribution of each modality to the classification result. However, regularization with the functional entropy is challenging. To address this, we develop a method based on the log-Sobolev inequality, which bounds the functional entropy with the functional-Fisher-information. Intuitively, this maximizes the amount of information that the modalities contribute. On the two challenging multi-modal datasets VQA-CPv2 and SocialIQ, we obtain state-of-the-art results while more uniformly exploiting the modalities. In addition, we demonstrate the efficacy of our method on Colored MNIST. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen 等NeurIPS 2021 · 被引用 404 次
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang 等ICML 2022 · 被引用 168 次
- ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic ArithmeticYoad Tewel, Yoav Shalev, Idan Schwartz, Lior WolfCVPR 2022 · 被引用 129 次
- Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural NetworksNan Wu, Stanislaw Jastrzebski, Kyunghyun Cho, Krzysztof J. GerasICML 2022 · 被引用 124 次
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu 等NeurIPS 2021 · 被引用 102 次
它引用的顶会 Paper3
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin 等ICCV 2019 · 被引用 288 次
- Sequential Latent Spaces for Modeling the Intention During Diverse Image CaptioningJyoti Aneja, Harsh Agrawal, Dhruv Batra, Alexander G. SchwingICCV 2019 · 被引用 71 次
- What Makes Training Multi-Modal Classification Networks Hard?Weiyao Wang, Du Tran, Matt FeiszliCVPR 2020
相关 Paper
- Reducing Unimodal Bias in Multi-Modal Semantic Segmentation With Multi-Scale Functional Entropy RegularizationXu Zheng, Yuanhuiyi Lyu, Lutao Jiang, Danda Pani Paudel 等ICCV 2025 · 被引用 2 次
- A Functional Information Perspective on Model InterpretationItai Gat, Nitay Calderon, Roi Reichart, Tamir HazanICML 2022 · 被引用 6 次
- Balancing Multimodal Training Through Game-Theoretic RegularizationKonstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang 等NeurIPS 2025 · 被引用 17 次
- Perceptual Score: What Data Modalities Does Your Model Perceive?Itai Gat, Idan Schwartz, Alexander G. SchwingNeurIPS 2021 · 被引用 56 次
- Cross-Modal Feature Distribution Calibration for Few-Shot Visual Question AnsweringJing Zhang, Xiaoqiang Liu, Mingzhe Chen, Zhe WangAAAI 2024 · 被引用 2 次
