Does the Data Processing Inequality Reflect Practice? On the Utility of Low-Level Tasks
Roy Turgeman, Tom Tirer
摘要
The data processing inequality is an information-theoretic principle stating that the information content of a signal cannot be increased by processing the observations. In particular, it suggests that there is no benefit in enhancing the signal or encoding it before addressing a classification problem. This assertion can be proven to be true for the case of the optimal Bayes classifier. However, in practice, it is common to perform "low-level" tasks before "high-level" downstream tasks despite the overwhelming capabilities of modern deep neural networks. In this paper, we aim to understand when and why low-level processing can be beneficial for classification. We present a comprehensive theoretical study of a binary classification setup, where we consider a classifier that is tightly connected to the optimal Bayes classifier and converges to it as the number of training samples increases. We prove that for any finite number of training samples, there exists a pre-classification processing that improves the classification accuracy. We also explore the effect of class separation, training set size, and class balance on the relative gain from this procedure. We support our theory with an empirical investigation of the theoretical setup. Finally, we conduct an empirical study where we investigate the effect of denoising and encoding on the performance of practical deep classifiers on benchmark datasets. Specifically, we vary the size and class distribution of the training set, and the noise level, and demonstrate trends that are consistent with our theoretical results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart 等ICLR 2020 · 被引用 211 次
- AutoBalance: Optimized Loss Functions for Imbalanced DataMingchen Li, Xuechen Zhang, Christos Thrampoulidis, Jiasi Chen 等NeurIPS 2021 · 被引用 89 次
- Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian MixturesYuan Cao, Quanquan Gu, Mikhail BelkinNeurIPS 2021 · 被引用 57 次
- Dual Directed Capsule Network for Very Low Resolution Image RecognitionManeet Singh, Shruti Nagpal, Richa Singh, Mayank VatsaICCV 2019 · 被引用 56 次
- Sliced Mutual Information: A Scalable Measure of Statistical DependenceZiv Goldfeld, Kristjan H. GreenewaldNeurIPS 2021 · 被引用 48 次
相关 Paper
- Error-Bounded Correction of Noisy LabelsSongzhu Zheng, Pengxiang Wu, Aman Goswami, Mayank Goswami 等ICML 2020 · 被引用 153 次
- Benign Overfitting in Two-Layer ReLU Convolutional Neural Networks for XOR DataXuran Meng, Difan Zou, Yuan CaoICML 2024 · 被引用 11 次
- A Near-Optimal Algorithm for Debiasing Trained Machine Learning ModelsIbrahim M. Alabdulmohsin, Mario LucicNeurIPS 2021 · 被引用 26 次
- LQF: Linear Quadratic Fine-TuningAlessandro Achille, Aditya Golatkar, Avinash Ravichandran, Marzia Polito 等CVPR 2021
- Towards Disentangling Information Paths with Coded ResNeXtApostolos Avranas, Marios KountourisNeurIPS 2022 · 被引用 1 次
