Does the Data Processing Inequality Reflect Practice? On the Utility of Low-Level Tasks
Roy Turgeman, Tom Tirer
Abstract
The data processing inequality is an information-theoretic principle stating that the information content of a signal cannot be increased by processing the observations. In particular, it suggests that there is no benefit in enhancing the signal or encoding it before addressing a classification problem. This assertion can be proven to be true for the case of the optimal Bayes classifier. However, in practice, it is common to perform "low-level" tasks before "high-level" downstream tasks despite the overwhelming capabilities of modern deep neural networks. In this paper, we aim to understand when and why low-level processing can be beneficial for classification. We present a comprehensive theoretical study of a binary classification setup, where we consider a classifier that is tightly connected to the optimal Bayes classifier and converges to it as the number of training samples increases. We prove that for any finite number of training samples, there exists a pre-classification processing that improves the classification accuracy. We also explore the effect of class separation, training set size, and class balance on the relative gain from this procedure. We support our theory with an empirical investigation of the theoretical setup. Finally, we conduct an empirical study where we investigate the effect of denoising and encoding on the performance of practical deep classifiers on benchmark datasets. Specifically, we vary the size and class distribution of the training set, and the noise level, and demonstrate trends that are consistent with our theoretical results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart et al.ICLR 2020 · 211 citations
- AutoBalance: Optimized Loss Functions for Imbalanced DataMingchen Li, Xuechen Zhang, Christos Thrampoulidis, Jiasi Chen et al.NeurIPS 2021 · 89 citations
- Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian MixturesYuan Cao, Quanquan Gu, Mikhail BelkinNeurIPS 2021 · 57 citations
- Dual Directed Capsule Network for Very Low Resolution Image RecognitionManeet Singh, Shruti Nagpal, Richa Singh, Mayank VatsaICCV 2019 · 56 citations
- Sliced Mutual Information: A Scalable Measure of Statistical DependenceZiv Goldfeld, Kristjan H. GreenewaldNeurIPS 2021 · 48 citations
Related papers
- Error-Bounded Correction of Noisy LabelsSongzhu Zheng, Pengxiang Wu, Aman Goswami, Mayank Goswami et al.ICML 2020 · 153 citations
- Benign Overfitting in Two-Layer ReLU Convolutional Neural Networks for XOR DataXuran Meng, Difan Zou, Yuan CaoICML 2024 · 11 citations
- A Near-Optimal Algorithm for Debiasing Trained Machine Learning ModelsIbrahim M. Alabdulmohsin, Mario LucicNeurIPS 2021 · 26 citations
- LQF: Linear Quadratic Fine-TuningAlessandro Achille, Aditya Golatkar, Avinash Ravichandran, Marzia Polito et al.CVPR 2021
- Towards Disentangling Information Paths with Coded ResNeXtApostolos Avranas, Marios KountourisNeurIPS 2022 · 1 citation
