Autoencoders that don't overfit towards the Identity
Harald Steck
Abstract
Autoencoders (AE) aim to reproduce the output from the input. They may hence tend to overfit towards learning the identity-function between the input and output, i.e., they may predict each feature in the output from itself in the input. This is not useful, however, when AEs are used for prediction tasks in the presence of noise in the data. It may seem intuitively evident that this kind of overfitting is prevented by training a denoising AE [36] , as the dropped-out features have to be predicted from the other features. In this paper, we consider linear autoencoders, as they facilitate analytic solutions, and first show that denoising / dropout actually prevents the overfitting towards the identity-function only to the degree that it is penalized by the induced L2-norm regularization. In the main theorem of this paper, we show that the emphasized denoising AE [37] is indeed capable of completely eliminating the overfitting towards the identity-function. Our derivations reveal several new insights, including the closed-form solution of the full-rank model, as well as a new (near-)orthogonality constraint in the low-rank model. While this constraint is conceptually very different from the regularizers recently proposed in [11, 42, 14] , their resulting effects on the learned embeddings are empirically similar. Our experiments on three well-known data-sets corroborate the various theoretical insights derived in this paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Towards a Better Understanding of Linear Models for RecommendationRuoming Jin, Dong Li, Jing Gao, Zhi Liu et al.KDD 2021 · 21 citations
- Understanding Representation Dynamics of Diffusion Models via Low-Dimensional ModelingXiao Li, Zekai Zhang, Xiang Li, Siyi Chen et al.NeurIPS 2025 · 19 citations
- Collaborative Residual Metric LearningTianjun Wei, Jianghong Ma, Tommy W. S. ChowSIGIR 2023 · 5 citations
- Fine-tuning Partition-aware Item Similarities for Efficient and Scalable RecommendationTianjun Wei, Jianghong Ma, Tommy W. S. ChowWWW 2023 · 5 citations
- Fast Offline Policy Optimization for Large Scale RecommendationOtmane Sakhi, David Rohde, Alexandre GilotteAAAI 2023 · 5 citations
Related papers
- It's Enough: Relaxing Diagonal Constraints in Linear Autoencoders for RecommendationJaewan Moon, Hye-young Kim, Jongwuk LeeSIGIR 2023 · 3 citations
- Mitigating Memorization of Noisy Labels via Regularization between RepresentationsHao Cheng, Zhaowei Zhu, Xing Sun, Yang LiuICLR 2023 · 8 citations
- High-dimensional Asymptotics of Denoising AutoencodersHugo Cui, Lenka ZdeborováNeurIPS 2023 · 26 citations
- Autoencoder Image Interpolation by Shaping the Latent SpaceAlon Oring, Zohar Yakhini, Yacov Hel-OrICML 2021 · 41 citations
- Regularized Autoencoders for Isometric Representation LearningYonghyeon Lee, Sangwoong Yoon, Minjun Son, Frank Chongwoo ParkICLR 2022 · 46 citations
