Hiding Data Helps: On the Benefits of Masking for Sparse Coding
Muthu Chidambaram, Chenwei Wu, Yu Cheng, Rong Ge
Abstract
Sparse coding, which refers to modeling a signal as sparse linear combinations of the elements of a learned dictionary, has proven to be a successful (and interpretable) approach in applications such as signal processing, computer vision, and medical imaging. While this success has spurred much work on provable guarantees for dictionary recovery when the learned dictionary is the same size as the ground-truth dictionary, work on the setting where the learned dictionary is larger (or over-realized) with respect to the ground truth is comparatively nascent. Existing theoretical results in this setting have been constrained to the case of noise-less data. We show in this work that, in the presence of noise, minimizing the standard dictionary learning objective can fail to recover the elements of the ground-truth dictionary in the over-realized regime, regardless of the magnitude of the signal in the data-generating process. Furthermore, drawing from the growing body of work on self-supervised learning, we propose a novel masking objective for which recovering the ground-truth dictionary is in fact optimal as the signal increases for a large class of data-generating processes. We corroborate our theoretical results with experiments across several parameter regimes showing that our proposed objective also enjoys better empirical performance than the standard reconstruction objective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86529866-52ac-45b4-be2a-d7dfd35c1d70Builds on5
- Self-supervised Learning from a Multi-view PerspectiveYao-Hung Hubert Tsai, Yue Wu, Ruslan Salakhutdinov, Louis-Philippe MorencyICLR 2021 · 232 citations
- Predicting What You Already Know Helps: Provable Self-Supervised LearningJason D. Lee, Qi Lei, Nikunj Saunshi, Jiacheng ZhuoNeurIPS 2021 · 219 citations
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 119 citations
- Towards Understanding Why Mask Reconstruction Pretraining Helps in Downstream TasksJiachun Pan, Pan Zhou, Shuicheng YanICLR 2023 · 6 citations
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li et al.CVPR 2022
Related papers
- Understanding Masked Autoencoders via Hierarchical Latent Variable ModelsLingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing et al.CVPR 2023
- Global Identifiability of Overcomplete Dictionary Learning via L1 and Volume MinimizationYuchen Sun, Kejun HuangICLR 2025
- Geometric Analysis of Nonconvex Optimization Landscapes for Overcomplete LearningQing Qu, Yuexiang Zhai, Xiao Li, Yuqian Zhang et al.ICLR 2020 · 29 citations
- A Random Matrix Theory of Masked Self-Supervised LearningArie Zurich, Federica Gerace, Bruno Loureiro, Yue LuICML 2026
- Empirical Study of the Benefits of Overparameterization in Learning Latent Variable ModelsRares-Darius Buhai, Yoni Halpern, Yoon Kim, Andrej Risteski et al.ICML 2020 · 33 citations
