Diverse Dictionary Learning
Yujia Zheng, Zijian Li, Shunxing Fan, Andrew Gordon Wilson, Kun Zhang
Abstract
Given only observational data , where both the latent variables and the generating process are unknown, recovering is ill-posed without additional assumptions. Existing methods often assume linearity or rely on auxiliary supervision and functional constraints. However, such assumptions are rarely verifiable in practice, and most theoretical guarantees break down under even mild violations, leaving uncertainty about how to reliably understand the hidden world. To make identifiability actionable in the real-world scenarios, we take a complementary view: in the general settings where full identifiability is unattainable, what can still be recovered with guarantees, and what biases could be universally adopted? We introduce the problem of diverse dictionary learning to formalize this view. Specifically, we show that intersections, complements, and symmetric differences of latent variables linked to arbitrary observations, along with the latent-to-observed dependency structure, are still identifiable up to appropriate indeterminacies even without strong assumptions. These set-theoretic results can be composed using set algebra to construct structured and essential views of the hidden world, such as genus-differentia definitions. When sufficient structural diversity is present, they further imply full identifiability of all latent variables. Notably, all identifiability benefits follow from a simple inductive bias during estimation that can be readily integrated into most models. We validate the theory and demonstrate the benefits of the bias on both synthetic and real-world data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c8896d8-b855-4e19-adfa-8e347ca60d13Builds on32
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel et al.NeurIPS 2021 · 421 citations
- Weakly supervised causal representation learningJohann Brehmer, Pim de Haan, Phillip Lippe, Taco S. CohenNeurIPS 2022 · 196 citations
- Disentanglement by Nonlinear ICA with General Incompressible-flow Networks (GIN)Peter Sorrenson, Carsten Rother, Ullrich KötheICLR 2020 · 132 citations
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele et al.NeurIPS 2023 · 127 citations
Related papers
- Synergy Between Sufficient Changes and Sparse Mixing Procedure for Disentangled Representation LearningZijian Li, Shunxing Fan, Yujia Zheng, Ignavier Ng et al.ICLR 2025
- Properties from mechanisms: an equivariance perspective on identifiable representation learningKartik Ahuja, Jason S. Hartford, Yoshua BengioICLR 2022 · 40 citations
- Differentiable Structure Learning and Causal Discovery for General Binary DataChang Deng, Bryon AragamNeurIPS 2025
- A Bayesian Nonparametric Framework For Learning Disentangled RepresentationsVaishnavi Patil, Siddhi Patil, Matthew Evanusa, Amit Kumar Kundu et al.ICLR 2026
- On Finite-Sample Identifiability of Contrastive Learning-Based Nonlinear Independent Component AnalysisQi Lyu, Xiao FuICML 2022 · 6 citations
