Interpretable Self-Supervised Learning via Representer Landmarks and Nyström Approximation
Maedeh Zarvandi, Michael Timothy, Theresa Wasserer, Debarghya Ghoshdastidar
Abstract
Self-supervised learning (SSL) learns representations from massive unlabeled data, yet the resulting models typically operate as black boxes, necessitating domain-specific explanations. We introduce KREPES, a unified framework to analytically interpret the learned representations of SSL objectives, including SimCLR, BYOL, and VICReg. By bridging empirical neural tangent kernel approximations of neural networks with the Representer Theorem for kernels, we express the learned latent space directly via "Representer Landmarks", which are the representations of influential unlabeled training examples. We introduce novel metrics, "Sample-Specific Influence Score", "Concept-Conditioned Influence Score" and "Feature Alignment Gap", to quantify the transparency of the learned representations. KREPES enables direct audit of the latent space without supervision, for example, revealing an algorithmic bias in the Adult-1M dataset where SSL uses demographic proxies for income. Finally, to ensure scalability to benchmarks with 1M+ samples (ImageNet-1K, Adult-1M), KREPES introduces a novel Nyström approximation-based analytical inference framework for SSL objectives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e07f018b-a2d6-416f-86f2-179c40d28f31Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
Related papers
- Contrastive and Non-Contrastive Self-Supervised Learning Recover Global and Local Spectral Embedding MethodsRandall Balestriero, Yann LeCunNeurIPS 2022 · 189 citations
- Measuring Self-Supervised Representation Quality for Downstream Classification Using Discriminative FeaturesNeha Mukund Kalibhat, Kanika Narang, Hamed Firooz, Maziar Sanjabi et al.AAAI 2024 · 12 citations
- Non-parametric Representation Learning with KernelsPascal Mattia Esser, Maximilian Fleissner, Debarghya GhoshdastidarAAAI 2024 · 11 citations
- On the duality between contrastive and non-contrastive self-supervised learningQuentin Garrido, Yubei Chen, Adrien Bardes, Laurent Najman et al.ICLR 2023 · 24 citations
- On the Stepwise Nature of Self-Supervised LearningJames B. Simon, Maksis Knutins, Liu Ziyin, Daniel Geisz et al.ICML 2023 · 45 citations
