Learning Representation from Neural Fisher Kernel with Low-rank Approximation
Ruixiang Zhang, Shuangfei Zhai, Etai Littwin, Joshua M. Susskind
Abstract
In this paper, we study the representation of neural networks from the view of kernels. We first define the Neural Fisher Kernel (NFK), which is the Fisher Kernel (Jaakkola and Haussler, 1998) applied to neural networks. We show that NFK can be computed for both supervised and unsupervised learning models, which can serve as a unified tool for representation extraction. Furthermore, we show that practical NFKs exhibit low-rank structures. We then propose an efficient algorithm that computes a low rank approximation of NFK, which scales to large datasets and networks. We show that the low-rank approximation of NFKs derived from unsupervised generative models and supervised learning models gives rise to high-quality compact representations of data, achieving competitive results on a variety of machine learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive FlowsRuixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang et al.NeurIPS 2025 · 8 citations
- KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and UnlearningYinyi Luo, Zhexian Zhou, Hao Chen, Kai Qiu et al.ICLR 2026 · 4 citations
- Supervision Complexity and its Role in Knowledge DistillationHrayr Harutyunyan, Ankit Singh Rawat, Aditya Krishna Menon, Seungyeon Kim et al.ICLR 2023 · 1 citation
- The Geometry of Updates: Fisher Alignment at Vocabulary ScaleJohn SweeneyICML 2026 · 1 citation
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
Related papers
- Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsMinshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao et al.NeurIPS 2020 · 61 citations
- Deterministic Bounds and Random Estimates of Metric Tensors on NeuromanifoldsKe SunICLR 2026
- Nuclear Norm Regularization for Deep LearningChristopher Scarvelis, Justin M. SolomonNeurIPS 2024 · 17 citations
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 62 citations
- Neural Fisher Discriminant Analysis: Optimal Neural Network Embeddings in Polynomial TimeBurak Bartan, Mert PilanciICML 2022 · 3 citations
