Learning Representation from Neural Fisher Kernel with Low-rank Approximation
Ruixiang Zhang, Shuangfei Zhai, Etai Littwin, Joshua M. Susskind
摘要
In this paper, we study the representation of neural networks from the view of kernels. We first define the Neural Fisher Kernel (NFK), which is the Fisher Kernel (Jaakkola and Haussler, 1998) applied to neural networks. We show that NFK can be computed for both supervised and unsupervised learning models, which can serve as a unified tool for representation extraction. Furthermore, we show that practical NFKs exhibit low-rank structures. We then propose an efficient algorithm that computes a low rank approximation of NFK, which scales to large datasets and networks. We show that the low-rank approximation of NFKs derived from unsupervised generative models and supervised learning models gives rise to high-quality compact representations of data, achieving competitive results on a variety of machine learning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive FlowsRuixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang 等NeurIPS 2025 · 被引用 8 次
- KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and UnlearningYinyi Luo, Zhexian Zhou, Hao Chen, Kai Qiu 等ICLR 2026 · 被引用 4 次
- Supervision Complexity and its Role in Knowledge DistillationHrayr Harutyunyan, Ankit Singh Rawat, Aditya Krishna Menon, Seungyeon Kim 等ICLR 2023 · 被引用 1 次
- The Geometry of Updates: Fisher Alignment at Vocabulary ScaleJohn SweeneyICML 2026 · 被引用 1 次
它引用的顶会 Paper16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
相关 Paper
- Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsMinshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao 等NeurIPS 2020 · 被引用 61 次
- Deterministic Bounds and Random Estimates of Metric Tensors on NeuromanifoldsKe SunICLR 2026
- Nuclear Norm Regularization for Deep LearningChristopher Scarvelis, Justin M. SolomonNeurIPS 2024 · 被引用 17 次
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 被引用 62 次
- Neural Fisher Discriminant Analysis: Optimal Neural Network Embeddings in Polynomial TimeBurak Bartan, Mert PilanciICML 2022 · 被引用 3 次
