The Geometry of Updates: Fisher Alignment at Vocabulary Scale
John Sweeney
摘要
Training-free source selection for LLM families with shared vocabularies arises in scientific string domains such as SMILES, protein, and genomic sequences, where candidate corpora share a tokenizer but differ in prediction targets. This creates an activation-dark regime: representation-similarity metrics can be uninformative without assumptions about label-conditioned error geometry, while classical update-geometry metrics are computationally prohibitive at vocabulary scale. We show that, in a shared-output head setting, representation metrics (e.g., CKA) are non-identifiable for transfer; models can share identical representations yet have orthogonal head updates. The key identity is that head Fisher alignment is exactly a cosine between kernel mean embeddings in the joint activation-error space, exposing activation, error, and coupling factors rather than requiring a materialized Fisher matrix. FisherSketch estimates this cosine directly in a single streaming pass, making K=128,256 head Fisher alignment practical with a 16 KB task signature (m=4096) and a 192 KB per-task streaming state–small enough to store next to a model hash, but encoding transfer-relevant update structure. Beyond source selection, the same signatures and marginals provide a diagnostic instrument for studying whether LLM task similarity is driven by activations, errors, or their coupling; shared-parameter and internal-layer validations, together with Llama-3.1-8B verbalizer-shift experiments, show that FisherSketch remains informative when activation similarity cannot distinguish tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas 等ICML 2020 · 被引用 651 次
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran 等ICCV 2019 · 被引用 359 次
- Efficiently Identifying Task Groupings for Multi-Task LearningChris Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu 等NeurIPS 2021 · 被引用 352 次
相关 Paper
- Beyond Variance: Knowledge-Aware LLM Compression via Fisher-Aligned Subspace DiagnosticsIbne Farabi Shihab, Sanjeda Akter, Anuj SharmaACL 2026 · 被引用 4 次
- Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language ModelsPatrick Leask, Neel Nanda, Noura Al MoubayedICML 2025
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMsXiaoyu Xu, Xiang Yue, Yang Liu, Qingqing Ye 等ICML 2026 · 被引用 36 次
- From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMsLouis Schiekiera, Max Zimmer, Christophe Roux, Sebastian Pokutta 等ICML 2026 · 被引用 2 次
- Leveraging Generative Models for Unsupervised Alignment of Neural Time Series DataAyesha Vermani, Il Memming Park, Josue NassarICLR 2024 · 被引用 5 次
