SubspacePath Pruner: Inference-time Pruning via Probe-based Representation–Parameter Coupling
Zhiren Gong, Yikun Hou, Fan Wu, CHE WANG, Fuyao Zhang, Tiantong Wu, Yurong Hao, Jiaming Zhang, Yiyang Duan, Tiantong Wang, Fei Huang, Chau Yuen, Wei Yang Bryan Lim
摘要
Large-scale dedicated application of LLMs in diverse scenarios increasingly demands specialized model inference behavior under strict constraints of accuracy, latency, and memory. However, the heterogeneous and long-tailed nature of real-world specialized scenarios makes it difficult to obtain training data and optimize models. We study a practical inference-time specialization setting: given an LLM base, we compile a reusable, budget-bounded pathway/subnetwork within a specific scenario. Our approach is motivated by an empirical coupling phenomenon: input scenario sets aligned with similar representation subspaces (e.g., domain) in embedding space tend to activate a consistent and sparse set of internal reasoning pathways in model parameter space. To build the bridge between them, we propose probe-based SUBSPACEPATH PRUNER with two core components: (1) Domain-Basis Synthesis (DBS) constructs a quasi-orthogonal basis of domain axes in embedding space, serving as a stable coordinate system. (2) Probe-based Scenario Pruning (PSP) uses efficient layer-wise linear probes to estimate axis alignment and compute budgeted head-wise pathways for a specific scenario. Experiments on LLaMA-2-13B show 29.3 average Recall on cross-domain tests (vs. 24.7 dense) and 21.6 on cross-dataset tests (vs. 25.5 dense) with 1.27× speedup at 30% pruning ratio. We have code available in GitHub repository and project page in SubspacePath Pruner.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh 等ICLR 2025 · 被引用 10 次
相关 Paper
- From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM AlignmentHao Chen, Qi Zhang, Liyao Li, Zhanming Shen 等ICML 2026
- SEAP: Sparse Expert Activation Pruning Unlocks the Brainpower of Large Language ModelsXun Liang, Hanyu Wang, Huayi Lai, Simin Niu 等AAAI 2026
- TrimLLM: Progressive Layer Dropping for Domain-Specific LLMsLanxiang Hu, Tajana Rosing, Hao ZhangACL 2025
- Persona-Pruner: Sculpting Lightweight Models for Role-PlayingJinsu Kim, Jihoon Tack, Noah Lee, Jongheon JeongICML 2026
- Computation and Memory-Efficient Model Compression with Gradient ReweightingZhiwei Li, Yuesen Liao, Binrui Wu, Yuquan Zhou 等NeurIPS 2025
