Mitigating Structural Knowledge Collapse in Domain-Specific LLMs via Morpheme-Aware KV-Aggregation
Yuxuan Si, Zheqi Lv, Chengxi Zang, Zhengyu Chen, Fei Wu
摘要
. Abstract Standard tokenizers over-fragment domain terms, disrupting morpheme semantics. We characterize this representational misalignment as Structural Knowledge Collapse (SKC), where attention mechanisms fail to reconstruct coherent concepts from fragmented inputs. While existing input-centric solutions like vocabulary expansion address this, they necessitate expensive embedding retraining and neglect internal attention compositionality. To this end, we introduce Morpheme-aware KV-aggregation Attention (MorphKA), a lightweight adapter that dynamically consolidates fragments without tokenizer changes. Bypassing tokenizer retraining, MorphKA employs a dual-phase strategy—Input-Level Mor-pheme Aggregation (IMA) and Context-Aware KV-Aggregation (AMRF)—to stabilize mor-pheme spans and synthesize higher-order concepts. Experiments on medical and legal benchmarks show MorphKA outperforms vocabulary adaptation baselines by 3.2–4.6%, reaching 7.9% on high-fragmentation terms. More-over, MorphKA reduces catastrophic interference on general capabilities by 18–22% with ∼ 80% fewer parameters than embedding re-training approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li 等CVPR 2022 · 被引用 364 次
- Continual Pre-training of Language ModelsZixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi 等ICLR 2023 · 被引用 15 次
相关 Paper
- Dialogue Without Limits: Constant-Sized KV Caches for Extended Response in LLMsRavi Ghadia, Avinash Kumar, Gaurav Jain, Prashant J. Nair 等ICML 2025
- RaSE-KGC: A Relation-Aware Segment Encoding Approach for Knowledge Graph CompletionChenxiao Lin, Ye Luo, Kunhong Liu, Qingqiang WuICDE 2026
- Knowledge-driven Augmentation and Retrieval for Integrative Temporal AdaptationWeisi Liu, Guangzeng Han, Xiaolei HuangACL 2026 · 被引用 1 次
- Semi-Supervised Knowledge Amalgamation for Sequence ClassificationJidapa Thadajarassiri, Thomas Hartvigsen, Xiangnan Kong, Elke A. RundensteinerAAAI 2021 · 被引用 14 次
- FiTs: Fine-Grained Two-Stage Training for Knowledge-Aware Question AnsweringQichen Ye, Bowen Cao, Nuo Chen, Weiyuan Xu 等AAAI 2023 · 被引用 23 次
