Mitigating Structural Knowledge Collapse in Domain-Specific LLMs via Morpheme-Aware KV-Aggregation
Yuxuan Si, Zheqi Lv, Chengxi Zang, Zhengyu Chen, Fei Wu
Abstract
. Abstract Standard tokenizers over-fragment domain terms, disrupting morpheme semantics. We characterize this representational misalignment as Structural Knowledge Collapse (SKC), where attention mechanisms fail to reconstruct coherent concepts from fragmented inputs. While existing input-centric solutions like vocabulary expansion address this, they necessitate expensive embedding retraining and neglect internal attention compositionality. To this end, we introduce Morpheme-aware KV-aggregation Attention (MorphKA), a lightweight adapter that dynamically consolidates fragments without tokenizer changes. Bypassing tokenizer retraining, MorphKA employs a dual-phase strategy—Input-Level Mor-pheme Aggregation (IMA) and Context-Aware KV-Aggregation (AMRF)—to stabilize mor-pheme spans and synthesize higher-order concepts. Experiments on medical and legal benchmarks show MorphKA outperforms vocabulary adaptation baselines by 3.2–4.6%, reaching 7.9% on high-fragmentation terms. More-over, MorphKA reduces catastrophic interference on general capabilities by 18–22% with ∼ 80% fewer parameters than embedding re-training approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
- Continual Pre-training of Language ModelsZixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi et al.ICLR 2023 · 15 citations
Related papers
- Dialogue Without Limits: Constant-Sized KV Caches for Extended Response in LLMsRavi Ghadia, Avinash Kumar, Gaurav Jain, Prashant J. Nair et al.ICML 2025
- RaSE-KGC: A Relation-Aware Segment Encoding Approach for Knowledge Graph CompletionChenxiao Lin, Ye Luo, Kunhong Liu, Qingqiang WuICDE 2026
- Knowledge-driven Augmentation and Retrieval for Integrative Temporal AdaptationWeisi Liu, Guangzeng Han, Xiaolei HuangACL 2026 · 1 citation
- Semi-Supervised Knowledge Amalgamation for Sequence ClassificationJidapa Thadajarassiri, Thomas Hartvigsen, Xiangnan Kong, Elke A. RundensteinerAAAI 2021 · 14 citations
- FiTs: Fine-Grained Two-Stage Training for Knowledge-Aware Question AnsweringQichen Ye, Bowen Cao, Nuo Chen, Weiyuan Xu et al.AAAI 2023 · 23 citations
