HanjaBridge: Resolving Semantic Ambiguity in Korean LLMs via Hanja-Augmented Pre-Training
Seungho Choi, Sihyun Park, Minsang Kim, Chansol Park, Bongsu Kim
摘要
Large language models (LLMs) often show poor performance in low-resource languages like Korean, partly due to unique linguistic challenges such as homophonous Sino-Korean words that are indistinguishable in Hangul script. To address this semantic ambiguity, we propose HanjaBridge, a novel meaning-injection technique integrated into a continual pre-training (CPT) framework. Instead of deterministically mapping a word to a single Hanja (Chinese character), HanjaBridge presents the model with all possible Hanja candidates for a given homograph, encouraging the model to learn contextual disambiguation. This process is paired with token-level knowledge distillation to prevent catastrophic forgetting. Experimental results show that HanjaBridge significantly improves Korean language understanding, achieving a 21% relative improvement on the KoBALT benchmark. Notably, by reinforcing semantic alignment between Korean and Chinese through shared Hanja, we observe a strong positive cross-lingual transfer. Furthermore, these gains persist even when Hanja augmentation is omitted at inference time, ensuring practical efficiency with no additional run-time cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- SEED: Self-supervised Distillation For Visual RepresentationZhiyuan Fang, Jianfeng Wang, Lijuan Wang, Lei Zhang 等ICLR 2021 · 被引用 213 次
- What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained TransformersBoseop Kim, HyoungSeok Kim, Sang-Woo Lee, Gichang Lee 等EMNLP 2021 · 被引用 11 次
- Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-trainingZheheng Luo, Xin Zhang, Xiao Liu, Haoling Li 等ACL 2025 · 被引用 8 次
- Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual TransferabilityYoshinari Fujinuma, Jordan L. Boyd-Graber, Katharina KannACL 2022
相关 Paper
- Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language ModelMingqi Li, Fei Ding, Dan Zhang, Long Cheng 等EMNLP 2022 · 被引用 3 次
- Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible MultilingualityMengyu Bu, Yang FengACL 2026 · 被引用 2 次
- TokAlign: Efficient Vocabulary Adaptation via Token AlignmentChong Li, Jiajun Zhang, Chengqing ZongACL 2025 · 被引用 7 次
- PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity RecognitionTao Zhang, Congying Xia, Philip S. Yu, Zhiwei Liu 等EMNLP 2021 · 被引用 22 次
- Toward Robust Multilingual Adaptation of LLMs for Low-Resource LanguagesHaolin Li, Haipeng Zhang, Mang Li, Yaohua Wang 等ICML 2026
