HanjaBridge: Resolving Semantic Ambiguity in Korean LLMs via Hanja-Augmented Pre-Training
Seungho Choi, Sihyun Park, Minsang Kim, Chansol Park, Bongsu Kim
Abstract
Large language models (LLMs) often show poor performance in low-resource languages like Korean, partly due to unique linguistic challenges such as homophonous Sino-Korean words that are indistinguishable in Hangul script. To address this semantic ambiguity, we propose HanjaBridge, a novel meaning-injection technique integrated into a continual pre-training (CPT) framework. Instead of deterministically mapping a word to a single Hanja (Chinese character), HanjaBridge presents the model with all possible Hanja candidates for a given homograph, encouraging the model to learn contextual disambiguation. This process is paired with token-level knowledge distillation to prevent catastrophic forgetting. Experimental results show that HanjaBridge significantly improves Korean language understanding, achieving a 21% relative improvement on the KoBALT benchmark. Notably, by reinforcing semantic alignment between Korean and Chinese through shared Hanja, we observe a strong positive cross-lingual transfer. Furthermore, these gains persist even when Hanja augmentation is omitted at inference time, ensuring practical efficiency with no additional run-time cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 455e8692-d797-45c2-92cb-a611a9ded655Builds on4
- SEED: Self-supervised Distillation For Visual RepresentationZhiyuan Fang, Jianfeng Wang, Lijuan Wang, Lei Zhang et al.ICLR 2021 · 213 citations
- What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained TransformersBoseop Kim, HyoungSeok Kim, Sang-Woo Lee, Gichang Lee et al.EMNLP 2021 · 11 citations
- Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-trainingZheheng Luo, Xin Zhang, Xiao Liu, Haoling Li et al.ACL 2025 · 8 citations
- Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual TransferabilityYoshinari Fujinuma, Jordan L. Boyd-Graber, Katharina KannACL 2022
Related papers
- Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language ModelMingqi Li, Fei Ding, Dan Zhang, Long Cheng et al.EMNLP 2022 · 3 citations
- Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible MultilingualityMengyu Bu, Yang FengACL 2026 · 2 citations
- TokAlign: Efficient Vocabulary Adaptation via Token AlignmentChong Li, Jiajun Zhang, Chengqing ZongACL 2025 · 7 citations
- PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity RecognitionTao Zhang, Congying Xia, Philip S. Yu, Zhiwei Liu et al.EMNLP 2021 · 22 citations
- Toward Robust Multilingual Adaptation of LLMs for Low-Resource LanguagesHaolin Li, Haipeng Zhang, Mang Li, Yaohua Wang et al.ICML 2026
