Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning
Zhu Xu, Zhiqiang Zhao, Zihan Zhang, Yuchi Liu, Quanwei Shen, Fei Liu, Yu Kuang, Jian He, Conglin Liu
Abstract
Tokenization methods like Byte-Pair Encoding (BPE) enhance computational efficiency in large language models (LLMs) but often obscure internal character structures within tokens. This limitation hinders LLMs' ability to predict precise character positions, which is crucial in tasks like Chinese Spelling Correction (CSC) where identifying the positions of misspelled characters accelerates correction processes. We propose Token Internal Position Awareness (TIPA), a method that significantly improves models' ability to capture character positions within tokens by training them on reverse character prediction tasks using the tokenizer's vocabulary. Experiments demonstrate that TIPA enhances position prediction accuracy in LLMs, enabling more precise identification of target characters in original text. Furthermore, when applied to downstream tasks that do not require exact position prediction, TIPA still boosts performance in tasks needing character-level information, validating its versatility and effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79b67760-d9ff-44ba-a422-9a03ae2f2708Cited by top-tier papers1
Ask how each one uses itBuilds on11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 213 citations
- Premise Order Matters in Reasoning with Large Language ModelsXinyun Chen, Ryan A. Chi, Xuezhi Wang, Denny ZhouICML 2024 · 59 citations
- Chinese Spelling Correction as Rephrasing Language ModelLinfeng Liu, Hongqiu Wu, Hai ZhaoAAAI 2024 · 36 citations
Related papers
- PLOME: Pre-training with Misspelled Knowledge for Chinese Spelling CorrectionShulin Liu, Tao Yang, Tianchi Yue, Feng Zhang et al.ACL 2021
- A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language ModelsHouquan Zhou, Zhenghua Li, Bo Zhang, Chen Li et al.EMNLP 2024 · 2 citations
- C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterKunting Li, Yong Hu, Liang He, Fandong Meng et al.EMNLP 2024 · 9 citations
- Spelling Error Correction with Soft-Masked BERTShaohua Zhang, Haoran Huang, Jicong Liu, Hang LiACL 2020 · 204 citations
- ARM: An Alignment-and-Replacement Module for Chinese Spelling Check Based on LLMsChangchun Liu, Kai Zhang, Junzhe Jiang, Zirui Liu et al.EMNLP 2024 · 3 citations
