"Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews
Ruyuan Wan, Changye Li, Ting-Hao 'Kenneth' Huang
Abstract
Coded language is an important part of human communication. It refers to cases where users intentionally encode meaning so that the surface text differs from the intended meaning and must be decoded to be understood. Current language models handle coded language poorly. Progress has been limited by the lack of real-world datasets and clear taxonomies. This paper introduces CODEDLANG, a dataset of 7,744 Chinese Google Maps reviews, including 900 reviews with span-level annotations of coded language. We developed a seven-class taxonomy that captures common encoding strategies, including phonetic, orthographic, and cross-lingual substitutions. We benchmarked language models on coded language detection, classification, and review rating prediction. Results show that even strong models can fail to identify or understand coded language. Because many coded expressions rely on pronunciation-based strategies, we further conducted a phonetic analysis of coded and decoded forms. Our code and dataset are publicly available 1 . Together, our results highlight coded language as an important and underexplored challenge for real-world NLP systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3c1d519-4520-4fd6-8c91-e8af28cfad91Builds on5
- UCTopic: Unsupervised Contrastive Learning for Phrase Representations and Topic MiningJiacheng Li, Jingbo Shang, Julian J. McAuleyACL 2022 · 68 citations
- Hashtag Re-Appropriation for Audience Control on Recommendation-Driven Social Media Xiaohongshu (rednote)Ruyuan Wan, Lingbo Tong, Tiffany Knearem, Toby Jia-Jun Li et al.CHI 2025 · 19 citations
- From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language ModelsJulia Mendelsohn, Ronan Le Bras, Yejin Choi, Maarten SapACL 2023 · 14 citations
- Trkic G00gle: Why and How Users Game Translation AlgorithmsSoomin Kim, Changhoon Oh, Won-Ik Cho, Donghoon Shin et al.CSCW 2021 · 10 citations
- ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking PerturbationsYunze Xiao, Yujia Hu, Kenny T. W. Choo, Roy Ka-Wei LeeEMNLP 2024 · 5 citations
Related papers
- Enhancing Chinese Offensive Language Detection with Homophonic PerturbationJunqi Wu, Shujie Ji, Kang Zhong, Huiling Peng et al.EMNLP 2025
- PunMemeCN: A Benchmark to Explore Vision-Language Models' Understanding of Chinese Pun MemesZhijun Xu, Siyu Yuan, Yiqiao Zhang, Jingyu Sun et al.EMNLP 2025
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisEnzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong et al.ACL 2026
- Towards Real-World Writing Assistance: A Chinese Character Checking Benchmark with Faked and Misspelled CharactersYinghui Li, Zishan Xu, Shaoshen Chen, Haojing Huang et al.ACL 2024 · 9 citations
- CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in ChinaBolun Sun, Charles Chang, Yuen Yuen Ang, Ruotong Mu et al.ACL 2026
