Can Large Language Models Understand Internet Buzzwords Through User-Generated Content
Chen Huang, Junkai Luo, Xinzuo Wang, Wenqiang Lei, Jiancheng Lv
Abstract
The massive user-generated content (UGC) available in Chinese social media is giving rise to the possibility of studying internet buzzwords. In this paper, we study if large language models (LLMs) can generate accurate definitions for these buzzwords based on UGC as examples. Our work serves a threefold contribution. First, we introduce CHEER, the first dataset of Chinese internet buzzwords, each annotated with a definition and relevant UGC. Second, we propose a novel method, called RESS, to effectively steer the comprehending process of LLMs to produce more accurate buzzword definitions, mirroring the skills of human language learning. Third, with CHEER, we benchmark the strengths and weaknesses of various off-the-shelf definition generation methods and our RESS. Our benchmark demonstrates the effectiveness of RESS while revealing crucial shared challenges: over-reliance on prior exposure, underdeveloped inferential abilities, and difficulty identifying high-quality UGC to facilitate comprehension. We believe our work lays the groundwork for future advancements in LLM-based definition generation. Our dataset and code are available at https://github.com/SCUNLP/Buzzword. # Buzzwords 1127.0 # UGC (Example Sentences) 34607.0 Avg. #examples per buzzword 30.7 Avg. length of description per buzzword 262.5 Avg. length of definition per buzzword 50.0 Avg. length of examples per buzzword 85.4
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e63c476-e0fe-4e73-bba9-ddb2195cbc63Cited by top-tier papers2
- New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMsShiyao Cui, Qinglin Zhang, Di Wang, Yida Lu et al.ACL 2026
- Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGCLinjuan Wu, Ruiqi Zhang, Xinze Lyu, Ye Guo et al.ICML 2026
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang et al.ICLR 2024 · 176 citations
Related papers
- Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme ExplanationYubo Xie, Chenkai Wang, Zongyang Ma, Fahui MiaoEMNLP 2025
- CSCD-NS: a Chinese Spelling Check Dataset for Native SpeakersYong Hu, Fandong Meng, Jie ZhouACL 2024 · 9 citations
- SNS-Bench: Defining, Building, and Assessing Capabilities of Large Language Models in Social Networking ServicesHongcheng Guo, Yue Wang, Shaosheng Cao, Fei Zhao et al.ICML 2025
- Can Language Models Make Fun? A Case Study in Chinese Comical CrosstalkJianquan Li, Xiangbo Wu, Xiaokang Liu, Qianqian Xie et al.ACL 2023 · 2 citations
- Talk Funny! A Large-Scale Humor Response Dataset with Chain-of-Humor InterpretationYuyan Chen, Yichen Yuan, Panjun Liu, Dayiheng Liu et al.AAAI 2024 · 34 citations
