FLUID QA: A Multilingual Benchmark for Figurative Language Usage in Dialogue across English, Chinese, and Korean
Seoyoon Park, Hyeji Choi, Minseon Kim, Subin An, Xiaonan Wang, Gyuri Choi, Hansaem Kim
Abstract
Figurative language conveys stance, emotion, and social nuance, making its appropriate use essential in dialogue. While large language models (LLMs) often succeed in recognizing figurative expressions at the sentence level, their ability to use them coherently in conversation remains uncertain. We introduce FLUID QA, the first multilingual benchmark that evaluates figurative usage in dialogue across English, Korean, and Chinese. Each item embeds figurative choices into multi-turn contexts. To support interpretation, we include FLUTE-bi, a sentence-level diagnostic task. Results reveal a persistent gap: models that perform well on FLUTE-bi frequently fail on FLUID QA, especially in sarcasm and metaphor. These errors reflect systematic rhetorical confusion and limited discourse reasoning. FLUID QA provides a scalable framework for assessing usage-level figurative competence across languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 834aa7b2-c804-45f6-806b-c0c46f437688Cited by top-tier papers1
Ask how each one uses itBuilds on4
- MEGA: Multilingual Evaluation of Generative AIKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng et al.EMNLP 2023 · 91 citations
- IMPLI: Investigating NLI Models' Performance on Figurative LanguageKevin Stowe, Prasetya Ajie Utama, Iryna GurevychACL 2022 · 52 citations
- FLUTE: Figurative Language Understanding through Textual ExplanationsTuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, Smaranda MuresanEMNLP 2022 · 35 citations
- Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse PromptsXuan-Phi Nguyen, Mahani Aljunied, Shafiq Joty, Lidong BingACL 2024
Related papers
- FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human FeedbackYouquan Li, Miao Zheng, Fan Yang, Guosheng Dong et al.EMNLP 2025
- PunchBench: Benchmarking MLLMs in Multimodal Punchline ComprehensionKun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu et al.ACL 2025 · 3 citations
- (RSA)²: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language UnderstandingCesare Spinoso Di Piano, David Eric Austin, Pablo Piantanida, Jackie CK CheungACL 2025 · 1 citation
- Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET BenchmarkMinje Choi, Jiaxin Pei, Sagar Kumar, Chang Shu et al.EMNLP 2023 · 36 citations
- EMODIS: A Benchmark for Context-Dependent Emoji Disambiguation in Large Language ModelsJiacheng Huang, Ning Yu, Xiaoyin YiAAAI 2026
