Is Sarcasm Detection a Step-by-Step Reasoning Process in Large Language Models?
Ben Yao, Yazhou Zhang, Qiuchi Li, Jing Qin
Abstract
Elaborating a series of intermediate reasoning steps significantly improves the ability of large language models (LLMs) to solve complex problems, as such steps would evoke LLMs to think sequentially. However, human sarcasm understanding is often considered an intuitive and holistic cognitive process, in which various linguistic, contextual, and emotional cues are integrated to form a comprehensive understanding, in a way that does not necessarily follow a step-bystep fashion. To verify the validity of this argument, we introduce a new prompting framework (called SarcasmCue) containing four sub-methods, viz. chain of contradiction (CoC), graph of cues (GoC), bagging of cues (BoC) and tensor of cues (ToC), which elicits LLMs to detect human sarcasm by considering sequential and non-sequential prompting methods. Through a comprehensive empirical comparison on four benchmarks, we highlight three key findings: (1) CoC and GoC show superior performance with more advanced models like GPT-4 and Claude 3.5, with an improvement of 3.5% ↑. (2) ToC significantly outperforms other methods when smaller LLMs are evaluated, boosting the F1 score by 29.7% ↑ over the best baseline. (3) Our proposed framework consistently pushes the state-of-the-art (i.e., ToT) by 4.2%, 2.0%, 29.7%, and 58.2% in F1 scores across four datasets. This demonstrates the effectiveness and stability of the proposed framework. 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- RAM-SD: Retrieval-Augmented Multi-agent framework for Sarcasm DetectionZiyang Zhou, Ziqi Liu, Yan Wang, Yiming Lin et al.ACL 2026 · 1 citation
- DRInQ: Evaluating Conversational Implicature with Controlled Context VariationHirona Jacqueline Arai, Xiang RenACL 2026
- Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis ModelsGuangyu Xie, Yice Zhang, Jianzhu Bao, Qianlong Wang et al.EMNLP 2025
- SEVADE: Self-Evolving Multi-Agent Analysis with Decoupled Evaluation for Hallucination-Resistant Sarcasm DetectionZiqi Liu, Ziyang Zhou, Yilin Li, Mingxuan Hu et al.AAAI 2026
- Reinforcement Learning-Guided Adaptive Tuning for Out-of-Distribution Harmful Text DetectionMengyu Xiang, Tinghao Chen, Boxu Han, Qiudan Li et al.ACL 2026
Builds on5
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 234 citations
- Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional NetworkBin Liang, Chenwei Lou, Xiang Li, Min Yang et al.ACL 2022 · 151 citations
- Humor Knowledge Enriched Transformer for Understanding Multimodal HumorMd. Kamrul Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh et al.AAAI 2021 · 98 citations
Related papers
- Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic LanguageXi Chen, Shuo WangEMNLP 2025 · 7 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?Nemika Tyagi, Mihir Parmar, Mohith Kulkarni, Aswin RRV et al.EMNLP 2024 · 3 citations
- Chain of Code: Reasoning with a Language Model-Augmented Code EmulatorChengshu Li, Jacky Liang, Andy Zeng, Xinyun Chen et al.ICML 2024 · 155 citations
- PunchBench: Benchmarking MLLMs in Multimodal Punchline ComprehensionKun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu et al.ACL 2025 · 3 citations
