LLMs in Sarcasm Detection? It's elementary! (Or is it?)
Priyanshu Mahato, Aniket Santosh Mishra, Kripabandhu Ghosh
Abstract
While Large Language Models (LLMs) are frequently cited for their sophisticated pragmatic reasoning (Wei et al., 2022;Bubeck et al., 2023), recent progress in sarcasm detection increasingly relies on synthetic benchmarks (Li et al., 2025;Anonymous, 2025). This study exposes a catastrophic generalization gap in this paradigm: we observe that models achieve near-perfect accuracy on synthetic data but collapse to random guessing on organic human speech. By triangulating hidden state geometry, entropy analysis, and causal interventions, we demonstrate that this disparity stems from shortcut learning (Geirhos et al., 2020)-models exploit the low-entropy statistical signatures of generated text while remaining "semantically blind" to the pragmatic cues essential for irony. Our findings indicate that high performance on synthetic leaderboards reflects forensic pattern matching rather than the genuine linguistic intelligence assumed in prior work, creating a statistical mirage of competence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a160858-a0e4-47b2-9872-ca4c18d7ce0fBuilds on3
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language ModelsJingyang Zhang, Jingwei Sun, Eric C. Yeats, Yang Ouyang et al.ICLR 2025
Related papers
- The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMsZihan Chen, Yiming Zhang, Wenxiang Geng, Zenghui Ding et al.ACL 2026
- LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal ModelsJunyan Ye, Baichuan Zhou, Zilong Huang, Junan Zhang et al.ICLR 2025
- On the Emotion Understanding of Synthesized SpeechYuan Ge, Haishu Zhao, Aokai Hao, Junxiang Zhang et al.ACL 2026 · 1 citation
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language ModelsYongding Tao, Tian Wang, Yihong Dong, Huanyu Liu et al.ICLR 2026 · 5 citations
- Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal InferenceJin Du, Li Chen, Xun Xian, An Luo et al.ICLR 2026 · 4 citations
