LLMs in Sarcasm Detection? It's elementary! (Or is it?)
Priyanshu Mahato, Aniket Santosh Mishra, Kripabandhu Ghosh
摘要
While Large Language Models (LLMs) are frequently cited for their sophisticated pragmatic reasoning (Wei et al., 2022;Bubeck et al., 2023), recent progress in sarcasm detection increasingly relies on synthetic benchmarks (Li et al., 2025;Anonymous, 2025). This study exposes a catastrophic generalization gap in this paradigm: we observe that models achieve near-perfect accuracy on synthetic data but collapse to random guessing on organic human speech. By triangulating hidden state geometry, entropy analysis, and causal interventions, we demonstrate that this disparity stems from shortcut learning (Geirhos et al., 2020)-models exploit the low-entropy statistical signatures of generated text while remaining "semantically blind" to the pragmatic cues essential for irony. Our findings indicate that high performance on synthetic leaderboards reflects forensic pattern matching rather than the genuine linguistic intelligence assumed in prior work, creating a statistical mirage of competence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language ModelsJingyang Zhang, Jingwei Sun, Eric C. Yeats, Yang Ouyang 等ICLR 2025
相关 Paper
- The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMsZihan Chen, Yiming Zhang, Wenxiang Geng, Zenghui Ding 等ACL 2026
- LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal ModelsJunyan Ye, Baichuan Zhou, Zilong Huang, Junan Zhang 等ICLR 2025
- On the Emotion Understanding of Synthesized SpeechYuan Ge, Haishu Zhao, Aokai Hao, Junxiang Zhang 等ACL 2026 · 被引用 1 次
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language ModelsYongding Tao, Tian Wang, Yihong Dong, Huanyu Liu 等ICLR 2026 · 被引用 5 次
- Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal InferenceJin Du, Li Chen, Xun Xian, An Luo 等ICLR 2026 · 被引用 4 次
