Lune

ACL2026顶会

LLMs in Sarcasm Detection? It's elementary! (Or is it?)

Priyanshu Mahato, Aniket Santosh Mishra, Kripabandhu Ghosh

2026年份

摘要

While Large Language Models (LLMs) are frequently cited for their sophisticated pragmatic reasoning (Wei et al., 2022;Bubeck et al., 2023), recent progress in sarcasm detection increasingly relies on synthetic benchmarks (Li et al., 2025;Anonymous, 2025). This study exposes a catastrophic generalization gap in this paradigm: we observe that models achieve near-perfect accuracy on synthetic data but collapse to random guessing on organic human speech. By triangulating hidden state geometry, entropy analysis, and causal interventions, we demonstrate that this disparity stems from shortcut learning (Geirhos et al., 2020)-models exploit the low-entropy statistical signatures of generated text while remaining "semantically blind" to the pragmatic cues essential for irony. Our findings indicate that high performance on synthetic leaderboards reflects forensic pattern matching rather than the genuine linguistic intelligence assumed in prior work, creating a statistical mirage of competence.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper3

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖