Lune

ACL2026Top-tier venue

LLMs in Sarcasm Detection? It's elementary! (Or is it?)

Priyanshu Mahato, Aniket Santosh Mishra, Kripabandhu Ghosh

2026Year

Abstract

While Large Language Models (LLMs) are frequently cited for their sophisticated pragmatic reasoning (Wei et al., 2022;Bubeck et al., 2023), recent progress in sarcasm detection increasingly relies on synthetic benchmarks (Li et al., 2025;Anonymous, 2025). This study exposes a catastrophic generalization gap in this paradigm: we observe that models achieve near-perfect accuracy on synthetic data but collapse to random guessing on organic human speech. By triangulating hidden state geometry, entropy analysis, and causal interventions, we demonstrate that this disparity stems from shortcut learning (Geirhos et al., 2020)-models exploit the low-entropy statistical signatures of generated text while remaining "semantically blind" to the pragmatic cues essential for irony. Our findings indicate that high performance on synthetic leaderboards reflects forensic pattern matching rather than the genuine linguistic intelligence assumed in prior work, creating a statistical mirage of competence.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9a160858-a0e4-47b2-9872-ca4c18d7ce0f

Builds on3

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines