Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
Yang Wang, Chenghao Xiao, Chia-Yi Hsiao, Zi Yan Chang, Chi-Li Chen, Tyler Loakman, Chenghua Lin
Abstract
We introduce Drivelology, a unique linguistic phenomenon characterised as "nonsense with depth" -utterances that are syntactically coherent yet pragmatically paradoxical, emotionally loaded, or rhetorically subversive. While such expressions may resemble surface-level nonsense, they encode implicit meaning requiring contextual inference, moral reasoning, or emotional interpretation. We find that current large language models (LLMs), despite excelling at many natural language processing (NLP) tasks, consistently fail to grasp the layered semantics of Drivelological text. To investigate this, we construct a benchmark dataset of over 1,200+ meticulously curated and diverse examples across English, Mandarin, Spanish, French, Japanese, and Korean. Each example underwent careful expert review to verify its Drivelological characteristics, involving multiple rounds of discussion and adjudication to address disagreements. Using this dataset, we evaluate a range of LLMs on classification, generation, and reasoning tasks. Our results reveal clear limitations of LLMs: models often confuse Drivelology with shallow nonsense, produce incoherent justifications, or miss implied rhetorical functions altogether. These findings highlight a deep representational gap in LLMs' pragmatic understanding and challenge the assumption that statistical fluency implies cognitive comprehension. We release our dataset 1 and code 2 to facilitate further research in modelling linguistic depth beyond surface-level coherence. * "The cat sat on the mat." (normal sentence) * "Colourless green ideas sleep furiously." (pure nonsense) ## Annotation Tasks • Drivelology Tagging -Task: Classify Drivelology samples into one or more categories only if the sample is Drivelology: * Misdirection: A rhetorical technique where the focus shifts but connects back to the original topic through indirect hints. * Paradox: A statement that combines ideas that do not logically fit together but conveys a deeper meaning. * Switchbait: A language trick that changes meaning based on cultural knowledge or idioms. * Inversion: Rearranging the usual order of words or ideas to create a surprising effect. * Wordplay: Creative use of language through puns or double meanings. -Instructions: * Identify the primary characteristics (i.e., the first strong impression) of the text. * Assign one or more categories based on the definitions above. • Implicit Narrative Writing
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 657f829a-26dc-4c6b-b4c0-ebb4a0cece55Builds on8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous ContradictionsZhe Hu, Tuo Liang, Jing Li, Yiren Lu et al.NeurIPS 2024 · 21 citations
Related papers
- This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsIker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios et al.EMNLP 2023 · 8 citations
- M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMsTianlong Zheng, Yating Yang, Rui Dong, Bo Ma et al.AAAI 2026
- ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language UnderstandingSayan Ghosh, Shashank SrivastavaACL 2022
- DRInQ: Evaluating Conversational Implicature with Controlled Context VariationHirona Jacqueline Arai, Xiang RenACL 2026
- Metaphor Understanding Challenge Dataset for LLMsXiaoyu Tong, Rochelle Choenni, Martha Lewis, Ekaterina ShutovaACL 2024
