Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura, Ketan Pravin More, Omkar Thawakar, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer
摘要
Arabic poetry is one of the richest and most culturally rooted forms of expression in the Arabic language, known for its layered meanings, stylistic diversity, and deep historical continuity. Although large language models (LLMs) have demonstrated strong performance across languages and tasks, their ability to understand Arabic poetry remains largely unexplored. In this work, we introduce Fann or Flop, the first benchmark designed to assess the comprehension of Arabic poetry by LLMs in 12 historical eras, covering 14 core poetic genres and a variety of metrical forms, from classical structures to contemporary free verse. The benchmark comprises a curated corpus of poems with explanations that assess semantic understanding, metaphor interpretation, prosodic awareness, and cultural context. We argue that poetic comprehension offers a strong indicator for testing how good the LLM understands classical Arabic through Arabic poetry. Unlike surfacelevel tasks, this domain demands deeper interpretive reasoning and cultural sensitivity. Our evaluation of state-of-the-art LLMs shows that most models struggle with poetic understanding despite strong results on standard Arabic benchmarks. We release Fann or Flop 1 along with the evaluation suite 2 as an open-source resource to enable rigorous evaluation and advancement for Arabic language models. Era Approx. Years Genres (Theme) Meter Notable Poets Pre-Islamic (Jahiliyyah)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- FLUTE: Figurative Language Understanding through Textual ExplanationsTuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, Smaranda MuresanEMNLP 2022 · 被引用 35 次
- ALLaM: Large Language Models for Arabic and EnglishM. Saiful Bari, Yazeed Alnumay, Norah A. Alzahrani, Nouf M. Alotaibi 等ICLR 2025 · 被引用 4 次
相关 Paper
- POEMetric: The Last Stanza of HumanityBingru Li, Han Wang, Hazel WilkinsonICLR 2026 · 被引用 2 次
- Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese PoetryHan Zhang, Zihan Gu, Zhiyuan Wang, Tianyi Ma 等ACL 2026
- Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and EleganceAndong Chen, Lianzhang Lou, Kehai Chen, Xuefeng Bai 等EMNLP 2025
- F-Eval: Asssessing Fundamental Abilities with Refined Evaluation MethodsYu Sun, Keyuchen Keyuchen, Shujie Wang, Peiji Li 等ACL 2024
- We Politely Insist: Your LLM Must Learn the Persian Art of TaarofNikta Gohari Sadr, Sahar Heidariasl, Karine Megerdoomian, Laleh Seyyed-Kalantari 等EMNLP 2025 · 被引用 2 次
