Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
Diana Galván-Sosa, Gabrielle Gaudeau, Pride Kavumba, Yunmeng Li, Hongyi Gu, Zheng Yuan, Keisuke Sakaguchi, Paula Buttery
摘要
The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations have been found to be unreliable, making it difficult for users to distinguish good from bad explanations. To address this issue, we present Rubrik's CUBE-an education-inspired rubric and a dataset of 26k explanations, written and later quality-annotated using the rubric by both humans and six open-and closedsource LLMs. The CUBE dataset focuses on two reasoning and two language tasks, providing the necessary diversity for us to effectively test our proposed rubric. Using Rubrik, we find that explanations are influenced by both task and perceived difficulty. Low quality stems primarily from a lack of conciseness in LLMgenerated explanations, rather than cohesion and word choice. The full dataset, rubric, and code are available at https://github.com/ RubriksCube/rubriks_cube . 1 See StackOverflow's policy on the use of ChatGPT and other LLMs: https:// COMPONENTS DIMENSIONS necessary parts of an explanation necessary qualities of a good explanation Typology of Explanations Language Content Typ1. COMMENTARY 1.a) Action Grammaticality Conciseness 1.b) Reason Word Choice Appropriateness Cohesion Coherence Typ2. JUSTIFICATION 2.a) Evidence Plausibility Typ3. ARGUMENT 3.a) Affective appeal(s) and Qualifier(s) Stance Clarity
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for Open-Ended LLM ReasoningYang Zhou, Sunzhu Li, Shunyu Liu, Wenkai Fang 等ICML 2026 · 被引用 44 次
- Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain ReasoningBaolong Bi, Shenghua Liu, Yiwei Wang, Siqian Tong 等ICML 2026 · 被引用 19 次
- RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model ReasoningYukun Chen, Jiaming Li, Longze Chen, Ze Gong 等ICML 2026 · 被引用 5 次
- Unbiased Principles, Robust RewardsQingnan Ren, Zhen Fang, Shiting Huang, Yu Zeng 等ICML 2026
它引用的顶会 Paper5
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu 等ICML 2024 · 被引用 406 次
- Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow QuestionsSamia Kabir, David N. Udo-Imeh, Bonan Kou, Tianyi ZhangCHI 2024 · 被引用 149 次
- Modeling Fluency and Faithfulness for Diverse Neural Machine TranslationYang Feng, Wanying Xie, Shuhao Gu, Chenze Shao 等AAAI 2020 · 被引用 28 次
- All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated TextElizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong 等ACL 2021
相关 Paper
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme 等ACL 2024 · 被引用 27 次
- iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for RevisionJingwen Bai, Wei Soon Cheong, Philippe Muller, Brian Y. LimCHI 2026 · 被引用 1 次
- XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMsZichen Chen, Jianda Chen, Ambuj K. Singh, Misha SraEMNLP 2024 · 被引用 5 次
- CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language ModelsTong Zhang, Peixin Qin, Yang Deng, Chen Huang 等ACL 2024
- RPC-Bench: A Fine-grained Benchmark for Research Paper ComprehensionYelin Chen, Fanjin Zhang, Suping Sun, Yunhe Pang 等ACL 2026 · 被引用 2 次
