Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
Diana Galván-Sosa, Gabrielle Gaudeau, Pride Kavumba, Yunmeng Li, Hongyi Gu, Zheng Yuan, Keisuke Sakaguchi, Paula Buttery
Abstract
The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations have been found to be unreliable, making it difficult for users to distinguish good from bad explanations. To address this issue, we present Rubrik's CUBE-an education-inspired rubric and a dataset of 26k explanations, written and later quality-annotated using the rubric by both humans and six open-and closedsource LLMs. The CUBE dataset focuses on two reasoning and two language tasks, providing the necessary diversity for us to effectively test our proposed rubric. Using Rubrik, we find that explanations are influenced by both task and perceived difficulty. Low quality stems primarily from a lack of conciseness in LLMgenerated explanations, rather than cohesion and word choice. The full dataset, rubric, and code are available at https://github.com/ RubriksCube/rubriks_cube . 1 See StackOverflow's policy on the use of ChatGPT and other LLMs: https:// COMPONENTS DIMENSIONS necessary parts of an explanation necessary qualities of a good explanation Typology of Explanations Language Content Typ1. COMMENTARY 1.a) Action Grammaticality Conciseness 1.b) Reason Word Choice Appropriateness Cohesion Coherence Typ2. JUSTIFICATION 2.a) Evidence Plausibility Typ3. ARGUMENT 3.a) Affective appeal(s) and Qualifier(s) Stance Clarity
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for Open-Ended LLM ReasoningYang Zhou, Sunzhu Li, Shunyu Liu, Wenkai Fang et al.ICML 2026 · 44 citations
- Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain ReasoningBaolong Bi, Shenghua Liu, Yiwei Wang, Siqian Tong et al.ICML 2026 · 19 citations
- RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model ReasoningYukun Chen, Jiaming Li, Longze Chen, Ze Gong et al.ICML 2026 · 5 citations
- Unbiased Principles, Robust RewardsQingnan Ren, Zhen Fang, Shiting Huang, Yu Zeng et al.ICML 2026
Builds on5
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu et al.ICML 2024 · 406 citations
- Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow QuestionsSamia Kabir, David N. Udo-Imeh, Bonan Kou, Tianyi ZhangCHI 2024 · 149 citations
- Modeling Fluency and Faithfulness for Diverse Neural Machine TranslationYang Feng, Wanying Xie, Shuhao Gu, Chenze Shao et al.AAAI 2020 · 28 citations
- All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated TextElizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong et al.ACL 2021
Related papers
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for RevisionJingwen Bai, Wei Soon Cheong, Philippe Muller, Brian Y. LimCHI 2026 · 1 citation
- XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMsZichen Chen, Jianda Chen, Ambuj K. Singh, Misha SraEMNLP 2024 · 5 citations
- CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language ModelsTong Zhang, Peixin Qin, Yang Deng, Chen Huang et al.ACL 2024
- RPC-Bench: A Fine-grained Benchmark for Research Paper ComprehensionYelin Chen, Fanjin Zhang, Suping Sun, Yunhe Pang et al.ACL 2026 · 2 citations
