iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for Revision
Jingwen Bai, Wei Soon Cheong, Philippe Muller, Brian Y. Lim
摘要
Large Language Models (LLMs) have become indispensable for evaluating writing. However, text feedback they provide is often unintelligible, generic, and not specific to user criteria. Inspired by structured rubrics in education and intelligible AI explanations, we propose iRULER following identified design guidelines to scaffold the review process by specific criteria, providing justification for score selection, and offering actionable revisions to target different quality levels. To qualify user-defined criteria, we recursively used iRULER with a rubric-of-rubrics to iteratively refine rubrics. In controlled experiments on writing revision and rubric creation, iRULER most improved validated LLM-judged review scores and was perceived as most helpful and aligned compared to read-only rubric and text-based LLM feedback. Qualitative findings further support how iRULER satisfies the design guidelines for user-defined feedback. This work contributes interactive rubric tools for intelligible LLM-based review and revision of writing, and user-defined rubric creation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper43
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 被引用 758 次
相关 Paper
- Intelligent Support Engages Writers Through Relevant Cognitive ProcessesAndreas Göldi, Thiemo Wambsganss, Seyed Parsa Neshaei, Roman RietscheCHI 2024 · 被引用 20 次
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim 等CHI 2024 · 被引用 81 次
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme 等ACL 2024 · 被引用 27 次
- ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated AgentsZhuofeng Li, Yi Lu, Dongfu Jiang, Haoxiang Zhang 等ACL 2026 · 被引用 1 次
- Friction: Deciphering Writing Feedback into Writing Revisions through LLM-Assisted ReflectionChao Zhang, Kexin Ju, Peter Bidoshi, Yu-Chun Grace Yen 等CHI 2025 · 被引用 25 次
