LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, Chris Kedzie
Abstract
This paper introduces a framework for the automated evaluation of natural language texts. A manually constructed rubric describes how to assess multiple dimensions of interest. To evaluate a text, a large language model (LLM) is prompted with each rubric question and produces a distribution over potential responses. The LLM predictions often fail to agree well with human judges-indeed, the humans do not fully agree with one another. However, the multiple LLM distributions can be combined to predict each human judge's annotations on all questions, including a summary question that assesses overall quality or relevance. LLM-RUBRIC accomplishes this by training a small feed-forward neural network that includes both judge-specific and judge-independent parameters. When evaluating dialogue systems in a human-AI information-seeking task, we find that LLM-RUBRIC with 9 questions (assessing dimensions such as naturalness, conciseness, and citation quality) predicts human judges' assessment of overall user satisfaction, on a scale of 1-4, with RMS error ă 0.5, a 2î mprovement over the uncalibrated baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6ccc2ba-501f-4515-9cc7-ebca38e2233dCited by top-tier papers19
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsAnisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath et al.ICLR 2026 · 340 citations
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM AlignmentTianci Liu, Ran Xu, Tony Yu, Ilgee Hong et al.ACL 2026 · 75 citations
- ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured ChecklistsJie Ruan, Inderjeet Nair, Shuyang Cao, Amy Liu et al.ICLR 2026 · 25 citations
- mR3: Multilingual Rubric-Agnostic Reward Reasoning ModelsDavid Anugraha, Shou-Yi Hung, Zilu Tang, En-Shiun Annie Lee et al.ICLR 2026 · 9 citations
- All Code, No Thought: Language Models Struggle to Reason in Ciphered LanguageShiyuan Guo, Henry Sleight, Fabien RogerICLR 2026 · 5 citations
Builds on22
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 338 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
Related papers
- HypoEval: Hypothesis-Guided Evaluation for Natural Language GenerationMingxuan Li, Hanchen Li, Chenhao TanACL 2026 · 1 citation
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai et al.ACL 2024
- SaMer: A Scenario-aware Multi-dimensional Evaluator for Large Language ModelsKehua Feng, Keyan Ding, Jing Yu, Yiwen Qu et al.ICLR 2025
- MENLO: From Preferences to Proficiency - Evaluating and Modeling Native-like Quality Across 47 LanguagesChenxi Whitehouse, Sebastian Ruder, Tony Lin, Oksana Kurylo et al.ICLR 2026 · 3 citations
- LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English WritingZhengxiang Wang, Veronika Makarova, Zhi Li, Jordan Kodner et al.ACL 2025 · 5 citations
