RubricRobustness: Evaluating the Sensitivity of Rubrics-Based Benchmarks to Simple Perturbations
Manasi Sharma
Abstract
The advancement of Large Language Models (LLMs) into higher-level reasoning domains has rendered traditional heuristic evaluators insufficient for long-form open-ended responses, precipitating the widespread adoption of rubric-based benchmarks. While these frameworks utilize expert-curated criteria and LLM-as-a-judge to assess open-ended generation, the intrinsic robustness of these evaluation harnesses to fundamental validity assessments remains critically under-investigated. To bridge this gap, we introduce RubricRobustness, a systematic sensitivity analysis framework that subjects these benchmarks to three common sense perturbations: semantic negation, stochastic deletion and irrelevant addition. We investigate the extent to which manipulating the semantic veracity of a model's response impacts its resulting score by applying the robustness framework to two of the most popular rubrics-based benchmarks: HealthBench and WildBench. Our findings reveal systematic vulnerabilities: while both benchmarks respond sharply to semantic negation (e.g., degradation slopes of approximately on HealthBench and on WildBench), they are substantially less responsive to irrelevant addition, often requiring over 35% of sentences to be perturbed before inducing even a 25% score drop. We argue that perturbation-based sensitivity analyses of this form are a necessary prerequisite for validating rubric coverage, ensuring that automated evaluation frameworks reliably penalize basic semantic failures. We will release our framework as an open-source tool for building more resilient benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4467a26-84a0-4a49-a9ca-3ec636641fafBuilds on12
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng et al.ICLR 2024 · 299 citations
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research AgentsManasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang et al.ICLR 2026 · 83 citations
- Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsSeungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin et al.EMNLP 2024 · 38 citations
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction FollowingYun He, Wenzhe Li, Hejia Zhang, Songlin Li et al.ACL 2026 · 36 citations
Related papers
- RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine GenerationSunzhu Li, Jiale Zhao, Huimin Ren, Zhenlin Wei et al.ACL 2026 · 22 citations
- RubricBench: Aligning Model-Generated Rubrics with Human StandardsJunyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu et al.ACL 2026 · 7 citations
- TabReX: Tabular Referenceless eXplainable EvaluationTejas Anvekar, Junha Park, Aparna Garimella, Vivek GuptaACL 2026
- MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMsHuiyi Chen, Jiawei Peng, Dehai Min, Changchang Sun et al.ICML 2026 · 18 citations
- VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language ModelsRohit Saxena, Alessandro Suglia, Pasquale MinerviniICML 2026 · 7 citations
