ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
Olga Golovneva, Moya Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, Asli Celikyilmaz
摘要
Large language models show improved downstream task performance when prompted to generate step-by-step reasoning to justify their final answers (Nye et al., 2021; Wei et al., 2022) . These reasoning steps greatly improve model interpretability and verification, but objectively studying their correctness (independent of the final answer) is difficult without reliable methods for automatic evaluation. We simply do not know how often the stated reasoning steps actually support the final end task predictions. In this work, we present ROSCOE, a suite of interpretable, unsupervised automatic scores that improve and extend previous text generation evaluation metrics. To evaluate ROSCOE against baseline metrics, we design a typology of reasoning errors and collect synthetic and human evaluation scores on commonly used reasoning datasets. In contrast with existing metrics, ROSCOE can measure semantic consistency, logicality, informativeness, fluency, and factuality -among other traits -by leveraging properties of step-by-step rationales. We empirically verify the strength of our metrics on five human annotated and six programmatically perturbed diagnostics datasets -covering a diverse set of tasks that require reasoning skills and show that ROSCOE can consistently outperform baseline metrics. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Deductive Verification of Chain-of-Thought ReasoningZhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang 等NeurIPS 2023 · 被引用 234 次
- Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What MattersBoshi Wang, Sewon Min, Xiang Deng, Jiaming Shen 等ACL 2023 · 被引用 100 次
- CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language ModelsZhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele 等NeurIPS 2023 · 被引用 74 次
- Decompose, Analyze and Rethink: Solving Intricate Problems with Human-like Reasoning CycleShangzi Xue, Zhenya Huang, Jiayu Liu, Xin Lin 等NeurIPS 2024 · 被引用 55 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- Are Machine Rationales (Not) Useful to Humans? Measuring and Improving Human Utility of Free-text RationalesBrihi Joshi, Ziyi Liu, Sahana Ramnath, Aaron Chan 等ACL 2023 · 被引用 6 次
- Understanding Chain-of-Thought in LLMs through Information TheoryJean-Francois Ton, Muhammad Faaiz Taufiq, Yang LiuICML 2025
- Calibrating Reasoning in Language Models with Internal ConsistencyZhihui Xie, Jizhou Guo, Tong Yu, Shuai LiNeurIPS 2024 · 被引用 37 次
- Large Language Models Have Intrinsic Meta-Cognition, but Need a Good LensZiyang Ma, Qingyue Yuan, Zhenglin Wang, Deyu ZhouEMNLP 2025 · 被引用 7 次
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 等ICML 2024 · 被引用 90 次
