ALERT: Adapt Language Models to Reasoning Tasks
Ping Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin, Gargi Ghosh, Mona T. Diab, Asli Celikyilmaz
摘要
Recent advancements in large language models have enabled them to perform well on complex tasks that require step-by-step reasoning with few-shot learning. However, it is unclear whether these models are applying reasoning skills they have learned during pre-training, or if they are simply memorizing their training corpus at finer granularity and have learned to better understand their context. To address this question, we introduce ALERT, a benchmark and suite of analyses for evaluating reasoning skills of language models. ALERT enables comparing pre-trained and finetuned models on complex tasks that require reasoning skills to solve them. Our benchmark provides a test bed to assess any language model on fine-grained reasoning skills, which spans over 20 datasets and covers 10 different reasoning skills. To prove the efficacy of ALERT we investigate the role of finetuning. Our extensive empirical analysis shows that language models acquire reasoning skills such as textual entailment, abductive reasoning, and analogical reasoning during the finetuning stage compared to pretraining stage. Another finding is when language models are finetuned they tend to overfit to the prompt template, which hurts the robustness of models resulting in generalization problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across LanguagesLibo Qin, Qiguang Chen, Fuxuan Wei, Shijue Huang 等EMNLP 2023 · 被引用 26 次
- SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit TokensYinhan He, Wendy Zheng, Yaochen Zhu, Zaiyi Zheng 等NeurIPS 2025 · 被引用 19 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningAntonia Creswell, Murray Shanahan, Irina HigginsICLR 2023 · 被引用 110 次
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?Zhanke Zhou, Xiao Feng, Zhaocheng Zhu, Jiangchao Yao 等ICML 2025
- ThinkSum: Probabilistic reasoning over sets using large language modelsBatu Ozturkler, Nikolay Malkin, Zhen Wang, Nebojsa JojicACL 2023 · 被引用 10 次
- Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?Neeladri Bhuiya, Viktor Schlegel, Stefan WinklerEMNLP 2024 · 被引用 2 次
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 等EMNLP 2023 · 被引用 17 次
