PIPER: Benchmarking and Prompting Event Reasoning Boundary of LLMs via Debiasing-Distillation Enhanced Tuning
Zhicong Lu, Changyuan Tian, PeiguangLi PeiguangLi, Li Jin, Sirui Wang, Wei Jia, Ying Shen, Guangluan Xu
摘要
While Large Language Models (LLMs) excel in diverse domains, their validity in event reasoning remains underexplored. Most existing works merely stagnate at assessing LLMs' event reasoning with a single event relational type or reasoning format, failing to conduct a complete evaluation and provide a practical solution for capability enhancement. In this paper, we propose PIPER, the first comprehensive benchmark for Probing Into the Performance boundary of LLMs in Event Reasoning. Motivated by our evaluation observations and error patterns analysis, we meticulously craft 10K diverse instruction-tuning demonstrations to alleviate event reasoning-oriented data scarcity. Additionally, a novel Debiasing and Distillation-Enhanced Supervised Fine-Tuning (D 2 E-SFT) strategy is presented, which facilitates adhering to context and fixating significant contextual event information to elevate the event reasoning capability. Specifically, D 2 E-SFT removes the given sample's context to construct an imagined sample, subtracting its logits to mitigate the bias of neglecting context and improve contextual faithfulness. To guide the model in emphasizing significant contextual event information, D 2 E-SFT employs a context-refined sample to achieve selfdistillation with the alignment of logits. Extensive experimental results demonstrate the effectiveness of our data and strategy in expanding the performance boundary of event reasoning. * Equal Contribution † Corresponding author. (a) (b) FC Barcelona's youth academy, La Masia. Did Messi's leadership in the 2022 World Cup final lead to Argentina's victory over France? (Causal) Where did Messi begin his football career before making his first-team debut in 2004? (Temporal) How would Messi's legacy be perceived if he hadn't won the 2022 World Cup? A. Unchanged B. Diminished C. Enhanced D. None. (Counterfactual) CRI CEI LLM QA Yes No NLI Question SCQ B D C A
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
相关 Paper
- A Comprehensive Evaluation on Event Reasoning of Large Language ModelsZhengwei Tao, Zhi Jin, Yifan Zhang, Xiancai Chen 等AAAI 2025 · 被引用 8 次
- Once Upon an Input: Reasoning via Per-Instance Program SynthesisAdam Stein, Neelay Velingker, Mayur Naik, Eric WongNeurIPS 2025
- PINTO: Faithful Language Reasoning Using Prompt-Generated RationalesPeifeng Wang, Aaron Chan, Filip Ilievski, Muhao Chen 等ICLR 2023 · 被引用 21 次
- DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM ReasoningWeize Liu, Yongchi Zhao, Yijia Luo, Mingyu Xu 等ICLR 2026 · 被引用 6 次
- Back to the Future: Towards Explainable Temporal Reasoning with Large Language ModelsChenhan Yuan, Qianqian Xie, Jimin Huang, Sophia AnaniadouWWW 2024
