PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization
Xu Sun, Lionel Delphin-Poulat, Christèle Tarnec, Anastasia Shimorina
摘要
Large language models (LLMs) are increasingly used for zero-shot conversation summarization, but often exhibit positional bias—tending to overemphasize content from the beginning or end of a conversation while neglecting the middle. To address this issue, we introduce PoSum-Bench, a comprehensive benchmark for evaluating positional bias in conversational summarization, featuring diverse English and French conversational datasets spanning formal meetings, casual conversations, and customer service interactions. We propose a novel semantic similarity-based sentence-level metric to quantify the direction and magnitude of positional bias in model-generated summaries, enabling systematic and reference-free evaluation across conversation positions, languages, and conversational contexts. Our benchmark and methodology thus provide the first systematic, cross-lingual framework for reference-free evaluation of positional bias in conversational summarization, laying the groundwork for developing more balanced and unbiased summarization models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2020 · 被引用 121 次
- SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri 等EMNLP 2023 · 被引用 29 次
- Leveraging Lead Bias for Zero-shot Abstractive News SummarizationChenguang Zhu, Ziyi Yang, Robert Gmyr, Michael Zeng 等SIGIR 2021 · 被引用 23 次
相关 Paper
- On Context Utilization in Summarization with Large Language ModelsMathieu Ravaut, Aixin Sun, Nancy F. Chen, Shafiq JotyACL 2024
- Towards Multi-dimensional Evaluation of LLM Summarization across Domains and LanguagesHyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng 等ACL 2025 · 被引用 8 次
- Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved AnnotationYulong Chen, Huajian Zhang, Yijie Zhou, Xuefeng Bai 等ACL 2023 · 被引用 5 次
- An Empirical Study of Many-to-Many Summarization with Large Language ModelsJiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang 等ACL 2025
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model ResponsesXin Xu, Xunzhi He, Churan Zhi, Ruizhe Chen 等ICLR 2026 · 被引用 4 次
