On Context Utilization in Summarization with Large Language Models
Mathieu Ravaut, Aixin Sun, Nancy F. Chen, Shafiq Joty
Abstract
Large language models (LLMs) excel in abstractive summarization tasks, delivering fluent and pertinent summaries. Recent advancements have extended their capabilities to handle long-input contexts, exceeding 100k tokens. However, in question answering, language models exhibit uneven utilization of their input context. They tend to favor the initial and final segments, resulting in a U-shaped performance pattern concerning where the answer is located within the input. This bias raises concerns, particularly in summarization where crucial content may be dispersed throughout the source document(s). Besides, in summarization, mapping facts from the source to the summary is not trivial as salient content is usually re-phrased. In this paper, we conduct the first comprehensive study on context utilization and position bias in summarization. Our analysis encompasses 6 LLMs, 10 datasets, and 5 evaluation metrics. We introduce a new evaluation benchmark called MiddleSum on the which we benchmark two alternative inference methods to alleviate position bias: hierarchical summarization and incremental summarization 1 . Metric Model CNN/DM XSum Reddit SAMSum Multi-X AVG Arxiv PubMed GovReport SummScreenFD Multi-N AVG ROUGE-2 Flan-UL2 -0.296 -0.124 0.048 -0.069 -0.201 -0.128 _ _ _ _ _ _ Llama-2-7B -0.160 -0.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 850daaec-043a-4fc6-beaf-398cfec4f0c5Cited by top-tier papers15
- Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsPhilippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng WuEMNLP 2024 · 19 citations
- Media Bias Detector: Designing and Implementing a Tool for Real-Time Selection and Framing Bias Analysis in News CoverageJenny S. Wang, Samar Haider, Amir Tohidi, Anushkaa Gupta et al.CHI 2025 · 11 citations
- Hierarchical Tree Search-based User Lifelong Behavior Modeling on Large Language ModelYu Xia, Rui Zhong, Hao Gu, Wei Yang et al.SIGIR 2025 · 5 citations
- PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational SummarizationXu Sun, Lionel Delphin-Poulat, Christèle Tarnec, Anastasia ShimorinaEMNLP 2025 · 5 citations
- Exploring the Trade-Off within Visual Information for MultiModal Sentence SummarizationMinghuan Yuan, Shiyao Cui, Xinghua Zhang, Shicheng Wang et al.SIGIR 2024 · 3 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
Related papers
- Where to show Demos in Your Prompt: A Positional Bias of In-Context LearningKwesi A. Cobbina, Tianyi ZhouEMNLP 2025
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee et al.ICLR 2024 · 131 citations
- Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length ContextsYuho Lee, Jiaqi Deng, Nicole Hee-Yeon Kim, Hyangsuk Min et al.EMNLP 2025
- Never Lost in the Middle: Mastering Long-Context Question Answering with Position-Agnostic Decompositional TrainingJunqing He, Kunhao Pan, Xiaoqun Dong, Zhuoyang Song et al.ACL 2024 · 3 citations
- Towards Improving Faithfulness in Abstractive SummarizationXiuying Chen, Mingzhe Li, Xin Gao, Xiangliang ZhangNeurIPS 2022 · 39 citations
