Dissecting Generation Modes for Abstractive Summarization Models via Ablation and Attribution
Jiacheng Xu, Greg Durrett
Abstract
Despite the prominence of neural abstractive summarization models, we know little about how they actually form summaries and how to understand where their decisions come from. We propose a two-step method to interpret summarization model decisions. We first analyze the model's behavior by ablating the full model to categorize each decoder decision into one of several generation modes: roughly, is the model behaving like a language model, is it relying heavily on the input, or is it somewhere in between? After isolating decisions that do depend on the input, we explore interpreting these decisions using several different attribution methods. We compare these techniques based on their ability to select content and reconstruct the model's predicted token from perturbations of the input, thus revealing whether highlighted attributions are truly important for the generation of the next token. While this machinery can be broadly useful even beyond summarization, we specifically demonstrate its capability to identify phrases the summarization model has memorized and determine where in the training pipeline this memorization happened, as well as study complex generation phenomena like sentence fusion on a per-instance basis. Conclusion: These doc tokens impacted prediction the most (according to int. grad.) Conclusion: Higher difference means higher dependence on context Ablation mayoral Cameron 0.01 0.01 0.56 0.99 Diff between LM and full model for Khan 0.96 0.99 0.99 0.99 Input Article Speaking at a rally for Tory candidate Zac Goldsmith, the prime minister warned about the dangers of a Labour victory for the capital's economy. Mr Goldsmith said his Labour rival was "Mr Corbyn's man" in City Hall. But Mr Khan said he was "no patsy" to Mr Corbyn and […] Predicted Summary David Cameron has urged Londoners to vote for the Conservatives in the mayoral election, saying Labour's Sadiq Khan is "Jeremy Corbyn's man". BART Compare decoder-only LM ( ) with full model ( ) Attribution When context matters, use attribution to find the content supporting the decision LM-like Contextual … the prime minister warned … David Cameron Encoder Decoder next word BART Figure 1: Our two-stage ablation-attribution framework. First, we compare a decoder-only language model (not fine-tuned on summarization task, and not conditioned on the input article) and a full summarization model. They are colored in gray and orange respectively. the The higher the difference, the more heavily model depends on the input context. For those context-dependent decisions, we conduct content attribution to find the relevant supporting content with methods like Integrated Gradient or Occlusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on16
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Discourse-Aware Neural Extractive Text SummarizationJiacheng Xu, Zhe Gan, Yu Cheng, Jingjing LiuACL 2020 · 264 citations
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsXiaochuang Han, Byron C. Wallace, Yulia TsvetkovACL 2020 · 91 citations
Related papers
- SARA: Salience-Aware Reinforced Adaptive Decoding for Large Language Models in Abstractive SummarizationNayu Liu, Junnan Zhu, Yiming Ma, Zhicong Lu et al.ACL 2025 · 7 citations
- Controlling the Amount of Verbatim Copying in Abstractive SummarizationKaiqiang Song, Bingqing Wang, Zhe Feng, Ren Liu et al.AAAI 2020 · 53 citations
- Joint Parsing and Generation for Abstractive SummarizationKaiqiang Song, Logan Lebanoff, Qipeng Guo, Xipeng Qiu et al.AAAI 2020 · 29 citations
- Towards Improving Faithfulness in Abstractive SummarizationXiuying Chen, Mingzhe Li, Xin Gao, Xiangliang ZhangNeurIPS 2022 · 39 citations
- Summarization Programs: Interpretable Abstractive Summarization with Neural Modular TreesSwarnadeep Saha, Shiyue Zhang, Peter Hase, Mohit BansalICLR 2023 · 7 citations
