HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization
Shuyang Cao, Lu Wang
Abstract
Document structure is critical for efficient information consumption. However, it is challenging to encode it efficiently into the modern Transformer architecture. In this work, we present HIBRIDS, which injects Hierarchical Biases foR Incorporating Document Structure into the calculation of attention scores. We further present a new task, hierarchical questionsummary generation, for summarizing salient content in the source document into a hierarchy of questions and summaries, where each follow-up question inquires about the content of its parent question-summary pair. We also annotate a new dataset with 6, 153 questionsummary hierarchies labeled on long government reports. Experiment results show that our model produces better question-summary hierarchies than comparisons on both hierarchy quality and content coverage, a finding also echoed by human judges. Additionally, our model improves the generation of longform summaries from lengthy government reports and Wikipedia articles, as measured by ROUGE scores.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02a51949-26ad-44fb-b012-202b6be48ffcCited by top-tier papers7
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna et al.ICLR 2024 · 460 citations
- SQuALITY: Building a Long-Document Summarization Dataset the Hard WayAlex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang et al.EMNLP 2022 · 19 citations
- Leveraging Locality in Abstractive Text SummarizationYixin Liu, Ansong Ni, Linyong Nan, Budhaditya Deb et al.EMNLP 2022 · 19 citations
- Factorizing Content and Budget Decisions in Abstractive Summarization of Long DocumentsMarcio Fonseca, Yftah Ziser, Shay B. CohenEMNLP 2022 · 14 citations
- Incorporating Distributions of Discourse Structure for Long Document Abstractive SummarizationDongqi Liu, Yifan Wang, Vera DembergACL 2023 · 12 citations
Builds on7
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Discourse-Aware Neural Extractive Text SummarizationJiacheng Xu, Zhe Gan, Yu Cheng, Jingjing LiuACL 2020 · 264 citations
- Asking Questions the Human Way: Scalable Question-Answer Generation from Text CorpusBang Liu, Haojie Wei, Di Niu, Haolan Chen et al.WWW 2020 · 100 citations
- Evaluating the Factual Consistency of Abstractive Text SummarizationWojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard SocherEMNLP 2020 · 67 citations
Related papers
- Stepwise Extractive Summarization and Planning with Structured TransformersShashi Narayan, Joshua Maynez, Jakub Adámek, Daniele Pighin et al.EMNLP 2020 · 3 citations
- HEGEL: Hypergraph Transformer for Long Document SummarizationHaopeng Zhang, Xiao Liu, Jiawei ZhangEMNLP 2022 · 33 citations
- HiCI: Hierarchical Construction–Integration for Long-Context AttentionXiangyu Zeng, Qi Xu, Yunke Wang, Chang XuICML 2026 · 3 citations
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM EmbeddingsXueying Ding, Xingyue Huang, Mingxuan Ju, Liam Collins et al.ACL 2026 · 3 citations
- HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language GenerationZhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia et al.ACL 2022
