Attributable and Scalable Opinion Summarization
Tom Hosking, Hao Tang, Mirella Lapata
Abstract
We propose a method for unsupervised opinion summarization that encodes sentences from customer reviews into a hierarchical discrete latent space, then identifies common opinions based on the frequency of their encodings. We are able to generate both abstractive summaries by decoding these frequent encodings, and extractive summaries by selecting the sentences assigned to the same frequent encodings. Our method is attributable, because the model identifies sentences used to generate the summary as part of the summarization process. It scales easily to many hundreds of input reviews, because aggregation is performed in the latent space rather than over long sequences of tokens. We also demonstrate that our appraoch enables a degree of control, generating aspect-specific summaries by restricting the model to parts of the encoding space that correspond to desired aspects (e.g., location or food). Automatic and human evaluation on two datasets from different domains demonstrates that our method generates summaries that are more informative than prior work and better grounded in the input reviews.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1984945e-2863-4390-96aa-bbb0a0fd9f05Cited by top-tier papers2
- Human Feedback is not Gold StandardTom Hosking, Phil Blunsom, Max BartoloICLR 2024 · 96 citations
- CoNewsReader: Supporting Comprehensive Understanding and Raising Critical Thoughts on Social Media News Through CommentsKangyu Yuan, Guanzheng Chen, Sizhe Liang, Hehai Lin et al.CSCW 2026
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationMax Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel et al.EMNLP 2021 · 68 citations
- Learning Opinion Summarizers by Selecting Informative ReviewsArthur Brazinskas, Mirella Lapata, Ivan TitovEMNLP 2021 · 26 citations
Related papers
- Unsupervised Opinion Summarization as Copycat-Review GenerationArthur Brazinskas, Mirella Lapata, Ivan TitovACL 2020 · 14 citations
- Unsupervised Extractive Opinion Summarization Using Sparse CodingSomnath Basu Roy Chowdhury, Chao Zhao, Snigdha ChaturvediACL 2022
- Aspect-Controllable Opinion SummarizationReinald Kim Amplayo, Stefanos Angelidis, Mirella LapataEMNLP 2021
- Few-Shot Learning for Opinion SummarizationArthur Brazinskas, Mirella Lapata, Ivan TitovEMNLP 2020 · 4 citations
- Cone: Unsupervised Contrastive Opinion ExtractionRuncong Zhao, Lin Gui, Yulan HeSIGIR 2023 · 1 citation
