EmailSum: Abstractive Email Thread Summarization
Shiyue Zhang, Asli Celikyilmaz, Jianfeng Gao, Mohit Bansal
Abstract
Recent years have brought about an interest in the challenging task of summarizing conversation threads (meetings, online discussions, etc.). Such summaries help analysis of the long text to quickly catch up with the decisions made and thus improve our work or communication efficiency. To spur research in thread summarization, we have developed an abstractive Email Thread Summarization (EMAILSUM) dataset, which contains humanannotated short (<30 words) and long (<100 words) summaries of 2,549 email threads (each containing 3 to 10 emails) over a wide variety of topics. We perform a comprehensive empirical study to explore different summarization techniques (including extractive and abstractive methods, single-document and hierarchical models, as well as transfer and semisupervised learning) and conduct human evaluations on both short and long summary generation tasks. Our results reveal the key challenges of current abstractive summarization models in this task, such as understanding the sender's intent and identifying the roles of sender and receiver. Furthermore, we find that widely used automatic evaluation metrics (ROUGE, BERTScore) are weakly correlated with human judgments on this email thread summarization task. Hence, we emphasize the importance of human evaluation and the development of better metrics by the community. 1 Table 8: Examples of high-quality summaries generated by model. Emails are separated by '|||' and some content are omit by '...'. (salience=xx, faithfulness=xx) gives the average human rating for that summary.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a100e7a-8910-4133-b583-ca13ddf4c0e5Cited by top-tier papers7
- Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error DetectorsLiyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban et al.ACL 2023 · 38 citations
- Dialogue Summarization with Static-Dynamic Structure Fusion GraphShen Gao, Xin Cheng, Mingzhe Li, Xiuying Chen et al.ACL 2023 · 14 citations
- Summarizing Community-based Question-Answer PairsTing-Yao Hsu, Yoshi Suhara, Xiaolan WangEMNLP 2022 · 5 citations
- DIONYSUS: A Pre-trained Model for Low-Resource Dialogue SummarizationYu Li, Baolin Peng, Pengcheng He, Michel Galley et al.ACL 2023 · 4 citations
- CHOIR: A Chatbot-mediated Organizational Memory Leveraging Communication in University Research LabsSangwook Lee, Adnan Abbas, Yan Chen, Young-Ho Kim et al.CHI 2026 · 3 citations
Builds on6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 90 citations
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
- Storytelling with Dialogue: A Critical Role Dungeons and Dragons DatasetRevanth Rameshkumar, Peter BaileyACL 2020 · 37 citations
Related papers
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu et al.EMNLP 2020 · 3 citations
- ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument MiningAlexander R. Fabbri, Faiaz Rahman, Imad Rizvi, Borui Wang et al.ACL 2021
- How Far are We from Robust Long Abstractive Summarization?Huan Yee Koh, Jiaxin Ju, He Zhang, Ming Liu et al.EMNLP 2022 · 16 citations
- CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1, 500+ Language PairsAbhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Yuan-Fang Li et al.ACL 2023 · 23 citations
- MailEx: Email Event and Argument ExtractionSaurabh Srivastava, Gaurav Singh, Shou Matsumoto, Ali K. Raz et al.EMNLP 2023 · 3 citations
