EmailSum: Abstractive Email Thread Summarization
Shiyue Zhang, Asli Celikyilmaz, Jianfeng Gao, Mohit Bansal
摘要
Recent years have brought about an interest in the challenging task of summarizing conversation threads (meetings, online discussions, etc.). Such summaries help analysis of the long text to quickly catch up with the decisions made and thus improve our work or communication efficiency. To spur research in thread summarization, we have developed an abstractive Email Thread Summarization (EMAILSUM) dataset, which contains humanannotated short (<30 words) and long (<100 words) summaries of 2,549 email threads (each containing 3 to 10 emails) over a wide variety of topics. We perform a comprehensive empirical study to explore different summarization techniques (including extractive and abstractive methods, single-document and hierarchical models, as well as transfer and semisupervised learning) and conduct human evaluations on both short and long summary generation tasks. Our results reveal the key challenges of current abstractive summarization models in this task, such as understanding the sender's intent and identifying the roles of sender and receiver. Furthermore, we find that widely used automatic evaluation metrics (ROUGE, BERTScore) are weakly correlated with human judgments on this email thread summarization task. Hence, we emphasize the importance of human evaluation and the development of better metrics by the community. 1 Table 8: Examples of high-quality summaries generated by model. Emails are separated by '|||' and some content are omit by '...'. (salience=xx, faithfulness=xx) gives the average human rating for that summary.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error DetectorsLiyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban 等ACL 2023 · 被引用 38 次
- Dialogue Summarization with Static-Dynamic Structure Fusion GraphShen Gao, Xin Cheng, Mingzhe Li, Xiuying Chen 等ACL 2023 · 被引用 14 次
- Summarizing Community-based Question-Answer PairsTing-Yao Hsu, Yoshi Suhara, Xiaolan WangEMNLP 2022 · 被引用 5 次
- DIONYSUS: A Pre-trained Model for Low-Resource Dialogue SummarizationYu Li, Baolin Peng, Pengcheng He, Michel Galley 等ACL 2023 · 被引用 4 次
- CHOIR: A Chatbot-mediated Organizational Memory Leveraging Communication in University Research LabsSangwook Lee, Adnan Abbas, Yan Chen, Young-Ho Kim 等CHI 2026 · 被引用 3 次
它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 被引用 317 次
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 被引用 90 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
- Storytelling with Dialogue: A Critical Role Dungeons and Dragons DatasetRevanth Rameshkumar, Peter BaileyACL 2020 · 被引用 37 次
相关 Paper
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu 等EMNLP 2020 · 被引用 3 次
- ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument MiningAlexander R. Fabbri, Faiaz Rahman, Imad Rizvi, Borui Wang 等ACL 2021
- How Far are We from Robust Long Abstractive Summarization?Huan Yee Koh, Jiaxin Ju, He Zhang, Ming Liu 等EMNLP 2022 · 被引用 16 次
- CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1, 500+ Language PairsAbhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Yuan-Fang Li 等ACL 2023 · 被引用 23 次
- MailEx: Email Event and Argument ExtractionSaurabh Srivastava, Gaurav Singh, Shou Matsumoto, Ali K. Raz 等EMNLP 2023 · 被引用 3 次
