Lune

ACL2023Top-tier venue

Annotating and Detecting Fine-grained Factual Errors for Dialogue Summarization

Rongxin Zhu, Jianzhong Qi, Jey Han Lau

2023Year
1Citations
5Top-tier citations

Abstract

A series of datasets and models have been proposed for summaries generated for wellformatted documents such as news articles. Dialogue summaries, however, have been under explored. In this paper, we present the first dataset with fine-grained factual error annotations named DIASUMFACT. We define finegrained factual error detection as a sentencelevel multi-label classification problem, and we evaluate two state-of-the-art (SOTA) models on our dataset. Both models yield sub-optimal results, with a macro-averaged F1 score of around 0.25 over 6 error classes. We further propose an unsupervised model ENDERANKER via candidate ranking using pretrained encoder-decoder models. Our model performs on par with the SOTA models while requiring fewer resources. These observations confirm the challenges in detecting factual errors from dialogue summaries, which call for further studies, for which our dataset and results offer a solid foundation. 1 Lilly: Wanna go out tonight? Marshall: can't :( money's low Lilly: my treat :) Marshall: I wouldn't let a woman pay for me.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers5

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines