ACCENT: An Automatic Event Commonsense Evaluation Metric for Open-Domain Dialogue Systems
Sarik Ghazarian, Yijia Shao, Rujun Han, Aram Galstyan, Nanyun Peng
摘要
Commonsense reasoning is omnipresent in human communications and thus is an important feature for open-domain dialogue systems. However, evaluating commonsense in dialogue systems is still an open challenge. We take the first step by focusing on event commonsense that considers events and their relations, and is crucial in both dialogues and general commonsense reasoning. We propose AC-CENT, an event commonsense evaluation metric empowered by commonsense knowledge bases (CSKBs). ACCENT first extracts eventrelation tuples from a dialogue, and then evaluates the response by scoring the tuples in terms of their compatibility with the CSKB. To evaluate ACCENT, we construct the first public event commonsense evaluation dataset for open-domain dialogues. Our experiments show that ACCENT is an efficient metric for event commonsense evaluation, which achieves higher correlations with human judgments than existing baselines. * Equal contribution † The work was done while the author was conducting a summer internship at UCLA. ‡ The collaboration started when the author was a graduate student at USC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Open-Domain Text Evaluation via Contrastive Distribution MethodsSidi Lu, Hongyi Liu, Asli Celikyilmaz, Tianlu Wang 等ICML 2024 · 被引用 2 次
- Harnessing Black-Box Control to Boost Commonsense in LM's GenerationYufei Tian, Felix Zhang, Nanyun PengEMNLP 2023
它引用的顶会 Paper10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da 等AAAI 2021 · 被引用 458 次
- ASER: A Large-scale Eventuality Knowledge GraphHongming Zhang, Xin Liu, Haojie Pan, Yangqiu Song 等WWW 2020 · 被引用 183 次
相关 Paper
- Benchmarking Commonsense Knowledge Base Population with an Effective Evaluation DatasetTianqing Fang, Weiqi Wang, Sehyun Choi, Shibo Hao 等EMNLP 2021 · 被引用 17 次
- CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language ModelsDan Shi, Chaobin You, Jiantao Huang, Taihao Li 等AAAI 2024 · 被引用 3 次
- CICERO: A Dataset for Contextualized Commonsense Inference in DialoguesDeepanway Ghosal, Siqi Shen, Navonil Majumder, Rada Mihalcea 等ACL 2022
- Dialogue-Based Relation ExtractionDian Yu, Kai Sun, Claire Cardie, Dong YuACL 2020 · 被引用 106 次
- TIMEDIAL: Temporal Commonsense Reasoning in DialogLianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He 等ACL 2021
