What's Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview Dialogs
Anna Wegmann, Tijs A. van den Broek, Dong Nguyen
Abstract
Best practices for high conflict conversations like counseling or customer support almost always include recommendations to paraphrase the previous speaker. Although paraphrase classification has received widespread attention in NLP, paraphrases are usually considered independent from context, and common models and datasets are not applicable to dialog settings. In this work, we investigate paraphrases across turns in dialog (e.g., Speaker 1: "That book is mine." becomes Speaker 2: "That book is yours."). We provide an operationalization of context-dependent paraphrases, and develop a training for crowd-workers to classify paraphrases in dialog. We introduce ContextDeP, a dataset with utterance pairs from NPR and CNN news interviews annotated for contextdependent paraphrases. To enable analysis on label variation, the dataset contains 5,581 annotations on 600 utterance pairs. We present promising results with in-context learning and with token classification models for automatic paraphrase detection in dialog. What? Shortened Examples Clear Contextual Equivalence ⊆ CP Guest: I know they are cruel. Host: You know they are cruel. G: We have been the punching bag of the president. H: The president has been using Chicago as a punching bag. 1 https://github.com/nlpsoc/ Paraphrases-in-News-Interviews 2 https://huggingface.co/datasets/AnnaWegmann/ Paraphrases-in-Interviews 3 This is in line with the license from the original data publication (Zhu et al., 2021) .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6101c862-2834-439f-8a27-6a507c066d68Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Interview: Large-scale Modeling of Media Dialog with Discourse Patterns and Knowledge GroundingBodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni, Julian J. McAuleyEMNLP 2020 · 12 citations
- I like fish, especially dolphins: Addressing Contradictions in Dialogue ModelingYixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela et al.ACL 2021
- Improving Large-scale Paraphrase Acquisition and GenerationYao Dou, Chao Jiang, Wei XuEMNLP 2022 · 11 citations
- Simple Conversational Data Augmentation for Semi-supervised Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2021 · 31 citations
- Red Teaming Language Models for Processing Contradictory DialoguesXiaofei Wen, Bangzheng Li, Tenghao Huang, Muhao ChenEMNLP 2024 · 1 citation
