I Wish I Would Have Loved This One, But I Didn't - A Multilingual Dataset for Counterfactual Detection in Product Review
James O'Neill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota, Danushka Bollegala
摘要
Counterfactual statements describe events that did not or cannot take place. We consider the problem of counterfactual detection (CFD) in product reviews. For this purpose, we annotate a multilingual CFD dataset from Amazon product reviews covering counterfactual statements written in English, German, and Japanese languages. The dataset is unique as it contains counterfactuals in multiple languages, covers a new application area of ecommerce reviews, and provides high quality professional annotations. We train CFD models using different text representation methods and classifiers. We find that these models are robust against the selectional biases introduced due to cue phrase-based sentence selection. Moreover, our CFD dataset is compatible with prior datasets and can be merged to learn accurate CFD models. Applying machine translation on English counterfactual examples to create multilingual data performs poorly, demonstrating the language-specificity of this problem, which has been ignored so far.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding ModelXinping Zhao, Xinshuo Hu, Zifei Shan, Shouzheng Huang 等ICLR 2026 · 被引用 47 次
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos 等ICLR 2025 · 被引用 10 次
- AdaSent: Efficient Domain-Adapted Sentence Embeddings for Few-Shot ClassificationYongxin Huang, Kexin Wang, Sourav Dutta, Raj Nath Patel 等EMNLP 2023 · 被引用 3 次
- Uncovering the Latent Potential of Deep Intermediate RepresentationsArnesh Batra, Arush Gumber, Aniket Khandelwal, Jashn Khemani 等ICML 2026 · 被引用 1 次
- Don't Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space TransformationYingchaojie Feng, Yiqun Sun, Yandong Sun, Minfeng Zhu 等ACL 2025
它引用的顶会 Paper4
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Interpreting Pretrained Contextualized Representations via Reductions to Static EmbeddingsRishi Bommasani, Kelly Davis, Claire CardieACL 2020 · 被引用 137 次
相关 Paper
- HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine TranslationDavid Dale, Elena Voita, Janice Lam, Prangthip Hansanti 等EMNLP 2023 · 被引用 8 次
- Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example GenerationQianli Wang, Van Bach Nguyen, Yihong Liu, Fedor Splitt 等ACL 2026
- ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in VideosPatrick Giedemann, Pius von Däniken, Jan Milan Deriu, Álvaro Rodrigo 等EMNLP 2025 · 被引用 1 次
- Cross-Lingual Low-Resource Set-to-Description Retrieval for Global E-CommerceJuntao Li, Chang Liu, Jian Wang, Lidong Bing 等AAAI 2020 · 被引用 14 次
- MAD-TSC: A Multilingual Aligned News Dataset for Target-dependent Sentiment ClassificationEvan Dufraisse, Adrian Popescu, Julien Tourille, Armelle Brun 等ACL 2023 · 被引用 3 次
