ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue Summarization
Xiutian Zhao, Ke Wang, Wei Peng
Abstract
Dialogue agents have been receiving increasing attention for years, and this trend has been further boosted by the recent progress of large language models (LLMs). Stance detection and dialogue summarization are two core tasks of dialogue agents in application scenarios that involve argumentative dialogues. However, research on these tasks is limited by the insufficiency of public datasets, especially for non-English languages. To address this language resource gap in Chinese, we present OR-CHID (Oral Chinese Debate), the first Chinese dataset for benchmarking target-independent stance detection and debate summarization. Our dataset consists of 1,218 real-world debates that were conducted in Chinese on 476 unique topics, containing 2,436 stance-specific summaries and 14,133 fully annotated utterances. Besides providing a versatile testbed for future research, we also conduct an empirical study on the dataset and propose an integrated task. The results show the challenging nature of the dataset and suggest a potential of incorporating stance detection in summarization for argumentative dialogue. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- DialogLM: Pre-trained Model for Long Dialogue Understanding and SummarizationMing Zhong, Yang Liu, Yichong Xu, Chenguang Zhu et al.AAAI 2022 · 150 citations
- Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic ModelingYicheng Zou, Lujun Zhao, Yangyang Kang, Jun Lin et al.AAAI 2021 · 63 citations
- CSDS: A Fine-Grained Chinese Dataset for Customer Service Dialogue SummarizationHaitao Lin, Liqun Ma, Junnan Zhu, Lu Xiang et al.EMNLP 2021 · 21 citations
- Zero-Shot Stance Detection: A Dataset and Model using Generalized Topic RepresentationsEmily Allaway, Kathleen R. McKeownEMNLP 2020 · 7 citations
- IAM: A Comprehensive and Large-Scale Dataset for Integrated Argument Mining TasksLiying Cheng, Lidong Bing, Ruidan He, Qian Yu et al.ACL 2022
Related papers
- C-STANCE: A Large Dataset for Chinese Zero-Shot Stance DetectionChenye Zhao, Yingjie Li, Cornelia CarageaACL 2023 · 12 citations
- Towards Multi-dimensional Evaluation of LLM Summarization across Domains and LanguagesHyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng et al.ACL 2025 · 8 citations
- Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining DatasetsBenjamin Schiller, Johannes Daxenberger, Andreas Waldis, Iryna GurevychEMNLP 2024 · 3 citations
- Bilingual Zero-Shot Stance DetectionChenye Zhao, Cornelia CarageaACL 2025 · 1 citation
- Towards Understanding Omission in Dialogue SummarizationYicheng Zou, Kaitao Song, Xu Tan, Zhongkai Fu et al.ACL 2023 · 3 citations
