CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China
Bolun Sun, Charles Chang, Yuen Yuen Ang, Ruotong Mu, Yuchen Xu, Zhengxin Zhang, Pingxu Hao
摘要
We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives annotated with a fivecolor typology of policy signals, capturing clarity and ambiguity, grounded in the theory of adaptive policy communication. Spanning 1949-2023, this corpus includes laws, regulations, and rules issued by Chinese central authorities, segmented into 3.3 million paragraph units. We further propose and validate an expert-directed LLM annotation method that integrates codebook design, structured training, a two-step workflow, and LLM-based scaling. Alongside the corpus, we release metadata and a gold-standard labeled set developed by trained coders. Inter-annotator agreement achieves a Fleiss' kappa of κ = 0.86 on directive labels, indicating high reliability. We provide baseline classification results with several large language models (LLMs), together with our codebook, and describe patterns from the data. This release enables downstream tasks and multilingual NLP research in communication strategies under complexity and uncertainty.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi 等EMNLP 2024 · 被引用 119 次
- Empowering Users in Digital Privacy Management through Interactive LLM-Based AgentsBolun Sun, Yifan Zhou, Haiyun JiangICLR 2025
相关 Paper
- Know Your Place: Diagnosing Implicit Social Adaptation Failures in Chinese Large Language ModelsYu Tian, Jie Xing, Ziming Li, Jiang Li 等ACL 2026
- CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language ModelsDan Shi, Chaobin You, Jiantao Huang, Taihao Li 等AAAI 2024 · 被引用 3 次
- "Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online ReviewsRuyuan Wan, Changye Li, Ting-Hao 'Kenneth' HuangACL 2026
- A Fine-grained Chinese Software Privacy Policy Dataset for Sequence Labeling and Regulation Compliant IdentificationKaifa Zhao, Le Yu, Shiyao Zhou, Jing Li 等EMNLP 2022 · 被引用 7 次
- AmbigNLG: Addressing Task Ambiguity in Instruction for NLGAyana Niwa, Hayate IsoEMNLP 2024 · 被引用 2 次
