Identifying Chinese Opinion Expressions with Extremely-Noisy Crowdsourcing Annotations
Xin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang, Xiaobin Wang, Min Zhang
Abstract
Recent works of opinion expression identification (OEI) rely heavily on the quality and scale of the manually-constructed training corpus, which could be extremely difficult to satisfy. Crowdsourcing is one practical solution for this problem, aiming to create a large-scale but quality-unguaranteed corpus. In this work, we investigate Chinese OEI with extremely-noisy crowdsourcing annotations, constructing a dataset at a very low cost. Following Zhang el al. (2021), we train the annotator-adapter model by regarding all annotations as gold-standard in terms of crowd annotators, and test the model by using a synthetic expert, which is a mixture of all annotators. As this annotator-mixture for testing is never modeled explicitly in the training phase, we propose to generate synthetic training samples by a pertinent mixup strategy to make the training and testing highly consistent. The simulation experiments on our constructed dataset show that crowdsourcing is highly promising for OEI, and our proposed annotator-mixup can further enhance the crowdsourcing modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36855042-9f2d-4a52-a806-e3cdeb21cae8Cited by top-tier papers4
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng et al.ICLR 2024 · 368 citations
- Unsupervised Entity Alignment for Temporal Knowledge GraphsXiaoze Liu, Junyang Wu, Tianyi Li, Lu Chen et al.WWW 2023 · 56 citations
- Character-level White-Box Adversarial Attacks against Transformers via Attachable Subwords SubstitutionAiwei Liu, Honghai Yu, Xuming Hu, Shu'ang Li et al.EMNLP 2022 · 20 citations
- Can Large Language Models be Effective Online Opinion Miners?Ryang Heo, Yongsik Seo, Junseong Lee, Dongha LeeEMNLP 2025 · 2 citations
Builds on2
- UDapter: Language Adaptation for Truly Universal Dependency ParsingAhmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van NoordEMNLP 2020 · 10 citations
- Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity RecognitionXin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang et al.ACL 2021
Related papers
- Chinese Opinion Role Labeling with Corpus Translation: A Pivot StudyRanran Zhen, Rui Wang, Guohong Fu, Chengguo Lv et al.EMNLP 2021 · 4 citations
- Learning from Crowds by Modeling Common ConfusionsZhendong Chu, Jing Ma, Hongning WangAAAI 2021 · 60 citations
- Coupled Confusion Correction: Learning from Crowds with Sparse AnnotationsHansong Zhang, Shikun Li, Dan Zeng, Chenggang Yan et al.AAAI 2024 · 23 citations
- A Probabilistic Graphical Model for Analyzing the Subjective Visual Quality Assessment Data from CrowdsourcingJing Li, Suiyi Ling, Junle Wang, Patrick Le CalletACM MM 2020 · 23 citations
- PerspectiveMod: A Perspectivist Resource for Deliberative ModerationEva Maria Vecchi, Neele Falk, Carlotta Quensel, Iman Jundi et al.EMNLP 2025
