CEMA - Cost-Efficient Machine-Assisted Document Annotations
Guowen Yuan, Ben Kao, Tien-Hsuan Wu
摘要
We study the problem of semantically annotating textual documents that are complex in the sense that the documents are long, feature rich, and domain specific. Due to their complexity, such annotation tasks require trained human workers, which are very expensive in both time and money. We propose CEMA, a method for deploying machine learning to assist humans in complex document annotation. CEMA estimates the human cost of annotating each document and selects the set of documents to be annotated that strike the best balance between model accuracy and human cost. We conduct experiments on complex annotation tasks in which we compare CEMA against other document selection and annotation strategies. Our results show that CEMA is the most cost-efficient solution for those tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Active Learning with Query Generation for Cost-Effective Text ClassificationYifan Yan, Sheng-Jun Huang, Shaoyi Chen, Meng Liao 等AAAI 2020 · 被引用 27 次
- COCA: Cost-Effective Collaborative Annotation System by Combining Experts and AmateursJiayu Lei, Zheng Zhang, Lan Zhang, Xiang-Yang LiICDE 2022 · 被引用 6 次
- MCAL: Minimum Cost Human-Machine Active LabelingHang Qiu, Krishna Chintalapudi, Ramesh GovindanICLR 2023
- Learning a Cost-Effective Annotation Policy for Question AnsweringBernhard Kratzwald, Stefan Feuerriegel, Huan SunEMNLP 2020 · 被引用 9 次
- Cost-aware LLM-based Online Dataset AnnotationEray Can Elumar, Cem Tekin, Osman YaganNeurIPS 2025 · 被引用 4 次
