CEMA - Cost-Efficient Machine-Assisted Document Annotations
Guowen Yuan, Ben Kao, Tien-Hsuan Wu
Abstract
We study the problem of semantically annotating textual documents that are complex in the sense that the documents are long, feature rich, and domain specific. Due to their complexity, such annotation tasks require trained human workers, which are very expensive in both time and money. We propose CEMA, a method for deploying machine learning to assist humans in complex document annotation. CEMA estimates the human cost of annotating each document and selects the set of documents to be annotated that strike the best balance between model accuracy and human cost. We conduct experiments on complex annotation tasks in which we compare CEMA against other document selection and annotation strategies. Our results show that CEMA is the most cost-efficient solution for those tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 899bf37e-e384-4f21-be41-ee6953594bbbBuilds on1
Related papers
- Active Learning with Query Generation for Cost-Effective Text ClassificationYifan Yan, Sheng-Jun Huang, Shaoyi Chen, Meng Liao et al.AAAI 2020 · 27 citations
- COCA: Cost-Effective Collaborative Annotation System by Combining Experts and AmateursJiayu Lei, Zheng Zhang, Lan Zhang, Xiang-Yang LiICDE 2022 · 6 citations
- MCAL: Minimum Cost Human-Machine Active LabelingHang Qiu, Krishna Chintalapudi, Ramesh GovindanICLR 2023
- Learning a Cost-Effective Annotation Policy for Question AnsweringBernhard Kratzwald, Stefan Feuerriegel, Huan SunEMNLP 2020 · 9 citations
- Cost-aware LLM-based Online Dataset AnnotationEray Can Elumar, Cem Tekin, Osman YaganNeurIPS 2025 · 4 citations
