ARAIDA: Analogical Reasoning-Augmented Interactive Data Annotation
Chen Huang, Yiping Jin, Ilija Ilievski, Wenqiang Lei, Jiancheng Lv
Abstract
Human annotation is a time-consuming task that requires a significant amount of effort. To address this issue, interactive data annotation utilizes an annotation model to provide suggestions for humans to approve or correct. However, annotation models trained with limited labeled data are prone to generating incorrect suggestions, leading to extra human correction effort. To tackle this challenge, we propose ARAIDA, an analogical reasoning-based approach that enhances automatic annotation accuracy in the interactive data annotation setting and reduces the need for human corrections. ARAIDA involves an error-aware integration strategy that dynamically coordinates an annotation model and a k-nearest neighbors (KNN) model, giving more importance to KNN's predictions when predictions from the annotation model are deemed inaccurate. Empirical studies demonstrate that ARAIDA is adaptable to different annotation tasks and models. On average, it reduces human correction labor by 11.02% compared to vanilla interactive data annotation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d181c60-e5a8-487e-b5b0-f5af8a6fe471Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Nearest Neighbor Machine TranslationUrvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2021 · 323 citations
- Reinforced active learning for image segmentationArantxa Casanova, Pedro O. Pinheiro, Negar Rostamzadeh, Christopher J. PalICLR 2020 · 127 citations
- Cody: An AI-Based System to Semi-Automate Coding for Qualitative ResearchTim Rietz, Alexander MaedcheCHI 2021 · 67 citations
Related papers
- Rapid Image Labeling via Neuro-Symbolic LearningYifeng Wang, Zhi Tu, Yiwen Xiang, Shiyuan Zhou et al.KDD 2023 · 3 citations
- Learning a Cost-Effective Annotation Policy for Question AnsweringBernhard Kratzwald, Stefan Feuerriegel, Huan SunEMNLP 2020 · 9 citations
- PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule SynthesisSimret Araya Gebreegziabher, Zheng Zhang, Xiaohang Tang, Yihao Meng et al.CHI 2023 · 70 citations
- CalCo: A Hierarchical Bayesian Framework for Scalable Human-LLM Hybrid LabelingViet-An Nguyen, Xu Chen, Udi WeinsbergKDD 2026
- Less is More: Attention Supervision with Counterfactuals for Text ClassificationSeungtaek Choi, Haeju Park, Jinyoung Yeo, Seung-won HwangEMNLP 2020 · 16 citations
