ARTICLE: Annotator Reliability Through In-Context Learning
Sujan Dutta, Deepak Pandita, Tharindu Cyril Weerasooriya, Marcos Zampieri, Christopher M. Homan, Ashiqur R. KhudaBukhsh
Abstract
This paper discusses and contains content that is offensive or disturbing. Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to distinguish disagreement due to poor work from that due to differences of opinions between sincere annotators. With the goal of increasing diverse perspectives in annotation while ensuring consistency, we propose ARTICLE, an in-context learning (ICL) framework to estimate annotation quality through self-consistency. We evaluate this framework on two offensive speech datasets using multiple LLMs and compare its performance with traditional methods. Our findings indicate that ARTICLE can be used as a robust method for identifying reliable annotators, hence improving data quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ea1e636-73f3-4fed-bb2d-bf525611f3e6Cited by top-tier papers3
- What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency in Closed Domains Via Adversarial NudgeArka Dutta, Sujan Dutta, Rijul Magu, Soumyajit Datta et al.ACL 2026
- Counterfactual-based Cognitive Alignment In-Context Learning for Relation ExtractionQibin Li, Shengyuan Bai, Nai Zhou, Nianmin YaoAAAI 2026
- Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope SpeechJonathan Pofcher, Christopher M. Homan, Randall Sell, Ashiqur R. KhudaBukhshEMNLP 2025
Builds on8
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?Rishi Bommasani, Kathleen A. Creel, Ananya Kumar, Dan Jurafsky et al.NeurIPS 2022 · 179 citations
- What Can We Learn from Collective Human Opinions on Natural Language Inference Data?Yixin Nie, Xiang Zhou, Mohit BansalEMNLP 2020 · 77 citations
- If in a Crowdsourced Data Annotation Pipeline, a GPT-4Zeyu He, Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Shaurya Rohatgi et al.CHI 2024 · 31 citations
Related papers
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini et al.EMNLP 2021 · 2 citations
- Data Quality Matters: A Case Study of Obsolete Comment DetectionShengbin Xu, Yuan Yao, Feng Xu, Tianxiao Gu et al.ICSE 2023 · 7 citations
- Quality Matters: Evaluating Synthetic Data for Tool-Using LLMsShadi Iskander, Sofia Tolmach, Ori Shapira, Nachshon Cohen et al.EMNLP 2024 · 2 citations
- From Granular Grief to Binary Belief: A Collaborative Optimization of Annotation Techniques for Anti-Autistic LanguageNaba Rizvi, Alexis Morales Flores, Mohammad Rizvi, Nedjma Ousidhoum et al.CSCW 2025 · 2 citations
- Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveTharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri et al.EMNLP 2023 · 12 citations
