Cascading Biases: Investigating the Effect of Heuristic Annotation Strategies on Data and Models
Chaitanya Malaviya, Sudeep Bhatia, Mark Yatskar
Abstract
Cognitive psychologists have documented that humans use cognitive heuristics, or mental shortcuts, to make quick decisions while expending less effort. While performing annotation work on crowdsourcing platforms, we hypothesize that such heuristic use among annotators cascades on to data quality and model robustness. In this work, we study cognitive heuristic use in the context of annotating multiple-choice reading comprehension datasets. We propose tracking annotator heuristic traces, where we tangibly measure low-effort annotation strategies that could indicate usage of various cognitive heuristics. We find evidence that annotators might be using multiple such heuristics, based on correlations with a battery of psychological tests. Importantly, heuristic use among annotators determines data quality along several dimensions: (1) known biased models, such as partial input models, more easily solve examples authored by annotators that rate highly on heuristic use, (2) models trained on annotators scoring highly on heuristic use don't generalize as well, and (3) heuristic-seeking annotators tend to create qualitatively less challenging examples. Our findings suggest that tracking heuristic usage among annotators can potentially help with collecting challenging datasets and diagnosing model biases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 332c4f4a-929f-4054-9da2-28e08a7155aeCited by top-tier papers3
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- Whose Facts Win? LLM Source Preferences under Knowledge ConflictsJakob Schuster, Vagrant Gautam, Katja MarkertACL 2026 · 3 citations
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
Builds on10
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 99 citations
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsXiaochuang Han, Byron C. Wallace, Yulia TsvetkovACL 2020 · 91 citations
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters et al.EMNLP 2021 · 72 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
Related papers
- Studying the Effects of Cognitive Biases in Evaluation of Conversational AgentsSashank Santhanam, Alireza Karduni, Samira ShaikhCHI 2020 · 18 citations
- Which Shortcut Solution Do Question Answering Models Prefer to Learn?Kazutoshi Shinoda, Saku Sugawara, Akiko AizawaAAAI 2023 · 10 citations
- Method for Appropriating the Brief Implicit Association Test to Elicit Biases in UsersTilman Dingler, Benjamin Tag, David A. Eccles, Niels van Berkel et al.CHI 2022 · 11 citations
- Improving User Behavior Prediction: Leveraging Annotator Metadata in Supervised Machine Learning ModelsLynnette Hui Xian Ng, Kokil Jaidka, Kai Yuan Tay, Niyati ChhayaCSCW 2025
- Who's in the Crowd Matters: Cognitive Factors and Beliefs Predict Misinformation Assessment AccuracyRobert A. Kaufman, Michael Robert Haupt, Steven P. DowCSCW 2022 · 27 citations
