KaLOS finds Consensus: A Meta-Algorithm for Evaluating Inter-Annotator Agreement in Complex Vision Tasks
David Tschirschwitz, Volker Rodehorst
Abstract
Progress in object detection benchmarks is stagnating. It is limited not by architectures but by the inability to distinguish model improvements from label noise. To restore trust in benchmarking the field requires rigorous quantification of annotation consistency to ensure the reliability of evaluation data. However, standard statistical metrics fail to handle the instance correspondence problem inherent to vision tasks. Furthermore, validating new agreement metrics remains circular because no objective ground truth for agreement exists. This forces reliance on unverifiable heuristics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be21ad8c-b236-42a2-8875-a9d6a0765fdcBuilds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- DETRs with Collaborative Hybrid Assignments TrainingZhuofan Zong, Guanglu Song, Yu LiuICCV 2023 · 594 citations
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 482 citations
- Towards Robust Adaptive Object Detection under Noisy AnnotationsXinyu Liu, Wuyang Li, Qiushi Yang, Baopu Li et al.CVPR 2022 · 40 citations
- Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation TasksAlexander Braylan, Omar Alonso, Matthew LeaseWWW 2022 · 34 citations
Related papers
- A Theory of Dynamic BenchmarksAli Shirali, Rediet Abebe, Moritz HardtICLR 2023 · 1 citation
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei et al.ICLR 2020 · 11 citations
- Pixel-level Quality Assessment for Oriented Object DetectionYunhui Zhu, Buliao HuangAAAI 2026
- Formally Exploring Visual Anomaly Detection Evaluation MetricsNasar Iqbal, Dennis Wagner, Philipp Liznerski, Nabeel Hussain Syed et al.ICML 2026
- Are We Overconfident in Models and Results for Semi-Supervised 3D Medical Image Segmentation?Jun Li, ZIWEI QINICML 2026
