Attribution Analysis-based Concept Alignment: A Human-in-the-loop Data Debugging Framework
Lei Chai, Lu Qi, Hailong Sun, Jing Zhang, Jingxuan Xu
Abstract
Ensuring consistently high-quality training data is essential for developing reliable machine learning systems. Recent research demonstrates that incorporating human supervision into training set debugging effectively improves model performance, especially for text classification tasks. However, such methods often prove inapplicable to image understanding tasks, where inherently unstructured pixel data presents challenges in understanding and correcting biases. Inspired by human-AI alignment, we introduce AACA (Attribution Analysis-based Concept Alignment), a human-in-the-loop framework that mitigates bias in the training set by aligning the concepts used by humans and AI during the decision-making process. Specifically, AACA comprises two primary stages: interpretable data bug discovery and targeted data augmentation. During the data bug discovery stage, AACA identifies confounded and valid concepts to explain why prediction failure occurs and what concept the model should focus, using interpretability methods and human annotation. In the stage of targeted data augmentation, AACA adopts these concept-level attributions as clues to synthesize debugging instances via text-to-image generative model. The initial model is then retrained on the augmented set to correct prediction failures. Comparative experiments conducted on crowdsourced annotations and real-world datasets demonstrate that AACA can accurately identifies data bugs and effectively repairs prediction failures, thereby significantly improving prediction performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdeb72af-afc7-40aa-890e-3c2a8365e66dBuilds on12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- DeepAccident: A Motion and Accident Prediction Benchmark for V2X Autonomous DrivingTianqi Wang, Sukmin Kim, Wenxuan Ji, Enze Xie et al.AAAI 2024 · 132 citations
- Training Data Debugging for the Fairness of Machine Learning SoftwareYanhui Li, Linghan Meng, Lin Chen, Li Yu et al.ICSE 2022 · 49 citations
- Missingness Bias in Model DebuggingSaachi Jain, Hadi Salman, Eric Wong, Pengchuan Zhang et al.ICLR 2022 · 45 citations
- Refining Language Models with Compositional ExplanationsHuihan Yao, Ying Chen, Qinyuan Ye, Xisen Jin et al.NeurIPS 2021 · 39 citations
Related papers
- RA3: A Human-in-the-loop Framework for Interpreting and Improving Image Captioning with Relation-Aware Attribution AnalysisLei Chai, Lu Qi, Hailong Sun, Jingzheng LiICDE 2024
- Debugging Concept Bottleneck Models through Removal and RetrainingEric Enouen, Sainyam GalhotraICLR 2026 · 2 citations
- ESSA: Explanation Iterative Supervision via Saliency-guided Data AugmentationSiyi Gu, Yifei Zhang, Yuyang Gao, Xiaofeng Yang et al.KDD 2023 · 7 citations
- Towards Trustable Skin Cancer Diagnosis via Rewriting Model's DecisionSiyuan Yan, Zhen Yu, Xuelin Zhang, Dwarikanath Mahapatra et al.CVPR 2023
- HiBug: On Human-Interpretable Model DebugMuxi Chen, Yu Li, Qiang XuNeurIPS 2023 · 22 citations
