Attribution Analysis-based Concept Alignment: A Human-in-the-loop Data Debugging Framework
Lei Chai, Lu Qi, Hailong Sun, Jing Zhang, Jingxuan Xu
摘要
Ensuring consistently high-quality training data is essential for developing reliable machine learning systems. Recent research demonstrates that incorporating human supervision into training set debugging effectively improves model performance, especially for text classification tasks. However, such methods often prove inapplicable to image understanding tasks, where inherently unstructured pixel data presents challenges in understanding and correcting biases. Inspired by human-AI alignment, we introduce AACA (Attribution Analysis-based Concept Alignment), a human-in-the-loop framework that mitigates bias in the training set by aligning the concepts used by humans and AI during the decision-making process. Specifically, AACA comprises two primary stages: interpretable data bug discovery and targeted data augmentation. During the data bug discovery stage, AACA identifies confounded and valid concepts to explain why prediction failure occurs and what concept the model should focus, using interpretability methods and human annotation. In the stage of targeted data augmentation, AACA adopts these concept-level attributions as clues to synthesize debugging instances via text-to-image generative model. The initial model is then retrained on the augmented set to correct prediction failures. Comparative experiments conducted on crowdsourced annotations and real-world datasets demonstrate that AACA can accurately identifies data bugs and effectively repairs prediction failures, thereby significantly improving prediction performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- DeepAccident: A Motion and Accident Prediction Benchmark for V2X Autonomous DrivingTianqi Wang, Sukmin Kim, Wenxuan Ji, Enze Xie 等AAAI 2024 · 被引用 132 次
- Training Data Debugging for the Fairness of Machine Learning SoftwareYanhui Li, Linghan Meng, Lin Chen, Li Yu 等ICSE 2022 · 被引用 49 次
- Missingness Bias in Model DebuggingSaachi Jain, Hadi Salman, Eric Wong, Pengchuan Zhang 等ICLR 2022 · 被引用 45 次
- Refining Language Models with Compositional ExplanationsHuihan Yao, Ying Chen, Qinyuan Ye, Xisen Jin 等NeurIPS 2021 · 被引用 39 次
相关 Paper
- RA3: A Human-in-the-loop Framework for Interpreting and Improving Image Captioning with Relation-Aware Attribution AnalysisLei Chai, Lu Qi, Hailong Sun, Jingzheng LiICDE 2024
- Debugging Concept Bottleneck Models through Removal and RetrainingEric Enouen, Sainyam GalhotraICLR 2026 · 被引用 2 次
- ESSA: Explanation Iterative Supervision via Saliency-guided Data AugmentationSiyi Gu, Yifei Zhang, Yuyang Gao, Xiaofeng Yang 等KDD 2023 · 被引用 7 次
- Towards Trustable Skin Cancer Diagnosis via Rewriting Model's DecisionSiyuan Yan, Zhen Yu, Xuelin Zhang, Dwarikanath Mahapatra 等CVPR 2023
- HiBug: On Human-Interpretable Model DebugMuxi Chen, Yu Li, Qiang XuNeurIPS 2023 · 被引用 22 次
