Towards Interpretable Object Detection by Unfolding Latent Structures
Tianfu Wu, Xi Song
Abstract
This paper first proposes a method of formulating model interpretability in visual understanding tasks based on the idea of unfolding latent structures. It then presents a case study in object detection using popular two-stage region-based convolutional network (i.e., R-CNN) detection systems. The proposed method focuses on weakly-supervised extractive rationale generation, that is learning to unfold latent discriminative part configurations of object instances automatically and simultaneously in detection without using any supervision for part configurations. It utilizes a top-down hierarchical and compositional grammar model embedded in a directed acyclic AND-OR Graph (AOG) to explore and unfold the space of latent part configurations of regions of interest (RoIs). It presents an AOGParsing operator that seamlessly integrates with the RoIPooling/RoIAlign operator widely used in R-CNN and is trained end-to-end. In object detection, a bounding box is interpreted by the best parse tree derived from the AOG on-the-fly, which is treated as the qualitatively extractive rationale generated for interpreting detection. In experiments, Faster R-CNN is used to test the proposed method on the PASCAL VOC 2007 and the COCO 2017 object detection datasets. The experimental results show that the proposed method can compute promising latent structures without hurting the performance. The code and pretrained models are available at https://github.com/iVMCL/iRCNN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0402671c-2d01-4764-877f-48b56595391eCited by top-tier papers5
- Towards Interpretable Deep Networks for Monocular Depth EstimationZunzhi You, Yi-Hsuan Tsai, Wei-Chen Chiu, Guanbin LiICCV 2021 · 19 citations
- FFAM: Feature Factorization Activation Map for Explanation of 3D DetectorsShuai Liu, Boyang Li, Zhiyu Fang, Mingyue Cui et al.NeurIPS 2024 · 4 citations
- ODAM: Gradient-based Instance-Specific Visual Explanations for Object DetectionChenyang Zhao, Antoni B. ChanICLR 2023 · 4 citations
- BBAM: Bounding Box Attribution Map for Weakly Supervised Semantic and Instance SegmentationJungbeom Lee, Jihun Yi, Chaehun Shin, Sungroh YoonCVPR 2021
- Towards White-Box Deep Wireless SensingXie Zhang, Yina Wang, Chenshu WuUbiComp 2026
Related papers
- A Probabilistic Graphical Model Based on Neural-symbolic Reasoning for Visual Relationship DetectionDongran Yu, Bo Yang, Qianhao Wei, Anchen Li et al.CVPR 2022 · 18 citations
- Robust Instance Segmentation Through Reasoning About Multi-Object OcclusionXiaoding Yuan, Adam Kortylewski, Yihong Sun, Alan L. YuilleCVPR 2021
- Relation Parsing Neural Network for Human-Object Interaction DetectionPenghao Zhou, Mingmin ChiICCV 2019 · 155 citations
- Inducing Hierarchical Compositional Model by Sparsifying Generator NetworkXianglei Xing, Tianfu Wu, Song-Chun Zhu, Ying Nian WuCVPR 2020
- Detection-Based Intermediate Supervision for Visual Question AnsweringYuhang Liu, Daowan Peng, Wei Wei, Yuanyuan Fu et al.AAAI 2024 · 3 citations
