Interpretable Generative Models through Post-hoc Concept Bottlenecks
Akshay R. Kulkarni, Ge Yan, Chung-En Sun, Tuomas P. Oikarinen, Tsui-Wei Weng
摘要
Concept bottleneck models (CBM) aim to produce inherently interpretable models that rely on human-understandable concepts for their predictions. However, existing approaches to design interpretable generative models based on CBMs are not yet efficient and scalable, as they require expensive generative model training from scratch as well as real images with labor-intensive concept supervision. To address these challenges, we present two novel and low-cost methods to build interpretable generative models through post-hoc techniques and we name our approaches: concept-bottleneck autoencoder (CB-AE) and concept controller (CC). Our proposed approaches enable efficient and scalable training without the need of real data and require only minimal to no concept supervision. Additionally, our methods generalize across modern generative model families including generative adversarial networks and diffusion models. We demonstrate the superior interpretability and steerability of our methods on numerous standard datasets like CelebA, CelebA-HQ, and CUB with large improvements (average ∼25%) over the prior work, while being 4-15× faster to train. Finally, a large-scale user study is performed to validate the interpretability and steerability of our methods. * Equal contribution 1 Code: github.com/Trustworthy-ML-Lab/posthoc-generative-cbm Efficiently train only CB-AE/CC with frozen generative model Concept Controller noise gen. image concept Concept Bottleneck AE B. Ours (Post-Hoc Interpretable Generative Models) noise Concept Bottleneck Layer
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Uncovering Conceptual Blindspots in Generative Image Models Using Sparse AutoencodersMatyas Bohacek, Thomas Fel, Maneesh Agrawala, Ekdeep Singh LubanaICLR 2026 · 被引用 7 次
- Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned InterventionsAda Görgün, Fawaz Sammani, Nikos Deligiannis, Bernt Schiele 等ICLR 2026 · 被引用 7 次
- Interpretable and Steerable Concept Bottleneck Sparse AutoencodersAkshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy, Shusen Liu 等CVPR 2026 · 被引用 6 次
- A Probabilistic Hard Concept Bottleneck for Steerable Generative ModelsMaría Martínez-García, Ricardo Vazquez Alvarez, Alejandro Lancho, Pablo M. Olmos 等ICLR 2026
- TRANSPORTER: Transferring Visual Semantics from VLM ManifoldsAlexandros StergiouCVPR 2026
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Concept Bottleneck Generative ModelsAya Abdelsalam Ismail, Julius Adebayo, Héctor Corrada Bravo, Stephen Ra 等ICLR 2024 · 被引用 43 次
- Label-free Concept Bottleneck ModelsTuomas P. Oikarinen, Subhro Das, Lam M. Nguyen, Tsui-Wei WengICLR 2023 · 被引用 17 次
- Post-hoc Concept Bottleneck ModelsMert Yüksekgönül, Maggie Wang, James ZouICLR 2023 · 被引用 37 次
- Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image ClassificationYue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin 等CVPR 2023
- Learning Concept Bottleneck Models from Mechanistic ExplanationsAntonio De Santis, Schrasing Tong, Marco Brambilla, Lalana KagalICLR 2026 · 被引用 6 次
