Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
Youngjun Jun, Seil Kang, Woojung Han, Seong Jae Hwang
Abstract
Video Diffusion Transformers (DiTs) have been synthesizing high-quality video with high fidelity to text descriptions involving motion. However, the understanding of how Video DiTs convert motion words into video remains lagging behind. Furthermore, prior studies on interpretable saliency maps primarily target objects, leaving it behind to observe how Video DiTs behave with respect to motion. In this paper, we inquire into concrete motion features that specify which object moves and at what time for a given motion concept. First, to spatially localize, we introduce GramCol, which adaptively renders per-frame saliency maps for any text concept, including both motion and non-motion. Second, we propose an automatic motion-feature selecting algorithm to obtain an Interpretable Motion-Attentive Map (IMAP) that localizes motions spatially and temporally. Our methods discover concept saliency maps without the need for any gradient-based training or parameters. Experimentally, our methods show standout localization capability in the motion localization task and zero-shot video semantic segmentation, providing interpretable and clearer saliency maps for both motion and non-motion concepts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72e9da1b-15a3-49d0-b34e-80563066aaabBuilds on52
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- ConceptAttention: Diffusion Transformers Learn Highly Interpretable FeaturesAlec Helbling, Tuna Han Salih Meral, Benjamin Hoover, Pinar Yanardag et al.ICML 2025
- Video Motion Transfer with Diffusion TransformersAlexander Pondaven, Aliaksandr Siarohin, Sergey Tulyakov, Philip Torr et al.CVPR 2025
- TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image GenerationVictor Shea-Jay Huang, Le Zhuo, Yi Xin, Zhaokai Wang et al.AAAI 2026 · 10 citations
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 7 citations
- Understanding Video Transformers via Universal Concept DiscoveryMatthew Kowal, Achal Dave, Rares Ambrus, Adrien Gaidon et al.CVPR 2024
