AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning
Binxiao Xu, Junyu Feng, Xiaopeng Lin, Haodong Li, ZhiYuan Feng, Bohan Zeng, Ruichuan An, Ming Lu, Qi She, Wentao Zhang
摘要
Multimodal understanding of advertising videos is essential for interpreting the intricate relationship between visual storytelling and abstract persuasion strategies. However, despite excelling at general search, existing agents often struggle to bridge the cognitive gap between pixel-level perception and high-level marketing logic. To address this challenge, we introduce AD-MIR , a framework designed to decode advertising intent via a two-stage architecture. First, in the Structure-Aware Memory Construction phase, the system converts raw video into a structured database by integrating semantic retrieval with exact keyword matching. This approach prioritizes fine-grained brand details, such as logos and on-screen text, while dynamically filtering out irrelevant background noise to isolate key protagonists. Second, the Structured Reasoning Agent mimics a marketing expert through an iterative inquiry loop, decomposing the narrative to deduce implicit persuasion tactics. Crucially, it employs an evidence-based self-correction mechanism that rigorously validates these insights against specific video frames, automatically backtracking when visual support is lacking. Evaluation on the AdsQA benchmark demonstrates that AD-MIR achieves state-of-the-art performance, surpassing the strongest general-purpose agent, DVD, by 1.8 and 9.5 percentage points in strict and relaxed accuracy, respectively. These results underscore that effective advertising understanding demands explicitly grounding abstract marketing strategies in pixel-level evidence. The code is available at https://github.com/Little-Fridge/AD-MIR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- Deep Video Discovery: Agentic Search with Tool Use for Long-form Video UnderstandingXiaoyi Zhang, Zhaoyang Jia, Zongyu Guo, Jiahao Li 等NeurIPS 2025 · 被引用 95 次
相关 Paper
- AdsQA: Towards Advertisement Video UnderstandingXinwei Long, Kai Tian, Peng Xu, Guoli Jia 等ICCV 2025 · 被引用 1 次
- From Detection to Understanding: Multi-Turn Reasoning for Video Misinformation AnalysisZhi Zeng, Jiaying Wu, Minnan Luo, Di Zhang 等ACL 2026
- SVAgent: Storyline-guided Long Video Understanding via Cross-Modal Multi-Agent CollaborationZhongyu Yang, Zuhao Yang, Shuo Zhan, Tan Yue 等CVPR 2026 · 被引用 5 次
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video UnderstandingYufei Yin, Qianke Meng, Minghao Chen, Jiajun Ding 等CVPR 2026 · 被引用 24 次
- Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual UnderstandingJunhan Chen, Zilu Zhou, Yujun Tong, Dongliang Chang 等CVPR 2026 · 被引用 2 次
