MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond
Duy-Kien Nguyen, Vedanuj Goswami, Xinlei Chen
摘要
This paper focuses on visual counting, which aims to predict the number of occurrences given a natural image and a query (e.g. a question or a category). Unlike most prior works that use explicit, symbolic models which can be computationally expensive and limited in generalization, we propose a simple and effective alternative by revisiting modulated convolutions that fuse the query and the image locally. Following the design of residual bottleneck, we call our method MoVie, short for Modulated conVolutional bottlenecks. Notably, MoVie reasons implicitly and holistically and only needs a single forward-pass during inference. Nevertheless, MoVie showcases strong performance for counting: 1) advancing the state-of-the-art on counting-specific VQA tasks while being more efficient; 2) outperforming prior-art on difficult benchmarks like COCO for common object counting; 3) helped us secure the first place of 2020 VQA challenge when integrated as a module for 'number' related questions in generic VQA models. Finally, we show evidence that modulated convolutions such as MoVie can serve as a general mechanism for reasoning tasks beyond counting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Entity-Focused Dense Passage Retrieval for Outside-Knowledge Visual Question AnsweringJialin Wu, Raymond J. MooneyEMNLP 2022 · 被引用 9 次
- Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning ScenariosShantanu Jaiswal, Debaditya Roy, Basura Fernando, Cheston TanNeurIPS 2024 · 被引用 9 次
- COMMA: Co-articulated Multi-Modal LearningLianyu Hu, Liqing Gao, Zekang Liu, Chi-Man Pun 等AAAI 2024 · 被引用 7 次
- Improving Selective Visual Question Answering by Learning from Your PeersCorentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam 等CVPR 2023
它引用的顶会 Paper1
相关 Paper
- Focal and Composed Vision-semantic Modeling for Visual Question AnsweringYudong Han, Yangyang Guo, Jianhua Yin, Meng Liu 等ACM MM 2021 · 被引用 14 次
- MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering BenchmarkShaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie 等CVPR 2026 · 被引用 4 次
- Decomposed Attention Fusion in MLLMs for Training-free Video Reasoning SegmentationSu Ho Han, Jeongseok Hyun, Pilhyeon Lee, Minho Shim 等ICLR 2026 · 被引用 2 次
- Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language ModelsIdo Cohen, Daniela Gottesman, Mor Geva, Raja GiryesACL 2025
- Core-to-Global Reasoning for Compositional Visual Question AnsweringHao Zhou, Tingjin Luo, Zhangqi JiangAAAI 2025 · 被引用 2 次
