Differentiable Adaptive Computation Time for Visual Reasoning
Cristóbal Eyzaguirre, Álvaro Soto
Abstract
This paper presents a novel attention-based algorithm for achieving adaptive computation called DACT, which, unlike existing ones, is end-to-end differentiable. Our method can be used in conjunction with many networks; in particular, we study its application to the widely known MAC architecture, obtaining a significant reduction in the number of recurrent steps needed to achieve similar accuracies, therefore improving its performance to computation ratio. Furthermore, we show that by increasing the maximum number of steps used, we surpass the accuracy of even our best non-adaptive MAC in the CLEVR dataset, demonstrating that our approach is able to control the number of steps without significant loss of performance. Additional advantages provided by our approach include considerably improving interpretability by discarding useless steps and providing more insights into the underlying reasoning process. Finally, we present adaptive computation as an equivalent to an ensemble of models, similar to a mixture of expert formulation. Both the code and the configuration files for our experiments are made available to support further research in this area 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 278a3a10-3bcf-473c-9e33-bbc82a5efa35Cited by top-tier papers7
- Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent NetworksAvi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang et al.NeurIPS 2021 · 133 citations
- Towards Scale-Invariant Graph-related Problem Solving by Iterative Homogeneous GNNsHao Tang, Zhiao Huang, Jiayuan Gu, Bao-Liang Lu et al.NeurIPS 2020 · 54 citations
- End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without OverthinkingArpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam et al.NeurIPS 2022 · 54 citations
- Else-Net: Elastic Semantic Network for Continual Action Recognition from Skeleton DataTianjiao Li, Qiuhong Ke, Hossein Rahmani, Rui En Ho et al.ICCV 2021 · 46 citations
- Path Independent Equilibrium Models Can Better Exploit Test-Time ComputationCem Anil, Ashwini Pokle, Kaiqu Liang, Johannes Treutlein et al.NeurIPS 2022 · 43 citations
Related papers
- Mixture of Attention Heads: Selecting Attention Heads Per TokenXiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou et al.EMNLP 2022 · 23 citations
- Towards Robust Image Classification Using Sequential Attention ModelsDaniel Zoran, Mike Chrzanowski, Po-Sen Huang, Sven Gowal et al.CVPR 2020
- Adaptive Computation Modules: Granular Conditional Computation for Efficient InferenceBartosz Wójcik, Alessio Devoto, Karol Pustelnik, Pasquale Minervini et al.AAAI 2025 · 8 citations
- From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round RefinementJianzhi Yan, Le Liu, Youcheng Pan, Shiwei Chen et al.EMNLP 2025
- ChainGPT: Dual-Reasoning Model with Recurrent Depth and Multi-Rank State UpdatesYunao Zheng, Xiaojie Wang, Lei Ren, Chen WeiICLR 2026
