Parametric Visual Program Induction with Function Modularization
Xuguang Duan, Xin Wang, Ziwei Zhang, Wenwu Zhu
摘要
Generating programs to describe visual observations has gained much research attention recently. However, most of the existing approaches are based on non-parametric primitive functions, making them unable to handle complex visual scenes involving many attributes and details. In this paper, we propose the concept of parametric visual program induction. Learning to generate parametric programs for visual scenes is challenging due to the huge number of function variants and the complex function correlations. To solve these challenges, we propose the method of function modularization, capable of dealing with numerous function variants and complex correlations. Specifically, we model each parametric function as a multi-head self-contained neural module to cover different function variants. Moreover, to eliminate the complex correlations between functions, we propose the hierarchical heterogeneous Monto-Carlo tree search (H2MCTS) algorithm which can provide high-quality uncorrelated supervision during training, and serve as an efficient searching technique during testing. We demonstrate the superiority of the proposed method on three visual program induction datasets involving parametric primitive functions. Experimental results show that our proposed model is able to significantly outperform the state-of-the-art baseline methods in terms of generating accurate programs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Synthesizing Tasks for Block-based ProgrammingUmair Z. Ahmed, Maria Christakis, Aleksandr Efremov, Nigel Fernandez 等NeurIPS 2020 · 被引用 26 次
- PLANS: Neuro-Symbolic Program Learning from VideosRaphaël Dang-NhuNeurIPS 2020 · 被引用 18 次
- De-fine: Decomposing and Refining Visual Programs with Auto-FeedbackMinghe Gao, Juncheng Li, Hao Fei, Liang Pang 等ACM MM 2024 · 被引用 1 次
- Learning to Infer Generative Template Programs for Visual ConceptsR. Kenny Jones, Siddhartha Chaudhuri, Daniel RitchieICML 2024 · 被引用 3 次
- Hierarchical Semantic Correspondence Networks for Video Paragraph GroundingChaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng 等CVPR 2023
