ProTo: Program-Guided Transformer for Program-Guided Tasks
Zelin Zhao, Karan Samel, Binghong Chen, Le Song
Abstract
Programs, consisting of semantic and structural information, play an important role in the communication between humans and agents. Towards learning general program executors to unify perception, reasoning, and decision making, we formulate program-guided tasks which require learning to execute a given program on the observed task specification. Furthermore, we propose the Program-guided Transformer (ProTo), which integrates both semantic and structural guidance of a program by leveraging cross-attention and masked self-attention to pass messages between the specification and routines in the program. ProTo executes a program in a learned latent space and enjoys stronger representation ability than previous neural-symbolic approaches. We demonstrate that ProTo significantly outperforms the previous state-of-the-art methods on GQA visual reasoning and 2D Minecraft policy learning datasets. Additionally, ProTo demonstrates better generalization to unseen, complex, and human-written programs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ced0bcc-d8e4-46bf-b35f-e5a0c76423f6Cited by top-tier papers11
- MST: Masked Self-Supervised Transformer for Visual RepresentationZhaowen Li, Zhiyang Chen, Fan Yang, Wei Li et al.NeurIPS 2021 · 194 citations
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 104 citations
- Dual-stream Network for Visual RecognitionMingyuan Mao, Peng Gao, Renrui Zhang, Honghui Zheng et al.NeurIPS 2021 · 88 citations
- Container: Context Aggregation NetworksPeng Gao, Jiasen Lu, Hongsheng Li, Roozbeh Mottaghi et al.NeurIPS 2021 · 86 citations
- Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal ReasoningWenhao Ding, Haohong Lin, Bo Li, Ding ZhaoNeurIPS 2022 · 59 citations
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- ResT: An Efficient Transformer for Visual RecognitionQinglong Zhang, Yu-Bin YangNeurIPS 2021 · 313 citations
Related papers
- HYSYNTH: Context-Free LLM Approximation for Guiding Program SynthesisShraddha Barke, Emmanuel Anaya Gonzalez, Saketh Ram Kasibatla, Taylor Berg-Kirkpatrick et al.NeurIPS 2024 · 34 citations
- Programmatically Grounded, Compositionally Generalizable Robotic ManipulationRenhao Wang, Jiayuan Mao, Joy Hsu, Hang Zhao et al.ICLR 2023 · 4 citations
- NePTune: A Neuro-Pythonic Framework for Tunable Compositional Reasoning on Vision-LanguageDanial Kamali, Parisa KordjamshidiICLR 2026 · 10 citations
- Program Guided AgentShao-Hua Sun, Te-Lin Wu, Joseph J. LimICLR 2020 · 63 citations
- Synthesizing Visual Concepts as Vision-Language ProgramsAntonia Wüst, Wolfgang Stammer, Hikaru Shindo, Lukas Helff et al.CVPR 2026 · 6 citations
