Program-Guided Image Manipulators
Xiuming Zhang, Jiayuan Mao, Yikai Li, William T. Freeman, Joshua B. Tenenbaum, Jiajun Wu
摘要
Humans are capable of building holistic representations for images at various levels, from local objects, to pairwise relations, to global structures. The interpretation of structures involves reasoning over repetition and symmetry of the objects in the image. In this paper, we present the Program-Guided Image Manipulator (PG-IM), inducing neuro-symbolic program-like representations to represent and manipulate images. Given an image, PG-IM detects repeated patterns, induces symbolic programs, and manipulates the image using a neural network that is guided by the program. PG-IM learns from a single image, exploiting its internal statistics. Despite trained only on image inpainting, PG-IM is directly capable of extrapolation and regularity editing in a unified framework. Extensive experiments show that PG-IM achieves superior performance on all the tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Is Programming by Example Solved by LLMs?Wen-Ding Li, Kevin EllisNeurIPS 2024 · 被引用 45 次
- ProTo: Program-Guided Transformer for Program-Guided TasksZelin Zhao, Karan Samel, Binghong Chen, Le SongNeurIPS 2021 · 被引用 37 次
- Text as Neural Operator: Image Manipulation by Text InstructionTianhao Zhang, Hung-Yu Tseng, Lu Jiang, Weilong Yang 等ACM MM 2021 · 被引用 28 次
- Unsupervised Learning of Shape Programs with Repeatable Implicit PartsBoyang Deng, Sumith Kulal, Zhengyang Dong, Congyue Deng 等NeurIPS 2022 · 被引用 18 次
- Multi-Plane Program Induction with 3D Box PriorsYikai Li, Jiayuan Mao, Xiuming Zhang, Bill Freeman 等NeurIPS 2020 · 被引用 18 次
它引用的顶会 Paper2
相关 Paper
- Generating Programmatic Referring Expressions via Program SynthesisJiani Huang, Calvin Smith, Osbert Bastani, Rishabh Singh 等ICML 2020 · 被引用 11 次
- Hierarchical Motion Understanding via Motion ProgramsSumith Kulal, Jiayuan Mao, Alex Aiken, Jiajun WuCVPR 2021
- Image Shape Manipulation from a Single Augmented Training SampleYael Vinker, Eliahu Horwitz, Nir Zabari, Yedid HoshenICCV 2021 · 被引用 24 次
- Image Translation as Diffusion Visual ProgrammersCheng Han, James Chenhao Liang, Qifan Wang, Majid Rabbani 等ICLR 2024 · 被引用 19 次
- Visual Programming: Compositional visual reasoning without trainingTanmay Gupta, Aniruddha KembhaviCVPR 2023
