Image Manipulation via Multi-Hop Instructions - A New Dataset and Weakly-Supervised Neuro-Symbolic Approach
Harman Singh, Poorva Garg, Mohit Gupta, Kevin Shah, Ashish Goswami, Satyam Modi, Arnab Kumar Mondal, Dinesh Khandelwal, Dinesh Garg, Parag Singla
摘要
We are interested in image manipulation via natural language text -a task that is useful for multiple AI applications but requires complex reasoning over multi-modal spaces. We extend recently proposed Neuro Symbolic Concept Learning (NSCL) (Mao et al., 2019) , which has been quite effective for the task of Visual Question Answering (VQA), for the task of image manipulation. Our system referred to as NEUROSIM can perform complex multi-hop reasoning over multi-object scenes and only requires weak supervision in the form of annotated data for VQA. NEUROSIM parses an instruction into a symbolic program, based on a Domain Specific Language (DSL) comprising of object attributes and manipulation operations, that guides its execution. We create a new dataset for the task, and extensive experiments demonstrate that NEUROSIM is highly competitive with or beats SOTA baselines that make use of supervised data for manipulation. † Equal Contribution, * Equal Contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- On Aliased Resizing and Surprising Subtleties in GAN EvaluationGaurav Parmar, Richard Zhang, Jun-Yan ZhuCVPR 2022 · 被引用 250 次
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm 等ICCV 2019 · 被引用 128 次
- DreamCoder: bootstrapping inductive program synthesis with wake-sleep library learningKevin Ellis, Catherine Wong, Maxwell I. Nye, Mathias Sablé-Meyer 等PLDI 2021 · 被引用 97 次
- Text as Neural Operator: Image Manipulation by Text InstructionTianhao Zhang, Hung-Yu Tseng, Lu Jiang, Weilong Yang 等ACM MM 2021 · 被引用 28 次
相关 Paper
- Visual Programming: Compositional visual reasoning without trainingTanmay Gupta, Aniruddha KembhaviCVPR 2023
- Motion Question Answering via Modular Motion ProgramsMark Endo, Joy Hsu, Jiaman Li, Jiajun WuICML 2023 · 被引用 28 次
- Synthesizing Visual Concepts as Vision-Language ProgramsAntonia Wüst, Wolfgang Stammer, Hikaru Shindo, Lukas Helff 等CVPR 2026 · 被引用 6 次
- ImageEye: Batch Image Processing using Program SynthesisCeleste Barnaby, Qiaochu Chen, Roopsha Samanta, Isil DilligPLDI 2023 · 被引用 13 次
- Naturally Supervised 3D Visual Grounding with Language-Regularized Concept LearnersChun Feng, Joy Hsu, Weiyu Liu, Jiajun WuCVPR 2024
