Synthetic Visual Genome
Jae Sung Park, Zixian Ma, Linjie Li, Chenhao Zheng, Cheng-Yu Hsieh, Ximing Lu, Khyathi Raghavi Chandu, Quan Kong, Norimasa Kobori, Ali Farhadi, Yejin Choi, Ranjay Krishna
Abstract
Reasoning over visual relationships-spatial, functional, interactional, social, etc.-is considered to be a fundamental component of human cognition. Yet, despite the major advances in visual comprehension in multimodal language models (MLMs), precise reasoning over relationships and their generations remains a challenge. We introduce ROBIN: an MLM instruction-tuned with densely annotated relationships capable of constructing high-quality dense scene graphs at scale. To train ROBIN, we curate SVG 1 , a synthetic scene graph dataset by completing the missing relations of selected objects in existing scene graphs using a teacher MLM and a carefully designed filtering process to ensure high-quality. To generate more accurate and rich scene graphs at scale for any image, we introduce SG-EDIT: a self-distillation framework where GPT-4o further refines ROBIN's predicted scene graphs by removing unlikely relations and/or suggesting relevant ones. In total, our dataset contains 146K images and 5.6M relationships for 2.6M objects. Results show that our ROBIN-3B model, despite being trained on less than 3 million instances, outperforms similar-size models trained on over 300 million instances on relationship understanding benchmarks, and even surpasses larger models up to 13B parameters. Notably, it achieves state-of-the-art performance in referring expression comprehension with a score of 88.2, surpassing the previous best of 87.4. Our results suggest that training on the refined scene graph data is crucial to maintaining high performance across diverse visual reasoning tasks 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation TrainingZiqi Gao, Weikai Huang, Jieyu Zhang, Aniruddha Kembhavi et al.ICLR 2026 · 1 citation
- Fine-Grained Multi Image Object Hallucination BenchmarkJoonki Min, Chaeyun Kim, Hyungwook Choi, Yejin Kim et al.CVPR 2026 · 1 citation
- BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship DetectionMelissa Schween, Mathis Kruse, Bodo RosenhahnCVPR 2026
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Robin3D Improving 3D Large Language Model via Robust Instruction TuningWeitai Kang, Haifeng Huang, Yuzhang Shang, Mubarak Shah et al.ICCV 2025 · 6 citations
- Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language ModelsZhiwei Yang, Yuanchen Wu, Nan Zhang, Yucong Meng et al.ICML 2026 · 1 citation
- Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene GraphsRoei Herzig, Alon Mendelson, Leonid Karlinsky, Assaf Arbelle et al.EMNLP 2023 · 15 citations
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen et al.CVPR 2024
