DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinement
Shaoqing Lin, Chong Teng, Fei Li, Donghong Ji, Lizhen Qu, Zhuang Li
Abstract
Vision-Language Models (VLMs) generate discourse-level, multi-sentence visual descriptions, challenging text scene graph parsers built for single-sentence caption-to-graph mapping. Current approaches typically merge sentencelevel parsing outputs for discourse input, often missing phenomena like cross-sentence coreference, resulting in fragmented graphs and degraded downstream VLM task performance. We introduce a new task, Discourse-level text Scene Graph parsing (DiscoSG), and release DiscoSG-DS, a dataset of 400 expert-annotated and 8,430 synthesised multi-sentence captiongraph pairs. Each caption averages 9 sentences, and each graph contains at least 3× more triples than those in existing datasets. Fine-tuning GPT-4o on DiscoSG-DS yields over 40% higher SPICE metric than the best sentence-merging baseline. However, its high inference cost and licensing restrict opensource use. Smaller fine-tuned open-source models (e.g., Flan-T5) perform well on simpler graphs yet degrade on denser, more complex graphs. To bridge this gap, we introduce DiscoSG-Refiner, a lightweight open-source parser that drafts a seed graph and iteratively refines it with a novel learned graph-editing model, achieving 30% higher SPICE than the baseline while delivering 86× faster inference than GPT-4o. It generalises from simple to dense graphs, thereby consistently improving downstream VLM tasks, including discourselevel caption evaluation and hallucination detection, outperforming alternative open-source parsers. Code and data are available at https: //github.com/ShaoqLin/DiscoSG .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1743ef93-aa23-4661-9ea8-dc301a9ba337Builds on12
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen et al.CVPR 2024 · 38 citations
- HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction DataQifan Yu, Juncheng Li, Longhui Wei, Liang Pang et al.CVPR 2024 · 29 citations
- Improving Scene Graph Classification by Exploiting Knowledge from TextsSahand Sharifzadeh, Sina Moayed Baharlou, Martin Schmitt, Hinrich Schütze et al.AAAI 2022 · 20 citations
Related papers
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen et al.CVPR 2024
- ZINA: Multimodal Fine-grained Hallucination Detection and EditingYuiga Wada, Kazuki Matsuda, Komei Sugiura, Graham NeubigCVPR 2026 · 5 citations
- Hallucination Mitigation in Natural Language Generation from Large-Scale Open-Domain Knowledge GraphsXiao Shi, Zhengyuan Zhu, Zeyu Zhang, Chengkai LiEMNLP 2023 · 7 citations
- Multimodal Contextualized Semantic Parsing from SpeechJordan Voas, David Harwath, Raymond MooneyACL 2024
- Synthetic Visual GenomeJae Sung Park, Zixian Ma, Linjie Li, Chenhao Zheng et al.CVPR 2025
