CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question Answering
Shaowei Wang, Lingling Zhang, Longji Zhu, Tao Qin, Kim-Hui Yap, Xinyu Zhang, Jun Liu
摘要
Diagram Question Answering (DQA) is a challenging task, requiring models to answer natural language questions based on visual diagram contexts. It serves as a crucial basis for academic tutoring, technical support, and more practical applications. DQA poses significant challenges, such as the demand for domain-specific knowledge and the scarcity of annotated data, which restrict the applicability of large-scale deep models. Previous approaches have explored external knowledge integration through pretraining, but these methods are costly and can be limited by domain disparities. While Large Language Models (LLMs) show promise in question-answering, there is still a gap in how to cooperate and interact with the diagram parsing process. In this paper, we introduce the Chain-of-Guiding Learning Model for Diagram Question Answering (CoG-DQA), a novel framework that effectively addresses DQA challenges. CoG-DQA leverages LLMs to guide diagram parsing tools (DPTs) through the guiding chains, enhancing the precision of diagram parsing while introducing rich background knowledge. Our experimental findings reveal that CoG-DQA surpasses all comparison models in various DQA scenarios, achieving an average accuracy enhancement exceeding 5% and peaking at 11% across four datasets. These results underscore CoG-DQA's capacity to advance the field of visual question answering and promote the integration of LLMs into specialized domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Sketch2Diagram: Generating Vector Diagrams from Hand-Drawn SketchesItsumi Saito, Haruto Yoshida, Keisuke SakaguchiICLR 2025
- Chain-of-region: Visual Language Models Need Details for Diagram AnalysisXue Li, Yiyou Sun, Wei Cheng, Yinglun Zhu 等ICLR 2025
- Diagram-Driven Course Questions GenerationXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang 等EMNLP 2025
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
相关 Paper
- GlFoMR: A Glance-then-Focus Multimodal Reasoning Framework for Diagram Question AnsweringYaxian Wang, Bifan Wei, Jun Liu, Lingling Zhang 等SIGIR 2025
- Hierarchical Multi-Task Learning for Diagram Question Answering with Multi-Modal TransformerZhaoquan Yuan, Xiao Peng, Xiao Wu, Changsheng XuACM MM 2021 · 被引用 10 次
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question AnsweringJingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao 等EMNLP 2025 · 被引用 6 次
- Union Is Strength! Unite the Power of LLMs and MLLMs for Chart Question AnsweringJiapeng Liu, Liang Li, Shihao Rao, Xiyan Gao 等AAAI 2025 · 被引用 3 次
- Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question AnsweringMahsa Mozaffari, Hitesh Sapkota, Xumin Liu, Qi YuICLR 2026
