DaVinci: Reinforcing Visual-Structural Syntax in MLLMs for Generalized Scientific Diagram Parsing
Xingchen Zeng, Zhewei Su, Hengming Zhang, Juyong Jiang, Jiazhi Xia, Wei Zeng
摘要
Parsing raster-based scientific diagrams into structured representations is critical for editability and reusability. However, existing multimodal LLMs (MLLMs) struggle with the diverse visual primitives, complex structural layouts, and strict syntax involved. To address this, we introduce DaVinci, a novel MLLM that learns diagram parsing based on a two-stage framework: supervised learning of visual primitives followed by reinforcement learning of their structural relationships. Our model learns visual-structural syntax through supervised training on TikZ30K, a newly curated dataset of high-quality diagram-TikZ code pairs that features abundant visual primitives and structurally optimized drawing sequences. We further refine the model via reinforcement learning, guided by a hybrid reward function that jointly optimizes for visual fidelity, structural consistency, and code correctness. Extensive experiments show that DaVinci significantly outperforms existing open-source MLLMs and surpasses leading proprietary models like GPT-5 and Claude-Sonnet-4. Code, datasets, and models are available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li 等AAAI 2020 · 被引用 4,823 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong 等NeurIPS 2025 · 被引用 181 次
- Think Only When You Need with Large Hybrid-Reasoning ModelsLingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong 等NeurIPS 2025 · 被引用 71 次
- AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZJonas Belouadi, Anne Lauscher, Steffen EgerICLR 2024 · 被引用 64 次
相关 Paper
- Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram GenerationZhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li 等ACM MM 2025 · 被引用 3 次
- From Model Diagram to Code: A Benchmark Dataset and Multi-Agent FrameworkMengzhen Wang, Xunbin Huang, Jiayuan Xie, Shukai Ma 等ACM MM 2025 · 被引用 1 次
- FlowGen: Synthesizing Diverse Flowcharts to Enhance and Benchmark MLLM ReasoningKaiwen Shi, Sichen Liu, Ziyue Lin, Hangrui Guo 等ICLR 2026
- TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement LearningChristian Greisinger, Steffen EgerICLR 2026 · 被引用 5 次
- Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided RefinementZhihan Zhang, Yixin Cao, Lizi LiaoACM MM 2025
