Twin-T & TwintVQA: A Reliable Structure–Detail Separating VLM and a Comprehensive Benchmark for Chart and Table Tasks
Jiahua Bao, Siyao Cheng, Jiaxing Du, Qingtao Xia, Changjiang He, Zeming Lang, Jie Liu
摘要
With the rapid development of Vision-Language Models (VLMs), there is a growing demand for automatic analysis of structured visual data. Charts and tables are primary carriers of quantitative information, with regular layouts and explicit numbers. However, current general VLMs and expert models make limited use of these chart-table features during training and inference. Another challenge is cross-format conversion in realistic settings, as chart and table outputs span Python and LaTeX and most VLMs struggle to handle this breadth reliably. These gaps often lead to analysis mistakes, and unreliable generation text. To overcome these limitations, we propose \underline{\texttt{\textbf{Twin-T}}}, a two-stage expert VLM for comprehensive char\underline{\texttt{\textbf{t}}}-\underline{\texttt{\textbf{t}}}able tasks across Image, LaTeX, and Python. In stage 1, we propose a novel dual-head image encoder that can separate structural cues and fine details from input images. In stage 2, we propose MINT, a preference learning method that emphasizes numbers and keywords fidelity and vision–text matching. Furthermore, we introduce a comprehensive TwintVQA benchmark with 17 chart types, 11 task types, 3 data formats and short / medium / long QA settings. Our model narrows the gap between open-source and closed-source models on mainstream chart–table benchmarks, outperforming open-source models and GLM-4.5V-106B while even remaining competitive with GPT-4o and Gemini-2.5-Pro. Our code and additional details are available in the Appendix.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyShusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye 等ICML 2024 · 被引用 274 次
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan 等NeurIPS 2024 · 被引用 214 次
- ChartCap: Mitigating Hallucination of Dense Chart CaptioningJunyoung Lim, Jaewoo Ahn, Gunhee KimICCV 2025 · 被引用 8 次
相关 Paper
- TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token MergingLiang Zhang, Anwen Hu, Haiyang Xu, Ming Yan 等EMNLP 2024 · 被引用 15 次
- RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo 等ACL 2026
- Same or Not? Enhancing Visual Perception in Vision-Language ModelsDamiano Marsili, Aditya Mehta, Ryan Y. Lin, Georgia GkioxariCVPR 2026 · 被引用 5 次
- From Charts to Code: A Hierarchical Benchmark for Multimodal ModelsJiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang 等ACL 2026 · 被引用 5 次
- Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data GenerationYue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta 等ACL 2025
