Twin-T & TwintVQA: A Reliable Structure–Detail Separating VLM and a Comprehensive Benchmark for Chart and Table Tasks
Jiahua Bao, Siyao Cheng, Jiaxing Du, Qingtao Xia, Changjiang He, Zeming Lang, Jie Liu
Abstract
With the rapid development of Vision-Language Models (VLMs), there is a growing demand for automatic analysis of structured visual data. Charts and tables are primary carriers of quantitative information, with regular layouts and explicit numbers. However, current general VLMs and expert models make limited use of these chart-table features during training and inference. Another challenge is cross-format conversion in realistic settings, as chart and table outputs span Python and LaTeX and most VLMs struggle to handle this breadth reliably. These gaps often lead to analysis mistakes, and unreliable generation text. To overcome these limitations, we propose \underline{\texttt{\textbf{Twin-T}}}, a two-stage expert VLM for comprehensive char\underline{\texttt{\textbf{t}}}-\underline{\texttt{\textbf{t}}}able tasks across Image, LaTeX, and Python. In stage 1, we propose a novel dual-head image encoder that can separate structural cues and fine details from input images. In stage 2, we propose MINT, a preference learning method that emphasizes numbers and keywords fidelity and vision–text matching. Furthermore, we introduce a comprehensive TwintVQA benchmark with 17 chart types, 11 task types, 3 data formats and short / medium / long QA settings. Our model narrows the gap between open-source and closed-source models on mainstream chart–table benchmarks, outperforming open-source models and GLM-4.5V-106B while even remaining competitive with GPT-4o and Gemini-2.5-Pro. Our code and additional details are available in the Appendix.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65dcdf50-71fa-46b9-adb7-8ae2dadc72c0Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyShusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye et al.ICML 2024 · 274 citations
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan et al.NeurIPS 2024 · 214 citations
- ChartCap: Mitigating Hallucination of Dense Chart CaptioningJunyoung Lim, Jaewoo Ahn, Gunhee KimICCV 2025 · 8 citations
Related papers
- TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token MergingLiang Zhang, Anwen Hu, Haiyang Xu, Ming Yan et al.EMNLP 2024 · 15 citations
- RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo et al.ACL 2026
- Same or Not? Enhancing Visual Perception in Vision-Language ModelsDamiano Marsili, Aditya Mehta, Ryan Y. Lin, Georgia GkioxariCVPR 2026 · 5 citations
- From Charts to Code: A Hierarchical Benchmark for Multimodal ModelsJiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang et al.ACL 2026 · 5 citations
- Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data GenerationYue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta et al.ACL 2025
