VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
贾 子怡, Zijian Cheng, Xinyue Zhang, Kun-Yang Yu, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo
Abstract
Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domains like healthcare and industry, remains underexplored. In this paper, we introduce VT-Bench, the first unified benchmark for standardizing vision-tabular discriminative prediction and generative reasoning tasks. VT-Bench aggregates 14 datasets across 9 domains (medical-centric, while covering pets, media, and transportation) with over 756K samples. We evaluate 23 representative models, including unimodal experts, specialized visual-tabular models, general-purpose vision-language models (VLMs), and tool-augmented methods, highlighting substantial challenges of visual-tabular learning. We believe VT-Bench will stimulate the community to build more powerful multi-modal vision-tabular foundation models. Benchmark: https://github.com/LAMDA-NeSy/VT-Bench
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78cfdb1b-8d73-4885-a6e5-40b7de0682d6Cited by top-tier papers4
- Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-TrainingHong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia et al.ICML 2026 · 3 citations
- On the Learnability of Test-Time Adaptation: A Recovery Complexity PerspectiveZhi Zhou, Ming Yang, Shi-Yu Tian, Kun-Yang Yu et al.ICML 2026 · 2 citations
- TopBench: A Benchmark for Implicit Predictive Reasoning in Tabular Question AnsweringAn-Yang Ji, Jun-Peng Jiang, De-Chuan Zhan, Han-Jia YeICML 2026 · 1 citation
- OvisOCR: End-to-End Document Parsing via Aligning Specialized Perception with General ReasoningJun-Peng Jiang, Shiyin Lu, An-Yang Ji, Yinglun Li et al.ICML 2026
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav et al.ICLR 2021 · 229 citations
- Multimodal Tabular Reasoning with Privileged Structured InformationJun-Peng Jiang, Yu Xia, Hai-Long Sun, Shiyin Lu et al.NeurIPS 2025 · 16 citations
- Tabular Insights, Visual Impacts: Transferring Expertise from Tables to ImagesJun-Peng Jiang, Han-Jia Ye, Leye Wang, Yang Yang et al.ICML 2024 · 15 citations
- The Golden Subspace: Where Efficiency Meets Generalization in Continual Test-Time AdaptationGuannan Lai, Da-Wei Zhou, Zhenguo Li, Han-Jia YeCVPR 2026 · 2 citations
Related papers
- Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image UnderstandingDexuan Xu, Jiayin Yuan, Jianing Wang, Yanyuan Chen et al.ACL 2026
- U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound UnderstandingAnjie Le, Henan Liu, Yue Wang, Zhenyu Liu et al.ICLR 2026 · 8 citations
- TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist ModelsFangxu Yu, Xingang Guo, Lingzhi Yuan, Haoqiang Kang et al.ICML 2026
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent EnvironmentsZelai Xu, Zhexuan Xu, Xiangmin Yi, Huining Yuan et al.CVPR 2026 · 3 citations
- MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning ModelsVanya Cohen, Ray MooneyICML 2026 · 2 citations
