VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
贾 子怡, Zijian Cheng, Xinyue Zhang, Kun-Yang Yu, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo
摘要
Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domains like healthcare and industry, remains underexplored. In this paper, we introduce VT-Bench, the first unified benchmark for standardizing vision-tabular discriminative prediction and generative reasoning tasks. VT-Bench aggregates 14 datasets across 9 domains (medical-centric, while covering pets, media, and transportation) with over 756K samples. We evaluate 23 representative models, including unimodal experts, specialized visual-tabular models, general-purpose vision-language models (VLMs), and tool-augmented methods, highlighting substantial challenges of visual-tabular learning. We believe VT-Bench will stimulate the community to build more powerful multi-modal vision-tabular foundation models. Benchmark: https://github.com/LAMDA-NeSy/VT-Bench
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-TrainingHong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia 等ICML 2026 · 被引用 3 次
- On the Learnability of Test-Time Adaptation: A Recovery Complexity PerspectiveZhi Zhou, Ming Yang, Shi-Yu Tian, Kun-Yang Yu 等ICML 2026 · 被引用 2 次
- TopBench: A Benchmark for Implicit Predictive Reasoning in Tabular Question AnsweringAn-Yang Ji, Jun-Peng Jiang, De-Chuan Zhan, Han-Jia YeICML 2026 · 被引用 1 次
- OvisOCR: End-to-End Document Parsing via Aligning Specialized Perception with General ReasoningJun-Peng Jiang, Shiyin Lu, An-Yang Ji, Yinglun Li 等ICML 2026
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav 等ICLR 2021 · 被引用 229 次
- Multimodal Tabular Reasoning with Privileged Structured InformationJun-Peng Jiang, Yu Xia, Hai-Long Sun, Shiyin Lu 等NeurIPS 2025 · 被引用 16 次
- Tabular Insights, Visual Impacts: Transferring Expertise from Tables to ImagesJun-Peng Jiang, Han-Jia Ye, Leye Wang, Yang Yang 等ICML 2024 · 被引用 15 次
- The Golden Subspace: Where Efficiency Meets Generalization in Continual Test-Time AdaptationGuannan Lai, Da-Wei Zhou, Zhenguo Li, Han-Jia YeCVPR 2026 · 被引用 2 次
相关 Paper
- Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image UnderstandingDexuan Xu, Jiayin Yuan, Jianing Wang, Yanyuan Chen 等ACL 2026
- U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound UnderstandingAnjie Le, Henan Liu, Yue Wang, Zhenyu Liu 等ICLR 2026 · 被引用 8 次
- TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist ModelsFangxu Yu, Xingang Guo, Lingzhi Yuan, Haoqiang Kang 等ICML 2026
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent EnvironmentsZelai Xu, Zhexuan Xu, Xiangmin Yi, Huining Yuan 等CVPR 2026 · 被引用 3 次
- MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning ModelsVanya Cohen, Ray MooneyICML 2026 · 被引用 2 次
