TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
Jongha Kim, Minseong Bae, Sanghyeok Lee, Jinsung Yoon, Hyunwoo J. Kim
Abstract
Table images present unique challenges for effective and efficient understanding due to the need for question-specific focus and the presence of redundant background regions. Existing Multimodal Large Language Model (MLLM) approaches often overlook these characteristics, resulting in uninformative and redundant visual representations. To address these issues, we aim to generate visual features that are both informative and compact for improved table understanding. We first propose progressive question conditioning, which injects the question into Vision Transformer layers with gradually increasing frequency, considering each layer’s capacity to handle additional information, to generate question-aware visual features. To reduce redundancy, we introduce a pruning strategy that discards background tokens, thereby improving efficiency. To mitigate information loss from pruning, we further propose token focusing, a training strategy that encourages the model to concentrate essential information in the retained tokens. By combining these approaches, we present TabFlash, an efficient and effective MLLM for table understanding. TabFlash achieves state-of-the-art performance, outperforming both open-source and proprietary MLLMs, while requiring 27% less FLOPs and 30% less memory usage compared to the second-best MLLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b57a2a0f-7d4e-4bf0-866b-cac4d51e13a4Cited by top-tier papers3
- DocPrune: Efficient Document Question Answering via Background, Question, and Comprehension-aware Token PruningJoonmyung Choi, Sanghyeok Lee, Jongha Kim, Sehyung Kim et al.CVPR 2026 · 4 citations
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and EvaluationWei Zhou, Bolei Ma, Annemarie Friedrich, Mohsen MesgarACL 2026 · 3 citations
- Improving Large Molecular Language Model via Relation-aware Multimodal CollaborationJinyoung Park, Minseong Bae, Jeehye Na, Hyunwoo J. KimAAAI 2026
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Efficient Large Multi-modal Models via Visual Context CompressionJieneng Chen, Luoxin Ye, Ju He, Zhaoyang Wang et al.NeurIPS 2024 · 49 citations
- Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models Via Adaptive Token SkippingWeili Zeng, Ziyuan Huang, Kaixiang Ji, Yichao YanICCV 2025
- Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text InformationYi Chen, Jian Xu, Xu-Yao Zhang, Wen-Zhuo Liu et al.AAAI 2025 · 18 citations
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 1 citation
- Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token SelectionZixun Sun, Yubo Dong, Hehe Fan, Yi YangCVPR 2026
