TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition
Junyuan Zhang, Bin Wang, Qintong Zhang, Fan Wu, Zichen Wen, Jialin Lu, Junjie Shan, Ziqi Zhao, Shuya Yang, Ziling Wang, Ziyang Miao, Huaping Zhong
Abstract
Table recognition (TR) aims to transform table images into semi-structured representations such as HTML or Markdown. As a core component of document parsing, TR has long relied on supervised learning, with recent efforts dominated by fine-tuning vision-language models (VLMs) using labeled data. While VLMs have brought TR to the next level, pushing performance further demands large-scale labeled data that is costly to obtain. Consequently, although proprietary models have continuously pushed the performance boundary, open-source models, often trained with limited resources and, in practice, the only viable option for many due to privacy regulations, still lag far behind. To bridge this gap, we introduce TRivia, a self-supervised fine-tuning method that enables pretrained VLMs to learn TR directly from unlabeled table images in the wild. Built upon Group Relative Policy Optimization, TRivia automatically identifies unlabeled samples that most effectively facilitate learning and eliminates the need for human annotations through a question-answering-based reward mechanism. An attention-guided module generates diverse questions for each table image, and the ability to interpret the recognition results and answer them correctly provides feedback to optimize the TR model. This closed-loop process allows the TR model to autonomously learn to recognize, structure, and reason over tables without labeled data. Leveraging this pipeline, we present TRivia-3B, an open-sourced, compact, and state-of-the-art TR model that surpasses existing systems (e.g., Gemini 2.5 Pro, MinerU2.5) on three popular benchmarks. Model and code are released at: https://github.com/HKU-TASR/TRivia
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e43ed9fe-9a4c-40d3-ba6f-a3f6bc26aecdBuilds on17
- Parsing Table Structures in the WildRujiao Long, Wen Wang, Nan Xue, Feiyu Gao et al.ICCV 2021 · 77 citations
- TSRFormer: Table Structure Recognition with TransformersWeihong Lin, Zheng Sun, Chixiang Ma, Mingze Li et al.ACM MM 2022 · 56 citations
- Prism: A Framework for Decoupling and Assessing the Capabilities of VLMsYuxuan Qiao, Haodong Duan, Xinyu Fang, Junming Yang et al.NeurIPS 2024 · 49 citations
- Neural Collaborative Graph Machines for Table Structure RecognitionHao Liu, Xin Li, Bing Liu, Deqiang Jiang et al.CVPR 2022 · 45 citations
- LORE: Logical Location Regression Network for Table Structure RecognitionHangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu et al.AAAI 2023 · 43 citations
Related papers
- TableVLM: Multi-modal Pre-training for Table Structure RecognitionLeiyuan Chen, Chengsong Huang, Xiaoqing Zheng, Jinshu Lin et al.ACL 2023 · 8 citations
- Can GRPO Boost Complex Multimodal Table Understanding?Xiaoqiang Kang, Shengen Wu, Zimu Wang, Yilin Liu et al.EMNLP 2025 · 1 citation
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Fine-tuningJunjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong et al.EMNLP 2025 · 2 citations
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
