Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate Modeling
Yongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li, Zecheng Xie, Shenggao Zhu, Liangcai Gao, Wei Peng
摘要
Table structure recognition aims to extract the logical and physical structure of unstructured table images into a machine-readable format. The latest end-to-end imageto-text approaches simultaneously predict the two structures by two decoders, where the prediction of the physical structure (the bounding boxes of the cells) is based on the representation of the logical structure. However, the previous methods struggle with imprecise bounding boxes as the logical representation lacks local visual information. To address this issue, we propose an end-to-end sequential modeling framework for table structure recognition called VAST. It contains a novel coordinate sequence decoder triggered by the representation of the non-empty cell from the logical structure decoder. In the coordinate sequence decoder, we model the bounding box coordinates as a language sequence, where the left, top, right and bottom coordinates are decoded sequentially to leverage the intercoordinate dependency. Furthermore, we propose an auxiliary visual-alignment loss to enforce the logical representation of the non-empty cells to contain more local visual details, which helps produce better cell bounding boxes. Extensive experiments demonstrate that our proposed method can achieve state-of-the-art results in both logical and physical structure recognition. The ablation study also validates that the proposed coordinate sequence decoder and the visual-alignment loss are the keys to the success of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table RecognitionJianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu 等CVPR 2024 · 被引用 29 次
- GridFormer: Towards Accurate Table Structure Recognition via Grid PredictionPengyuan Lyu, Weihong Ma, Hongyi Wang, Yuechen Yu 等ACM MM 2023 · 被引用 17 次
- Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language ModelsJun Ling, Yao Qi, Tao Huang, Shibo Zhou 等NeurIPS 2025 · 被引用 9 次
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen 等CVPR 2026 · 被引用 9 次
- TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table RecognitionJunyuan Zhang, Bin Wang, Qintong Zhang, Fan Wu 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper6
- Pix2seq: A Language Modeling Framework for Object DetectionTing Chen, Saurabh Saxena, Lala Li, David J. Fleet 等ICLR 2022 · 被引用 435 次
- PubTables-1M: Towards comprehensive table extraction from unstructured documentsBrandon Smock, Rohith Pesala, Robin AbrahamCVPR 2022 · 被引用 125 次
- Parsing Table Structures in the WildRujiao Long, Wen Wang, Nan Xue, Feiyu Gao 等ICCV 2021 · 被引用 77 次
- TSRFormer: Table Structure Recognition with TransformersWeihong Lin, Zheng Sun, Chixiang Ma, Mingze Li 等ACM MM 2022 · 被引用 56 次
- Neural Collaborative Graph Machines for Table Structure RecognitionHao Liu, Xin Li, Bing Liu, Deqiang Jiang 等CVPR 2022 · 被引用 45 次
相关 Paper
- G2LFormer: Global-to-Local Query Enhancement for Robust Table Structure RecognitionHaosheng Cai, Yang XueACM MM 2025
- LORE: Logical Location Regression Network for Table Structure RecognitionHangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu 等AAAI 2023 · 被引用 43 次
- TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual AlignmentChunxia Qin, Chenyu Liu, Pengcheng Xia, Jun Du 等CVPR 2026 · 被引用 2 次
- TableVLM: Multi-modal Pre-training for Table Structure RecognitionLeiyuan Chen, Chengsong Huang, Xiaoqing Zheng, Jinshu Lin 等ACL 2023 · 被引用 8 次
- TGRNet: A Table Graph Reconstruction Network for Table Structure RecognitionWenyuan Xue, Baosheng Yu, Wen Wang, Dacheng Tao 等ICCV 2021 · 被引用 65 次
