PIX-TAB: Efficient PIXel-Precise TABle Structure Recognition Approach with Speculative Decoding and Region-Based Image Segmentation
Viktor Zaytsev, Olena Vynokurova, Pavlo Tytarchuk, Dmytro Kozii, Vitalii Pohribnyi, Olga Radyvonenko, Artem Shcherbina
摘要
Table structure recognition in document AI faces significant challenges due to layout inconsistencies, merged cells, and complex nested structures, which is further exacerbated by the scarcity of large, diverse annotated datasets. In this paper, we present PIX-TAB (Efficient PIXel-Precise TABle Structure Recognition Approach) that provides exact, pixellevel structure using a small, lightweight model that can run on-device. The approach is language-agnostic, as it allows adding support for a new languages simply by replacing the Optical Character Recognition (OCR) model without modifying the core structure recognition model. Key innovations include: position-aware pixel-precise tokens for deterministic cell reconstruction; speculative decoding for faster sequence generation, and training-only box supervision to stabilize spatial grounding; region-based image segmentation. To mitigate data scarcity we propose a pipeline for generating a large synthetic table dataset. Experimental results validate each component. To address the limitations of existing evaluation methods we introduce TEDS struct100 and TEDS 100 metrics. Speculative decoding approach significantly improves recognition speed while maintaining accuracy. Finally, the combined techniques enable a mobileoptimized model that is more than three times faster than the full-size version.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- TSRFormer: Table Structure Recognition with TransformersWeihong Lin, Zheng Sun, Chixiang Ma, Mingze Li 等ACM MM 2022 · 被引用 56 次
- Neural Collaborative Graph Machines for Table Structure RecognitionHao Liu, Xin Li, Bing Liu, Deqiang Jiang 等CVPR 2022 · 被引用 45 次
- OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table RecognitionJianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu 等CVPR 2024 · 被引用 29 次
- GridFormer: Towards Accurate Table Structure Recognition via Grid PredictionPengyuan Lyu, Weihong Ma, Hongyi Wang, Yuechen Yu 等ACM MM 2023 · 被引用 17 次
相关 Paper
- Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate ModelingYongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li 等CVPR 2023
- TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual AlignmentChunxia Qin, Chenyu Liu, Pengcheng Xia, Jun Du 等CVPR 2026 · 被引用 2 次
- LORE: Logical Location Regression Network for Table Structure RecognitionHangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu 等AAAI 2023 · 被引用 43 次
- PixT3: Pixel-based Table-To-Text GenerationIñigo Alonso, Eneko Agirre, Mirella LapataACL 2024 · 被引用 2 次
- TGRNet: A Table Graph Reconstruction Network for Table Structure RecognitionWenyuan Xue, Baosheng Yu, Wen Wang, Dacheng Tao 等ICCV 2021 · 被引用 65 次
