G2LFormer: Global-to-Local Query Enhancement for Robust Table Structure Recognition
Haosheng Cai, Yang Xue
摘要
Table structure recognition (TSR), the task of extracting logical and physical structures from table images, is critical for document understanding. Current end-to-end image-to-text methods typically employ a top-down strategy where physical structure prediction depends on the logical decoder's output sequence. However, this process often suffers from training instability and misalignment between predicted bounding boxes and ground-truth cell positions. To address this issue, we propose G2LFormer, a novel transformer-based framework that employs a ''Global-to-Local'' query enhancement strategy. Specifically, G2LFormer introduces a Vision-guided Query Enhancer to integrate both textual and visual modalities, significantly improving the overall query representation capability and boosting prediction accuracy. Additionally, we design a Multi-scale Manhattan Vision-guider that leverages a spatial attenuation matrix to guide each query towards its corresponding cell location, effectively balancing local and global information for more precise bounding box generation. Extensive experiments on benchmark datasets demonstrate G2LFormer's superior performance, while ablation studies confirming the significant contribution of each proposed module in achieving state-of-the-art results. The source code and model have been released at: https://github.com/Hzbupahaozi/G2LFormer.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate ModelingYongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li 等CVPR 2023
- TSRFormer: Table Structure Recognition with TransformersWeihong Lin, Zheng Sun, Chixiang Ma, Mingze Li 等ACM MM 2022 · 被引用 56 次
- LORE: Logical Location Regression Network for Table Structure RecognitionHangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu 等AAAI 2023 · 被引用 43 次
- TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual AlignmentChunxia Qin, Chenyu Liu, Pengcheng Xia, Jun Du 等CVPR 2026 · 被引用 2 次
- Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware GuidanceYingjie Zhu, Xuefeng Bai, Kehai Chen, Yang Xiang 等ICML 2026 · 被引用 3 次
