Neural Collaborative Graph Machines for Table Structure Recognition
Hao Liu, Xin Li, Bing Liu, Deqiang Jiang, Yinsong Liu, Bo Ren
Abstract
Recently, table structure recognition has achieved impressive progress with the help of deep graph models. Most of them exploit single visual cues of tabular elements or simply combine visual cues with other modalities via early fusion to reason their graph relationships. However, neither early fusion nor individually reasoning in terms of multiple modalities can be appropriate for all varieties of table structures with great diversity. Instead, different modalities are expected to collaborate with each other in different patterns for different table cases. In the community, the importance of intra-inter modality interactions for table structure reasoning is still unexplored. In this paper, we define it as heterogeneous table structure recognition (Hetero-TSR) problem. With the aim of filling this gap, we present a novel Neural Collaborative Graph Machines (NCGM) equipped with stacked collaborative blocks, which alternatively extracts intra-modality context and models inter-modality interactions in a hierarchical way. It can represent the intrainter modality relationships of tabular elements more robustly, which significantly improves the recognition performance. We also show that the proposed NCGM can modulate collaborative pattern of different modalities conditioned on the context of intra-modality cues, which is vital for diversified table cases. Experimental results on benchmarks demonstrate our proposed NCGM achieves state-ofthe-art performance and beats other contemporary methods by a large margin especially under challenging scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61ca7021-90cf-4753-8e1a-be50ba623bc3Cited by top-tier papers10
- TabPedia: Towards Comprehensive Visual Table Understanding with Concept SynergyWeichao Zhao, Hao Feng, Qi Liu, Jingqun Tang et al.NeurIPS 2024 · 97 citations
- LORE: Logical Location Regression Network for Table Structure RecognitionHangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu et al.AAAI 2023 · 43 citations
- GridFormer: Towards Accurate Table Structure Recognition via Grid PredictionPengyuan Lyu, Weihong Ma, Hongyi Wang, Yuechen Yu et al.ACM MM 2023 · 17 citations
- Grab What You Need: Rethinking Complex Table Structure Recognition with Flexible Components DeliberationHao Liu, Xin Li, Mingming Gong, Bing Liu et al.AAAI 2024 · 11 citations
- Relational Representation Learning in Visually-Rich DocumentsXin Li, Yan Zheng, Yiqing Hu, Haoyu Cao et al.ACM MM 2022 · 7 citations
Builds on7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
Related papers
- Web Table Retrieval using Multimodal Deep LearningRoee Shraga, Haggai Roitman, Guy Feigenblat, Mustafa CanimSIGIR 2020 · 45 citations
- TabGLM: Tabular Graph Language Model for Learning Transferable Representations Through Multi-Modal Consistency MinimizationAnay Majee, Maria Xenochristou, Wei-Peng ChenAAAI 2025 · 3 citations
- TableVLM: Multi-modal Pre-training for Table Structure RecognitionLeiyuan Chen, Chengsong Huang, Xiaoqing Zheng, Jinshu Lin et al.ACL 2023 · 8 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- M3TR: Multi-modal Multi-label Recognition with TransformerJiawei Zhao, Yifan Zhao, Jia LiACM MM 2021 · 45 citations
