End-to-End Hyper-Relational Information Extraction for Engineering Diagrams via Dynamically Tokenized Relation Transformer
Tianyou Bai, Yan-Ming Zhang, Zixiang Zhang, Jibin Zhou, Fei Yin, Cheng-Lin Liu
摘要
Engineering diagrams are the core carriers of technical information in industrial contexts, where the pressing demand for their digitization from industrial sectors has driven great advancements in related research domains. However, existing research still suffers from three limitations. Firstly, the detection of symbols, lines, and texts typically involves multiple independent models, resulting in cumbersome workflows. In addition, high-resolution diagrams often impose an excessive computational cost on existing models. Moreover, parsing frameworks solely based on object detection can merely localize component positions, yet fail to capture the topological connection semantics and structured knowledge among components, thus offering limited convenience for industrial applications. To address these issues, we propose an end-to-end information extraction framework based on the Dynamically Tokenized Relation Transformer (DTRT), which can dynamically reduce received image tokens, filter redundant information, and efficiently extract structural knowledge to construct hyper-relational knowledge graphs. We practiced our model on piping and instrumentation diagrams (P&IDs) and electrical diagrams (EDs): the former are widely used in chemical engineering enterprises, while the latter are employed to describe circuit systems. DTRT achieves an R@1000 accuracy of 94.84% on PIDs and R@200 accuracy of 92.52% on EDs with a significantly reduced computational cost. Our code is available at https://github.com/Tianyou-Bai/ DTRT-Diagram-Parsing-in-Scene-Graph.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen 等NeurIPS 2020 · 被引用 2,118 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
相关 Paper
- GPTR: Gestalt-Perception Transformer for Diagram Object DetectionXin Hu, Lingling Zhang, Jun Liu, Jinfu Fan 等AAAI 2023 · 被引用 9 次
- A Tree-Based Structure-Aware Transformer Decoder for Image-To-Markup GenerationShuhan Zhong, Sizhe Song, Guanyao Li, S.-H. Gary ChanACM MM 2022 · 被引用 17 次
- Hierarchical Multi-Task Learning for Diagram Question Answering with Multi-Modal TransformerZhaoquan Yuan, Xiao Peng, Xiao Wu, Changsheng XuACM MM 2021 · 被引用 10 次
- GAT-CADNet: Graph Attention Network for Panoptic Symbol Spotting in CAD DrawingsZhaohua Zheng, Jianfang Li, Lingjie Zhu, Honghua Li 等CVPR 2022 · 被引用 19 次
- TableFormer: Table Structure Understanding with TransformersAhmed S. Nassar, Nikolaos Livathinos, Maksym Lysak, Peter W. J. StaarCVPR 2022 · 被引用 5 次
