WIERT: Web Information Extraction via Render Tree
Zimeng Li, Bo Shao, Linjun Shou, Ming Gong, Gen Li, Daxin Jiang
Abstract
Web information extraction (WIE) is a fundamental problem in web document understanding, with a significant impact on various applications. Visual information plays a crucial role in WIE tasks as the nodes containing relevant information are often visually distinct, such as being in a larger font size or having a brighter color, from the other nodes. However, rendering visual information of a web page can be computationally expensive. Previous works have mainly focused on the Document Object Model (DOM) tree, which lacks visual information. To efficiently exploit visual information, we propose leveraging the render tree, which combines the DOM tree and Cascading Style Sheets Object Model (CS-SOM) tree, and contains not only content and layout information but also rich visual information at a little additional acquisition cost compared to the DOM tree. In this paper, we present WIERT, a method that effectively utilizes the render tree of a web page based on a pretrained language model. We evaluate WIERT on the Klarna product page dataset, a manually labeled dataset of renderable e-commerce web pages, demonstrating its effectiveness and robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Semantic Constraint Inference for Web Form Test GenerationParsa Alian, Noor Nashid, Mobina Shahbandeh, Ali MesbahISSTA 2024 · 1 citation
- LiveWeb-IE: A Benchmark For Online Web Information ExtractionSeungbin Yang, Jihwan Kim, Jaemin Choi, Dongjin Kim et al.ICLR 2026 · 1 citation
- Approximation to Smooth Functions by Low-Rank Swish NetworksZimeng Li, Hongjun Li, Jingyuan Wang, Ke TangICML 2025
Builds on5
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng et al.WWW 2022 · 88 citations
- MarkupLM: Pre-training of Text and Markup Language for Visually Rich Document UnderstandingJunlong Li, Yiheng Xu, Lei Cui, Furu WeiACL 2022 · 75 citations
- WebSRC: A Dataset for Web-Based Structural Reading ComprehensionXingyu Chen, Zihan Zhao, Lu Chen, Jiabao Ji et al.EMNLP 2021 · 42 citations
- FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web DocumentsBill Yuchen Lin, Ying Sheng, Nguyen Vo, Sandeep TataKDD 2020 · 31 citations
Related papers
- GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render TreeZirui Shao, Feiyu Gao, Zhongda Qi, Hangdi Xing et al.EMNLP 2023 · 3 citations
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian et al.SIGIR 2022 · 30 citations
- Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource SettingsZhenyu Zhang, Bowen Yu, Tingwen Liu, Tianyun Liu et al.WWW 2023 · 7 citations
- MUSTIE: Multimodal Structural Transformer for Web Information ExtractionQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng et al.ACL 2023 · 16 citations
- Enhancing Vision-Language Pre-Training with Rich SupervisionsYuan Gao, Kunyu Shi, Pengkai Zhu, Edouard Belval et al.CVPR 2024
