WIERT: Web Information Extraction via Render Tree
Zimeng Li, Bo Shao, Linjun Shou, Ming Gong, Gen Li, Daxin Jiang
摘要
Web information extraction (WIE) is a fundamental problem in web document understanding, with a significant impact on various applications. Visual information plays a crucial role in WIE tasks as the nodes containing relevant information are often visually distinct, such as being in a larger font size or having a brighter color, from the other nodes. However, rendering visual information of a web page can be computationally expensive. Previous works have mainly focused on the Document Object Model (DOM) tree, which lacks visual information. To efficiently exploit visual information, we propose leveraging the render tree, which combines the DOM tree and Cascading Style Sheets Object Model (CS-SOM) tree, and contains not only content and layout information but also rich visual information at a little additional acquisition cost compared to the DOM tree. In this paper, we present WIERT, a method that effectively utilizes the render tree of a web page based on a pretrained language model. We evaluate WIERT on the Klarna product page dataset, a manually labeled dataset of renderable e-commerce web pages, demonstrating its effectiveness and robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Semantic Constraint Inference for Web Form Test GenerationParsa Alian, Noor Nashid, Mobina Shahbandeh, Ali MesbahISSTA 2024 · 被引用 1 次
- LiveWeb-IE: A Benchmark For Online Web Information ExtractionSeungbin Yang, Jihwan Kim, Jaemin Choi, Dongjin Kim 等ICLR 2026 · 被引用 1 次
- Approximation to Smooth Functions by Low-Rank Swish NetworksZimeng Li, Hongjun Li, Jingyuan Wang, Ke TangICML 2025
它引用的顶会 Paper5
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng 等WWW 2022 · 被引用 88 次
- MarkupLM: Pre-training of Text and Markup Language for Visually Rich Document UnderstandingJunlong Li, Yiheng Xu, Lei Cui, Furu WeiACL 2022 · 被引用 75 次
- WebSRC: A Dataset for Web-Based Structural Reading ComprehensionXingyu Chen, Zihan Zhao, Lu Chen, Jiabao Ji 等EMNLP 2021 · 被引用 42 次
- FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web DocumentsBill Yuchen Lin, Ying Sheng, Nguyen Vo, Sandeep TataKDD 2020 · 被引用 31 次
相关 Paper
- GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render TreeZirui Shao, Feiyu Gao, Zhongda Qi, Hangdi Xing 等EMNLP 2023 · 被引用 3 次
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian 等SIGIR 2022 · 被引用 30 次
- Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource SettingsZhenyu Zhang, Bowen Yu, Tingwen Liu, Tianyun Liu 等WWW 2023 · 被引用 7 次
- MUSTIE: Multimodal Structural Transformer for Web Information ExtractionQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng 等ACL 2023 · 被引用 16 次
- Enhancing Vision-Language Pre-Training with Rich SupervisionsYuan Gao, Kunyu Shi, Pengkai Zhu, Edouard Belval 等CVPR 2024
