FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web Documents
Bill Yuchen Lin, Ying Sheng, Nguyen Vo, Sandeep Tata
Abstract
Extracting structured data from HTML documents is a long-studied problem with a broad range of applications like augmenting knowledge bases, supporting faceted search, and providing domain-specific experiences for key verticals like shopping and movies. Previous approaches have either required a small number of examples for each target site or relied on carefully handcrafted heuristics built over visual renderings of websites. In this paper, we present a novel two-stage neural approach, named FreeDOM, which overcomes both these limitations. The first stage learns a representation for each DOM node in the page by combining both the text and markup information. The second stage captures longer range distance and semantic relatedness using a relational neural network. By combining these stages, FreeDOM is able to generalize to unseen sites after training on a small number of seed sites from that vertical without requiring expensive hand-crafted features over visual renderings of the page. Through experiments on a public dataset with 8 different verticals, we show that FreeDOM beats the previous state of the art by nearly 3.7 F1 points on average without requiring features over rendered pages or expensive hand-crafted features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c7a73ca-02ec-42c7-a66d-7e640ad9f0c3Cited by top-tier papers14
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng et al.WWW 2022 · 88 citations
- MarkupLM: Pre-training of Text and Markup Language for Visually Rich Document UnderstandingJunlong Li, Yiheng Xu, Lei Cui, Furu WeiACL 2022 · 75 citations
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture SearchJin Xu, Xu Tan, Renqian Luo, Kaitao Song et al.KDD 2021 · 49 citations
- Web question answering with neurosymbolic program synthesisQiaochu Chen, Aaron Lamoreaux, Xinyu Wang, Greg Durrett et al.PLDI 2021 · 25 citations
- Data Extraction via Semantic Regular Expression SynthesisQiaochu Chen, Arko Banerjee, Çagatay Demiralp, Greg Durrett et al.OOPSLA 2023 · 24 citations
Builds on2
- NERO: A Neural Rule Grounding Framework for Label-Efficient Relation ExtractionWenxuan Zhou, Hongtao Lin, Bill Yuchen Lin, Ziqi Wang et al.WWW 2020 · 57 citations
- ZeroShotCeres: Zero-Shot Relation Extraction from Semi-Structured WebpagesColin Lockard, Prashant Shiralkar, Xin Luna Dong, Hannaneh HajishirziACL 2020 · 2 citations
Related papers
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt et al.ACL 2020 · 111 citations
- MUSTIE: Multimodal Structural Transformer for Web Information ExtractionQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng et al.ACL 2023 · 16 citations
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian et al.SIGIR 2022 · 30 citations
- Reward-based Input Construction for Cross-document Relation ExtractionByeonghu Na, Suhyeon Jo, Yeongmin Kim, Il-Chul MoonACL 2024 · 3 citations
- SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement LearningShicheng Liu, Kai Sun, Lisheng Fu, Xilun Chen et al.ICLR 2026 · 2 citations
