FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web Documents
Bill Yuchen Lin, Ying Sheng, Nguyen Vo, Sandeep Tata
摘要
Extracting structured data from HTML documents is a long-studied problem with a broad range of applications like augmenting knowledge bases, supporting faceted search, and providing domain-specific experiences for key verticals like shopping and movies. Previous approaches have either required a small number of examples for each target site or relied on carefully handcrafted heuristics built over visual renderings of websites. In this paper, we present a novel two-stage neural approach, named FreeDOM, which overcomes both these limitations. The first stage learns a representation for each DOM node in the page by combining both the text and markup information. The second stage captures longer range distance and semantic relatedness using a relational neural network. By combining these stages, FreeDOM is able to generalize to unseen sites after training on a small number of seed sites from that vertical without requiring expensive hand-crafted features over visual renderings of the page. Through experiments on a public dataset with 8 different verticals, we show that FreeDOM beats the previous state of the art by nearly 3.7 F1 points on average without requiring features over rendered pages or expensive hand-crafted features.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng 等WWW 2022 · 被引用 88 次
- MarkupLM: Pre-training of Text and Markup Language for Visually Rich Document UnderstandingJunlong Li, Yiheng Xu, Lei Cui, Furu WeiACL 2022 · 被引用 75 次
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture SearchJin Xu, Xu Tan, Renqian Luo, Kaitao Song 等KDD 2021 · 被引用 49 次
- Web question answering with neurosymbolic program synthesisQiaochu Chen, Aaron Lamoreaux, Xinyu Wang, Greg Durrett 等PLDI 2021 · 被引用 25 次
- Data Extraction via Semantic Regular Expression SynthesisQiaochu Chen, Arko Banerjee, Çagatay Demiralp, Greg Durrett 等OOPSLA 2023 · 被引用 24 次
它引用的顶会 Paper2
- NERO: A Neural Rule Grounding Framework for Label-Efficient Relation ExtractionWenxuan Zhou, Hongtao Lin, Bill Yuchen Lin, Ziqi Wang 等WWW 2020 · 被引用 57 次
- ZeroShotCeres: Zero-Shot Relation Extraction from Semi-Structured WebpagesColin Lockard, Prashant Shiralkar, Xin Luna Dong, Hannaneh HajishirziACL 2020 · 被引用 2 次
相关 Paper
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt 等ACL 2020 · 被引用 111 次
- MUSTIE: Multimodal Structural Transformer for Web Information ExtractionQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng 等ACL 2023 · 被引用 16 次
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian 等SIGIR 2022 · 被引用 30 次
- Reward-based Input Construction for Cross-document Relation ExtractionByeonghu Na, Suhyeon Jo, Yeongmin Kim, Il-Chul MoonACL 2024 · 被引用 3 次
- SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement LearningShicheng Liu, Kai Sun, Lisheng Fu, Xilun Chen 等ICLR 2026 · 被引用 2 次
