Webformer: Pre-training with Web Pages for Information Retrieval
Yu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian, Xinyu Zhang, Hao Jiang, Zhao Cao, Zhicheng Dou
摘要
Pre-trained language models (PLMs) have achieved great success in the area of Information Retrieval. Studies show that applying these models to ad-hoc document ranking can achieve better retrieval effectiveness. However, on the Web, most information is organized in the form of HTML web pages. In addition to the pure text content, the structure of the content organized by HTML tags is also an important part of the information delivered on a web page. Currently, such structured information is totally ignored by pre-trained models which are trained solely based on text content. In this paper, we propose to leverage large-scale web pages and their DOM (Document Object Model) tree structures to pre-train models for information retrieval. We argue that using the hierarchical structure contained in web pages, we can get richer contextual information for training better language models. To exploit this kind of information, we devise four pre-training objectives based on the structure of web pages, then pre-train a Transformer model towards these tasks jointly with traditional masked language model objective. Experimental results on two authoritative ad-hoc retrieval datasets prove that our model can significantly improve ranking performance compared to existing pre-trained models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG SystemsJiejun Tan, Zhicheng Dou, Wen Wang, Mang Wang 等WWW 2025 · 被引用 42 次
- Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-PagesRitesh Sarkhel, Binxuan Huang, Colin Lockard, Prashant ShiralkarVLDB 2023 · 被引用 11 次
- SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine OptimizationSunghwan Kim, Wooseok Jeong, Serin Kim, Sangam Lee 等KDD 2026 · 被引用 8 次
- PSLOG: Pretraining with Search Logs for Document RankingZhan Su, Zhicheng Dou, Yujia Zhou, Ziyuan Zhao 等KDD 2023 · 被引用 2 次
- Semantic Constraint Inference for Web Form Test GenerationParsa Alian, Noor Nashid, Mobina Shahbandeh, Ali MesbahISSTA 2024 · 被引用 1 次
它引用的顶会 Paper10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang 等ICLR 2020 · 被引用 325 次
相关 Paper
- H-ERNIE: A Multi-Granularity Pre-Trained Language Model for Web SearchXiaokai Chu, Jiashu Zhao, Lixin Zou, Dawei YinSIGIR 2022 · 被引用 11 次
- Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource SettingsZhenyu Zhang, Bowen Yu, Tingwen Liu, Tianyun Liu 等WWW 2023 · 被引用 7 次
- Wikiformer: Pre-training with Structured Information of Wikipedia for Ad-Hoc RetrievalWeihang Su, Qingyao Ai, Xiangsheng Li, Jia Chen 等AAAI 2024
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng 等WWW 2022 · 被引用 88 次
- Table Search Using a Deep Contextualized Language ModelZhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu 等SIGIR 2020 · 被引用 48 次
