Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM
Mengjie Liu, Jiahui Peng, Wenchang Ning, Pei Chu, Jiantao Qiu, Ren Ma, He Zhu, Rui Min, Lindong Lu, Linfeng Hou, Kaiwen Liu, Yuan Qu
Abstract
High-quality main content extraction from web pages is a critical prerequisite for constructing large-scale training corpora. While traditional heuristic extractors are efficient, they lack the semantic reasoning required to handle the structural heterogeneity of the modern web. Conversely, well-pretrained generative Large Language Models (LLMs) offer superior document comprehension but remain impractical at web scale due to excessive computational costs, limited context windows, and hallucination risks. We present Dripper, a lightweight framework that resolves these bottlenecks through four contributions: (1) We reformulate extraction as a constrained sequence labeling task using SLMs (Small Language Models). This paradigm eliminates generative hallucinations and achieves exceptional efficiency, reaching a throughput of 3.08 pages per second on a single A100 GPU. (2) We construct WebMainBench, a rigorous benchmark of 7,809 human-annotated pages covering 5,434 unique domains and multiple languages. Evaluations show our Dripper-0.6B model outperforms heuristics like Trafilatura and rivals massive models like DeepSeek-V3.2 (685B), GPT-5, and Gemini-2.5-Pro, offering an optimal efficiency-accuracy trade-off. (3) We demonstrate infrastructural value by pre-training a 1B model on a Dripper-curated corpus (63B tokens). This model significantly outperforms baselines in downstream tasks, providing evidence for the critical role of extraction quality and the effectiveness of our framework. (4) We open-source the Dripper-0.6B weights and codebase to facilitate the construction of high-quality datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb2f74bf-6fc3-448b-a794-c23960d18d82Builds on5
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng et al.WWW 2022 · 88 citations
- Organize the Web: Constructing Domains Enhances Pre-Training Data CurationAlexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi et al.ICML 2025
Related papers
- GneissWeb: Preparing High Quality Data for LLMs at ScaleHajar Emami Gohari, Swanand Ravindra Kadhe, Yousaf Shah, Constantin M Adam et al.ICLR 2026 · 7 citations
- Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-PagesRitesh Sarkhel, Binxuan Huang, Colin Lockard, Prashant ShiralkarVLDB 2023 · 11 citations
- SWEb: A Large Web Dataset for the Scandinavian LanguagesTobias Norlund, Tim Isbister, Amaru Cuba Gyllensten, Paul Gabriel dos Santos et al.ICLR 2025
- Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk et al.ACL 2024
- GTA: Generating Long-horizon Tasks for Web Agents at ScaleTenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou et al.ACL 2026
