SWEb: A Large Web Dataset for the Scandinavian Languages
Tobias Norlund, Tim Isbister, Amaru Cuba Gyllensten, Paul Gabriel dos Santos, Danila Petrelli, Ariel Ekgren, Magnus Sahlgren
摘要
This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel model-based text extractor that significantly reduces complexity in comparison with rule-based approaches. We also introduce a new cloze-style benchmark for evaluating language models in Swedish, and use this test to compare models trained on the SWEb data to models trained on FineWeb, with competitive results. All data, models and code are shared openly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in NorwegianPeng Liu, Lemei Zhang, Terje Nissen Farup, Even W. Lauvrak 等EMNLP 2024 · 被引用 1 次
- Superlim: A Swedish Language Understanding Evaluation BenchmarkAleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger 等EMNLP 2023 · 被引用 2 次
- FinGPT: Large Generative Models for a Small LanguageRisto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen 等EMNLP 2023 · 被引用 9 次
- GneissWeb: Preparing High Quality Data for LLMs at ScaleHajar Emami Gohari, Swanand Ravindra Kadhe, Yousaf Shah, Constantin M Adam 等ICLR 2026 · 被引用 7 次
- Multilingual Language Model Pretraining using Machine-translated DataJiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin 等EMNLP 2025
