NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in Norwegian
Peng Liu, Lemei Zhang, Terje Nissen Farup, Even W. Lauvrak, Jon Espen Ingvaldsen, Simen Eide, Jon Atle Gulla, Zhirong Yang
摘要
Norwegian, spoken by only 5 million population, is under-representative within the most impressive breakthroughs in NLP tasks. To the best of our knowledge, there has not yet been a comprehensive evaluation of the existing language models (LMs) on Norwegian generation tasks during the article writing process. To fill this gap, we 1) compiled the existing Norwegian dataset and pre-trained 4 Norwegian Open Language Models varied from parameter scales and architectures, collectively called NorGLM; 2) introduced a comprehensive benchmark, NLEBench, for evaluating natural language generation capabilities in Norwegian, encompassing translation and human annotation. Based on the investigation, we find that: 1) the mainstream, English-dominated LM GPT-3.5 has limited capability in understanding the Norwegian context; 2) the increase in model parameter scales demonstrates limited impact on the performance of downstream tasks when the pre-training dataset is constrained in size; 3) smaller models also demonstrate the reasoning capability through Chain-of-Thought; 4) a multi-task dataset that includes synergy tasks can be used to verify the generalizability of LLMs on natural language understanding and, meanwhile, test the interconnectedness of these NLP tasks. We share our resources and code for reproducibility 1 under a CC BY-NC 4.0 license.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai 等ICLR 2024 · 被引用 254 次
相关 Paper
- FinGPT: Large Generative Models for a Small LanguageRisto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen 等EMNLP 2023 · 被引用 9 次
- SWEb: A Large Web Dataset for the Scandinavian LanguagesTobias Norlund, Tim Isbister, Amaru Cuba Gyllensten, Paul Gabriel dos Santos 等ICLR 2025
- Benchmarking Large Language Model Capabilities for Conditional GenerationJoshua Maynez, Priyanka Agrawal, Sebastian GehrmannACL 2023 · 被引用 6 次
- VMLU Benchmarks: A comprehensive benchmark toolkit for Vietnamese LLMsCuc Thi Bui, Nguyen Truong Son, Trang Van Truong, Viet Lam Phung 等ACL 2025
- CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 LanguagesYilun Yang, Yekun ChaiEMNLP 2025 · 被引用 1 次
