NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
Sergei Bogdanov, Alexandre Constantin, Timothée Bernard, Benoît Crabbé, Etienne Bernard
Abstract
Large Language Models (LLMs) have shown impressive abilities in data annotation, opening the way for new approaches to solve classic NLP problems. In this paper, we show how to use LLMs to create NuNER, a compact language representation model specialized in the Named Entity Recognition (NER) task. NuNER can be fine-tuned to solve downstream NER problems in a data-efficient way, outperforming similar-sized foundation models in the few-shot regime and competing with much larger LLMs. We find that the size and entity-type diversity of the pre-training dataset are key to achieving good performance. We view NuNER as a member of the broader family of task-specific foundation models, recently unlocked by LLMs. NuNER and NuNER's dataset are open-sourced with MIT License 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- ReflectDiffu: Reflect between Emotion-intent Contagion and Mimicry for Empathetic Response Generation via a RL-Diffusion FrameworkJiahao Yuan, Zixiang Di, Zhiqing Cui, Guisong Yang et al.ACL 2025 · 6 citations
- Cuckoo: An IE Free Rider Hatched by Massive Nutrition in LLM's NestLetian Peng, Zilong Wang, Feng Yao, Jingbo ShangACL 2025 · 1 citation
- Can LLMs be Literary Companions?: Analysing LLMs on Bengali Figures of Speech IdentificationSourav Das, Kripabandhu GhoshEMNLP 2025
- DiZiNER: Disagreement-guided Instruction Refinement via Simulating Pilot Annotation for Zero-shot Named Entity RecognitionSiun Kim, Hyung-Jin YoonACL 2026
- Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NERAhmed Ewais, Ahmed Hashish, Amr AliACL 2026
Builds on6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- GoLLIE: Annotation Guidelines improve Zero-Shot Information-ExtractionOscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle et al.ICLR 2024 · 168 citations
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen et al.ICLR 2024 · 118 citations
- Coarse-to-Fine Pre-training for Named Entity RecognitionMengge Xue, Bowen Yu, Zhenyu Zhang, Tingwen Liu et al.EMNLP 2020 · 49 citations
- Optimizing Bi-Encoder for Named Entity Recognition via Contrastive LearningSheng Zhang, Hao Cheng, Jianfeng Gao, Hoifung PoonICLR 2023 · 22 citations
Related papers
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose et al.EMNLP 2021 · 97 citations
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä et al.EMNLP 2025 · 2 citations
- OneNet: A Fine-Tuning Free Framework for Few-Shot Entity Linking via Large Language Model PromptingXukai Liu, Ye Liu, Kai Zhang, Kehang Wang et al.EMNLP 2024 · 7 citations
- DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity RecognitionHanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li et al.EMNLP 2025
- GuideNER: Annotation Guidelines Are Better than Examples for In-Context Named Entity RecognitionShizhou Huang, Bo Xu, Yang Yu, Changqun Li et al.AAAI 2025 · 1 citation
