NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
Sergei Bogdanov, Alexandre Constantin, Timothée Bernard, Benoît Crabbé, Etienne Bernard
摘要
Large Language Models (LLMs) have shown impressive abilities in data annotation, opening the way for new approaches to solve classic NLP problems. In this paper, we show how to use LLMs to create NuNER, a compact language representation model specialized in the Named Entity Recognition (NER) task. NuNER can be fine-tuned to solve downstream NER problems in a data-efficient way, outperforming similar-sized foundation models in the few-shot regime and competing with much larger LLMs. We find that the size and entity-type diversity of the pre-training dataset are key to achieving good performance. We view NuNER as a member of the broader family of task-specific foundation models, recently unlocked by LLMs. NuNER and NuNER's dataset are open-sourced with MIT License 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ReflectDiffu: Reflect between Emotion-intent Contagion and Mimicry for Empathetic Response Generation via a RL-Diffusion FrameworkJiahao Yuan, Zixiang Di, Zhiqing Cui, Guisong Yang 等ACL 2025 · 被引用 6 次
- Cuckoo: An IE Free Rider Hatched by Massive Nutrition in LLM's NestLetian Peng, Zilong Wang, Feng Yao, Jingbo ShangACL 2025 · 被引用 1 次
- Can LLMs be Literary Companions?: Analysing LLMs on Bengali Figures of Speech IdentificationSourav Das, Kripabandhu GhoshEMNLP 2025
- DiZiNER: Disagreement-guided Instruction Refinement via Simulating Pilot Annotation for Zero-shot Named Entity RecognitionSiun Kim, Hyung-Jin YoonACL 2026
- Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NERAhmed Ewais, Ahmed Hashish, Amr AliACL 2026
它引用的顶会 Paper6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- GoLLIE: Annotation Guidelines improve Zero-Shot Information-ExtractionOscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle 等ICLR 2024 · 被引用 168 次
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen 等ICLR 2024 · 被引用 118 次
- Coarse-to-Fine Pre-training for Named Entity RecognitionMengge Xue, Bowen Yu, Zhenyu Zhang, Tingwen Liu 等EMNLP 2020 · 被引用 49 次
- Optimizing Bi-Encoder for Named Entity Recognition via Contrastive LearningSheng Zhang, Hao Cheng, Jianfeng Gao, Hoifung PoonICLR 2023 · 被引用 22 次
相关 Paper
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose 等EMNLP 2021 · 被引用 97 次
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 等EMNLP 2025 · 被引用 2 次
- OneNet: A Fine-Tuning Free Framework for Few-Shot Entity Linking via Large Language Model PromptingXukai Liu, Ye Liu, Kai Zhang, Kehang Wang 等EMNLP 2024 · 被引用 7 次
- DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity RecognitionHanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li 等EMNLP 2025
- GuideNER: Annotation Guidelines Are Better than Examples for In-Context Named Entity RecognitionShizhou Huang, Bo Xu, Yang Yu, Changqun Li 等AAAI 2025 · 被引用 1 次
