ToMMeR - Efficient Entity Mention Detection from Large Language Models
Victor Morand, Nadi Tomeh, Josiane Mothe, Benjamin Piwowarski
Abstract
Identifying which text spans refer to entities - mention detection - is both foundational for information extraction and a known performance bottleneck. We introduce ToMMeR, a lightweight model (<300K parameters) probing mention detection capabilities from early LLM layers. Across 13 NER benchmarks, ToMMeR achieves 93% recall zero-shot, with an estimated 90% precision under a human-calibrated LLM-judge protocol, showing that ToMMeR rarely produces spurious predictions despite high recall. Cross-model analysis reveals that diverse architectures (14M-15B parameters) converge on similar mention boundaries (DICE>75%), confirming that mention detection emerges naturally from language modeling. When extended with span classification heads, ToMMeR achieves competitive NER performance (80-87% F1 on standard benchmarks). Our work provides evidence that structured entity representations exist in early transformer layers and can be efficiently recovered with minimal parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a89c8138-cc15-43f6-8da7-0419109cf931Builds on14
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- CrossNER: Evaluating Cross-Domain Named Entity RecognitionZihan Liu, Yan Xu, Tiezheng Yu, Wenliang Dai et al.AAAI 2021 · 201 citations
- Universal Information Extraction as Unified Semantic MatchingJie Lou, Yaojie Lu, Dai Dai, Wei Jia et al.AAAI 2023 · 96 citations
- How do Language Models Bind Entities in Context?Jiahai Feng, Jacob SteinhardtICLR 2024 · 81 citations
Related papers
- Multimodal Language Models See Better When They Look ShallowerHaoran Chen, Junyan Lin, Xinghao Chen, Yue Fan et al.EMNLP 2025
- Seq2seq is All You Need for Coreference ResolutionWenzheng Zhang, Sam Wiseman, Karl StratosEMNLP 2023 · 5 citations
- Mention Memory: incorporating textual knowledge into Transformers through entity mention attentionMichiel de Jong, Yury Zemlyanskiy, Nicholas FitzGerald, Fei Sha et al.ICLR 2022 · 55 citations
- ExtEnD: Extractive Entity DisambiguationEdoardo Barba, Luigi Procopio, Roberto NavigliACL 2022
- NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated DataSergei Bogdanov, Alexandre Constantin, Timothée Bernard, Benoît Crabbé et al.EMNLP 2024 · 29 citations
