Code and Named Entity Recognition in StackOverflow
Jeniya Tabassum, Mounica Maddela, Wei Xu, Alan Ritter
Abstract
There is an increasing interest in studying natural language and computer code together, as large corpora of programming texts become readily available on the Internet. For example, StackOverflow currently has over 15 million programming related questions written by 8.5 million users. Meanwhile, there is still a lack of fundamental NLP techniques for identifying code tokens or software-related named entities that appear within natural language sentences. In this paper, we introduce a new named entity recognition (NER) corpus for the computer programming domain, consisting of 15,372 sentences annotated with 20 fine-grained entity types. We trained indomain BERT representations (BERTOverflow) on 152 million sentences from Stack-Overflow, which lead to an absolute increase of +10 F 1 score over off-the-shelf BERT. We also present the SoftNER model which achieves an overall 79.10 F 1 score for code and named entity recognition on StackOverflow data. Our SoftNER model incorporates a context-independent code token classifier with corpus-level features to improve the BERTbased tagging model. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen et al.ICLR 2024 · 118 citations
- AUGER: automatically generating review comments with pre-training modelsLingwei Li, Li Yang, Huaxi Jiang, Jun Yan et al.FSE 2022 · 56 citations
- Fast Changeset-based Bug Localization with BERTAgnieszka Ciborowska, Kostadin DamevskiICSE 2022 · 53 citations
- SemParser: A Semantic Parser for Log AnalyticsYintong Huo, Yuxin Su, Cheryl Lee, Michael R. LyuICSE 2023 · 51 citations
- MosaicBERT: A Bidirectional Encoder Optimized for Fast PretrainingJacob P. Portes, Alexander Trott, Sam Havens, Daniel King et al.NeurIPS 2023 · 46 citations
Related papers
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 438 citations
- Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra et al.ACL 2023 · 24 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä et al.EMNLP 2025 · 2 citations
- Software Entity Recognition with Noise-Robust LearningTai Nguyen, Yifeng Di, Joohan Lee, Muhao Chen et al.ASE 2023 · 2 citations
