FiNER: Financial Numeric Entity Recognition for XBRL Tagging
Lefteris Loukas, Manos Fergadiotis, Ilias Chalkidis, Eirini Spyropoulou, Prodromos Malakasiotis, Ion Androutsopoulos, Georgios Paliouras
Abstract
Publicly traded companies are required to submit periodic reports with eXtensive Business Reporting Language (XBRL) word-level tags. Manually tagging the reports is tedious and costly. We, therefore, introduce XBRL tagging as a new entity extraction task for the financial domain and release FiNER-139, a dataset of 1.1M sentences with gold XBRL tags. Unlike typical entity extraction datasets, FiNER-139 uses a much larger label set of 139 entity types. Most annotated tokens are numeric, with the correct tag per token depending mostly on context, rather than the token itself. We show that subword fragmentation of numeric expressions harms BERT’s performance, allowing word-level BILSTMs to perform better. To improve BERT’s performance, we propose two simple and effective solutions that replace numeric expressions with pseudo-tokens reflecting original token shapes and numeric magnitudes. We also experiment with FIN-BERT, an existing BERT model for the financial domain, and release our own BERT (SEC-BERT), pre-trained on financial filings, which performs best. Through data and error analysis, we finally identify possible limitations to inspire future work on XBRL tagging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language ModelsQizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma et al.ICLR 2026 · 374 citations
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy et al.ACL 2024 · 10 citations
- Linking Industry Sectors and Financial Statements: A Hybrid Approach for Company ClassificationGuy Stephane Waffo Dzuyo, Gaël Guibon, Christophe Cerisara, Luis Belmar-LetelierAAAI 2025 · 1 citation
- RankGuess: Password Guessing Using Adversarial RankingTao Yang, Ding WangS&P 2025
Builds on2
- Injecting Numerical Reasoning Skills into Language ModelsMor Geva, Ankit Gupta, Jonathan BerantACL 2020 · 12 citations
- An Empirical Study on Large-Scale Multi-Label Text Classification Including Few and Zero-Shot LabelsIlias Chalkidis, Manos Fergadiotis, Sotiris Kotitsas, Prodromos Malakasiotis et al.EMNLP 2020 · 2 citations
Related papers
- Numerical Tuple Extraction from Tables with Pre-trainingQingping Yang, Yixuan Cao, Ping LuoKDD 2022 · 1 citation
- DILBERT: Customized Pre-Training for Domain Adaptation with Category Shift, with an Application to Aspect ExtractionEntony Lekhtman, Yftah Ziser, Roi ReichartEMNLP 2021 · 25 citations
- Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERTAkshay Smit, Saahil Jain, Pranav Rajpurkar, Anuj Pareek et al.EMNLP 2020 · 212 citations
- When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial DomainRaj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah et al.EMNLP 2022 · 63 citations
- Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive MimickingTimo Schick, Hinrich SchützeAAAI 2020 · 106 citations
