FiNER: Financial Numeric Entity Recognition for XBRL Tagging
Lefteris Loukas, Manos Fergadiotis, Ilias Chalkidis, Eirini Spyropoulou, Prodromos Malakasiotis, Ion Androutsopoulos, Georgios Paliouras
摘要
Publicly traded companies are required to submit periodic reports with eXtensive Business Reporting Language (XBRL) word-level tags. Manually tagging the reports is tedious and costly. We, therefore, introduce XBRL tagging as a new entity extraction task for the financial domain and release FiNER-139, a dataset of 1.1M sentences with gold XBRL tags. Unlike typical entity extraction datasets, FiNER-139 uses a much larger label set of 139 entity types. Most annotated tokens are numeric, with the correct tag per token depending mostly on context, rather than the token itself. We show that subword fragmentation of numeric expressions harms BERT’s performance, allowing word-level BILSTMs to perform better. To improve BERT’s performance, we propose two simple and effective solutions that replace numeric expressions with pseudo-tokens reflecting original token shapes and numeric magnitudes. We also experiment with FIN-BERT, an existing BERT model for the financial domain, and release our own BERT (SEC-BERT), pre-trained on financial filings, which performs best. Through data and error analysis, we finally identify possible limitations to inspire future work on XBRL tagging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language ModelsQizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma 等ICLR 2026 · 被引用 374 次
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy 等ACL 2024 · 被引用 10 次
- Linking Industry Sectors and Financial Statements: A Hybrid Approach for Company ClassificationGuy Stephane Waffo Dzuyo, Gaël Guibon, Christophe Cerisara, Luis Belmar-LetelierAAAI 2025 · 被引用 1 次
- RankGuess: Password Guessing Using Adversarial RankingTao Yang, Ding WangS&P 2025
它引用的顶会 Paper2
- Injecting Numerical Reasoning Skills into Language ModelsMor Geva, Ankit Gupta, Jonathan BerantACL 2020 · 被引用 12 次
- An Empirical Study on Large-Scale Multi-Label Text Classification Including Few and Zero-Shot LabelsIlias Chalkidis, Manos Fergadiotis, Sotiris Kotitsas, Prodromos Malakasiotis 等EMNLP 2020 · 被引用 2 次
相关 Paper
- Numerical Tuple Extraction from Tables with Pre-trainingQingping Yang, Yixuan Cao, Ping LuoKDD 2022 · 被引用 1 次
- DILBERT: Customized Pre-Training for Domain Adaptation with Category Shift, with an Application to Aspect ExtractionEntony Lekhtman, Yftah Ziser, Roi ReichartEMNLP 2021 · 被引用 25 次
- Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERTAkshay Smit, Saahil Jain, Pranav Rajpurkar, Anuj Pareek 等EMNLP 2020 · 被引用 212 次
- When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial DomainRaj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah 等EMNLP 2022 · 被引用 63 次
- Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive MimickingTimo Schick, Hinrich SchützeAAAI 2020 · 被引用 106 次
