Linking Industry Sectors and Financial Statements: A Hybrid Approach for Company Classification
Guy Stephane Waffo Dzuyo, Gaël Guibon, Christophe Cerisara, Luis Belmar-Letelier
Abstract
The identification of the financial characteristics of industry sectors has a large importance in accounting audit, allowing auditors to prioritize the most important area during audit. Existing company classification standards such as the Standard Industry Classification (SIC) code allow to map a company to a category based on its activity and products. In this paper, we explore the potential of machine learning algorithms and language models to analyze the relationship between those categories and companies' financial statements. We propose a supervised company classification methodology and analyze several types of representations for financial statements. Existing works address this task using solely numerical information in financial records. Our findings show that beyond numbers, textual information occurring in financial records can be leveraged by language models to match the performance of dedicated decision tree-based classifiers, while providing better explainability and more generic accounting representations. We think this work can serve as a preliminary work towards semi-automatic auditing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9759871f-c66d-4999-9d37-329f9c358350Cited by top-tier papers1
Ask how each one uses itBuilds on6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- FiNER: Financial Numeric Entity Recognition for XBRL TaggingLefteris Loukas, Manos Fergadiotis, Ilias Chalkidis, Eirini Spyropoulou et al.ACL 2022 · 90 citations
- NumHTML: Numeric-Oriented Hierarchical Transformer Model for Multi-Task Financial ForecastingLinyi Yang, Jiazheng Li, Ruihai Dong, Yue Zhang et al.AAAI 2022 · 54 citations
- Instruction Pre-Training: Language Models are Supervised Multitask LearnersDaixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi et al.EMNLP 2024 · 13 citations
Related papers
- Multi-perspective Analysis of Large Language Model Domain Specialization: An Experiment in Accounting Audit Procedures GenerationYusuke NoroEMNLP 2025
- EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial StatementsIssa Sugiura, Takashi Ishida, Taro Makino, Chieko Tazuke et al.ICLR 2026 · 9 citations
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong et al.EMNLP 2020 · 203 citations
- AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate StatementsAdriana Eufrosina Bora, Pierre-Luc St-Charles, Mirko Bronzi, Arsène Fansi Tchango et al.ICLR 2025
- Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary LearningJohn Wu, David Wu, Jimeng SunEMNLP 2024 · 4 citations
