MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems
Arda Yüksel, Gabriel Thiem, Susanne Walter, Patrick Felka, Gabriela Alves Werb, Ivan Habernal
摘要
Industry classification schemes are integral parts of public and corporate databases as they classify businesses based on economic activity. Due to the size of the company registers, manual annotation is costly, and fine-tuning models with every update in industry classification schemes requires significant data collection. We replicate the manual expert verification by using existing or easily retrievable multimodal resources for industry classification. We present MONETA, the first multimodal industry classification benchmark with text (Website, Wikipedia, Wikidata) and geospatial sources (OpenStreetMap and satellite imagery). Our dataset enlists 1,000 businesses in Europe with 20 economic activity labels according to EU guidelines (NACE). Our training-free baseline reaches 62.10% and 74.10% with open and closed-source Multimodal Large Language Models (MLLM). We observe an increase of up to 22.80% with the combination of multi-turn design, context enrichment, and classification explanations. We will release our dataset and the enhanced guidelines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Granular Privacy Control for Geolocation with Vision Language ModelsEthan Mendes, Yang Chen, James Hays, Sauvik Das 等EMNLP 2024 · 被引用 3 次
- Identifying Spatio-Temporal Drivers of Extreme EventsMohamad Hakam Shams Eddin, Jürgen GallNeurIPS 2024 · 被引用 2 次
- Linking Industry Sectors and Financial Statements: A Hybrid Approach for Company ClassificationGuy Stephane Waffo Dzuyo, Gaël Guibon, Christophe Cerisara, Luis Belmar-LetelierAAAI 2025 · 被引用 1 次
相关 Paper
- Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly DetectionDahu Shi, Chengshen He, Shaochen Zhang, Bo Qian 等CVPR 2026
- Web-Scale Visual Entity Recognition: An LLM-Driven Data ApproachMathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet IscenNeurIPS 2024 · 被引用 5 次
- On Large Multimodal Models as Open-World Image ClassifiersAlessandro Conti, Massimiliano Mancini, Enrico Fini, Yiming Wang 等ICCV 2025 · 被引用 3 次
- MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly DetectionXi Jiang, Jian Li, Hanqiu Deng, Yong Liu 等ICLR 2025 · 被引用 3 次
- IPdb: A High-Precision IP Level Industry Categorization of Web ServicesHongxu Chen, Guanglei Song, Zhiliang Wang, Jiahai Yang 等WWW 2025 · 被引用 5 次
