Towards Temporal-Aware Multi-Modal Retrieval Augemented Generation in Finance
Fengbin Zhu, Junfeng Li, Liangming Pan, Wenjie Wang, Fuli Feng, Chao Wang, Huanbo Luan, Tat-Seng Chua
Abstract
Finance decision-making often relies on in-depth data analysis across various data sources, including financial tables, news articles, stock prices, etc. In this work, we introduce FinTMMBench, the first comprehensive benchmark for evaluating temporal-aware multi-modal Retrieval-Augmented Generation (RAG) systems in finance. Built from heterologous data of NASDAQ 100 companies, FinTMMBench offers three significant advantages. 1) Multi-modal Corpus: It encompasses a hybrid of financial tables, news articles, daily stock prices, and visual technical charts as the corpus. 2) Temporal-aware Questions: Each question requires the retrieval and interpretation of its relevant data over a specific time period, including daily, weekly, monthly, quarterly, and annual periods. 3) Diverse Financial Analysis Tasks: The questions involve 10 different financial analysis tasks designed by domain experts, including information extraction, trend analysis, sentiment analysis and event detection, etc. We further propose a novel TMMHybridRAG method, which first leverages a multi-modal LLM to convert data from other modalities (e.g., tabular, visual and time-series data) into textual format and then incorporates temporal information in each node when constructing graphs and dense indexes. Its effectiveness has been validated in extensive experiments, but notable gaps remain, highlighting the challenges presented by our FinTMMBench. The benchmark and source code will be made publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext becfa22d-54b5-4cfa-879c-4864a9a124b5Builds on10
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav et al.ICLR 2021 · 229 citations
- MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual DataYilun Zhao, Yunxiang Li, Chenying Li, Rui ZhangACL 2022 · 168 citations
- Towards Complex Document Understanding By Discrete ReasoningFengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang et al.ACM MM 2022 · 40 citations
- Learning to Imagine: Integrating Counterfactual Thinking in Neural Discrete ReasoningMoxin Li, Fuli Feng, Hanwang Zhang, Xiangnan He et al.ACL 2022 · 39 citations
Related papers
- FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial DomainSuifeng Zhao, Zhuoran Jin, Sujian Li, Jun GaoEMNLP 2025 · 1 citation
- REAL-MM-RAG: A Real-World Multi-Modal Retrieval BenchmarkNavve Wasserman, Roi Pony, Oshri Naparstek, Adi Raz Goldfarb et al.ACL 2025 · 33 citations
- OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial DomainShuting Wang, Jiejun Tan, Zhicheng Dou, Ji-Rong WenEMNLP 2025 · 6 citations
- FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step ComputationZichen Tang, Haihong E, Rongjin Li, Jiacheng Liu et al.AAAI 2026
- FCMR: Robust Evaluation of Financial Cross-Modal Multi-Hop ReasoningSeunghee Kim, Changhyeon Kim, Taeuk KimACL 2025
