GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
Shikhhar Siingh, Abhinav Rawat, Chitta Baral, Vivek Gupta
Abstract
Publicly significant images from events carry valuable contextual information with applications in domains such as journalism and education. However, existing methodologies often struggle to accurately extract this contextual relevance from images. To address this challenge, we introduce GETREASON (Geospatial Event Temporal Reasoning), a framework designed to go beyond surfacelevel image descriptions and infer deeper contextual meaning. We hypothesize that extracting global event, temporal, and geospatial information from an image enables a more accurate understanding of its contextual significance. We also introduce a new metric GREAT (Geospatial, Reasoning and Event Accuracy with Temporal alignment) for a reasoning capturing evaluation. Our layered multi-agentic approach, evaluated using a reasoning-weighted metric, demonstrates that meaningful information can be inferred from images, allowing them to be effectively linked to their corresponding events and broader contextual background.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15af93df-3304-42c5-9eb9-4836cbccae4cCited by top-tier papers1
Ask how each one uses itBuilds on6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- CrAM: Credibility-Aware Attention Modification in LLMs for Combating Misinformation in RAGBoyi Deng, Wenjie Wang, Fengbin Zhu, Qifan Wang et al.AAAI 2025 · 21 citations
- Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMsHuichen Will Wang, Larry Birnbaum, Vidya SetlurCHI 2025 · 11 citations
- "Image, Tell me your story!" Predicting the original meta-context of visual misinformationJonathan Tonglet, Marie-Francine Moens, Iryna GurevychEMNLP 2024 · 6 citations
Related papers
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning ChainsChun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan et al.NeurIPS 2025 · 18 citations
- TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World SettingsAzmine Toushik Wasi, Shahriyar Zaman Ridoy, Koushik Ahamed Tonmoy, Kinga Tshering et al.ICML 2026
- Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense KnowledgeFan Li, Jianxing Yu, Jielong Tang, Wenqing Chen et al.ACL 2025 · 3 citations
- There's a Time and Place for Reasoning Beyond the ImageXingyu Fu, Ben Zhou, Ishaan Preetam Chandratreya, Carl Vondrick et al.ACL 2022
- GT-Loc: Unifying When and Where in Images Through a Joint Embedding SpaceDavid G. Shatwell, Ishan Rajendrakumar Dave, Sirnam Swetha, Mubarak ShahICCV 2025 · 1 citation
