There's a Time and Place for Reasoning Beyond the Image
Xingyu Fu, Ben Zhou, Ishaan Preetam Chandratreya, Carl Vondrick, Dan Roth
摘要
Images are often more significant than only the pixels to human eyes, as we can infer, associate, and reason with contextual information from other sources to establish a more complete picture. For example, in Figure 1 , we can find a way to identify the news articles related to the picture through segment-wise understandings of the signs, the buildings, the crowds, and more. This reasoning could provide the time and place the image was taken, which will help us in subsequent tasks, such as automatic storyline construction, correction of image source in intended effect photographs, and upper-stream processing such as image clustering for certain location or time. In this work, we formulate this problem and introduce TARA: a dataset with 16k images with their associated news, time, and location, automatically extracted from New York Times 1 (NYT), and an additional 61k examples as distant supervision from WIT (Srinivasan et al., 2021) . On top of the extractions, we present a crowdsourced subset in which we believe it is possible to find the images' spatiotemporal information for evaluation purpose. We show that there exists a 70% gap between a state-of-the-art joint model and human performance, which is slightly filled by our proposed model that uses segment-wise reasoning, motivating higher-level vision-language joint models that can conduct open-ended reasoning with world knowledge. The data and code are publicly available at https://github. com/zeyofu/TARA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- EDIS: Entity-Driven Image Search over Multimodal Web ContentSiqi Liu, Weixi Feng, Tsu-Jui Fu, Wenhu Chen 等EMNLP 2023 · 被引用 6 次
- Understanding Task Transfer in Vision-Language ModelsBhuvan Sachdeva, Karan Uppal, Abhinav Java, Vineeth N. BalasubramanianCVPR 2026 · 被引用 2 次
- GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent ReasoningShikhhar Siingh, Abhinav Rawat, Chitta Baral, Vivek GuptaACL 2025 · 被引用 1 次
- VIEWS: Entity-Aware News Video CaptioningHammad A. Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin 等EMNLP 2024 · 被引用 1 次
- NL-Eye: Abductive NLI For ImagesMor Ventura, Michael Toker, Nitay Calderon, Zorik Gekhman 等ICLR 2025
它引用的顶会 Paper6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Sample and Computation Redistribution for Efficient Face DetectionJia Guo, Jiankang Deng, Alexandros Lattas, Stefanos ZafeiriouICLR 2022 · 被引用 173 次
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
- Temporal Common Sense Acquisition with Minimal SupervisionBen Zhou, Qiang Ning, Daniel Khashabi, Dan RothACL 2020 · 被引用 76 次
- Who's Waldo? Linking People Across Text and ImagesClaire Yuqing Cui, Apoorv Khandelwal, Yoav Artzi, Noah Snavely 等ICCV 2021 · 被引用 21 次
相关 Paper
- Urban Socio-Semantic Segmentation with Vision-Language ReasoningYu Wang, Yi Wang, Rui Dai, Yujie Wang 等ICLR 2026 · 被引用 4 次
- Visual Abductive ReasoningChen Liang, Wenguan Wang, Tianfei Zhou, Yi YangCVPR 2022 · 被引用 50 次
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 被引用 67 次
- Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual CluesQingxiu Dong, Ziwei Qin, Heming Xia, Tian Feng 等ACL 2022
- TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World SettingsAzmine Toushik Wasi, Shahriyar Zaman Ridoy, Koushik Ahamed Tonmoy, Kinga Tshering 等ICML 2026
