TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data
Jeremy Andrew Irvin, Emily Ruoyu Liu, Joyce Chuyi Chen, Ines Dormoy, Jinyoung Kim, Samar Khanna, Zhuo Zheng, Stefano Ermon
摘要
Large vision and language assistants have enabled new capabilities for interpreting natural images. These approaches have recently been adapted to earth observation data, but they are only able to handle single image inputs, limiting their use for many real-world tasks. In this work, we develop a new vision and language assistant called TEOChat that can engage in conversations about temporal sequences of earth observation data. To train TEOChat, we curate an instruction-following dataset composed of many single image and temporal tasks including building change and damage assessment, semantic change detection, and temporal scene classification. We show that TEOChat can perform a wide variety of spatial and temporal reasoning tasks, substantially outperforming previous vision and language assistants, and even achieving comparable or better performance than several specialist models trained to perform specific tasks. Furthermore, TEOChat achieves impressive zero-shot performance on a change detection and change question answering dataset, outperforms GPT-4o and Gemini 1.5 Pro on multiple temporal tasks, and exhibits stronger single image capabilities than a comparable single image instruction-following model on scene classification, visual question answering, and captioning. We publicly release our data, model, and code at https://github.com/ermongroup/TEOChat .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring ModelsDilxat Muhtar, Enzhuo Zhang, Zhenshi Li, Feng Gu 等NeurIPS 2025 · 被引用 16 次
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing ImagesZepeng Xin, Kaiyu Li, Luodi Chen, Wanchen Li 等CVPR 2026 · 被引用 14 次
- InstructSAM: A Training-free Framework for Instruction-Oriented Remote Sensing Object RecognitionYijie Zheng, Weijie Wu, Qingyun Li, Xuehui Wang 等NeurIPS 2025 · 被引用 12 次
- TerraScope: Pixel-Grounded Visual Reasoning for Earth ObservationYan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu 等CVPR 2026 · 被引用 9 次
- GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote SensingAoran Xiao, Shihao Cheng, Yonghao Xu, Yexian Ren 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu 等NeurIPS 2022 · 被引用 707 次
相关 Paper
- EarthDial: Turning Multi-sensory Earth Observations to Interactive DialoguesSagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz 等CVPR 2025
- GeoChat: Grounded Large Vision-Language Model for Remote SensingKartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das 等CVPR 2024
- SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image UnderstandingYing Chen, Guoan Wang, Yuanfeng Ji, Yanjun Li 等CVPR 2025
- Streaming Video Instruction TuningJiaer Xia, Peixian Chen, Mengdan Zhang, Xing Sun 等CVPR 2026 · 被引用 28 次
- ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and ReasoningZhe Xie, Zeyan Li, Xiao He, Longlong Xu 等VLDB 2025 · 被引用 87 次
