MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision Transformer
Fudong Lin, Summer Crawford, Kaleb Guillot, Yihe Zhang, Yan Chen, Xu Yuan, Li Chen, Shelby Williams, Robert Minvielle, Xiangming Xiao, Drew Gholson, Nicolas Ashwell
Abstract
Precise crop yield prediction provides valuable information for agricultural planning and decision-making processes. However, timely predicting crop yields remains challenging as crop growth is sensitive to growing season weather variation and climate change. In this work, we develop a deep learning-based solution, namely Multi-Modal Spatial-Temporal Vision Transformer (MMST-ViT), for predicting crop yields at the county level across the United States, by considering the effects of short-term meteorological variations during the growing season and the longterm climate change on crops. Specifically, our MMST-ViT consists of a Multi-Modal Transformer, a Spatial Transformer, and a Temporal Transformer. The Multi-Modal Transformer leverages both visual remote sensing data and short-term meteorological data for modeling the effect of growing season weather variations on crop growth. The Spatial Transformer learns the high-resolution spatial dependency among counties for accurate agricultural tracking. The Temporal Transformer captures the long-range temporal dependency for learning the impact of long-term climate change on crops. Meanwhile, we also devise a novel multi-modal contrastive learning technique to pre-train our model without extensive human supervision. Hence, our MMST-ViT captures the impacts of both short-term weather variations and long-term climate change on crops by leveraging both satellite images and meteorological data. We have conducted extensive experiments on over 200 counties in the United States, with the experimental results exhibiting that our MMST-ViT outperforms its counterparts under three performance metrics of interest. Our dataset and code are available at https://github.com/fudong03/ MMST-ViT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Medformer: A Multi-Granularity Patching Transformer for Medical Time-Series ClassificationYihe Wang, Nan Huang, Taida Li, Yujun Yan et al.NeurIPS 2024 · 158 citations
- Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal ApproachYuting Huang, Ziquan Fang, Zhihao Zeng, Lu Chen et al.NeurIPS 2025 · 6 citations
- Long-Tailed Recognition via Information-Preservable Two-Stage LearningFudong Lin, Xu YuanNeurIPS 2025 · 3 citations
- PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield PredictionYu Luo, Xiaogang Zhu, Shan Zeng, Wei Xiang et al.CVPR 2026 · 1 citation
- BoneMet: An Open Large-Scale Multi-Modal Murine Dataset for Breast Cancer Bone Metastasis Diagnosis and PrognosisTiankuo Chu, Fudong Lin, Shubo Wang, Jason Jiang et al.ICLR 2025
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- ViTs for SITS: Vision Transformers for Satellite Image Time SeriesMichail Tarasiou, Erik Chavez, Stefanos ZafeiriouCVPR 2023
- Efficient Representation Learning of Satellite Image Time Series and Their Fusion for Spatiotemporal ApplicationsPoonam Goyal, Arshveer Kaur, Arvind Ram, Navneet GoyalAAAI 2024 · 3 citations
- ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language ModelTong Zhao, Junping Du, Zhe Xue, Meiyu Liang et al.AAAI 2026
- Long-Short Temporal Contrastive Learning of Video TransformersJue Wang, Gedas Bertasius, Du Tran, Lorenzo TorresaniCVPR 2022 · 44 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
