Any2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched Description
Xu Zhang, Jianzhong Huang, Lefei Zhang
Abstract
Recent advances in controllable text-to-image (T2I) generation have achieved impressive results in natural images, but remote sensing (RS) T2I generation remains challenging due to the unique nature of geospatial data. Existing methods struggle to integrate diverse spatial controls and model complex spatial relationships, often failing to maintain semantic consistency with typically vague or incomplete textual descriptions. Moreover, limited by small-scale, low-quality datasets, these models produce outputs with inconsistent layouts and unrealistic content. To address these issues, we propose Any2RSI, a flexible framework for controllable RS T2I generation. It features a Cross-Modal Multi-Control Adapter that extracts modality-agnostic embeddings from heterogeneous spatial inputs, enabling precise spatial guidance. To compensate for sparse or ambiguous text prompts, we introduce a VLM-Empowered Enriched Description Generation module that enhances input descriptions with cross-modal semantics for more coherent T2I generation. Furthermore, we present RST2I-110K, a dataset of over 115,000 high-quality RS image-text pairs across diverse scenes, addressing the current lack of semantically grounded textual annotations needed for RS T2I generation. Any2RSI achieves state-of-the-art performance on both existing and new datasets, improving the realism and structural accuracy of generated RS imagery.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2cf1934-69ad-42b2-8422-5f47c73bea4aBuilds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- RS2-SAM2: Customized SAM2 for Referring Remote Sensing Image SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangAAAI 2026 · 1 citation
- CF-IPT: Cross-Modal Fusion Interactive Prompt Tuning of Vision-Language Pre-Trained Model for Multisource Remote Sensing Data ClassificationJinheng Ji, Jiahui Qu, Wenqian Dong, Yunsong LiCVPR 2026
- Any2Any: Unified Arbitrary Modality Translation for Remote SensingHaoyang Chen, Jing Zhang, Di Wang, Hebaixu Wang et al.ICML 2026 · 5 citations
- SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingZhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu et al.AAAI 2024 · 167 citations
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image GenerationCihang Peng, Qiming Hou, Zhong Ren, Kun ZhouICCV 2025
