A Unified Image-Dense Annotation Generation Model for Underwater Scenes
Hongkai Lin, Dingkang Liang, Zhenghao Qi, Xiang Bai
摘要
Underwater dense prediction, especially depth estimation and semantic segmentation, is crucial for gaining a comprehensive understanding of underwater scenes. Nevertheless, high-quality and large-scale underwater datasets with dense annotations remain scarce because of the complex environment and the exorbitant data collection costs. This paper proposes a unified Text-to-Image and DEnse annotation generation method (TIDE) for underwater scenes. It relies solely on text as input to simultaneously generate realistic underwater images and multiple highly consistent dense annotations. Specifically, we unify the generation of text-to-image and text-to-dense annotations within a single model. The Implicit Layout Sharing mechanism (ILS) and cross-modal interaction method called Time Adaptive Normalization (TAN) are introduced to jointly optimize the consistency between image and dense annotations. We synthesize a large-scale underwater dataset using TIDE to validate the effectiveness of our method in underwater dense prediction tasks. The results demonstrate that our method effectively improves the performance of existing underwater dense prediction models and mitigates the scarcity of underwater data with dense annotations. We hope our method can offer new perspectives on alleviating data scarcity issues in other fields. The code is available at https://github.com/HongkLin/TIDE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- NAUTILUS: A Large Multimodal Model for Underwater Scene UnderstandingWei Xu, Cheng Wang, Dingkang Liang, Zongchuang Zhao 等NeurIPS 2025 · 被引用 16 次
- More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion ModelsHongkai Lin, Dingkang Liang, Mingyang Du, Xin Zhou 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Atlantis: Enabling Underwater Depth Estimation with Stable DiffusionFan Zhang, Shaodi You, Yu Li, Ying FuCVPR 2024 · 被引用 22 次
- SEA-PACE: Semi-Supervised Underwater Image Enhancement via Gaussian Process-Assisted Self-Paced LearningJingyang Wang, Hengyue Bi, Jingchao Cao, Feng Gao 等AAAI 2026
- Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale DatasetShijie Lian, Ziyi Zhang, Hua Li, Wenjie Li 等ICML 2024 · 被引用 52 次
- Empowering Semantic-Sensitive Underwater Image Enhancement with VLMGuodong Fan, Shengning Zhou, Genji Yuan, Huiyu Li 等AAAI 2026
- BiPA: Bilevel Prompt Adaptation for Underwater Instance SegmentationLong Ma, Haoze Zheng, Yuhang Mao, Jinyuan Liu 等CVPR 2026
