Style Nursing with Spatial and Semantic Guidance for Zero-Shot Traffic Scene Style Transfer
Zhen Wang, Zihang Lin, Meng Yuan, Yuehu Liu, Chi Zhang
摘要
Recent advances in text-to-image diffusion models have shown an outstanding ability in zero-shot style transfer. However, existing methods often struggle to balance preserving the semantic content of the input image and faithfully transferring the target style in line with the edit prompt. Especially when applied to complex traffic scenes with diverse objects, layouts, and stylistic variations, current diffusion models tend to exhibit Style Neglection, i.e., failing to generate the required style in the prompt. To address this issue, we propose Style Nursing, which directs the model to focus on style subject tokens in the text prompt and excites their corresponding visual activations. Moreover, we introduce Spatial and Semantic Guidance to guide the preservation of content after editing, which utilizes spatial features from the DDIM sampling process together with attention maps from the semantic reconstruction. To evaluate the performance of zero-shot style transfer methods in traffic scenes, we present STREET-6K, a new benchmark dataset comprising 6000 images showcasing diverse traffic scenes and style transfer variations, accompanied by comprehensive annotations and evaluation metrics. Our approach beats state-of-the-art image translation methods in comprehensive quantitative metrics and human evaluations on traffic scene image synthesis while seamlessly generalizing to various other types of images without training or fine-tuning. Further experiments on detection and segmentation tasks show that fine-tuning perception models on our synthesized images improves Recall and mean Intersection over Union (mIoU) by over 10% and 3% respectively in rarely-seen traffic scenes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu 等AAAI 2024 · 被引用 1,641 次
- Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion ModelsHila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf 等SIGGRAPH 2023 · 被引用 438 次
相关 Paper
- Zero-Shot Contrastive Loss for Text-Guided Diffusion Image Style TransferSerin Yang, Hyunmin Hwang, Jong Chul YeICCV 2023 · 被引用 94 次
- Traffic Scene Parsing Through the TSP6K DatasetPeng-Tao Jiang, Yuqi Yang, Yang Cao, Qibin Hou 等CVPR 2024
- SVDGNet: Shapley Value-Based Weight Adjustment for Unsupervised Image Style TransferYi Han, Yaochen Li, Peijun Chen, Wenlong Zhou 等ACM MM 2025
- Semantix: An Energy-guided Sampler for Semantic Style TransferHuiang He, Minghui Hu, Chuanxia Zheng, Chaoyue Wang 等ICLR 2025
- Diffusion-based Image Translation using disentangled style and content representationGihyun Kwon, Jong Chul YeICLR 2023 · 被引用 46 次
