Affective Image Filter: Reflecting Emotions from Text to Images
Shuchen Weng, Peixuan Zhang, Zheng Chang, Xinlong Wang, Si Li, Boxin Shi
摘要
Understanding the emotions in text and presenting them visually is a very challenging problem that requires a deep understanding of natural language and high-quality image synthesis simultaneously. In this work, we propose Affective Image Filter (AIF), a novel model that is able to understand the visually-abstract emotions from the text and reflect them to visually-concrete images with appropriate colors and textures. We build our model based on the multi-modal transformer architecture, which unifies both images and texts into tokens and encodes the emotional prior knowledge. Various loss functions are proposed to understand complex emotions and produce appropriate visualization. In addition, we collect and contribute a new dataset with abundant aesthetic images and emotional texts for training and evaluating the AIF model. We carefully design four quantitative metrics and conduct a user study to comprehensively evaluate the performance, which demonstrates our AIF model outperforms state-of-the-art methods and could evoke specific emotional responses from human observers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang 等ICLR 2024 · 被引用 173 次
- CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content GenerationKaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu 等ICLR 2026 · 被引用 10 次
- LuminAIRe: Illumination-Aware Conditional Image Repainting for Lighting-Realistic GenerationJiajun Tang, Haofeng Zhong, Shuchen Weng, Boxin ShiNeurIPS 2023 · 被引用 6 次
- Make Me Happier: Evoking Emotions through Image Diffusion ModelsQing Lin, Jingfeng Zhang, Yew-Soon Ong, Mengmi ZhangICCV 2025 · 被引用 4 次
- EmoStyle: Emotion-Driven Image StylizationJingyuan Yang, Zihuan Bai, Hui HuangCVPR 2026 · 被引用 2 次
它引用的顶会 Paper28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
相关 Paper
- EmIT: Emotional Interaction control in Text-to-image diffusion modelsHaofan Zhang, Shangfei WangACM MM 2025
- MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion RecognitionJian Chen, Yuxuan Hu, Haifeng Lu, Wei Wang 等ACM MM 2025 · 被引用 5 次
- EmotiCrafter: Text-to-Emotional-Image Generation Based on Valence-Arousal ModelShengqi Dang, Yi He, Long Ling, Ziqing Qian 等ICCV 2025 · 被引用 2 次
- EmoEdit: Evoking Emotions through Image ManipulationJingyuan Yang, Jiawei Feng, Weibin Luo, Dani Lischinski 等CVPR 2025
- Affection: Learning Affective Explanations for Real-World Visual DataPanos Achlioptas, Maks Ovsjanikov, Leonidas J. Guibas, Sergey TulyakovCVPR 2023
