UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in RL
Rui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu, Zhe Gan, Yinfei Yang, Zuxuan Wu, Afshin Dehghan
2026年份
摘要
Image Editing transfer the image into a loose, flowing watercolor-wash style extract the white and black puppy including its red collar and the surrounding green grass remove the white cat curled up and sleeping on the chair replace the pomegranate seeds in the image with a handful of ripe, dark purple grape change the clear blue sky background to a dramatic sunset with vibrant orange and purple hue add a bouquet of white roses in the opening of the black vase, gently spilling over the top Text-to-Image Figure 1. Examples of images generated by UniGen-1.5.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper40
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
相关 Paper
- Exploring Sparse MoE in GANs for Text-conditioned Image SynthesisJiapeng Zhu, Ceyuan Yang, Kecheng Zheng, Yinghao Xu 等CVPR 2025
- UniTune: Text-Driven Image Editing by Fine Tuning a Diffusion Model on a Single ImageDani Valevski, Matan Kalman, Eyal Molad, Eyal Segalis 等SIGGRAPH 2023 · 被引用 61 次
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
- CCEdit: Creative and Controllable Video Editing via Diffusion ModelsRuoyu Feng, Wenming Weng, Yanhui Wang, Yuhui Yuan 等CVPR 2024
- Multi-Concept Customization of Text-to-Image DiffusionNupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman 等CVPR 2023
