MoEdit: On Learning Quantity Perception for Multi-object Image Editing
Yanfeng Li, Ka-Hou Chan, Yue Sun, Chan-Tong Lam, Tong Tong, Zitong Yu, Keren Fu, Xiaohong Liu, Tao Tan
Abstract
Ten koalas "...steampunck style" "...vibrant portrait painting of Salvador Dalí" "...with blanket" "→mice, by the sea" "→foxes, futuristic metropolis style" Reference TurboEdit MoEdit (Ours) Three rabbits and two foxes "...dark horror style" "...with smilling faces" "→bears, in a natural field" "→corgis, by the sea" "→foxes, in a natural field" Figure 1. Visual comparisons of our MoEdit with TurboEdit [52]. Reference represents input images. Five different images edited by each method are based on five distinct text prompts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b784ba35-d67f-412f-807c-bdcb1c998552Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and EditingDongxu Li, Junnan Li, Steven C. H. HoiNeurIPS 2023 · 587 citations
- Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsShihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao et al.NeurIPS 2023 · 505 citations
Related papers
- CCEdit: Creative and Controllable Video Editing via Diffusion ModelsRuoyu Feng, Wenming Weng, Yanhui Wang, Yuhui Yuan et al.CVPR 2024
- HQ-Edit: A High-Quality Dataset for Instruction-based Image EditingMude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi et al.ICLR 2025
- DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity PreservationJiwook Kim, Seonho Lee, Jaeyo Shin, Jiho Choi et al.ICLR 2025
- RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion ModelsOzgur Kara, Bariscan Kurtkaya, Hidir Yesiltepe, James M. Rehg et al.CVPR 2024
- Style-Editor: Text-driven Object-centric Style EditingJihun Park, Jongmin Gim, Kyoungmin Lee, Seunghun Lee et al.CVPR 2025
