PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models
Mingde Yao, Zhiyuan You, King-Man Tam, Menglu Wang, Tianfan Xue
Abstract
With the recent fast development of generative models, instruction-based image editing has shown great potential in generating high-quality images. However, the quality of editing highly depends on carefully designed instructions, placing the burden of task decomposition and sequencing entirely on the user. To achieve autonomous image editing, we present PhotoAgent, a system that advances image editing through explicit aesthetic planning. Specifically, PhotoAgent formulates autonomous image editing as a long-horizon decision-making problem. It reasons over user aesthetic intent, plans multi-step editing actions via tree search, and iteratively refines results through closedloop execution with memory and visual feedback, without requiring step-by-step user prompts. To support reliable evaluation in real-world scenarios, we introduce UGC-Edit, an aesthetic evaluation benchmark consisting of 7,000 photos and a learned aesthetic reward model. We also construct a test set containing 1,017 photos to systematically assess autonomous photo editing performance. Extensive experiments demonstrate that PhotoAgent consistently improves both instruction adherence and visual quality compared with baseline methods. The project page is https://mdyao.github.io/PhotoAgent .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd3f58a1-4e34-405e-8fbd-66454836029dBuilds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
Related papers
- CAISE: Conversational Agent for Image Search and EditingHyounghun Kim, Doo Soon Kim, Seunghyun Yoon, Franck Dernoncourt et al.AAAI 2022 · 6 citations
- PSBench: Editing Image via GUI Agents in PhotoshopYinuo Zhang, Zian Cheng, Ziya Zhao, Zongyu Li et al.ICML 2026
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn EditingTianyu Chen, Yasi Zhang, Zhi Zhang, Peiyu Yu et al.ICLR 2026 · 11 citations
- PosterAgent: Agentic Poster Generation via Stage-Aware Reinforcement LearningZhuocheng Yu, Feng Zhang, Sujian Li, Kai JiaICML 2026
- PerTouch: VLM-Driven Agent for Personalized and Semantic Image RetouchingZewei Chang, Zheng-Peng Duan, Jianxing Zhang, Chun-Le Guo et al.AAAI 2026 · 3 citations
