PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching
Zewei Chang, Zheng-Peng Duan, Jianxing Zhang, Chun-Le Guo, Siyu Liu, Hyungju Chun, Hyunhee Park, Zikun Liu, Chongyi Li
Abstract
Image retouching aims to enhance visual quality while aligning with users' personalized aesthetic preferences. To address the challenge of balancing controllability and subjectivity, we propose a unified diffusion-based image retouching framework called PerTouch. Our method supports semantic-level image retouching while maintaining global aesthetics. Using parameter maps containing attribute values in specific semantic regions as input, PerTouch constructs an explicit parameter-to-image mapping for fine-grained image retouching. To improve semantic boundary perception, we introduce semantic replacement and parameter perturbation mechanisms during training. To connect natural language instructions with visual control, we develop a VLM-driven agent to handle both strong and weak user instructions. Equipped with mechanisms of feedback-driven rethinking and scene-aware memory, PerTouch better aligns with user intent and captures long-term preferences. Extensive experiments demonstrate each component’s effectiveness and the superior performance of PerTouch in personalized image retouching.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9e3c3b2-2a94-4e8d-9223-2ffd92927e54Cited by top-tier papers1
Ask how each one uses itBuilds on14
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Representative Color Transform for Image EnhancementHanul Kim, Su-Min Choi, Chang-Su Kim, Yeong Jun KohICCV 2021 · 89 citations
- StarEnhancer: Learning Real-Time and Style-Aware Image EnhancementYuda Song, Hui Qian, Xin DuICCV 2021 · 59 citations
- AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-time Image EnhancementCanqian Yang, Meiguang Jin, Xu Jia, Yi Xu et al.CVPR 2022 · 57 citations
Related papers
- DiffRetouch: Using Diffusion to Retouch on the Shoulder of ExpertsZheng-Peng Duan, Jiawei Zhang, Zheng Lin, Xin Jin et al.AAAI 2025
- RetouchAgent: Towards Interactive and Explainable Image Retouching with MLLM AgentsShuo Zhang, Xinyu YangAAAI 2026
- VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo RetouchingYihong Guo, Youwei Lyu, Jiajun Tang, Yizhuo Zhou et al.SIGGRAPH 2026
- AttriCtrl: A Generalizable Framework for Controlling Semantic Attribute Intensity in Diffusion ModelsDie Chen, Zhongjie Duan, Zhiwen Li, Cen Chen et al.ICLR 2026
- Agentic Retoucher for Text-To-Image GenerationShaocheng Shen, Jianfeng Liang, Chunlei Cai, Cong Geng et al.CVPR 2026 · 9 citations
