PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching
Zewei Chang, Zheng-Peng Duan, Jianxing Zhang, Chun-Le Guo, Siyu Liu, Hyungju Chun, Hyunhee Park, Zikun Liu, Chongyi Li
摘要
Image retouching aims to enhance visual quality while aligning with users' personalized aesthetic preferences. To address the challenge of balancing controllability and subjectivity, we propose a unified diffusion-based image retouching framework called PerTouch. Our method supports semantic-level image retouching while maintaining global aesthetics. Using parameter maps containing attribute values in specific semantic regions as input, PerTouch constructs an explicit parameter-to-image mapping for fine-grained image retouching. To improve semantic boundary perception, we introduce semantic replacement and parameter perturbation mechanisms during training. To connect natural language instructions with visual control, we develop a VLM-driven agent to handle both strong and weak user instructions. Equipped with mechanisms of feedback-driven rethinking and scene-aware memory, PerTouch better aligns with user intent and captures long-term preferences. Extensive experiments demonstrate each component’s effectiveness and the superior performance of PerTouch in personalized image retouching.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Representative Color Transform for Image EnhancementHanul Kim, Su-Min Choi, Chang-Su Kim, Yeong Jun KohICCV 2021 · 被引用 89 次
- StarEnhancer: Learning Real-Time and Style-Aware Image EnhancementYuda Song, Hui Qian, Xin DuICCV 2021 · 被引用 59 次
- AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-time Image EnhancementCanqian Yang, Meiguang Jin, Xu Jia, Yi Xu 等CVPR 2022 · 被引用 57 次
相关 Paper
- DiffRetouch: Using Diffusion to Retouch on the Shoulder of ExpertsZheng-Peng Duan, Jiawei Zhang, Zheng Lin, Xin Jin 等AAAI 2025
- RetouchAgent: Towards Interactive and Explainable Image Retouching with MLLM AgentsShuo Zhang, Xinyu YangAAAI 2026
- VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo RetouchingYihong Guo, Youwei Lyu, Jiajun Tang, Yizhuo Zhou 等SIGGRAPH 2026
- AttriCtrl: A Generalizable Framework for Controlling Semantic Attribute Intensity in Diffusion ModelsDie Chen, Zhongjie Duan, Zhiwen Li, Cen Chen 等ICLR 2026
- Agentic Retoucher for Text-To-Image GenerationShaocheng Shen, Jianfeng Liang, Chunlei Cai, Cong Geng 等CVPR 2026 · 被引用 9 次
