VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo Retouching
Yihong Guo, Youwei Lyu, Jiajun Tang, Yizhuo Zhou, Hongliang Wang, Jinwei Chen, Changqing Zou, Qingnan Fan
摘要
Reasoning photo retouching has gained significant traction, requiring models to analyze image defects, give reasoning processes, and execute precise retouching enhancements. However, existing approaches often rely on non-differentiable external software, creating optimization barriers and suffering from high parameter redundancy and limited generalization. To address these challenges, we propose VeraRetouch, a lightweight and fully differentiable framework for multi-task photo retouching. We employ a 0.5B Vision-Language Model (VLM) as the central intelligence to formulate retouching plans based on instructions and scene semantics. Furthermore, we develop a fully differentiable Retouch Renderer that replaces external tools, enabling direct end-to-end pixel-level training through decoupled control latents for lighting, global color, and specific color adjustments. To overcome data scarcity, we introduce AetherRetouch-1M+, the first million-scale dataset for professional retouching, constructed via a new inverse degradation workflow. Furthermore, we propose DAPO-AE, a reinforcement learning post-training strategy that enhances autonomous aesthetic cognition. Extensive experiments demonstrate that VeraRetouch achieves state-of-the-art performance across multiple benchmarks while maintaining a significantly smaller footprint, enabling mobile deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen 等ICML 2024 · 被引用 499 次
- Unpaired Image Enhancement Featuring Reinforcement-Learning-Controlled Image Editing SoftwareSatoshi Kosugi, Toshihiko YamasakiAAAI 2020 · 被引用 104 次
- Representative Color Transform for Image EnhancementHanul Kim, Su-Min Choi, Chang-Su Kim, Yeong Jun KohICCV 2021 · 被引用 89 次
- AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-time Image EnhancementCanqian Yang, Meiguang Jin, Xu Jia, Yi Xu 等CVPR 2022 · 被引用 57 次
相关 Paper
- PerTouch: VLM-Driven Agent for Personalized and Semantic Image RetouchingZewei Chang, Zheng-Peng Duan, Jianxing Zhang, Chun-Le Guo 等AAAI 2026 · 被引用 3 次
- MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching SkillsNiladri Shekhar Dutt, Duygu Ceylan, Niloy J. MitraSIGGRAPH 2025 · 被引用 2 次
- RetouchAgent: Towards Interactive and Explainable Image Retouching with MLLM AgentsShuo Zhang, Xinyu YangAAAI 2026
- DiffRetouch: Using Diffusion to Retouch on the Shoulder of ExpertsZheng-Peng Duan, Jiawei Zhang, Zheng Lin, Xin Jin 等AAAI 2025
- Agentic Retoucher for Text-To-Image GenerationShaocheng Shen, Jianfeng Liang, Chunlei Cai, Cong Geng 等CVPR 2026 · 被引用 9 次
