AceTone: Bridging Words and Colors for Conditional Image Grading
Tianren Ma, Mingxiang Liao, Xijin Zhang, Qixiang Ye
Abstract
Color affects how we interpret image style and emotion. Previous color grading methods rely on patch-wise recoloring or fixed filter banks, struggling to generalize across creative intents or align with human aesthetic preferences. In this study, we propose AceTone, the first approach that supports multimodal conditioned color grading within a unified framework. AceTone formulates grading as a generative color transformation task, where a model directly produces 3D-LUTs conditioned on text prompts or reference images. We develop a VQ-VAE based tokenizer which compresses a LUT vector to 64 discrete tokens with fidelity. We further build a large-scale dataset, AceTone-800K, and train a vision-language model to predict LUT tokens, followed by reinforcement learning to align outputs with perceptual fidelity and aesthetics. Experiments show that AceTone achieves state-of-the-art performance on both text-guided and reference-guided grading tasks, improving LPIPS by up to 50% over existing methods. Human evaluations confirm that AceTone's results are visually pleasing and stylistically coherent, demonstrating a new pathway toward language-driven, aesthetic-aligned color grading.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63f036b8-f2c8-4de2-9c7c-44466057c8bbBuilds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
- Photorealistic Style Transfer via Wavelet TransformsJaejun Yoo, Youngjung Uh, Sanghyuk Chun, Byeongkyu Kang et al.ICCV 2019 · 412 citations
Related papers
- HDR-VLM: HDR-Domain Adaptation of VLMs and Preference-Aligned Quality Assessment for HDR Video Color GradingHao Yuan, Jiabin Zhang, Yajing Wu, Ruixuan Pang et al.CVPR 2026
- Video Color Grading via Look-Up Table GenerationSeunghyun Shin, Dongmin Shin, Jisu Shin, Hae-Gon Jeon et al.ICCV 2025 · 2 citations
- Textual Aesthetics in Large Language ModelsLingjie Jiang, Shaohan Huang, Xun Wu, Furu WeiEMNLP 2025
- Probabilistic Prompt Adaptation for Unified Image Aesthetics and Quality AssessmentTakayuki Hara, Yuya OtsukaCVPR 2026
- Vinedresser3D: Towards Agentic Text-guided 3D EditingYankuan Chi, Xiang Li, Zixuan Huang, James M.CVPR 2026
