CamEdit: Continuous Camera Parameter Control for Photorealistic Image Editing
Xinran Qin, Zhixin Wang, Fan Li, Haoyu Chen, Renjing Pei, Wenbo Li, Xiaochun Cao
Abstract
Recent advances in diffusion models have substantially improved text-driven image editing. However, existing frameworks based on discrete textual tokens struggle to support continuous control over camera parameters and smooth transitions in visual effects. These limitations hinder their applications to realistic, camera-aware, and fine-grained editing tasks. In this paper, we present CamEdit, a diffusionbased framework for photorealistic image editing that enables continuous and semantically meaningful manipulation of common camera parameters such as aperture and shutter speed. CamEdit incorporates a continuous parameter prompting mechanism and a parameter-aware modulation module that guides the model in smoothly adjusting focal plane, aperture, and shutter speed, reflecting the effects of varying camera settings within the diffusion process. To support supervised learning in this setting, we introduce CamEdit50K, a dataset specifically designed for photorealistic image editing with continuous camera parameter settings. It contains over 50k image pairs combining real and synthetic data with dense camera * Equal Contribution † Corresponding Author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
parameter variations across diverse scenes. Extensive experiments demonstrate that CamEdit enables flexible, consistent, and high-fidelity image editing, achieving state-of-the-art performance in camera-aware visual manipulation and fine-grained photographic control.
Recently, diffusion models [1, 2,20,49,50,52,54,51,32] have become powerful tools for both image generation and editing. They usually apply a pre-trained text encoder such as CLIP [48] and T5 [66] to inject manual textual prompts information into the generation process, enabling better generation quality and more precise control. Meanwhile, as social media platforms grow and smartphone cameras continue to advance, editing images to reflect photorealistic optical effects has become practically valuable. This highlights the need for editing methods that can directly manipulate camera parameters. However, most existing image editing methods [4,5,19,23,24,29,35,58,62] focus mainly on three main tasks: semantic editing, stylistic editing and structural editing.
Few prior works target photorealistic image editing, which edits indistinguishable from real photographs through precise control of camera parameters. In this work, we focus on the precise adjustments of focal plane 1 , aperture and shutter speed in camera parameters, which play a fundamental role respectively in determining focal range, background defocus degree, and exposure time [46,57] during the photo-taking process.
Diffusion models capture strong spatial priors and scene geometry [55], making them suited for photorealistic editing. Recent works encode camera settings as discrete tokens within text-toimage (T2I) or text-to-video (T2V) generation frameworks [11,70]. However, such discrete textual token-based approaches are difficult to directly apply to editing tasks involving continuous camera parameter control through textual prompt input (e.g. "Adjust the image with aperture f /2.8", etc.). This mismatch hampers smooth parameter adjustment and limits applicability to photographic editing.
To overcome these challenges, we introduce CamEdit, a diffusion-based framework for photorealistic image editing that allows continuous control of camera settings using text prompts. Instead of turning parameter values into separate tokens, we propose a continuous parameter prompting method, which interpolates between predefined anchor embeddings in the text space. This preserves alignment with representation distribution of the pre-trained model while enabling fine-grained control over a wide range of settings (such as "aperture f /[2, 10]", "shutter speed [0, 1]"). As diffusion backbones lack explicit camera priors and fail to capture parameter-specific effects, we further propose a parameteraware modulation module that conditions spatial and channel features throughout the diffusion transformer, making explicit both local and global effects that text embeddings alone miss.
Given the lack of high-quality datasets for photorealistic camera-aware editing, we construct a hybrid dataset named CamEdit50K, which includes real-world photographs with extracted or estimated EXIF metadata 2 , along with synthetic image pairs rendered under controlled variations in focal plane, aperture, and shutter speed. This dataset provides a strong foundation for learning models that are physically consistent and aware of the effect of varying camera parameters.
In summary, our main contributions can be summarized as follows:
• We propose CamEdit, a diffusion-based framework for photorealistic image editing that enables continuous and fine-grained control over intrinsic camera parameters such as aperture, focal plane, and shutter speed, entirely through manual textual prompts.
• We design a continu
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token GenerationRuoyu Chen, Xiaoqing Guo, Kangwei Liu, Siyuan Liang et al.CVPR 2026 · 19 citations
- YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object RemovalChenyang Wu, Lina Lei, Fan Li, Chunle Guo et al.CVPR 2026 · 4 citations
- HP-Edit: A Human-Preference Post-Training Framework for Image EditingFan Li, Chonghuinan Wang, Lina Lei, Yuping Qiu et al.CVPR 2026 · 4 citations
- CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex InstructionsChonghuinan Wang, Zihan Chen, Yuxiang Wei, Tianyi Jiang et al.CVPR 2026 · 3 citations
- ColorFLUX: A Structure-Color Decoupling Framework for Old Photo ColorizationBingchen Li, Zhixin Wang, Fan Li, Jiaqi Xu et al.CVPR 2026 · 1 citation
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- An Item Is Worth a Prompt: Versatile Image Editing with Disentangled ControlAosong Feng, Weikang Qiu, Jinbin Bai, Zhen Dong et al.AAAI 2025 · 9 citations
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 439 citations
- Modular-Cam: Modular Dynamic Camera-view Video Generation with LLMZirui Pan, Xin Wang, Yipeng Zhang, Hong Chen et al.AAAI 2025 · 6 citations
- AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image EditingZhiyuan Ma, Guoli Jia, Bowen ZhouAAAI 2024 · 13 citations
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera ControlSherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace et al.ICLR 2025
