Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
Kang Liao, Size Wu, Zhonghua Wu, Linyi Jin, Chao Wang, Yikai Wang, Fei Wang, Wei Li, Chen Change Loy
Abstract
Ground Truth UVP GeoCalib Puffin (Ours) (a) Camera Controllable Generation (b) Camera Understanding Horizon Lines: <think> (Roll) The image exhibits a large clockwise Dutch angle, as evidenced by the tilted perspective where the tree and roofline are not aligned with the natural horizontal plane, suggesting a significant rotation of the camera. (Pitch) The vertical orientation indicates a large tilt-up, as the sky occupies a substantial portion of the frame, with the tree and roofline positioned lower, emphasizing an upward view. (FoV) The image exhibits a medium shot, capturing a portion of the tree and roof while maintaining a balanced composition that includes both the sky and the elements in the foreground. </think> <answer>0.5107, 0.5550, 0.7558</answer> Prompt: Reason the spatial distribution of this image in a thinking mode, and then estimate its camera parameters (roll, pitch, and FoV, in radius). Prompt: Given a scene description and corresponding camera parameters, merge them into a coherent prompt and generate an accurate visualization that highlights visual cues. Caption: A sleek, modern interior with a large, minimalistic window framed in light beige, offering a view of a textured blue stone wall reflecting glimpses of a nearby building and lush greenery. The camera parameters (roll, pitch, and FoV) are: -0.2761, -0.1348, 0.9415. <think> The slight tilt of the window frame relative to the vertical lines of the wall suggests a small counterclockwise Dutch angle in the horizontal orientation. The visible portion of the window and the wall indicates a medium shot, capturing a portion of the interior and exterior without showing the entire scene. The subtle downward angle of the camera, evident from the alignment of the windowsill and the wall, suggests a small tilt-down in the vertical orientation. </think> Error Map Generated Image Error=7.5°Error=2.7°L atitude Camera Map
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou et al.CVPR 2026 · 33 citations
- Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and CroppingTianxiang Du, Hulingxiao He, Yuxin PengCVPR 2026 · 3 citations
- AesFormer: Transform Everyday Photos into Beautiful MemoriesTianxiang Du, Hulingxiao He, Yuxin PengICML 2026
- CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise WarpingHaoyu Zhao, Jiaxi Gu, Haoran Chen, Qingping Zheng et al.ICML 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
Related papers
- SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual GroundingRong Li, Shijie Li, Lingdong Kong, Xulei Yang et al.CVPR 2025
- Perspective Fields for Single Image Camera CalibrationLinyi Jin, Jianming Zhang, Yannick Hold-Geoffroy, Oliver Wang et al.CVPR 2023
- Camera Control for Text-to-Image Generation via Learning Viewpoint TokensXinxuan Lu, Charless Fowlkes, Alexander C. BergCVPR 2026 · 2 citations
- Unified Camera Positional Encoding for Controlled Video GenerationCheng Zhang, Boying Li, Meng Wei, Yan-Pei Cao et al.CVPR 2026 · 38 citations
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
