Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
Kang Liao, Size Wu, Zhonghua Wu, Linyi Jin, Chao Wang, Yikai Wang, Fei Wang, Wei Li, Chen Change Loy
摘要
Ground Truth UVP GeoCalib Puffin (Ours) (a) Camera Controllable Generation (b) Camera Understanding Horizon Lines: <think> (Roll) The image exhibits a large clockwise Dutch angle, as evidenced by the tilted perspective where the tree and roofline are not aligned with the natural horizontal plane, suggesting a significant rotation of the camera. (Pitch) The vertical orientation indicates a large tilt-up, as the sky occupies a substantial portion of the frame, with the tree and roofline positioned lower, emphasizing an upward view. (FoV) The image exhibits a medium shot, capturing a portion of the tree and roof while maintaining a balanced composition that includes both the sky and the elements in the foreground. </think> <answer>0.5107, 0.5550, 0.7558</answer> Prompt: Reason the spatial distribution of this image in a thinking mode, and then estimate its camera parameters (roll, pitch, and FoV, in radius). Prompt: Given a scene description and corresponding camera parameters, merge them into a coherent prompt and generate an accurate visualization that highlights visual cues. Caption: A sleek, modern interior with a large, minimalistic window framed in light beige, offering a view of a textured blue stone wall reflecting glimpses of a nearby building and lush greenery. The camera parameters (roll, pitch, and FoV) are: -0.2761, -0.1348, 0.9415. <think> The slight tilt of the window frame relative to the vertical lines of the wall suggests a small counterclockwise Dutch angle in the horizontal orientation. The visible portion of the window and the wall indicates a medium shot, capturing a portion of the interior and exterior without showing the entire scene. The subtle downward angle of the camera, evident from the alignment of the windowsill and the wall, suggests a small tilt-down in the vertical orientation. </think> Error Map Generated Image Error=7.5°Error=2.7°L atitude Camera Map
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou 等CVPR 2026 · 被引用 33 次
- Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and CroppingTianxiang Du, Hulingxiao He, Yuxin PengCVPR 2026 · 被引用 3 次
- AesFormer: Transform Everyday Photos into Beautiful MemoriesTianxiang Du, Hulingxiao He, Yuxin PengICML 2026
- CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise WarpingHaoyu Zhao, Jiaxi Gu, Haoran Chen, Qingping Zheng 等ICML 2026
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
相关 Paper
- SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual GroundingRong Li, Shijie Li, Lingdong Kong, Xulei Yang 等CVPR 2025
- Perspective Fields for Single Image Camera CalibrationLinyi Jin, Jianming Zhang, Yannick Hold-Geoffroy, Oliver Wang 等CVPR 2023
- Camera Control for Text-to-Image Generation via Learning Viewpoint TokensXinxuan Lu, Charless Fowlkes, Alexander C. BergCVPR 2026 · 被引用 2 次
- Unified Camera Positional Encoding for Controlled Video GenerationCheng Zhang, Boying Li, Meng Wei, Yan-Pei Cao 等CVPR 2026 · 被引用 38 次
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
