PreciseCam: Precise Camera Control for Text-to-Image Generation
Edurne Bernal-Berdun, Ana Serrano, Belén Masiá, Matheus Gadelha, Yannick Hold-Geoffroy, Xin Sun, Diego Gutierrez
摘要
A photograph of an inviting reading room with a large armchair, a low wooden table, and floor-to-ceiling bookshelves packed with books, softly lit by a standing lamp. An impressionist painting of a serene lakeside at dawn. Soft brushstrokes blend the colors of the sky, water, and distant mountains. The lake reflects the pastel hues of the rising sun, and small details of trees and boats blur into the overall mood of tranquility. Pitch Figure 1. Our approach enhances the artistic expression of text-to-image generative models by incorporating precise control over camera angles and lens distortion effects. Left: Our input consists of a standard text prompt along with extrinsic (roll and pitch) and intrinsic (vertical field of view and distortion ξ) camera parameters, which are translated into a suitable and efficient representation for learning camera views. Right: Examples varying roll (top) and pitch (bottom) with the same prompt, while keeping the remaining camera parameters fixed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Unified Camera Positional Encoding for Controlled Video GenerationCheng Zhang, Boying Li, Meng Wei, Yan-Pei Cao 等CVPR 2026 · 被引用 38 次
- EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video GuidanceZun Wang, Jaemin Cho, Jialu Li, Han Lin 等ICML 2026 · 被引用 19 次
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and GenerationKang Liao, Size Wu, Zhonghua Wu, Linyi Jin 等ICLR 2026 · 被引用 19 次
- SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image GenerationVaibhav Agrawal, Rishubh Parihar, Pradhaan Bhat, Ravi Kiran Sarvadevabhatla 等CVPR 2026 · 被引用 5 次
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept PersonalizationTsai-Shien Chen, Aliaksandr Siarohin, Gordon Guocheng Qian, Kuan-Chieh Jackson Wang 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper18
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- ImageSet2Text: Describing Sets of Images Through TextPiera Riccio, Francesco Galati, Kajetan Schweighofer, Noa Garcia 等AAAI 2026 · 被引用 1 次
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
- UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in RLRui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu 等CVPR 2026
- Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image SynthesisYu Yuan, Xijun Wang, Yichen Sheng, Prateek Chennuri 等CVPR 2025
- Guided Score identity Distillation for Data-Free One-Step Text-to-Image GenerationMingyuan Zhou, Zhendong Wang, Huangjie Zheng, Hai HuangICLR 2025
