GAIT: Generating Aesthetic Indoor Tours with Deep Reinforcement Learning
Desai Xie, Ping Hu, Xin Sun, Sören Pirk, Jianming Zhang, Radomír Mech, Arie E. Kaufman
Abstract
Placing and orienting a camera to compose aesthetically meaningful shots of a scene is not only a key objective in real-world photography and cinematography but also for virtual content creation. The framing of a camera often significantly contributes to the story telling in movies, games, and mixed reality applications. Generating single camera poses or even contiguous trajectories either requires a significant amount of manual labor or requires solving high-dimensional optimization problems, which can be computationally demanding and error-prone. In this paper, we introduce GAIT, a framework for training a Deep Reinforcement Learning (DRL) agent, that learns to automatically control a camera to generate a sequence of aesthetically meaningful views for synthetic 3D indoor scenes. To generate sequences of frames with high aesthetic value, GAIT relies on a neural aesthetics estimator, which is trained on a crowed-sourced dataset. Additionally, we introduce regularization techniques for diversity and smoothness to generate visually interesting trajectories for a 3D environment, and to constrain agent acceleration in the reward function to generate a smooth sequence of camera frames. We validated our method by comparing it to baseline algorithms, based on a perceptual user study, and through ablation studies. Code and visual results are available on the project website: https://desaixie.github.io/gait-rl
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 068aab9f-3f78-4759-922d-50c0345cdefdCited by top-tier papers3
- Pulp Motion: Framing-aware multimodal camera and human motion generationRobin Courant, Xi WANG, David Loiseaux, Marc Christie et al.ICLR 2026 · 8 citations
- InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene ComplexityHaoming Wang, Qiyao Xue, Wei GaoCVPR 2026 · 6 citations
- Aesthetic Camera Viewpoint Suggestion with 3D Aesthetic FieldSheyang Tang, Armin Shafiee Sarvestani, Jialu Xu, Xiaoyu Xu et al.CVPR 2026 · 1 citation
Builds on10
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos et al.AAAI 2021 · 506 citations
Related papers
- Optimization-based User Support for Cinematographic Quadrotor Camera Target FramingChristoph Gebhardt, Otmar HilligesCHI 2021 · 7 citations
- Deep Reinforcement Learning for Active Human Pose EstimationErik Gärtner, Aleksis Pirinen, Cristian SminchisescuAAAI 2020 · 27 citations
- Aesthetics-Driven Virtual Time-Lapse Photography GenerationLihua Lu, Hui Wei, Xin Jin, Yihao Zhang et al.ACM MM 2023 · 1 citation
- An Interactive System for Supporting Creative Exploration of Cinematic Composition DesignsRui He, Huaxin Wei, Ying CaoUIST 2024 · 10 citations
- ScenePhotographer: Object-Oriented Photography for Residential ScenesShao-Kui Zhang, Hanxi Zhu, Xuebin Chen, Jinghuan Chen et al.ACM MM 2024
