Towards Smart Point-and-Shoot Photography
Jiawan Li, Fei Zhou, Zhipeng Zhong, Jiongzhi Lin, Guoping Qiu
摘要
Hundreds of millions of people routinely take photos using their smartphones as point and shoot (PAS) cameras, yet very few would have the photography skills to compose a good shot of a scene. While traditional PAS cameras have built-in functions to ensure a photo is well focused and has the right brightness, they cannot tell the users how to compose the best shot of a scene. In this paper, we present a first of its kind smart point and shoot (SPAS) system to help users to take good photos. Our SPAS proposes to help users to compose a good shot of a scene by automatically guiding the users to adjust the camera pose live on the scene. We first constructed a large dataset containing 320K images with camera pose information from 4000 scenes. We then developed an innovative CLIP-based Composition Quality Assessment (CCQA) model to assign pseudo labels to these images. The CCQA introduces a unique learnable text embedding technique to learn continuous word embeddings capable of discerning subtle visual quality differences in the range covered by five levels of quality description words bad, poor, f air, good, perf ect. And finally we have developed a camera pose adjustment model (CPAM) which first determines if the current view can be further improved and if so it outputs the adjust suggestion in the form of two camera pose adjustment angles. The two tasks of CPAM make decisions in a sequential manner and each involves different sets of training samples, we have developed a mixture-of-experts model with a gated loss function to train the CPAM in an end-to-end manner. We will present extensive results to demonstrate the performances of our SPAS system using publicly available image composition datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PhotoFramer: Multi-modal Image Composition InstructionZhiyuan You, Ke Wang, He Zhang, Xin Cai 等CVPR 2026 · 被引用 8 次
- Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and CroppingTianxiang Du, Hulingxiao He, Yuxin PengCVPR 2026 · 被引用 3 次
- Aesthetic Camera Viewpoint Suggestion with 3D Aesthetic FieldSheyang Tang, Armin Shafiee Sarvestani, Jialu Xu, Xiaoyu Xu 等CVPR 2026 · 被引用 1 次
- AesFormer: Transform Everyday Photos into Beautiful MemoriesTianxiang Du, Hulingxiao He, Yuxin PengICML 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen 等ICML 2024 · 被引用 499 次
- Image Cropping with Composition and Saliency Aware Aesthetic Score MapYi Tu, Li Niu, Weijie Zhao, Dawei Cheng 等AAAI 2020 · 被引用 55 次
- Spherical Criteria for Fast and Accurate 360° Object DetectionPengyu Zhao, Ansheng You, Yuanxing Zhang, Jiaying Liu 等AAAI 2020 · 被引用 35 次
- TransView: Inside, Outside, and Across the Cropping View BoundariesZhiyu Pan, Zhiguo Cao, Kewei Wang, Hao Lu 等ICCV 2021 · 被引用 21 次
相关 Paper
- Perceptual Quality Assessment of Smartphone PhotographyYuming Fang, Hanwei Zhu, Yan Zeng, Kede Ma 等CVPR 2020
- SQAD: Automatic Smartphone Camera Quality Assessment and BenchmarkingZilin Fang, Andrey Ignatov, Eduard Zamfir, Radu TimofteICCV 2023 · 被引用 6 次
- Photography Perspective Composition: Towards Aesthetic Perspective RecommendationLujian Yao, Siming Zheng, Xinbin Yuan, Zhuoxuan Cai 等NeurIPS 2025 · 被引用 2 次
- CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality AssessmentYating Liu, Yujie Zhang, Ziyu Shan, Yiling XuAAAI 2025 · 被引用 9 次
- Time-Aware Auto White Balance in Mobile PhotographyMahmoud Afifi, Luxi Zhao, Abhijith Punnappurath, Mohammed A. Abdelsalam 等ICCV 2025 · 被引用 12 次
