DISeR: Designing Imaging Systems with Reinforcement Learning
Tzofi Klinghoffer, Kushagra Tiwary, Nikhil Behari, Bhavya Agrawalla, Ramesh Raskar
摘要
Imaging systems consist of cameras to encode visual information about the world and perception models to interpret this encoding. Cameras contain (1) illumination sources, (2) optical elements, and (3) sensors, while perception models use (4) algorithms. Directly searching over all combinations of these four building blocks to design an imaging system is challenging due to the size of the search space. Moreover, cameras and perception models are often designed independently, leading to sub-optimal task performance. In this paper, we formulate these four building blocks of imaging systems as a context-free grammar (CFG), which can be automatically searched over with a learned camera designer to jointly optimize the imaging system with task-specific perception models. By transforming the CFG to a state-action space, we then show how the camera designer can be implemented with reinforcement learning to intelligently search over the combinatorial space of possible imaging system configurations. We demonstrate our approach on two tasks, depth estimation and camera rig design for autonomous vehicles, showing that our method yields rigs that outperform industry-wide standards. We believe that our proposed approach is an important step towards automating imaging system design. Our project page is https://tzofi.github.io/diser .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Task-Driven Implicit Representations for Automated Design of LiDAR SystemsNikhil Behari, Aaron Young, Tzofi Klinghoffer, Akshat Dave 等CVPR 2026
- Grammar Reinforcement Learning: path and cycle counting in graphs with a Context-Free Grammar and Transformer approachJason Piquenot, Maxime Berar, Romain Raveaux, Pierre Héroux 等ICLR 2025
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Cross-view Transformers for real-time Map-view Semantic SegmentationBrady Zhou, Philipp KrähenbühlCVPR 2022 · 被引用 279 次
- Deep Optics for Monocular Depth Estimation and 3D Object DetectionJulie Chang, Gordon WetzsteinICCV 2019 · 被引用 219 次
- Data-Efficient Graph Grammar Learning for Molecular GenerationMinghao Guo, Veronika Thost, Beichen Li, Payel Das 等ICLR 2022 · 被引用 46 次
相关 Paper
- AdaptiveISP: Learning an Adaptive Image Signal Processor for Object DetectionYujin Wang, Tianyi Xu, Zhang Fan, Tianfan Xue 等NeurIPS 2024 · 被引用 38 次
- MAS-ISP: A Proxy-Free Online Hyperparameter Optimization Framework for ISP Hardware SystemJiaming Liu, Xuan Huang, Zhijian Hao, Ruoxi Zhu 等DAC 2025 · 被引用 3 次
- Parameterized Decision-Making with Multi-Modality Perception for Autonomous DrivingYuyang Xia, Shuncheng Liu, Quanlin Yu, Liwei Deng 等ICDE 2024 · 被引用 26 次
- RL-SeqISP: Reinforcement Learning-Based Sequential Optimization for Image Signal ProcessingXinyu Sun, Zhikun Zhao, Lili Wei, Congyan Lang 等AAAI 2024 · 被引用 13 次
- Rig3R: Rig-Aware Conditioning and Discovery for 3D ReconstructionSamuel Li, Pujith Kachana, Prajwal Chidananda, Saurabh Nair 等NeurIPS 2025 · 被引用 7 次
