DISeR: Designing Imaging Systems with Reinforcement Learning
Tzofi Klinghoffer, Kushagra Tiwary, Nikhil Behari, Bhavya Agrawalla, Ramesh Raskar
Abstract
Imaging systems consist of cameras to encode visual information about the world and perception models to interpret this encoding. Cameras contain (1) illumination sources, (2) optical elements, and (3) sensors, while perception models use (4) algorithms. Directly searching over all combinations of these four building blocks to design an imaging system is challenging due to the size of the search space. Moreover, cameras and perception models are often designed independently, leading to sub-optimal task performance. In this paper, we formulate these four building blocks of imaging systems as a context-free grammar (CFG), which can be automatically searched over with a learned camera designer to jointly optimize the imaging system with task-specific perception models. By transforming the CFG to a state-action space, we then show how the camera designer can be implemented with reinforcement learning to intelligently search over the combinatorial space of possible imaging system configurations. We demonstrate our approach on two tasks, depth estimation and camera rig design for autonomous vehicles, showing that our method yields rigs that outperform industry-wide standards. We believe that our proposed approach is an important step towards automating imaging system design. Our project page is https://tzofi.github.io/diser .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5cad714-c48a-4de4-91b0-2888e6d4ed84Cited by top-tier papers2
- Task-Driven Implicit Representations for Automated Design of LiDAR SystemsNikhil Behari, Aaron Young, Tzofi Klinghoffer, Akshat Dave et al.CVPR 2026
- Grammar Reinforcement Learning: path and cycle counting in graphs with a Context-Free Grammar and Transformer approachJason Piquenot, Maxime Berar, Romain Raveaux, Pierre Héroux et al.ICLR 2025
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Cross-view Transformers for real-time Map-view Semantic SegmentationBrady Zhou, Philipp KrähenbühlCVPR 2022 · 279 citations
- Deep Optics for Monocular Depth Estimation and 3D Object DetectionJulie Chang, Gordon WetzsteinICCV 2019 · 219 citations
- Data-Efficient Graph Grammar Learning for Molecular GenerationMinghao Guo, Veronika Thost, Beichen Li, Payel Das et al.ICLR 2022 · 46 citations
Related papers
- AdaptiveISP: Learning an Adaptive Image Signal Processor for Object DetectionYujin Wang, Tianyi Xu, Zhang Fan, Tianfan Xue et al.NeurIPS 2024 · 38 citations
- MAS-ISP: A Proxy-Free Online Hyperparameter Optimization Framework for ISP Hardware SystemJiaming Liu, Xuan Huang, Zhijian Hao, Ruoxi Zhu et al.DAC 2025 · 3 citations
- Parameterized Decision-Making with Multi-Modality Perception for Autonomous DrivingYuyang Xia, Shuncheng Liu, Quanlin Yu, Liwei Deng et al.ICDE 2024 · 26 citations
- RL-SeqISP: Reinforcement Learning-Based Sequential Optimization for Image Signal ProcessingXinyu Sun, Zhikun Zhao, Lili Wei, Congyan Lang et al.AAAI 2024 · 13 citations
- Rig3R: Rig-Aware Conditioning and Discovery for 3D ReconstructionSamuel Li, Pujith Kachana, Prajwal Chidananda, Saurabh Nair et al.NeurIPS 2025 · 7 citations
