SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
Binyuan Huang, Yuqing Wen, Yucheng Zhao, Yaosi Hu, Yingfei Liu, Fan Jia, Weixin Mao, Tiancai Wang, Chi Zhang, Chang Wen Chen, Zhenzhong Chen, Xiangyu Zhang
Abstract
Autonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data for autonomous driving applications and present SubjectDrive, the first model proven to scale generative data production in a way that could continuously improve autonomous driving applications. We investigate the impact of scaling up the quantity of generative data on the performance of downstream perception models and find that enhancing data diversity plays a crucial role in effectively scaling generative data production. Therefore, we have developed a novel model equipped with a subject control mechanism, which allows the generative model to leverage diverse external data sources for producing varied and useful data. Extensive evaluations confirm SubjectDrive's efficacy in generating scalable autonomous driving training data, marking a significant step toward revolutionizing data production methods in this field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef6d9336-e67e-4ccd-beb3-98f30e532ee6Cited by top-tier papers9
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real WorldAo Liang, Lingdong Kong, Tianyi Yan, Hongsi Liu et al.CVPR 2026 · 28 citations
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilityYu Yang, Alan Liang, Jianbiao Mei, Yukai Ma et al.NeurIPS 2025 · 22 citations
- DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationJiazhe Guo, Yikang Ding, Xiwu Chen, Shuo Chen et al.ICCV 2025 · 5 citations
- Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video GenerationBinyuan Huang, Yuning Lu, Weinan Jia, Hualiang Wang et al.CVPR 2026 · 3 citations
- LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge PointsXuemiao Zhang, Can Ren, Chengying Tu, Rongxiang Weng et al.ACL 2026 · 3 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
Related papers
- PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion ModelJinhua Zhang, Hualian Sheng, Sijia Cai, Bing Deng et al.ICCV 2025 · 2 citations
- Panacea: Panoramic and Controllable Video Generation for Autonomous DrivingYuqing Wen, Yucheng Zhao, Yingfei Liu, Fan Jia et al.CVPR 2024 · 26 citations
- SimGen: Simulator-conditioned Driving Scene GenerationYunsong Zhou, Michael Simon, Zhenghao Mark Peng, Sicheng Mo et al.NeurIPS 2024 · 44 citations
- Rethinking Driving World Model as Synthetic Data Generator for Perception TasksKai Zeng, Zhanqian Wu, Kaixin Xiong, Xiaobao Wei et al.ICLR 2026 · 14 citations
- Generalized Predictive Model for Autonomous DrivingJiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen et al.CVPR 2024 · 32 citations
