Reward Finetuning for Faster and More Accurate Unsupervised Object Discovery
Katie Luo, Zhenzhen Liu, Xiangyu Chen, Yurong You, Sagie Benaim, Cheng Perng Phoo, Mark E. Campbell, Wen Sun, Bharath Hariharan, Kilian Q. Weinberger
摘要
Recent advances in machine learning have shown that Reinforcement Learning from Human Feedback (RLHF) can improve machine learning models and align them with human preferences. Although very successful for Large Language Models (LLMs), these advancements have not had a comparable impact in research for autonomous vehicles-where alignment with human expectations can be imperative. In this paper, we propose to adapt similar RL-based methods to unsupervised object discovery, i.e. learning to detect objects from LiDAR points without any training labels. Instead of labels, we use simple heuristics to mimic human feedback. More explicitly, we combine multiple heuristics into a simple reward function that positively correlates its score with bounding box accuracy, i.e., boxes containing objects are scored higher than those without. We start from the detector's own predictions to explore the space and reinforce boxes with high rewards through gradient updates. Empirically, we demonstrate that our approach is not only more accurate, but also orders of magnitudes faster to train compared to prior works on object discovery. Code is available at https://github.com/katieluo88/DRIFT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous DrivingShuyao Shang, Yuntao Chen, Yuqi Wang, Yingyan Li 等NeurIPS 2025 · 被引用 49 次
- UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-ClassesTed de Vries Lentsch, Holger Caesar, Dariu GavrilaNeurIPS 2024 · 被引用 30 次
- DiffuBox: Refining 3D Object Detection with Point DiffusionXiangyu Chen, Zhenzhen Liu, Katie Luo, Siddhartha Datta 等NeurIPS 2024 · 被引用 10 次
- Seg2Box: 3D Object Detection by Point-Wise Semantics SupervisionMaoji Zheng, Ziyu Xu, Qiming Xia, Hai Wu 等AAAI 2025 · 被引用 3 次
- MonoSOWA: Scalable Monocular 3D Object Detector Without Human AnnotationsJan Skvrna, Lukás NeumannICCV 2025 · 被引用 3 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
相关 Paper
- Learning to Detect Mobile Objects from LiDAR Scans Without LabelsYurong You, Katie Luo, Cheng Perng Phoo, Wei-Lun Chao 等CVPR 2022 · 被引用 33 次
- Unsupervised Learning of Object Landmarks via Self-Training CorrespondenceDimitrios Mallis, Enrique Sanchez, Matthew Bell, Georgios TzimiropoulosNeurIPS 2020 · 被引用 20 次
- BAR - A Reinforcement Learning Agent for Bounding-Box Automated RefinementMorgane Ayle, Jimmy Tekli, Julia El Zini, Boulos El Asmar 等AAAI 2020 · 被引用 9 次
- Unsupervised Adaptation from Repeated Traversals for Autonomous DrivingYurong You, Cheng Perng Phoo, Katie Luo, Travis Zhang 等NeurIPS 2022 · 被引用 16 次
- Learning to Detect Objects from Multi-Agent LiDAR Scans without Manual LabelsQiming Xia, Wenkai Lin, Haoen Xiang, Xun Huang 等CVPR 2025
