Online 3D Bin Packing with Constrained Deep Reinforcement Learning
Hang Zhao, Qijin She, Chenyang Zhu, Yin Yang, Kai Xu
摘要
We solve a challenging yet practically useful variant of 3D Bin Packing Problem (3D-BPP). In our problem, the agent has limited information about the items to be packed into a single bin, and an item must be packed immediately after its arrival without buffering or readjusting. The item's placement also subjects to the constraints of order dependence and physical stability. We formulate this online 3D-BPP as a constrained Markov decision process (CMDP). To solve the problem, we propose an effective and easy-to-implement constrained deep reinforcement learning (DRL) method under the actor-critic framework. In particular, we introduce a prediction-and-projection scheme: The agent first predicts a feasibility mask for the placement actions as an auxiliary task and then uses the mask to modulate the action probabilities output by the actor during training. Such supervision and projection facilitate the agent to learn feasible policies very efficiently. Our method can be easily extended to handle lookahead items, multi-bin packing, and item re-orienting. We have conducted extensive evaluation showing that the learned policy significantly outperforms the state-of-the-art methods. A preliminary user study even suggests that our method might attain a human-level performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Learning to Search Feasible and Infeasible Regions of Routing Problems with Flexible Neural k-OptYining Ma, Zhiguang Cao, Yeow Meng CheeNeurIPS 2023 · 被引用 129 次
- AoI-minimal UAV Crowdsensing by Model-based Graph Convolutional Reinforcement LearningZipeng Dai, Chi Harold Liu, Yuxiao Ye, Rui Han 等INFOCOM 2022 · 被引用 72 次
- Learning to Handle Complex Constraints for Vehicle Routing ProblemsJieyi Bi, Yining Ma, Jianan Zhou, Wen Song 等NeurIPS 2024 · 被引用 62 次
- WISK: A Workload-aware Learned Index for Spatial Keyword QueriesYufan Sheng, Xin Cao, Yixiang Fang, Kaiqi Zhao 等SIGMOD 2023 · 被引用 24 次
- Virne: A Comprehensive Benchmark for RL-based Network Resource Allocation in NFVTianfu Wang, Liwei Deng, Xi Chen, Junyang Wang 等ICLR 2026 · 被引用 2 次
相关 Paper
- Learning Efficient Online 3D Bin Packing on Packing Configuration TreesHang Zhao, Yang Yu, Kai XuICLR 2022 · 被引用 56 次
- Adjustable Robust Reinforcement Learning for Online 3D Bin PackingYuxin Pan, Yize Chen, Fangzhen LinNeurIPS 2023 · 被引用 23 次
- Deep Reinforcement Learning for Scalable Offline Three-Dimensional PackingHao Yin, Hongjie He, Fan ChenAAAI 2026
- ASAP: Exploiting the Satisficing Generalization Edge in Neural Combinatorial OptimizationHan Fang, Paul Weng, Yutong BanICML 2026 · 被引用 1 次
- Learning to solve Class-Constrained Bin Packing Problems via Encoder-Decoder ModelHanni Cheng, Ya Cong, Weihao Jiang, Shiliang PuICLR 2024 · 被引用 2 次
