Online 3D Bin Packing with Constrained Deep Reinforcement Learning
Hang Zhao, Qijin She, Chenyang Zhu, Yin Yang, Kai Xu
Abstract
We solve a challenging yet practically useful variant of 3D Bin Packing Problem (3D-BPP). In our problem, the agent has limited information about the items to be packed into a single bin, and an item must be packed immediately after its arrival without buffering or readjusting. The item's placement also subjects to the constraints of order dependence and physical stability. We formulate this online 3D-BPP as a constrained Markov decision process (CMDP). To solve the problem, we propose an effective and easy-to-implement constrained deep reinforcement learning (DRL) method under the actor-critic framework. In particular, we introduce a prediction-and-projection scheme: The agent first predicts a feasibility mask for the placement actions as an auxiliary task and then uses the mask to modulate the action probabilities output by the actor during training. Such supervision and projection facilitate the agent to learn feasible policies very efficiently. Our method can be easily extended to handle lookahead items, multi-bin packing, and item re-orienting. We have conducted extensive evaluation showing that the learned policy significantly outperforms the state-of-the-art methods. A preliminary user study even suggests that our method might attain a human-level performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b09d8041-649c-4502-aff9-dc8d7c3a026aCited by top-tier papers9
- Learning to Search Feasible and Infeasible Regions of Routing Problems with Flexible Neural k-OptYining Ma, Zhiguang Cao, Yeow Meng CheeNeurIPS 2023 · 129 citations
- AoI-minimal UAV Crowdsensing by Model-based Graph Convolutional Reinforcement LearningZipeng Dai, Chi Harold Liu, Yuxiao Ye, Rui Han et al.INFOCOM 2022 · 72 citations
- Learning to Handle Complex Constraints for Vehicle Routing ProblemsJieyi Bi, Yining Ma, Jianan Zhou, Wen Song et al.NeurIPS 2024 · 62 citations
- WISK: A Workload-aware Learned Index for Spatial Keyword QueriesYufan Sheng, Xin Cao, Yixiang Fang, Kaiqi Zhao et al.SIGMOD 2023 · 24 citations
- Virne: A Comprehensive Benchmark for RL-based Network Resource Allocation in NFVTianfu Wang, Liwei Deng, Xi Chen, Junyang Wang et al.ICLR 2026 · 2 citations
Related papers
- Learning Efficient Online 3D Bin Packing on Packing Configuration TreesHang Zhao, Yang Yu, Kai XuICLR 2022 · 56 citations
- Adjustable Robust Reinforcement Learning for Online 3D Bin PackingYuxin Pan, Yize Chen, Fangzhen LinNeurIPS 2023 · 23 citations
- Deep Reinforcement Learning for Scalable Offline Three-Dimensional PackingHao Yin, Hongjie He, Fan ChenAAAI 2026
- ASAP: Exploiting the Satisficing Generalization Edge in Neural Combinatorial OptimizationHan Fang, Paul Weng, Yutong BanICML 2026 · 1 citation
- Learning to solve Class-Constrained Bin Packing Problems via Encoder-Decoder ModelHanni Cheng, Ya Cong, Weihao Jiang, Shiliang PuICLR 2024 · 2 citations
