Lightweight and Faithful Visual Condition Checking in Behavior Trees via Expert-Regularized Reinforcement Learning
Hyosik Moon, Eldan Cohen
Abstract
Behavior trees provide a transparent and modular structure for encoding expert-designed policies, enabling interpretable decision-making in complex tasks. Yet, applying behavior trees to high-dimensional perceptual inputs such as images or language is challenging as defining symbolic predicates over raw perceptual data is non-trivial. While state-of-the-art large multimodal models (such as vision-language models) can overcome this issue by utilizing natural language queries over perceptual inputs, they incur high computational cost, making them unsuitable for many applications. Imitation learning offers a way to distill these expert models into compact models, though it requires extensive supervision. In contrast, reinforcement learning reduces the need for costly supervision but risks misalignment of condition nodes with their intended semantics as well as poor credit assignment. To address these challenges, we introduce CERL (Condition-node Expert-regularized Reinforcement Learning), a framework that leverages expert-regularized reinforcement learning to preserve semantic faithfulness, while employing a factorized policy that aggregates sequential condition-node decisions into a single decision unit to alleviate credit assignment challenges. Experiments across seven tasks from the GymCards, FrozenLake, and BabyAIText suites demonstrate that our framework outperforms pure imitation learning or reinforcement learning baselines, retains strong agreement with expert decisions, and achieves substantial gains in inference speed and model size over expert models. Our implementation is available in https://github.com/HyosikMoon/CERL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 466 citations
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong et al.NeurIPS 2021 · 207 citations
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma et al.ICLR 2024 · 31 citations
- Guarded Policy Optimization with Imperfect Online DemonstrationsZhenghai Xue, Zhenghao Peng, Quanyi Li, Zhihan Liu et al.ICLR 2023 · 4 citations
Related papers
- VisRL: Intention-Driven Visual Perception via Reinforced ReasoningZhangquan Chen, Xufang Luo, Dongsheng LiICCV 2025 · 2 citations
- Embodied Navigation with Auxiliary Task of Action Description PredictionHaru Kondoh, Asako KanezakiICCV 2025 · 3 citations
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language ReasoningXiaojun Guo, Runyu Zhou, Yifei Wang, Qi Zhang et al.ICML 2026 · 6 citations
- Object-Aware Regularization for Addressing Causal Confusion in Imitation LearningJongjin Park, Younggyo Seo, Chang Liu, Li Zhao et al.NeurIPS 2021 · 31 citations
- ELEMENTAL: Interactive Learning from Demonstrations and Vision-Language Models for Reward Design in RoboticsLetian Chen, Nina Marie Moorman, Matthew Craig GombolayICML 2025
