ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
Pei-An Chen, Yong-Ching Liang, Jia-Fong Yeh, Hung-Ting Su, Yi-Ting Chen, Min Sun, Winston H. Hsu
Abstract
Intelligent embodied agents should not simply follow instructions, as real-world environments often involve unexpected conditions and exceptions. However, existing methods usually focus on directly executing instructions, without considering whether the target objects can actually be manipulated, meaning they fail to assess available affordances. To address this limitation, we introduce DynAfford, a benchmark that evaluates embodied agents in dynamic environments where object affordances may change over time and are not specified in the instruction. DynAfford requires agents to perceive object states, infer implicit preconditions, and adapt their actions accordingly. To enable this capability, we introduce ADAPT (Affordance-Driven Adaptive Planning and Task execution), a plug-and-play module that augments existing planners with explicit affordance reasoning. Experiments demonstrate that incorporating ADAPT significantly improves robustness and task success across both seen and unseen environments. We also show that a domain-adapted, LoRAfinetuned vision-language model used as the affordance inference backend outperforms a commercial LLM (GPT-4o), highlighting the importance of task-aligned affordance grounding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a0e3a7f-4e8b-4e4b-8137-fed00575206fCited by top-tier papers1
Ask how each one uses itBuilds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 228 citations
- Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied AgentsByeonghwi Kim, Jinyeon Kim, Yuyeong Kim, Cheolhong Min et al.ICCV 2023 · 46 citations
- Factorizing Perception and Policy for Interactive Instruction FollowingKunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi et al.ICCV 2021 · 39 citations
Related papers
- DetGPT: Detect What You Need via ReasoningRenjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan et al.EMNLP 2023 · 57 citations
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language ModelsHaowen Liu, Shaoxiong Yao, Haonan Chen, Jiawei Gao et al.CVPR 2026 · 8 citations
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao et al.ICML 2025
- VirtualEnv: A Platform for Embodied AI ResearchKabir Swain, Sijie Han, Ayush Raina, Jin Zhang et al.AAAI 2026
- ForeAct: Steering Your VLA with Efficient Visual Foresight PlanningZhuoyang Zhang, Shang Yang, Qinghao Hu, Luke J. Huang et al.CVPR 2026 · 7 citations
