Speeding up Inference with User Simulators throughPolicy Modulation
Hee-Seung Moon, Seungwon Do, Wonjae Kim, Jiwon Seo, Minsuk Chang, Byungjoo Lee
摘要
The simulation of user behavior with deep reinforcement learning agents has shown some recent success. However, the inverse problem, that is, inferring the free parameters of the simulator from observed user behaviors, remains challenging to solve. This is because the optimization of the new action policy of the simulated agent, which is required whenever the model parameters change, is computationally impractical. In this study, we introduce a network modulation technique that can obtain a generalized policy that immediately adapts to the given model parameters. Further, we demonstrate that the proposed technique improves the efficiency of user simulator-based inference by eliminating the need to obtain an action policy for novel model parameters. We validated our approach using the latest user simulator for point-and-click behavior. Consequently, we succeeded in inferring the user’s cognitive parameters and intrinsic reward settings with less than 1/1000 computational power to those of existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Amortized Inference with User SimulationsHee-Seung Moon, Antti Oulasvirta, Byungjoo LeeCHI 2023 · 被引用 20 次
- Real-time 3D Target Inference via Biomechanical SimulationHee-Seung Moon, Yi-Chi Liao, Chenyu Li, Byungjoo Lee 等CHI 2024 · 被引用 15 次
- Amortised Experimental Design and Parameter Estimation for User Models of PointingAntti Keurulainen, Isak Rafael Westerlund, Oskar Keurulainen, Andrew HowesCHI 2023 · 被引用 7 次
- DisMouse: Disentangling Information from Mouse Movement DataGuanhua Zhang, Zhiming Hu, Andreas BullingUIST 2024 · 被引用 4 次
- Efficient Human-in-the-Loop Optimization via Priors Learned from User ModelsYi-Chi Liao, João Marcelo Evangelista Belo, Hee-Seung Moon, Jürgen Steimle 等CHI 2026 · 被引用 3 次
它引用的顶会 Paper6
- Multi-Task Reinforcement Learning with Soft ModularizationRuihan Yang, Huazhe Xu, Yi Wu, Xiaolong WangNeurIPS 2020 · 被引用 247 次
- Touchscreen Typing As Optimal Supervisory ControlJussi Jokinen, Aditya Acharya, Mohammad Uzair, Xinhui Jiang 等CHI 2021 · 被引用 105 次
- Predicting Mid-Air Interaction Movements and Fatigue Using Deep Reinforcement LearningNoshaba Cheema, Laura A. Frey-Law, Kourosh Naderi, Jaakko Lehtinen 等CHI 2020 · 被引用 66 次
- An Intermittent Click Planning ModelEunji Park, Byungjoo LeeCHI 2020 · 被引用 33 次
- A Simulation Model of Intermittently Controlled Point-and-Click BehaviourSeungwon Do, Minsuk Chang, Byungjoo LeeCHI 2021 · 被引用 31 次
相关 Paper
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg 等ICML 2020 · 被引用 81 次
- Inverse decision-making using neural amortized Bayesian actorsDominik Straub, Tobias F. Niehues, Jan Peters, Constantin A. RothkopfICLR 2025
- CogReact: A Reinforced Framework to Model Human Cognitive Reaction Modulated by Dynamic InterventionSonglin Xu, Xinyu ZhangICML 2025
- UserSim: User Simulation via Supervised GenerativeAdversarial NetworkXiangyu Zhao, Long Xia, Lixin Zou, Hui Liu 等WWW 2021 · 被引用 31 次
- Multi-Agent Task-Oriented Dialog Policy Learning with Role-Aware Reward DecompositionRyuichi Takanobu, Runze Liang, Minlie HuangACL 2020 · 被引用 47 次
