Learning Dialog Policies from Weak Demonstrations
Gabriel Gordon-Hall, Philip John Gorinski, Shay B. Cohen
Abstract
Deep reinforcement learning is a promising approach to training a dialog manager, but current methods struggle with the large state and action spaces of multi-domain dialog systems. Building upon Deep Q-learning from Demonstrations (DQfD), an algorithm that scores highly in difficult Atari games, we leverage dialog data to guide the agent to successfully respond to a user's requests. We make progressively fewer assumptions about the data needed, using labeled, reduced-labeled, and even unlabeled data to train expert demonstrators. We introduce Reinforced Fine-tune Learning, an extension to DQfD, enabling us to overcome the domain gap between the datasets and the environment. Experiments in a challenging multi-domain dialog system framework validate our approaches, and get high success rates even when trained on out-of-domain data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d184adb2-9cf7-4f89-ab5e-99731e3e3b1aCited by top-tier papers1
Ask how each one uses itBuilds on2
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta et al.AAAI 2020 · 707 citations
Related papers
- Transferable Dialogue Systems and User SimulatorsBo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, Bill ByrneACL 2021
- Meta-Reinforced Multi-Domain State Generator for Dialogue SystemsYi Huang, Junlan Feng, Min Hu, Xiaoting Wu et al.ACL 2020 · 30 citations
- Learning Efficient Dialogue Policy from Demonstrations through ShapingHuimin Wang, Baolin Peng, Kam-Fai WongACL 2020 · 18 citations
- Multi-Agent Task-Oriented Dialog Policy Learning with Role-Aware Reward DecompositionRyuichi Takanobu, Runze Liang, Minlie HuangACL 2020 · 47 citations
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin et al.ICML 2022 · 32 citations
