Chain-of-Questions Training with Latent Answers for Robust Multistep Question Answering
Wang Zhu, Jesse Thomason, Robin Jia
Abstract
We propose Chain-of-Questions, a framework that trains a model to robustly answer multistep questions by generating and answering sub-questions. We obtain supervision for subquestions from human-annotated question decomposition meaning representation (QDMR), but QDMR does not include annotated answers to sub-questions. To overcome this technical challenge, we treat sub-answers as latent variables and infer them with a novel dynamic mixture of Hard-EM and MAPO. Chain-of-Questions is effective and robust, greatly outperforming strong neuro-symbolic methods by 9.0 F1 on a DROP contrast set and GPT-3.5 by 24.3 F1 on a HOTPOTQA adversarial set. Question Context Ground-truth QDMR Generated QDMR & Answers How many years after Pegu fell did the king die? (DROP) After the fall of Pegu in December 1599 ... but the king died during the campaign on 3 March 1606.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d0a20c8-109f-4967-9e2e-9b32b6abc982Cited by top-tier papers2
- Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting LayersYoumin Ko, Sungjong Seo, Hyunjoon KimNeurIPS 2025 · 2 citations
- POQD: Performance-Oriented Query Decomposer for Multi-vector retrievalYaoyang Liu, Junlin Li, Yinjun Wu, Zhen ChenICML 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- Neural Symbolic Reader: Scalable Integration of Distributed and Symbolic Representations for Reading ComprehensionXinyun Chen, Chen Liang, Adams Wei Yu, Denny Zhou et al.ICLR 2020 · 109 citations
Related papers
- Unsupervised Question Decomposition for Question AnsweringEthan Perez, Patrick Lewis, Wen-tau Yih, Kyunghyun Cho et al.EMNLP 2020 · 6 citations
- Low-Resource Generation of Multi-hop Reasoning QuestionsJianxing Yu, Wei Liu, Shuang Qiu, Qinliang Su et al.ACL 2020 · 11 citations
- Generating Multi-hop Reasoning Questions to Improve Machine Reading ComprehensionJianxing Yu, Xiaojun Quan, Qinliang Su, Jian YinWWW 2020 · 25 citations
- Mitigating Lost-in-Retrieval Problems in Retrieval Augmented Multi-Hop Question AnsweringRongzhi Zhu, Xiangyu Liu, Zequn Sun, Yiwei Wang et al.ACL 2025 · 14 citations
- Robustifying Multi-hop QA through Pseudo-Evidentiality TrainingKyungjae Lee, Seung-won Hwang, Sang-eun Han, Dohyeon LeeACL 2021
