ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering
Yindong Zhang, Wenmian Yang, Yiquan Zhang, Weijia Jia
Abstract
Developing agents capable of navigating fragmented, multi-source information remains challenging, primarily due to the scarcity of benchmarks reflecting hybrid workflows combining database querying with external APIs. To bridge this gap, we introduce ReCoQA, a largescale benchmark of 29,270 real-estate instances featuring machine-verifiable supervision for intermediate steps, including structured intent labels, SQL queries, and API calls. Complementarily, we propose HIRE-Agent, a hierarchical framework instantiating an understand-plan-execute architecture as a strong baseline. By orchestrating a Front-end parser, a planning Supervisor, and execution Specialists, HIRE-Agent effectively integrates heterogeneous evidence. Extensive experiments demonstrate that HIRE-Agent constitutes a strong baseline and substantiates the necessity of hierarchical collaboration for complex, real-world reasoning tasks. The benchmark and source code are available at: https://github.com/ Husky-989/ReCoQA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85fe4e3c-90f5-42a6-9a80-054f259ca351Builds on6
- Co-GAT: A Co-Interactive Graph Attention Network for Joint Dialog Act Recognition and Sentiment ClassificationLibo Qin, Zhouyang Li, Wanxiang Che, Minheng Ni et al.AAAI 2021 · 77 citations
- Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear QueriesXinyi He, Mengyu Zhou, Xinrun Xu, Xiaojun Ma et al.AAAI 2024 · 48 citations
- Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical ReasoningPan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu et al.ICLR 2023 · 41 citations
- Toward a Unified Framework for Unsupervised Complex Tabular ReasoningZhenyu Li, Xiuxing Li, Zhichao Duan, Bowen Dong et al.ICDE 2023 · 4 citations
- RETQA: A Large-Scale Open-Domain Tabular Question Answering Dataset for Real Estate SectorZhensheng Wang, Wenmian Yang, Kun Zhou, Yiquan Zhang et al.AAAI 2025 · 3 citations
Related papers
- HiRA: Decoupling Planning and Execution with Hierarchical Reasoning in Deep SearchJiajie Jin, Xiaoxi Li, Yuyao Zhang, Guanting Dong et al.SIGIR 2026
- GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository LeveragingZiyi Ni, Huacan Wang, Shuo Zhang, Shuo Lu et al.AAAI 2026 · 13 citations
- LakeQA: An Exploratory QA Benchmark over a Million-Scale Data LakeHaonan Wang, Jiaxiang Liu, Yurong Liu, Austin Wijaya et al.ICML 2026
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and ReasoningLiang Hu, Jianpeng Jiao, Jiashuo Liu, Dongyuan Mutu et al.ICLR 2026 · 29 citations
- WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation ModelsRui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang et al.ACL 2026 · 4 citations
