ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering
Yindong Zhang, Wenmian Yang, Yiquan Zhang, Weijia Jia
摘要
Developing agents capable of navigating fragmented, multi-source information remains challenging, primarily due to the scarcity of benchmarks reflecting hybrid workflows combining database querying with external APIs. To bridge this gap, we introduce ReCoQA, a largescale benchmark of 29,270 real-estate instances featuring machine-verifiable supervision for intermediate steps, including structured intent labels, SQL queries, and API calls. Complementarily, we propose HIRE-Agent, a hierarchical framework instantiating an understand-plan-execute architecture as a strong baseline. By orchestrating a Front-end parser, a planning Supervisor, and execution Specialists, HIRE-Agent effectively integrates heterogeneous evidence. Extensive experiments demonstrate that HIRE-Agent constitutes a strong baseline and substantiates the necessity of hierarchical collaboration for complex, real-world reasoning tasks. The benchmark and source code are available at: https://github.com/ Husky-989/ReCoQA
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Co-GAT: A Co-Interactive Graph Attention Network for Joint Dialog Act Recognition and Sentiment ClassificationLibo Qin, Zhouyang Li, Wanxiang Che, Minheng Ni 等AAAI 2021 · 被引用 77 次
- Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear QueriesXinyi He, Mengyu Zhou, Xinrun Xu, Xiaojun Ma 等AAAI 2024 · 被引用 48 次
- Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical ReasoningPan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu 等ICLR 2023 · 被引用 41 次
- Toward a Unified Framework for Unsupervised Complex Tabular ReasoningZhenyu Li, Xiuxing Li, Zhichao Duan, Bowen Dong 等ICDE 2023 · 被引用 4 次
- RETQA: A Large-Scale Open-Domain Tabular Question Answering Dataset for Real Estate SectorZhensheng Wang, Wenmian Yang, Kun Zhou, Yiquan Zhang 等AAAI 2025 · 被引用 3 次
相关 Paper
- HiRA: Decoupling Planning and Execution with Hierarchical Reasoning in Deep SearchJiajie Jin, Xiaoxi Li, Yuyao Zhang, Guanting Dong 等SIGIR 2026
- GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository LeveragingZiyi Ni, Huacan Wang, Shuo Zhang, Shuo Lu 等AAAI 2026 · 被引用 13 次
- LakeQA: An Exploratory QA Benchmark over a Million-Scale Data LakeHaonan Wang, Jiaxiang Liu, Yurong Liu, Austin Wijaya 等ICML 2026
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and ReasoningLiang Hu, Jianpeng Jiao, Jiashuo Liu, Dongyuan Mutu 等ICLR 2026 · 被引用 29 次
- WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation ModelsRui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang 等ACL 2026 · 被引用 4 次
