ICML2026

Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models

Wei Liu, Peijie Yu, Michele Orini, Yali Du, Yulan He

被引用 2 次

摘要

The agency expected of Agentic Large Language Models goes beyond answering correctly, requiring autonomy to set goals and decide what to explore. We term this investigatory intelligence , distinguishing it from executional intelligence , which merely completes assigned tasks. Data Science provides a natural testbed, as real-world analysis starts from raw data rather than explicit queries, yet few benchmarks focus on it. To address this, we introduce Deep Data Research (DDR) , an open-ended task where LLMs autonomously extract key insights from databases, and DDR-Bench , a large-scale, checklist-based benchmark that enables verifiable evaluation. Results show that while frontier models display emerging agency, long-horizon exploration remains challenging. Our analysis highlights that effective investigatory intelligence depends not only on agent scaffolding or merely scaling, but also on intrinsic strategies of agentic models.