Cognitive Scaffold: From Fluid Context to Crystallized Memory for Long-Horizon DeepResearch Agents
Qiuyuan Ai, Zenghuang Fu, Zhaoyang Li, Ping Jiang, Haoyu Wu, Jie Song, Guannan He
Abstract
Scaling LLM-based agents to long-horizon deep research is constrained by the “context-noise trade-off,” where linear history accumulation degrades reasoning and dilutes fine-grained evidence. To address this, we introduce the Cognitive Scaffold , a factorized memory architecture that decouples the cognitive state into a Fluid Working Context for immediate reasoning and a persistent Knowledge Graph for long-term retention. Unlike unstructured summarization, our framework employs a Re-jection Sampling Fine-Tuning (RFT) pipeline to crystallize saturated context into structured “event snapshots,” strictly enforcing atomic constraints to preserve numerical values and entities. During reasoning, a thought-driven dual-path retrieval mechanism enables the agent to proactively recover precise evidence. Empirical evaluations on Xbench-DeepSearch, BrowseComp-ZH, and GAIA demonstrate that Cognitive Scaffold consistently outperforms baselines, achieving 74.7% Avg@3 and 87.0% Pass@3 on Xbench-DeepSearch, 48.5% Avg@3 and 65.9% Pass@3 on BrowseComp-ZH, and 72.8% Avg@3 and 88.3% Pass@3 on GAIA, while reducing compression hallucinations to 5.3%. We open-source our codebase to facilitate future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0f21086-b1c2-4eb5-98fb-e1b69f38a2e3Builds on6
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 90 citations
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
Related papers
- Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement LearningZhuoen Chen, Dongfang Li, Meishan Zhang, Baotian Hu et al.ACL 2026 · 2 citations
- AgentFold: Long-Horizon Web Agents with Proactive Context FoldingRui Ye, Zhongwang Zhang, Kuan Li, Huifeng Yin et al.ICLR 2026 · 45 citations
- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving ContextGuangya Wan, Mingyang Ling, Xiaoqi Ren, Rujun Han et al.ACL 2026 · 11 citations
- Chain-of-Memory: Lightweight Memory Construction with Dynamic Evolution for LLM AgentsXiucheng Xu, Bingbing Xu, Tian Xueyun, Zihe Huang et al.ACL 2026 · 4 citations
- GAM: Hierarchical Graph-based Agentic Memory for LLM AgentsZhaofen Wu, Hanrong Zhang, Fulin Lin, Wujiang Xu et al.ACL 2026 · 8 citations
