Analyzing and Internalizing Complex Policy Documents for LLM Agents
Jiateng Liu, Zhenhailong Wang, Xiaojiang Huang, Yingjie Li, Xiang Li, Chenlei Guo, Xing Fan, Ruhi Sarikaya, Heng Ji
摘要
Large Language Model (LLM) based agentic systems rely heavily on in-context policy documents that encode diverse business rules. As business requirements expand, these documents grow substantially, creating significant computational overhead. This motivates the need for internalization methods that embed policy documents into model priors while preserving performance. While prior prompt compression research primarily targets generic prompts, we find that agentic policy documents span multiple levels of complexity and demand more intensive reasoning, presenting greater internalization challenges. We first introduce CC-Gen, an agentic benchmark generator with Controllable Complexity defined across four levels, enabling systematic benchmarking of how well agents handle complexities and provides a framework for comprehensive evaluation of policy internalization algorithms. Our initial analysis reveals that complex policy specifications governing agent workflows may pose the most significant reasoning challenges. When supporting internalization with gold user-agent interaction trajectories containing chain-of-thought (CoT) annotations through supervised finetuning (SFT), we find that this baseline is highly data-intensive and its effectiveness deteriorates markedly as policy document complexity increases. To mitigate data burden and reasoning challenges, we propose Category-Aware Policy Continued Pretraining(CAP-CPT). Our automated pipeline analyzes policy documents to extract key specifications, grouping them into factual, behavioral, and conditional types. We further isolate complex conditions, which introduce high workflow complexity and drive core reasoning difficulty. This categorization guides a targeted therapy, synthesizing specialized training data for each specification type and enabling agents to internalize policy information more effectively through an autoregressive pretraining loss. Our extensive experiments demonstrate the effectiveness of the curated data and training objective. Combined with SFT, our approach improves baseline across all data scenarios. It is especially effective in data-sparse settings and under high policy complexity, yielding gains of up to 41% and 22% on Qwen-3-32B. Overall, we achieve up to 97.3% prompt length reduction in our benchmark. Applied to τ -Bench, our approach further improves performance and reduces input length with very limited SFT data. 1 * Work done during an internship at Alexa AI 1 All data and code will be publicly released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 被引用 488 次
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen 等ICLR 2024 · 被引用 308 次
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng 等ICLR 2024 · 被引用 299 次
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less ForgettingSanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che 等EMNLP 2020 · 被引用 152 次
相关 Paper
- Multimodal Policy Internalization for Conversational AgentsZhenhailong Wang, Jiateng Liu, Amin Fazel, Ritesh Sarkhel 等ICLR 2026 · 被引用 2 次
- AutoTool: Dynamic Tool Selection and Integration for Agentic ReasoningJiaru Zou, Ling Yang, Yunzhe Qi, Sirui Chen 等ICML 2026 · 被引用 4 次
- Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought CompressionYuntian Tang, Bohan Jia, Wenxuan Huang, Lianyue Zhang 等ICML 2026 · 被引用 5 次
- Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step EntropyZeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen 等ICLR 2026 · 被引用 35 次
- Scaling Agents via Continual Pre-trainingLiangcai Su, Zhen Zhang, Guangyu Li, Zhuo Chen 等ICLR 2026 · 被引用 46 次
