A Multi-Agent Framework for High-Interaction Terminal Simulation
Kai Wei, Yuwen Cui, Kehan Shen, Hua Wei, Guangjing Wang
摘要
Terminal simulation, framed as a terminal command-level Turing test, is a long-standing symbolic language generation problem in dialogue and interactive systems. Prior scripted simulators lack flexibility for complex, multiturn interactions, while LLM-based approaches often misinterpret commands, break output formats, drift from system state, and remain vulnerable to prompt injection. In this work, we propose MANTIS, a terminal simulation framework that improves realism, consistency, and robustness for command language generation. MANTIS integrates a multi-agent architecture with a filter-based routing model that safely dispatches commands to external tools or an LLM-based agent to support interactive commands and defend against prompt injection attacks. In addition, we design an agentic file system with history memory pruning for long-term state consistency. We release three datasets: 28,045 real terminal input-output pairs, a 1,000 multi-turn interaction session dataset, and a 25,849 labeled classification dataset. MANTIS outperforms stateof-the-art baselines by more than 9%, achieving over 95% accuracy on multi-turn terminal simulation. The dataset and source code are available at https://github.com/kaiwei666a/ MANTIS_Terminal_Simulation .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller 等ACL 2025 · 被引用 552 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Large Language Models for Code: Security Hardening and Adversarial TestingJingxuan He, Martin T. VechevCCS 2023 · 被引用 98 次
- Dynamic Context Pruning for Efficient and Interpretable Autoregressive TransformersSotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci 等NeurIPS 2023 · 被引用 95 次
相关 Paper
- ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn InteractionXingshan Zeng, Weiwen Liu, Lingzhi Wang, Liangyou Li 等ICLR 2026 · 被引用 15 次
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM AgentsHwan Chang, Yonghyun Jun, Hwanhee LeeICLR 2026 · 被引用 32 次
- NL Schedule: Evaluate Multitask Scheduling Capability of Large Language ModelsWenrui Liao, Weihong Du, Yi Li, Hongru Liang 等ACL 2026
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI AgentsKaijie Zhu, Xianjun Yang, Jindong Wang, Wenbo Guo 等ICML 2025
- AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse EnvironmentsZhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong 等ACL 2025 · 被引用 20 次
