Lune

ICML2026顶会

Persistent Semantic Entities in Tool-Augmented LLM Systems

Zhaohui Wang

2026年份

摘要

Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries-largely invisible to standard debugging. We formalize this as Persistent Semantic Entities (PSE): constructs defined by name binding, event triggering, and cross-boundary propagation, and evaluate them across 24 models from 11 families (1.5B-1T parameters). First, every tested model is susceptible (20-100% on the 20model susceptibility panel), with name binding as the necessary and dominant mechanism: without it, contamination is 0%. Second, persistence depends on contamination type rather than scale or deployment: preference contamination persists undecayed on every model probed (100% at t=10) and instruction contamination persists wherever adopted, persona-style injection decays partially (90%→10%), while factual injection is model-dependent-self-corrected on Llama-3.1-8B and GPT-4o-mini but held at ceiling on both Qwen2.5-coder variants, so we do not claim it selfcorrects in general. The preference and instruction results hold across providers in our controlled setting. Third, context-isolated self-verification achieves 20-79% reduction (median 36.5%) without oracle references while keyword-based detection produces systematic false positives, and contamination compounds 1.9× along a fourstage agent pipeline (40%→75%). Preference and instruction contamination-persistent, lacking self-correction, and poorly captured by standard monitoring-represent a particularly concerning attack surface for deployed agent systems.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖