Lune

ACL2025顶会

Adversarial Tokenization

Renato Lui Geh, Zilei Shao, Guy Van den Broeck

2025年份
4顶会引用

摘要

Current LLM pipelines account for only one possible tokenization for a given string, ignoring exponentially many alternative tokenizations during training and inference. For example, the standard Llama3 tokenization of penguin is [p,enguin], yet [peng,uin] is another perfectly valid alternative. In this paper, we show that despite LLMs being trained solely on one tokenization, they still retain semantic understanding of other tokenizations, raising questions about their implications in LLM safety. Put succinctly, we answer the following question: can we adversarially tokenize an obviously malicious string to evade safety and alignment restrictions? We show that not only is adversarial tokenization an effective yet previously neglected axis of attack, but it is also competitive against existing state-of-the-art adversarial approaches without changing the text of the harmful request. We empirically validate this exploit across three state-of-the-art LLMs and adversarial datasets, revealing a previously unknown vulnerability in subword models. github.com/RenatoGeh/advtok Algorithm 1 Compilation Input string x, upper bound k, reference v Output MRMDD M 0..k 1 Compile MDD M from x 2 Create k + 1 copies M 0 , M 1 , . . . , M k 3 for each edge e = (i, j) ∈ M do 4 if e |= v then 5 Mark edges M (i) l , M (j) l , ∀l ∈ [0..k] 6 else 7 Add edges M (i) l , M (j) l-1 , ∀l ∈ [1..k] 8 Prune all paths that are unmarked or do not end at a terminal node in M 0 9 return M 0..k := (M 0 , M 1 , . . . , M k )

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖