Lune

CCS2026顶会

Selective Fine-Tuning for Targeted and Robust Concept Unlearning

Mansi Mansi, Avinash Kori, Francesca Toni, Soteris Demetriou

2026年份
2被引次数

摘要

Text-guided diffusion models are used by millions of users, but can be easily exploited to produce harmful content. Concept unlearning methods have emerged as a principled model-level defense, but state-of-the-art methods remain brittle: they are vulnerable to adversarial prompts, degrade generation quality for benign concepts, and lack interpretability. Concept localization approaches improve on concept unlearning through more targeted and interpretable selective fine-tuning, but current techniques are still limited: they are static, which results in suboptimal utility and weak robustness. In order to tackle these challenges, we propose TRuST (Targeted Robust Selective fine-Tuning), a novel approach for dynamically localizing target concept neurons and unlearning them demonstrated with two new optimization objectives and a selective fine-tuning method. We evaluate TRuST under both static and adaptive adversaries, and show experimentally, against a number of state-of-the-art baselines, that TRuST achieves up to 14× reduction in Attack Success Rate (ASR) across both adversary types, preserves generation quality (ΔFID = 0.016, CLIP = 30.95), and converges in up to 22× fewer fine-tuning steps than the state-of-the-art. Our method achieves robust and interpretable unlearning of not only individual concepts but also combinations of concepts and conditional concepts, without any task-specific regularization. Content Warning: This paper contains examples of unsafe and explicit visual content and language included solely for research and evaluation purposes.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper37

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖