Lune

NeurIPS2025顶会

Knowledge Distillation Detection for Open-weights Models

Qin Shi, Amber Yijia Zheng, Qifan Song, Raymond A. Yeh

2025年份
4被引次数

摘要

We propose the task of knowledge distillation detection, which aims to determine whether a student model has been distilled from a given teacher, under a practical setting where only the student's weights and the teacher's API are available. This problem is motivated by growing concerns about model provenance and unauthorized replication through distillation. To address this task, we introduce a model-agnostic framework that combines data-free input synthesis and statistical score computation for detecting distillation. Our approach is applicable to both classification and generative models. Experiments on diverse architectures for image classification and text-to-image generation show that our method improves detection accuracy over the strongest baselines by 59.6% on CIFAR-10, 71.2% on ImageNet, and 20.0% for text-to-image generation. The code is available at https://github.com/shqii1j/distillation_detection.

Prior works have considered related problems such as membership inference [5,22] and out-ofdistribution detection [29]. These works demonstrate that models may retain statistical traces of their training data or source model, even after distillation or fine-tuning. In particular, recent work on auditing data usage [22] and memorization in generative models [6,47] shows that it is possible to identify some aspects of model origin through probing-based methods. However, these approaches typically focus on detecting training data membership, rely on access to the teacher model, or assume architectural similarity. They do not directly tackle the problem of determining whether a model is distilled from another model that we are interested in.

To address knowledge distillation detection, we propose a model-agnostic approach that only requires access to the student model's weights, without requiring the teacher model's weights or training data. Our approach consists of three stages: input construction, score computation, and decision making. In the first stage, we generate synthetic queries that probe model behavior, using data-free synthesis techniques tailored to the model modality. In the second stage, we extract statistical scores from the model's responses, capturing either point-wise discrepancies or set-level distributional alignment. Finally, we apply a decision rule to determine whether the model has been distilled, based on these scores. This design is general and thus supports a wide range of models and tasks.

We conduct extensive experiments on both image classification and text-to-image generation tasks. For classification, we test a variety of architectures (ResNet18, DLA, DPN, DenseNet, ResNet50, and ResNeXt50) and distillation methods (KD, RKD, OFA-KD). For text-to-image generation, we report on student models distilled from SD-v2.1, SDXL, and PixArt, including BK-SDM, AMD, DMD2, and SDXL-Lightning. Across both domains, our method improves detection accuracy over the strongest baselines by 59.6% on CIFAR-10, 71.2% on ImageNet, and 20.0% for text-to-image generation.

Our contributions are summarized as follows:

• We introduce the task of knowledge distillation detection of open-weights student models, without access to the teacher model or training data.

• We propose a model-agnostic approach consisting of input construction, score computation, and decision rule stages, adaptable across both classification and generative settings.

• Empirical results on both classification and text-to-image tasks shows the effectiveness and generality of our method.

2 Related Work Knowledge distillation (KD) is an effective technique for compressing information from large pre-trained models into lightweight student models by aligning either the outputs or intermediate representations of the teacher and student [8,37,43,48,61,63,68]. The concept was first introduced by Buciluǎ et al. [3] and later popularized by Hinton et al. [20]. For classification tasks, knowledge distillation methods are commonly grouped into three categories: logit-based [20], feature-based [18, 43], and relation-based approaches [10, 35].

While classification-focused KD emphasizes improving accuracy or reducing model size, distillation for diffusion models primarily aims to accelerate generation [23, 28, 31, 32, 42, 44-46, 53, 56]. A typical approach is to train a student model to approximate the teacher's sampling trajectory, often formulated as an ODE, with significantly fewer inference steps. For instance, AMD [45] distills generative features from the teacher; BK-SDM [23] uses block pruning and feature distillation; and DMD [57] matches the output distribution directly via KL divergence without step-wise supervision. Unlike prior work that focuses on improving the performance of distillation, we study the inverse problem to detect whether one model is distilled from the other.

We review two different tasks that use sample synthesis to train a quantized/student model. One task is data-free model quantization [52], which targ

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext fdd53c61-5d55-42a6-9e29-04ce6c00a0d1

它引用的顶会 Paper34

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖