Lune

NeurIPS2025Top-tier venue

Knowledge Distillation Detection for Open-weights Models

Qin Shi, Amber Yijia Zheng, Qifan Song, Raymond A. Yeh

2025Year
4Citations

Abstract

We propose the task of knowledge distillation detection, which aims to determine whether a student model has been distilled from a given teacher, under a practical setting where only the student's weights and the teacher's API are available. This problem is motivated by growing concerns about model provenance and unauthorized replication through distillation. To address this task, we introduce a model-agnostic framework that combines data-free input synthesis and statistical score computation for detecting distillation. Our approach is applicable to both classification and generative models. Experiments on diverse architectures for image classification and text-to-image generation show that our method improves detection accuracy over the strongest baselines by 59.6% on CIFAR-10, 71.2% on ImageNet, and 20.0% for text-to-image generation. The code is available at https://github.com/shqii1j/distillation_detection.

Prior works have considered related problems such as membership inference [5,22] and out-ofdistribution detection [29]. These works demonstrate that models may retain statistical traces of their training data or source model, even after distillation or fine-tuning. In particular, recent work on auditing data usage [22] and memorization in generative models [6,47] shows that it is possible to identify some aspects of model origin through probing-based methods. However, these approaches typically focus on detecting training data membership, rely on access to the teacher model, or assume architectural similarity. They do not directly tackle the problem of determining whether a model is distilled from another model that we are interested in.

To address knowledge distillation detection, we propose a model-agnostic approach that only requires access to the student model's weights, without requiring the teacher model's weights or training data. Our approach consists of three stages: input construction, score computation, and decision making. In the first stage, we generate synthetic queries that probe model behavior, using data-free synthesis techniques tailored to the model modality. In the second stage, we extract statistical scores from the model's responses, capturing either point-wise discrepancies or set-level distributional alignment. Finally, we apply a decision rule to determine whether the model has been distilled, based on these scores. This design is general and thus supports a wide range of models and tasks.

We conduct extensive experiments on both image classification and text-to-image generation tasks. For classification, we test a variety of architectures (ResNet18, DLA, DPN, DenseNet, ResNet50, and ResNeXt50) and distillation methods (KD, RKD, OFA-KD). For text-to-image generation, we report on student models distilled from SD-v2.1, SDXL, and PixArt, including BK-SDM, AMD, DMD2, and SDXL-Lightning. Across both domains, our method improves detection accuracy over the strongest baselines by 59.6% on CIFAR-10, 71.2% on ImageNet, and 20.0% for text-to-image generation.

Our contributions are summarized as follows:

• We introduce the task of knowledge distillation detection of open-weights student models, without access to the teacher model or training data.

• We propose a model-agnostic approach consisting of input construction, score computation, and decision rule stages, adaptable across both classification and generative settings.

• Empirical results on both classification and text-to-image tasks shows the effectiveness and generality of our method.

2 Related Work Knowledge distillation (KD) is an effective technique for compressing information from large pre-trained models into lightweight student models by aligning either the outputs or intermediate representations of the teacher and student [8,37,43,48,61,63,68]. The concept was first introduced by Buciluǎ et al. [3] and later popularized by Hinton et al. [20]. For classification tasks, knowledge distillation methods are commonly grouped into three categories: logit-based [20], feature-based [18, 43], and relation-based approaches [10, 35].

While classification-focused KD emphasizes improving accuracy or reducing model size, distillation for diffusion models primarily aims to accelerate generation [23, 28, 31, 32, 42, 44-46, 53, 56]. A typical approach is to train a student model to approximate the teacher's sampling trajectory, often formulated as an ODE, with significantly fewer inference steps. For instance, AMD [45] distills generative features from the teacher; BK-SDM [23] uses block pruning and feature distillation; and DMD [57] matches the output distribution directly via KL divergence without step-wise supervision. Unlike prior work that focuses on improving the performance of distillation, we study the inverse problem to detect whether one model is distilled from the other.

We review two different tasks that use sample synthesis to train a quantized/student model. One task is data-free model quantization [52], which targ

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext fdd53c61-5d55-42a6-9e29-04ce6c00a0d1

Builds on34

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines