Adversarial Knowledge Extraction via Steering Diffusion Models

Conference Paper (2025)
Author(s)

Chi Hong (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Jiyue Huang (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Lydia Chen (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Robert Birke (University of Turin)

Research Group
Data-Intensive Systems
DOI related publication
https://doi.org/10.1007/978-981-96-7005-5_23 Final published version
More Info
expand_more
Publication Year
2025
Language
English
Research Group
Data-Intensive Systems
Pages (from-to)
336-350
Publisher
Springer
ISBN (print)
9789819670048
Event
31st International Conference on Neural Information Processing, ICONIP 2024 (2024-12-02 - 2024-12-06), Auckland, New Zealand
Downloads counter
50
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Model stealing allows to extract the knowledge of a deployed target machine learning model by sending query images and training a substitute model via the inference results. However, accessing the same data used to train the target model is often impractical. Recent data-free model stealing methods overcome this limit using specifically crafted noise as queries. However, two major flaws limit their effectiveness. First, data-free methods suffer from low query efficiency, requiring high query budgets, such as over 20 million queries to train a substitute for CIFAR-10. Second, crafted queries lack any perceptual semantics. Hence, they can easily be filtered out from legitimate requests. To address these issues, we propose AEDM, a framework for Adversarial knowledge Extraction via steering Diffusion Models. AEDM leverages publicly available pre-trained diffusion models to craft adversarial query images, controlled through the latent variable fed into the diffusion models, to maximize the knowledge transferred to the substitute model. These queries not only resemble in all aspects legitimate requests of images with a semantic meaning, but also significantly lower the number of required queries. Results on three datasets demonstrate that, given a budget as limited as 1/200 queries of the baselines, the accuracy of our trained substitute model can outperform that of state-of-the-art data-free stealing methods.

Files

978-981-96-7005-5_23.pdf
(pdf | 1.85 Mb)
- Embargo expired in 24-12-2025
– Personal use only – Dutch Copyright Act (Article 25fa)