Loading...
HODA: Hardness-oriented detection of model extraction attacks
Sadeghzadeh, A. M ; Sharif University of Technology | 2024
90
Viewed
- Type of Document: Article
- DOI: 10.1109/TIFS.2023.3320609
- Publisher: 2024
- Abstract:
- Model extraction attacks exploit the target model's prediction API to create a surrogate model, allowing the adversary to steal or reconnoiter the functionality of the target model in the black-box setting. Several recent studies have shown that a data-limited adversaries with no or limited access to the samples from the target model's training data distribution, can employ synthesized or semantically similar samples to conduct model extraction attacks. In this paper, we introduce the concept of hardness degree to characterize sample difficulty based on the concept of learning speed. The hardness degree of a sample depends on the epoch number at which the predicted label for that sample converges. We investigate the hardness degree of samples and demonstrate that the hardness degree histogram of a data-limited adversary's sample sequence is differs significantly from that of benign users' sample sequences. We propose Hardness-Oriented Detection Approach (HODA) to detect the sample sequences of model extraction attacks. Our results indicate that HODA can effectively detect model extraction attack sequences with a high success rate, using only 100 monitored samples. It outperforms all previously proposed methods for model extraction detection. © 2005-2012 IEEE
- Keywords:
- Adversarial machine learning ; Hardness of samples ; Extraction ; Graphic methods ; Learning systems ; Closed box ; Computational modelling ; Histogram ; Machine-learning ; Model extraction ; Model stealing ; Predictive models ; Training data ; Hardness
- Source: IEEE Transactions on Information Forensics and Security ; Volume 19 , 2024 , Pages 1429-1439 ; 15566013 (ISSN)
- URL: https://ieeexplore.ieee.org/abstract/document/10266376?signout=success
