Leveraging Imperfect Restoration for Data Availability Attack
Abstract
The abundance of online data is at risk of unauthorized usage in training deep learning models. To counter this, various Data Availability Attacks (DAAs) have been devised to make data unlearnable for such models by subtly perturbing the training data. However, existing attacks often excel against either Supervised Learning (SL) or Self-Supervised Learning (SSL) scenarios. Among these, a model-free approach that generates a Convolution-based Unlearnable Dataset (CUDA) stands out as the most robust DAA across both SSL and SL. Nonetheless, CUDA’s effectiveness against SSL is underwhelming and it faces a severe trade-off between image quality and its poisoning effect. In this paper, we conduct a theoretical analysis of CUDA, uncovering the sub-optimal gradients it introduces and elucidating the strategy it employs to induce class-wise bias for data poisoning. Building on this, we propose a novel poisoning method named Imperfect Restoration Poisoning (IRP), aiming to preserve high image quality while achieving strong poisoning effects. Through extensive comparisons of IRP with eight baselines across SL and SSL, coupled with evaluations alongside five representative defense methods, we showcase the superiority of IRP. Code: https://github.com/lyumingzhi/IRP
Keywords:
Data Availability Attacks Supervised Learning Self-Supervised Learning1 Introduction
The proliferation of online data has proven an indispensable resource for the advancement of deep learning models. However, the collection of certain datasets without explicit consent poses a possible threat to personal privacy [2]. Additionally, the emergence of generative AI presents new challenges, as training with unauthorized data could infringe upon the owners’ copyright [34].
DAAs have emerged as a promising strategy to address the issue of unauthorized data usage[12, 42, 20, 13, 33, 29, 41, 40, 17, 27, 14]. These attacks introduce subtle perturbations into training data, hindering the model’s ability to effectively learn useful information, which subsequently leads to poor performance on unseen data. This poor performance manifests as a substantial decrease in accuracy when the model is tested on clean data in classification tasks. Currently, the majority of DAAs are developed within the framework of SL. However, a recent study by He et al. [17] reveals that Adversarial Poison (AP) [13], a cutting-edge DAA targeting SL, lacks effectiveness when applied to SSL. In response, they introduce a novel approach, termed Contrastive Poison (CP) [17], to counter SSL. Their research leads to several questions: (1) Are other DAAs designed for SL also ineffective against SSL (2) Does the proposed CP method maintain similar effectiveness in SL scenarios? (3) Can we develop a DAA capable of conducting effective attacks in both SL and SSL scenarios?


To answer these questions, we experimentally evaluate seven representative DAAs designed for SL on SSL. SSL is a strong defense against most DAAs due to its augmentation invariance objective; DAA features that are distorted or obfuscated by augmentations are ignored by the SSL algorithm and cannot affect downstream tasks. Accordingly, we find in Fig. 2 that most poisons for CIFAR-10 [23] fail to generalize to SSL. Only CUDA [27] and CP [17] exhibit significant impact on SSL. CUDA filters inject poison features throughout images that are resistant to cropping, flipping, and color shifts, and CP is designed specifically to counter SSL. We note a significant loss of performance for CP in SL scenarios. Therefore, CUDA emerges as the most potent DAA in both SSL and SL. However, the clean test accuracy of SSL trained on CUDA still hovers around 70%, suggesting that models trained on poisoned data retain some usability. As suggested by Sadasivan et al. [27], increasing the size of the convolution kernels in CUDA could provide a stronger poisoning effect, but it also degrades image quality. Fig. 2 shows that increased kernel size results in obviously blurred images. This compromise is often unacceptable in practical scenarios, particularly for sharing images and artwork online.
To comprehend CUDA’s efficacy, Sadasivan et al. provide a theoretical analysis under assumptions of Gaussian-distributed data with independent elements and a two-class scenario. However, these assumptions may not hold in real-world scenarios. In addition, their analysis is not based on deep learning, making it challenging to derive insights for designing a more effective poison method. To gain a deeper understanding of CUDA’s impact on model training, we first conduct a theoretical analysis of CUDA using a deep learning model. Our analysis reveals how CUDA generates sub-optimal gradients for clean data and introduces class-wise bias through random filters. Building upon this analysis, we propose a new poisoning method, named Imperfect Restoration Poisoning (IRP), which maintains high image quality while achieving a stronger poisoning effect. We compare IRP to eight representative DAAs in both SL and SSL scenarios. In the SL scenario, we additionally compare IRP with other DAAs under various defense mechanisms, including adversarial training [33], Image Shortcut Squeezing [24], Mixup [44], Cutout [9], and CutMix [43]. All experimental results indicate that IRP significantly outperforms previous DAAs. Our main contributions are:
- •
We experimentally assess eight representative DAAs in both supervised and self-supervised learning scenarios, revealing that existing DAAs fail to achieve satisfactory effectiveness simultaneously in both contexts.
- •
We conduct a theoretical analysis of CUDA on deep learning models, revealing how it generates sub-optimal gradients for clean data and introduces class-wise bias through random filters.
- •
We propose a novel DAA named IRP based on imperfect restoration, which achieves high effectiveness in both SL and SSL and maintains superior image quality compared to CUDA.
- •
We perform an exhaustive comparison between IRP and eight other DAAs across SL, SSL, and five defense techniques on CIFAR-10 and a subset of ImageNet [8]. IRP outperforms baseline poisons by a large margin in all cases. We further verify IRP across five additional architectures and two supplementary datasets. IRP displays high effectiveness across all experiments.
2 Related Works
DAAs have garnered significant attention in recent years for their application to data protection. Various methodologies have been proposed to deliberately induce malfunctions in the target model in order to safeguard against unauthorized data usage. DAAs can be broadly categorized into two types, model-reliant methods [12, 42, 20, 13, 14, 17] and model-free methods [29, 40, 41, 27].
Model-Reliant Methods: Early approaches to DAAs often frame the poisoning problem as a bilevel optimization task, balancing loss minimization with respect to model parameters with loss maximization with respect to perturbed inputs. While these methods [22, 1] initially showed promise, their applicability to deep neural networks is limited due to intractability in obtaining exact solutions. Efforts to address optimization challenges in perturbation generation, as seen in works by Feng et al. [12] and Yuan et al. [42], often introduce significant computational overhead, limiting scalability for real-world applications. Alternatively, Huang et al.[20] propose another bi-level Error Minimizing (EM) approach, deliberately optimizing perturbations to minimize training loss. By alternately training a surrogate model and optimizing perturbations, the efficacy of the perturbations is attained through multiple rounds of optimization. In contrast to traditional bi-level approaches, Fowl et al. [13] demonstrate the effectiveness of utilizing common objectives from adversarial examples for generating potent Adversarial Poisons (AP). However, Tao et al. [33] pinpoint that these DAAs can be easily broken by adversarial training. In response, Fu et al. [14] introduce Robust Error-Minimizing (REM) which enhances the poisoning effects against adversarial training by replacing the surrogate in EM with an adversarially-trained model. Recently, He et al. [17] assert that all aforementioned methods are tailored for supervised scenarios, thereby making them unsuitable for self-supervised learning. To address this gap, they introduce Contrastive Poison (CP), specifically designed for self-supervised learning. However, their assessment of DAAs tailored for SL is exclusively centered on AP, with other DAAs left unexplored. Additionally, while they show some effectiveness of their class-wise CP on both SL and SSL, the test accuracies on SSL remain high, ranging from 60.7% to 68.0%. In addition, class-wise additive perturbations can be recovered by the average image of a class, making them relatively easy to remove [29]. Therefore, we do not consider the class-wise poisons that apply the same additive noise to images, including class-wise attacks by EM, AP, and CP.
Model-Free Methods: Model-free DAAs tackle the data protection issue by leveraging shortcut learning. These methods are grounded in the observation that Deep Neural Networks (DNNs) often prioritize easily learnable shortcuts over semantic features, allowing them to distinguish between examples from different classes [3, 21]. Expanding on this concept, Yu et al. [41] propose incorporating artificially generated Linearly Separable Patterns (LSP) sampled from high-dimensional Gaussian distributions into images as an effective method for data poisoning. Sandoval-Segura et al. [29] discover that applying additive perturbations derived from autoregressive (AR) processes to clean data can serve as an effective shortcut. Wu et al. [40] examine the model’s susceptibility to sparse poisons and demonstrate that consistently perturbing only one pixel is adequate to generate potent poisons. In contrast to applying additive perturbations, Sadasivan et al. [27] introduce a Convolution-based Unlearnable Dataset (CUDA) generation method, which utilizes randomly generated class-wise convolutional filters. In our evaluation, CUDA demonstrates moderate robustness to SSL. However, its effectiveness on SSL remains insufficient and CUDA grapples with a severe trade-off between image quality and poison efficacy.
DAA Mitigation Strategies: Previous studies have demonstrated that Adversarial Training (AT) [15] is an effective method to mitigate the potency of DAAs [33, 39]. However, AT is impeded by its significant computational demands. Alternatively, various image preprocessing techniques, such as data augmentation (e.g., Mixup [44], Cutout [9], and CutMix [43]), have been explored. While these augmentation-based techniques show some impact, they are not as effective as AT [13]. Recently, Liu et al. [24] introduce Image Shortcut Squeezing (ISS) [24], a compression-based approach to counter DAAs. They demonstrate that ISS surpasses previously studied data augmentation countermeasures and achieves results comparable to AT. We evaluate the proposed IRP against multiple countermeasures, including AT, ISS, Mixup, Cutout, and CutMix.
3 Analysis
3.1 Threat Model
Like CUDA, most DAA papers define an “attacker” that applies a DAA on a labeled dataset that they want to protect yet publish. The attacker assumes no knowledge of any downstream DNN training process. It is therefore crucial to demonstrate that DAAs are effective across multiple training methods, including SL, SSL, and AT. The attacker generates class-based perturbations to their image set in order to inject poison features. The attacker publishes their poisoned images, optionally with captions or labels. Model trainers collect these poisoned images into datasets and may further alter or label the images.
3.2 Notation and Preliminaries
For clarity, we define a set of notations. represents a Convolution Neural Network (CNN), including its training loss function but excluding the first layer convolution 2D filter . denotes an input image and represents an image belonging to class , where is total number of the classes. Given that cross-correlation is predominantly used in place of convolutions in practice, the first layer operation is denoted as , where signifies cross-correlation. Note that is a part of the CNN and is the loss value of . To retain readability, we focus on gradient analysis of the first layer filter, . CUDA employs class-specific convolutional filters for data poisoning, and we provide a brief overview of its operation. Let be a CUDA filter for class images. According to [27], the generation procedure for involves setting one random parameter among the parameters to 1 within each . Concurrently, the remaining parameters are initialized randomly, drawn from a uniform distribution with support . An image belonging to class is poisoned by cross-correlation with the corresponding CUDA filter , denoted as .
3.3 The Non-Optimal Gradient
To understand how CUDA lowers the accuracy of clean data, we analyze the gradient of for both clean and CUDA-poisoned data. Without loss of generality, the training set consists of clean images, each belonging to a different class, with corresponding poisoned images denoted as , . The gradient of the clean training data is
| (1) |
where is an intermediate variable equal to . The gradient of the poisoned training data is
| (2) |
where the subscript indicates that the intermediate variable is from poisoned data. Using these gradients to update , we obtain
| (3) |
where is learning rate. Inputting all clean training images, to and , we obtain clean and poison loss values and , respectively. Given that in Eq. 1 is the optimal gradient descent direction for the clean data, any deviation from this direction will result in increased loss when is small. Thus, for clean data. In other words, CUDA employs a non-optimal gradient to disrupt training, rendering clean data unrecognizable.
3.4 Poisoning through Class-Wise Bias
Though the previous subsection uncovers how CUDA filters increase training loss, it does not indicate how these filters poison the network. To further investigate its mechanism, we input to , whose first layer output is
| (4) |
Assuming is not poisoned by CUDA filters, we ignore the term . Since cross-correlation is equivalent to a rotated convolution (i.e., , where represents convolution and represents rotating the 2D filter by ), the second term of Eq. 4 can be rewritten as
| (5) |
Eq. 5 shows that the output of the first layer is influenced by . Since , we analyze . Theorem 1 describes the statistical probabilities of and . The proof is provided in Appendix 0.A.1.
Theorem 3.1
Given two different filters, and , generated by CUDA, if , then and have the following properties:
- 1.
The peak of is located at the center.
- 2.
The peak of occurs at the position where two ones in and appear at the same position in .
- 3.
The probability that the peak of lies in the center is .
- 4.
The expected peak of is higher than the expected peak of .
Figure 3 presents two CUDA filters along with the corresponding and to illustrate these statistical properties. In practice, CUDA utilizes and for ImageNet, which fulfills . Similar properties also apply on CIFAR-10 (see Appendix 0.A.3). These properties indicate that in Eq. 5 retains more information in but shifts the information in to a different position with a probability of . For , the probability is . Consequently, the upper layers receive class information in the correct position, but information from other classes is shifted to incorrect positions. For , we note the peak value of is smaller than the sum of its non-peak values (estimated mean ratio of 0.099). This means that non-peak values also play an important role in degrading other class information. This blurring effect also affect the -class, though to a lesser extent. More clearly, the estimated mean ratio of the peak value of over the sum of its non-peak values is 0.106. Combining these two effects, CUDA filters utilize peak shifts and unique blur patterns to introduce class-wise bias in the poisoned class.
3.5 Poisoning with Imperfect Restoration
The previous analysis shows that CUDA filters create class-wise bias, and that retains the training information of class . When is closer to , it is expected to preserve the class information better and enhance the image quality. However, if , then there is no poisoning effect. Thus, we propose the Imperfect Restoration Poison (IRP). Let be one of the class images, where . IRP first applies CUDA filter on all class training images and obtains . Then, the filtered images are cropped into patches with the size same as the CUDA filter i.e., , denoted as . The pixel value of corresponding to the center of the patch is denoted , and the patch is reshaped into a column vector denoted as . Finally, a class-wise linear model and the mean square loss are used to estimate the value of . More precisely, the class-wise model is obtained by minimizing
| (6) |
Once is obtained, it is reshaped back to a filter, denoted as . Finally, IRP uses to perform poisoning, where . Theorem 0.A.2 establishes that IRP is incapable of achieving perfect restoration (proof in Appendix 0.A.2).
Theorem 3.2
Given a CUDA filter size larger than 1 and as the vector form of the patch in involved in producing (i.e., , where is a matrix), if the column vectors of , excluding the one corresponding to , span , then, such that .
Fig. 4 shows two IRP filters with the corresponding self-correlation and cross-correlation. Similar to CUDA, we observe that exhibits a peak at the center, while has a peak in a different position. However, they have more complex patterns around the peaks. Therefore, like CUDA, IRP exploits peak shifts to create class-wise bias. However, IRP filters have turbulent patterns around the peaks, creating unique high-frequency attack patterns that are more complex than the low-frequency patterns from CUDA. Due to their dependence on the training data, it would be challenging to theoretically analyze IRP filters, as was done for CUDA, without proper assumptions on the data distribution. We refrain from imposing unrealistic assumptions on the data distribution (e.g., Gaussian) as this could lead to misleading analytic results.
Fig. 5 displays some images generated by IRP and CUDA. IRP images are visually clearer and maintain higher quality. Notably, the man’s face is sharper and the spots of the salamander and stingray are easily distinguishable under IRP protection. The backgrounds of all three images are sharper under IRP.
4 Experiments
In this section, we begin by evaluating the effectiveness of IRP alongside eight DAA methods in standard SL and SSL settings, as well as against representative defense methods. We then investigate the effectiveness of different DAAs under partial poisoning scenarios. For model-reliant methods, our experiments primarily focus on the nominal setup, wherein a surrogate model generates poison images. To assess the transferability of IRP, we evaluate IRP across six different architectures on three datasets. Finally, we compare the image quality generated by CUDA and IRP and ablate different filter types across various training methods.
4.1 Experimental Setup
Datasets: We examine poisoning on CIFAR-10 [23], CIFAR-100 [23], STL-10 [7], and ImageNet-100, a subset of ImageNet [8] consisting of 100 classes. By default, we use 50,000 images for training and 10,000 images for testing on CIFAR-10 and CIFAR-100. For STL-10, we use their default 5000 training images and 8000 test images for training and testing, respectively. For ImageNet-100, we train on the images from the first 100 classes of the official training set and test on all corresponding images from the official validation set.
Poison Settings: We compare IRP with eight state-of-the-art DAAs, including four model-reliant methods, including EM [20], TAP [13], REM [14], and CP [17], and four model-free methods, including CUDA [27], LSP [41], OPS [40], and AR [29]. As generating poison data using model-reliant methods is time-consuming on ImageNet, we generate a 20% subset of ImageNet-100 for CP, TAP, and EM. We maintain all default settings and retain the norm of the defensive perturbation at 8/255. We utilize the full REM-poisoned ImageNet-100 dataset shared by the REM authors. According to [14], the norm of the defensive perturbation is set to , and the adversarial perturbation radius is set to for controlling the level of protection of the noise against adversarial training. For model-free methods, we maintain the settings as per their original papers, except for LSP and OPS on ImageNet-100. As the original settings could not yield a satisfactory poisoning effect, we amplify the perturbation of LSP and OPS to and , respectively. As AR does not report its effectiveness on ImageNet-100 and has shown to be vulnerable to SSL and AT on CIFAR-10, we opt not to include AR in further testing on ImageNet-100.
Models: Unless stated otherwise, we employ ResNet-18 [18] as the surrogate for model-reliant methods and as the target models for all the testing methods in both SL and SSL scenarios. Various SSL algorithms are employed, including SimCLR [4], SimSiam[5], MoCoV3[6] and BYOL[16]. Linear probing accuracy, consistent with CP [17], is utilized to evaluate the effectiveness of DAAs in SSL. To assess transferability, we evaluate IRP on target models with various architectures, including ResNet-18, ResNet-34 [18], VGG-19 [32], DenseNet-121 [19], MobileNet-V2 [28], and ViT [11]. Training details are deferred to Appendix 0.B.1.
4.2 Comparisons to Baseline DAAs
| DAA | SL | SSL | AT | Cutout | Cutmix | Mixup | ISS | Max |
|---|---|---|---|---|---|---|---|---|
| Clean | 93.86 | 90.55 | 89.57 | 95.62 | 95.78 | 95.46 | 82.40 | 95.62 |
| EM | 32.53 | 88.45 | 89.31 | 36.18 | 38.40 | 54.19 | 82.12 | 89.31 |
| REM | 18.91 | 86.70 | 34.45 | 19.67 | 27.46 | 23.27 | 80.04 | 86.70 |
| AP | 12.45 | 75.77 | 85.67 | 9.70 | 8.38 | 33.27 | 81.10 | 85.67 |
| CP | 93.59 | 73.08 | 88.43 | 94.51 | 93.77 | 93.36 | 82.21 | 94.51 |
| LSP | 21.47 | 88.87 | 85.67 | 18.47 | 21.68 | 22.76 | 79.33 | 88.87 |
| OPS | 27.24 | 87.79 | 20.10 | 65.12 | 85.27 | 37.56 | 72.14 | 87.79 |
| AR | 10.95 | 86.86 | 72.94 | 12.17 | 12.95 | 14.15 | 82.83 | 86.86 |
| CUDA | 21.89 | 66.58 | 48.58 | 23.46 | 24.04 | 21.75 | 21.94 | 66.58 |
| IRP | 10.39 | 43.24 | 32.21 | 10.85 | 15.21 | 16.03 | 29.42 | 43.24 |
Effectiveness on SL and SSL: We first evaluate IRP and the other DAAs on standard SL and SSL training scenarios on CIFAR-10 and ImageNet-100. SimCLR is utilized for model pretraining in the SSL evaluation. Columns 2-3 of Tables 1 and 2 show that IRP achieves state-of-the-art poisoning, enforcing the lowest clean test accuracies in both SL and SSL and for both CIFAR-10 and ImageNet-100.
We begin by analyzing results for CIFAR-10 in Table 1. For SL with CIFAR-10, AR and IRP demonstrate a similar level of poisoning effect under SL, reducing the test accuracy to around 10%. However, AR underperforms on SSL, permitting a significantly higher accuracy of 86.86% compared to that of IRP at 43.24%. CUDA is the second-best DAA for SSL with an accuracy of 66.58%, which is still drastically higher than that of IRP.
For ImageNet-100 results in Table 2, we begin by nothing the drop in effectiveness of OPS relative to CIFAR-10. This is attributed to the RandomResizeCrop augmentation utilized for training models on ImageNet-100. We chose to use this augmentation on ImageNet-100 as it is commonly used for SL and SSL on ImageNet-100 in literature and it enhances the effectiveness of all the defense methods. For SL with ImageNet-100, CUDA, EM, AP, and LSP demonstrate strong poisoning capabilities, reducing the clean test accuracy to below 10%. Even so, IRP outperforms all others with a test accuracy of 1.98%. For SSL with ImageNet-100, the test accuracies of most DAAs remain high, except for CUDA at 21.89% and CP at 18.76%. IRP outperforms both CUDA and CP, achieving a clean test accuracy of 9.30%. These results establish IRP as the most effective DAA under both SL and SSL scenarios.
| DAA | SL | SSL | AT | Cutout | Cutmix | Mixup | ISS | Max |
|---|---|---|---|---|---|---|---|---|
| Clean | 77.66 | 70.96 | 69.76 | 78.04 | 81.08 | 80.38 | 71.58 | 80.38 |
| EM | 6.32 | 50.02 | 43.26 | 6.42 | 5.80 | 12.38 | 41.18 | 50.02 |
| REM | 14.18 | 65.30 | 57.34 | 15.60 | 16.10 | 33.08 | 67.52 | 67.52 |
| AP | 8.18 | 41.80 | 42.88 | 7.68 | 8.38 | 9.68 | 24.22 | 42.88 |
| CP | 57.46 | 18.76 | 48.76 | 58.20 | 62.28 | 61.76 | 50.18 | 62.28 |
| LSP | 6.62 | 62.06 | 28.92 | 5.06 | 7.84 | 5.50 | 33.70 | 62.06 |
| OPS | 51.00 | 62.22 | 52.56 | 53.86 | 65.12 | 48.82 | 57.36 | 65.12 |
| CUDA | 6.10 | 26.12 | 36.34 | 8.66 | 5.70 | 5.16 | 3.58 | 36.34 |
| IRP | 1.98 | 9.30 | 22.10 | 2.18 | 1.52 | 2.02 | 4.14 | 22.10 |
Effectiveness under Common Countermeasures: To further evaluate the effectiveness of IRP under defensive scenarios, we test IRP and the baselines on representative defense methods, including Adversarial Training (AT) [15], Image Shortcut Squeezing (ISS) [24], and common data augmentation techniques such as Cutout, CutMix, and Mixup. Following [14], we employ the norm for AT and set the norm bound to 4/255. 10-step PGD with a step size of 0.6/255 is used for AT. As suggested by [24], we employ a combination of grayscale and JPEG with the default JPEG quality factor (JPEG-10) for ISS to ensure global effectiveness against all DAAs. The results on CIFAR-10 and ImageNet-100 are presented in columns 4-8 of Tables 1 and 2, respectively. The last column displays the maximum clean test accuracy of the corresponding methods across all testing scenarios, representing the worst performance among all the test cases.
DAA results on CIFAR-10 in Table 1 vary across defenses. For instance, OPS shows high effectiveness on AT but is susceptible to ISS and CutMix, while AP performs well against CutMix and Cutout but is less effective against AT and ISS. Only CUDA and IRP demonstrate effectiveness across all countermeasures. However, CUDA permits maximum clean accuracies of 48.58% across all countermeasures and 66.58% across all testing cases, which are 16.37% and 23.24% higher, respectively, than the 43.24% maximum accuracy allowed by IRP.
Similarly, CUDA and IRP emerge as the two most effective methods across all defenses on ImageNet-100 in Table 2. However, CUDA permits a maximum clean accuracy of 36.34%, significantly higher than that of IRP at 22.10%. The results demonstrate that IRP is the most effective DAA in all testing scenarios.
Evaluation on Other SSL Algorithms: To further assess the efficacy of IRP on SSL, we test IRP-poisoned ImageNet-100 on three additional SSL algorithms: SimSiam[5], MoCoV3[6], and BYOL[16]. As shown in Table 3, IRP lowers the test accuracies of all SSL algorithms to around 10%. Even in its the worst case performance with BYOL, IRP reduces clean accuracy to 11.6%.
| SimCLR | MoCoV3 | SimSiam | BYOL | Max | |
|---|---|---|---|---|---|
| Clean | 70.96 | 75.54 | 70.68 | 75.34 | 75.54 |
| IRP | 9.3 | 10.58 | 7.82 | 11.6 | 11.6 |
4.3 Additional Training Scenarios and Datasets
Different Poison Percentages: In this section, we first evaluate IRP on a more challenging and realistic learning scenario, where only a portion of the data is shielded by IRP. We randomly select a portion of the training data from the entire training dataset of CIFAR-10 and apply IRP to the selected subset. The poisoned data is denoted as and the clean data in the training dataset that has not been used for generating poison data is denoted as . Then, we conduct standard supervised training with ResNet-18 on the mixed data () consisting of the and , as well as on the clean data subset . The results in Table 4 show that the effectiveness drops quickly when the data are not 100% poisoned. This phenomenon applies for all other DAAs as well. Models trained only on demonstrate a similar level of performance as the models trained with . This suggests that the high performance for <100% poisoning is not due to a failure of the DAAs. Finally, works on DAAs for artwork protection [30, 31] have noted similar trends but find that individuals can still protect their own data via DAA even if the majority of artworks are unprotected. That is, it’s still possible to make subsets of the training data unlearnable.
| 0% | 20% | 40% | 60% | 80% | 100% | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| EM | 93.86 | 94.02 | 92.38 | 92.72 | 91.74 | 90.71 | 87.84 | 85.75 | 77.36 | 32.53 |
| REM | 92.85 | 92.15 | 90.05 | 86.29 | 18.91 | |||||
| AP | 92.57 | 91.16 | 90.72 | 88.11 | 12.45 | |||||
| LSP | 92.95 | 92.40 | 90.19 | 85.76 | 21.47 | |||||
| OPS | 93.40 | 92.25 | 90.63 | 86.68 | 27.24 | |||||
| AR | 92.98 | 92.66 | 90.94 | 87.93 | 10.95 | |||||
| CUDA | 93.27 | 92.27 | 90.51 | 86.90 | 21.89 | |||||
| IRP | 93.07 | 92.56 | 89.89 | 85.70 | 10.39 | |||||
| Dataset | RN18 | RN34 | VGG19 | DN121 | MNv2 | ViT-S | Max | |
|---|---|---|---|---|---|---|---|---|
| CIFAR10 | Clean | 89.57 | 89.32 | 78.57 | 79.97 | 75.94 | 47.84 | 89.57 |
| CUDA | 48.58 | 43.63 | 42.02 | 49.44 | 26.33 | 44.93 | 49.44 | |
| IRP | 32.21 | 17.37 | 19.06 | 17.87 | 14.13 | 35.39 | 35.39 | |
| CIFAR100 | Clean | 65.30 | 66.43 | 48.50 | 58.92 | 47.00 | 27.19 | 66.43 |
| CUDA | 37.21 | 37.95 | 27.44 | 23.07 | 16.00 | 24.88 | 37.95 | |
| IRP | 27.13 | 28.85 | 22.50 | 25.24 | 20.70 | 20.82 | 28.85 | |
| STL10 | Clean | 71.94 | 69.73 | 64.35 | 72.41 | 58.93 | 43.86 | 72.41 |
| CUDA | 44.70 | 44.35 | 45.99 | 39.39 | 44.49 | 42.14 | 45.99 | |
| IRP | 34.10 | 34.33 | 33.89 | 33.19 | 33.81 | 35.58 | 35.58 |
Effectiveness on Different Models: The previous evaluations are all conducted on ResNet-18 (RN18). To evaluate the effectiveness of IRP on different architectures, we evaluate ResNet-34 (RN34), VGG19, DenseNet121 (DN121), MobileNetv2 (MNv2), and ViT-S on IRP-poisoned CIFAR-10. All models are trained with AT ( constraint of , as before), which emerged as the strongest defense against different DAAs from our experiments in Section 4.2. We include CUDA in this evaluation as a reference. On all the examined networks, IRP outperforms CUDA by a significant margin. In their respective worst cases, CUDA permits a test accuracy of 49.44%, while IRP allows a test accuracy of 35.39%. Table 5 shows that even under AT defense, IRP is still effective on all the examined architectures.
Evaluation on Different Datasets: We extend our evaluation to the CIFAR-100 and STL-10 datasets. The same set of models and AT settings are utilized. As shown in Table 5, both IRP and CUDA demonstrate effectiveness on CIFAR-100 and STL-10 across all the network architectures under AT. IRP consistently outperforms CUDA by a large margin on most networks for both CIFAR-100 and STL-10. In the two cases where IRP is weaker than CUDA, the clean test accuracies of IRP are still within 5% of those of CUDA. The maximum clean test accuracies of IRP on CIFAR-100 and STL-10 are 28.85% and 35.58%, respectively. Both of these values are significantly lower than those of CUDA, which are 37.95% and 45.99% on CIFAR-100 and STL-10, respectively. This experiment demonstrates that IRP is universally effective across networks and across different datasets.
4.4 Image Quality Comparison against CUDA
We compare the quality of ImageNet-100 images poisoned by IRP and CUDA. We quantify image quality with three common full-reference quality indices, including LPIPS [45], SSIM [37], and MS-SSIM [38], and two no-reference quality indices, including CLIP-IQA [35] and BRISQUE [25]. The original image is used as the reference for LPIPS, SSIM, and MS-SSIM. Table 6 presents the evaluation results of all quality indices. Both Table 6 and Fig. 5 demonstrate that IRP generates poisoned images with better image quality than CUDA.
| Metric | Clean | CUDA | IRP |
|---|---|---|---|
| LPIPS | - | 0.272 | 0.142 |
| SSIM | - | 0.561 | 0.775 |
| MS-SSIM | - | 0.831 | 0.903 |
| CLIP-IQA | 0.787 | 0.528 | 0.706 |
| BRISQUE | 16.625 | 29.570 | 17.318 |
4.5 Filter Ablation
The IRP filter is a kernel with a larger size of , where and are both filters of size . To verify that the high effectiveness of IRP is not caused by the bigger kernel size, we generate filters of the same size as the IRP under two different filter settings. For ‘Random Sharpness’, we generate a random sharpness filter of size for each class. Within each , one random parameter out of the parameters is set to 1, while the remaining parameters are randomly initialized from a uniform distribution with support [, 0]. After testing several values for , we find that 0.01 is suitable for generating poisoned images with good visual quality. We then poison class images by . For ‘Double Blur’, we generate another random CUDA filter of size for each class, and then poison class c images by .
We compare the effectiveness of the two additional filter settings with CUDA and IRP on SL, SSL, and AT for ImageNet-100. The results in Fig. 6 show that Random Sharpness and Double Blur exhibit slightly better performance than CUDA in SSL. However, the effectiveness of Random Sharpness and Double Blur cannot match that of IRP. Additionally, we provide examples of images generated by Double Blur and Random Sharpness in Appendix 0.C.1. These samples highlight their inferior quality compared to those generated by IRP.
5 Conclusion
In this study, we conduct a theoretical analysis of CUDA, uncovering the sub-optimal gradients it introduces and elucidating the strategy it employs to induce class-wise bias for data poisoning. Building on these insights, we introduce IRP, which demonstrates high effectiveness in both SL and SSL as well as under state-of-the-art defense scenarios. IRP also maintains high image quality, rendering it suitable for data protection in real-world applications.
Acknowledgements
This research is supported by the National Research Foundation, Singapore and Infocomm Media Development Authority under its Trust Tech Funding Initiative and Strategic Capability Research Centres Funding Initiative. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not reflect the views of National Research Foundation, Singapore and Infocomm Media Development Authority.
Mingzhi Lyu and Fan Wang are supported by ROSE @ NTU, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore.
References
- [1] (2012) Poisoning attacks against support vector machines. arXiv preprint arXiv:1206.6389. Cited by: §2.
- [2] (2021) Large image datasets: a pyrrhic win for computer vision?. In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1536–1546. Cited by: §1.
- [3] (2020) A finite sample analysis of the benign overfitting phenomenon for ridge function estimation. arXiv preprint arXiv:2007.12882. Cited by: §2.
- [4] (2020) A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709. Cited by: §0.B.1, §4.1.
- [5] (2021) Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 15750–15758. Cited by: §0.B.1, §4.1, §4.2.
- [6] (2021) An empirical study of training self-supervised vision transformers. arXiv preprint arXiv:2104.02057. Cited by: §0.B.1, §4.1, §4.2.
- [7] (2011) An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pp. 215–223. Cited by: §4.1.
- [8] (2009) Imagenet: a large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Cited by: 4th item, §4.1.
- [9] (2017) Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552. Cited by: §1, §2.
- [10] (2024) The devil’s advocate: shattering the illusion of unexploitable data using diffusion models. External Links: 2303.08500, Link Cited by: §0.C.2.
- [11] (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §4.1.
- [12] (2019) Learning to confuse: generating training time adversarial data with auto-encoder. Advances in Neural Information Processing Systems 32. Cited by: §1, §2, §2.
- [13] (2021) Adversarial examples make strong poisons. Advances in Neural Information Processing Systems 34, pp. 30339–30351. Cited by: §1, §2, §2, §2, §4.1.
- [14] (2022) Robust unlearnable examples: protecting data privacy against adversarial learning. In International Conference on Learning Representations, External Links: Link Cited by: §1, §2, §2, §4.1, §4.2.
- [15] (2014) Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572. Cited by: §2, §4.2.
- [16] (2020) Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems 33, pp. 21271–21284. Cited by: §0.B.1, §4.1, §4.2.
- [17] (2023) Indiscriminate poisoning attacks on unsupervised contrastive learning. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: §1, §1, §2, §2, §4.1, §4.1.
- [18] (2016) Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778. Cited by: §4.1.
- [19] (2017) Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700–4708. Cited by: §4.1.
- [20] (2021) Unlearnable examples: making personal data unexploitable. arXiv preprint arXiv:2101.04898. Cited by: §1, §2, §2, §4.1.
- [21] (2019) Adversarial examples are not bugs, they are features. Advances in neural information processing systems 32. Cited by: §2.
- [22] (2018) Manipulating machine learning: poisoning attacks and countermeasures for regression learning. In 2018 IEEE symposium on security and privacy (SP), pp. 19–35. Cited by: §2.
- [23] (2009) Learning multiple layers of features from tiny images. Cited by: Figure 2, Figure 2, §1, §4.1.
- [24] (2023) Image shortcut squeezing: countering perturbative availability poisons with compression. arXiv preprint arXiv:2301.13838. Cited by: §1, §2, §4.2.
- [25] (2012) No-reference image quality assessment in the spatial domain. IEEE Transactions on image processing 21 (12), pp. 4695–4708. Cited by: §4.4.
- [26] (2023) Learning the unlearnable: adversarial augmentations suppress unlearnable example attacks. External Links: 2303.15127, Link Cited by: §0.C.2.
- [27] (2023) Cuda: convolution-based unlearnable datasets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3862–3871. Cited by: §1, §1, §2, §2, §3.2, §4.1.
- [28] (2018) Mobilenetv2: inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4510–4520. Cited by: §4.1.
- [29] (2022) Autoregressive perturbations for data poisoning. Advances in Neural Information Processing Systems 35, pp. 27374–27386. Cited by: §1, §2, §2, §2, §4.1.
- [30] (2023) Glaze: protecting artists from style mimicry by text-to-image models. External Links: 2302.04222, Link Cited by: §4.3.
- [31] (2024) Nightshade: prompt-specific poisoning attacks on text-to-image generative models. External Links: 2310.13828, Link Cited by: §4.3.
- [32] (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556. Cited by: §4.1.
- [33] (2021) Better safe than sorry: preventing delusive adversaries with adversarial training. Advances in Neural Information Processing Systems 34, pp. 16209–16225. Cited by: §1, §1, §2, §2.
- [34] (2023)Getty images sues ai art generator stable diffusion in the us for copyright infringement(Website) External Links: Link Cited by: §1.
- [35] (2023) Exploring clip for assessing the look and feel of images. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, pp. 2555–2563. Cited by: §4.4.
- [36] (2024) Corrupting convolution-based unlearnable datasets with pixel-based image transformations. External Links: 2311.18403, Link Cited by: §0.C.2.
- [37] (2004) Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13 (4), pp. 600–612. Cited by: §4.4.
- [38] (2003) Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, Vol. 2, pp. 1398–1402. Cited by: §4.4.
- [39] (2023) Is adversarial training really a silver bullet for mitigating data poisoning?. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: §2.
- [40] (2023) One-pixel shortcut: on the learning preference of deep neural networks. In The Eleventh International Conference on Learning Representations, Cited by: §1, §2, §2, §4.1.
- [41] (2022) Availability attacks create shortcuts. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 2367–2376. Cited by: §1, §2, §2, §4.1.
- [42] (2021) Neural tangent generalization attacks. In International Conference on Machine Learning, pp. 12230–12240. Cited by: §1, §2, §2.
- [43] (2019) Cutmix: regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 6023–6032. Cited by: §1, §2.
- [44] (2018) Mixup: beyond empirical risk minimization. In International Conference on Learning Representations, Cited by: §1, §2.
- [45] (2018) The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595. Cited by: §4.4.
Appendix 0.A Proofs
0.A.1 Proof of Theorem 1
Lemma 1
Given a random variable, follows a uniform distribution with support . The p.d.f. of is and
Proof
Considering
| (7) |
which is the c.d.f. of . To obtain the p.d.f. of , we compute . Note that the support of is not . To obtain
| (8) |
q.e.d.
Theorem 0.A.1
Given two different filters, and , generated by CUDA, if the condition is satisfied, and have the following properties.
- 1.
The peak of is located at the center.
- 2.
The peak of occurs at the position where two ones in and appear at the same position in .
- 3.
The probability of peak of at the center is .
- 4.
The expected peak of is higher than the expected peak of .
Proof
Property 1: The center of is at , where two filters completely overlap, and its minimum value is larger than 1. The maximum value at other positions is less than . Using the assumption, we can deduce that , for all .
Property 2: When the one in and the one in appear at the same position in , the minimum value of at the location is larger than one. The maximum value of other positions is less than . Based on the assumption, we achieve property 2.
Property 3: Based on property 2, the peak of is at the position where the one in and the one in appear at the same position in . If the peak is at the center of , it means that the one in and the one in must be at the position. The probability that they are in the position is .
Property 4: Utilizing Lemma 1 and the preceding properties, we can derive the expected peak value of :
| (9) |
When two ones are at the same position in and , achieves its maximum value. Thus the expected maximum peak value of is
| (10) |
0.A.2 Proof of Theorem 2
Theorem 0.A.2
Given a CUDA filter size being larger than 1 and being the vector form of the patch in involved in producing , i.e., , where is a matrix, if the column vectors of , excluding the one corresponding to , span , then, such that .
Proof
Let the patch in involved in computing the patch be . Since the patch size of and the size of are both , has the size of . When , is larger than . Reshaping and as vectors, and , their relationship can be represented as a linear equation, i.e., , where is a matrix. The matrix is constructed by . To achieve perfect reconstruction for all , must be , where the location of the one corresponds to in . Since the column vectors of , excluding the one corresponding to , span , cannot be orthogonal to all these column vectors. Thus, cannot be obtained.
0.A.3 Extending CUDA to CIFAR-10
Since CUDA uses and for CIFAR-10, it fulfils the condition . According to Theorem 0.A.3, CUDA filters exhibit similar properties shown in Theorem 0.A.1 on CIFAR-10.
Theorem 0.A.3
Given two different filters, and generated by CUDA, if the condition is satisfied, and have the following properties.
- 1.
The expected peak of is located at the center.
- 2.
The expected peak of occurs at the position where two ones in and appear at the same position in .
- 3.
The probability of the expected peak of at the centre is .
- 4.
The expected peak of is higher than the expected peak of .
Proof
Property 1: The center of is at , where two filters completely overlap, and its expected value is
| (11) |
Let be the number of pixels overlapped when computing , where . We have the following cases:
Case 1: If both ones are not inside the overlapping region,
Case 2: If both ones are inside the overlapping region but at different locations,
Case 3: If both ones are inside the overlapping region and at the same location,
Case 4: If only one is inside the overlapping region,
Among the four cases, the third case is the largest one, since . Clearly is larger than in the third case. Thus, the expected peak of is at the center.
Property 2: When two ones overlap, the minimum value of is larger than 1. The expected maximum value of when two ones do not overlap is
Using the assumption, the expected peak appears at the position when two ones overlap.
Property 3: Based on property 2, the expected peak appears at the position when two ones are at the same position. Thus, the probability of the expected peak appearing at the center is
Property 4: The proof is the same as the proof of property 4 in Theorem 0.A.1.
Appendix 0.B Experiment Setup
0.B.1 Training details
| Setting | SSL Pretrain | SSL Linear Probing | SL |
|---|---|---|---|
| Augmentations | Pretrain | LinProbe | LinProbe |
| Epochs | 400 | 200 | 100 | 100 |
| Batch Size | 512 | 128 | 512 | 256 |
| LR x BatchSize/256 | 1.0 | 0.25 | 1.0 | 10.0 | 0.1 |
| LR Warmup | 10 | 0 | 0 |
| LR Decay | Cosine | Step (0.2x at 60,75,90) | Cosine |
| Optimizer | SGD | SGD | SGD |
| Momentum | 0.9 | 0.9 | 0.9 |
| Weight Decay | 1e-4 | 0 | 5e-4 |
The training details are provided in Table 7. However, there are two exceptions: for VGG-19 and ViT models, the learning rate in SL is set to 0.01, and the training epoch for ViT is adjusted to 200. For the SL training, we utilize the standard linear probing augmentation of SSL as it has demonstrated an enhancement in clean test accuracy. This implies that the linear probing augmentation acts as an effective defense against certain DAAs. Specifically, in ImageNet training, the data augmentation comprises RandomResizedCrop with an expected output size of 224 along with RandomHorizontalFlip. For CIFAR-10 training, RandomCrop with an output size of 32 and padding set to 4 is employed, alongside RandomHorizontalFlip. For SSL pretraining and linear probing, we follow the standard augmentations used in [4, 5, 6, 16].
Appendix 0.C Additional Results
0.C.1 Images with Double Blur and Random Sharpness
The examples presented in Figure 7 depict outputs generated through Double Blur, Random Sharpness, and IRP. The images produced by Double Blur exhibit the highest level of blurriness. The degree of blurriness in the images generated by Random Sharpness varies, as the sharpness filters are randomly generated class-wise. Conversely, the images produced by IRP are the clearest, owing to the optimization process customized for the blurriness filter in IRP.
0.C.2 Additional Defenses
We further evaluate IRP on three recently invented poison defenses. AVATAR [10] adds Gaussian noise to poisoned images and ‘purifies’ them with a pretrained diffusion model. UEraser [26] applies multiple augmentations to poisoned images and selects the one that maximizes loss. COIN [36] applies random pixel interpolations to reduce the impact of filter-based poisons.
Figure 8 shows that AVATAR and COIN are useful defenses against both CUDA and IRP poisons. Under COIN purification, CUDA-poisoned CIFAR-10 permits a clean test accuracy of 71.90%. AVATAR raises the max accuracy of IRP-poisoned CIFAR-10 to 54.78%, though this value is still severely low compared to training on clean data. UEraser is less effective as a defense than standard methods like SSL and AT. Across all defenses, IRP is a more potent DAA than CUDA.
| DAA | SL | SSL | AT | AVATAR | UEraser | COIN | Max |
|---|---|---|---|---|---|---|---|
| Clean | 93.86 | 90.55 | 89.57 | - | - | - | 93.86 |
| CUDA | 21.89 | 66.58 | 48.58 | 55.93 | 40.56 | 71.90 | 71.90 |
| IRP | 10.39 | 43.24 | 32.21 | 54.78 | 26.55 | 53.32 | 54.78 |