arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.04627v1 [cs.AI] 04 Sep 2026

Leveraging Imperfect Restoration for Data Availability Attack

Yi Huang∗ Affiliation: College of Computing and Data Science, Nanyang Technological University, Singapore    Jeremy Styborski∗ Affiliation: College of Computing and Data Science, Nanyang Technological University, Singapore    Mingzhi Lyu∗ Affiliation: Rapid-Rich Object Search (ROSE) Lab, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore E-mail shellbyhuang@gmail.com, {styb0001,lyum0002,fan005,AdamsKong}@ntu.edu.sg    Fan Wang Affiliation: Rapid-Rich Object Search (ROSE) Lab, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore E-mail shellbyhuang@gmail.com, {styb0001,lyum0002,fan005,AdamsKong}@ntu.edu.sg    Adams Kong Affiliation: College of Computing and Data Science, Nanyang Technological University, Singapore
Abstract

The abundance of online data is at risk of unauthorized usage in training deep learning models. To counter this, various Data Availability Attacks (DAAs) have been devised to make data unlearnable for such models by subtly perturbing the training data. However, existing attacks often excel against either Supervised Learning (SL) or Self-Supervised Learning (SSL) scenarios. Among these, a model-free approach that generates a Convolution-based Unlearnable Dataset (CUDA) stands out as the most robust DAA across both SSL and SL. Nonetheless, CUDA’s effectiveness against SSL is underwhelming and it faces a severe trade-off between image quality and its poisoning effect. In this paper, we conduct a theoretical analysis of CUDA, uncovering the sub-optimal gradients it introduces and elucidating the strategy it employs to induce class-wise bias for data poisoning. Building on this, we propose a novel poisoning method named Imperfect Restoration Poisoning (IRP), aiming to preserve high image quality while achieving strong poisoning effects. Through extensive comparisons of IRP with eight baselines across SL and SSL, coupled with evaluations alongside five representative defense methods, we showcase the superiority of IRP. Code: https://github.com/lyumingzhi/IRP

Keywords: 
Data Availability Attacks Supervised Learning Self-Supervised Learning
**footnotetext: Equal contribution

1 Introduction

The proliferation of online data has proven an indispensable resource for the advancement of deep learning models. However, the collection of certain datasets without explicit consent poses a possible threat to personal privacy [2]. Additionally, the emergence of generative AI presents new challenges, as training with unauthorized data could infringe upon the owners’ copyright [34].

DAAs have emerged as a promising strategy to address the issue of unauthorized data usage[12, 42, 20, 13, 33, 29, 41, 40, 17, 27, 14]. These attacks introduce subtle perturbations into training data, hindering the model’s ability to effectively learn useful information, which subsequently leads to poor performance on unseen data. This poor performance manifests as a substantial decrease in accuracy when the model is tested on clean data in classification tasks. Currently, the majority of DAAs are developed within the framework of SL. However, a recent study by He et al. [17] reveals that Adversarial Poison (AP) [13], a cutting-edge DAA targeting SL, lacks effectiveness when applied to SSL. In response, they introduce a novel approach, termed Contrastive Poison (CP) [17], to counter SSL. Their research leads to several questions: (1) Are other DAAs designed for SL also ineffective against SSL (2) Does the proposed CP method maintain similar effectiveness in SL scenarios? (3) Can we develop a DAA capable of conducting effective attacks in both SL and SSL scenarios?

Refer to caption
Figure 1: The clean test accuracies of different DAAs on CIFAR-10 [23]. Lower accuracy represents better attack performance.
Refer to caption
Figure 2: The poisoned images generated by CUDA with different kernel sizes.

To answer these questions, we experimentally evaluate seven representative DAAs designed for SL on SSL. SSL is a strong defense against most DAAs due to its augmentation invariance objective; DAA features that are distorted or obfuscated by augmentations are ignored by the SSL algorithm and cannot affect downstream tasks. Accordingly, we find in Fig. 2 that most poisons for CIFAR-10 [23] fail to generalize to SSL. Only CUDA [27] and CP [17] exhibit significant impact on SSL. CUDA filters inject poison features throughout images that are resistant to cropping, flipping, and color shifts, and CP is designed specifically to counter SSL. We note a significant loss of performance for CP in SL scenarios. Therefore, CUDA emerges as the most potent DAA in both SSL and SL. However, the clean test accuracy of SSL trained on CUDA still hovers around 70%, suggesting that models trained on poisoned data retain some usability. As suggested by Sadasivan et al. [27], increasing the size of the convolution kernels in CUDA could provide a stronger poisoning effect, but it also degrades image quality. Fig. 2 shows that increased kernel size results in obviously blurred images. This compromise is often unacceptable in practical scenarios, particularly for sharing images and artwork online.

To comprehend CUDA’s efficacy, Sadasivan et al. provide a theoretical analysis under assumptions of Gaussian-distributed data with independent elements and a two-class scenario. However, these assumptions may not hold in real-world scenarios. In addition, their analysis is not based on deep learning, making it challenging to derive insights for designing a more effective poison method. To gain a deeper understanding of CUDA’s impact on model training, we first conduct a theoretical analysis of CUDA using a deep learning model. Our analysis reveals how CUDA generates sub-optimal gradients for clean data and introduces class-wise bias through random filters. Building upon this analysis, we propose a new poisoning method, named Imperfect Restoration Poisoning (IRP), which maintains high image quality while achieving a stronger poisoning effect. We compare IRP to eight representative DAAs in both SL and SSL scenarios. In the SL scenario, we additionally compare IRP with other DAAs under various defense mechanisms, including adversarial training [33], Image Shortcut Squeezing [24], Mixup [44], Cutout [9], and CutMix [43]. All experimental results indicate that IRP significantly outperforms previous DAAs. Our main contributions are:

  • •

    We experimentally assess eight representative DAAs in both supervised and self-supervised learning scenarios, revealing that existing DAAs fail to achieve satisfactory effectiveness simultaneously in both contexts.

  • •

    We conduct a theoretical analysis of CUDA on deep learning models, revealing how it generates sub-optimal gradients for clean data and introduces class-wise bias through random filters.

  • •

    We propose a novel DAA named IRP based on imperfect restoration, which achieves high effectiveness in both SL and SSL and maintains superior image quality compared to CUDA.

  • •

    We perform an exhaustive comparison between IRP and eight other DAAs across SL, SSL, and five defense techniques on CIFAR-10 and a subset of ImageNet [8]. IRP outperforms baseline poisons by a large margin in all cases. We further verify IRP across five additional architectures and two supplementary datasets. IRP displays high effectiveness across all experiments.

2 Related Works

DAAs have garnered significant attention in recent years for their application to data protection. Various methodologies have been proposed to deliberately induce malfunctions in the target model in order to safeguard against unauthorized data usage. DAAs can be broadly categorized into two types, model-reliant methods [12, 42, 20, 13, 14, 17] and model-free methods [29, 40, 41, 27].

Model-Reliant Methods: Early approaches to DAAs often frame the poisoning problem as a bilevel optimization task, balancing loss minimization with respect to model parameters with loss maximization with respect to perturbed inputs. While these methods [22, 1] initially showed promise, their applicability to deep neural networks is limited due to intractability in obtaining exact solutions. Efforts to address optimization challenges in perturbation generation, as seen in works by Feng et al. [12] and Yuan et al. [42], often introduce significant computational overhead, limiting scalability for real-world applications. Alternatively, Huang et al.[20] propose another bi-level Error Minimizing (EM) approach, deliberately optimizing perturbations to minimize training loss. By alternately training a surrogate model and optimizing perturbations, the efficacy of the perturbations is attained through multiple rounds of optimization. In contrast to traditional bi-level approaches, Fowl et al. [13] demonstrate the effectiveness of utilizing common objectives from adversarial examples for generating potent Adversarial Poisons (AP). However, Tao et al. [33] pinpoint that these DAAs can be easily broken by adversarial training. In response, Fu et al. [14] introduce Robust Error-Minimizing (REM) which enhances the poisoning effects against adversarial training by replacing the surrogate in EM with an adversarially-trained model. Recently, He et al. [17] assert that all aforementioned methods are tailored for supervised scenarios, thereby making them unsuitable for self-supervised learning. To address this gap, they introduce Contrastive Poison (CP), specifically designed for self-supervised learning. However, their assessment of DAAs tailored for SL is exclusively centered on AP, with other DAAs left unexplored. Additionally, while they show some effectiveness of their class-wise CP on both SL and SSL, the test accuracies on SSL remain high, ranging from 60.7% to 68.0%. In addition, class-wise additive perturbations can be recovered by the average image of a class, making them relatively easy to remove [29]. Therefore, we do not consider the class-wise poisons that apply the same additive noise to images, including class-wise attacks by EM, AP, and CP.

Model-Free Methods: Model-free DAAs tackle the data protection issue by leveraging shortcut learning. These methods are grounded in the observation that Deep Neural Networks (DNNs) often prioritize easily learnable shortcuts over semantic features, allowing them to distinguish between examples from different classes [3, 21]. Expanding on this concept, Yu et al. [41] propose incorporating artificially generated Linearly Separable Patterns (LSP) sampled from high-dimensional Gaussian distributions into images as an effective method for data poisoning. Sandoval-Segura et al. [29] discover that applying additive perturbations derived from autoregressive (AR) processes to clean data can serve as an effective shortcut. Wu et al. [40] examine the model’s susceptibility to sparse poisons and demonstrate that consistently perturbing only one pixel is adequate to generate potent poisons. In contrast to applying additive perturbations, Sadasivan et al. [27] introduce a Convolution-based Unlearnable Dataset (CUDA) generation method, which utilizes randomly generated class-wise convolutional filters. In our evaluation, CUDA demonstrates moderate robustness to SSL. However, its effectiveness on SSL remains insufficient and CUDA grapples with a severe trade-off between image quality and poison efficacy.

DAA Mitigation Strategies: Previous studies have demonstrated that Adversarial Training (AT) [15] is an effective method to mitigate the potency of DAAs [33, 39]. However, AT is impeded by its significant computational demands. Alternatively, various image preprocessing techniques, such as data augmentation (e.g., Mixup [44], Cutout [9], and CutMix [43]), have been explored. While these augmentation-based techniques show some impact, they are not as effective as AT [13]. Recently, Liu et al. [24] introduce Image Shortcut Squeezing (ISS) [24], a compression-based approach to counter DAAs. They demonstrate that ISS surpasses previously studied data augmentation countermeasures and achieves results comparable to AT. We evaluate the proposed IRP against multiple countermeasures, including AT, ISS, Mixup, Cutout, and CutMix.

3 Analysis

3.1 Threat Model

Like CUDA, most DAA papers define an “attacker” that applies a DAA on a labeled dataset that they want to protect yet publish. The attacker assumes no knowledge of any downstream DNN training process. It is therefore crucial to demonstrate that DAAs are effective across multiple training methods, including SL, SSL, and AT. The attacker generates class-based perturbations to their image set in order to inject poison features. The attacker publishes their poisoned images, optionally with captions or labels. Model trainers collect these poisoned images into datasets and may further alter or label the images.

3.2 Notation and Preliminaries

For clarity, we define a set of notations. ff represents a Convolution Neural Network (CNN), including its training loss function but excluding the first layer convolution 2D filter W∈ℝd1×d2W\in\mathbb{R}^{d_{1}\times d_{2}}. X∈ℝm1×m2X\in\mathbb{R}^{m_{1}\times m_{2}} denotes an input image and XcX_{c} represents an image belonging to class c∈[1,…,ℂ]c\in[1,...,\mathbb{C}], where ℂ\mathbb{C} is total number of the classes. Given that cross-correlation is predominantly used in place of convolutions in practice, the first layer operation is denoted as X⋆WX\star W, where ⋆\star signifies cross-correlation. Note that WW is a part of the CNN and f⁡(X⋆W)f(X\star W) is the loss value of XX. To retain readability, we focus on gradient analysis of the first layer filter, WW. CUDA employs class-specific convolutional filters for data poisoning, and we provide a brief overview of its operation. Let Rc∈ℝκ×κR_{c}\in\mathbb{R}^{\kappa\times\kappa} be a CUDA filter for class cc images. According to [27], the generation procedure for RcR_{c} involves setting one random parameter among the κ2\kappa^{2} parameters to 1 within each RcR_{c}. Concurrently, the remaining parameters are initialized randomly, drawn from a uniform distribution with support [0,pu][0,p_{u}]. An image XcX_{c} belonging to class cc is poisoned by cross-correlation with the corresponding CUDA filter RcR_{c}, denoted as Xc⋆RcX_{c}\star R_{c}.

3.3 The Non-Optimal Gradient

To understand how CUDA lowers the accuracy of clean data, we analyze the gradient of ff for both clean and CUDA-poisoned data. Without loss of generality, the training set consists of ℂ\mathbb{C} clean images, each belonging to a different class, with corresponding poisoned images denoted as Xi⋆RiX_{i}\star R_{i}, i∈[1,…,ℂ]i\in[1,...,\mathbb{C}]. The gradient of the clean training data is

∂∑i=1ℂf⁡(Xi⋆W)∂W=∑i=1ℂXi⋆∂f⁡(ui)∂ui,\frac{\partial\sum_{i=1}^{\mathbb{C}}f(X_{i}\star W)}{\partial W}=\sum_{i=1}^{\mathbb{C}}X_{i}\star\frac{\partial f(u_{i})}{\partial u_{i}}, (1)

where uiu_{i} is an intermediate variable equal to Xi⋆WX_{i}\star W. The gradient of the poisoned training data is

∂∑i=1ℂf⁡((Xi⋆Ri)⋆W)∂W=∑i=1ℂ(Xi⋆Ri)⋆∂f⁡(up,i)∂up,i,\frac{\partial\sum_{i=1}^{\mathbb{C}}f((X_{i}\star R_{i})\star W)}{\partial W}=\sum_{i=1}^{\mathbb{C}}(X_{i}\star R_{i})\star\frac{\partial f(u_{p,i})}{\partial u_{p,i}}, (2)

where the subscript pp indicates that the intermediate variable is from poisoned data. Using these gradients to update WW, we obtain

Wt+1=Wt−λ​∑i=1ℂXi⋆∂f⁡(ui)∂uiWp,t+1=Wt−λ​∑i=1ℂ(Xi⋆Ri)⋆∂f⁡(up,i)∂up,i\begin{split}W_{t+1}=W_{t}-\lambda\sum_{i=1}^{\mathbb{C}}X_{i}\star\frac{\partial f(u_{i})}{\partial u_{i}}\\ W_{p,t+1}=W_{t}-\lambda\sum_{i=1}^{\mathbb{C}}(X_{i}\star R_{i})\star\frac{\partial f(u_{p,i})}{\partial u_{p,i}}\end{split} (3)

where λ\lambda is learning rate. Inputting all clean training images, X1,⋯,XCX_{1},\cdots,X_{C} to f(⋅⋆Wt+1)f(\cdot\star W_{t+1}) and f(⋅⋆Wp,t+1)f(\cdot\star W_{p,t+1}), we obtain clean and poison loss values ∑i=1ℂf⁡(Xi⋆Wt+1)\sum_{i=1}^{\mathbb{C}}f\left(X_{i}\star W_{t+1}\right) and ∑i=1ℂf⁡(Xi⋆Wp,t+1)\sum_{i=1}^{\mathbb{C}}f\left(X_{i}\star W_{p,t+1}\right), respectively. Given that ∑i=1ℂXi⋆∂f⁡(oi)∂i\sum_{i=1}^{\mathbb{C}}X_{i}\star\frac{\partial f(o_{i})}{\partial i} in Eq. 1 is the optimal gradient descent direction for the clean data, any deviation from this direction will result in increased loss when λ\lambda is small. Thus, ∑i=1ℂf⁡(Xi⋆Wp,t+1)>∑i=1ℂf⁡(Xi⋆Wt+1)\sum_{i=1}^{\mathbb{C}}f\left(X_{i}\star W_{p,t+1}\right)>\sum_{i=1}^{\mathbb{C}}f\left(X_{i}\star W_{t+1}\right) for clean data. In other words, CUDA employs a non-optimal gradient to disrupt training, rendering clean data unrecognizable.

3.4 Poisoning through Class-Wise Bias

Though the previous subsection uncovers how CUDA filters increase training loss, it does not indicate how these filters poison the network. To further investigate its mechanism, we input Xc⋆RcX_{c}\star R_{c} to f(⋅⋆Wp,t+1)f(\cdot\star W_{p,t+1}), whose first layer output is

(Xc⋆Rc)⋆Wp,t+1=(Xc⋆Rc)⋆Wt−λ⁡(Xc⋆Rc)⋆∑i=1C((Xi⋆Ri)⋆∂f⁡(up,i)∂up,i).(X_{c}\star R_{c})\star W_{p,t+1}=(X_{c}\star R_{c})\star W_{t}-\lambda(X_{c}\star R_{c})\star\sum_{i=1}^{C}\left((X_{i}\star R_{i})\star\frac{\partial f(u_{p,i})}{\partial u_{p,i}}\right). (4)

Assuming WtW_{t} is not poisoned by CUDA filters, we ignore the term (Xc⋆Rc)⋆Wt(X_{c}\star R_{c})\star W_{t}. Since cross-correlation is equivalent to a rotated convolution (i.e., Xc⋆Rc=Xc∗Rc​πX_{c}\star R_{c}=X_{c}\ast R_{c\pi}, where ∗\ast represents convolution and Rc​πR_{c\pi} represents rotating the 2D filter RcR_{c} by π\pi), the second term of Eq. 4 can be rewritten as

λ​Xc∗∑i=1CRc​π∗Ri∗Xi​π∗∂f⁡(up,i)∂up,i.\lambda X_{c}*\sum_{i=1}^{C}R_{c\pi}*R_{i}*X_{i\pi}*\frac{\partial f(u_{p,i})}{\partial u_{p,i}}. (5)

Eq. 5 shows that the output of the first layer is influenced by Rc​π∗RiR_{c\pi}*R_{i}. Since Rc​π∗Ri=Rc​π⋆Ri​πR_{c\pi}*R_{i}=R_{c\pi}\star R_{i\pi}, we analyze Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi}. Theorem 1 describes the statistical probabilities of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} and Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi}. The proof is provided in Appendix 0.A.1.

Theorem 3.1

Given two different filters, RiR_{i} and RcR_{c}, generated by CUDA, if 1>(κ2−2)​pu2+2​pu1>(\kappa^{2}-2)p_{u}^{2}+2p_{u}, then Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} and Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} have the following properties:

  1. 1.

    The peak of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi}is located at the center.

  2. 2.

    The peak of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} occurs at the position where two ones in Rc​πR_{c\pi} and Ri​πR_{i\pi} appear at the same position in Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi}.

  3. 3.

    The probability that the peak of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} lies in the center is 1κ2\frac{1}{\kappa^{2}}.

  4. 4.

    The expected peak of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} is higher than the expected peak of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi}.

Figure 3 presents two CUDA filters along with the corresponding Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} and Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} to illustrate these statistical properties. In practice, CUDA utilizes pu=0.06p_{u}=0.06 and κ=9\kappa=9 for ImageNet, which fulfills 1>(κ2−2)​pu2+2​pu1>(\kappa^{2}-2)p_{u}^{2}+2p_{u}. Similar properties also apply on CIFAR-10 (see Appendix 0.A.3). These properties indicate that Rc​π∗RcR_{c\pi}*R_{c} in Eq. 5 retains more information in Xc​π∗∂f⁡(up,c)∂up,cX_{c\pi}*\frac{\partial f(u_{p,c})}{\partial u_{p,c}} but Rc​π∗RiR_{c\pi}*R_{i} shifts the information in Xi​π∗∂f⁡(up,i)∂up,iX_{i\pi}*\frac{\partial f(u_{p,i})}{\partial u_{p,i}} to a different position with a probability of (κ2−1)/κ2(\kappa^{2}-1)/\kappa^{2}. For κ=9\kappa=9, the probability is 0.9880.988. Consequently, the upper layers receive class cc information in the correct position, but information from other classes is shifted to incorrect positions. For pu=0.06p_{u}=0.06, we note the peak value of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} is smaller than the sum of its non-peak values (estimated mean ratio of 0.099). This means that non-peak values also play an important role in degrading other class information. This blurring effect also affect the cc-class, though to a lesser extent. More clearly, the estimated mean ratio of the peak value of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} over the sum of its non-peak values is 0.106. Combining these two effects, CUDA filters utilize peak shifts and unique blur patterns to introduce class-wise bias in the poisoned class.

Refer to caption
Figure 3: Two CUDA filters Rc​πR_{c\pi} and Ri​πR_{i\pi}(the first and second plots) along with the corresponding Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} (the third plot) and Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} (the last plot).

3.5 Poisoning with Imperfect Restoration

The previous analysis shows that CUDA filters create class-wise bias, and that Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} retains the training information of class cc. When Xc⋆RcX_{c}\star R_{c} is closer to XcX_{c}, it is expected to preserve the class cc information better and enhance the image quality. However, if Xc⋆Rc=XcX_{c}\star R_{c}=X_{c}, then there is no poisoning effect. Thus, we propose the Imperfect Restoration Poison (IRP). Let Xc​iX_{ci} be one of the class cc images, where i∈1,⋯,ni\in{1,\cdots,n}. IRP first applies CUDA filter RcR_{c} on all class cc training images and obtains Xc​i⋆RcX_{ci}\star R_{c}. Then, the filtered images are cropped into patches with the size same as the CUDA filter i.e., κ×κ\kappa\times\kappa, denoted as (Xc​i⋆Rc)j(X_{ci}\star R_{c})_{j}. The pixel value of Xc​iX_{ci} corresponding to the center of the patch is denoted yc​i​jy_{cij}, and the patch is reshaped into a column vector denoted as ηc​i​j\eta_{cij}. Finally, a class-wise linear model αc∈ℝκ2×1\alpha_{c}\in\mathbb{R}^{\kappa^{2}\times 1} and the mean square loss are used to estimate the value of yc​i​jy_{cij}. More precisely, the class-wise model αc^\widehat{\alpha_{c}} is obtained by minimizing

αc^=argminαc​∑i∑j‖yc​i​j−αcT​ηc​i​j‖22\widehat{\alpha_{c}}=\underset{\alpha_{c}}{\text{argmin}}\sum_{i}\sum_{j}\|y_{cij}-\alpha_{c}^{T}\eta_{cij}\|_{2}^{2} (6)

Once αc^\widehat{\alpha_{c}} is obtained, it is reshaped back to a κ×κ\kappa\times\kappa filter, denoted as AcA_{c}. Finally, IRP uses Xc​i⋆PcX_{ci}\star P_{c} to perform poisoning, where Pc=Rc⋆AcP_{c}=R_{c}\star A_{c}. Theorem 0.A.2 establishes that IRP is incapable of achieving perfect restoration (proof in Appendix 0.A.2).

Theorem 3.2

Given a CUDA filter size larger than 1 and xc​i​jx_{cij} as the vector form of the patch in Xc​iX_{ci} involved in producing ηc​i​j\eta_{cij} (i.e., ηc​i​j=℘​xc​i​j\eta_{cij}=\wp x_{cij}, where ℘\wp is a κ2×(3​κ−2)2\kappa^{2}\times(3\kappa-2)^{2} matrix), if the column vectors of ℘\wp, excluding the one corresponding to yc​i​jy_{cij}, span ℝκ2\mathbb{R}^{\kappa^{2}}, then, ∃xc​i​j∈ℝ(3​κ−2)2\exists x_{cij}\in\mathbb{R}^{(3\kappa-2)^{2}} such that yc​i​j≠acT​℘​xc​i​jy_{cij}\neq a_{c}^{T}\wp x_{cij}.

Refer to caption
Figure 4: Two IRP filters Pc​πP_{c\pi} and Pi​πP_{i\pi} (the first and second plots) along with the corresponding Pc​π⋆Pc​πP_{c\pi}\star P_{c\pi} (the third plot) and Pc​π⋆Pi​πP_{c\pi}\star P_{i\pi} (the last plot).

Fig. 4 shows two IRP filters with the corresponding self-correlation and cross-correlation. Similar to CUDA, we observe that Pc​π⋆Pc​πP_{c\pi}\star P_{c\pi} exhibits a peak at the center, while Pc​π⋆Pi​πP_{c\pi}\star P_{i\pi} has a peak in a different position. However, they have more complex patterns around the peaks. Therefore, like CUDA, IRP exploits peak shifts to create class-wise bias. However, IRP filters have turbulent patterns around the peaks, creating unique high-frequency attack patterns that are more complex than the low-frequency patterns from CUDA. Due to their dependence on the training data, it would be challenging to theoretically analyze IRP filters, as was done for CUDA, without proper assumptions on the data distribution. We refrain from imposing unrealistic assumptions on the data distribution (e.g., Gaussian) as this could lead to misleading analytic results.

Fig. 5 displays some images generated by IRP and CUDA. IRP images are visually clearer and maintain higher quality. Notably, the man’s face is sharper and the spots of the salamander and stingray are easily distinguishable under IRP protection. The backgrounds of all three images are sharper under IRP.

Refer to caption
(a) CUDA
Refer to caption
(b) IRP
Figure 5: Poisoned examples generated by CUDA (left) and IRP (right).

4 Experiments

In this section, we begin by evaluating the effectiveness of IRP alongside eight DAA methods in standard SL and SSL settings, as well as against representative defense methods. We then investigate the effectiveness of different DAAs under partial poisoning scenarios. For model-reliant methods, our experiments primarily focus on the nominal setup, wherein a surrogate model generates poison images. To assess the transferability of IRP, we evaluate IRP across six different architectures on three datasets. Finally, we compare the image quality generated by CUDA and IRP and ablate different filter types across various training methods.

4.1 Experimental Setup

Datasets: We examine poisoning on CIFAR-10 [23], CIFAR-100 [23], STL-10 [7], and ImageNet-100, a subset of ImageNet [8] consisting of 100 classes. By default, we use 50,000 images for training and 10,000 images for testing on CIFAR-10 and CIFAR-100. For STL-10, we use their default 5000 training images and 8000 test images for training and testing, respectively. For ImageNet-100, we train on the images from the first 100 classes of the official training set and test on all corresponding images from the official validation set.

Poison Settings: We compare IRP with eight state-of-the-art DAAs, including four model-reliant methods, including EM [20], TAP [13], REM [14], and CP [17], and four model-free methods, including CUDA [27], LSP [41], OPS [40], and AR [29]. As generating poison data using model-reliant methods is time-consuming on ImageNet, we generate a 20% subset of ImageNet-100 for CP, TAP, and EM. We maintain all default settings and retain the L∞L_{\infty} norm of the defensive perturbation at 8/255. We utilize the full REM-poisoned ImageNet-100 dataset shared by the REM authors. According to [14], the L∞L_{\infty} norm of the defensive perturbation ρu\rho_{u} is set to 8/2558/255, and the adversarial perturbation radius ρa\rho_{a} is set to 4/2554/255 for controlling the level of protection of the noise against adversarial training. For model-free methods, we maintain the settings as per their original papers, except for LSP and OPS on ImageNet-100. As the original settings could not yield a satisfactory poisoning effect, we amplify the perturbation of LSP and OPS to L∞=16/255L_{\infty}=16/255 and L0=5L_{0}=5, respectively. As AR does not report its effectiveness on ImageNet-100 and has shown to be vulnerable to SSL and AT on CIFAR-10, we opt not to include AR in further testing on ImageNet-100.

Models: Unless stated otherwise, we employ ResNet-18 [18] as the surrogate for model-reliant methods and as the target models for all the testing methods in both SL and SSL scenarios. Various SSL algorithms are employed, including SimCLR [4], SimSiam[5], MoCoV3[6] and BYOL[16]. Linear probing accuracy, consistent with CP [17], is utilized to evaluate the effectiveness of DAAs in SSL. To assess transferability, we evaluate IRP on target models with various architectures, including ResNet-18, ResNet-34 [18], VGG-19 [32], DenseNet-121 [19], MobileNet-V2 [28], and ViT [11]. Training details are deferred to Appendix 0.B.1.

4.2 Comparisons to Baseline DAAs

Table 1: Top-1 clean test accuracies (%) of ResNet-18 trained with data poisoned by various DAAs on CIFAR-10. The “Max” column represents the highest clean test accuracy, which corresponds to the worst performance of each DAA under different training scenarios. Bold indicates the best performance, and underline denotes the second-best performance.
DAA SL SSL AT Cutout Cutmix Mixup ISS Max
Clean 93.86 90.55 89.57 95.62 95.78 95.46 82.40 95.62
EM 32.53 88.45 89.31 36.18 38.40 54.19 82.12 89.31
REM 18.91 86.70 34.45 19.67 27.46 23.27 80.04 86.70
AP 12.45 75.77 85.67 9.70 8.38 33.27 81.10 85.67
CP 93.59 73.08 88.43 94.51 93.77 93.36 82.21 94.51
LSP 21.47 88.87 85.67 18.47 21.68 22.76 79.33 88.87
OPS 27.24 87.79 20.10 65.12 85.27 37.56 72.14 87.79
AR 10.95 86.86 72.94 12.17 12.95 14.15 82.83 86.86
CUDA 21.89 66.58 48.58 23.46 24.04 21.75 21.94 66.58
IRP 10.39 43.24 32.21 10.85 15.21 16.03 29.42 43.24

Effectiveness on SL and SSL: We first evaluate IRP and the other DAAs on standard SL and SSL training scenarios on CIFAR-10 and ImageNet-100. SimCLR is utilized for model pretraining in the SSL evaluation. Columns 2-3 of Tables 1 and 2 show that IRP achieves state-of-the-art poisoning, enforcing the lowest clean test accuracies in both SL and SSL and for both CIFAR-10 and ImageNet-100.

We begin by analyzing results for CIFAR-10 in Table 1. For SL with CIFAR-10, AR and IRP demonstrate a similar level of poisoning effect under SL, reducing the test accuracy to around 10%. However, AR underperforms on SSL, permitting a significantly higher accuracy of 86.86% compared to that of IRP at 43.24%. CUDA is the second-best DAA for SSL with an accuracy of 66.58%, which is still drastically higher than that of IRP.

For ImageNet-100 results in Table 2, we begin by nothing the drop in effectiveness of OPS relative to CIFAR-10. This is attributed to the RandomResizeCrop augmentation utilized for training models on ImageNet-100. We chose to use this augmentation on ImageNet-100 as it is commonly used for SL and SSL on ImageNet-100 in literature and it enhances the effectiveness of all the defense methods. For SL with ImageNet-100, CUDA, EM, AP, and LSP demonstrate strong poisoning capabilities, reducing the clean test accuracy to below 10%. Even so, IRP outperforms all others with a test accuracy of 1.98%. For SSL with ImageNet-100, the test accuracies of most DAAs remain high, except for CUDA at 21.89% and CP at 18.76%. IRP outperforms both CUDA and CP, achieving a clean test accuracy of 9.30%. These results establish IRP as the most effective DAA under both SL and SSL scenarios.

Table 2: Top-1 clean test accuracies (%) of ResNet-18 trained with data poisoned by various DAAs on ImageNet-100. The “Max” column represents the highest clean test accuracy, which corresponds to the worst performance of each DAA under different training scenarios. Bold indicates the best performance, and underline denotes the second-best performance.
DAA SL SSL AT Cutout Cutmix Mixup ISS Max
Clean 77.66 70.96 69.76 78.04 81.08 80.38 71.58 80.38
EM 6.32 50.02 43.26 6.42 5.80 12.38 41.18 50.02
REM 14.18 65.30 57.34 15.60 16.10 33.08 67.52 67.52
AP 8.18 41.80 42.88 7.68 8.38 9.68 24.22 42.88
CP 57.46 18.76 48.76 58.20 62.28 61.76 50.18 62.28
LSP 6.62 62.06 28.92 5.06 7.84 5.50 33.70 62.06
OPS 51.00 62.22 52.56 53.86 65.12 48.82 57.36 65.12
CUDA 6.10 26.12 36.34 8.66 5.70 5.16 3.58 36.34
IRP 1.98 9.30 22.10 2.18 1.52 2.02 4.14 22.10

Effectiveness under Common Countermeasures: To further evaluate the effectiveness of IRP under defensive scenarios, we test IRP and the baselines on representative defense methods, including Adversarial Training (AT) [15], Image Shortcut Squeezing (ISS) [24], and common data augmentation techniques such as Cutout, CutMix, and Mixup. Following [14], we employ the L∞L_{\infty} norm for AT and set the norm bound to 4/255. 10-step PGD with a step size of 0.6/255 is used for AT. As suggested by [24], we employ a combination of grayscale and JPEG with the default JPEG quality factor (JPEG-10) for ISS to ensure global effectiveness against all DAAs. The results on CIFAR-10 and ImageNet-100 are presented in columns 4-8 of Tables 1 and 2, respectively. The last column displays the maximum clean test accuracy of the corresponding methods across all testing scenarios, representing the worst performance among all the test cases.

DAA results on CIFAR-10 in Table 1 vary across defenses. For instance, OPS shows high effectiveness on AT but is susceptible to ISS and CutMix, while AP performs well against CutMix and Cutout but is less effective against AT and ISS. Only CUDA and IRP demonstrate effectiveness across all countermeasures. However, CUDA permits maximum clean accuracies of 48.58% across all countermeasures and 66.58% across all testing cases, which are 16.37% and 23.24% higher, respectively, than the 43.24% maximum accuracy allowed by IRP.

Similarly, CUDA and IRP emerge as the two most effective methods across all defenses on ImageNet-100 in Table 2. However, CUDA permits a maximum clean accuracy of 36.34%, significantly higher than that of IRP at 22.10%. The results demonstrate that IRP is the most effective DAA in all testing scenarios.

Evaluation on Other SSL Algorithms: To further assess the efficacy of IRP on SSL, we test IRP-poisoned ImageNet-100 on three additional SSL algorithms: SimSiam[5], MoCoV3[6], and BYOL[16]. As shown in Table 3, IRP lowers the test accuracies of all SSL algorithms to around 10%. Even in its the worst case performance with BYOL, IRP reduces clean accuracy to 11.6%.

Table 3: Top-1 accuracies (%) of ResNet-18 trained with different SSL algorithms on ImageNet-100.
SimCLR MoCoV3 SimSiam BYOL Max
Clean 70.96 75.54 70.68 75.34 75.54
IRP 9.3 10.58 7.82 11.6 11.6

4.3 Additional Training Scenarios and Datasets

Different Poison Percentages: In this section, we first evaluate IRP on a more challenging and realistic learning scenario, where only a portion of the data is shielded by IRP. We randomly select a portion of the training data from the entire training dataset of CIFAR-10 and apply IRP to the selected subset. The poisoned data is denoted as DpD_{p} and the clean data in the training dataset that has not been used for generating poison data is denoted as DcD_{c}. Then, we conduct standard supervised training with ResNet-18 on the mixed data (Dp+cD_{p+c}) consisting of the DpD_{p} and DcD_{c}, as well as on the clean data subset DcD_{c}. The results in Table 4 show that the effectiveness drops quickly when the data are not 100% poisoned. This phenomenon applies for all other DAAs as well. Models trained only on DcD_{c} demonstrate a similar level of performance as the models trained with Dp+cD_{p+c}. This suggests that the high performance for <100% poisoning is not due to a failure of the DAAs. Finally, works on DAAs for artwork protection [30, 31] have noted similar trends but find that individuals can still protect their own data via DAA even if the majority of artworks are unprotected. That is, it’s still possible to make subsets of the training data unlearnable.

Table 4: The top-1 clean test accuracies (%) of ResNet-18 trained with partial poisoned data on CIFAR-10. The percentage in the first row indicates the percentage of poison data, and DcD_{c} indicates the remaining cleaning data. 0% and 100% signify training scenarios where the entire dataset remains either clean or fully poisoned, respectively.
0% 20% 40% 60% 80% 100%
Dp+cD_{p+c} DcD_{c} Dp+cD_{p+c} DcD_{c} Dp+cD_{p+c} DcD_{c} Dp+cD_{p+c} DcD_{c}
EM 93.86 94.02 92.38 92.72 91.74 90.71 87.84 85.75 77.36 32.53
REM 92.85 92.15 90.05 86.29 18.91
AP 92.57 91.16 90.72 88.11 12.45
LSP 92.95 92.40 90.19 85.76 21.47
OPS 93.40 92.25 90.63 86.68 27.24
AR 92.98 92.66 90.94 87.93 10.95
CUDA 93.27 92.27 90.51 86.90 21.89
IRP 93.07 92.56 89.89 85.70 10.39
Table 5: The top-1 clean test accuracies (%) of different models trained on data poisoned by IRP and CUDA with AT.
Dataset RN18 RN34 VGG19 DN121 MNv2 ViT-S Max
CIFAR10 Clean 89.57 89.32 78.57 79.97 75.94 47.84 89.57
CUDA 48.58 43.63 42.02 49.44 26.33 44.93 49.44
IRP 32.21 17.37 19.06 17.87 14.13 35.39 35.39
CIFAR100 Clean 65.30 66.43 48.50 58.92 47.00 27.19 66.43
CUDA 37.21 37.95 27.44 23.07 16.00 24.88 37.95
IRP 27.13 28.85 22.50 25.24 20.70 20.82 28.85
STL10 Clean 71.94 69.73 64.35 72.41 58.93 43.86 72.41
CUDA 44.70 44.35 45.99 39.39 44.49 42.14 45.99
IRP 34.10 34.33 33.89 33.19 33.81 35.58 35.58

Effectiveness on Different Models: The previous evaluations are all conducted on ResNet-18 (RN18). To evaluate the effectiveness of IRP on different architectures, we evaluate ResNet-34 (RN34), VGG19, DenseNet121 (DN121), MobileNetv2 (MNv2), and ViT-S on IRP-poisoned CIFAR-10. All models are trained with AT (L∞L_{\infty} constraint of 4/2554/255, as before), which emerged as the strongest defense against different DAAs from our experiments in Section 4.2. We include CUDA in this evaluation as a reference. On all the examined networks, IRP outperforms CUDA by a significant margin. In their respective worst cases, CUDA permits a test accuracy of 49.44%, while IRP allows a test accuracy of 35.39%. Table 5 shows that even under AT defense, IRP is still effective on all the examined architectures.

Evaluation on Different Datasets: We extend our evaluation to the CIFAR-100 and STL-10 datasets. The same set of models and AT settings are utilized. As shown in Table 5, both IRP and CUDA demonstrate effectiveness on CIFAR-100 and STL-10 across all the network architectures under AT. IRP consistently outperforms CUDA by a large margin on most networks for both CIFAR-100 and STL-10. In the two cases where IRP is weaker than CUDA, the clean test accuracies of IRP are still within 5% of those of CUDA. The maximum clean test accuracies of IRP on CIFAR-100 and STL-10 are 28.85% and 35.58%, respectively. Both of these values are significantly lower than those of CUDA, which are 37.95% and 45.99% on CIFAR-100 and STL-10, respectively. This experiment demonstrates that IRP is universally effective across networks and across different datasets.

4.4 Image Quality Comparison against CUDA

We compare the quality of ImageNet-100 images poisoned by IRP and CUDA. We quantify image quality with three common full-reference quality indices, including LPIPS [45], SSIM [37], and MS-SSIM [38], and two no-reference quality indices, including CLIP-IQA [35] and BRISQUE [25]. The original image is used as the reference for LPIPS, SSIM, and MS-SSIM. Table 6 presents the evaluation results of all quality indices. Both Table 6 and Fig. 5 demonstrate that IRP generates poisoned images with better image quality than CUDA.

Metric Clean CUDA IRP
LPIPS ↓\downarrow - 0.272 0.142
SSIM ↑\uparrow - 0.561 0.775
MS-SSIM↑\uparrow - 0.831 0.903
CLIP-IQA↑\uparrow 0.787 0.528 0.706
BRISQUE↓\downarrow 16.625 29.570 17.318
Table 6: ImageNet-100 image quality under different quality indices.
Refer to caption
Figure 6: The Top-1 accuracy of different setting on ImageNet-100.

4.5 Filter Ablation

The IRP filter Pc=Rc⋆AcP_{c}=R_{c}\star A_{c} is a kernel with a larger size of (2​κ+1)×(2​κ+1)(2\kappa+1)\times(2\kappa+1), where RcR_{c} and AcA_{c} are both filters of size κ×κ\kappa\times\kappa. To verify that the high effectiveness of IRP is not caused by the bigger kernel size, we generate filters of the same size as the IRP under two different filter settings. For ‘Random Sharpness’, we generate a random sharpness filter Ac​1A_{c1} of size κ×κ\kappa\times\kappa for each class. Within each Ac​1A_{c1}, one random parameter out of the κ×κ\kappa\times\kappa parameters is set to 1, while the remaining parameters are randomly initialized from a uniform distribution with support [−ps-p_{s}, 0]. After testing several values for psp_{s}, we find that 0.01 is suitable for generating poisoned images with good visual quality. We then poison class cc images by Xc⋆Rc⋆Ac​1X_{c}\star R_{c}\star A_{c1}. For ‘Double Blur’, we generate another random CUDA filter Rc​1R_{c1} of size κ×κ\kappa\times\kappa for each class, and then poison class c images by Xc⋆Rc⋆Rc​1X_{c}\star R_{c}\star R_{c1}.

We compare the effectiveness of the two additional filter settings with CUDA and IRP on SL, SSL, and AT for ImageNet-100. The results in Fig. 6 show that Random Sharpness and Double Blur exhibit slightly better performance than CUDA in SSL. However, the effectiveness of Random Sharpness and Double Blur cannot match that of IRP. Additionally, we provide examples of images generated by Double Blur and Random Sharpness in Appendix 0.C.1. These samples highlight their inferior quality compared to those generated by IRP.

5 Conclusion

In this study, we conduct a theoretical analysis of CUDA, uncovering the sub-optimal gradients it introduces and elucidating the strategy it employs to induce class-wise bias for data poisoning. Building on these insights, we introduce IRP, which demonstrates high effectiveness in both SL and SSL as well as under state-of-the-art defense scenarios. IRP also maintains high image quality, rendering it suitable for data protection in real-world applications.

Acknowledgements

This research is supported by the National Research Foundation, Singapore and Infocomm Media Development Authority under its Trust Tech Funding Initiative and Strategic Capability Research Centres Funding Initiative. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not reflect the views of National Research Foundation, Singapore and Infocomm Media Development Authority.

Mingzhi Lyu and Fan Wang are supported by ROSE @ NTU, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore.

References

  • [1] B. Biggio, B. Nelson, and P. Laskov (2012) Poisoning attacks against support vector machines. arXiv preprint arXiv:1206.6389. Cited by: §2.
  • [2] A. Birhane and V. U. Prabhu (2021) Large image datasets: a pyrrhic win for computer vision?. In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1536–1546. Cited by: §1.
  • [3] E. Caron and S. Chrétien (2020) A finite sample analysis of the benign overfitting phenomenon for ridge function estimation. arXiv preprint arXiv:2007.12882. Cited by: §2.
  • [4] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton (2020) A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709. Cited by: §0.B.1, §4.1.
  • [5] X. Chen and K. He (2021) Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 15750–15758. Cited by: §0.B.1, §4.1, §4.2.
  • [6] X. Chen*, S. Xie*, and K. He (2021) An empirical study of training self-supervised vision transformers. arXiv preprint arXiv:2104.02057. Cited by: §0.B.1, §4.1, §4.2.
  • [7] A. Coates, A. Ng, and H. Lee (2011) An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pp. 215–223. Cited by: §4.1.
  • [8] J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei (2009) Imagenet: a large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Cited by: 4th item, §4.1.
  • [9] T. DeVries and G. W. Taylor (2017) Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552. Cited by: §1, §2.
  • [10] H. M. Dolatabadi, S. Erfani, and C. Leckie (2024) The devil’s advocate: shattering the illusion of unexploitable data using diffusion models. External Links: 2303.08500, Link Cited by: §0.C.2.
  • [11] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §4.1.
  • [12] J. Feng, Q. Cai, and Z. Zhou (2019) Learning to confuse: generating training time adversarial data with auto-encoder. Advances in Neural Information Processing Systems 32. Cited by: §1, §2, §2.
  • [13] L. Fowl, M. Goldblum, P. Chiang, J. Geiping, W. Czaja, and T. Goldstein (2021) Adversarial examples make strong poisons. Advances in Neural Information Processing Systems 34, pp. 30339–30351. Cited by: §1, §2, §2, §2, §4.1.
  • [14] S. Fu, F. He, Y. Liu, L. Shen, and D. Tao (2022) Robust unlearnable examples: protecting data privacy against adversarial learning. In International Conference on Learning Representations, External Links: Link Cited by: §1, §2, §2, §4.1, §4.2.
  • [15] I. J. Goodfellow, J. Shlens, and C. Szegedy (2014) Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572. Cited by: §2, §4.2.
  • [16] J. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al. (2020) Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems 33, pp. 21271–21284. Cited by: §0.B.1, §4.1, §4.2.
  • [17] H. He, K. Zha, and D. Katabi (2023) Indiscriminate poisoning attacks on unsupervised contrastive learning. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: §1, §1, §2, §2, §4.1, §4.1.
  • [18] K. He, X. Zhang, S. Ren, and J. Sun (2016) Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778. Cited by: §4.1.
  • [19] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger (2017) Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700–4708. Cited by: §4.1.
  • [20] H. Huang, X. Ma, S. M. Erfani, J. Bailey, and Y. Wang (2021) Unlearnable examples: making personal data unexploitable. arXiv preprint arXiv:2101.04898. Cited by: §1, §2, §2, §4.1.
  • [21] A. Ilyas, S. Santurkar, D. Tsipras, L. Engstrom, B. Tran, and A. Madry (2019) Adversarial examples are not bugs, they are features. Advances in neural information processing systems 32. Cited by: §2.
  • [22] M. Jagielski, A. Oprea, B. Biggio, C. Liu, C. Nita-Rotaru, and B. Li (2018) Manipulating machine learning: poisoning attacks and countermeasures for regression learning. In 2018 IEEE symposium on security and privacy (SP), pp. 19–35. Cited by: §2.
  • [23] A. Krizhevsky G. Hinton et al. (2009) Learning multiple layers of features from tiny images. Cited by: Figure 2, Figure 2, §1, §4.1.
  • [24] Z. Liu, Z. Zhao, and M. Larson (2023) Image shortcut squeezing: countering perturbative availability poisons with compression. arXiv preprint arXiv:2301.13838. Cited by: §1, §2, §4.2.
  • [25] A. Mittal, A. K. Moorthy, and A. C. Bovik (2012) No-reference image quality assessment in the spatial domain. IEEE Transactions on image processing 21 (12), pp. 4695–4708. Cited by: §4.4.
  • [26] T. Qin, X. Gao, J. Zhao, K. Ye, and C. Xu (2023) Learning the unlearnable: adversarial augmentations suppress unlearnable example attacks. External Links: 2303.15127, Link Cited by: §0.C.2.
  • [27] V. S. Sadasivan, M. Soltanolkotabi, and S. Feizi (2023) Cuda: convolution-based unlearnable datasets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3862–3871. Cited by: §1, §1, §2, §2, §3.2, §4.1.
  • [28] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L. Chen (2018) Mobilenetv2: inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4510–4520. Cited by: §4.1.
  • [29] P. Sandoval-Segura, V. Singla, J. Geiping, M. Goldblum, T. Goldstein, and D. Jacobs (2022) Autoregressive perturbations for data poisoning. Advances in Neural Information Processing Systems 35, pp. 27374–27386. Cited by: §1, §2, §2, §2, §4.1.
  • [30] S. Shan, J. Cryan, E. Wenger, H. Zheng, R. Hanocka, and B. Y. Zhao (2023) Glaze: protecting artists from style mimicry by text-to-image models. External Links: 2302.04222, Link Cited by: §4.3.
  • [31] S. Shan, W. Ding, J. Passananti, S. Wu, H. Zheng, and B. Y. Zhao (2024) Nightshade: prompt-specific poisoning attacks on text-to-image generative models. External Links: 2310.13828, Link Cited by: §4.3.
  • [32] K. Simonyan and A. Zisserman (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556. Cited by: §4.1.
  • [33] L. Tao, L. Feng, J. Yi, S. Huang, and S. Chen (2021) Better safe than sorry: preventing delusive adversaries with adversarial training. Advances in Neural Information Processing Systems 34, pp. 16209–16225. Cited by: §1, §1, §2, §2.
  • [34] J. Vincent (2023)Getty images sues ai art generator stable diffusion in the us for copyright infringement(Website) External Links: Link Cited by: §1.
  • [35] J. Wang, K. C. Chan, and C. C. Loy (2023) Exploring clip for assessing the look and feel of images. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, pp. 2555–2563. Cited by: §4.4.
  • [36] X. Wang, S. Hu, M. Li, Z. Yu, Z. Zhou, and L. Y. Zhang (2024) Corrupting convolution-based unlearnable datasets with pixel-based image transformations. External Links: 2311.18403, Link Cited by: §0.C.2.
  • [37] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli (2004) Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13 (4), pp. 600–612. Cited by: §4.4.
  • [38] Z. Wang, E. P. Simoncelli, and A. C. Bovik (2003) Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, Vol. 2, pp. 1398–1402. Cited by: §4.4.
  • [39] R. Wen, Z. Zhao, Z. Liu, M. Backes, T. Wang, and Y. Zhang (2023) Is adversarial training really a silver bullet for mitigating data poisoning?. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: §2.
  • [40] S. Wu, S. Chen, C. Xie, and X. Huang (2023) One-pixel shortcut: on the learning preference of deep neural networks. In The Eleventh International Conference on Learning Representations, Cited by: §1, §2, §2, §4.1.
  • [41] D. Yu, H. Zhang, W. Chen, J. Yin, and T. Liu (2022) Availability attacks create shortcuts. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 2367–2376. Cited by: §1, §2, §2, §4.1.
  • [42] C. Yuan and S. Wu (2021) Neural tangent generalization attacks. In International Conference on Machine Learning, pp. 12230–12240. Cited by: §1, §2, §2.
  • [43] S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo (2019) Cutmix: regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 6023–6032. Cited by: §1, §2.
  • [44] H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz (2018) Mixup: beyond empirical risk minimization. In International Conference on Learning Representations, Cited by: §1, §2.
  • [45] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang (2018) The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595. Cited by: §4.4.

Appendix 0.A Proofs

0.A.1 Proof of Theorem 1

Lemma 1

Given a random variable, XX follows a uniform distribution with support U⁡(0,pu)U(0,p_{u}). The p.d.f. of X2X^{2} is 12​pu​x\frac{1}{2p_{u}\sqrt{x}} and E⁡(X2)=pu23E(X^{2})=\frac{p_{u}^{2}}{3}

Proof

Considering

Pr⁡(X2<x)=Pr⁡(X<x)=∫0x1pu​𝑑x=1pu​x,\Pr(X^{2}<x)=\Pr(X<\sqrt{x})=\int_{0}^{\sqrt{x}}\frac{1}{p_{u}}\,dx=\frac{1}{p_{u}}\sqrt{x}, (7)

which is the c.d.f. of X2X^{2}. To obtain the p.d.f. of X2X^{2}, we compute dd​x​(1pu​x)=12​pu​x\frac{d}{dx}\left(\frac{1}{p_{u}}\sqrt{x}\right)=\frac{1}{2p_{u}\sqrt{x}}. Note that the support of X2X^{2} is [0,pu2][0,p_{u}^{2}] not [0,pu][0,p_{u}]. To obtain

E⁡(X2)=∫0pu212​pu​x​x​𝑑x=12​pu​23​x3/2|0pu2=13​pu​pu3=pu23E(X^{2})=\int_{0}^{p_{u}^{2}}\frac{1}{2p_{u}\sqrt{x}}\,x\,dx=\frac{1}{2p_{u}}\frac{2}{3}x^{3/2}\bigg|_{0}^{p_{u}^{2}}=\frac{1}{3p_{u}}p_{u}^{3}=\frac{p_{u}^{2}}{3} (8)

q.e.d.

Theorem 0.A.1

Given two different filters, RiR_{i} and RcR_{c}, generated by CUDA, if the condition 1>(κ2−2)​pu2+2​pu1>(\kappa^{2}-2)p_{u}^{2}+2p_{u} is satisfied, Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} and Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} have the following properties.

  1. 1.

    The peak of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} is located at the center.

  2. 2.

    The peak of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} occurs at the position where two ones in Rc​πR_{c\pi} and Ri​πR_{i\pi} appear at the same position in Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi}.

  3. 3.

    The probability of peak of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} at the center is 1κ2\frac{1}{\kappa^{2}}.

  4. 4.

    The expected peak of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} is higher than the expected peak of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi}.

Proof

Property 1: The center of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} is at (κ,κ)(\kappa,\kappa), where two filters completely overlap, and its minimum value is larger than 1. The maximum value at other positions is less than (κ2−κ−2)​pu2+2​pu(\kappa^{2}-\kappa-2)p_{u}^{2}+2p_{u}. Using the assumption, we can deduce that Rc​π⋆Rc​π​(κ,κ)>Rc​π⋆Rc​π​(i,j)R_{c\pi}\star R_{c\pi}(\kappa,\kappa)>R_{c\pi}\star R_{c\pi}(i,j), for all (i,j)≠(κ,κ)(i,j)\neq(\kappa,\kappa).

Property 2: When the one in Rc​πR_{c\pi} and the one in Rj​πR_{j\pi} appear at the same position in Rc​π⋆Rj​πR_{c\pi}\star R_{j\pi}, the minimum value of Rc​π⋆Rj​πR_{c\pi}\star R_{j\pi} at the location is larger than one. The maximum value of other positions is less than (κ2−2)​pu2+2​pu(\kappa^{2}-2)p_{u}^{2}+2p_{u}. Based on the assumption, we achieve property 2.

Property 3: Based on property 2, the peak of Rc​π⋆Rj​πR_{c\pi}\star R_{j\pi} is at the position where the one in Rc​πR_{c\pi} and the one in Rj​πR_{j\pi} appear at the same position in Rc​π⋆Rj​πR_{c\pi}\star R_{j\pi}. If the peak is at the center of Rc​π⋆Rj​πR_{c\pi}\star R_{j\pi}, it means that the one in Rc​πR_{c\pi} and the one in Rj​πR_{j\pi} must be at the position. The probability that they are in the position is 1κ2\frac{1}{\kappa^{2}}.

Property 4: Utilizing Lemma 1 and the preceding properties, we can derive the expected peak value of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi}:

E⁡(Rc​π⋆Rc​π​(κ,κ))=1+(κ2−1)​(pu23).E(R_{c\pi}\star R_{c\pi}(\kappa,\kappa))=1+(\kappa^{2}-1)\left(\frac{p_{u}^{2}}{3}\right). (9)

When two ones are at the same position in Rc​πR_{c\pi} and Ri​πR_{i\pi}, Rc​π⋆Ri​π​(κ,κ)R_{c\pi}\star R_{i\pi}(\kappa,\kappa) achieves its maximum value. Thus the expected maximum peak value of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} is

E⁡(Rc​π⋆Ri​π​(κ,κ))=1+(κ2−1)​(pu24).E(R_{c\pi}\star R_{i\pi}(\kappa,\kappa))=1+(\kappa^{2}-1)\left(\frac{p_{u}^{2}}{4}\right). (10)

0.A.2 Proof of Theorem 2

Theorem 0.A.2

Given a CUDA filter size being larger than 1 and xc​i​jx_{cij} being the vector form of the patch in Xc​iX_{ci} involved in producing ηc​i​j\eta_{cij}, i.e., ηc​i​j=℘​xc​i​j\eta_{cij}=\wp x_{cij} , where ℘\wp is a κ2×(2​κ−1)2\kappa^{2}\times(2\kappa-1)^{2} matrix, if the column vectors of ℘\wp, excluding the one corresponding to yc​i​jy_{cij}, span ℝκ2\mathbb{R}^{\kappa^{2}}, then, ∃xc​i​j∈ℝ(2​κ−1)2\exists x_{cij}\in\mathbb{R}^{(2\kappa-1)^{2}} such that yc​i​j≠acT​℘​xc​i​jy_{cij}\neq a_{c}^{T}\wp x_{cij}.

Proof

Let the patch in Xc​iX_{ci} involved in computing the patch (Xc​i⋆Rc)j(X_{ci}\star R_{c})_{j} be (Xc​i)j(X_{ci})_{j}. Since the patch size of (Xc​i⋆Rc)j(X_{ci}\star R_{c})_{j} and the size of RcR_{c} are both κ×κ\kappa\times\kappa, ((Xc​i)j)((X_{ci})_{j}) has the size of (2​κ−1)×(2​κ−1)(2\kappa-1)\times(2\kappa-1). When κ>1\kappa>1, (Xc​i)j(X_{ci})_{j} is larger than (Xc​i⋆Rc)j(X_{ci}\star R_{c})_{j}. Reshaping (Xc​i)j(X_{ci})_{j} and (Xc​i⋆Rc)j(X_{ci}\star R_{c})_{j} as vectors, xc​i​jx_{cij} and ηc​i​j\eta_{cij}, their relationship can be represented as a linear equation, i.e., ηc​i​j=℘​xc​i​j\eta_{cij}=\wp x_{cij}, where ℘\wp is a κ2×(2​κ−1)2\kappa^{2}\times(2\kappa-1)^{2} matrix. The matrix ℘\wp is constructed by RcR_{c}. To achieve perfect reconstruction yc​i​j=(αcT​℘)​xc​i​jy_{cij}=(\alpha_{c}^{T}\wp)x_{cij} for all xc​i​j∈ℝ(2​κ−1)2x_{cij}\in\mathbb{R}^{(2\kappa-1)^{2}}, αcT​℘\alpha_{c}^{T}\wp must be [0⋯0 1 0⋯0][0\cdots 0\,1\,0\cdots 0], where the location of the one corresponds to yc​i​jy_{cij} in (Xc​i)j(X_{ci})_{j}. Since the column vectors of ℘\wp, excluding the one corresponding to yc​i​jy_{cij}, span ℝκ2\mathbb{R}^{\kappa^{2}}, αc\alpha_{c} cannot be orthogonal to all these column vectors. Thus, αcT℘=[0⋯0 1 0⋯0]\alpha_{c}^{T}\wp=[0\cdots 0\,1\,0\cdots 0] cannot be obtained.

0.A.3 Extending CUDA to CIFAR-10

Since CUDA uses pu=0.3p_{u}=0.3 and κ=3\kappa=3 for CIFAR-10, it fulfils the condition 1>14​(κ2−2)​pu2+pu1>\frac{1}{4}(\kappa^{2}-2)p_{u}^{2}+p_{u}. According to Theorem 0.A.3, CUDA filters exhibit similar properties shown in Theorem 0.A.1 on CIFAR-10.

Theorem 0.A.3

Given two different filters, RiR_{i} and RcR_{c} generated by CUDA, if the condition 1>14​(κ2−2)​pu2+pu1>\frac{1}{4}(\kappa^{2}-2)p_{u}^{2}+p_{u} is satisfied, Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} and Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} have the following properties.

  1. 1.

    The expected peak of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} is located at the center.

  2. 2.

    The expected peak of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} occurs at the position where two ones in Rc​πR_{c\pi} and Ri​πR_{i\pi} appear at the same position in Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi}.

  3. 3.

    The probability of the expected peak of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi} at the centre is 1k2\frac{1}{k^{2}}.

  4. 4.

    The expected peak of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} is higher than the expected peak of Rc​π⋆Ri​πR_{c\pi}\star R_{i\pi}.

Proof

Property 1: The center of Rc​π⋆Rc​πR_{c\pi}\star R_{c\pi} is at (κ,κ)(\kappa,\kappa), where two filters completely overlap, and its expected value is

E⁡(Rc​π⋆Rc​π​(κ,κ))=1+13​(κ2−1)​pu2.E(R_{c\pi}\star R_{c\pi}(\kappa,\kappa))=1+\frac{1}{3}(\kappa^{2}-1)p_{u}^{2}. (11)

Let ω\omega be the number of pixels overlapped when computing E⁡(Rc​π⋆Rc​π​(i,j))E(R_{c\pi}\star R_{c\pi}(i,j)), where (i,j)≠(κ,κ)(i,j)\neq(\kappa,\kappa). We have the following cases:

Case 1: If both ones are not inside the overlapping region,

E⁡(Rc​π⋆Rc​π​(i,j))=14​ω​pu2E(R_{c\pi}\star R_{c\pi}(i,j))=\frac{1}{4}\omega p_{u}^{2}

Case 2: If both ones are inside the overlapping region but at different locations,

E⁡(Rc​π⋆Rc​π​(i,j))=ρu+14​(ω−2)​pu2E(R_{c\pi}\star R_{c\pi}(i,j))=\rho_{u}+\frac{1}{4}(\omega-2)p_{u}^{2}

Case 3: If both ones are inside the overlapping region and at the same location,

E⁡(Rc​π⋆Rc​π​(i,j))=1+14​(ω−1)​pu2E(R_{c\pi}\star R_{c\pi}(i,j))=1+\frac{1}{4}(\omega-1)p_{u}^{2}

Case 4: If only one is inside the overlapping region,

E⁡(Rc​π⋆Rc​π​(i,j))=12​ρu+14​(ω−1)​pu2E(R_{c\pi}\star R_{c\pi}(i,j))=\frac{1}{2}\rho_{u}+\frac{1}{4}(\omega-1)p_{u}^{2}

Among the four cases, the third case is the largest one, since ρu<1\rho_{u}<1. Clearly E⁡(Rc​π⋆Rc​π​(κ,κ))E(R_{c\pi}\star R_{c\pi}(\kappa,\kappa)) is larger than E⁡(Rc​π⋆Rc​π​(i,j))E(R_{c\pi}\star R_{c\pi}(i,j)) in the third case. Thus, the expected peak of E⁡(Rc​π⋆Rc​π​(κ,κ))E(R_{c\pi}\star R_{c\pi}(\kappa,\kappa)) is at the center.

Property 2: When two ones overlap, the minimum value of E⁡(Rc​π⋆Ri​π​(i,j))E(R_{c\pi}\star R_{i\pi}(i,j)) is larger than 1. The expected maximum value of E⁡(Rc​π⋆Ri​π​(i,j))E(R_{c\pi}\star R_{i\pi}(i,j)) when two ones do not overlap is

E⁡(Rc​π⋆Ri​π​(i,j))=ρu+14​(κ2−2)​pu2.E(R_{c\pi}\star R_{i\pi}(i,j))=\rho_{u}+\frac{1}{4}(\kappa^{2}-2)p_{u}^{2}.

Using the assumption, the expected peak appears at the position when two ones overlap.

Property 3: Based on property 2, the expected peak appears at the position when two ones are at the same position. Thus, the probability of the expected peak appearing at the center is 1κ2\frac{1}{\kappa^{2}}

Property 4: The proof is the same as the proof of property 4 in Theorem 0.A.1.

Appendix 0.B Experiment Setup

0.B.1 Training details

Table 7: Training settings for SL and SSL. For cells with two values, the left value is for CIFAR-10 and the right value is for ImageNet-100. ‘Pretrain’ is the standard augmentation utilized in SSL pretraining, while ‘LinProbe’ is the standard augmentation employed in SSL linear probing.
Setting SSL Pretrain SSL Linear Probing SL
Augmentations Pretrain LinProbe LinProbe
Epochs 400 | 200 100 100
Batch Size 512 | 128 512 256
LR x BatchSize/256 1.0 | 0.25 1.0 | 10.0 0.1
LR Warmup 10 0 0
LR Decay Cosine Step (0.2x at 60,75,90) Cosine
Optimizer SGD SGD SGD
Momentum 0.9 0.9 0.9
Weight Decay 1e-4 0 5e-4

The training details are provided in Table 7. However, there are two exceptions: for VGG-19 and ViT models, the learning rate in SL is set to 0.01, and the training epoch for ViT is adjusted to 200. For the SL training, we utilize the standard linear probing augmentation of SSL as it has demonstrated an enhancement in clean test accuracy. This implies that the linear probing augmentation acts as an effective defense against certain DAAs. Specifically, in ImageNet training, the data augmentation comprises RandomResizedCrop with an expected output size of 224 along with RandomHorizontalFlip. For CIFAR-10 training, RandomCrop with an output size of 32 and padding set to 4 is employed, alongside RandomHorizontalFlip. For SSL pretraining and linear probing, we follow the standard augmentations used in [4, 5, 6, 16].

Appendix 0.C Additional Results

0.C.1 Images with Double Blur and Random Sharpness

The examples presented in Figure 7 depict outputs generated through Double Blur, Random Sharpness, and IRP. The images produced by Double Blur exhibit the highest level of blurriness. The degree of blurriness in the images generated by Random Sharpness varies, as the sharpness filters are randomly generated class-wise. Conversely, the images produced by IRP are the clearest, owing to the optimization process customized for the blurriness filter in IRP.

Refer to caption
Figure 7: Example images generated by Double Blur (the first row), Random Sharpness (the second row), and IRP (the third row) on ImageNet-100.

0.C.2 Additional Defenses

We further evaluate IRP on three recently invented poison defenses. AVATAR [10] adds Gaussian noise to poisoned images and ‘purifies’ them with a pretrained diffusion model. UEraser [26] applies multiple augmentations to poisoned images and selects the one that maximizes loss. COIN [36] applies random pixel interpolations to reduce the impact of filter-based poisons.

Figure 8 shows that AVATAR and COIN are useful defenses against both CUDA and IRP poisons. Under COIN purification, CUDA-poisoned CIFAR-10 permits a clean test accuracy of 71.90%. AVATAR raises the max accuracy of IRP-poisoned CIFAR-10 to 54.78%, though this value is still severely low compared to training on clean data. UEraser is less effective as a defense than standard methods like SSL and AT. Across all defenses, IRP is a more potent DAA than CUDA.

Table 8: Top-1 clean test accuracies (%) of ResNet-18 trained with data poisoned by various DAAs on CIFAR-10. The “Max” column represents the highest clean test accuracy, which corresponds to the worst performance of each DAA under different training scenarios. Bold indicates the best performance.
DAA SL SSL AT AVATAR UEraser COIN Max
Clean 93.86 90.55 89.57 - - - 93.86
CUDA 21.89 66.58 48.58 55.93 40.56 71.90 71.90
IRP 10.39 43.24 32.21 54.78 26.55 53.32 54.78