Learning Normal Diffusion Dynamics for
Backdoor Defense in Text-to-Image Models
Abstract
Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this paper, we study backdoor defense of T2I diffusion models from a transition-dynamics perspective. We observe that benign diffusion trajectories exhibit structured and timestep-dependent transition patterns from cross-attention, latent and noise spaces, whereas backdoor attacks tend to induce deviations from such normal evolution. Motivated by these observations, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework that learns the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL constructs compact multi-space trajectory representations and trains a timestep-conditioned dynamics model to predict the diffusion evolution. In the inference phase, deviations between the observed and predicted transitions are exploited to quantify dynamics inconsistency for backdoor detection. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.
1 INTRODUCTION
Text-to-image (T2I) diffusion models have achieved remarkable success in high-quality image synthesis, facilitating widespread real-world applications (Rombach et al., 2022b; Huang et al., 2023; Wang et al., 2025; Liu et al., 2026). However, the increasing prevalence of publicly available models introduces substantial security risks (Qu et al., 2023; Truong et al., 2025; Liu et al., 2025; Yan et al., 2025; Gao et al., 2026b; Chen et al., 2026). In particular, backdoor attacks can embed hidden behaviors in T2I diffusion models (Chou et al., 2023a; Chou et al., 2023b). Thus, a backdoored model behaves normally on benign prompts while generating the attacker-specified contents once the trigger is present (Zhai et al., 2023; Wang et al., 2024a; Lyu et al., 2026; Struppek et al., 2023). Since the downstream users typically have no prior knowledge of the attack mechanism, the reliable backdoor defense is essential for the secure deployment of T2I diffusion models (Zhang et al., 2025; Zhang et al., 2026).
Existing defense strategies typically exploit abnormal behaviors induced by backdoors, including distinct patterns in attention, noise prediction, neuron activations (Wang et al., 2024b; Mo et al., 2024; Zhai et al., 2025). While these methods indicate that backdoor attacks can leave detectable traces, the resulting detection criteria are usually associated with particular representations or abnormal phenomena. This presents a fundamental challenge for general backdoor defense. Since different types of attacks may rely on different mechanisms or objectives, the internal activation patterns of backdoored models vary across representation spaces and denoising stages (Zhai et al., 2025; Pan et al., 2026). Consequently, the defenses based on attack-specific characteristics limit the generalizability with the emergence of increasingly diverse backdoor attacks.
We therefore reconsider the defense problem from a different perspective: Can we shift the focus from how backdoor behaviors appear abnormal to how benign diffusion normally behaves? This perspective is natural for T2I diffusion models, whose generation process is based on a sequence of timestep-dependent denoising transitions (Ho et al., 2020; Li et al., 2024; Rombach et al., 2022b; Xu et al., 2023a; Xu et al., 2023b; Song et al., 2021).
Instead of inducing an obvious anomaly at the particular timestep, backdoor may perturb the denoising transitions and gradually deviate the generation from the normal evolution. To this end, we analyze the model diffusion trajectories from cross-attention, latent and noise spaces. Our empirical study reveals two key observations: (1) Observation I: As shown in Figure 1, benign trajectories present structured and timestep-dependent transition patterns; (2) Observation II: As shown in Figure 2, backdoors consistently introduce deviations to the normal evolution across different attacks. These observations suggest that benign diffusion exhibits learnable transition rule, while backdoor tends to introduce perturbations. Motivated by the above analysis, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework for learning the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL first constructs compact trajectory representations from cross-attention, latent and noise spaces. Then, a dynamics model is trained on benign trajectories to predict the next-step representation. During inference, the deviations between the observed and predicted transitions are utilized to quantify dynamics inconsistency, with large deviations indicating potential backdoor attack. Furthermore, NDDL can localize the potential trigger tokens by performing substitution with low-semantic words, enabling trigger identification without requiring prior knowledge of the embedded backdoor.
Our contributions are summarized as follows:
- •
We introduce a new transition-dynamics perspective for backdoor attacks in T2I diffusion models, showing that trigger effects can be viewed as the deviations from the normal evolution of diffusion trajectories rather than the anomalies in the particular spaces.
- •
We propose a backdoor defense framework NDDL for learning normal diffusion dynamics, leveraging compact multi-space trajectory representations for backdoor detection and localizing suspicious trigger tokens performing substitution with low-semantic words.
- •
We conduct extensive experiments against diverse backdoor attacks, demonstrating that NDDL achieves effective and generalizable performance in both backdoor detection and trigger localization compared with existing defense methods.
2 Related Work
Backdoor attacks in T2I diffusion models. Backdoor attacks on deep neural networks inject triggers into inputs to hijack model behavior (Doan et al., 2021; Li et al., 2022; Khaddaj et al., 2023; Gu et al., 2017; Zhao et al., 2022; Gao et al., 2026a). Recently, such attacks have extended to generative models, particularly T2I diffusion models. Rickrolling (Struppek et al., 2023) aligns the feature representations of backdoor and target prompts within the text embedding space while preserving the original feature embedding of benign samples. BadT2I (Zhai et al., 2023) pioneers prompt-based backdoor attacks through data poisoning. By leveraging a regularization loss, T2I diffusion models can be efficiently backdoored with only a few fine-tuning steps. EvilEdit (Wang et al., 2024a) directly edits the projection matrices in the cross-attention layers to achieve projection alignment between a trigger and the corresponding backdoor target. MasqLoRA (Lyu et al., 2026) leverages an independent LoRA module as the attack vehicle to stealthily inject malicious behavior into T2I diffusion models. STEBA (Pan et al., 2026) proposes a spatio-temporally acceleration strategy for backdoor injection, improving computational efficiency and reducing memory overhead.
Backdoor defenses in T2I diffusion models. Backdoor defense mechanisms in discriminative models often rely on input perturbations or behavioral monitoring (Wang et al., 2019; Zhu et al., 2023; Wei et al., 2024; Yu et al., 2025), and similar ideas have been extended to T2I diffusion models. T2IShield (Wang et al., 2024b) identifies backdoor behavior via assimilation patterns in cross-attention maps. UFID (Guan et al., 2025) is a black-box method, using image level similarity to separate benign from backdoored outputs without internal access. NaviT2I (Zhai et al., 2025) navigates T2I diffusion models to prevent malicious inputs by analyzing neuron activation variations caused by input tokens. STEDF (Pan et al., 2026) formulates backdoor detection as a spatio-temporal feature analysis problem, exploiting weight enrichment patterns and temporal anisotropy to distinguish malicious models from benign ones. However, existing backdoor defenses for T2I diffusion models primarily rely on detecting specific abnormal signatures associated with backdoor activation. These approaches often struggle to generalize against increasingly diverse attack mechanisms due to the lack of a unified characterization of normal generation processes. By learning normal transition dynamics exclusively from benign samples, NDDL does not rely on specific attack artifacts, thereby offering a more generalizable and principled defense paradigm against unknown and diverse backdoor threats.
3 Transition Dynamics Analysis
3.1 Transition Dynamics Formulation
Let denote the diffusion state at timestep . Rather than directly modeling the raw high-dimensional diffusion states, we map them into a compact trajectory representation , where is the representation mapping. Motivated by Observation I, from a dynamical-system perspective, we model benign diffusion evolution as a timestep-conditioned transition process , where represents the normal transition rule at timestep . Observation II demonstrates that backdoor trajectories present obvious deviations from their benign counterparts for various backdoor attacks. Thus, we model the backdoor diffusion evolution as , where represents the mapping trajectory representation related to the backdoor attack and is the perturbation induced by backdoor. Thus, we obtain:
Hypothesis 1 (Transition Dynamics Deviation Hypothesis).
Backdoor attacks perturb the normal transition dynamics of the diffusion evolution, yielding the trajectory that deviates from the normal one.
This hypothesis provides a unified perspective on the deviations observed from the three representations. Based on it, we aim to investigate two questions: (i) How does the backdoor perturbation affect the trajectory? (ii) Can backdoor-induced deviations be revealed by normal transition prediction?
3.2 How Does The Backdoor Perturbation Affect the Trajectory?
To investigate how a transition perturbation affects the subsequent trajectory, we first present a mild regularization assumption on normal diffusion dynamics.
Assumption 1 (Time-Conditioned Local Smoothness).
For each timestep , the normal transition rule is locally Lipschitz around normal diffusion trajectory:
| (1) |
where varies with timestep.
Assumption 1 does not require the diffusion trajectory to evolve smoothly over time. Instead, it emphasizes that the adjacent states at the same timestep exhibit the locally bounded differences after transition. Then, we characterize how the backdoor-induced transition perturbation affects subsequent trajectory.
Proposition 1 (Propagation of Trigger-Induced Deviations).
Let denote the difference between the backdoor and benign trajectories. From Assumption 1, can be obtained. Recursively, for any , there is:
| (2) |
3.3 Can Backdoor-Induced Deviations Be Revealed by Normal Transition Prediction?
Furthermore, we investigate whether the backdoor-induced transition deviations can be detected through normal transition predictions. Given a predictor from benign trajectories to approximate the normal transition rule , the approximation error can be bounded as:
| (3) |
where is the prediction error at timestep .
Next, we elaborate the correlation between backdoor-induced perturbation and the transition prediction error.
Proposition 2 (Backdoor-Induced Prediction Inconsistency).
and are the transition prediction errors of benign and backdoor trajectories, respectively. Then, the following two bounds hold:
| (4) |
Proposition 2 shows that backdoor-induced perturbation is reflected in the inconsistency between the perturbed transitions and the learned normal dynamics. In particular, perturbation that substantially exceeds the normal prediction error become more distinguishable from the benign transition. This motivates us to employ the transition prediction errors as the criterion for backdoor defense. The proof of Proposition 2 is provided in Appendix B.2.
4 Method
4.1 Threat Model
Scenario and defender capability. We consider a realistic deployment scenario where T2I diffusion models are obtained from potentially untrusted third-party providers. An adversary implants a hidden backdoor into the model and distributes it as a seemingly benign model, while the downstream user remains unaware of the compromise. The defender is assumed to have white-box access to the deployed model but no prior knowledge of the embedded backdoor. A limited set of benign prompts is available and utilized to model the normal diffusion dynamics.
Defense goals. Our defense includes two objectives: (1) Detection: distinguish the backdoor prompts from the benign ones; (2) Localization: identify the tokens that induce the backdoor behaviors.
4.2 The details of NDDL
We propose a defense framework NDDL that views the backdoor behaviors as the deviations of the normal diffusion evolution. The overview of NDDL is illustrated in Figure 3. NDDL first constructs a compact multi-space representation of the diffusion trajectory and then learns the normal transition dynamics using only benign samples. During inference, NDDL identifies backdoor prompts through transition prediction inconsistency, while localizing the suspicious tokens based on the anomaly-score reduction induced by low-semantic token substitution. The pseudocode can be found in Appendix C.5.
4.2.1 Stage I: Multi-Space Trajectory Representation
Rather than directly modeling the raw states , we construct the compact representation that summarize the structural and temporal dynamics.
Cross-attention representation. For the cross-attention weight , we extract four descriptors including attention entropy , effective rank , token importance and head diversity . Attention entropy presents the concentration of token-wise attention distribution, effective rank captures the structural complexity of attention maps, token importance shows the relative contribution of text tokens and head diversity measures variation among different attention heads. Notably, we extract the cross-attention representation of the earliest cross-attention layer along the forward pass.
Latent representation. For the latent state , we extract the descriptors depicting the instantaneous structure and local temporal evolution. Specially, we consider channel norm , trajectory curvature , frequency-domain energy and temporal variation . These descriptors elucidate the geometric and spectral dynamics of the latent representation across the diffusion process.
Noise representation. Similarly, we construct utilizing channel norm , channel variance , frequency-domain energy and temporal variation . These features characterize both the distribution properties of the predicted noises and the temporal evolution across diffusion process.
Notably, considering the heterogeneous scales of different trajectory descriptors, we apply robust normalization based on benign training statistics followed by block-wise scaling before feature concatenation. Implementation details of representation extractions and normalizations can be seen in Appendixes C.1 and C.2.
4.2.2 Stage II: Normal Diffusion Dynamics Learning
Section 3 suggests that backdoor attack can be viewed as the deviation from the normal diffusion evolution. Thus, we consider to model the evolution process of the benign trajectories. We define the transition increment of two consecutive trajectory representations as . Instead of directly predicting , we train a model to obtain this transition increment conditioned on timestep . The details of is presented in Appendix C.3. The prediction of the transition increment is denoted as . Then, we can obtain the prediction of the next trajectory representation:
| (5) |
Thus, the learned approximation of the normal transition rule is:
| (6) |
The predictor is trained only using trajectories from the benign prompts. Since the three representations exhibit different statistical characteristics, we employ a block weighted loss during model training:
| (7) |
where
| (8) |
, and are the balance weights. After training, serves as the approximation of the normal diffusion transition dynamics.
4.2.3 Stage III: Backdoor Detection
For an unseen prompt , we extract its compact trajectory and compute the transition inconsistency as:
| (9) |
To evaluate the dynamics consistency, we select a temporal interval of the denoising process and partition it into equal-length short windows . This strategy prevents the transition inconsistencies from being diluted by averaging the entire diffusion trajectory. For each window , we obtain:
| (10) |
We define the final anomaly score as and consider the prompt as suspicious when . can be utilized on benign validation trajectories, which is similar to the method in (Zhai et al., 2025). Specially, to achieve backdoor-agnostic thresholding, we perform a Gaussian fitting on the prediction errors of the benign validation trajectories, i.e., . Thus, can be set as:
| (11) |
where is a balance weight.
4.2.4 Stage IV: Trigger Localization
For a suspicious prompt , we further localize the tokens inducing backdoor behavior. A single predefined substitute may introduce replacement-dependent bias and interact with the embedded backdoor. Therefore, we adopt a corrected multi-substitution strategy, where multiple low-semantic words are selected based on their consistency over benign reference prompts.
Specially, we first collect an initial candidate set including low-semantic words, which are presented in Appendix C.4. Utilizing the benign reference prompts , we estimate the change introduced by each candidate replacement as:
| (12) |
where represents the prompt by replacing the token at position with the word . Candidates obtaining only small variations are preserved to form the corrected set .
Given a prompt , each token is replaced by the token . The score of the replacement is defined as:
| (13) |
We classify as a trigger token if . The setting of is the same as that of .
5 Experiments
5.1 Experimental settings
Attack and defense methods. We consider the following backdoor attack methods for T2I models: (1) BadT2I Zhai et al. (2023) with one token trigger ‘\u200b’ and the sentence trigger ‘I like this photo.’; (2) EvilEdit Wang et al. (2024a) with the trigger tokens ‘beautiful cat’; (3) MasqLoRA Lyu et al. (2026) with the trigger ‘cool car’; (4) RickRolling Struppek et al. (2023) with the special character ‘o (U+043E)’ as the trigger; (5) STEBA Pan et al. (2026) with the trigger ‘A Object:’. Four defense methods are considered as baselines: UFID Guan et al. (2025), T2IShield Wang et al. (2024b), NaviT2I Zhai et al. (2025) and STEDF Pan et al. (2026).
Dataset and models. We utilize DiffusionDB Wang et al. (2023) to sample the prompts in our experiments. For each attack, we sample 1,000 benign prompts and 1,000 backdoor prompts with the triggers. We conduct main experiments on Stable Diffusion v1.5 Rombach et al. (2022a), Stable Diffusion XL Blattmann et al. (2023). Moreover, we also validate our method on the diffusion transformer (DiT) based Stable Diffusion v3.5 Esser et al. (2024) and Pixart- Chen et al. (2024).
Evaluation metric. For the results of backdoor detection, we calculate the detection accuracy (ACC). Meanwhile, to eliminate the impact of varying thresholds, we also adopt the area under receiver operating curve (AUROC). We evaluate trigger localization using exact trigger recovery (ETR) and AUROC, where ETR measures the proportion of backdoor prompts that all trigger tokens are correctly identified.
5.2 Defense Results
We comprehensively evaluate NDDL across three critical dimensions: backdoor detection, trigger localization, and generalization to DiT architectures. Our results demonstrate that learning normal diffusion dynamics provides a unified, attack-agnostic defense mechanism that consistently outperforms existing methods.
| Method | BadT2I | EvilEdit | MasqLoRA | Rickrolling | STEBA | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| ACC | AUROC | ACC | AUROC | ACC | AUROC | ACC | AUROC | ACC | AUROC | |
| UFID | 71.5 | 71.4 | 62.5 | 62.2 | 68.1 | 68.2 | 62 | 61.7 | 52.7 | 51.9 |
| T2IShield | 84.8 | 84.6 | 85.3 | 85.8 | 81.3 | 81.9 | 85.5 | 86.1 | 73.3 | 73.5 |
| NaviT2I | 96.3 | 96.9 | 94.8 | 95.5 | 91.3 | 91.8 | 88 | 89.1 | 79.8 | 80.4 |
| STEDF | 99.3 | 99.4 | 98 | 98.1 | 97.5 | 97.7 | 96.2 | 96.6 | 89.3 | 89.6 |
| NDDL (Ours) | 99 | 99.2 | 98.5 | 98.6 | 98.2 | 98.3 | 98.2 | 98.1 | 96.8 | 97 |
| Method | One-token | Multi-token | Special-character | Sentence | ||||
|---|---|---|---|---|---|---|---|---|
| ETR | AUROC | ETR | AUROC | ETR | AUROC | ETR | AUROC | |
| T2IShield | 86.8 | 87.1 | 81.3 | 80.5 | 89.2 | 88.8 | 73.8 | 75.5 |
| NaviT2I | 96.8 | 96.3 | 96.8 | 96.5 | 93.3 | 93.5 | 81.5 | 81.8 |
| NDDL (Ours) | 98.8 | 99.1 | 98.5 | 98.9 | 96.0 | 95.6 | 89.2 | 88.8 |
Evaluation results on backdoor detection. Table 1 reports the performance of different defense methods in distinguishing benign and backdoor prompts on Stable Diffusion v1.5. Specifically, NDDL demonstrates excellent performance, maintaining high ACC and AUROC across diverse backdoor settings. NDDL achieves the best results on EvilEdit, MasqLoRA, Rickrolling, and STEBA, while remaining highly competitive on BadT2I. The improvement is particularly pronounced on STEBA, where NDDL increases ACC from 89.3 to 96.8 and AUROC from 89.6 to 97.0 compared with the best baseline. These results demonstrate that the deviations from normal diffusion dynamics provide the stable and attack-agnostic indicators for identifying the benign and backdoor samples. Detection results on Stable Diffusion XL are reported in Appendix D.1, where NDDL still presents the superior performance.
Evaluation results on trigger localization. To comprehensively evaluate trigger localization for diverse trigger forms, we implement one-token and sentence-level triggers with BadT2I, multi-token triggers with MasqLoRA and special-character triggers using Rickrolling. As shown in Table 2, NDDL outperforms existing localization methods for all four trigger forms. NDDL obtains high localization accuracy for both one-token and multiple-token triggers, while preserving robust performance on special-character triggers. Notably, NDDL demonstrates its greatest superiority on the more challenging sentence-level triggers, improving ETR and AUROC by 7.7 and 7.0 over NaviT2I. More results on different diffusion models are provided in Appendix D.2. Overall, these results demonstrate that NDDL can reliably localize triggers with diverse trigger forms, including long and structurally complex triggers.
Evaluation results on DiT structure model. Most existing backdoor attack and defense methods for T2I models focus on U-Net based structure. To demonstrate the generalizability of NDDL, we adopt Rickrolling on Stable Diffusion v3.5, which is based on DiT structure. As shown in Table 3, NDDL still presents the best defense performance, demonstrating its effectiveness and generalizability beyond U-Net based diffusion models. The more similar results on Pixart- also based on DiT structure are shown in Appendix D.3.
| Method | Detection | Localization | ||
|---|---|---|---|---|
| ACC | AUROC | ETR | AUROC | |
| UFID | 42.2 | 40.1 | – | – |
| NaviT2I | 83.8 | 82.5 | 75.5 | 74.3 |
| NDDL | 91.2 | 92.7 | 87.5 | 86.9 |
5.3 Ablation Study
We conduct ablation studies on Stable Diffusion v1.5 using BadT2I as the representative backdoor attack. Unless otherwise specified, all variants are evaluated under the same experimental setting.
Effect of multi-space representations. Table 4 systematically investigates the impact of varying diffusion representation spaces. While single-space representations exhibit constrained detection capacity due to their partial view of the data manifold, combining multiple representations substantially boosts performance. The full representation setting obtains the best results, indicating that the three diffusion spaces offer mutually complementary information. Specifically, while one space may predominantly capture semantic inconsistencies, others might reveal structural or noise-level anomalies, thereby enabling a more robust detection of backdoor-induced transition deviations.
| Variant | ACC | AUROC |
|---|---|---|
| Attention only | 63.5 | 62.3 |
| Latent only | 59.5 | 59.5 |
| Noise only | 58.7 | 62.9 |
| Attention + Noise | 84.5 | 86.9 |
| Latent + Noise | 73.6 | 70.9 |
| Attention + Latent | 89.8 | 88.8 |
| Full | 99 | 99.2 |
Effect of different sampling methods. We evaluate the performance of NDDL utilizing different sampling methods. In Table 5, we test the four representative methods. Results indicates that NDDL presents the consistent defense performance for the evaluated samplers, demonstrating its robustness and generality.
| Sample type | ACC | AUROC |
|---|---|---|
| DDIM | 99.0 | 99.2 |
| DDPM | 97.7 | 97 |
| DPM | 98.2 | 98.6 |
| PLMS | 98.6 | 98.5 |
Effect of denoising stages. We investigate the effect of different denoising stages for the detection results. For the 50 sampling steps, we divide them into three stages: early stage (1-15 steps), middle stage (16-30 steps) and late stage (31-50 steps). Moreover, the window length is set as 5. As can be seen from Figure 4, the middle stage presents the best performance, followed by the late stage, while the early stage performs the worst. The results indicate that the dynamics deviations induced by backdoor are not equally discriminative over the whole denoising process. The weak performance in the early stage is likely related to the insufficient exhibition of trigger effects. The large transition variations of benign trajectories in the late denoising stage may obscure backdoor-induced deviations to some extent, leading to slightly degraded detection performance. Overall, these results suggest that the middle denoising stage provide the most distinguishable dynamics for the detection performance.
Effect of normal dynamics modeling and window length: We investigate the contribution of normal dynamics modeling through two variants: (1) Raw Diff: directly using adjacent-step representation differences without learning a transition model; (2) w/o Time: retaining the dynamics predictor but removes timestep conditioning. As shown in Figure 5(a), directly using raw transition differences results in the worst performance. Learning normal dynamics improves defense performance, especially when timestep conditioning is incorporated. We also study the effect of window length by varying , ranging from single step to full trajectory. In Figure 5(b), utilizing the single step shows limited detection performance, while aggregating residuals of short windows improves the results. However, as the window becomes longer, performance gradually decreases. These results suggest that short-window aggregation better preserves stage-specific transition inconsistencies, while overly long windows dilute the anomalies and reduce the detection performance.
6 Conclusion
In this work, we study backdoor defense for T2I diffusion models from a transition-dynamics perspective. Rather than relying on attack-specific abnormal indicators, we view backdoor attacks as the deviations from the normal evolution of diffusion trajectories. Our empirical analysis shows that benign trajectories exhibit structured and timestep-dependent transition regularities, while backdoor attack induces deviations from such normal dynamics. Based on this observation, we propose NDDL, a novel backdoor defense framework that learns normal diffusion transitions from benign trajectories and detects backdoor prompts through dynamics inconsistency. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.
7 AI Use Statement
In this work, we used generative AI tools (specifically ChatGPT) for language editing and polishing to improve the readability and grammatical accuracy of the manuscript. We have not used generative AI tools for generating research ideas, conducting data analysis, or writing original technical content, and AI-assisted figure generation or code synthesis are not applicable to this work. Additionally, we used generative AI tools for refining sentence structure and word choice during the revision process. We have reviewed all AI-assisted work. Specifically, we manually verified every AI-suggested modification against our original draft to ensure that no scientific meaning was altered, hallucinated, or misrepresented; all edits were strictly limited to linguistic improvements and were approved by all authors. We take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI.
8 Ethics Statement
This work adheres to the ICLR Code of Ethics. In this study, no human subjects or animal experimentation was involved. All datasets used, were sourced in compliance with relevant usage guidelines, ensuring no violation of privacy. We have taken care to avoid any biases or discriminatory outcomes in our research process. No personally identifiable information was used, and no experiments were conducted that could raise privacy or security concerns. We are committed to maintaining transparency and integrity throughout the research process.
9 Reproducibility Statement
We have made every effort to ensure that the results presented in this paper are reproducible. All code and datasets have been made publicly available in an anonymous repository to facilitate replication and verification. The experimental setup, including training steps, model configurations, and hardware details, is described in detail in the paper. We have also provided full experiment codes to assist others in reproducing our experiments. Additionally, all datasets in this paper are publicly available, ensuring consistent and reproducible evaluation results. We believe these measures will enable other researchers to reproduce our work and further advance the field.
References
- Blattmann et al. (2023) Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22563–22575, 2023.
- Chen et al. (2024) Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-: Fast training of diffusion transformer for photorealistic text-to-image synthesis. In Proceedings of the International Conference on Learning Representations, 2024.
- Chen et al. (2026) Linghan Chen, Yudong Gao, Jiyao Wang, Kaiyan Ji, and Honglong Chen. Do system prompts leave behavioral fingerprints? a large-scale empirical study of clone detection via output similarity. arXiv preprint arXiv:2608.24461, 2026.
- Chou et al. (2023a) Sheng-Yen Chou, Pin-Yu Chen, and Tsung-Yi Ho. How to backdoor diffusion models? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4015–4024, 2023a.
- Chou et al. (2023b) Sheng-Yen Chou, Pin-Yu Chen, and Tsung-Yi Ho. Villandiffusion: A unified backdoor attack framework for diffusion models. Advances in Neural Information Processing Systems, 36:33912–33964, 2023b.
- Doan et al. (2021) Khoa Doan, Yingjie Lao, Weijie Zhao, and Ping Li. Lira: Learnable, imperceptible and robust backdoor attacks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11946–11956, 2021.
- Esser et al. (2024) Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. Scaling rectified flow transformers for high-resolution image synthesis. In Proceedings of the International Conference on Machine Learning, pp. 12606–12633, 2024.
- Gao et al. (2026a) Yudong Gao, Linghan Chen, Wenhan Wu, Mia Zhou, Jiyao Wang, Kaiyan Ji, Mingyu Guo, and Honglong Chen. Bit-flip attacks on vision-language-action models: Action-decoding architecture shapes the vulnerability. arXiv preprint arXiv:2608.15475, 2026a.
- Gao et al. (2026b) Yudong Gao, Qingyue Wang, Yuanyuan Yuan, Ruixuan Huang, Linghan Chen, Zimo Ji, and Shuai Wang. Pathmark: Protecting intellectual property of mixture-of-expert llms via path watermarks. arXiv preprint arXiv:2607.03688, 2026b.
- Gu et al. (2017) Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733, 2017.
- Guan et al. (2025) Zihan Guan, Mengxuan Hu, Sheng Li, and Anil Kumar Vullikanti. Ufid: A unified framework for black-box input-level backdoor detection on diffusion models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, 2025.
- Ho et al. (2020) Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
- Huang et al. (2023) Ziqi Huang, Kelvin CK Chan, Yuming Jiang, and Ziwei Liu. Collaborative diffusion for multi-modal face generation and editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6080–6090, 2023.
- Khaddaj et al. (2023) Alaa Khaddaj, Guillaume Leclerc, Aleksandar Makelov, Kristian Georgiev, Hadi Salman, Andrew Ilyas, and Aleksander Madry. Rethinking backdoor attacks. In Proceedings of the International Conference on Machine Learning, pp. 16216–16236, 2023.
- Li et al. (2024) Mingxiao Li, Tingyu Qu, Ruicong Yao, Wei Sun, and Marie-Francine Moens. Alleviating exposure bias in diffusion models through sampling with shifted time steps. In Proceedings of the International Conference on Learning Representations, 2024.
- Li et al. (2022) Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. Backdoor learning: A survey. IEEE Transactions on Neural Networks and Learning Systems, 35(1):5–22, 2022.
- Liu et al. (2025) Tong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang, Shuo Chen, Philip Torr, Vera Demberg, Volker Tresp, and Jindong Gu. Multimodal pragmatic jailbreak on text-to-image models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pp. 4681–4720, 2025.
- Liu et al. (2026) Xiaolong Liu, Junjian Li, Yuan Xiao, Jiaqi Deng, Dayong Ye, Tianqing Zhu, and Huan Huo. Dual inversion for text-to-image diffusion models: From both prompt and noise perspectives. arXiv preprint arXiv:2607.26735, 2026.
- Lyu et al. (2026) Liangwei Lyu, Jiaqi Xu, Jianwei Ding, and Qiyao Deng. When lora betrays: Backdooring text-to-image models by masquerading as benign adapters. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8577–8586, 2026.
- Mo et al. (2024) Yichuan Mo, Hui Huang, Mingjie Li, Ang Li, and Yisen Wang. Terd: a unified framework for safeguarding diffusion models against backdoors. In Proceedings of International Conference on Machine Learning, pp. 35892–35909, 2024.
- Pan et al. (2026) Yu Pan, Jiahao Chen, Lin Wang, Bingrong Dai, and Wenjie Wang. Stediff: Revealing the spatial and temporal redundancy of backdoor attacks in text-to-image diffusion models. In Proceedings of the International Conference on Learning Representations, 2026.
- Qu et al. (2023) Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, and Yang Zhang. Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models. In Proceedings of the ACM SIGSAC Conference on Computer and Communications Security, pp. 3403–3417, 2023.
- Rombach et al. (2022a) Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022a.
- Rombach et al. (2022b) Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10674–10685, 2022b.
- Song et al. (2021) Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In Proceedings of the International Conference on Learning Representations, 2021.
- Struppek et al. (2023) Lukas Struppek, Dominik Hintersdorf, and Kristian Kersting. Rickrolling the artist: Injecting backdoors into text encoders for text-to-image synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4584–4596, 2023.
- Truong et al. (2025) Vu Tuan Truong, Luan Ba Dang, and Long Bao Le. Attacks and defenses for generative diffusion models: A comprehensive survey. ACM Computing Surveys, 57(8):1–44, 2025.
- Wang et al. (2019) Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao. Neural cleanse: Identifying and mitigating backdoor attacks in neural networks. In Proceedings of the IEEE Symposium on Security and Privacy, pp. 707–723, 2019.
- Wang et al. (2024a) Hao Wang, Shangwei Guo, Jialing He, Kangjie Chen, Shudong Zhang, Tianwei Zhang, and Tao Xiang. Eviledit: Backdooring text-to-image diffusion models in one second. In Proceedings of the ACM International Conference on Multimedia, pp. 3657–3665, 2024a.
- Wang et al. (2025) Zhendong Wang, Jianmin Bao, Shuyang Gu, Dong Chen, Wengang Zhou, and Houqiang Li. Designdiffusion: High-quality text-to-design image generation with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20906–20915, 2025.
- Wang et al. (2024b) Zhongqi Wang, Jie Zhang, Shiguang Shan, and Xilin Chen. T2ishield: Defending against backdoors on text-to-image diffusion models. European Conference on Computer Vision, pp. 107–124, 2024b.
- Wang et al. (2023) Zijie J Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau. Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models. In Proceedings of the Annual Meeting of Association for Computational Linguistics, pp. 893–911, 2023.
- Wei et al. (2024) Shaokui Wei, Hongyuan Zha, and Baoyuan Wu. Mitigating backdoor attack by injecting proactive defensive backdoor. Advances in Neural Information Processing Systems, 37:80674–80705, 2024.
- Xu et al. (2023a) Jiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao, Ying Shan, Xiaohu Qie, and Shenghua Gao. Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20908–20918, 2023a.
- Xu et al. (2023b) Xingqian Xu, Zhangyang Wang, Gong Zhang, Kai Wang, and Humphrey Shi. Versatile diffusion: Text, images and variations all in one diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7754–7765, 2023b.
- Yan et al. (2025) Song Yan, Hui Wei, Jinlong Fei, Guoliang Yang, Zhengyu Zhao, and Zheng Wang. Universally unfiltered and unseen: Input-agnostic multimodal jailbreaks against text-to-image model safeguards. In Proceedings of the ACM International Conference on Multimedia, pp. 11279–11287, 2025.
- Yu et al. (2025) Jimiao Yu, Honglong Chen, Junjian Li, Linghan Chen, Yudong Gao, Weifeng Liu, and Lei Zhang. Black-box adversarial defense based on image decomposition and reconstruction. IEEE Transactions on Multimedia, 27:5909–5921, 2025.
- Zhai et al. (2023) Shengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu, Yuejian Fang, and Hang Su. Text-to-image diffusion models can be easily backdoored through multimodal data poisoning. In Proceedings of the ACM International Conference on Multimedia, pp. 1577–1587, 2023.
- Zhai et al. (2025) Shengfang Zhai, Jiajun Li, Yue Liu, Huanran Chen, Zhihua Tian, Wenjie Qu, Qingni Shen, Ruoxi Jia, Yinpeng Dong, and Jiaheng Zhang. Efficient input-level backdoor defense on text-to-image synthesis via neuron activation variation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15182–15193, 2025.
- Zhang et al. (2025) Chenyu Zhang, Mingwang Hu, Wenhui Li, and Lanjun Wang. Adversarial attacks and defenses on text-to-image diffusion models: A survey. Information Fusion, 114:102701, 2025.
- Zhang et al. (2026) Yi Zhang, Zhen Chen, Chih-Hong Cheng, Wenjie Ruan, Xiaowei Huang, Dezong Zhao, David Flynn, Siddartha Khastgir, and Xingyu Zhao. Trustworthy text-to-image diffusion models: A timely and focused survey. Information Fusion, pp. 104264, 2026.
- Zhao et al. (2022) Zhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong, Dakui Wang, and Kaitai Liang. Defeat: Deep hidden feature backdoor attacks by imperceptible perturbation and latent representation constraints. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15192–15201, 2022.
- Zhu et al. (2023) Mingli Zhu, Shaokui Wei, Li Shen, Yanbo Fan, and Baoyuan Wu. Enhancing fine-tuning based backdoor defense with sharpness-aware minimization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4443–4454, 2023.
Appendix A The Details of Observation Results on Diffusion Dynamics
A.1 Observation I
For the benign prompts, we measure the transition differences between the consecutive denoising steps: , where is JS divergence and represents mean squared error. To study the consistency of benign diffusion dynamics, we randomly sample 10 benign prompts and analyze the temporal evolution of each representation over the denoising process. Specially, for each prompt, we track the consecutive transition differences in cross-attention weights, latent and noise spaces. As shown in Figure 6, despite semantic differences for the sampled prompts, the overall evolution remains highly consistent in each representation. Then, for each representation, we compute the transition differences by averaging 1000 benign trajectories. As shown in Figure 1(a), the averaged curves present clear timestep-dependent evolution patterns for the three representations. To quantify this consistency, we compute the Pearson correlation and cosine similarity between each individual transition trajectory and the corresponding mean results. High similarities in Figure 1(b) indicate that, despite substantial semantic diversity, the benign prompts follow a common temporal evolution pattern. In addition, we evaluate the results of different models. As shown in Figures 7, 8 and 9, although different models are employed, similar evolution patterns can be observed. Overall, these results lead to our first empirical observation: Benign diffusion trajectories exhibit structured and time-dependent transition dynamics.
A.2 Observation II
We further investigate how backdoor triggers affect the evolution of diffusion trajectories. We compare the representation trajectories of benign and backdoor prompts, calculating their discrepancy at each denoising timestep: . For different backdoor attack methods, we randomly sample 500 prompts and their corresponding trigger-injected ones. As shown in Figure 2, the backdoor-related trajectories exhibit obvious discrepancies from their benign ones for the three representations. Although the temporal patterns of these discrepancies vary in different attack methods, the separation between benign and backdoor trajectories remains consistently observable over the diffusion process. To further validate Observation II, we provide additional results of different backdoor attacks. Specifically, we compare the diffusion trajectories of benign prompts with their corresponding backdoor ones in the cross-attention, latent and noise spaces. As shown in Figures 10, 11, 12, 13 and 14, obvious discrepancies between benign and backdoor trajectories can be consistently observed in different attacks and representation spaces. We also provide quantitative similarity analysis. As can be seen, backdoor prompts of different attack methods exhibit highly similar temporal patterns of trajectory discrepancy. These results lead to our second empirical observation: Backdoor attacks induce distinct deviations from benign diffusion trajectories.
Appendix B Transition Dynamics Analysis
B.1 Proof of Proposition 1
The proof of Proposition 1 is as follows:
Proof.
The benign and backdoor transition dynamics are given by:
| (14) |
Let denote the deviation between the backdoor and benign trajectories. Then, we obtain:
| (15) |
To use the norm and the triangle inequality, we can further obtain:
| (16) |
Recursively, we obtain:
| (19) |
Since the backdoor prompts includes the triggers, holds. Thus, we have:
| (20) |
∎
B.2 Proof of Proposition 2
The proof of Proposition 2 is as follows:
Proof.
We recall that the benign and backdoor transition dynamics are defined as:
| (21) |
Assume that the learned predictor approximates the normal transition function with the bounded error:
| (22) |
Given a benign trajectory, the transition prediction error is:
| (23) |
Given a backdoor trajectory, we have:
| (24) |
This decomposition separates the backdoor-induced perturbation from the approximation error of the learned normal dynamics predictor. Applying the reverse triangle inequality, we obtain:
| (25) |
Therefore, the following two inequations can be obtained:
| (26) |
Obviously, sufficiently large backdoor-induced deviations can be distinguished from benign transitions through prediction inconsistency. ∎
Appendix C The Details of NDDL
C.1 Calculation of Multi-space Representations
C.1.1 Cross-Attention Weight
At each diffusion timestep , we denote the cross-attention weight as , where denotes the number of attention heads, denotes the number of spatial query positions and denotes the number of tokens. Specifically, represents the attention weight associated with the -th attention head at spatial position to the -th token. Since cross-attention weights are normalized along the token dimension, we can obtain . Then, we extract to form four complementary descriptors: attention entropy, effective rank, token importance and head diversity. These descriptors characterize the concentration, structural complexity, token-level contribution and inter-head variation of cross-attention, respectively. It should be particularly noted that we capture joint attention weight for Stable Diffusion v3.5.
Attention entropy: Attention entropy presents the concentration of token-wise attention distribution. For the -th head at spatial position , we calculate:
| (27) |
Then, we average the entropy over all spatial query positions:
| (28) |
Thus, the entropy descriptor can be formed. A larger entropy indicates that attention is distributed over a broader set of tokens, whereas a smaller entropy indicates that attention is concentrated on fewer tokens.
Effective rank: Effective rank depicts the structural complexity of each attention map. For the -th head, we view as a two-dimensional attention matrix and perform singular value decomposition:
| (29) |
Let represent the singular values of . We normalize the singular values as:
| (30) |
The entropy of the normalized singular-value distribution is:
| (31) |
Thus, we can obtain the effective rank:
| (32) |
The descriptor of effective rank is denoted as . A larger effective rank indicates greater structural complexity of the attention map, whereas a smaller value indicates that its structure is dominated by fewer components.
Token importance: Token importance measures the overall attention assigned to each token. For the -th token, we average its attention weights over all heads and spatial positions:
| (33) |
This descriptor is denoted as , representing the average relative attention assigned to the -th token at timestep .
Head diversity: Head diversity reflects the variation in token-level attention distribution among different heads. For each head, we first average the attention weights over all spatial positions:
| (34) |
We then normalize the averaged attention weights over the token dimension:
| (35) |
For each pair of attention heads , we define their mixture distribution as:
| (36) |
JS divergence is calculated as:
| (37) |
where KL divergence is obtained from:
| (38) |
Finally, we average JS divergences over all distinct attention-head pairs:
| (39) |
The head diversity descriptor is . A larger value indicates greater variation in token-level attention distribution among different heads, whereas a smaller value indicates more consistent attention behaviors.
C.1.2 Latent
At each diffusion timestep , the latent state is denoted as , where denotes the number of latent channels, and denote the spatial resolution. We construct the compact representation including four descriptors: channel norm, trajectory curvature, frequency energy and temporal variation. These descriptors describe the magnitude, variation degree, spectral structure and local transition behavior of the latent trajectory. Notably, for the boundary steps where the preceding states are unavailable, we apply zero-padding.
Channel norm: Latent norm reflects the overall magnitude of each latent channel. For the -th latent channel , we obtain:
| (40) |
The descriptor of latent norm is . A larger latent norm indicates a larger overall magnitude of the corresponding latent channel, providing a compact representation of the instantaneous latent state.
Trajectory curvature: Latent curvature is calculated from the second-order temporal variation of the latent trajectory. We first obtain the consecutive transition difference of the -th latent channel as:
| (41) |
The curvature is then calculated as the variation between two consecutive transition differences:
| (42) |
Equivalently, we have:
| (43) |
Thus, the descriptor of latent curvature is . A larger curvature indicates a stronger change in the local transition of the latent trajectory, while a smaller value suggests a more consistent evolution between consecutive denoising steps.
Frequency-domain energy: Latent frequency energy reflects the spectral structure of each latent channel. For the -th latent channel, we implement a two-dimensional Fourier transform:
| (44) |
where denotes the coordinate. The corresponding power spectrum is defined as:
| (45) |
We divide the frequency domain into frequency regions and compute the energy for each region:
| (46) |
To reduce the dynamic range of frequency energy, we exploit a logarithmic transformation:
| (47) |
Thus, the descriptor of frequency energy is We set corresponding to low-frequency, middle-frequency and high-frequency regions. This descriptor reveals how latent representation is distributed over different frequency components during denoising.
Temporal variation: Latent temporal variation reflects the magnitude of the local transition between two consecutive denoising steps. For the -th latent channel, we compute:
| (48) |
The descriptor of temporal variation is . A larger norm indicates a stronger transition between consecutive latent states, whereas a smaller value represents relatively mild local evolution.
C.1.3 Noise
At each diffusion timestep , we denote the predicted noise as , where denotes the number of noise channels, and denote the spatial resolution. We obtain the descriptor of noise predictions includes channel norm, channel variance, frequency energy and temporal variation. Similarly, we utilize zero-padding for the unavailable preceding states of the boundary steps.
Channel norm: Noise channel norm calculates the overall magnitude of each noise channel. For the -th channel , we have:
| (49) |
The descriptor of channel norm is defined as .
Channel variance: Noise channel variance represents the spatial dispersion of the predicted values of each noise channel. Given the -th predicted-noise channel , we first compute its spatial mean as:
| (50) |
The variance of is then obtained:
| (51) |
The descriptor of channel variance at timestep is denoted as . A larger variance indicates a more dispersed spatial distribution of the noise values, whereas a smaller variance indicates that the values are more concentrated around the channel mean.
Frequency-domain energy: Noise frequency energy manifests the spectral distribution of each noise channel. For the -th channel , we transform the spatial representation into the frequency domain using the two-dimensional discrete Fourier transform:
| (52) |
The corresponding power spectrum is then computed as:
| (53) |
The energy of the -th channel for each region is obtained:
| (54) |
Also, we further apply a logarithmic transformation:
| (55) |
The descriptor of frequency energy is .
Temporal variation: Noise temporal variation reflects the magnitude of the local transition between two consecutive noise predictions. For the -th noise channel, we compute:
| (56) |
The descriptor of temporal variation is .
| Representation | Descriptor | Characterized Property |
|---|---|---|
| Cross-Attention Weight | Attention Entropy | Distribution concentration |
| Effective Rank | Structural complexity | |
| Token Importance | Token-level contribution | |
| Head Diversity | Inter-head variation | |
| Latent | Channel Norm | State magnitude |
| Curvature | Second-order temporal variation | |
| Frequency Energy | Spectral structure | |
| Temporal Variation | First-order temporal variation | |
| Noise Prediction | Channel Norm | Prediction magnitude |
| Channel Variance | Spatial dispersion | |
| Frequency Energy | Spectral structure | |
| Temporal Variation | First-order temporal variation |
Table 6 provides the descriptors of the trajectory representation mapping. We do not claim that these descriptors are exhaustive or uniquely optimal. We aim to demonstrate that learning normal transition dynamics in a compact multi-space representation provides an effective strategy for backdoor defense of T2I diffusion models.
C.2 Normalization Details
We first normalize each feature dimension independently using the statistics computed exclusively from the benign training trajectories. Let denote the value of the -th feature dimension in the -th benign training sample. For each feature dimension , we compute its median as and the median absolute deviation (MAD) as . The normalized feature is then calculated as:
| (57) |
where prevents extreme case when the MAD approaches zero. The constant rescales the MAD to provide a robust estimate comparable to the standard deviation under a Gaussian distribution.
After feature-wise normalization, the obtained descriptors are grouped into three representation blocks: cross-attention, latent and noise prediction. The normalized feature block for representation is represented as:
| (58) |
where is its dimensionality.
To mitigate the dimensionality-induced imbalance, we then scale each block by the square root of its dimensionality:
| (59) |
C.3 The architecture of normal dynamics network
The normal dynamics network is a residual MLP conditioned on timestep . Given the current state , the network predicts the state increment between adjacent steps. Specially, the timestep is first mapped to a 32-dimensional vector via sinusoidal positional encoding, and then processed by a two-layer MLP to obtain the time representation . The state and time representation are concatenated and projected into a 512-dimensional hidden space via a linear layer followed by LayerNorm and GELU. Then, the hidden output is input into 4 residual blocks. Finally, the output head maps the hidden result to the state space, i.e., the increment . The details of normal dynamics network can be seen in Tables 7 and 8.
| Stage | Operation | Output shape |
|---|---|---|
| State input | ||
| Step input | ||
| Time embedding | MLP | |
| Concatenation | ||
| Input projection | Linear LayerNorm GELU | |
| Dynamics core | ResidualBlock | |
| Output head | Linear GELU Linear | |
| Output |
| Sub-layer | Operation | Shape |
|---|---|---|
| Input | ||
| Linear (up) | Linear GELU | |
| Dropout | Dropout | |
| Linear (down) | Linear | |
| Normalize | LayerNorm | |
| Add (skip) |
C.4 Details of Low-Semantic Words
For trigger localization, we construct an initial substitution set using common function words with low semantics, such as a, an, the, this, that, some, and, with and of. These words typically introduce less semantic perturbations.
C.5 Pseudocode of NDDL
Algorithm 1 presents the training process of normal dynamics model. Algorithms 2 and 3 show the details of defense process, including both backdoor detection and trigger localization.
Appendix D Experimental results
D.1 Detection Results
Table 9 summarizes the backdoor detection results on Stable Diffusion XL. Remarkably, NDDL consistently achieves superior performance, exceeding 97% in both ACC and AUROC across all evaluated attacks. The most significant improvement is observed on STEBA, where NDDL surpasses the strongest baseline by a substantial margin, increasing ACC from 83.3% to 97.2% and AUROC from 83.1% to 97.0%. These results further corroborate the effectiveness and generalizability of NDDL across diverse diffusion architectures.
D.2 Localization Results
As presented in Table 10, NDDL also attains the best trigger-localization performance on Stable Diffusion XL. In addition to one-token and multi-token triggers, NDDL further yields substantial improvements in localizing both special-character and sentence-level triggers. These results further underscore the effectiveness and generalizability of NDDL in localizing diverse trigger forms across different diffusion architectures.
| Method | BadT2I | EvilEdit | MasqLoRA | Rickrolling | STEBA | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| ACC | AUROC | ACC | AUROC | ACC | AUROC | ACC | AUROC | ACC | AUROC | |
| UFID | 64.8 | 65.7 | 65.3 | 67.0 | 65.8 | 66.6 | 55.8 | 56.5 | 52.8 | 53.3 |
| T2IShield | 84.3 | 84.9 | 83.5 | 84.5 | 86.8 | 87.4 | 81.3 | 81.8 | 75.8 | 76.9 |
| NaviT2I | 93.2 | 92.9 | 96.2 | 96.4 | 90.7 | 90.9 | 84.1 | 84.9 | 70.8 | 71.6 |
| STEDF | 98.2 | 98.3 | 98.8 | 99.1 | 96.3 | 96.1 | 96.5 | 96.9 | 83.3 | 83.1 |
| NDDL (Ours) | 98.7 | 99.0 | 98.5 | 98.7 | 98.1 | 98.2 | 97.3 | 97.1 | 97.2 | 97.0 |
| Method | One-token | Multi-token | Special-character | Sentence | ||||
|---|---|---|---|---|---|---|---|---|
| ETR | AUROC | ETR | AUROC | ETR | AUROC | ETR | AUROC | |
| T2IShield | 90.0 | 90.4 | 80.6 | 80.9 | 83.2 | 82.6 | 71.8 | 72.2 |
| NaviT2I | 97.1 | 97.6 | 96.5 | 96.5 | 90.2 | 88.1 | 80.7 | 79.2 |
| NDDL (Ours) | 98.2 | 98.1 | 97.2 | 96.6 | 94.5 | 96.0 | 84.3 | 84.6 |
D.3 Evaluation results on Pixart-
Table 11 presents a comprehensive comparison of various defense methods on Pixart-. When evaluated against the attack on Pixart-, NDDL achieves the best overall performance in both backdoor detection and trigger localization, outperforming all competing defenses. These results further substantiate the generalizability of NDDL to the DiT architecture.
| Method | Detection | Localization | ||
|---|---|---|---|---|
| ACC | AUROC | ETR | AUROC | |
| UFID | 52.5 | 50.6 | – | – |
| NaviT2I | 85.5 | 87.0 | 78.2 | 78.2 |
| NDDL (Ours) | 92.3 | 91.7 | 87.1 | 87.5 |