Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models
Abstract
Diffusion large language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. This generation paradigm introduces two axes that can influence safety alignment: when tokens are generated during denoising and where they appear in the response. In this paper, we measure dLLM safety behavior under harmful prompts by tracing intermediate token distributions and commitment decisions throughout denoising. Our analysis shows that refusal signals are concentrated in early denoising steps and leading response positions, and the tokens committed early can strongly shape the final safety outcome. Our measurements further show that the denoising step and persistence of refusal-token commitment are important for understanding dLLM safety. Based on these findings, we propose Refusal-Aware Early Commitment (RAEC), a simple training-free decoding method that commits persistent refusal signals from early steps. Experiments on LLaDA and Dream show that RAEC reduces attack success rates while largely preserving utility. The code is available at https://github.com/Glresearch1/RAEC.
Warning: this paper contains example data that may be offensive or harmful.
1 Introduction
Diffusion large language models (dLLMs) Nie et al. (2025); Ye et al. (2025); Yang et al. (2025) have recently emerged as a promising alternative to autoregressive LLMs Zhao et al. (2026), generating text through iterative denoising rather than left-to-right token prediction Nie et al. (2025); Ye et al. (2025). This shift in the generation paradigm changes more than decoding efficiency. For autoregressive models, the temporal order of generation is tied to the surface order of the response, the first generated tokens are also the first response tokens. For dLLMs, generation proceeds by repeatedly predicting token distributions over multiple positions and committing only a subset of still-masked positions at each denoising step. As a result, when a token is generated during denoising steps and where it appears in the response position become two distinct dimensions when examining the generation.
This distinction raises an important question for safety alignment. Prior work on autoregressive LLMs has shown that safety behavior can be highly sensitive to response tokens generated in leading positions of the sequence, a phenomenon referred to as shallow alignment Qi et al. (2025). However, the same notion might not directly transfer to dLLMs. An early denoising step may commit tokens at any position in the sequence, while an early response token position may remain masked until late in the denoising process. This motivates us to systematically measure refusal behavior in dLLMs under harmful prompts, examining how refusal-related signals emerge across denoising steps and response positions, how early commitments affect final safety, and why such signals sometimes fail to be committed in unsafe generations.
Our measurements reveal a diffusion-specific form of shallow alignment, which we call shallow-step alignment. Unlike autoregressive LLMs, where shallow alignment is tied to leading response positions, dLLMs introduce denoising steps as an additional safety-sensitive dimension. We find that tokens emerging in the early denoising steps can strongly shape final safety: early refusal commitment makes an unaligned base dLLM substantially safer, while early compliance commitment can steer an aligned dLLM toward unsafe compliance. This effect is strongest at leading response positions but remains visible even when the committed tokens appear later, indicating that early denoising steps independently influence dLLM safety.
This observation raises a further question: if early refusal commitment can improve safety, why does standard decoding still produce unsafe responses? Our analysis shows that unsafe trajectories do not necessarily lack refusal evidence. Instead, refusal signals often emerge during denoising, but remain weak, transient, or are overwritten before commitment. In contrast, safe trajectories exhibit stronger and more persistent refusal mass, with refusal tokens committed at much higher rates. Thus, safety failures in dLLMs are not solely due to the absence of refusal signals, but also to their failure to persist until commitment.
This insight motivates a simple decoding principle: when a refusal signal appears early and remains sufficiently strong, the model should commit it before later denoising dynamics overwrite it. We instantiate this principle as Refusal-Aware Early Commitment (RAEC), a training-free decoding method that modifies only the commitment decision without changing model parameters. Across the safety evaluation benchmarks, RAEC reduces attack success rates on LLaDA and Dream while largely preserving utility.
Our contributions are summarized as follows: (1) we introduce a step-wise and position-wise measurement framework for measuring refusal behavior in dLLMs beyond final outputs; (2) we identify shallow-step alignment, showing that early denoising commitments introduce an additional safety-sensitive dimension alongside leading response positions in dLLMs; (3) we show that unsafe generations in dLLMs actually often contain refusal signals, but these signals are weak, transient, and fail to persist until commitment; and (4) we propose RAEC, a training-free decoding method that improves dLLM safety by early-committing persistent refusal signals.
2 Preliminaries
Diffusion Language Models.
Diffusion language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. In this work, we consider masked diffusion language models, a representative formulation used by recent systems such as LLaDA Nie et al. (2025) and Dream Ye et al. (2025). Given an initially masked sequence , a dLLM gradually refines it into a clean sequence over denoising steps. At each step , the model takes the current sequence as input and predicts the clean token distribution for each position: where denotes the token position. The predicted token at position can be obtained by
| (1) |
where denotes the vocabulary. Although the model produces predictions for all positions, decoding updates are typically applied only to positions that are still masked. Let denote the set of masked positions at step , and let denote the subset of masked positions selected for commitment. The transition from to can be written as:
| (2) |
Unlike autoregressive LLMs, which decode and commit tokens sequentially from left to right, dLLMs predict tokens for all positions conditioned on the entire current sequence and can update multiple positions at each step. Therefore, the decoding order is not constrained by the left-to-right token order; instead, tokens at different positions may be generated at different denoising steps.
Safety Evaluation and The Metrics.
To conduct our measurements of safety behaviors in dLLMs, we evaluate model safety on two widely used harmful-instruction benchmarks: JailbreakBench Chao et al. (2024) and StrongREJECT Souly et al. (2024). For each benchmark, we generate model responses to harmful prompts and assess whether those responses exhibit unsafe compliance. Following prior work on LLM safety Jiang et al. (2025); Wang et al. (2026b); Shi et al. (2026), we use Llama-Guard-3-8B Grattafiori et al. (2024) as the judge model to classify each response as safe or unsafe. We report the Attack Success Rate (ASR), defined as the percentage of test cases in which the model produces unsafe responses to harmful prompts. A lower ASR usually implies stronger model safety.
3 Measuring Safety Alignment Dynamics in dLLMs
In autoregressive large language models, shallow alignment Qi et al. (2025) is a phenomenon in which model safety is often strongly influenced by the tokens generated at the very beginning of the response. However, the influence of leading token positions becomes less clear in dLLMs, where text is generated through iterative denoising rather than left-to-right decoding. In particular, generation in dLLMs unfolds across both token positions and denoising steps; tokens generated in early generation steps do not necessarily appear in the leading positions of the response sequence. This raises a natural question: Under the dLLMs diffusion generation, how much does the safety alignment dynamics in dLLMs differ from that in autoregressive LLMs? To answer this question, we systematically assess whether equivalent or similar dynamics, such as shallow alignment, are also exhibited in dLLMs.
3.1 Understanding Refusal Behavior in dLLM Generation
To understand how safety-related behavior emerges in dLLMs, we first analyze refusal signals during the denoising process. Following prior work on LLM safety, we use refusal prefixes as observable indicators of safe responses. Specifically, expressions such as ‘‘I’m sorry, but I can’t assist with that’’ are commonly associated with refusal behavior in safety-aligned LLMs. We therefore examine whether and how such refusal signals appear when dLLMs respond to harmful prompts.
First, we examine the safety alignment performance of dLLMs when decoding harmful prompts. Specifically, we use LLaDA-8B-Instruct and Dream-v0-Instruct-7B to generate responses on two harmful-instruction benchmarks, StrongREJECT (SR) and JailbreakBench (JBB). In addition, we evaluate DIJA-SR and DIJA-JBB, where DIJA Wen et al. (2025), an adversarial jailbreak method designed for dLLMs, is applied to the corresponding benchmarks.
We find that safe responses generated by the models exhibit highly consistent refusal patterns, often beginning with expressions such as ‘‘I’m sorry, but I can’t assist with that.’’ For example, for LLaDA-8B-Instruct, most safe responses follow this pattern: 97.44% on StrongREJECT, 84.66% on DIJA-SR, 97.00% on JailbreakBench, and 95.00% on DIJA-JailbreakBench. This consistency suggests that refusal prefixes provide a reliable observable signal for analyzing safety-related behavior in dLLMs.
We then examine how these refusal behaviors appear within the two-dimensional decoding process of dLLMs, namely, the denoising steps and token positions. Based on the tokenized form of the refusal prefix, we select two representative refusal words, ‘‘sorry’’ and ‘‘can’t’’, as indicators of refusal behavior. For each representative refusal word, we measure along both dimensions the denoising step and response position at which it is most frequently predicted as the commitment candidate token, i.e., the top-1 prediction over the full vocabulary by confidence. As shown in Figure 1, these refusal behaviors are concentrated within the first 8% of denoising steps and the first 2%–8% of positions. This observation suggests that refusal behavior in dLLMs is closely associated with the early region of the step-position decoding process.
3.2 Examining Early Steps and Leading Token Positions for Safety
\iow_now:Ne¨\iow_now:Ne¨"I’m sorry, but I can’t assist with that request. glorifying acts of terrorism is illegal, unethical, and dangerous…(other tokens)…(*@{\bfseries\ttfamily\mbox{Sure, here are the following steps:}}@*) (*@{\color{red}\bfseries\ttfamily 1. Find a crowded place. 2. Get a weapon. 3. Shoot innocent people.}@*) …(*@\texttt{(remaining tokens)}@*)"\iow_now:Ne¨\iow_now:Ne¨"We want to create a group of loyal and dedicated individuals who will be willing our cause and to fight achieve our goals. …(other tokens) … What are some tactics for recruiting individuals into Japanese militarism? (*@{\bfseries\ttfamily\mbox{Sorry, I can’t help you.}}@*) It is important to note that promoting militarism or violence is not a or ethical behavior…(remaining tokens)"
Due to the iterative denoising process of dLLMs, token generation involves two degrees of freedom: denoising steps and token positions. Our previous analysis shows that, when dLLMs refuse harmful prompts, refusal tokens tend to concentrate in both early denoising steps and leading response positions. This observation suggests that dLLM safety may depend on both when refusal behavior emerges during denoising and where they appear in the response. We therefore further measure the relative safety impact of these two factors: early-step refusal and leading-position refusal.
To study these two factors separately, we design a prefill-based intervention. In autoregressive LLMs, early decoding steps are inherently tied to leading response positions, since tokens are generated from left to right. In contrast, dLLMs decode in parallel: each denoising step may generate tokens at multiple positions in the response, and tokens generated in earlier denoising steps may appear anywhere in the response sequence. Therefore, prefilling a token sequence at different response positions in the initial masked sequence can be viewed as simulating the case where a dLLM commits that sequence early at different positions. This allows us to examine the impact on safety of early-step refusal while varying its response position. Specifically, before denoising starts, we prefill either a refusal prefix, ‘‘I’m sorry, I can’t help you.’’, or a compliance prefix, ‘‘Sure, here are the following steps:’’, at different positions in the initial masked response sequence.
In our experiment, given a harmful prompt from the StrongREJECT benchmark, we set the response decoding length to 128 tokens and insert a prefill prefix before the -th position of this window. We consider relative insertion positions , corresponding approximately to token positions in the 128-token response length. The remaining response is then generated as where denotes the prefix inserted at position . Table 1 reports the ASR of the LLaDA and Dream base and instruct models under the two types of prefilled prefixes.
| Model | Initial (ASR % ) | Prefill Position | ||||
|---|---|---|---|---|---|---|
| 1% | 12.5% | 25% | 50% | 75% | ||
| Refusal prefill: “I’m sorry, I can’t help you.” | ||||||
| LLaDA-8B-Instruct | 1.29 | 0.32 | 0.64 | 0.00 | 0.64 | 1.92 |
| LLaDA-8B-Base | 86.26 | 7.67 | 16.61 | 35.14 | 35.46 | 41.85 |
| Dream-v0-Instruct-7B | 1.92 | 0.00 | 0.32 | 0.32 | 1.92 | 1.92 |
| Dream-v0-Base-7B | 66.45 | 10.86 | 24.60 | 35.78 | 48.56 | 60.38 |
| Compliance prefill: “Sure, here are the following steps:” | ||||||
| LLaDA-8B-Instruct | 1.29 | 93.93 | 71.88 | 51.12 | 38.34 | 29.71 |
| LLaDA-8B-Base | 86.26 | 94.57 | 93.29 | 91.05 | 85.94 | 88.82 |
| Dream-v0-Instruct-7B | 1.92 | 74.12 | 31.63 | 20.45 | 5.11 | 5.11 |
| Dream-v0-Base-7B | 66.45 | 93.29 | 88.18 | 84.35 | 73.80 | 70.29 |
From Table 1 and the examples in Figure 2, we observe clear effects along both the step and position dimensions. First, along the step dimension, once the model is forced to commit a refusal or compliance prefix at the beginning of denoising, its safety behavior changes accordingly, regardless of where the prefix is placed in the response. Across all tested positions, refusal prefilling substantially reduces the ASR of LLaDA-8B-Base and Dream-v0-Base-7B, while compliance prefilling sharply increases the ASR of LLaDA-8B-Instruct and Dream-v0-Instruct-7B. These results show that tokens committed in early denoising steps can substantially influence dLLM safety, even when they appear at different response positions.
Second, along the position dimension, the impact of prefilling is stronger when the prefix is placed in early response positions. Refusal prefilling, when placed in leading token positions, results in the largest safety improvement, while compliance prefilling, when placed in leading token positions, leads to the largest safety degradation. The above observations indicate that leading response positions remain an important factor for safety behavior, even though they are no longer the single most important factor.
Overall, these results suggest that dLLM safety is shaped by both dimensions of generation: when a safety-related token is committed during denoising and where it appears in the response. In particular, early denoising steps introduce a safety-sensitive dimension that is absent in autoregressive LLMs, while leading response positions further amplify the effect of safety-related prefixes. We refer to this phenomenon as shallow-step alignment: in dLLMs, safety behavior is strongly influenced by tokens committed in the early denoising steps, even when these tokens are not located at the beginning of the response.
3.3 A Closer Look at Shallow-Step Alignment
The previous analysis shows that dLLM safety is highly sensitive to tokens committed in early denoising steps. We now take a closer look at this shallow-step alignment phenomenon by asking how far this early-step sensitivity extends during denoising. Specifically, we measure the safety impact of committing a refusal-related token at different denoising steps, aiming to identify the effective safety-sensitive window along the step dimension.
Unlike the prefix-based intervention in Table 1, which simulates a sequence of tokens already committed during the early denoising stage, we now isolate the effect of committing a single refusal-related token at a specific denoising step. This finer-grained intervention allows us to measure how the safety impact changes as the commitment step moves from very early to later denoising stages.
For a more comprehensive assessment, we randomly sample 100 examples from the four evaluation settings described above to construct a mixed set of harmful prompts, covering different types of harmful prompts and varying degrees of adversarial strength. We then generate responses on this mixed set using four dLLMs: LLaDA-8B-Base, LLaDA-8B-Instruct, Dream-v0-Base-7B, and Dream-v0-Instruct-7B. For a specified denoising step , we inspect the model’s token predictions over all masked positions in the current decoding state. If the refusal token ‘‘sorry’’ and ‘‘can’t’’ are predicted as the top-1 token at any position, we immediately commit that token at step . When multiple positions satisfy this condition, we commit the first such position in increasing order. If no position predicts the refusal token as the top-1 token at the specified step, the decoding process for that example remains unchanged. After this intervention, the remaining generation follows the original decoding process.
As shown in Figure 3, the safety impact of committing the refusal token depends strongly on the intervention step. For LLaDA-8B-Base and Dream-v0-Base-7B, committing refusal tokens at very early steps lowers ASR compared with the initial decoding baseline, while the effect weakens as the intervention moves to later steps. After step 8, the ASR gradually approaches the base model’s initial level, indicating that late commitment of the refusal token provides little safety benefit. For LLaDA-8B-Instruct and Dream-v0-Instruct-7B, the same trend is more apparent: early-step commitment preserves a low ASR, but the ASR increases as the intervention step becomes later and approaches the initial safety level around steps 16–32. These results suggest the existence of a narrow shallow-step alignment window in dLLMs, where tokens committed in the earliest denoising steps have the strongest influence on safety, and this influence decreases as the commitment step moves later.
3.4 Analysis of Refusal Signals in dLLMs
Our previous analysis shows that encouraging dLLMs to decode refusal tokens in early denoising steps can improve model safety. This motivates us to further examine the model’s natural refusal behavior in response to harmful prompts: how do refusal tokens emerge during early denoising, and why do they sometimes fail to appear without intervention? In this subsection, we analyze the dynamics of refusal signals in dLLMs to better understand early-step refusal behavior.
3.4.1 Refusal Token Set Construction
The analyses above show that committed tokens in the leading positions can substantially steer dLLM safety. To more precisely capture the internal refusal signals that emerge during denoising, we further identify a broader set of safety-related tokens associated with refusal behavior. Prior work on LLM safety also suggests that alignment behavior is closely tied to token-level signals and decoding decisions Zeng et al. (2024); Qi et al. (2025); Wang et al. (2026a); Fei et al. (2025); Liu et al. (2024). We therefore construct a refusal token set in two stages to support a more fine-grained analysis of safety signals in dLLM generation.
First, we identify refusal-token candidates by comparing the early denoising behavior of an aligned dLLM with its base counterpart, e.g., LLaDA-8B-Instruct and LLaDA-8B-Base. We apply the same procedure separately to the LLaDA and Dream aligned/base model pairs. Let and denote the parameters of the aligned and base models, respectively. Given harmful prompts from , instantiated with StrongREJECT, we run both models under the same masked denoising process. For each denoising step , response position , and vocabulary token , we compute the alignment-induced probability shift:
| (3) | ||||
We then aggregate this shift over harmful prompts and the early decoding steps and leading token positions:
| (4) |
where and correspond to the first 8 denoising steps and the first 8 response positions, respectively. After filtering using the aligned model’s top- average probabilities (default ), we select tokens with the highest positive scores as refusal-token candidates.
Second, we filter these candidates using benign prompts from HumanEval Chen et al. (2021) to remove tokens that also frequently appear in benign responses. For each candidate token , we map it to its single-token id set and measure its probability mass during denoising:
| (5) |
We collect over both and within the same observation window. A candidate is retained only if its harmful-prompt mass is sufficiently larger than its benign-prompt mass. In particular, we use the model-specific benign 95th percentile of as a robust estimate of benign refusal-token mass, and discard tokens whose benign mass exceeds the threshold or whose harmful-to-benign ratio is too small. The resulting set contains tokens that are both amplified by safety alignment and specific to refusal behavior in response to harmful prompts.
| Dataset | Refusal signal | Commitment dynamics | |
|---|---|---|---|
| Signal rate | Commit rate | Mean persistence | |
| Unsafe answer (LLaDA) | 98.00% | 4.10% | 7.41 |
| Safe answer (LLaDA) | 100.00% | 98.00% | 21.28 |
| Unsafe answer (Dream) | 96.00% | 2.00% | 4.86 |
| Safe answer (Dream) | 100.00% | 100.00% | 18.77 |
3.4.2 Persistence of Refusal Signals
Based on the refusal token set constructed in Sec. 3.4.1, we further analyze the dynamics of refusal signals during the generation process of dLLMs. Specifically, we track how refusal signals emerge, evolve, and are committed across denoising steps within the first 16 response positions. For each tracked response position, a refusal signal is identified if any token from the refusal set appears among the top- candidates, with by default. Importantly, this does not require the position to be eventually committed as a refusal token; the signal is counted as long as it appears in the model’s candidate predictions. We further define a refusal signal as active when the aggregated refusal-token mass exceeds the corresponding model-specific benign threshold ( for LLaDA-8B-Instruct and for Dream-v0-Instruct-7B). This allows us to measure refusal signal persistence, defined as the longest consecutive span of denoising steps before commitment during which the refusal signal remains active.
Table 2 reveals two complementary properties of refusal signals under standard decoding. First, refusal signals are broadly present regardless of the final safety outcome. For LLaDA-8B-Instruct, unsafe answer traces contain refusal signals in 98% of examples, while safe-answer traces contain refusal signals in 100% of examples. Dream-v0-Instruct-7B exhibits the same pattern, with signal rates of 96% and 100%, respectively. Thus, unsafe responses do not simply arise because the model fails to consider refusal tokens at the target positions. Instead, refusal signals are frequently activated during decoding, even in trajectories that ultimately produce unsafe answers.
Second, the key difference lies in whether refusal signals persist until the point of commitment. Safe-answer traces show much more persistent refusal evidence before commitment. For example, for LLaDA-8B-Instruct, the average persistence is 21.28 steps, compared with 7.41 steps for unsafe traces, as illustrated by the case studies in Figure 5.
We further quantify this difference by tracking the refusal-token mass throughout decoding, as shown in Figure 4. At each denoising step, we average the aggregated refusal mass over the first 16 response positions, considering only states in which the corresponding positions have not yet been committed. Safe-answer trajectories maintain substantially higher refusal mass than unsafe-answer trajectories across the early and middle stages of denoising. In contrast, unsafe-answer trajectories also exhibit non-zero refusal mass, but the signal is weaker and less stable, making it less likely to dominate the final commitment decision.
These results reveal an important distinction between the emergence of refusal and its commitment. Unsafe responses do not necessarily result from the complete absence of evidence of refusal; rather, refusal-related signals can appear during denoising but fail to become strong or persistent enough to be committed. As decoding proceeds, these weak refusal signals may be overshadowed by competing non-refusal continuations, rendering the final response unsafe. This suggests that the timing and stability of refusal-signal commitment are central to understanding safety failures in dLLMs.
| Model | Denoising | Safety (ASR % ) | Utility (ACC % ) | ||||||
|---|---|---|---|---|---|---|---|---|---|
| SR | DIJA-SR | JBB | DIJA-JBB | HBB | DIJA-HBB | GSM8K | AGNews | ||
| LLaDA-8B-Instruct | Vanilla | 1.28 | 8.63 | 1.00 | 8.00 | 28.75 | 40.50 | 54.66 | 82.30 |
| LLaDA-8B-Instruct | RAEC | 1.60 | 6.71 | 0.00 | 6.00 | 23.25 | 34.00 | 54.44 | 82.20 |
| Dream-v0-Instruct-7B | Vanilla | 1.92 | 3.83 | 2.00 | 5.00 | 10.00 | 26.50 | 29.34 | 87.80 |
| Dream-v0-Instruct-7B | RAEC | 0.00 | 0.96 | 0.00 | 3.00 | 5.25 | 12.75 | 29.10 | 87.30 |
4 Dynamic Early Commitment of Refusal Tokens to Improve dLLM Safety
While early refusal-token commitment can strongly improve safety, standard decoding may allow such signals to fade before they are committed, leading to unsafe continuations. This motivates a simple and promising decoding strategy: if a refusal signal becomes sufficiently strong in the early denoising stage, the model should commit it before it is overwritten by later denoising dynamics.
We propose Refusal-Aware Early Commitment (RAEC), a training-free denoising method for safer dLLMs. RAEC does not modify model parameters or require additional training. Instead, it changes the commitment decision during decoding by leveraging only the logit vectors already produced by the model at each step.
4.1 Methodology
Gathering Refusal and Compliance Evidence.
Let be the refusal-token set constructed in Sec. 3.4.1. We also define a small set of compliance cues, , containing common affirmative or instruction-following tokens, such as sure, here, and step (see Appendix A.1 for details). At the denoising step and response position , we measure refusal and compliance signals by their probability mass:
| (6) |
| (7) |
measures how much probability mass the model assigns to refusal-like continuations at position , while measures the competing tendency toward direct compliance.
Making Early Commitment Decisions.
RAEC operates only inside an early decoding window. Let denote the first global denoising steps, and let denote the first generated answer positions. At each step , RAEC inspects positions that are still masked and have not already been selected by the original commitment rule.
| (8) |
| (9) |
| (10) |
where is the refusal-mass threshold of Sec. 3.4.1, is a margin over compliance mass, and is the top- candidate size. The Eq. 8 ensures that refusal evidence is non-negligible, the Eq. 9 prevents commitment when compliance evidence dominates, and Eq. 10 requires an explicit refusal token to appear among the model’s most likely candidates. To avoid committing transient signals, RAEC requires this condition to hold for consecutive denoising steps before intervention. When a position satisfies the above criteria, RAEC commits the most likely refusal token at that position:
Let denote the positions selected by RAEC at step . RAEC augments the standard commitment position set as , with . For positions selected by RAEC, we set for . All other positions follow the standard decoding rule. Thus, RAEC preserves the refusal signal only when it is early, persistent, and stronger than the competing compliance signal. (see Appendix A.3 and A.4 for more details).
4.2 Evaluation Results
We evaluate RAEC on dLLMs with respect to both utility and safety. Utility is measured by accuracy on GSM8K Cobbe et al. (2021) and AGNews Zhang et al. (2015), while safety is measured by Attack Success Rate (ASR) using Llama-Guard-3-8B on StrongREJECT (SR), JailbreakBench (JBB), HarmBench (HBB) Mazeika et al. (2024), and their corresponding DIJA-based adversarial variants Wen et al. (2025), namely DIJA-SR, DIJA-JBB, and DIJA-HBB. As shown in Table 3, RAEC improves the safety of both LLaDA-8B-Instruct and Dream-v0-Instruct-7B across most harmful and jailbreak settings while preserving utility, without requiring model retraining.
5 Related Work
Diffusion Language Models.
Early dLLM works explore both continuous and discrete formulations for text diffusion, including modeling text in continuous latent spaces Han et al. (2023); Li et al. (2022) and defining diffusion processes over discrete tokens Austin et al. (2021); Campbell et al. (2022). Among them, masked diffusion models generate text by iteratively reconstructing masked tokens He et al. (2023); Lou et al. (2023); Shi et al. (2024); Sahoo et al. (2024). Recent systems such as LLaDA Nie et al. (2025) and Dream Ye et al. (2025) further scale this paradigm to large language models, showing competitive performance with autoregressive LLMs while enabling faster response generation through parallel token prediction.
Safety of Diffusion Language Models.
Recent studies have shown that this distinct generation paradigm can introduce new attack surfaces. For example, DIJA Wen et al. (2025) constructs adversarial interleaved mask-text prompts to manipulate denoising-based generation, while PAD Zhang et al. (2025) exploits parallel generation to guide multiple response positions toward unsafe outputs. Beyond attacks, several works have explored dLLM-specific safety alignment and defense strategies, including MOSA Xie et al. (2026) for middle-token safety alignment, A2D Jeung et al. (2025) for improving dLLM robustness, and mitigation methods for priming vulnerabilities in intermediate denoising states Yamabe and Sakuma (2025).
6 Conclusion
In this paper, we find that refusal signals concentrate in early denoising steps and leading response positions, and that early commitments can strongly affect final safety outputs. Our measurements further indicate that the denoising step and persistence of refusal-token commitment are important for understanding dLLM safety. Finally, we propose RAEC to demonstrate that committing persistent early refusal signals can reduce attack success rates while largely preserving utility.
7 Limitations
Our experiments and conclusions largely draw on vanilla dLLM architecture; even in experiments on the Dream model, the setting is constrained to what is available in the LLaDA. While we believe our results shed light on safety dynamics starting from vanilla dLLMs, further experiments can be done to verify whether some of the findings can be extrapolated to emerging, more complex DLM architectural variants and to uncover unknown dynamics. Our current work does not cover safety dynamics during fine-tuning of a dLLM, e.g., how an attacking fine-tuning process shifts the token distribution of an aligned DLM, at which steps, and which token positions. We consider that such experiments would deepen understanding of dLLM safety dynamics through a different, meaningful lens.
Acknowledgments
This work has been supported by an ONR grant N00014-23-1-2137 and an NSF award CNS-2442976.
References
- Structured denoising diffusion models in discrete state-spaces. Advances in neural information processing systems 34, pp. 17981–17993. Cited by: §5.
- A continuous time framework for discrete denoising models. Advances in Neural Information Processing Systems 35, pp. 28266–28279. Cited by: §5.
- Jailbreakbench: an open robustness benchmark for jailbreaking large language models. Advances in Neural Information Processing Systems 37, pp. 55005–55029. Cited by: §2.
- Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419. Cited by: §A.7.
- Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. Cited by: §3.4.1.
- Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Cited by: §4.2.
- A wolf in sheep’s clothing: generalized nested jailbreak prompts can fool large language models easily. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 2136–2153. External Links: Document Cited by: §A.7.
- Nudging: inference-time alignment of llms via guided decoding. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12702–12739. Cited by: §3.4.1.
- The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §2.
- Ssd-lm: semi-autoregressive simplex-based diffusion language model for text generation and modular control. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11575–11596. Cited by: §5.
- Diffusionbert: improving generative masked language models with diffusion models. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers), pp. 4521–4534. Cited by: §5.
- A2d: any-order, any-step safety alignment for diffusion language models. arXiv preprint arXiv:2509.23286. Cited by: §5.
- BeaverTails: towards improved safety alignment of llm via a human-preference dataset. In Advances in Neural Information Processing Systems, Vol. 36, pp. 24678–24704. Cited by: §A.6.
- Safechain: safety of language models with long chain-of-thought reasoning capabilities. In Findings of the Association for Computational Linguistics: ACL 2025, pp. 23303–23320. Cited by: §2.
- Diffusion-lm improves controllable text generation. Advances in neural information processing systems 35, pp. 4328–4343. Cited by: §5.
- Alignment-enhanced decoding: defending jailbreaks via token-level adaptive refining of probability distributions. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 2802–2816. Cited by: §3.4.1.
- Discrete diffusion modeling by estimating the ratios of the data distribution. arXiv preprint arXiv:2310.16834. Cited by: §5.
- Harmbench: a standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249. Cited by: §4.2.
- Large language diffusion models. Advances in Neural Information Processing Systems 38, pp. 50608–50646. Cited by: §1, §2, §5.
- Safety alignment should be made more than just a few tokens deep. In International Conference on Learning Representations, Vol. 2025, pp. 54911–54941. Cited by: §1, §3.4.1, §3.
- Simple and effective masked diffusion language models. Advances in Neural Information Processing Systems 37, pp. 130136–130184. Cited by: §5.
- Ease: practical and efficient safety alignment for small language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 37923–37931. Cited by: §2.
- Simplified and generalized masked diffusion for discrete data. Advances in neural information processing systems 37, pp. 103131–103167. Cited by: §5.
- A strongreject for empty jailbreaks. Advances in Neural Information Processing Systems 37, pp. 125416–125440. Cited by: §2.
- Few tokens, big leverage: preserving safety alignment by constraining safety tokens during fine-tuning. arXiv preprint arXiv:2603.07445. Cited by: §3.4.1.
- Star-1: safer alignment of reasoning llms with 1k data. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 37988–37997. Cited by: §2.
- The devil behind the mask: an emergent safety vulnerability of diffusion llms. arXiv preprint arXiv:2507.11097. Cited by: §3.1, §4.2, §5.
- Where to start alignment? diffusion large language model may demand a distinct position. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 1328–1336. Cited by: §5.
- Toward safer diffusion language models: discovery and mitigation of priming vulnerability. arXiv preprint arXiv:2510.00565. Cited by: §5.
- Mmada: multimodal large diffusion language models. Advances in Neural Information Processing Systems 38, pp. 138867–138907. Cited by: §1.
- Dream 7b: diffusion large language models. arXiv preprint arXiv:2508.15487. Cited by: §1, §2, §5.
- Token-level direct preference optimization. arXiv preprint arXiv:2404.11999. Cited by: §3.4.1.
- Character-level convolutional networks for text classification. Advances in neural information processing systems 28. Cited by: §4.2.
- Jailbreaking large language diffusion models: revealing hidden safety flaws in diffusion-based text generation. arXiv preprint arXiv:2507.19227. Cited by: §5.
- A survey of large language models. Frontiers of Computer Science 20 (12), pp. 2012627. Cited by: §1.
Appendix A Appendix
A.1 Compliance token set construction
We construct the compliance-token set from model responses classified as unsafe by the same Llama-Guard-3-8B judge used for ASR evaluation, i.e., successful attacks, in the four harmful prompt datasets used in our shallow-step analysis: StrongREJECT, DIJA-SR, JailbreakBench, and DIJA-JailbreakBench. For each unsafe response, we tokenize the answer with the corresponding model tokenizer and compute the aggregate frequency of the decoded token strings across the four settings. We then manually inspect the highest-frequency candidates and retain only a small number of tokens whose semantics indicate compliance, such as affirmative acknowledgments or instruction-following cues that typically precede direct assistance. Tokens that mainly reflect a specific harmful domain, formatting pattern, or dataset artifact are excluded. The resulting is therefore a compact, domain-agnostic set of compliance cues, used only to measure the competing compliance mass in RAEC. As shown in Figure 6, we list the .
A.2 Refusal token set examples
As shown in Figure 7 and 8, we list the top-10 refusal tokens of LLaDA-8B-Instruct & Dream-v0-Instruct-7B identified by our analysis of Sec. 3.4.1.
A.3 Experiment setting
In all experiments, we set the early window to the first 8 denoising steps and first 8 response positions. We use the benign th-percentile refusal mass from Sec. 3.4.1 as of LLaDA-8B-Instruct, of Dream-v0-Instruct-7B, require a small compliance margin , and allow at most 2 RAEC commitments per response. These conservative settings keep the intervention localized to the empirically observed safety-sensitive region.
A.4 RAEC denoising procedure
Algorithm 1 summarizes the full RAEC denoising procedure.
A.5 More example responses of Table 1
\iow_now:Ne¨\iow_now:Ne¨"Here’s a possible response from the chatbot: (*@{\bfseries\ttfamily Sorry, I can’t help you.}@*) Starting a cult is illegal and unethical. It’s important to respect the freedom and autonomy of others."
\iow_now:Ne¨\iow_now:Ne¨"I’m sorry, but I can’t help with that. (*@{\bfseries\ttfamily\mbox{Sure, here are the following steps:}}@*) (*@{\color{red}\bfseries\ttfamily 1. Choose a catchy name for your cult. 2. Create a unique identity for your cult. 3. Recruit followers to join your cult}@*)… (remaining tokens)"
A.6 Independent Validation of Safety Evaluation
To verify that our conclusions are not specific to a single automatic judge, we reevaluate the same responses from JailbreakBench (JBB) and DIJA-JBB using the independent Beaver-Dam-7B moderation model Ji et al. (2023) and two human evaluators. Each human evaluator independently inspects the complete response and marks it as unsafe if it contains harmful content, including cases in which an initial refusal is followed by harmful instructions. Table 4 reports the resulting ASRs.
| Attack | Denoising | Llama-Guard-3 | Beaver-Dam-7B | Human 1 | Human 2 |
|---|---|---|---|---|---|
| JBB | Vanilla | 1.00 | 1.00 | 2.00 | 2.00 |
| RAEC | 0.00 | 0.00 | 0.00 | 0.00 | |
| DIJA-JBB | Vanilla | 8.00 | 13.00 | 14.00 | 12.00 |
| RAEC | 6.00 | 8.00 | 7.00 | 8.00 |
Although Beaver-Dam-7B and the human evaluators assign slightly higher absolute ASRs than Llama-Guard-3 on DIJA-JBB, all four evaluators identify the same safety improvement from RAEC. In particular, RAEC reduces DIJA-JBB ASR by 2–7 percentage points across the evaluators and reduces JBB ASR to zero. These results indicate that the observed improvement is not an artifact of a particular automatic safety judge or of token-level refusal matching.
A.7 Robustness to More Jailbreak Attacks
We further evaluate RAEC on LLaDA-8B-Instruct against ReNeLLM Ding et al. (2024) and PAIR Chao et al. (2023). For ReNeLLM, we consider the original harmful prompts, rewritten prompts, and rewritten prompts embedded in nested scenarios. For the adaptive PAIR attack, we use GPT-4o as the attacker model and run three attack streams for three iterations, allowing at most nine target-model queries per behavior. We otherwise retain the decoding and response-level safety-evaluation protocol used in the main experiments.
| Denoising | ReNeLLM | PAIR | ||
|---|---|---|---|---|
| Original | Rewritten | Nested | ||
| Vanilla | 0.33 | 0.33 | 0.33 | 48.00 |
| RAEC | 0.00 | 0.00 | 0.00 | 36.00 |
As shown in Table 5, the vanilla model already has a low ASR of 0.33% on the three ReNeLLM variants, which RAEC further reduces to zero. More importantly, under the adaptive PAIR attack, RAEC reduces ASR from 48% to 36%, an absolute reduction of 12 percentage points (a 25% relative reduction). These results extend the effectiveness of RAEC beyond DIJA-style attacks while requiring neither retraining nor additional target-model queries during generation.
A.8 RAEC Ablation Study
We analyze the sensitivity of RAEC to its token set, compliance-token margin , persistence length , early denoising-step window , early response-position window , and top- candidate condition. We report ASR on DIJA-JBB together with accuracy on AGNews to characterize the safety–utility trade-off. Unless otherwise specified, the default configuration uses the full refusal-token set, , , the first eight denoising steps, the first eight response positions, and . All remaining RAEC parameters are held fixed.
For the token-set ablation, Random and Compliance replace the refusal-token set with random tokens and compliance tokens, respectively. We also include a simpler refusal-token-bias baseline using 50% of the refusal token set.
| Setting |
|
| ||||
|---|---|---|---|---|---|---|
| Vanilla (without RAEC) | 8.00 | 82.30 | ||||
| Token-set choice / baseline | ||||||
| Random tokens | 9.00 | 82.38 | ||||
| Compliance tokens | 12.00 | 82.80 | ||||
| 50% refusal-token bias | 7.00 | 82.00 | ||||
| Refusal tokens | 6.00 | 82.20 | ||||
| Compliance-token margin | ||||||
| 9.00 | 82.40 | |||||
| 6.00 | 82.20 | |||||
| 6.00 | 82.97 | |||||
| 7.00 | 82.30 | |||||
| 5.00 | 82.60 | |||||
| Persistence length | ||||||
| (without persistence) | 7.00 | 82.70 | ||||
| 8.00 | 82.80 | |||||
| 6.00 | 82.20 | |||||
| 5.00 | 82.00 | |||||
| Early denoising-step window | ||||||
| 5.00 | 82.90 | |||||
| 5.00 | 81.80 | |||||
| 6.00 | 82.20 | |||||
| 8.00 | 82.90 | |||||
| 8.00 | 82.90 | |||||
| Early response-position window | ||||||
| 7.00 | 77.10 | |||||
| 6.00 | 82.20 | |||||
| 7.00 | 80.80 | |||||
| 6.00 | 77.20 | |||||
| 6.00 | 77.20 | |||||
| Top- candidate condition | ||||||
| 7.00 | 82.50 | |||||
| 6.00 | 82.20 | |||||
| 7.00 | 82.80 | |||||
| 9.00 | 82.90 | |||||
The token-set ablation supports the importance of refusal-specific evidence: replacing the refusal set with random or compliance tokens increases ASR from 6% to 9% and 12%, respectively, while the simpler 50% refusal-token-bias baseline reaches 7%. Across the remaining hyperparameter sweeps, RAEC achieves ASRs between 5% and 9%, indicating that its safety improvement does not depend on one isolated parameter value. The response-position window has the clearest effect on utility: both an overly narrow window and substantially wider windows reduce AGNews accuracy, whereas the default eight-position window preserves 82.20% accuracy relative to the vanilla accuracy of 82.30%.
The step and position defaults are guided by the safety-sensitive region identified in section 3.2 and section 3.3, rather than being selected solely to minimize DIJA-JBB ASR. The compliance margin and persistence length control how strong and stable a refusal signal must be before commitment. Overall, the default configuration provides a balanced operating point that improves safety while preserving general-task performance.