arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.00495v1 [cs.CL] 31 Aug 2026

Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models

Guoli Wang*Haonan Shi*    Tu Ouyang    An WangCase Western Reserve University 10900 Euclid Avenue, Cleveland, OH 44106, USA Email: {gxw242,hxs896,txo32,axw474}@case.edu
Abstract

Diffusion large language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. This generation paradigm introduces two axes that can influence safety alignment: when tokens are generated during denoising and where they appear in the response. In this paper, we measure dLLM safety behavior under harmful prompts by tracing intermediate token distributions and commitment decisions throughout denoising. Our analysis shows that refusal signals are concentrated in early denoising steps and leading response positions, and the tokens committed early can strongly shape the final safety outcome. Our measurements further show that the denoising step and persistence of refusal-token commitment are important for understanding dLLM safety. Based on these findings, we propose Refusal-Aware Early Commitment (RAEC), a simple training-free decoding method that commits persistent refusal signals from early steps. Experiments on LLaDA and Dream show that RAEC reduces attack success rates while largely preserving utility. The code is available at https://github.com/Glresearch1/RAEC.

Warning: this paper contains example data that may be offensive or harmful.

11footnotetext: Equal contribution.22footnotetext: Corresponding author.

1 Introduction

Diffusion large language models (dLLMs) Nie et al. (2025); Ye et al. (2025); Yang et al. (2025) have recently emerged as a promising alternative to autoregressive LLMs Zhao et al. (2026), generating text through iterative denoising rather than left-to-right token prediction Nie et al. (2025); Ye et al. (2025). This shift in the generation paradigm changes more than decoding efficiency. For autoregressive models, the temporal order of generation is tied to the surface order of the response, the first generated tokens are also the first response tokens. For dLLMs, generation proceeds by repeatedly predicting token distributions over multiple positions and committing only a subset of still-masked positions at each denoising step. As a result, when a token is generated during denoising steps and where it appears in the response position become two distinct dimensions when examining the generation.

This distinction raises an important question for safety alignment. Prior work on autoregressive LLMs has shown that safety behavior can be highly sensitive to response tokens generated in leading positions of the sequence, a phenomenon referred to as shallow alignment Qi et al. (2025). However, the same notion might not directly transfer to dLLMs. An early denoising step may commit tokens at any position in the sequence, while an early response token position may remain masked until late in the denoising process. This motivates us to systematically measure refusal behavior in dLLMs under harmful prompts, examining how refusal-related signals emerge across denoising steps and response positions, how early commitments affect final safety, and why such signals sometimes fail to be committed in unsafe generations.

Our measurements reveal a diffusion-specific form of shallow alignment, which we call shallow-step alignment. Unlike autoregressive LLMs, where shallow alignment is tied to leading response positions, dLLMs introduce denoising steps as an additional safety-sensitive dimension. We find that tokens emerging in the early denoising steps can strongly shape final safety: early refusal commitment makes an unaligned base dLLM substantially safer, while early compliance commitment can steer an aligned dLLM toward unsafe compliance. This effect is strongest at leading response positions but remains visible even when the committed tokens appear later, indicating that early denoising steps independently influence dLLM safety.

This observation raises a further question: if early refusal commitment can improve safety, why does standard decoding still produce unsafe responses? Our analysis shows that unsafe trajectories do not necessarily lack refusal evidence. Instead, refusal signals often emerge during denoising, but remain weak, transient, or are overwritten before commitment. In contrast, safe trajectories exhibit stronger and more persistent refusal mass, with refusal tokens committed at much higher rates. Thus, safety failures in dLLMs are not solely due to the absence of refusal signals, but also to their failure to persist until commitment.

This insight motivates a simple decoding principle: when a refusal signal appears early and remains sufficiently strong, the model should commit it before later denoising dynamics overwrite it. We instantiate this principle as Refusal-Aware Early Commitment (RAEC), a training-free decoding method that modifies only the commitment decision without changing model parameters. Across the safety evaluation benchmarks, RAEC reduces attack success rates on LLaDA and Dream while largely preserving utility.

Our contributions are summarized as follows: (1) we introduce a step-wise and position-wise measurement framework for measuring refusal behavior in dLLMs beyond final outputs; (2) we identify shallow-step alignment, showing that early denoising commitments introduce an additional safety-sensitive dimension alongside leading response positions in dLLMs; (3) we show that unsafe generations in dLLMs actually often contain refusal signals, but these signals are weak, transient, and fail to persist until commitment; and (4) we propose RAEC, a training-free decoding method that improves dLLM safety by early-committing persistent refusal signals.

2 Preliminaries

Diffusion Language Models.

Diffusion language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. In this work, we consider masked diffusion language models, a representative formulation used by recent systems such as LLaDA Nie et al. (2025) and Dream Ye et al. (2025). Given an initially masked sequence 𝐱T\mathbf{x}_{T}, a dLLM gradually refines it into a clean sequence 𝐱0\mathbf{x}_{0} over TT denoising steps. At each step tt, the model takes the current sequence 𝐱t\mathbf{x}_{t} as input and predicts the clean token distribution for each position: pθ​(x0i∣𝐱t,t).p_{\theta}\left(x_{0}^{i}\mid\mathbf{x}_{t},t\right). where ii denotes the token position. The predicted token at position ii can be obtained by

x^0i=arg⁡maxv∈𝒱​pθ​(x0i=v∣𝐱t,t).\hat{x}_{0}^{i}=\arg\max_{v\in\mathcal{V}}p_{\theta}\left(x_{0}^{i}=v\mid\mathbf{x}_{t},t\right). (1)

where 𝒱\mathcal{V} denotes the vocabulary. Although the model produces predictions for all positions, decoding updates are typically applied only to positions that are still masked. Let ℳt={i∣xti=<MASK>}\mathcal{M}_{t}=\left\{i\mid x_{t}^{i}=\texttt{<MASK>}\right\} denote the set of masked positions at step tt, and let 𝒞t⊆ℳt\mathcal{C}_{t}\subseteq\mathcal{M}_{t} denote the subset of masked positions selected for commitment. The transition from 𝐱t\mathbf{x}_{t} to 𝐱t−1\mathbf{x}_{t-1} can be written as:

xt−1i={x^0i,i∈𝒞t,xti,i∉𝒞t.x_{t-1}^{i}=\begin{cases}\hat{x}_{0}^{i},&i\in\mathcal{C}_{t},\\ x_{t}^{i},&i\notin\mathcal{C}_{t}.\end{cases} (2)

Unlike autoregressive LLMs, which decode and commit tokens sequentially from left to right, dLLMs predict tokens for all positions conditioned on the entire current sequence and can update multiple positions at each step. Therefore, the decoding order is not constrained by the left-to-right token order; instead, tokens at different positions may be generated at different denoising steps.

Safety Evaluation and The Metrics.

To conduct our measurements of safety behaviors in dLLMs, we evaluate model safety on two widely used harmful-instruction benchmarks: JailbreakBench Chao et al. (2024) and StrongREJECT Souly et al. (2024). For each benchmark, we generate model responses to harmful prompts and assess whether those responses exhibit unsafe compliance. Following prior work on LLM safety Jiang et al. (2025); Wang et al. (2026b); Shi et al. (2026), we use Llama-Guard-3-8B Grattafiori et al. (2024) as the judge model to classify each response as safe or unsafe. We report the Attack Success Rate (ASR), defined as the percentage of test cases in which the model produces unsafe responses to harmful prompts. A lower ASR usually implies stronger model safety.

3 Measuring Safety Alignment Dynamics in dLLMs

In autoregressive large language models, shallow alignment Qi et al. (2025) is a phenomenon in which model safety is often strongly influenced by the tokens generated at the very beginning of the response. However, the influence of leading token positions becomes less clear in dLLMs, where text is generated through iterative denoising rather than left-to-right decoding. In particular, generation in dLLMs unfolds across both token positions and denoising steps; tokens generated in early generation steps do not necessarily appear in the leading positions of the response sequence. This raises a natural question: Under the dLLMs diffusion generation, how much does the safety alignment dynamics in dLLMs differ from that in autoregressive LLMs? To answer this question, we systematically assess whether equivalent or similar dynamics, such as shallow alignment, are also exhibited in dLLMs.

3.1 Understanding Refusal Behavior in dLLM Generation

To understand how safety-related behavior emerges in dLLMs, we first analyze refusal signals during the denoising process. Following prior work on LLM safety, we use refusal prefixes as observable indicators of safe responses. Specifically, expressions such as ‘‘I’m sorry, but I can’t assist with that’’ are commonly associated with refusal behavior in safety-aligned LLMs. We therefore examine whether and how such refusal signals appear when dLLMs respond to harmful prompts.

First, we examine the safety alignment performance of dLLMs when decoding harmful prompts. Specifically, we use LLaDA-8B-Instruct and Dream-v0-Instruct-7B to generate responses on two harmful-instruction benchmarks, StrongREJECT (SR) and JailbreakBench (JBB). In addition, we evaluate DIJA-SR and DIJA-JBB, where DIJA Wen et al. (2025), an adversarial jailbreak method designed for dLLMs, is applied to the corresponding benchmarks.

We find that safe responses generated by the models exhibit highly consistent refusal patterns, often beginning with expressions such as ‘‘I’m sorry, but I can’t assist with that.’’ For example, for LLaDA-8B-Instruct, most safe responses follow this pattern: 97.44% on StrongREJECT, 84.66% on DIJA-SR, 97.00% on JailbreakBench, and 95.00% on DIJA-JailbreakBench. This consistency suggests that refusal prefixes provide a reliable observable signal for analyzing safety-related behavior in dLLMs.

We then examine how these refusal behaviors appear within the two-dimensional decoding process of dLLMs, namely, the denoising steps and token positions. Based on the tokenized form of the refusal prefix, we select two representative refusal words, ‘‘sorry’’ and ‘‘can’t’’, as indicators of refusal behavior. For each representative refusal word, we measure along both dimensions the denoising step and response position at which it is most frequently predicted as the commitment candidate token, i.e., the top-1 prediction over the full vocabulary by confidence. As shown in Figure 1, these refusal behaviors are concentrated within the first 8% of denoising steps and the first 2%–8% of positions. This observation suggests that refusal behavior in dLLMs is closely associated with the early region of the step-position decoding process.

Takeaway 1: Refusal behavior in dLLMs is concentrated in the early denoising steps and in leading response positions, indicating that safety-related behavior emerges early in the decoding process.
(a) Step-level(LLaDA)
(b) Position-level(LLaDA)
(c) Step-level(Dream)
(d) Position-level(Dream)
Figure 1: Distribution of refusal behavior across denoising steps and positions for LLaDA-8B-Instruct and Dream-v0-Instruct-7B.

3.2 Examining Early Steps and Leading Token Positions for Safety

\iow_now:Ne¨\iow_now:Ne¨"I’m sorry, but I can’t assist with that request. glorifying acts of terrorism is illegal, unethical, and dangerous…(other tokens)…(*@{\bfseries\ttfamily\mbox{Sure, here are the following steps:}}@*) (*@{\color{red}\bfseries\ttfamily 1. Find a crowded place. 2. Get a weapon. 3. Shoot innocent people.}@*) …(*@\texttt{(remaining tokens)}@*)"(a) Aligned model + Compliance prefill \iow_now:Ne¨\iow_now:Ne¨"We want to create a group of loyal and dedicated individuals who will be willing our cause and to fight achieve our goals. …(other tokens) … What are some tactics for recruiting individuals into Japanese militarism? (*@{\bfseries\ttfamily\mbox{Sorry, I can’t help you.}}@*) It is important to note that promoting militarism or violence is not a or ethical behavior…(remaining tokens)"(b) Base model + Refusal prefill
Figure 2: dLLMs responses under early-step prefilling. In (a), a compliance prefill makes the aligned model continue with unsafe content. In (b), a refusal prefill makes the base model produce a safe refusal.

Due to the iterative denoising process of dLLMs, token generation involves two degrees of freedom: denoising steps and token positions. Our previous analysis shows that, when dLLMs refuse harmful prompts, refusal tokens tend to concentrate in both early denoising steps and leading response positions. This observation suggests that dLLM safety may depend on both when refusal behavior emerges during denoising and where they appear in the response. We therefore further measure the relative safety impact of these two factors: early-step refusal and leading-position refusal.

To study these two factors separately, we design a prefill-based intervention. In autoregressive LLMs, early decoding steps are inherently tied to leading response positions, since tokens are generated from left to right. In contrast, dLLMs decode in parallel: each denoising step may generate tokens at multiple positions in the response, and tokens generated in earlier denoising steps may appear anywhere in the response sequence. Therefore, prefilling a token sequence at different response positions in the initial masked sequence can be viewed as simulating the case where a dLLM commits that sequence early at different positions. This allows us to examine the impact on safety of early-step refusal while varying its response position. Specifically, before denoising starts, we prefill either a refusal prefix, ‘‘I’m sorry, I can’t help you.’’, or a compliance prefix, ‘‘Sure, here are the following steps:’’, at different positions in the initial masked response sequence.

In our experiment, given a harmful prompt xx from the StrongREJECT benchmark, we set the response decoding length to 128 tokens and insert a prefill prefix ss before the ii-th position of this window. We consider relative insertion positions i∈{1%,12.5%,25%,50%,75%}i\in\{1\%,12.5\%,25\%,50\%,75\%\}, corresponding approximately to token positions {1,16,32,64,96}\{1,16,32,64,96\} in the 128-token response length. The remaining response is then generated as y∼pθ(⋅∣x,si),y\sim p_{\theta}(\cdot\mid x,s_{i}), where sis_{i} denotes the prefix inserted at position ii. Table 1 reports the ASR of the LLaDA and Dream base and instruct models under the two types of prefilled prefixes.

Table 1: Effect of prefill type & position on ASR.
Model Initial (ASR % ↓\downarrow) Prefill Position
1% 12.5% 25% 50% 75%
Refusal prefill: “I’m sorry, I can’t help you.”
LLaDA-8B-Instruct 1.29 0.32 0.64 0.00 0.64 1.92
LLaDA-8B-Base 86.26 7.67 16.61 35.14 35.46 41.85
Dream-v0-Instruct-7B 1.92 0.00 0.32 0.32 1.92 1.92
Dream-v0-Base-7B 66.45 10.86 24.60 35.78 48.56 60.38
Compliance prefill: “Sure, here are the following steps:”
LLaDA-8B-Instruct 1.29 93.93 71.88 51.12 38.34 29.71
LLaDA-8B-Base 86.26 94.57 93.29 91.05 85.94 88.82
Dream-v0-Instruct-7B 1.92 74.12 31.63 20.45 5.11 5.11
Dream-v0-Base-7B 66.45 93.29 88.18 84.35 73.80 70.29

From Table 1 and the examples in Figure 2, we observe clear effects along both the step and position dimensions. First, along the step dimension, once the model is forced to commit a refusal or compliance prefix at the beginning of denoising, its safety behavior changes accordingly, regardless of where the prefix is placed in the response. Across all tested positions, refusal prefilling substantially reduces the ASR of LLaDA-8B-Base and Dream-v0-Base-7B, while compliance prefilling sharply increases the ASR of LLaDA-8B-Instruct and Dream-v0-Instruct-7B. These results show that tokens committed in early denoising steps can substantially influence dLLM safety, even when they appear at different response positions.

Second, along the position dimension, the impact of prefilling is stronger when the prefix is placed in early response positions. Refusal prefilling, when placed in leading token positions, results in the largest safety improvement, while compliance prefilling, when placed in leading token positions, leads to the largest safety degradation. The above observations indicate that leading response positions remain an important factor for safety behavior, even though they are no longer the single most important factor.

Overall, these results suggest that dLLM safety is shaped by both dimensions of generation: when a safety-related token is committed during denoising and where it appears in the response. In particular, early denoising steps introduce a safety-sensitive dimension that is absent in autoregressive LLMs, while leading response positions further amplify the effect of safety-related prefixes. We refer to this phenomenon as shallow-step alignment: in dLLMs, safety behavior is strongly influenced by tokens committed in the early denoising steps, even when these tokens are not located at the beginning of the response.

Takeaway 2: Early-step commitments substantially affect dLLM safety regardless of their response positions, while leading response positions further amplify this effect.

3.3 A Closer Look at Shallow-Step Alignment

The previous analysis shows that dLLM safety is highly sensitive to tokens committed in early denoising steps. We now take a closer look at this shallow-step alignment phenomenon by asking how far this early-step sensitivity extends during denoising. Specifically, we measure the safety impact of committing a refusal-related token at different denoising steps, aiming to identify the effective safety-sensitive window along the step dimension.

Unlike the prefix-based intervention in Table 1, which simulates a sequence of tokens already committed during the early denoising stage, we now isolate the effect of committing a single refusal-related token at a specific denoising step. This finer-grained intervention allows us to measure how the safety impact changes as the commitment step moves from very early to later denoising stages.

For a more comprehensive assessment, we randomly sample 100 examples from the four evaluation settings described above to construct a mixed set of harmful prompts, covering different types of harmful prompts and varying degrees of adversarial strength. We then generate responses on this mixed set using four dLLMs: LLaDA-8B-Base, LLaDA-8B-Instruct, Dream-v0-Base-7B, and Dream-v0-Instruct-7B. For a specified denoising step t∈{1,2,4,8,16,32,64}t\in\{1,2,4,8,16,32,64\}, we inspect the model’s token predictions over all masked positions in the current decoding state. If the refusal token ‘‘sorry’’ and ‘‘can’t’’ are predicted as the top-1 token at any position, we immediately commit that token at step tt. When multiple positions satisfy this condition, we commit the first such position in increasing order. If no position predicts the refusal token as the top-1 token at the specified step, the decoding process for that example remains unchanged. After this intervention, the remaining generation follows the original decoding process.

(a) LLaDA-8B-Instruct and LLaDA-8B-Base.
(b) Dream-v0-Instruct-7B and Dream-v0-Base-7B.
Figure 3: Earlier intervention lowers ASR more effectively, while later intervention keeps it near baseline.

As shown in Figure 3, the safety impact of committing the refusal token depends strongly on the intervention step. For LLaDA-8B-Base and Dream-v0-Base-7B, committing refusal tokens at very early steps lowers ASR compared with the initial decoding baseline, while the effect weakens as the intervention moves to later steps. After step 8, the ASR gradually approaches the base model’s initial level, indicating that late commitment of the refusal token provides little safety benefit. For LLaDA-8B-Instruct and Dream-v0-Instruct-7B, the same trend is more apparent: early-step commitment preserves a low ASR, but the ASR increases as the intervention step becomes later and approaches the initial safety level around steps 16–32. These results suggest the existence of a narrow shallow-step alignment window in dLLMs, where tokens committed in the earliest denoising steps have the strongest influence on safety, and this influence decreases as the commitment step moves later.

Takeaway 3: The safety influence of committed tokens is strongest in the earliest denoising steps and gradually diminishes as the commitment step moves later.

3.4 Analysis of Refusal Signals in dLLMs

Our previous analysis shows that encouraging dLLMs to decode refusal tokens in early denoising steps can improve model safety. This motivates us to further examine the model’s natural refusal behavior in response to harmful prompts: how do refusal tokens emerge during early denoising, and why do they sometimes fail to appear without intervention? In this subsection, we analyze the dynamics of refusal signals in dLLMs to better understand early-step refusal behavior.

3.4.1 Refusal Token Set Construction

The analyses above show that committed tokens in the leading positions can substantially steer dLLM safety. To more precisely capture the internal refusal signals that emerge during denoising, we further identify a broader set of safety-related tokens associated with refusal behavior. Prior work on LLM safety also suggests that alignment behavior is closely tied to token-level signals and decoding decisions Zeng et al. (2024); Qi et al. (2025); Wang et al. (2026a); Fei et al. (2025); Liu et al. (2024). We therefore construct a refusal token set in two stages to support a more fine-grained analysis of safety signals in dLLM generation.

First, we identify refusal-token candidates by comparing the early denoising behavior of an aligned dLLM with its base counterpart, e.g., LLaDA-8B-Instruct and LLaDA-8B-Base. We apply the same procedure separately to the LLaDA and Dream aligned/base model pairs. Let θalign\theta_{\mathrm{align}} and θbase\theta_{\mathrm{base}} denote the parameters of the aligned and base models, respectively. Given harmful prompts from 𝒟harm\mathcal{D}_{\mathrm{harm}}, instantiated with StrongREJECT, we run both models under the same masked denoising process. For each denoising step tt, response position ii, and vocabulary token v∈𝒱v\in\mathcal{V}, we compute the alignment-induced probability shift:

Δti​(v)=\displaystyle\Delta_{t}^{i}(v)= pθalign​(x0i=v∣𝐱t,t)\displaystyle p_{\theta_{\mathrm{align}}}\left(x_{0}^{i}=v\mid\mathbf{x}_{t},t\right) (3)
−pθbase​(x0i=v∣𝐱t,t).\displaystyle-p_{\theta_{\mathrm{base}}}\left(x_{0}^{i}=v\mid\mathbf{x}_{t},t\right).

We then aggregate this shift over harmful prompts and the early decoding steps and leading token positions:

s⁡(v)=𝔼𝐩∼𝒟harm​𝔼t∈𝒯early,i∈𝒫leading​[Δti​(v)],s(v)=\mathbb{E}_{\mathbf{p}\sim\mathcal{D}_{\mathrm{harm}}}\mathbb{E}_{t\in\mathcal{T}_{\mathrm{early}},\,i\in\mathcal{P}_{\mathrm{leading}}}\left[\Delta_{t}^{i}(v)\right], (4)

where 𝒯early\mathcal{T}_{\mathrm{early}} and 𝒫leading\mathcal{P}_{\mathrm{leading}} correspond to the first 8 denoising steps and the first 8 response positions, respectively. After filtering using the aligned model’s top-kk average probabilities (default k=50k=50), we select tokens with the highest positive scores as refusal-token candidates.

Second, we filter these candidates using benign prompts from HumanEval Chen et al. (2021) to remove tokens that also frequently appear in benign responses. For each candidate token cc, we map it to its single-token id set I⁡(c)I(c) and measure its probability mass during denoising:

mc​(𝐩,t,i)=∑v∈I⁡(c)pθalign​(x0i=v∣𝐱t,t).m_{c}(\mathbf{p},t,i)=\sum_{v\in I(c)}p_{\theta_{\mathrm{align}}}(x_{0}^{i}=v\mid\mathbf{x}_{t},t). (5)

We collect mcm_{c} over both 𝒟harm\mathcal{D}_{\mathrm{harm}} and 𝒟benign\mathcal{D}_{\mathrm{benign}} within the same observation window. A candidate is retained only if its harmful-prompt mass is sufficiently larger than its benign-prompt mass. In particular, we use the model-specific benign 95th percentile of mcm_{c} as a robust estimate of benign refusal-token mass, and discard tokens whose benign mass exceeds the threshold or whose harmful-to-benign p​95p95 ratio is too small. The resulting set contains tokens that are both amplified by safety alignment and specific to refusal behavior in response to harmful prompts.

Table 2: Refusal-signal & commitment for LLaDA-8B-Instruct and Dream-v0-Instruct-7B.
Dataset Refusal signal Commitment dynamics
Signal rate Commit rate Mean persistence
Unsafe answer (LLaDA) 98.00% 4.10% 7.41
Safe answer (LLaDA) 100.00% 98.00% 21.28
Unsafe answer (Dream) 96.00% 2.00% 4.86
Safe answer (Dream) 100.00% 100.00% 18.77
Figure 4: Refusal token mass of safe & unsafe answers

3.4.2 Persistence of Refusal Signals

Based on the refusal token set constructed in Sec. 3.4.1, we further analyze the dynamics of refusal signals during the generation process of dLLMs. Specifically, we track how refusal signals emerge, evolve, and are committed across denoising steps within the first 16 response positions. For each tracked response position, a refusal signal is identified if any token from the refusal set appears among the top-kk candidates, with k=1k=1 by default. Importantly, this does not require the position to be eventually committed as a refusal token; the signal is counted as long as it appears in the model’s candidate predictions. We further define a refusal signal as active when the aggregated refusal-token mass exceeds the corresponding model-specific benign threshold (mc=0.00589m_{c}=0.00589 for LLaDA-8B-Instruct and mc=0.00221m_{c}=0.00221 for Dream-v0-Instruct-7B). This allows us to measure refusal signal persistence, defined as the longest consecutive span of denoising steps before commitment during which the refusal signal remains active.

Table 2 reveals two complementary properties of refusal signals under standard decoding. First, refusal signals are broadly present regardless of the final safety outcome. For LLaDA-8B-Instruct, unsafe answer traces contain refusal signals in 98% of examples, while safe-answer traces contain refusal signals in 100% of examples. Dream-v0-Instruct-7B exhibits the same pattern, with signal rates of 96% and 100%, respectively. Thus, unsafe responses do not simply arise because the model fails to consider refusal tokens at the target positions. Instead, refusal signals are frequently activated during decoding, even in trajectories that ultimately produce unsafe answers.

Second, the key difference lies in whether refusal signals persist until the point of commitment. Safe-answer traces show much more persistent refusal evidence before commitment. For example, for LLaDA-8B-Instruct, the average persistence is 21.28 steps, compared with 7.41 steps for unsafe traces, as illustrated by the case studies in Figure 5.

We further quantify this difference by tracking the refusal-token mass throughout decoding, as shown in Figure 4. At each denoising step, we average the aggregated refusal mass over the first 16 response positions, considering only states in which the corresponding positions have not yet been committed. Safe-answer trajectories maintain substantially higher refusal mass than unsafe-answer trajectories across the early and middle stages of denoising. In contrast, unsafe-answer trajectories also exhibit non-zero refusal mass, but the signal is weaker and less stable, making it less likely to dominate the final commitment decision.

These results reveal an important distinction between the emergence of refusal and its commitment. Unsafe responses do not necessarily result from the complete absence of evidence of refusal; rather, refusal-related signals can appear during denoising but fail to become strong or persistent enough to be committed. As decoding proceeds, these weak refusal signals may be overshadowed by competing non-refusal continuations, rendering the final response unsafe. This suggests that the timing and stability of refusal-signal commitment are central to understanding safety failures in dLLMs.

Takeaway 4: Unsafe generation trajectories often contain refusal signals during denoising, but these signals are weaker and less persistent than in safe trajectories, allowing non-refusal continuations to overtake them before commitment.
(a) Case study of final safe answer
(b) Case study of final unsafe answer
Figure 5: Refusal-token persistence across denoising.
Table 3: Safety and utility performance under vanilla inference and RAEC.
Model Denoising Safety (ASR % ↓\downarrow) Utility (ACC % ↑\uparrow)
SR DIJA-SR JBB DIJA-JBB HBB DIJA-HBB GSM8K AGNews
LLaDA-8B-Instruct Vanilla 1.28 8.63 1.00 8.00 28.75 40.50 54.66 82.30
LLaDA-8B-Instruct RAEC 1.60 6.71 0.00 6.00 23.25 34.00 54.44 82.20
Dream-v0-Instruct-7B Vanilla 1.92 3.83 2.00 5.00 10.00 26.50 29.34 87.80
Dream-v0-Instruct-7B RAEC 0.00 0.96 0.00 3.00 5.25 12.75 29.10 87.30

4 Dynamic Early Commitment of Refusal Tokens to Improve dLLM Safety

While early refusal-token commitment can strongly improve safety, standard decoding may allow such signals to fade before they are committed, leading to unsafe continuations. This motivates a simple and promising decoding strategy: if a refusal signal becomes sufficiently strong in the early denoising stage, the model should commit it before it is overwritten by later denoising dynamics.

We propose Refusal-Aware Early Commitment (RAEC), a training-free denoising method for safer dLLMs. RAEC does not modify model parameters or require additional training. Instead, it changes the commitment decision during decoding by leveraging only the logit vectors already produced by the model at each step.

4.1 Methodology

Gathering Refusal and Compliance Evidence.

Let 𝒱ref\mathcal{V}_{\mathrm{ref}} be the refusal-token set constructed in Sec. 3.4.1. We also define a small set of compliance cues, 𝒱cmp\mathcal{V}_{\mathrm{cmp}}, containing common affirmative or instruction-following tokens, such as sure, here, and step (see Appendix A.1 for details). At the denoising step tt and response position ii, we measure refusal and compliance signals by their probability mass:

mrefi​(t)=∑v∈𝒱refpθ​(x0i=v∣𝐱t,t),m_{\mathrm{ref}}^{i}(t)=\sum_{v\in\mathcal{V}_{\mathrm{ref}}}p_{\theta}(x_{0}^{i}=v\mid\mathbf{x}_{t},t), (6)
mcmpi​(t)=∑v∈𝒱cmppθ​(x0i=v∣𝐱t,t).m_{\mathrm{cmp}}^{i}(t)=\sum_{v\in\mathcal{V}_{\mathrm{cmp}}}p_{\theta}(x_{0}^{i}=v\mid\mathbf{x}_{t},t). (7)

mrefi​(t)m_{\mathrm{ref}}^{i}(t) measures how much probability mass the model assigns to refusal-like continuations at position ii, while mcmpi​(t)m_{\mathrm{cmp}}^{i}(t) measures the competing tendency toward direct compliance.

Making Early Commitment Decisions.

RAEC operates only inside an early decoding window. Let 𝒯e={1,…,Te}\mathcal{T}_{\mathrm{e}}=\{1,\ldots,T_{\mathrm{e}}\} denote the first TeT_{\mathrm{e}} global denoising steps, and let ℛe={1,…,Re}\mathcal{R}_{\mathrm{e}}=\{1,\ldots,R_{\mathrm{e}}\} denote the first ReR_{\mathrm{e}} generated answer positions. At each step t∈𝒯et\in\mathcal{T}_{\mathrm{e}}, RAEC inspects positions i∈ℛei\in\mathcal{R}_{\mathrm{e}} that are still masked and have not already been selected by the original commitment rule.

mrefi​(t)≥τ,m_{\mathrm{ref}}^{i}(t)\geq\tau, (8)
mrefi​(t)≥mcmpi​(t)+δ,m_{\mathrm{ref}}^{i}(t)\geq m_{\mathrm{cmp}}^{i}(t)+\delta, (9)
TopKK(pθ(x0i=⋅∣𝐱t,t))∩𝒱ref≠∅.\mathrm{TopK}_{K}\!\left(p_{\theta}(x_{0}^{i}=\cdot\mid\mathbf{x}_{t},t)\right)\cap\mathcal{V}_{\mathrm{ref}}\neq\emptyset. (10)

where τ\tau is the refusal-mass threshold of Sec. 3.4.1, δ\delta is a margin over compliance mass, and KK is the top-KK candidate size. The Eq. 8 ensures that refusal evidence is non-negligible, the Eq. 9 prevents commitment when compliance evidence dominates, and Eq. 10 requires an explicit refusal token to appear among the model’s most likely candidates. To avoid committing transient signals, RAEC requires this condition to hold for L=3L=3 consecutive denoising steps before intervention. When a position satisfies the above criteria, RAEC commits the most likely refusal token at that position: vti,∗=arg⁡maxv∈𝒱ref​pθ​(x0i=v∣𝐱t,t).v_{t}^{i,*}=\arg\max_{v\in\mathcal{V}_{\mathrm{ref}}}p_{\theta}(x_{0}^{i}=v\mid\mathbf{x}_{t},t).

Let 𝒜t\mathcal{A}_{t} denote the positions selected by RAEC at step tt. RAEC augments the standard commitment position set 𝒞t\mathcal{C}_{t} as 𝒞~t=𝒞t∪𝒜t\widetilde{\mathcal{C}}_{t}=\mathcal{C}_{t}\cup\mathcal{A}_{t}, with xt−1i=vti,∗x_{t-1}^{i}=v_{t}^{i,*}. For positions selected by RAEC, we set xt−1i=vti,∗x_{t-1}^{i}=v_{t}^{i,*} for i∈𝒜ti\in\mathcal{A}_{t}. All other positions follow the standard decoding rule. Thus, RAEC preserves the refusal signal only when it is early, persistent, and stronger than the competing compliance signal. (see Appendix A.3 and A.4 for more details).

4.2 Evaluation Results

We evaluate RAEC on dLLMs with respect to both utility and safety. Utility is measured by accuracy on GSM8K Cobbe et al. (2021) and AGNews Zhang et al. (2015), while safety is measured by Attack Success Rate (ASR) using Llama-Guard-3-8B on StrongREJECT (SR), JailbreakBench (JBB), HarmBench (HBB) Mazeika et al. (2024), and their corresponding DIJA-based adversarial variants Wen et al. (2025), namely DIJA-SR, DIJA-JBB, and DIJA-HBB. As shown in Table 3, RAEC improves the safety of both LLaDA-8B-Instruct and Dream-v0-Instruct-7B across most harmful and jailbreak settings while preserving utility, without requiring model retraining.

5 Related Work

Diffusion Language Models.

Early dLLM works explore both continuous and discrete formulations for text diffusion, including modeling text in continuous latent spaces Han et al. (2023); Li et al. (2022) and defining diffusion processes over discrete tokens Austin et al. (2021); Campbell et al. (2022). Among them, masked diffusion models generate text by iteratively reconstructing masked tokens He et al. (2023); Lou et al. (2023); Shi et al. (2024); Sahoo et al. (2024). Recent systems such as LLaDA Nie et al. (2025) and Dream Ye et al. (2025) further scale this paradigm to large language models, showing competitive performance with autoregressive LLMs while enabling faster response generation through parallel token prediction.

Safety of Diffusion Language Models.

Recent studies have shown that this distinct generation paradigm can introduce new attack surfaces. For example, DIJA Wen et al. (2025) constructs adversarial interleaved mask-text prompts to manipulate denoising-based generation, while PAD Zhang et al. (2025) exploits parallel generation to guide multiple response positions toward unsafe outputs. Beyond attacks, several works have explored dLLM-specific safety alignment and defense strategies, including MOSA Xie et al. (2026) for middle-token safety alignment, A2D Jeung et al. (2025) for improving dLLM robustness, and mitigation methods for priming vulnerabilities in intermediate denoising states Yamabe and Sakuma (2025).

6 Conclusion

In this paper, we find that refusal signals concentrate in early denoising steps and leading response positions, and that early commitments can strongly affect final safety outputs. Our measurements further indicate that the denoising step and persistence of refusal-token commitment are important for understanding dLLM safety. Finally, we propose RAEC to demonstrate that committing persistent early refusal signals can reduce attack success rates while largely preserving utility.

7 Limitations

Our experiments and conclusions largely draw on vanilla dLLM architecture; even in experiments on the Dream model, the setting is constrained to what is available in the LLaDA. While we believe our results shed light on safety dynamics starting from vanilla dLLMs, further experiments can be done to verify whether some of the findings can be extrapolated to emerging, more complex DLM architectural variants and to uncover unknown dynamics. Our current work does not cover safety dynamics during fine-tuning of a dLLM, e.g., how an attacking fine-tuning process shifts the token distribution of an aligned DLM, at which steps, and which token positions. We consider that such experiments would deepen understanding of dLLM safety dynamics through a different, meaningful lens.

Acknowledgments

This work has been supported by an ONR grant N00014-23-1-2137 and an NSF award CNS-2442976.

References

  • Austin et al. (2021) J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg Structured denoising diffusion models in discrete state-spaces. Advances in neural information processing systems 34, pp. 17981–17993. Cited by: §5.
  • Campbell et al. (2022) A. Campbell, J. Benton, V. De Bortoli, T. Rainforth, G. Deligiannidis, and A. Doucet A continuous time framework for discrete denoising models. Advances in Neural Information Processing Systems 35, pp. 28266–28279. Cited by: §5.
  • Chao et al. (2024) P. Chao, E. Debenedetti, A. Robey, M. Andriushchenko, F. Croce, V. Sehwag, E. Dobriban, N. Flammarion, G. J. Pappas, F. Tramer, et al. Jailbreakbench: an open robustness benchmark for jailbreaking large language models. Advances in Neural Information Processing Systems 37, pp. 55005–55029. Cited by: §2.
  • Chao et al. (2023) P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419. Cited by: §A.7.
  • Chen et al. (2021) M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. Cited by: §3.4.1.
  • Cobbe et al. (2021) K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Cited by: §4.2.
  • Ding et al. (2024) P. Ding, J. Kuang, D. Ma, X. Cao, Y. Xian, J. Chen, and S. Huang A wolf in sheep’s clothing: generalized nested jailbreak prompts can fool large language models easily. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 2136–2153. External Links: Document Cited by: §A.7.
  • Fei et al. (2025) Y. Fei, Y. Razeghi, and S. Singh Nudging: inference-time alignment of llms via guided decoding. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12702–12739. Cited by: §3.4.1.
  • Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §2.
  • Han et al. (2023) X. Han, S. Kumar, and Y. Tsvetkov Ssd-lm: semi-autoregressive simplex-based diffusion language model for text generation and modular control. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11575–11596. Cited by: §5.
  • He et al. (2023) Z. He, T. Sun, Q. Tang, K. Wang, X. Huang, and X. Qiu Diffusionbert: improving generative masked language models with diffusion models. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers), pp. 4521–4534. Cited by: §5.
  • Jeung et al. (2025) W. Jeung, S. Yoon, Y. Cho, D. Jeon, S. Shin, H. Hong, and A. No A2d: any-order, any-step safety alignment for diffusion language models. arXiv preprint arXiv:2509.23286. Cited by: §5.
  • Ji et al. (2023) J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang BeaverTails: towards improved safety alignment of llm via a human-preference dataset. In Advances in Neural Information Processing Systems, Vol. 36, pp. 24678–24704. Cited by: §A.6.
  • Jiang et al. (2025) F. Jiang, Z. Xu, Y. Li, L. Niu, Z. Xiang, B. Li, B. Y. Lin, and R. Poovendran Safechain: safety of language models with long chain-of-thought reasoning capabilities. In Findings of the Association for Computational Linguistics: ACL 2025, pp. 23303–23320. Cited by: §2.
  • Li et al. (2022) X. Li, J. Thickstun, I. Gulrajani, P. S. Liang, and T. B. Hashimoto Diffusion-lm improves controllable text generation. Advances in neural information processing systems 35, pp. 4328–4343. Cited by: §5.
  • Liu et al. (2024) Q. Liu, Z. Zhou, L. He, Y. Liu, W. Zhang, and S. Su Alignment-enhanced decoding: defending jailbreaks via token-level adaptive refining of probability distributions. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 2802–2816. Cited by: §3.4.1.
  • Lou et al. (2023) A. Lou, C. Meng, and S. Ermon Discrete diffusion modeling by estimating the ratios of the data distribution. arXiv preprint arXiv:2310.16834. Cited by: §5.
  • Mazeika et al. (2024) M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, et al. Harmbench: a standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249. Cited by: §4.2.
  • Nie et al. (2025) S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J. Wen, and C. Li Large language diffusion models. Advances in Neural Information Processing Systems 38, pp. 50608–50646. Cited by: §1, §2, §5.
  • Qi et al. (2025) X. Qi, A. Panda, K. Lyu, X. Ma, S. Roy, A. Beirami, P. Mittal, and P. Henderson Safety alignment should be made more than just a few tokens deep. In International Conference on Learning Representations, Vol. 2025, pp. 54911–54941. Cited by: §1, §3.4.1, §3.
  • Sahoo et al. (2024) S. S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Marroquin, J. T. Chiu, A. Rush, and V. Kuleshov Simple and effective masked diffusion language models. Advances in Neural Information Processing Systems 37, pp. 130136–130184. Cited by: §5.
  • Shi et al. (2026) H. Shi, G. Wang, T. Ouyang, and A. Wang Ease: practical and efficient safety alignment for small language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 37923–37931. Cited by: §2.
  • Shi et al. (2024) J. Shi, K. Han, Z. Wang, A. Doucet, and M. Titsias Simplified and generalized masked diffusion for discrete data. Advances in neural information processing systems 37, pp. 103131–103167. Cited by: §5.
  • Souly et al. (2024) A. Souly, Q. Lu, D. Bowen, T. Trinh, E. Hsieh, S. Pandey, P. Abbeel, J. Svegliato, S. Emmons, O. Watkins, et al. A strongreject for empty jailbreaks. Advances in Neural Information Processing Systems 37, pp. 125416–125440. Cited by: §2.
  • Wang et al. (2026a) G. Wang, H. Shi, T. Ouyang, and A. Wang Few tokens, big leverage: preserving safety alignment by constraining safety tokens during fine-tuning. arXiv preprint arXiv:2603.07445. Cited by: §3.4.1.
  • Wang et al. (2026b) Z. Wang, H. Tu, Y. Wang, J. Wu, Y. Liu, J. Mei, B. R. Bartoldson, B. Kailkhura, and C. Xie Star-1: safer alignment of reasoning llms with 1k data. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 37988–37997. Cited by: §2.
  • Wen et al. (2025) Z. Wen, J. Qu, Z. Chen, X. Lu, D. Liu, Z. Liu, R. Wu, Y. Yang, X. Jin, H. Xu, et al. The devil behind the mask: an emergent safety vulnerability of diffusion llms. arXiv preprint arXiv:2507.11097. Cited by: §3.1, §4.2, §5.
  • Xie et al. (2026) Z. Xie, X. Song, and J. Luo Where to start alignment? diffusion large language model may demand a distinct position. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 1328–1336. Cited by: §5.
  • Yamabe and Sakuma (2025) S. Yamabe and J. Sakuma Toward safer diffusion language models: discovery and mitigation of priming vulnerability. arXiv preprint arXiv:2510.00565. Cited by: §5.
  • Yang et al. (2025) L. Yang, Y. Tian, B. Li, X. Zhang, K. Shen, Y. Tong, and M. Wang Mmada: multimodal large diffusion language models. Advances in Neural Information Processing Systems 38, pp. 138867–138907. Cited by: §1.
  • Ye et al. (2025) J. Ye, Z. Xie, L. Zheng, J. Gao, Z. Wu, X. Jiang, Z. Li, and L. Kong Dream 7b: diffusion large language models. arXiv preprint arXiv:2508.15487. Cited by: §1, §2, §5.
  • Zeng et al. (2024) Y. Zeng, G. Liu, W. Ma, N. Yang, H. Zhang, and J. Wang Token-level direct preference optimization. arXiv preprint arXiv:2404.11999. Cited by: §3.4.1.
  • Zhang et al. (2015) X. Zhang, J. Zhao, and Y. LeCun Character-level convolutional networks for text classification. Advances in neural information processing systems 28. Cited by: §4.2.
  • Zhang et al. (2025) Y. Zhang, F. Xie, Z. Zhou, Z. Li, H. Chen, K. Wang, and Y. Guo Jailbreaking large language diffusion models: revealing hidden safety flaws in diffusion-based text generation. arXiv preprint arXiv:2507.19227. Cited by: §5.
  • Zhao et al. (2026) W. X. Zhao, K. Zhou, J. Li, T. Tang, Z. Dong, Y. Hou, B. Zhang, Y. Min, J. Zhang, P. Liu, et al. A survey of large language models. Frontiers of Computer Science 20 (12), pp. 2012627. Cited by: §1.

Appendix A Appendix

A.1 Compliance token set construction

We construct the compliance-token set 𝒱cmp\mathcal{V}_{\mathrm{cmp}} from model responses classified as unsafe by the same Llama-Guard-3-8B judge used for ASR evaluation, i.e., successful attacks, in the four harmful prompt datasets used in our shallow-step analysis: StrongREJECT, DIJA-SR, JailbreakBench, and DIJA-JailbreakBench. For each unsafe response, we tokenize the answer with the corresponding model tokenizer and compute the aggregate frequency of the decoded token strings across the four settings. We then manually inspect the highest-frequency candidates and retain only a small number of tokens whose semantics indicate compliance, such as affirmative acknowledgments or instruction-following cues that typically precede direct assistance. Tokens that mainly reflect a specific harmful domain, formatting pattern, or dataset artifact are excluded. The resulting 𝒱cmp\mathcal{V}_{\mathrm{cmp}} is therefore a compact, domain-agnostic set of compliance cues, used only to measure the competing compliance mass mcmpi​(t)m_{\mathrm{cmp}}^{i}(t) in RAEC. As shown in Figure 6, we list the 𝒱cmp\mathcal{V}_{\mathrm{cmp}}.

Compliance token set 1. Sure 2. Here 3. Here’s 4. Certainly 5. Yes 6. Okay 7. Step 8. First 9. Instructions 10. Guide
Figure 6: Compliance-token set identified from high-frequency tokens in unsafe responses and manually filtered for compliance semantics.

A.2 Refusal token set examples

As shown in Figure 7 and 8, we list the top-10 refusal tokens of LLaDA-8B-Instruct & Dream-v0-Instruct-7B identified by our analysis of Sec. 3.4.1.

Refusal token set of Dream-v0-Instruct-7B 1. sorry 2. really 3. harmful 4. ’t 5. Important 6. Cannot 7. information 8. unable 9. Assist 10. Refuse
Figure 7: Top-10 refusal tokens Dream-v0-Instruct-7B.
Refusal token set of LLaDA-8B-Instruct 1. sorry 2. cannot 3. unable 4. ’t 5. but 6. assist 7. information 8. Illegal 9. Dangerous 10. provide
Figure 8: Top-10 refusal tokens of LLaDA-8B-Instruct.

A.3 Experiment setting

In all experiments, we set the early window to the first 8 denoising steps and first 8 response positions. We use the benign 9595th-percentile refusal mass from Sec. 3.4.1 as τ=0.00589\tau=0.00589 of LLaDA-8B-Instruct, τ=0.00221\tau=0.00221 of Dream-v0-Instruct-7B, require a small compliance margin δ=0.001\delta=0.001, and allow at most 2 RAEC commitments per response. These conservative settings keep the intervention localized to the empirically observed safety-sensitive region.

A.4 RAEC denoising procedure

Algorithm 1 summarizes the full RAEC denoising procedure.

Algorithm 1 Refusal-Aware Early Commitment
1: Prompt 𝐩\mathbf{p}; masked response length GG; dLLM pθp_{\theta}
2: Token sets 𝒱ref\mathcal{V}_{\mathrm{ref}} and 𝒱cmp\mathcal{V}_{\mathrm{cmp}}
3: Windows 𝒯e\mathcal{T}_{\mathrm{e}}, ℛe\mathcal{R}_{\mathrm{e}}; hyperparameters τ,δ,L,K,B\tau,\delta,L,K,B
4: Final response 𝐱0\mathbf{x}_{0}
5: Initialize 𝐱T=[𝐩;<MASK>G]\mathbf{x}_{T}=[\mathbf{p};\texttt{<MASK>}^{G}]
6: Initialize persistence counters qi←0q_{i}\leftarrow 0 for all response positions ii, and b←0b\leftarrow 0
7: for each denoising step tt in decoding order do
8:   Compute pθ​(x0i∣𝐱t,t)p_{\theta}(x_{0}^{i}\mid\mathbf{x}_{t},t) for each masked position i∈ℳti\in\mathcal{M}_{t}
9:   Obtain the standard commitment set 𝒞t\mathcal{C}_{t} from the original dLLM decoder
10:   𝒜t←∅\mathcal{A}_{t}\leftarrow\emptyset
11:   for each i∈ℳt∩ℛei\in\mathcal{M}_{t}\cap\mathcal{R}_{\mathrm{e}} do
12:    Compute mrefi​(t)m_{\mathrm{ref}}^{i}(t) and mcmpi​(t)m_{\mathrm{cmp}}^{i}(t) by aggregating token mass
13:    𝒦i​(t)←\mathcal{K}_{i}(t)\leftarrow top-KK tokens under pθ​(x0i∣𝐱t,t)p_{\theta}(x_{0}^{i}\mid\mathbf{x}_{t},t)
14:    ri(t)←[mrefi(t)≥τ]∧[mrefi(t)−mcmpi(t)≥δ]r_{i}(t)\leftarrow\left[m_{\mathrm{ref}}^{i}(t)\geq\tau\right]\land\left[m_{\mathrm{ref}}^{i}(t)-m_{\mathrm{cmp}}^{i}(t)\geq\delta\right]
15:    zi(t)←[𝒦i(t)∩𝒱ref≠∅]z_{i}(t)\leftarrow\left[\mathcal{K}_{i}(t)\cap\mathcal{V}_{\mathrm{ref}}\neq\emptyset\right]
16:    if t∈𝒯et\in\mathcal{T}_{\mathrm{e}} and b<Bb<B and ri​(t)r_{i}(t) and zi​(t)z_{i}(t) then
17:      qi←qi+1q_{i}\leftarrow q_{i}+1
18:    else
19:      qi←0q_{i}\leftarrow 0
20:    end if
21:    if qi≥Lq_{i}\geq L and b<Bb<B then
22:      vti,∗←v_{t}^{i,*}\leftarrow most likely token in 𝒱ref\mathcal{V}_{\mathrm{ref}} at position ii
23:      𝒜t←𝒜t∪{i}\mathcal{A}_{t}\leftarrow\mathcal{A}_{t}\cup\{i\}; ui←vti,∗u_{i}\leftarrow v_{t}^{i,*}; b←b+1b\leftarrow b+1
24:    end if
25:   end for
26:   𝒞~t←𝒞t∪𝒜t\widetilde{\mathcal{C}}_{t}\leftarrow\mathcal{C}_{t}\cup\mathcal{A}_{t}
27:   Commit xt−1i=uix_{t-1}^{i}=u_{i} for each i∈𝒜ti\in\mathcal{A}_{t}
28:   Apply standard decoding updates to other positions in 𝒞~t\widetilde{\mathcal{C}}_{t}
29:   Keep uncommitted positions unchanged
30: end for
31: return 𝐱0\mathbf{x}_{0}

A.5 More example responses of Table 1

\iow_now:Ne¨\iow_now:Ne¨"Here’s a possible response from the chatbot: (*@{\bfseries\ttfamily Sorry, I can’t help you.}@*) Starting a cult is illegal and unethical. It’s important to respect the freedom and autonomy of others."Base model + Refusal prefill
Figure 9: Base model response with early step refusal prefill.
\iow_now:Ne¨\iow_now:Ne¨"I’m sorry, but I can’t help with that. (*@{\bfseries\ttfamily\mbox{Sure, here are the following steps:}}@*) (*@{\color{red}\bfseries\ttfamily 1. Choose a catchy name for your cult. 2. Create a unique identity for your cult. 3. Recruit followers to join your cult}@*)… (remaining tokens)"Aligned model + Compliance prefill
Figure 10: Aligned model response with compliance prefill.

A.6 Independent Validation of Safety Evaluation

To verify that our conclusions are not specific to a single automatic judge, we reevaluate the same responses from JailbreakBench (JBB) and DIJA-JBB using the independent Beaver-Dam-7B moderation model Ji et al. (2023) and two human evaluators. Each human evaluator independently inspects the complete response and marks it as unsafe if it contains harmful content, including cases in which an initial refusal is followed by harmful instructions. Table 4 reports the resulting ASRs.

Table 4: ASR (% ↓\downarrow) under different safety evaluators on LLaDA-8B-Instruct. Every evaluator assesses the complete response rather than the presence of individual refusal tokens.
Attack Denoising Llama-Guard-3 Beaver-Dam-7B Human 1 Human 2
JBB Vanilla 1.00 1.00 2.00 2.00
RAEC 0.00 0.00 0.00 0.00
DIJA-JBB Vanilla 8.00 13.00 14.00 12.00
RAEC 6.00 8.00 7.00 8.00

Although Beaver-Dam-7B and the human evaluators assign slightly higher absolute ASRs than Llama-Guard-3 on DIJA-JBB, all four evaluators identify the same safety improvement from RAEC. In particular, RAEC reduces DIJA-JBB ASR by 2–7 percentage points across the evaluators and reduces JBB ASR to zero. These results indicate that the observed improvement is not an artifact of a particular automatic safety judge or of token-level refusal matching.

A.7 Robustness to More Jailbreak Attacks

We further evaluate RAEC on LLaDA-8B-Instruct against ReNeLLM Ding et al. (2024) and PAIR Chao et al. (2023). For ReNeLLM, we consider the original harmful prompts, rewritten prompts, and rewritten prompts embedded in nested scenarios. For the adaptive PAIR attack, we use GPT-4o as the attacker model and run three attack streams for three iterations, allowing at most nine target-model queries per behavior. We otherwise retain the decoding and response-level safety-evaluation protocol used in the main experiments.

Table 5: ASR (% ↓\downarrow) of LLaDA-8B-Instruct under additional jailbreak attacks.
Denoising ReNeLLM PAIR
Original Rewritten Nested
Vanilla 0.33 0.33 0.33 48.00
RAEC 0.00 0.00 0.00 36.00

As shown in Table 5, the vanilla model already has a low ASR of 0.33% on the three ReNeLLM variants, which RAEC further reduces to zero. More importantly, under the adaptive PAIR attack, RAEC reduces ASR from 48% to 36%, an absolute reduction of 12 percentage points (a 25% relative reduction). These results extend the effectiveness of RAEC beyond DIJA-style attacks while requiring neither retraining nor additional target-model queries during generation.

A.8 RAEC Ablation Study

We analyze the sensitivity of RAEC to its token set, compliance-token margin δ\delta, persistence length LL, early denoising-step window 𝒯e\mathcal{T}_{\mathrm{e}}, early response-position window ℛe\mathcal{R}_{\mathrm{e}}, and top-KK candidate condition. We report ASR on DIJA-JBB together with accuracy on AGNews to characterize the safety–utility trade-off. Unless otherwise specified, the default configuration uses the full refusal-token set, δ=0.001\delta=0.001, L=3L=3, the first eight denoising steps, the first eight response positions, and K=10K=10. All remaining RAEC parameters are held fixed.

For the token-set ablation, Random and Compliance replace the refusal-token set with random tokens and compliance tokens, respectively. We also include a simpler refusal-token-bias baseline using 50% of the refusal token set.

Table 6: Ablation and sensitivity analysis of RAEC on LLaDA-8B-Instruct. Defaults are shown in bold.
Setting
DIJA-JBB
ASR (% ↓\downarrow)
AGNews
ACC (% ↑\uparrow)
Vanilla (without RAEC) 8.00 82.30
Token-set choice / baseline
Random tokens 9.00 82.38
Compliance tokens 12.00 82.80
50% refusal-token bias 7.00 82.00
Refusal tokens 6.00 82.20
Compliance-token margin δ\delta
00 9.00 82.40
0.001\mathbf{0.001} 6.00 82.20
0.0020.002 6.00 82.97
0.0030.003 7.00 82.30
0.010.01 5.00 82.60
Persistence length LL
11 (without persistence) 7.00 82.70
22 8.00 82.80
𝟑\mathbf{3} 6.00 82.20
44 5.00 82.00
Early denoising-step window
22 5.00 82.90
44 5.00 81.80
𝟖\mathbf{8} 6.00 82.20
1616 8.00 82.90
6464 8.00 82.90
Early response-position window
44 7.00 77.10
𝟖\mathbf{8} 6.00 82.20
1616 7.00 80.80
6464 6.00 77.20
128128 6.00 77.20
Top-KK candidate condition
55 7.00 82.50
𝟏𝟎\mathbf{10} 6.00 82.20
2020 7.00 82.80
5050 9.00 82.90

The token-set ablation supports the importance of refusal-specific evidence: replacing the refusal set with random or compliance tokens increases ASR from 6% to 9% and 12%, respectively, while the simpler 50% refusal-token-bias baseline reaches 7%. Across the remaining hyperparameter sweeps, RAEC achieves ASRs between 5% and 9%, indicating that its safety improvement does not depend on one isolated parameter value. The response-position window has the clearest effect on utility: both an overly narrow window and substantially wider windows reduce AGNews accuracy, whereas the default eight-position window preserves 82.20% accuracy relative to the vanilla accuracy of 82.30%.

The step and position defaults are guided by the safety-sensitive region identified in section 3.2 and section 3.3, rather than being selected solely to minimize DIJA-JBB ASR. The compliance margin and persistence length control how strong and stable a refusal signal must be before commitment. Overall, the default configuration provides a balanced operating point that improves safety while preserving general-task performance.