arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.00061v2 [cs.LG] 05 Sep 2026

ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration

Yuchen Bao    Chao Wen    Haowei Wang    Ruoxin Chen    Donghao Luo    Jiahui Zhan    Wenjian Huang    Shen Chen    Yiting Wang    Taiping Yao    Chengjie Wang    Shouhong Ding    Jianguo Zhang ††thanks: Corresponding author.
Abstract

Reward post-training of diffusion generators inevitably concentrates probability mass on a few reward-favored modes, a mode collapse that erases within-prompt diversity. Existing methods for mitigating collapse rely on external signals or interfaces, augmenting the reward with perceptual objectives, adjusting reference regularization, or modifying the text encoder, but none repairs an adapter that has already collapsed while preserving the acquired reward. We observe that online post-training primarily reallocates probability mass over capabilities inherited from pretraining rather than learning new visual content. Collapse is therefore suppression, not deletion, and can be reversed from within the generator. We propose ReNFT, which repairs a high-reward, low-diversity adapter through internal probability-mass recalibration. Unconditional probes first prioritize “anti-hub” prompts where the prompt-independent bias is easiest to expose. Two policy-dominated mixed routes then generate matched counterfactual proposals from the same prompt and initial noise, one probing the frozen base direction for suppressed alternatives and the other exposing the post-trained unconditional tendency. Reward ranking with an adaptive flipping guard assigns pull and push roles, and a joint-and-paired NFT update realizes the repair. On PickScore and GenEval, ReNFT retains 98.9% and 99.0% of NFT’s reward while improving DreamSim-Div by 58.8% and 55.0%, respectively, offering a complementary alternative to external interventions.

1Southern University of Science and Technology

2Tencent Youtu Lab

1 Introduction

Reward-based post-training has become a practical way to align diffusion and flow generators beyond pretraining. Preference optimization and policy-gradient methods, represented by DiffusionDPO and Flow-GRPO, improve alignment by increasing the probability of reward-preferred outputs (Wallace et al. 2024; Liu et al. 2025). DiffusionNFT (Negative-aware Fine-Tuning, NFT) (Zheng et al. 2026a) makes this alignment markedly more efficient: it replaces likelihood-ratio policy gradients with forward-process reconstruction, requires no classifier-free guidance, and converges much faster. The same feedback loop, however, inevitably concentrates probability mass; the more efficient the optimization, the more severe the concentration: whenever a narrow visual mode repeatedly receives high reward, different initial noises under the same prompt are mapped to similar outputs. Reward differences among such semantically equivalent samples encode the preference of the reward model rather than genuine quality, so optimizing them only sharpens the preference. The result is a high-reward adapter whose within-prompt structural and stylistic diversity has collapsed, most severely under the fastest optimizer, NFT.

Refer to caption
Figure 1: Conceptual view of probability-mass recalibration. Pretraining spreads probability across the modes of each prompt, while unconditional behavior spans their marginal range. NFT collapses each prompt onto a reward-near mode and unconditional behavior onto a fixed mode, suppressing alternatives. ReNFT restores unconditional coverage and reward-boundary modes while leaving low-quality modes outside suppressed. Gray dashed lines compare the prompts’ highest-reward mode probabilities.

Existing methods for mitigating collapse mainly intervene at three positions. Reward-shaping approaches augment the reward with external perceptual signals (Liu et al. 2026a; Tan et al. 2026; Liu et al. 2026b), but can only reweight among samples the policy still generates. Regularization-based approaches align the policy toward the base model (Liu et al. 2026a; He et al. 2025), tying achievable reward to the base ceiling. Interface-level methods modify text representations (Hu et al. 2026; Chen et al. 2026), but depend on the architecture of the reward model. We therefore ask: can an already-collapsed adapter be repaired using only distributions already inside the generator, without external diversity objectives or text-encoder modification, and without sacrificing its acquired reward?

Our answer begins with a distinction: pretraining learns from external images, whereas online post-training supervises the generator with its own scored samples, so it reallocates probability mass over capabilities inherited from the base model, amplifying reward-favored modes while starving alternatives. Collapse is therefore suppression, not deletion: suppressed modes remain reachable through internal routes of the generator, and restoring their probability mass requires no new visual knowledge. We refer to this operational repair view as probability-mass recalibration.

Recalibration first requires an internal signal that exposes where the mass has gone. The post-trained unconditional route provides such a readout: it shares the updated parameters but drops the prompt condition, exposing what the adapter injects without being asked. This readout is diagnostic rather than a training target: the repair must act on conditional generations, and optimizing the unconditional direction directly offers no handle on per-prompt structure. We therefore use it in two indirect ways: locating prompts far from the exposed tendency, and constructing counterfactual samples inside ordinary conditional trajectories, where the tendency becomes testable by the original reward model.

These observations motivate ReNFT (Repair NFT), which repairs high-reward, low-diversity adapters through internal probability-mass recalibration. To expose the bias where it is most distinguishable, unconditional probes first prioritize “anti-hub” prompts, those farthest from the unconditional hub, the shared mode toward which the unconditional distribution concentrates under reward hacking. To obtain candidates without any external signal, two policy-dominated mixed routes generate matched proposals from the same prompt and initial noise: one probes the frozen base direction for suppressed alternatives, while the other exposes the post-trained unconditional tendency. The original reward then ranks the pair into pull and mirrored push targets, and an adaptive flipping guard keeps at least half of the pull targets on the base-probing route. A native-noise joint update then anchors the preference at the shared trajectory origin, and fresh-noise paired updates propagate it across intermediate noise levels. A frozen image–text encoder is used only to prioritize prompts; it never enters the reward, advantage, or repair loss.

In summary, our contributions are as follows:

  • •

    We formulate reward-induced mode collapse as an internal probability-mass reallocation problem and study the repair of a severely reward-hacked, high-reward adapter rather than diversity-aware training from scratch.

  • •

    We identify the post-trained unconditional route as a free internal readout of reward bias and exploit it in two ways: prioritizing “anti-hub” prompts farthest from the exposed tendency, and injecting the tendency into conditional trajectories to turn a hidden bias into an explicit, reward-verifiable candidate.

  • •

    We realize probability-mass recalibration through two internal routes of the same generator: branched from the same prompt and initial noise, the frozen base route expands probability mass back toward suppressed alternatives, while the unconditional route exposes the learned bias in a mixed endpoint as a mirrored-push candidate; pull and push follow the reward gap of each pair rather than route identity. This pair construction is training-only; the repaired adapter samples with the standard conditional forward pass, preserving 98.9–99.0% of NFT’s reward while improving DreamSim-Div (Fu et al. 2023) by 58.8% and 55.0% on PickScore (Kirstain et al. 2023) and GenEval (Ghosh, Hajishirzi, and Schmidt 2023), respectively.

2 Related Work

Refer to caption
Figure 2: Training dynamics and ablations. (a)(b): PickScore reward and DreamSim-Div; (c)(d): GenEval reward and DreamSim-Div. ReNFT branches from the hacked checkpoint HH (star) and repairs for 50 steps; the dashed line marks the base generator. (e)(f): mixed-route pattern and anti-hub ablations after 50 repair steps from the same HH; the dashed line marks NFT at HH.

DiffusionDPO (Wallace et al. 2024) adapts preference optimization to reward post-training of diffusion and flow generators; Flow-GRPO (Liu et al. 2025), DenseGRPO (Deng et al. 2026), and AWM (Xue et al. 2026) improve group-relative optimization, reward density, or stability; and DiffusionNFT (Zheng et al. 2026a) transfers supervision to forward-process reconstruction; Liu, He, and Li (2026) survey this line. Subsequent variants refine clipping, credit assignment, self-correction, and distillation (Ping et al. 2026; Tong et al. 2026; Qin et al. 2026; Li et al. 2026b; Fang et al. 2026; Go et al. 2026; Li et al. 2025). Without a mechanism preserving broad support, repeatedly reinforcing self-generated high-reward samples concentrates probability mass on a few modes, i.e., mode collapse, as also observed for on-policy, reverse-divergence objectives (Zheng et al. 2026b). Interventions differ mainly in where they inject the diversity signal.

Reward-shaping approaches. A first line augments the reward or advantage with perceptual diversity signals from external models. DiverseGRPO (Liu et al. 2026a) clusters same-prompt samples in CLIP space and adds an exploration bonus inversely proportional to cluster size. PEC (Tan et al. 2026) replaces policy entropy with a perceptual-entropy proxy from rollout states. DRIFT (Liu et al. 2026b) combines reward-concentrated rollout selection, prompt perturbation, and potential-based shaping. Related work replaces the reward signal itself with pairwise win rates, gated or adversarial advantages, or set-level statistics (Wang et al. 2025; Mao et al. 2025; Li et al. 2026c).

Regularization-based approaches. A second line modifies reference regularization. DiverseGRPO (Liu et al. 2026a) applies stronger KL regularization during early denoising steps and relaxes it later. GARDO (He et al. 2025) gates penalization to high-uncertainty samples, periodically updates the EMA reference, and amplifies rewards for high-quality diverse samples.

Interface-level approaches. A third line modifies the semantic interface through which prompts or rewards enter optimization. D2-Align (Chen et al. 2026) learns a directional correction in the text-embedding space of a frozen reward model. E2PO (Hu et al. 2026) perturbs content-token embeddings on the generator side and anneals the perturbation during denoising.

ReNFT is complementary to all three lines: its repair signal comes from components already inside the post-trained generator (a frozen encoder serves only to prioritize prompts). A natural question is whether these methods can also repair an already-collapsed adapter: reward-shaping and interface-level methods are designed for prevention during training, while KL-style anchoring can restore suppressed modes but ties the attainable reward to the base model. ReNFT realizes this repair through two internal routes of the same generator: a frozen base route that probes suppressed alternatives and an unconditional route that exposes prompt-independent bias. Trajectory-routing work (Soboleva et al. 2025; Cao et al. 2026; Yin et al. 2026; Jin, Shi, and Gu 2026) likewise shows that route identity and timestep matter, but does not construct reward-ranked, shared-noise counterfactuals.

3 Preliminaries

Refer to caption
Figure 3: ReNFT pipeline. (a) Unconditional probes prioritize anti-hub prompts; (b) two policy-dominated mixed routes generate matched proposals from the same prompt and initial noise; (c) reward ranking and the adaptive flipping guard assign pull/push targets for joint-and-paired NFT repair. The VAE decoder is omitted for clarity.

We introduce the flow-matching notation and the NFT-style forward-process post-training update that form the basis of our method.

3.1 Flow-Matching Generation

Let yy denote a text prompt, x0x_{0} a clean latent, and ε∼𝒩⁡(0,I)\varepsilon\sim\mathcal{N}(0,I) the initial noise. Under rectified-flow interpolation,

xt=(1−t)​x0+t​ε,t∈[0,1],x_{t}=(1-t)x_{0}+t\varepsilon,\quad t\in[0,1], (1)

the target transport velocity is v=ε−x0v=\varepsilon-x_{0}, and the model predicts vθ​(xt,t,y)v_{\theta}(x_{t},t,y) trained with the standard flow-matching regression objective. We denote three routes through this sampler: bb (frozen base), θ\theta (conditional policy), and uu (unconditional route of the current policy); an EMA-smoothed copy of the policy provides a stable reference, with its velocity written voldv_{\mathrm{old}}.

3.2 NFT-Style Forward-Process Post-Training

Reward-based post-training reshapes the distribution of a pretrained generator using a reward model R⁡(x0,y)R(x_{0},y). DiffusionNFT (Zheng et al. 2026a) avoids likelihood-ratio policy gradients by applying reward supervision through forward-process regression. Each generated endpoint x0x_{0} (from a rollout of NiN_{i} ODE steps) is re-noised at NtN_{t} training timesteps to xt=(1−σt)​x0+σt​ξx_{t}=(1-\sigma_{t})x_{0}+\sigma_{t}\xi, with fresh ξ∼𝒩⁡(0,I)\xi\sim\mathcal{N}(0,I) and noise level σt=t\sigma_{t}{=}t, and the model regresses the target velocity v=ξ−x0v=\xi-x_{0}. High-reward endpoints are reconstructed through the ordinary prediction

vθ+​(xt,t,y)=vθ​(xt,t,y),v_{\theta}^{+}(x_{t},t,y)=v_{\theta}(x_{t},t,y), (2)

while low-reward endpoints are pushed away through a mirrored prediction around the EMA reference. With the official default β=1\beta=1 in (1+β)​vold−β​vθ(1+\beta)v_{\mathrm{old}}-\beta v_{\theta}, this reduces to

vθ−​(xt,t,y)=2​vold​(xt,t,y)−vθ​(xt,t,y).v_{\theta}^{-}(x_{t},t,y)=2\,v_{\mathrm{old}}(x_{t},t,y)-v_{\theta}(x_{t},t,y). (3)

The combined objective over a rollout group is

ℒNFT=‖vθ+​(xt,t,y)−v‖22⏟pull+‖vθ−​(xt,t,y)−v‖22⏟push.\mathcal{L}_{\mathrm{NFT}}=\underbrace{\left\|v_{\theta}^{+}(x_{t},t,y)-v\right\|_{2}^{2}}_{\text{pull}}+\underbrace{\left\|v_{\theta}^{-}(x_{t},t,y)-v\right\|_{2}^{2}}_{\text{push}}. (4)

This MSE formulation trains quickly without likelihood ratios or SDE rollouts, and the EMA reference in vθ−v_{\theta}^{-} anchors the update and stabilizes post-training (ablated in the appendix). Equivalently, the loss can be written in endpoint space as ℓ⁡(x^0​(v),x0)\ell(\hat{x}_{0}(v),x_{0}) with x^0​(v)=xt−σt​v\hat{x}_{0}(v)=x_{t}-\sigma_{t}v; our implementation uses this form with a stop-gradient factor for loss-order adjustment. In our repair setting, the pull and push branches will operate on different endpoints rather than the same x0x_{0}; only at the pure-noise step σt=1\sigma_{t}=1 do the two trajectories share an identical state.

Online post-training supervises the model with its own sampled endpoints; when reward concentrates in a narrow mode, this acts as self-distillation that sharpens it. ReNFT retains the NFT parameterization but replaces this one-sided supervision with a matched local comparison (Section 4).

4 Methodology

We address the repair of a high-reward, low-diversity adapter without adding an external diversity objective or modifying text representations of the generator. ReNFT instead constructs candidate endpoints from available routes of the generator and ranks them with the same reward used for post-training. Section 4.1 first explains why an already-collapsed adapter can be repaired through these internal routes. Figure 3 summarizes the ensuing pipeline: unconditional probes prioritize anti-hub prompts where the prompt-independent bias is easiest to expose (Section 4.2); two policy-dominated mixed routes generate matched counterfactuals from the same prompt and initial noise (Section 4.3); and reward ranking assigns pull/push targets, optimized by native-noise joint and fresh-noise paired NFT updates (Section 4.4).

4.1 Repair View: Post-Training Reweights an Inherited Distribution

We study a repair setting in which an adapter θH\theta_{H} already achieves high reward but maps different initial noises under the same prompt to a narrow set of outputs. The frozen base generator and θH\theta_{H} share the same pretrained support; reward adaptation changes which parts of that support are easy to reach, but suppressed modes may remain accessible through routes closer to the base model. Collapse is therefore probability-mass compression rather than deletion, and repair should be possible from within.

A coarse mode-level view makes the feedback loop explicit. If mm indexes a visual mode, reward-only adaptation can be heuristically written as

pk+1​(m∣y)∝pk​(m∣y)​exp⁡(η​R¯k​(m,y)),p_{k+1}(m\mid y)\propto p_{k}(m\mid y)\exp\!\left(\eta\,\bar{R}_{k}(m,y)\right), (5)

where R¯k​(m,y)\bar{R}_{k}(m,y) is the average reward of sampled outputs in mode mm and η>0\eta>0 scales the reward-induced reweighting. Equation (5) is an interpretation of this feedback, not the exact update law. Figure 1 illustrates the view: per-prompt modes contract into narrow spikes as the loop reinforces the favored mode, then reopen during repair. Conditional modes that had lost probability regain output mass, and the unconditional distribution range expands accordingly.

This feedback also explains why re-optimizing reward cannot repair collapse: among semantically equivalent samples, reward differences encode preference rather than quality, so the reweighting only sharpens that preference. Repair instead needs reward gaps that reflect genuine quality differences. ReNFT achieves this by constructing matched counterfactuals from different internal routes (Sections 4.2–4.4), making the reward gap between paired endpoints informative rather than a restatement of the model bias. We validate this reversibility in Section 5.

4.2 Reading Post-Training Bias from the Unconditional Route

The unconditional route is the empty-prompt velocity uθ​(x,t)=vθ​(x,t,∅)u_{\theta}(x,t)=v_{\theta}(x,t,\varnothing). Although it shares θ\theta’s updated parameters, it has no separate reward target; its changes therefore expose prompt-independent side effects of conditional post-training: recurring structures, styles, or textures the adapter injects even when not requested, rather than directed optimization. Unconditional samples from a fixed noise grid contract from a diverse base into a narrow hub after post-training and progressively reopen during repair (see the appendix for unconditional samples across training steps). We use these tendencies as an internal bias probe, whose diagnostic validity is evaluated rather than assumed.

We first use this readout to prioritize prompts on which the bias is easiest to expose, as shown in Figure 3a. Offline, the prompt bank is encoded once by the SigLIP2 text encoder (Tschannen et al. 2025) and cached. Online, MM empty-prompt probes {uj}j=1M\{u_{j}\}_{j=1}^{M} are generated from the current policy and encoded by the SigLIP2 image encoder. Each candidate prompt yy is scored by

s⁡(y)=1M​∑j=1M⟨etext​(y),eimg​(uj)⟩,s(y)=\frac{1}{M}\sum_{j=1}^{M}\langle e_{\mathrm{text}}(y),e_{\mathrm{img}}(u_{j})\rangle, (6)

and the bottom-KK prompts (lowest mean similarity) are selected as anti-hub prompts, those farthest from the current unconditional tendency, where the exposed bias is most likely to yield distinguishable proposals. The frozen encoder serves only prompt prioritization; it never enters the reward, advantage, or repair loss. The same route also serves constructively in the mixed sampling of Section 4.3, where it exposes the bias inside conditional trajectories for reward ranking.

4.3 Constructing Internal Counterfactual Proposals

The frozen base route and the unconditional route play complementary roles. A base step introduces a base-like vector-field direction and can propose alternatives that are sampled less often by the post-trained policy; an unconditional step exposes the prompt-independent bias identified in Section 4.2. The conditional route θ\theta occupies most steps by design, limiting deviation from the high-reward policy. Crucially, neither route is hard-coded as a pull or push branch: mixed routing generates nearby alternatives, and the reward decides their roles (Figure 3b).

Refer to caption
Figure 4: Qualitative comparison on GenEval and OCR prompts. Each prompt group shows four samples from the same method under different initial noises.
Refer to caption
Figure 5: Qualitative comparison on PickScore prompts. Each prompt group shows four samples from the same method under different initial noises.

For a routing pattern PP, let TP,yT_{P,y} map initial noise to a terminal sample. The final configuration repeats two three-step blocks across the rollout,

PA=(bθθ)⋯,PB=(θθu)⋯.P_{A}=(b\theta\theta)\cdots,\qquad P_{B}=(\theta\theta u)\cdots. (7)

Both routes diverge from the same initial noise ε\varepsilon, keeping two thirds of each route conditional and avoiding long b​bbb or u​uuu blocks that stray too far from the policy (cf. stage-wise analyses (Jin, Shi, and Gu 2026)).

For a selected prompt yy and shared initial noise ε\varepsilon, the matched endpoints are x0A=TPA,y​(ε)x_{0}^{A}=T_{P_{A},y}(\varepsilon) and x0B=TPB,y​(ε)x_{0}^{B}=T_{P_{B},y}(\varepsilon). Sharing (y,ε)(y,\varepsilon) makes routing the main changed factor: two pure-θ\theta trajectories would coincide, while a direct base-versus-policy comparison is too one-sided; the two mixed routes instead perturb the policy in different directions, and either may win for a particular (y,ε)(y,\varepsilon). Trajectory-level branching has likewise been exploited for alignment and credit assignment in TMPO and DenseGRPO (Li et al. 2026a; Deng et al. 2026); we instead branch two route mixtures from identical noise to construct a reward-testable counterfactual pair.

We rank the endpoints with the original reward (Figure 3b, right) and obtain (x0H,x0L)(x_{0}^{H},x_{0}^{L}); when rewards differ, the higher-reward endpoint is always the pull target and the lower the push target, regardless of route. The matched reward gap

Δ​R​(y,ε)=R⁡(x0A,y)−R⁡(x0B,y)\Delta R(y,\varepsilon)=R(x_{0}^{A},y)-R(x_{0}^{B},y) (8)

provides a local preference without claiming global superiority of either route. For continuous rewards (PickScore), ties are rare; an adaptive flipping guard (Figure 3c) requires at least half of the pull targets to come from route AA, flipping the lowest-margin route-BB win when the ratio falls below 0.50.5. This guard keeps the coverage-restoring b​θ​θb\theta\theta branch dominant and prevents the hacking-prone θ​θ​u\theta\theta u branch from re-hacking the pull distribution. For rule-based discrete rewards (GenEval), ties are frequent; we break them with a near 50/50 split that slightly favors route AA, without activating the flipping guard.

4.4 Reward-Ranked Joint-and-Paired Repair

We realize the matched preference with the NFT mirror mechanism from Section 3.2. The loss in v-space is

ℒrepair=‖vθ−​(xtL,t,y)−vL‖22⏟mirrored push+‖vθ+​(xtH,t,y)−vH‖22⏟ordinary pull.\mathcal{L}_{\mathrm{repair}}=\underbrace{\left\|v_{\theta}^{-}(x_{t}^{L},t,y)-v^{L}\right\|_{2}^{2}}_{\text{mirrored push}}+\underbrace{\left\|v_{\theta}^{+}(x_{t}^{H},t,y)-v^{H}\right\|_{2}^{2}}_{\text{ordinary pull}}. (9)

The mirror reference vθ−v_{\theta}^{-} uses an EMA-smoothed copy of the policy with constant decay, since repair starts from a late-stage checkpoint.

For each paired timestep, we sample a fresh perturbation ξ\xi and construct

xtH=(1−σt)​x0H+σt​ξ,xtL=(1−σt)​x0L+σt​ξ.x_{t}^{H}=(1-\sigma_{t})x_{0}^{H}+\sigma_{t}\xi,\quad x_{t}^{L}=(1-\sigma_{t})x_{0}^{L}+\sigma_{t}\xi. (10)

with targets vH=ξ−x0Hv^{H}=\xi-x_{0}^{H} and vL=ξ−x0Lv^{L}=\xi-x_{0}^{L}. The pull and push share ξ\xi, treating corruption as a matched nuisance; sampling ξ\xi independently of ε\varepsilon avoids a per-(prompt, noise) fitting shortcut (Wen et al. 2024).

At σt=1\sigma_{t}=1, a fresh ξ\xi would break the coupling, so we use the native rollout noise ε\varepsilon for a joint update (Figure 3c), with xtH=xtL=εx_{t}^{H}=x_{t}^{L}=\varepsilon and targets vH=ε−x0Hv^{H}=\varepsilon-x_{0}^{H}, vL=ε−x0Lv^{L}=\varepsilon-x_{0}^{L}. This anchor matters because early denoising decisions strongly influence global structure (Jin, Shi, and Gu 2026). Each repair step applies one joint update with ε\varepsilon at σt=1\sigma_{t}=1, followed by an even number of paired updates with fresh ξ\xi at σt<1\sigma_{t}<1 (see the appendix for ablations on the joint-noise choice and rollout length). Mixed routes only construct training supervision; inference uses the standard conditional forward pass with no routing overhead.

5 Experiments

We evaluate ReNFT on PickScore and GenEval: standard reward post-training produces the collapsed checkpoint that motivates repair (Section 5.2), ReNFT recovers diversity while retaining reward both quantitatively and qualitatively (Sections 5.2 and 5.3), and controlled ablations identify the responsible design decisions (Section 5.4).

Method Model-Based Reward Diversity
PickScore ↑\uparrow Aesthetic ↑\uparrow ImageReward ↑\uparrow HPSv2 ↑\uparrow LPIPS-Div ↑\uparrow DreamSim-Div ↑\uparrow DINO-Div ↑\uparrow
SD3.5-M (21.73) (6.026) (1.06) (0.295) (0.599) (0.242) (0.251)
+ Flow-GRPO 22.90 6.277 1.31 0.305 0.463 0.122 0.134
+ DiffusionNFT 23.51 6.592 1.44 0.327 0.430 0.119 0.112
+ DiverseGRPO 22.99 6.234 1.27 0.332 0.506 0.163 0.145
+ E2PO† 23.38 6.538 1.29 0.325 – – –
+ Ours 23.26 6.344 1.41 0.323 0.565 0.189 0.182
Table 1: Quantitative results on the PickScore test set. Shaded: in-domain training reward. Parenthesized row: reference only, excluded from ranking. Bold: best; underline: second best. †See Section 5.1.
Method Rule-Based Reward Diversity
GenEval ↑\uparrow LPIPS-Div ↑\uparrow DreamSim-Div ↑\uparrow DINO-Div ↑\uparrow
SD3.5-M (0.654) (0.687) (0.339) (0.379)
+ Flow-GRPO 0.802 0.343 0.176 0.205
+ DiffusionNFT 0.946 0.292 0.149 0.186
+ DiverseGRPO 0.869 0.454 0.198 0.246
+ E2PO† 0.932 – – –
+ Ours 0.937 0.496 0.231 0.295
Table 2: Quantitative results on the GenEval test set. Shaded: in-domain training reward. Parenthesized row: reference only, excluded from ranking. Bold: best; underline: second best. †See Section 5.1.

5.1 Experimental Setup

We use SD3.5-M (Esser et al. 2024) with LoRA adapters and train separate PickScore (Kirstain et al. 2023) and GenEval (Ghosh, Hajishirzi, and Schmidt 2023) checkpoints. ReNFT reloads the DiffusionNFT LoRA at step 700 (PickScore) or 250 (GenEval) and repairs for 50 steps with the default configuration in the appendix. Model-based rewards use all 1,024 PickScore test prompts; the rule-based reward uses all 553 GenEval prompts; diversity uses 100 fixed prompts with 12 fixed-noise samples each. NFT, E2PO, and ReNFT use CFG=1; base and GRPO-style checkpoints use CFG=4.5, so cross-CFG comparisons are end-to-end method comparisons, not controlled CFG ablations.

We compare the untrained SD3.5-M reference, Flow-GRPO (Liu et al. 2025), DiffusionNFT (NFT) (Zheng et al. 2026a), DiverseGRPO (Liu et al. 2026a), and E2PO (Hu et al. 2026). E2PO is not open-sourced; we match its official training budget and reuse its paper-reported reward values (marked †), omitting it from the diversity ranking. Quality metrics are PickScore-v1, Aesthetic (Radford et al. 2021; Schuhmann et al. 2022), ImageReward-v1.0 (Xu et al. 2023), and HPSv2.1 (Wu et al. 2023). Diversity metrics are LPIPS-Div (Zhang et al. 2018), DreamSim-Div (Fu et al. 2023), and DINOv3-Div (Siméoni et al. 2025), each computed as the mean pairwise distance among 12 samples per prompt. Anti-hub uses 24 probes and 1,000 candidates; further details are in the appendix.

5.2 Quantitative Results

Reward post-training produces the repair setting. Tables 1 and 2 confirm the reward–diversity tension hypothesized in Section 4.1: the untrained base (shown in parentheses as a reference row, excluded from ranking) has the largest within-prompt diversity but the weakest task reward, while the continued DiffusionNFT run reaches the highest PickScore and GenEval values at steps 750 and 300, with DreamSim-Div falling from 0.242 to 0.119 and from 0.339 to 0.149, respectively. ReNFT instead branches from the earlier NFT checkpoints HH at steps 700 and 250 under the same remaining 50-step budget. Among the compared post-trained methods, ReNFT attains the highest diversity in all three metrics under both protocols.

ReNFT recovers diversity while retaining reward. Relative to NFT at the protocol endpoint, ReNFT retains 98.9% of PickScore (23.26 vs. 23.51) and 99.0% of GenEval (0.937 vs. 0.946), while improving DreamSim-Div by 58.8% (0.189 vs. 0.119) and 55.0% (0.231 vs. 0.149), respectively. The trade-off is asymmetric: roughly 1% reward cost versus 55–59% diversity gain. ReNFT ranks second on GenEval reward and PickScore ImageReward; all three diversity metrics are best among compared post-trained methods under both protocols. This asymmetry reflects how susceptible each reward is to hacking: model-based scores of PickScore depend heavily on exploiting the model preference, so an anti-hacking method faces a harder reward landscape, whereas GenEval’s rule-based scoring rewards correct color, position, and count through segmentation, leaving less room for hacking-specific artifacts.

Repair dynamics. In Figure 2(a)–(d), under PickScore reward decreases with local oscillations toward the table value, while diversity rises rapidly, peaks around step 35, and decays mildly, remaining above every other post-trained baseline at step 50. The rise–peak–decay pattern reflects the dual-route win ratio: b​θ​θb\theta\theta initially dominates θ​θ​u\theta\theta u, but as the routes converge the ratio oscillates near 0.50.5; adaptive flipping maintains a minimum b​θ​θb\theta\theta pull share, yet the unconditional branch gradually re-hacks, causing diversity to recede and reward to climb again along a hacking trajectory. GenEval shows no turning point within 50 steps: its wider acceptable reward range produces more stable curves: reward drops less, diversity rises more, and the dual-branch balance point is reached later.

Comparison with diversity-oriented baselines. DiverseGRPO preserves more diversity than NFT but at a larger reward cost: 97.8% of PickScore (22.99 vs. 23.51) and 91.9% of GenEval (0.869 vs. 0.946), compared to ReNFT’s 98.9% and 99.0%. E2PO’s reused reward enters the ranking with provenance marked †; its interface-level bias correction is complementary to our internal-route approach.

5.3 Qualitative Results

Figures 4 and 5 compare Base, NFT, DiverseGRPO, and ReNFT on GenEval+OCR and PickScore prompts, respectively, with a shared noise set.

PickScore prompts. NFT and DiverseGRPO exhibit recognizable hacking tendencies: a recurring color tone, densely packed elements, and a preference for single female subjects. NFT’s late-stage outputs respond primarily to prompt changes, becoming insensitive to the initial noise. ReNFT restores scene- and style-level variation across seeds without reverting to the base: dog statues vary in pose and setting, wedding scenes recover diverse compositions, and the emperor-and-robots prompt produces distinct arrangements.

GenEval and OCR prompts. NFT collapses counting-related prompts onto pure white backgrounds, a hacking shortcut that simplifies segmentation and counting. ReNFT removes this bias: count, position, and color remain correct while backgrounds and styles vary. The OCR prompts are out-of-distribution for our PickScore-trained repair: NFT and DiverseGRPO retain some text-rendering ability from OCR-adjacent prompts, but hacking distorts style and layout. ReNFT produces more legible and varied text, approaching base-level diversity with higher text fidelity. These observations confirm the repair-view hypothesis (Section 4.1): suppressed modes were compressed, not deleted. A diagnostic view from unconditional (CFG=0) samples is provided in the appendix.

5.4 Ablation Study

The ablations examine two design decisions in the main text: anti-hub prompt prioritization and internal route construction. All variants branch from the same checkpoint HH under the same 50-step budget and checkpoint-evaluation protocol.

Anti-hub prompt prioritization. The controlled variant replaces anti-hub prompts with random prompts from the same candidate pool, preserving the prompt count, sampling budget, and update rule. The two variants track nearly identical reward trajectories, whereas diversity diverges sharply: without anti-hub selection, diversity rises only briefly and then decays back below its starting level, while the full configuration sustains the recovery. The PickScore training distribution is itself biased toward the hacking preferences of the reward model, so θ​θ​u\theta\theta u may not be substantially worse than b​θ​θb\theta\theta from the start; without anti-hub prompts that expose the bias most distinguishably, there is little diversity-recovery momentum.

Internal route construction. Figure 2(e,f) compares (b​θ,θ​u)(b\theta,\theta u), (b​θ​θ,θ​θ​u)(b\theta\theta,\theta\theta u), and longer periodic variants under the same budget. With pattern length 2, the base and unconditional content is too high: θ​u\theta u consistently produces low-reward outputs, the win ratio stays near 1.01.0, and saturation/contrast anomalies cause reward to collapse while diversity rises from degenerate outputs. Longer blocks approach the NFT direction. The three-step pattern (b​θ​θ,θ​θ​u)(b\theta\theta,\theta\theta u) achieves the clearest diversity recovery with the least reward sacrifice.

Appendix ablations support the default rollout length, joint-noise choice, and EMA decay; sensitivity to the anti-hub pool size and flipping threshold remains future work.

6 Conclusion

We studied reward-induced mode collapse as internal probability-mass reallocation, where post-training concentrates mass on reward-favored modes and suppressed alternatives remain accessible through internal routes of the generator. ReNFT operationalizes this view by prioritizing anti-hub prompts with unconditional probes, generating matched counterfactuals from the same prompt and noise through two policy-dominated mixed routes, and assigning pull and push roles through reward ranking with an adaptive flipping guard, all realized through joint-and-paired NFT updates. Across PickScore and GenEval, ReNFT retains 98.9% and 99.0% of NFT’s reward while improving DreamSim-Div by 58.8% and 55.0%, confirming that collapsed adapters can be repaired from within without external diversity objectives, text-encoder modification, or sacrificing the acquired reward. This internal repair perspective opens a complementary direction to external interventions, extending naturally to other post-training paradigms and backbones.

References

  • Black Forest Labs (2025) Black Forest Labs. 2025. FLUX.2: Frontier Visual Intelligence. https://bfl.ai/blog/flux-2.
  • Cao et al. (2026) Cao, Q.; Chen, Y.; Ma, C.; and Yang, X. 2026. Dynamic Training-Free Fusion of Subject and Style LoRAs. https://qinglongcao.xyz/TVML-Diffusion.github.io/. Manuscript.
  • Chen et al. (2026) Chen, C.; Hu, S.; Zhu, J.; Wu, M.; Chen, J.; Li, Y.; Huang, N.; Fang, C.; Wu, J.; Chu, X.; and Li, X. 2026. Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  • Deng et al. (2026) Deng, H.; Yan, K.; Mao, C.; Wang, X.; Liu, Y.; Gao, C.; and Sang, N. 2026. DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment. In International Conference on Learning Representations (ICLR).
  • Esser et al. (2024) Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. 2024. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. In International Conference on Machine Learning (ICML), 12675–12702.
  • Fang et al. (2026) Fang, Z.; Huang, W.; Zeng, Y.; Zhao, Y.; Chen, S.; Feng, K.; Lin, Y.; Chen, L.; Chen, Z.; Cao, S.; and Zhao, F. 2026. Flow-OPD: On-Policy Distillation for Flow Matching Models. ArXiv preprint arXiv:2605.08063.
  • Fu et al. (2023) Fu, S.; Tamir, N.; Sundaram, S.; Chai, L.; Zhang, R.; Dekel, T.; and Isola, P. 2023. DreamSim: Learning New Dimensions of Human Visual Similarity Using Synthetic Data. In Advances in Neural Information Processing Systems (NeurIPS), 50742–50768.
  • Ghosh, Hajishirzi, and Schmidt (2023) Ghosh, D.; Hajishirzi, H.; and Schmidt, L. 2023. GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment. In Advances in Neural Information Processing Systems (NeurIPS).
  • Go et al. (2026) Go, H.; Chung, H.; Truong, P.; Bhat, G.; Mi, L.; An, Z.; Zhao, Z.; Narnhofer, D.; Belongie, S.; Tombari, F.; and Schindler, K. 2026. Stitched Value Model for Diffusion Alignment. ArXiv preprint arXiv:2605.19804.
  • He et al. (2025) He, H.; Ye, Y.; Liu, J.; Liang, J.; Wang, Z.; Yuan, Z.; Wang, X.; Mao, H.; Wan, P.; and Pan, L. 2025. GARDO: Reinforcing Diffusion Models without Reward Hacking. ArXiv preprint arXiv:2512.24138.
  • Hu et al. (2026) Hu, S.; Chen, C.; Zhu, J.; Wu, J.; Chu, X.; and Li, X. 2026. E2PO: Embedding-perturbed Exploration Preference Optimization for Flow Models. In International Conference on Machine Learning (ICML).
  • Jin, Shi, and Gu (2026) Jin, C.; Shi, Q.; and Gu, Y. 2026. Stage-wise Dynamics of Classifier-Free Guidance in Diffusion Models. In International Conference on Learning Representations (ICLR).
  • Kirstain et al. (2023) Kirstain, Y.; Polyak, A.; Singer, U.; Matiana, S.; Penna, J.; and Levy, O. 2023. Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation. In Advances in Neural Information Processing Systems (NeurIPS), 36652–36663.
  • Li et al. (2026a) Li, J.; Zhu, C.; Yi, N.; Bao, Y.; Sun, L.; Lv, Q.; Fang, X.; Liu, D.; Li, J.; He, K.; Zhou, B.; and Ma, Z. 2026a. TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment. ArXiv preprint arXiv:2605.10983.
  • Li et al. (2026b) Li, Q.; Yu, J.; Jiang, K.; Wei, Y.; Xing, Z.; Li, P.; Chu, R.; Zhang, S.; Liu, Y.; and Wu, Z. 2026b. DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models. ArXiv preprint arXiv:2605.15055.
  • Li et al. (2026c) Li, R.; Xu, M.; Gu, S.; Qu, L.; Feng, F.; Hu, H.; and Wang, W. 2026c. Optimizing Visual Generative Models via Distribution-wise Rewards. ArXiv preprint arXiv:2607.02291.
  • Li et al. (2025) Li, Z.; Liu, Z.; Zhang, Q.; Lin, B.; Wu, F.; Yuan, S.; Yan, Z.; Ye, Y.; Yu, W.; Niu, Y.; Wang, S.; Cheng, X.; and Yuan, L. 2025. Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback. ArXiv preprint arXiv:2510.16888.
  • Liu et al. (2026a) Liu, H.; Huang, H.; Wang, J.; Liu, C.; Li, X.; and Ji, X. 2026a. DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  • Liu et al. (2026b) Liu, J.; Li, H.; Sun, Z.; Chen, C.; Bian, Y.; Wang, B.; Dong, D.; Chen, C.; and Wang, Z. 2026b. Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation. ArXiv preprint arXiv:2601.12401.
  • Liu et al. (2025) Liu, J.; Liu, G.; Liang, J.; Li, Y.; Liu, J.; Wang, X.; Wan, P.; Zhang, D.; and Ouyang, W. 2025. Flow-GRPO: Training Flow Matching Models via Online RL. In Advances in Neural Information Processing Systems (NeurIPS).
  • Liu, He, and Li (2026) Liu, Z.; He, X.; and Li, Y. 2026. Advances in GRPO for Generation Models: A Survey. ArXiv preprint arXiv:2603.06623.
  • Mao et al. (2025) Mao, W.; Chen, H.; Yang, Z.; and Shou, M. Z. 2025. The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation. ArXiv preprint arXiv:2511.20256.
  • Ping et al. (2026) Ping, B.; Zhou, X.; Qi, P.; Luo, M.; Bo, L.; and Pang, T. 2026. Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models. ArXiv preprint arXiv:2606.11025.
  • Qin et al. (2026) Qin, Y.; Wang, L.; Fei, H.; Zimmermann, R.; Bo, L.; Lu, Q.; and Wang, C. 2026. SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models. ArXiv preprint arXiv:2604.12617.
  • Radford et al. (2021) Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models From Natural Language Supervision. In International Conference on Machine Learning (ICML), 8748–8763.
  • Schuhmann et al. (2022) Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022. LAION-5B: An Open Large-scale Dataset for Training Next Generation Image-Text Models. In Advances in Neural Information Processing Systems (NeurIPS).
  • Siméoni et al. (2025) Siméoni, O.; Vo, H. V.; Seitzer, M.; Baldassarre, F.; Oquab, M.; Jose, C.; Khalidov, V.; Szafraniec, M.; Yi, S.; Ramamonjisoa, M.; Massa, F.; Haziza, D.; Wehrstedt, L.; Wang, J.; Darcet, T.; Moutakanni, T.; Sentana, L.; Roberts, C.; Vedaldi, A.; Tolan, J.; Brandt, J.; Couprie, C.; Mairal, J.; Jégou, H.; Labatut, P.; and Bojanowski, P. 2025. DINOv3. Transactions on Machine Learning Research (TMLR). ArXiv:2508.10104.
  • Soboleva et al. (2025) Soboleva, V.; Alanov, A.; Kuznetsov, A.; and Sobolev, K. 2025. T-LoRA: Single Image Diffusion Model Customization Without Overfitting. ArXiv preprint arXiv:2507.05964.
  • Tan et al. (2026) Tan, X.; Liu, J.; Gao, B.-B.; Fan, Y.; Jiang, X.; Wang, C.; Wang, H.; and Zheng, F. 2026. When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy. ArXiv preprint arXiv:2605.12112.
  • Tong et al. (2026) Tong, Y.; Liu, M.; Zhao, C.; He, W.; Zhang, S.; Zhang, H.; Zhang, P.; Liu, J.; Huang, J.; Wang, J.; Jiang, H.; and Huang, P. 2026. Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO. ArXiv preprint arXiv:2602.06422.
  • Tschannen et al. (2025) Tschannen, M.; Gritsenko, A.; Wang, X.; Naeem, M. F.; Alabdulmohsin, I.; Parthasarathy, N.; Evans, T.; Beyer, L.; Xia, Y.; Mustafa, B.; Hénaff, O.; Harmsen, J.; Steiner, A.; and Zhai, X. 2025. SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features. ArXiv preprint arXiv:2502.14786.
  • Wallace et al. (2024) Wallace, B.; Dang, M.; Rafailov, R.; Zhou, L.; Lou, A.; Purushwalkam, S.; Ermon, S.; Xiong, C.; Joty, S.; and Naik, N. 2024. Diffusion Model Alignment Using Direct Preference Optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8228–8238.
  • Wang et al. (2025) Wang, Y.; Li, Z.; Zang, Y.; Zhou, Y.; Bu, J.; Wang, C.; Lu, Q.; Jin, C.; and Wang, J. 2025. Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning. ArXiv preprint arXiv:2508.20751.
  • Wen et al. (2024) Wen, Y.; Liu, Y.; Chen, C.; and Lyu, L. 2024. Detecting, Explaining, and Mitigating Memorization in Diffusion Models. In International Conference on Learning Representations (ICLR).
  • Wu et al. (2023) Wu, X.; Hao, Y.; Sun, K.; Chen, Y.; Zhu, F.; Zhao, R.; and Li, H. 2023. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis. ArXiv preprint arXiv:2306.09341.
  • Xu et al. (2023) Xu, J.; Liu, X.; Wu, Y.; Tong, Y.; Li, Q.; Ding, M.; Tang, J.; and Dong, Y. 2023. ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation. In Advances in Neural Information Processing Systems (NeurIPS), 56730–56748.
  • Xue et al. (2026) Xue, S.; Ge, C.; Zhang, S.; Li, Y.; and Ma, Z.-M. 2026. Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models. In International Conference on Learning Representations (ICLR).
  • Yin et al. (2026) Yin, B.; Hu, X.; Zhou, X.; He, Y.; Jiang, P.-T.; Liao, Y.; Zhu, J.; Zhang, J.; Tai, Y.; and Yan, S. 2026. FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning. https://github.com/YinBo0927/FeRA. Manuscript.
  • Zhang et al. (2018) Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 586–595.
  • Zheng et al. (2026a) Zheng, K.; Chen, H.; Ye, H.; Wang, H.; Zhang, Q.; Jiang, K.; Su, H.; Ermon, S.; Zhu, J.; and Liu, M.-Y. 2026a. DiffusionNFT: Online Diffusion Reinforcement with Forward Process. In International Conference on Learning Representations (ICLR).
  • Zheng et al. (2026b) Zheng, K.; He, G.; Zhao, M.; Zhang, J.; Chen, H.; Chen, J.; Lin, C.-H.; Liu, M.-Y.; Zhu, J.; and Ma, Q. 2026b. Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models. ArXiv preprint arXiv:2606.25473.

Appendix A Experimental Details

Training configuration. SD3.5-M (Esser et al. 2024) runs train LoRA adapters (rank 32, alpha 64) with learning rate 10−410^{-4}, 48 unique samples per epoch, and group size 24. All runs use seed 42, the same seed as prior work (Hu et al. 2026). Repair uses off-policy training with a constant EMA decay of 0.250.25, Ni=36N_{i}{=}36 rollout steps, and Nt=7N_{t}{=}7 training timesteps per update (one native-noise joint anchor and six fresh-noise paired updates). The minimum route-AA pull ratio is 0.50.5.

Prompt splits and tuning protocol. PickScore (Kirstain et al. 2023) provides 15,471 training prompts and GenEval (Ghosh, Hajishirzi, and Schmidt 2023) provides 33,199; evaluation uses the full held-out test sets as in the main paper. Anti-hub candidate pools are drawn from the training bank, and all hyperparameter selection reported in Section B was based on evaluation trajectories of saved checkpoints on the training prompts (not training-metric logs), with the held-out prompts never used for tuning.

Training budget and checkpoint choice. The total budgets (750 steps for PickScore, 300 for GenEval) follow the official configuration of E2PO (Hu et al. 2026). All trained baselines share these budgets and are trained to the same endpoints. ReNFT branches from the NFT checkpoint 50 steps before the endpoint and uses the remaining 50 steps for repair, matching the continued-training baselines step for step. The 50-step budget is motivated by the pull-ratio analysis in Section C.1.

Evaluation settings. On SD3.5-M, evaluation uses 512×\times512 resolution with 40-step ODE sampling. NFT-family checkpoints (NFT, E2PO, and ReNFT) are evaluated at CFG==1 (i.e., CFG-free), following the protocol of DiffusionNFT (Zheng et al. 2026a).

Diversity metrics. For each held-out prompt, we generate 12 samples from fixed initial noises. LPIPS (Zhang et al. 2018) uses the AlexNet backbone with bicubic resizing to 256×\times256; DreamSim (Fu et al. 2023) uses the official pretrained model and cosine distance between its embeddings; DINOv3 (Siméoni et al. 2025) uses the ViT-B/16 LVD-1689M CLS embedding with Resize(256)–CenterCrop(224) preprocessing and cosine distance. For every metric, we average all (122)=66\binom{12}{2}=66 pairwise distances within a prompt, then average the resulting prompt-level scores over the 100 held-out prompts.

Quality metrics. PickScore (Kirstain et al. 2023), Aesthetic (Radford et al. 2021; Schuhmann et al. 2022), ImageReward (Xu et al. 2023), and HPSv2 (Wu et al. 2023) are computed as in the main paper.

Algorithm summary. Algorithm 1 summarizes the repair loop; all symbols follow Sections 4.2–4.4 of the main paper, and concrete hyperparameter values are those given above.

Algorithm 1 Repair NFT (ReNFT).

Require: Hacked checkpoint θH\theta_{H}, frozen base route bb, reward model R⁡(⋅)R(\cdot), frozen SigLIP2 encoders etext,eimge_{\mathrm{text}},e_{\mathrm{img}}, training prompt bank 𝒴\mathcal{Y}, routing patterns PA,PBP_{A},P_{B} with rollout operator 𝒯\mathcal{T}, probe count MM, candidate number CC, training timesteps NtN_{t}, learning rate λ\lambda, EMA decay η\eta, minimum route-AA pull ratio τ\tau.
Initialize: θ←θH\theta\leftarrow\theta_{H}; θold←θH\theta^{\mathrm{old}}\leftarrow\theta_{H}; batch buffer ℬ←∅\mathcal{B}\leftarrow\varnothing.

1:  for each optimizer step do
2:   Sample MM unconditional probes {uj}j=1M\{u_{j}\}_{j=1}^{M} // anti-hub selection
3:   Score CC candidate prompts by s⁡(y)←1M​∑j=1M⟨etext​(y),eimg​(uj)⟩s(y)\leftarrow\frac{1}{M}\sum_{j=1}^{M}\langle e_{\mathrm{text}}(y),e_{\mathrm{img}}(u_{j})\rangle
4:   Select anti-hub set: 𝒴ah←BottomC,y∈𝒴s​(y)\mathcal{Y}_{\mathrm{ah}}\leftarrow\operatorname*{Bottom}_{C,\,y\in\mathcal{Y}}s(y) Eq. (6)
5:   for y∈𝒴ahy\in\mathcal{Y}_{\mathrm{ah}} do
6:    Draw ε∼𝒩⁡(0,I)\varepsilon\sim\mathcal{N}(0,I) // matched rollouts
7:    Compute matched endpoints x0A←𝒯PA,y​(ε)x_{0}^{A}\leftarrow\mathcal{T}_{P_{A},y}(\varepsilon), x0B←𝒯PB,y​(ε)x_{0}^{B}\leftarrow\mathcal{T}_{P_{B},y}(\varepsilon)
8:    Rank by reward: Hy←arg⁡maxk∈{A,B}⁡R⁡(x0k,y)H_{y}\leftarrow\arg\max_{k\in\{A,B\}}R(x_{0}^{k},y); Ly←arg⁡mink∈{A,B}⁡R⁡(x0k,y)L_{y}\leftarrow\arg\min_{k\in\{A,B\}}R(x_{0}^{k},y) Eq. (8)
9:    Collect (y,ε,x0Hy,x0Ly)(y,\varepsilon,x_{0}^{H_{y}},x_{0}^{L_{y}}) into ℬ\mathcal{B}
10:   end for
11:   ℬ←GuardFlip⁡(ℬ,τ)\mathcal{B}\leftarrow\operatorname{GuardFlip}(\mathcal{B},\tau) // minimum route-AA pull share
12:   for (y,ε,x0H,x0L)∈ℬ(y,\varepsilon,x_{0}^{H},x_{0}^{L})\in\mathcal{B} do
13:    for t∈{1,…,Nt}t\in\{1,\ldots,N_{t}\} do
14:     ξ←ε\xi\leftarrow\varepsilon if t=1t{=}1, else ξ∼𝒩⁡(0,I)\xi\sim\mathcal{N}(0,I) // self / fresh noise
15:     Re-noise high endpoint: xtH←(1−σt)​x0H+σt​ξx_{t}^{H}\leftarrow(1{-}\sigma_{t})x_{0}^{H}{+}\sigma_{t}\xi
16:     Re-noise low endpoint: xtL←(1−σt)​x0L+σt​ξx_{t}^{L}\leftarrow(1{-}\sigma_{t})x_{0}^{L}{+}\sigma_{t}\xi
17:     Targets: vH←ξ−x0Hv^{H}\leftarrow\xi{-}x_{0}^{H}; vL←ξ−x0Lv^{L}\leftarrow\xi{-}x_{0}^{L} Eq. (10)
18:     Implicit positive velocity: vθ+←vθ​(xtH,t,y)v_{\theta}^{+}\leftarrow v_{\theta}(x_{t}^{H},t,y) Eq. (2)
19:     Implicit negative velocity: vθ−←2​vold​(xtL,t,y)−vθ​(xtL,t,y)v_{\theta}^{-}\leftarrow 2\,v_{\mathrm{old}}(x_{t}^{L},t,y)-v_{\theta}(x_{t}^{L},t,y) Eq. (3)
20:     Positive loss: ℒ+←‖vθ+−vH‖22\mathcal{L}^{+}\leftarrow\|v_{\theta}^{+}-v^{H}\|_{2}^{2}
21:     Negative loss: ℒ−←‖vθ−−vL‖22\mathcal{L}^{-}\leftarrow\|v_{\theta}^{-}-v^{L}\|_{2}^{2}
22:     θ←θ−λ​∇θ(ℒ++ℒ−)\theta\leftarrow\theta-\lambda\nabla_{\theta}(\mathcal{L}^{+}{+}\mathcal{L}^{-}) Eq. (9)
23:    end for
24:   end for
25:   θold←η​θold+(1−η)​θ\theta^{\mathrm{old}}\leftarrow\eta\,\theta^{\mathrm{old}}+(1{-}\eta)\,\theta // EMA update
26:  end for

Output: vθv_{\theta}

Appendix B Additional Ablations

All ablations in this section use the SD3.5-M backbone under the PickScore protocol, branching from the same hacked checkpoint HH as the main experiments.

Figure 6: Ablation of Anti-Hub Candidate Pool Size on SD3.5-M (PickScore). The green trajectory (1,000 candidates) is identical to the default configuration in all other ablation figures.
Figure 7: Ablation of EMA Decay on SD3.5-M (PickScore). Larger decays preserve reward but barely recover diversity; smaller decays recover more diversity at larger reward cost, with on-policy training (decay 00) reaching the highest diversity but the lowest reward. The 0.250.25 curve is Ours.
Figure 8: Ablation of Rollout Length and Joint-Step Self-Noise on SD3.5-M (PickScore). The 36-step curve is Ours; all variants branch from the same checkpoint HH.

B.1 Ablation of Anti-Hub Selection and Pool Size

Anti-hub selection uses 24 unconditional probe generations to score each of 1,000 candidate prompts with SigLIP2 (Tschannen et al. 2025) text–image embeddings, selecting the lowest-mean-similarity prompts for repair. The on/off ablation in the main paper confirms that replacing anti-hub-selected prompts with random ones collapses diversity recovery while reward tracks closely, isolating prompt prioritization from prompt coverage.

Candidate pool size.

Figure 6 varies the candidate pool size {500,1000,2000,3000}\{500,1000,2000,3000\} under the same 50-step repair budget. Pools of 1,000 or more candidates yield nearly identical reward and diversity trajectories, indicating that coverage saturates quickly; the smallest pool (500) retains reward best but recovers noticeably less diversity due to insufficient prompt coverage. The default 1,000 candidates are sufficient.

B.2 Ablation of EMA Decay

We compare constant decays {0,0.1,0.25,0.5}\{0,0.1,0.25,0.5\} under the same hacked checkpoint, 50-step repair budget, and checkpoint-evaluation protocol, where 00 corresponds to on-policy training without the off-policy reference. Large decay (0.50.5) slows the reference so much that diversity barely improves while reward instead rises slightly, reproducing an NFT-like hacking direction. Smaller decays move along the trade-off in the opposite direction: decay 0.10.1 recovers substantially more diversity at a visible reward cost, and fully on-policy training (decay 00) reaches the highest diversity overall but the lowest reward, showing that without an off-policy reference the mirrored push drifts too far from the current policy. Decay 0.250.25 gives the best reward–diversity trade-off and is used in all main experiments.

B.3 Ablation of Rollout Length and Joint-Step Self-Noise

Rollout length. Figure 8 varies the rollout length. With 12 steps, reward drops quickly but diversity rises the highest, indicating that a short, coarse rollout leaves samples under-converged toward the high-reward mode. The 24- and 36-step variants are close on both metrics, with the 36-step rollout ending marginally higher on reward; since repair needs only about 50 optimizer steps, the extra sampling overhead is modest. We use 36 steps in the main configuration.

Self-noise at the joint step. At the joint (first) training timestep, ReNFT reuses the initial noise of the rollout ε\varepsilon rather than drawing fresh noise. Replacing this self-noise with fresh random noise at the joint step keeps reward higher but markedly weakens diversity recovery, consistent with aggravated hacking: fresh noise decouples the pull and push endpoints from the shared counterfactual trajectory. We retain self-noise at σt=1\sigma_{t}{=}1. Loading self-sampled noise at every training timestep was unstable and diverged, so we exclude that variant from the controlled comparison.

B.4 Ablation of Route Position and Contiguity

The main paper uses fixed route patterns in which the base step bb occupies the first position of each block and the unconditional step uu occupies the last. This section explains the design rationale.

Why bb at the block start.

Early denoising steps have the strongest influence on global structure (Jin, Shi, and Gu 2026): the vector field at high noise determines layout and composition before local details. Placing bb at the block start injects a base-like direction where it can most effectively steer toward suppressed structural alternatives. A base step at the block end, by contrast, would leave no subsequent θ\theta step to correct base-quality artifacts (blur, incomplete objects), which the reward ranking should filter out rather than recover.

Why uu at the block end.

The unconditional route exposes prompt-independent bias (recurring textures, palettes, stylistic artifacts). Placing uu at the block end lets it act on a nearly-formed image where this bias is observable. An earlier uu would steer toward the unconditional hacking mode before the conditional route establishes prompt-specific structure, yielding a pull ratio locked near 1.01.0; the pattern-length-2 ablation in the main paper shows this leads to degenerate outputs.

Why no consecutive b​bbb or u​uuu blocks.

The core principle is that bb and uu serve to inject diversity or expose bias, but they must not dominate the trajectory. Consecutive base steps (b​bbb) push the sampling manifold toward the base distribution for an extended interval, effectively reverting to base-like sampling and losing the acquired reward. Consecutive unconditional steps (u​uuu) amplify the prompt-independent bias for an extended interval, strengthening the very hacking direction we aim to suppress. In both cases the trajectory deviates too far from the current policy, destabilizing training and degrading both reward and diversity.

Toward a minimal repeating pattern.

The above considerations converge on a simple design: place bb at the start and uu at the end, avoid consecutive b​bbb or u​uuu blocks, and keep the majority of steps on the conditional route θ\theta. The most decoupled realization is to define the shortest pattern that satisfies these constraints and repeat it uniformly across the inference timesteps, which is exactly what the default (b​θ​θ,θ​θ​u)(b\theta\theta,\theta\theta u) configuration does. A systematic ablation over alternative positions (e.g., θ​θ​b\theta\theta b, θ​u​θ\theta u\theta) and consecutive blocks (e.g., b​b​θ​θbb\theta\theta) would further validate these choices and is left to future work.

Appendix C Repair Dynamics and Diagnostics

C.1 Dual-Route Pull-Ratio Dynamics

Figure 9: Dual-Route Pull Ratios during SD3.5-M PickScore Repair. (a) The raw route-AA (b​θ​θb\theta\theta, base-probing) pull ratio, which equals route AA’s win rate in the per-pair reward ranking before the adaptive flipping guard; (b) the guarded ratio after the guard enforces the ≥0.5\geq 0.5 per-group floor. The vertical dotted line marks the reported 50-step budget.

Figure 9 reports the dual-route pull ratios over 100 repair steps, extending beyond the reported 50-step budget. The raw route-AA pull ratio (route AA’s win rate before any flipping) starts at 0.840.84, drops below 0.50.5 around step 2020, and settles in the 0.320.32–0.420.42 band from step 5050 onward, ending at 0.400.40. This trajectory matches the diversity peak near step 3535 (Figure 2 of the main paper): as repair progresses, the unconditional route increasingly wins the ranking, i.e., it gradually re-hacks the reward model. The guarded ratio holds within 0.520.52–0.620.62 from step 1010 onward and never breaches the 0.50.5 floor, with the gap to the raw curve widening as the guard flips more low-margin pairs. The same pull-mass reasoning beyond step 5050 (continued repair draws pull mass from guard-flipped, lower-margin pairs while diversity has already passed its peak) directly motivates terminating repair at 5050 steps.

C.2 Unconditional Sampling Analysis

Refer to caption
Figure 10: Unconditional Sample Display across NFT Training and ReNFT Repair on SD3.5-M. Empty-prompt (CFG=0) samples from a fixed noise grid. The base generator branches into PickScore (top) and GenEval (bottom) repair trajectories, shown at the hacked checkpoint HH and repair steps 20 and 40.

Figure 10 tracks the same fixed noise indices across training and repair on SD3.5-M. The base generator spans unrelated subjects, scenes, styles, and palettes (Figure 1 of the main paper), whereas checkpoint HH maps the grid to protocol-specific hubs. Under PickScore, the nine samples concentrate on warm-toned, highly detailed female portraits with similar framing and backgrounds; under GenEval, they concentrate on isolated, full-body figures against plain backgrounds. Thus the two rewards induce different visual signatures, but both sharply reduce unconditional diversity. Repair progressively reopens the grid: step 2020 broadens appearance and context, and step 4040 produces more varied subjects, compositions, backgrounds, and palettes under the same noise indices. The repaired grids remain between the narrow HH hubs and the wider base range rather than simply reverting to Base. Because the same noise indices are shared across stages, each repaired sample partially retains the layout, style, and palette of the corresponding base sample (e.g., a blue background in Base tends to stay blue through HH and repair), and the correspondence strengthens as diversity recovers: the more varied grids at step 4040 align more closely with Base than the collapsed grids at HH. This same-noise contraction-and-reopening pattern supports the hub-reading use of the unconditional route (Section 4.2 of the main paper) and the suppression-not-deletion hypothesis (Section 4.1 of the main paper).

Refer to caption
Figure 11: Unconditional Sample Display on FLUX.2-klein-base. Empty-prompt samples from a fixed noise grid for the base model, the NFT checkpoint, and ReNFT (4B backbone on top, 9B at the bottom).

Figure 11 repeats the diagnostic on the two FLUX.2-klein-base backbones. The unconditional distribution again contracts from a diverse base onto a narrow hub (both backbones converge onto recurring portrait scenes after NFT) and reopens under repair with varied subjects, styles, and rendering modes, confirming that the contraction-and-reopening pattern, and hence the hub-reading use of the unconditional route, is not specific to SD3.5-M.

Appendix D Extended Comparisons

PickScore Test Set GenEval Test Set
Method PickScore ↑\uparrow Aesthetic ↑\uparrow ImageReward ↑\uparrow HPSv2 ↑\uparrow LPIPS-Div ↑\uparrow DreamSim-Div ↑\uparrow DINO-Div ↑\uparrow GenEval ↑\uparrow LPIPS-Div ↑\uparrow DreamSim-Div ↑\uparrow DINO-Div ↑\uparrow
SD3.5-M (21.73) (6.026) (1.06) (0.295) (0.599) (0.242) (0.251) (0.654) (0.687) (0.339) (0.379)
+ DiffusionNFT 23.51 6.592 1.44 0.327 0.430 0.119 0.112 0.946 0.292 0.149 0.186
+ E2PO (reproduced) 23.17 6.312 1.34 0.327 0.498 0.166 0.159 0.921 0.463 0.204 0.258
+ Ours 23.26 6.344 1.41 0.323 0.565 0.189 0.182 0.937 0.496 0.231 0.295
Table 3: Comparison with Locally Reproduced E2PO on SD3.5-M (PickScore and GenEval Test Sets). All metrics are computed under the same evaluation protocol. Unlike the main paper, which reuses E2PO’s paper-reported reward values, this table uses our local reproduction throughout. Shaded: in-domain training reward.
Method PickScore ↑\uparrow Aesthetic ↑\uparrow ImageReward ↑\uparrow HPSv2 ↑\uparrow LPIPS-Div ↑\uparrow DreamSim-Div ↑\uparrow DINO-Div ↑\uparrow
FLUX.2-klein-base-4B
Base (21.26) (5.995) (0.99) (0.289) (0.533) (0.246) (0.274)
+ DiffusionNFT 23.38 6.723 1.35 0.320 0.433 0.127 0.122
+ Ours 23.13 6.489 1.39 0.318 0.485 0.169 0.157
FLUX.2-klein-base-9B
Base (22.04) (5.921) (1.24) (0.298) (0.526) (0.232) (0.253)
+ DiffusionNFT 23.35 6.699 1.42 0.331 0.425 0.130 0.126
+ Ours 23.04 6.442 1.45 0.322 0.493 0.175 0.173
Table 4: Quantitative Results on Two FLUX.2-klein-base Backbones (PickScore Test Set). Shaded: in-domain training reward. Parenthesized rows: reference only, excluded from comparison.

D.1 Full Comparison with E2PO

E2PO (Hu et al. 2026) is not open-sourced, so its published checkpoint and evaluation code are unavailable for direct comparison. The main paper therefore reuses the paper-reported reward values of E2PO and omits diversity metrics. We have since reproduced E2PO locally, matching its official training budget and following its published configuration as closely as possible; Table 3 reports the full comparison using these reproduced values throughout. The reproduced reward (23.17 PickScore, 0.921 GenEval) is slightly lower than the paper-reported values (23.38, 0.932) used in the main paper, likely due to implementation differences. The main paper conservatively retains these stronger paper-reported values; Table 3 instead evaluates all methods under one protocol.

Compared to the reproduced E2PO, ReNFT achieves higher diversity on all six diversity metrics across both protocols while also retaining more reward (98.9% vs. 98.6% PickScore, 99.0% vs. 97.4% GenEval relative to the NFT endpoint). The interface-level approach does recover meaningful diversity, but at a larger reward cost and to a lesser extent than our internal-route repair.

D.2 Cross-Backbone Generalization

Table 4 extends the evaluation to two FLUX.2-klein-base (Black Forest Labs 2025) backbones (4B and 9B) on the PickScore test set. The NFT adapter of each backbone is trained for 200 steps and repaired for another 50 steps, with the same hyperparameters as on SD3.5-M except for the larger LoRA (rank 64, alpha 128); evaluation uses 20-step sampling, with the base model at CFG 4 and the NFT and ReNFT checkpoints CFG-free. The collapse pattern reproduces on both backbones: DiffusionNFT raises PickScore by 2.12 and 1.31 points over the base model while DreamSim-Div falls by 48% and 44%, respectively, matching the reward–diversity tension on SD3.5-M. ReNFT transfers without retuning: it retains 98.9% (4B) and 98.7% (9B) of NFT’s reward while improving DreamSim-Div by 33% and 35%, with consistent gains on LPIPS-Div and DINO-Div as well. The repair mechanism therefore generalizes beyond SD3.5-M: across a different backbone, LoRA capacity, training budget, and evaluation protocol, the same internal routes exist and the same 50-step budget recalibrates them without retuning.

Appendix E Extended Qualitative Results

Figures 12 and 13 compare Base, NFT, and ReNFT on five PickScore prompts per backbone, with four samples per method under different initial noises. We discuss three representative prompts per backbone below; the shared axolotl prompt enables a cross-backbone comparison at the end.

FLUX.2-klein-base-4B. On the axolotl-in-the-style-of-Minecraft prompt, all three methods adopt a side-view subject structure. Base has clean, simple backgrounds; NFT fixes the structure across samples with visible noise artifacts; ReNFT maintains the side-view structure but with fuller backgrounds and textures than Base, without NFT’s artifacts. On the storefront-with-AAAI-2027 prompt, ReNFT achieves higher text-spelling accuracy than both Base and NFT, with diversity approaching Base and far exceeding NFT. On the squirrel-gives-an-apple-to-a-bird prompt, Base shows high diversity in composition and style, though some samples have implausible object relationships; ReNFT produces more structurally coherent scenes with style diversity markedly higher than NFT.

FLUX.2-klein-base-9B. On the same axolotl-in-the-style-of-Minecraft prompt, the 9B base adopts a front-view subject structure, distinct from the side view of the 4B base, and NFT correspondingly collapses to a different hub, with ReNFT reopening within the 9B’s own range. On the fantasy-pastel-Wes-Anderson-pineapple-character prompt, Base already produces prompt-relevant pineapple characters; both NFT and ReNFT converge on pineapple-head with humanoid-body forms, likely reflecting PickScore’s preference for human-like subjects. ReNFT has cleaner and more varied backgrounds compared to NFT’s dense, ornate settings. On the sheep-holding-a-sign-that-says-play-chess prompt, Base shows high diversity; NFT fixes to a narrow mode with only minor structural differences across samples; ReNFT maintains multiple compositional styles and subject structures.

Cross-backbone axolotl comparison. Under the same prompt, the two backbones’ Base, NFT, and ReNFT samples remain visually distinct: 4B and 9B each concentrate on a different high-detail NFT hub and reopen toward different backbone-specific ranges. This provides qualitative evidence consistent with Section 4.1 of the main paper: post-training reweights the inherited distribution of each model rather than introducing a shared new mode (Figure 1 of the main paper). The backbone-specific reopening is also consistent with the suppression-not-deletion hypothesis.

Common NFT artifacts. Beyond the per-prompt structural collapse described above, the NFT columns in both grids share recurring visual symptoms: a warm, oversaturated color palette; densely textured or ornate backgrounds; and fine noise artifacts that give the images an over-processed appearance. These symptoms are consistent across both FLUX.2-klein-base backbones and match the hacking artifacts observed in the SD3.5-M unconditional display (Section 3.2), suggesting a shared reward-hacking signature rather than a backbone-specific artifact. These visual artifacts (over-processed textures, unnatural detail, and palette fixation) are not genuine quality improvements but shortcuts that overfit the preferences of the reward model, exploiting evaluation blind spots rather than producing subjectively higher-quality images.

Across the illustrated prompts, ReNFT is visibly more varied than NFT in composition, color, and rendering style, while Table 4 shows that this recovery retains most of NFT’s reward.

Refer to caption
Figure 12: Qualitative Comparison on FLUX.2-klein-base-4B. Each row is one PickScore prompt; each cell shows four samples from the same method under different initial noises.
Refer to caption
Figure 13: Qualitative Comparison on FLUX.2-klein-base-9B. Same layout as Figure 12.