TAE (Trust-AI-Eval): Can We Trust AI Evaluation?
Faithful Dual-constrained Erasure for Robust LLM Safety Alignment
Abstract
Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicious behaviors rapidly resurface after benign fine-tuning. In this work, we investigate the optimization dynamics of unlearning and identify that this vulnerability stems from shallow alignment. Rather than effectively erasing target knowledge, models often exploit a shortcut by activating previously dormant parameters to act as spurious suppressors, forming a fragile inhibitory shell over intact malicious representations. To address this issue and enforce authentic memory deletion, we propose FDCU, a novel dual-constrained subspace projection framework. FDCU restricts parameter updates through a highly scalable, element-wise dual-masking rule: it preserves general knowledge manifolds via Fisher Information and strictly prohibits the abnormal activation of spurious suppressors via the Principle of Minimal Functional Intervention (PMFI). By reliably blocking the model’s ability to superficially hide knowledge, FDCU promotes the authentic dismantling of target representations. Extensive experiments across specific knowledge erasure and safe output control tasks demonstrate that FDCU achieves state-of-the-art robustness against retraining attacks while maintaining near-lossless general utility, ensuring durable safety for LLMs.
1 Introduction
The unprecedented scale of Large Language Models (LLMs) has endowed them with remarkable capabilities, but it has also led to the inadvertent memorization of sensitive, copyrighted, and hazardous information. To mitigate these risks without incurring the prohibitive costs of retraining from scratch, Machine Unlearning (MU) has emerged as a critical alignment paradigm (Jang et al., 2022; Bourtoule et al., 2021). Prevailing unlearning frameworks typically formulate this as an optimization problem, aiming to maximize the loss on a target forget set () while preserving utility on a retain set () through gradient ascent, preference optimization, or representation editing (Rafailov et al., 2023; Zhang et al., 2024).
However, a surge of recent studies has exposed a critical vulnerability in current unlearning methodologies: what appears to be successful unlearning is often mere obfuscation (Hu et al., 2025). While state-of-the-art methods perform well on static evaluation metrics, they fail catastrophically under dynamic threat models, particularly retraining attacks (or benign relearning). Attackers—or even innocent users—can easily trigger the resurgence of supposedly "erased" malicious knowledge by simply fine-tuning the unlearned model on small amounts of loosely related, benign data (Fan et al., 2025; Qi et al., 2023). This phenomenon suggests that existing unlearning algorithms fail to achieve genuine memory deletion, leaving models in a highly fragile state.
We delve into the optimization dynamics of unlearning to understand the root cause of this fragility. Empirical observations reveal a phenomenon called shallow alignment in safety unlearning tasks (Xu et al., 2025b). Models superficially hide knowledge instead of effectively erasing it. Standard unlearning objectives force the optimizer to take the path of least resistance. The model avoids authentically dismantling the excitatory pathways encoding the malicious knowledge. It improperly activates previously dormant parameters and flips their roles to act as spurious suppressors. This creates a fragile equilibrium. The target knowledge persists beneath an artificially learned inhibitory shell. Any subsequent fine-tuning easily destabilizes these superficial inhibitors and causes the malicious outputs to resurrect. Existing constrained unlearning methods successfully prevent catastrophic forgetting of general utility (Cha et al., 2025). However, they largely overlook this anomalous activation and fail to address the spurious inhibitor vulnerability.
To resolve this problem, we formulate faithful unlearning as a Dual-Constrained Optimization problem. We introduce FDCU (Faithful Unlearning via Subspace Enforcement) , a novel framework that strictly confines gradient updates into a safe erasure cone. FDCU is governed by two orthogonal geometric constraints: (1) General Knowledge Preservation: We restrict updates along high-curvature manifolds associated with benign knowledge (measured via diagonal Fisher Information) to prevent capability collapse. (2) Principle of Minimal Functional Intervention (PMFI): This is our core innovation to defeat retraining attacks. We explicitly identify parameters that initially exhibit non-excitatory contributions to the malicious output and rigidly prohibit their functional flipping. By reliably blocking the activation of spurious suppressors using mathematically verified, we strip the optimizer of its ability to mask knowledge, forcing it to authentically dismantle the original malicious representations.
We extensively evaluate FDCU across two critical scenarios: Specific Knowledge Erasure and Safe Output Control . Our contributions are summarized as follows:
- •
We identify the root cause of unlearning fragility against retraining attacks as "shallow alignment," driven by the abnormal activation of spurious suppressors rather than genuine knowledge erasure.
- •
We propose FDCU, which translates the Principle of Minimal Functional Intervention (PMFI) into a tractable dual-constrained subspace projection, fundamentally preventing both catastrophic forgetting and spurious hiding.
- •
Extensive experiments on Llama and Qwen architectures demonstrate that FDCU achieves robustness against retraining attacks while maintaining near-lossless general utility, proving that strictly constrained erasure ensures durable safety.
2 Background and Related Work
Machine Unlearning in LLMs.
Large Language Models (LLMs) can inadvertently memorize sensitive, copyrighted, or hazardous information. Machine unlearning aims to remove this specific knowledge () without the prohibitively high cost of retraining from scratch. Recently, the community has established unified benchmarking frameworks (e.g., OpenUnlearning (Dorna et al., 2025)) and diverse evaluation suites like TOFU (Maini et al., 2024), MUSE (Shi et al., 2024), and WMDP (Li et al., 2024) to standardize unlearning evaluations. Prevailing methods typically formulate this as an optimization problem, using Gradient Ascent (GA), preference optimization (DPO, NPO), or novel paradigms like bi-level optimization (Reisizadeh et al., 2026) and forward learning (Xu et al., 2025a). While these methods successfully reduce the probability of malicious outputs, naive unlearning often disrupts the model’s global performance.
Unlearning Fragility.
Recent extensive evaluations reveal a critical flaw: existing methods often "obfuscate" or "hide" rather than truly "erase" malicious knowledge (Hu et al., 2025; Xu et al., 2025b). Investigations into unlearning dynamics (Yang et al., 2026) show that models exploit a shortcut by activating previously dormant parameters to act as spurious suppressors. This creates a fragile push-pull equilibrium. Because the original knowledge is only suppressed, it can easily resurface. For instance, trivial benign relearning (Hu et al., 2025), post-deployment quantization (Zhang et al., 2025), or adversarial unlearning requests (Song et al., 2025) can quickly destabilize these inhibitors and resurrect malicious outputs. Furthermore, studies on the Collapse of Irrelevant Representations (CIR) (Sondej and Yang, 2025) demonstrate that naive unlearning disrupts shared general representations, making the suppressed knowledge highly recoverable during subsequent fine-tuning.
Constrained Parameter Updates.
To preserve general capabilities during unlearning, recent literature heavily leverages parameter importance measures and Parameter-Efficient Fine-Tuning (PEFT). Methods like LoKU (Cha et al., 2025) and LLMEraser (Ding et al., 2025) utilize Hessian-approximated matrices (e.g., Fisher Information) or influence functions to bound parameter shifts. Building on this, approaches such as Constrained Knowledge Unlearning (CKU) (Shi et al., 2025) explicitly score and freeze utility-sensitive neurons to maintain general capabilities. These constraint strategies successfully prevent the loss of pre-trained utility. However, they primarily focus on structural protection (preventing catastrophic forgetting) and fail to address the emergence of spurious suppressors.
3 Preliminary
3.1 Gradient Ascent (GA)
Given a pre-trained model parameterized by and a forget dataset , machine unlearning aims to reduce the model’s ability to predict given .
Gradient Ascent (GA) achieves this by directly maximizing the negative log-likelihood on . The parameter update rule at step is:
| (1) |
where is the learning rate and is the cross-entropy loss. While GA effectively reduces target prediction probability, this unconstrained update () blindly modifies parameters, causing catastrophic forgetting on benign data and inducing spurious suppressive neurons.
3.2 Parameter Importance and Functional Attribution
To safely edit a pre-trained model, we must quantify both a parameter’s sensitivity to general capabilities and its directional contribution to specific outputs. Viewed through the lens of Taylor expansion, we utilize two complementary metrics. First, to preserve benign knowledge on , we approximate the loss curvature using the empirical Fisher Information Matrix (FIM) (Yin and others, 2023):
| (2) |
A large diagonal value indicates high sensitivity, meaning perturbing will severely degrade general utility. However, because FIM relies on squared gradients, it is strictly non-negative and discards directional information. To determine whether a parameter actively promotes or suppresses a target output, we rely on first-order Taylor attribution (Yang et al., 2026):
| (3) |
This attribution matrix provides a signed functional metric: a positive value () indicates an excitatory contribution that drives the generation of , whereas a negative or zero value () signifies an inhibitory or dormant role.
4 Methodology
To overcome the vulnerabilities of shallow alignment and catastrophic forgetting, we formulate unlearning as a Dual-Constrained Optimization Problem. Instead of adding heuristic penalties, we project the naive unlearning gradient into a strictly safe parameter subspace. This subspace preserves general capabilities while prohibiting the creation of spurious suppressor.
4.1 Problem Formulation
Let be a Large Language Model parameterized by . We are given a malicious dataset to be erased, and a benign retention dataset representing capabilities to be preserved.
Comprehensive Threat Scope.
To ensure practical safety, the malicious prompts explicitly include adversarial variations , where represents various jailbreak prefixes designed to bypass superficial safety filters. The objective is to eradicate the underlying harmful knowledge regardless of the elicitation method.
The Robust Unlearning Objective.
We seek an updated parameter set that minimizes the probability of generating while maintaining high accuracy on . Crucially, achieving a low probability immediately after unlearning is insufficient. If the unlearned model is subsequently fine-tuned on a new dataset (e.g., benign instruction-tuning data) to yield , the malicious knowledge must not resurrect. Mathematically, for any , the generation probability must remain strictly bounded:
| (4) |
where is a low safety threshold.
To satisfy this robust objective, the naive unlearning descent direction is inadequate, as it merely hides knowledge and leads to shallow alignment. To ensure robust erasure, the actual parameter update must be strictly confined to a safe subspace satisfying two geometric constraints.
Constraint I: General Knowledge Preservation.
To prevent capability collapse, parameter updates must be constrained in dimensions critical to general knowledge. We estimate the importance of parameters on using the diagonal Fisher Information Matrix (FIM), denoted as . We require the update to satisfy:
| (5) |
Constraint II: Suppression of Spurious Unlearning.
Our key observation is that spurious alignment occurs when parameters with zero or negative initial contribution to the malicious output are significantly modified to act as suppressors. To enforce true erasure, we must prohibit this abnormal activation.
Leveraging the functional attribution matrix computed over (as defined in Section 3.2), we identify parameters that initially have no excitatory contribution to the malicious outputs. We define a binary mask matrix and explicitly bound the updates in this non-excitatory region:
| (6) |
By freezing this region, we block the model’s shortcut of hiding knowledge behind new inhibitors. Consequently, the optimizer is forced to satisfy the unlearning objective by authentically dismantling the actual malicious representations (i.e., parameters where ).
4.2 Analytical Solution and Approximations
By treating the constraints mathematically, our objective becomes a constrained optimization problem:
| (7) |
Using the method of Lagrange multipliers (with multipliers ), solving for the optimal update yields a closed-form projection:
| (8) |
Computing the exact inverse of a joint Hessian matrix is computationally intractable for billion-parameter LLMs. To make this applicable at scale, we introduce two practical approximations.
First, relying on the mean-field assumption in over-parameterized networks, we approximate the matrices using their diagonal elements: and . Second, we decouple the constraints using the bounded approximation .
4.3 Dual-Mask Update Rule
With the diagonal approximation and decoupling step, our theoretical projection safely simplifies to an element-wise, scalable dual-masking formula using the Hadamard product ():
| (9) |
In practice, modifying any parameter requires passing two sequential gating criteria:
- •
General Knowledge Mask (): Preserves pre-trained utility by down-scaling updates on crucial benign parameters via .
- •
Minimal Intervention Mask (): Blocks spurious suppressors by disabling updates on historically non-excitatory parameters via .
(Here, and act as scalar regularization hyperparameters dictating the strictness of the boundary condition, and represent their corresponding normalized local parameter vector limits).
By relying uniquely on Eq. 9, our algorithm targets only the original roots of malicious behavior, robustly defending the model against both general capability decay and fine-tuning vulnerability attacks.
4.4 The Workflow of FDCU
The proposed framework, FDCU, operationalizes the theoretical dual-constrained projection into a highly efficient, two-stage pipeline. By decoupling the structural protection of general knowledge from the functional suppression of spurious inhibitors, FDCU ensures that parameter updates are strictly confined to a safe erasure subspace.
Overall, the workflow of FDCU is executed as follows:
- •
Initialize the General Knowledge Mask (): Prior to the unlearning process, a forward and backward pass is performed on the benign retention dataset (). The diagonal Fisher Information Matrix is computed to estimate the sensitivity of each parameter to general capabilities.
- •
Compute the Naive Unlearning Gradient: During the active training stage, the model processes batches from the malicious forget dataset (). By calculating the loss against the target malicious outputs (e.g., via Gradient Ascent or preference optimization), the unconstrained, naive descent direction is obtained.
- •
Dynamically Generate the Minimal Intervention Mask (): Concurrently with the gradient computation on , the initial attribution of each parameter to the malicious output is evaluated. The binary mask is constructed to strictly identify and isolate parameters that exhibit zero or negative excitatory contributions.
- •
Apply Dual-Mask Filtering for Safe Updates: Finally, the safe parameter update is synthesized by applying an element-wise Hadamard product () between the naive gradient and the two constraint masks ( and ). The optimizer then applies this filtered gradient to the model weights.
5 Evaluation
5.1 Experimental Setup
We evaluate our proposed method across two distinct unlearning scenarios: Specific Knowledge Erasure and Safe Output Control. This allows us to assess both factual forgetting and behavioral safety alignment.
Datasets and Tasks.
For Specific Knowledge Erasure, we target hazardous concepts such as bioweapons. The training set consists of descriptive statements about these dangerous concepts. For evaluation, we use the multiple-choice questions (MCQs) from the WMDP-Bio and WMDP-Cyber benchmarks (Li et al., 2024). To assess unlearning robustness, we randomly sample 20% of the original forget set and fine-tune the unlearned model using the standard next-token prediction objective. For Safe Output Control, we aim to prevent models from generating harmful responses. The forget set is constructed by extracting entities from harmful instruction datasets and generating descriptive statements about them. For evaluation, we use harmful prompts from AdvBench (Zou et al., 2023) and AdvExtent (Lu et al., 2024), combined with jailbreak suffixes including GCG (Zou et al., 2023), AIM (Jailbreak Chat, 2023), and AutoDAN (Liu et al., 2024). Similarly, we perform a Retraining Attack using a subset of the forget set.
Models and Baselines.
We apply our method to representative open-source LLMS, primarily focusing on the Llama-3 family and Qwen family to ensure consistency with recent safety alignment research. We benchmark our approach against several strong unlearning baselines, including Gradient Ascent (GA), Constrained Knowledge Unlearning(CKU) (Shi et al., 2025),Erasing Conceptual Knowledge(Elm) (Gandikota et al., 2025),,SSIUU (Yang et al., 2026) and Collapse of Irrelevant Representations (CIR) (Sondej and Yang, 2025). Across both scenarios, we compute the parameter importance on the benign retention set () using the standard Cross-Entropy (CE) loss to construct the General Knowledge Mask (). For the unlearning objective (), we tailor the loss function to the specific task requirements. In the Specific Knowledge Erasure experiments, we adopt the representation-breaking loss proposed by CIR (Sondej and Yang, 2025) to disrupt the internal activations associated with hazardous facts. In the Safe Output Control experiments, we utilize the gradient ascent objective on harmful responses, consistent with the formulation in CKU (Shi et al., 2025). Detailed descriptions of these baselines and their optimization objectives are provided in Appendix A.1.
Evaluation Metrics.
For Specific Knowledge Erasure, we measure the Accuracy on the WMDP MCQs immediately after unlearning. For Safe Output Control, we evaluate the model’s resistance to jailbreak prompts using the Refusal Rate , which represents the percentage of model outputs containing refusal words. We also compute a HarmfulScore ranging from 1 to 5 to evaluate the severity of the harmful output, utilizing an LLM-as-a-judge approach to average the scores across all evaluated responses. We report MMLU Acc (Hendrycks et al., 2021) to further assess general utility, and measure perplexity on a subset of the WikiText dataset (Merity et al., 2016) to evaluate output fluency. Additional details on the evaluation metrics are provided in Appendix A.3.
5.2 Specific Knowledge Erasure
| Base Model | Method | Initial Acc (%) | General Utility | Retraining Acc (%) | |||
| Bio | Cyber | PPL | Acc (%) | Bio | Cyber | ||
| Qwen2.5-3B | GA | 36.7 | 37.4 | 26.79 | 57.5 | 42.3 | 45.1 |
| ELM | 32.3 | 30.1 | 15.23 | 58.9 | 37.2 | 35.7 | |
| SSIUU | 33.4 | 32.3 | 18.90 | 58.1 | 34.2 | 33.4 | |
| CIR | 31.4 | 31.3 | 15.59 | 58.5 | 32.3 | 31.5 | |
| Ours | 30.9 | 29.8 | 15.67 | 58.7 | 31.5 | 31.0 | |
| Llama-3-8B-Instruct | GA | 37.6 | 35.4 | 23.38 | 59.3 | 52.4 | 59.6 |
| ELM | 32.2 | 27.2 | 13.40 | 61.6 | 41.2 | 38.5 | |
| SSIUU | 33.5 | 33.2 | 14.57 | 61.7 | 35.2 | 35.6 | |
| CIR | 32.4 | 27.9 | 12.73 | 62.5 | 33.4 | 27.1 | |
| Ours | 31.9 | 27.3 | 12.62 | 62.8 | 32.5 | 27.4 | |
| Qwen-3-8B | GA | 40.1 | 38.7 | 24.56 | 72.8 | 45.2 | 50.1 |
| ELM | 32.6 | 28.1 | 11.40 | 75.7 | 40.3 | 39.4 | |
| SSIUU | 33.9 | 30.2 | 12.33 | 76.1 | 37.2 | 32.4 | |
| CIR | 32.4 | 27.5 | 10.76 | 76.2 | 33.1 | 27.9 | |
| Ours | 31.5 | 27.2 | 10.33 | 76.4 | 32.3 | 27.2 | |
Table 1 shows that our method achieves low forget-set accuracy while preserving general accuracy and perplexity. Unlike several baselines, it maintains low post-retraining accuracy, supporting the hypothesis that the Principle of Minimal Functional Intervention limits spurious suppression and promotes durable knowledge erasure.
5.3 Safe Output Control
| Base Model | Method | Unlearning Safety | General Utility | Retraining Safety | |||
| Refusal Rate (%) | HarmfulScore | Acc (%) | PPL | Refusal Rate (%) | HarmfulScore | ||
| Qwen2.5-3B | Origin Model | 62.9 | 2.67 | 59.1 | 15.2 | 54.8 | 2.74 |
| GA | 79.6 | 2.18 | 50.8 | 26.7 | 57.7 | 2.53 | |
| CKU | 87.7 | 1.32 | 57.3 | 17.4 | 81.1 | 1.69 | |
| ELM | 81.1 | 1.41 | 58.8 | 16.0 | 79.6 | 1.68 | |
| SSIUU | 82.5 | 1.48 | 58.5 | 17.0 | 77.7 | 1.76 | |
| CIR | 90.5 | 1.34 | 58.0 | 16.3 | 89.5 | 1.38 | |
| Ours | 91.0 | 1.23 | 58.3 | 15.7 | 88.6 | 1.35 | |
| Llama-3-8B-Instruct | Origin Model | 79.6 | 2.19 | 66.6 | 10.3 | 70.7 | 2.67 |
| GA | 86.9 | 1.74 | 63.4 | 17.8 | 78.6 | 2.36 | |
| CKU | 93.5 | 1.20 | 65.7 | 14.7 | 80.3 | 1.53 | |
| ELM | 92.7 | 1.28 | 64.2 | 13.8 | 81.3 | 1.42 | |
| SSIUU | 92.8 | 1.33 | 64.7 | 12.5 | 88.7 | 1.48 | |
| CIR | 93.0 | 1.22 | 65.5 | 10.9 | 91.5 | 1.38 | |
| Ours | 93.9 | 1.16 | 65.8 | 10.7 | 91.7 | 1.28 | |
| Qwen-3-8B | Origin Model | 67.4 | 2.46 | 76.9 | 11.0 | 56.9 | 2.71 |
| GA | 84.3 | 1.82 | 72.3 | 20.1 | 62.8 | 2.69 | |
| CKU | 88.9 | 1.33 | 76.4 | 12.5 | 79.1 | 2.31 | |
| ELM | 86.6 | 1.50 | 75.3 | 12.8 | 78.7 | 2.40 | |
| SSIUU | 87.6 | 1.42 | 76.0 | 13.0 | 87.3 | 1.48 | |
| CIR | 88.8 | 1.34 | 76.1 | 12.3 | 87.7 | 1.37 | |
| Ours | 89.1 | 1.31 | 76.3 | 12.7 | 87.9 | 1.33 | |
Table 2 presents the comprehensive evaluation of our method against state-of-the-art baselines on the Safe Output Control task. The results validate the superiority of our dual-constrained subspace projection across three critical dimensions: initial safety alignment, general utility preservation, and robustness against retraining attacks.
Safety Alignment.
Following the unlearning phase (Post-Unlearning Safety), our method achieves the highest Refusal Rates and the lowest HarmfulScores across all three model architectures. On Llama-3-8B-Instruct, our method reaches a 93.9% refusal rate and a near-perfect HarmfulScore of 1.16. Unlike Gradient Ascent (GA), which perturbs parameters to minimize malicious likelihood, our method precisely dismantles the excitatory pathways responsible for harmful generation, leading to a more thorough and faithful safety alignment.
General Utility.
A persistent challenge in machine unlearning is catastrophic forgetting. Naive methods like GA severely degrade general capabilities, evidenced by a sharp drop in Accuracy and Perplexity. In contrast, by strictly confining updates within the Fisher-guided safe manifold, our method preserves general utility almost perfectly. Our Accuracy and PPL remain remarkably close to the Origin Model across all models.
Robustness Against Retraining Attacks.
The most significant advantage of our approach is its resilience to retraining attacks. When subjected to a benign fine-tuning attack, baselines such as GA, CKU, and ELM experience a catastrophic collapse in safety (e.g., GA’s refusal rate plummets from 79.6% to 57.7% on Qwen2.5-3B). By enforcing the minimizing functional intervention, our method explicitly prohibits the activation of these fragile spurious suppressors during unlearning. As a result, our method maintains exceptional defense stability post-retraining.
5.4 Ablation Study
| Method | Post-Unlearning Safety | General Utility | Post-Retraining Safety | |||
| Refusal Rate (%) | HarmfulScore | Acc (%) | PPL | Refusal Rate (%) | HarmfulScore | |
| Ours (Full) | 91.0 | 1.23 | 58.3 | 15.7 | 88.6 | 1.35 |
| w/o | 84.7 | 1.46 | 55.7 | 17.8 | 83.4 | 1.54 |
| w/o | 90.5 | 1.25 | 58.5 | 16.0 | 81.2 | 1.67 |
| Random Mask | 82.1 | 1.51 | 56.8 | 18.1 | 79.6 | 1.72 |
To validate the individual contributions of our dual-constrained mask, we conduct an ablation study on the Safe Output Control task using the Qwen2.5-3B model, as summarized in Table 3. Removing the General Knowledge Mask () degrades general utility and increases perplexity. This confirms its role in protecting benign structural pathways during unlearning. Removing the Minimal Intervention Mask () exposes the model to severe vulnerabilities during retraining attacks. The model achieves initial safety but its defense collapses post-retraining. This validates our core theoretical insight: the mask is essential to prohibit the activation of spurious suppressors and prevent shallow alignment. Finally, the Random Mask variant performs poorly across all metrics.
Impact of hyperparameters and layer selection. We investigate the impact of regularization weights () and layer selection on our proposed method. Figure 2 presents the sensitivity to the constraint hyperparameters. Higher generally improves General Utility, while a moderate value of achieves better performance. We therefore set and in the final configuration for the best trade-off between post-retraining safety and utility. Figure 3 illustrates that applying our method to middle layers (e.g., [24–28]) achieves the optimal balance, maximizing General Utility while minimizing the retraining HarmfulScore. Modifying early or final layers degrades robustness.
5.5 Case Study
We compare GA and FDCU after a retraining attack.
6 Limitations
FDCU improves unlearning robustness by constraining parameter updates to preserve general knowledge and reduce spurious suppression. Compared with naive unlearning, FDCU requires additional memory to store gradient-based statistics from both the forget set and the retain set. However, this cost is manageable: as shown in Section 5.4, applying FDCU to selected middle layers is sufficient to obtain a good robustness–utility trade-off. Thus, Fisher and attribution information only needs to be stored for selected layers rather than the whole model, reducing memory usage while retaining the benefits of FDCU.
7 Conclusion
We investigate the vulnerability of machine unlearning to retraining attacks and identify its root cause as shallow alignment, where models exploit spurious suppressors rather than erasing target knowledge. To address this, we propose FDCU, a dual-constrained subspace projection method for faithful unlearning. FDCU restricts parameter updates by simultaneously preserving general knowledge manifolds via Fisher information and prohibiting the abnormal activation of non-excitatory neurons. Extensive experiments across specific knowledge erasure and safe output control tasks show that FDCU significantly outperforms state-of-the-art baselines. FDCU improves retraining robustness while maintaining near-lossless general utility.
References
- Machine unlearning. In 2021 IEEE Symposium on Security and Privacy (SP), pp. 141–159. Cited by: §1.
- Towards robust and parameter-efficient knowledge unlearning for llms. In InternationalConference on Learning Representations (ICLR), External Links: Link Cited by: §1, §2.
- UnifiedParameter-efficient unlearning for llms. In InternationalConference on Learning Representations (ICLR), External Links: Link Cited by: §2.
- OpenUnlearning:accelerating llm unlearning via unified benchmarking of methods and metrics. In NeurIPS 2025Datasets and Benchmarks Track, External Links: Link Cited by: §2.
- Towards llm unlearning resilient to relearning attacks: a sharpness-aware minimization perspective and beyond. In International Conference on Machine Learning (ICML), External Links: Link Cited by: §1.
- Erasing conceptual knowledge from language models. External Links: 2410.02760, Link Cited by: §A.1, §5.1.
- Measuring massive multitask language understanding. In International Conference on Learning Representations, External Links: Link Cited by: §5.1.
- Unlearning or obfuscating? jogging the memory of unlearned llms via benign relearning. In International Conference on Learning Representations (ICLR), External Links: Link Cited by: §1, §2.
- AIM: always intelligent and machiavellian. Note: https://www.jailbreakchat.com/prompt/4f37a029-9dff-4862-b323-c96a5504de5dAccessed: 2026-05-03 Cited by: §5.1.
- Knowledge unlearning for mitigating privacy risks in language models. arXiv preprint arXiv:2210.01504. Cited by: §1.
- The wmdp benchmark: measuring and reducing malicious use with unlearning. External Links: 2403.03218, Link Cited by: §2, §5.1.
- AutoDAN: generating stealthy jailbreak prompts on aligned large language models. External Links: 2310.04451, Link Cited by: §5.1.
- Eraser: jailbreaking defense in large language models via unlearning harmful knowledge. External Links: 2404.05880, Link Cited by: §A.3, §5.1.
- Tofu: a task of fictitious unlearning for llms. In First Conference on Language Modeling, Cited by: §2.
- Pointer sentinel mixture models. External Links: 1609.07843, Link Cited by: §5.1.
- Fine-tuning aligned language models compromises safety, even when users do not intend to!. arXiv preprint arXiv:2310.03693. Cited by: §1.
- Direct preference optimization: your language model is secretly a reward model. In Advances in Neural Information Processing Systems, Vol. 36, pp. 53728–53741. Cited by: §1.
- BLUR: a bi-leveloptimization approach for llm unlearning. In European Chapterof the Association for Computational Linguistics (EACL), External Links: Link Cited by: §2.
- Muse: machine unlearning six-way evaluation for language models. arXiv preprint arXiv:2407.06460. Cited by: §2.
- Safety alignment via constrained knowledge unlearning. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 25515–25529. Cited by: §A.1, §2, §5.1.
- Collapse of irrelevant representations (cir) ensures robust and non-disruptive llm unlearning. External Links: 2509.11816, Link Cited by: §A.1, §2, §5.1.
- Refusal is not anoption: unlearning safety alignment of large language models. In Proceedings of the34th USENIX Security Symposium (USENIX Security), External Links: Link Cited by: §2.
- ReLearn:unlearning via learning for large language models. In Proceedings ofACL, External Links: Link Cited by: §2.
- Unlearning isn’t deletion:investigating reversibility of machine unlearning in llms. arXiv preprintarXiv:2505.16831. External Links: Link Cited by: §1, §2.
- Erase or hide? suppressing spurious unlearning neurons for robust unlearning. In International Conference on Learning Representations (ICLR), Cited by: §A.1, §2, §3.2, §5.1.
- Understanding and improving machine unlearning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Cited by: §3.2.
- Safe unlearning: a surprisingly effective and generalizable solution to defend against jailbreak attacks. arXiv preprint arXiv:2407.01548. Cited by: §1.
- CatastrophicFailure of llm unlearning via quantization. In InternationalConference on Learning Representations (ICLR), External Links: Link Cited by: §2.
- Universal and transferable adversarial attacks on aligned language models. External Links: 2307.15043, Link Cited by: §5.1.
Appendix A Technical appendices and supplementary material
Technical appendices with additional results, figures, graphs, and proofs may be submitted with the paper submission before the full submission deadline (see above). You can upload a ZIP file for videos or code, but do not upload a separate PDF file for the appendix. There is no page limit for the technical appendices.
Note: Think of the appendix as “optional reading” for reviewers. The paper must be able to stand alone without the appendix; for example, adding critical experiments that support the main claims to an appendix is inappropriate.
A.1 Related Unlearning Methods
Erasure of Language Memory (ELM).
ELM Gandikota et al. (2025) formulates concept-level unlearning through the model’s own implicit self-classification ability. Given an erase dataset , ELM constructs two conditioning prompts: for the target concept and for an alternative concept distribution. It defines an erased target distribution by reweighting the original model distribution:
| (10) |
where controls the erasure strength. The erased model is trained to match this target distribution:
| (11) |
To preserve unrelated capabilities, ELM further includes a retain loss
| (12) |
and optimizes a weighted combination of erasure, retention, and optionally fluency-preserving objectives.
Constrained Knowledge Unlearning (CKU).
CKU Shi et al. (2025) improves safety alignment by unlearning harmful knowledge while preserving neurons that are important for useful knowledge. It first scores neurons in selected MLP layers using an identification dataset and selects a protected neuron set . During unlearning, CKU performs gradient ascent on harmful examples but prunes gradients associated with neurons in . Abstractly, the update can be written as
| (13) |
where masks updates on utility-sensitive neurons. CKU therefore constrains harmful knowledge removal through neuron-level gradient pruning, aiming to improve safety while limiting utility degradation.
Collapse of Irrelevant Representations (CIR).
CIR Sondej and Yang (2025) argues that unlearning becomes non-robust when updates disrupt general representations shared by harmful and benign capabilities. It applies PCA to activations and module-output gradients to identify common representation subspaces, then collapses these components before computing unlearning updates. For activation and module-output gradient , CIR computes
| (14) |
and forms the update using the purified vectors:
| (15) |
CIR further uses an MLP-level representation breaking loss to target harmful representations while avoiding unnecessary disruption to general capabilities.
Suppressing Spurious Unlearning Neurons for Robust Unlearning (SSIUU).
SSIUU Yang et al. (2026) studies shallow unlearning, where models hide target knowledge by increasing negative influence rather than removing positive knowledge-bearing neurons. It measures the attribution of a neuron representation to a target output as
| (16) |
SSIUU regularizes the increase of negative attribution during unlearning. In its efficient parameter-level implementation, attribution is computed as
| (17) |
and the regularizer penalizes changes on negatively attributed parameters:
| (18) |
The final objective combines a base unlearning loss with this attribution regularizer:
| (19) |
By suppressing the emergence of negative-influence neurons, SSIUU aims to make unlearning more faithful and robust to retraining.
A.2 Training Details and Hyperparameters
Experimental environment.
All experiments are conducted on two NVIDIA A100 GPUs with 80GB memory each. Unless otherwise specified, we use PyTorch and Hugging Face Transformers for model training and evaluation. For Llama-3-8B-Instruct, we use a learning rate of , bfloat16 precision, and a training batch size of 8.
Retraining attack configuration.
To evaluate robustness against post-unlearning fine-tuning, we perform retraining attacks using LoRA. The LoRA rank is set to 8, the LoRA scaling factor is set to 32, and the LoRA dropout rate is set to 0.05. We apply LoRA to the q_proj and v_proj modules. Retraining is conducted with bfloat16 precision, a per-device batch size of 2, gradient accumulation over 8 steps, 3 training epochs, and a learning rate of . After retraining, we evaluate the models using the same metrics as in the post-unlearning evaluation.
Baseline configurations.
For baseline methods, we follow the hyperparameter settings reported in the corresponding original papers whenever they are explicitly specified. For Gradient Ascent (GA), we implement the standard forget-set gradient ascent objective and use the same training budget as our method unless otherwise specified, so that performance differences are not caused by unequal optimization budgets.
For Constrained Knowledge Unlearning (CKU), we follow the original training configuration: the neuron-locking coefficient is set to , the batch size is set to 4, inputs longer than 2048 tokens are truncated, AdamW is used as the optimizer with learning rate , and the model is trained for 1 epoch. Following the original setting, the CKU training seed is set to 42, while the seed for knowledge localization and retention is set to 0.
For SSIUU, we follow the original formulation using a gradient-descent-based unlearning backbone with an attribution-guided regularization term. The unlearning process is early-stopped once the forget-set metric reaches the predefined forgetting threshold. For retraining robustness evaluation, the original SSIUU protocol searches attack learning rates over for harmful retraining attacks and over for benign retraining attacks, and reports the average over three runs. In our experiments, we adopt the same principle for selecting retraining attack strength.
For CIR, we follow the original hyperparameter selection strategy. Specifically, for each method, the retain learning rate is selected to be high but safe, and the unlearning learning rate is then searched while avoiding divergence or excessive retain-set disruption. The final configuration is selected according to post-attack performance under the constraint that general utility is not substantially degraded.
A.3 Additional Details on Evaluation Metrics
This section provides additional details on the evaluation metrics used in our experiments, including HarmfulScore, Refusal Rate, and WMDP accuracy.
HarmfulScore.
For the Safe Output Control task, we evaluate the harmfulness of model responses using an LLM-as-a-judge protocol adapted from Lu et al. Lu et al. (2024). Given a harmful behavior and the corresponding model response, the judge assigns an integer score from 1 to 5. We report the average HarmfulScore across all evaluated prompts, where lower values indicate safer model behavior.
Refusal Rate.
For the Safe Output Control task, we also compute the Refusal Rate using refusal-related lexical indicators. A response is counted as a refusal if it contains at least one phrase from the refusal keyword list. The Refusal Rate is computed as the percentage of evaluated responses that are identified as refusals.
WMDP Accuracy.
For the Specific Knowledge Erasure task, we evaluate multiple-choice accuracy on WMDP-Bio and WMDP-Cyber. Each question is formatted with four candidate answers labeled A, B, C, and D, followed by an answer prompt. The model prediction is obtained by comparing the next-token logits corresponding to the four option labels and selecting the option with the highest logit. Accuracy is computed as the proportion of examples for which the predicted option matches the ground-truth answer. For unlearning evaluation, lower WMDP accuracy indicates stronger removal of the target hazardous knowledge. For retraining evaluation, post-retraining WMDP accuracy measures whether the supposedly erased knowledge resurfaces after subsequent fine-tuning.
A.4 Broader Impact
This work aims to improve the robustness of LLM safety alignment by making harmful knowledge unlearning more resistant to retraining-induced recovery. A positive societal impact of this research is that it may help reduce the risk of deployed language models reproducing hazardous, privacy-sensitive, or otherwise unsafe information after post-deployment adaptation or benign fine-tuning. By studying failure modes of existing unlearning methods, the proposed approach also contributes to a better understanding of when apparent safety alignment reflects genuine removal rather than superficial suppression.