arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.00064v1 [cs.LG] 30 Aug 2026

Attention Sensitivity Is Not Enough:
Dissociating Attention-Level and Behavioural
In-Context Learning under Fine-Tuning

Jinyuan Zhang Peng He∗
202621116012480@stu.hubu.edu.cn penghe@hubu.edu.cn
Yin Yuan He Hu
202521120012766@stu.hubu.edu.cn 202521120012751@stu.hubu.edu.cn
ShengShuo Jiao
202621120012764@stu.hubu.edu.cn

Hubei University, Wuhan, China

∗Corresponding author: penghe@hubu.edu.cn

Abstract

In-context learning (ICL) lets large language models adapt to new tasks from demonstrations, and fine-tuning can erode this behaviour. Many preservation diagnostics inspect attention: if attention changes when demonstrations change, the model is treated as context-sensitive. This paper asks how far that proxy can be trusted once it is optimised. We formalise In-Context Sensitivity (ICS), the average row distance between last-token attention on matched and mismatched demonstration prefixes, and pair it with ICL-GAP, the behavioural accuracy gap between the same prefixes. In a controlled four-arm ablation on Llama-2-7B, an ICS-maximising regulariser (F_ICS-Max) drives ICS to 1.4131.413, within 0.5%0.5\% of its geometric ceiling. The behavioural readout tells a different story: ICL-GAP stays near zero and MMLU accuracy moves from 0.3710.371 to 0.2790.279, a Goodhart dissociation of the bounded attention proxy. Endpoint statistics locate the mechanism: attention grows sharp and near-disjoint across prefixes yet routes to formatting and demonstration-body tokens rather than labels. A random-label protocol confirms that the behavioural probe family retains dynamic range at the same checkpoints. In a constructive sweep, behaviour gating partially mitigates the effect, while objectives anchored to pretrained computation hold the high-MMLU, moderate-ICS region that divergence maximisers leave. The main lesson is diagnostic: attention-level ICL proxies earn their place as training targets only after validation against behavioural gaps.

1 Introduction

Large language models can solve new tasks from demonstrations in their input, a capability known as in-context learning (ICL) (Brown et al., 2020; Wei et al., 2022). ICL requires no parameter updates, but subsequent fine-tuning can alter the behaviour on which this adaptation relies. Continued pretraining and instruction tuning have been observed to weaken ICL on held-out tasks (Luo et al., 2023; Shi et al., 2024; Wang et al., 2023b), motivating methods that measure and limit this drift.

Many diagnostics inspect attention because ICL has been associated with mechanisms such as induction heads and copying circuits (Olsson et al., 2022; Akyürek et al., 2023). A common approach compares attention maps under a matched prefix containing correct demonstrations and a mismatched prefix containing shuffled or random demonstrations. The resulting divergence measures whether attention responds to the demonstrations. It is inexpensive and differentiable, which makes it suitable for monitoring and, potentially, regularisation—though on its own it says nothing about whether that response improves task behaviour.

Refer to caption
Figure 1: Under the ICS-maximising stress test, the attention proxy approaches its ceiling while ICL-GAP stays near zero and MMLU declines. This divergence motivates testing the proxy against behavioural ICL under optimisation.

We test whether an attention-level ICL proxy remains faithful when it becomes an optimisation target. Attention maps are already used in distillation, alignment, and faithfulness analysis (Jiao et al., 2020; Tropeano et al., 2026; Yao et al., 2026), making differentiable attention probes plausible auxiliary preservation signals. Without attributing the exact ICS objective to prior work, we use a matched-vs.-mismatched attention diagnostic to test the broader assumption that increasing context-sensitive attention preserves behavioural ICL.

Setup and central finding.

We formalise the proxy as In-Context Sensitivity (ICS): the mean distance, over last-token attention rows in a band of mid-network layers, between a matched and a mismatched demonstration prefix. Because each attention row is a probability distribution, ICS has a fixed geometric ceiling, attained only when the two prefixes induce attention concentrated on disjoint single tokens. Its behavioural counterpart is ICL-GAP: the held-out task accuracy under the matched prefix minus the held-out accuracy under the mismatched prefix. ICL-GAP is the behavioural quantity any ICL-preservation method actually wants to keep above zero.

On Llama-2-7B we run a four-arm controlled ablation, with all arms starting from the same checkpoint and training for 5,000 steps on a fixed instruction-tuning mixture. The first three arms freeze attention, freeze MLPs, or fine-tune all parameters without an auxiliary loss; the fourth, F_ICS-Max, adds an ICS-maximising term to the cross-entropy loss. The unregularised arms remain close to the pretrained ICS of 0.516. F_ICS-Max reaches 1.413, within 0.5 percent of the theoretical ceiling of about 1.414, while ICL-GAP on a held-out QA probe remains near zero and MMLU accuracy falls from 0.371 to 0.279.

Main empirical finding.

Considered alone, the 2.74×2.74\times increase in attention divergence could be read as stronger context sensitivity. The behavioural measurement points the other way, and the gap between the two readings is exactly why attention-level ICL diagnostics need validation against a behavioural gap before they are used under optimisation pressure.

Contributions.

We make four contributions. First, we validate a matched-vs.-mismatched behavioural probe and define ICS as a geometrically bounded attention diagnostic paired with behavioural ICL-GAP. Second, a controlled Llama-2-7B ablation shows that explicit ICS maximisation saturates the proxy while leaving ICL-GAP near zero and degrading MMLU. Third, attention endpoint statistics and a structural construction show how sharp, disjoint routing can maximise ICS without selecting answer-relevant tokens. Finally, we test behaviour-gated B-ICS and pretrained anchoring as constructive responses to this dissociation. The supplement gives the contamination postmortem, full trajectories, layer ablations, and constructive sweeps.

2 Related Work

What ICL is, and where it lives.

In-context learning emerged as an empirical property of large transformers (Brown et al., 2020; Wei et al., 2022), and mechanistic work has since tied it to specific circuitry: induction heads that copy an earlier pattern’s continuation forward (Olsson et al., 2022), and attention layers that implement gradient-descent-like or Bayesian updates over the prefix (Akyürek et al., 2023; von Oswald et al., 2023; Xie et al., 2022). Recent probes cut the other way: in tabular and structured-data settings, apparent ICL can reduce to memorisation or label priors (Capano and Böhler, 2026; Pelusi et al., 2026). These accounts disagree about what ICL is, but share one implication: intact ICL should respond to demonstrations in attention. We do not contest that direction; we put the converse under optimisation pressure—that context-sensitive attention implies intact ICL—once the response becomes a training signal.

ICL under fine-tuning.

Continued training moves this behaviour, usually tracked at the output. Instruction tuning, supervised fine-tuning, and continual pretraining can all reduce held-out ICL accuracy (Luo et al., 2023; Shi et al., 2024; Wang et al., 2023b), and finer-grained work follows how fine-tuning reshapes in-context factual recall (Huang et al., 2026). The measurement is behavioural throughout: ICL accuracy, or accuracy gaps between matched and mismatched prefixes. Our ICL-GAP belongs to that family; we depart by instrumenting an attention-level diagnostic, ICS, alongside it and studying the relationship between the two under training.

Attention as evidence.

Attention maps are tempting evidence: cheap, differentiable, already load-bearing. Distillation transfers them between models (Jiao et al., 2020); pruning attention layers changes explanation faithfulness and confidence (Tropeano et al., 2026); alignment work optimises attention dynamics as a preference signal (Yao et al., 2026). Yet attention weights need not track the computation behind a prediction. Our stress test sharpens that caution: not whether attention explains a fixed model, but whether an attention proxy stays meaningful while being optimised.

Component-level attribution.

A parallel line localises capabilities to subnetworks. Mechanistic interpretability isolates circuits at the level of heads and MLP neurons (Elhage et al., 2021; Wang et al., 2023a); parameter-efficient tuning trains only attention or only MLP blocks (Houlsby et al., 2019; He et al., 2022). Our F_A and F_M arms repurpose the second design for a stability question: which subnetwork’s update carries ICL drift? In our setup neither does alone—both preserve ICS near its pretrained value—motivating an explicit proxy regulariser instead of a structural constraint.

Auxiliary objectives under Goodhart pressure.

Regularisers that protect old behaviour assume the auxiliary signal tracks the capability. EWC (Kirkpatrick et al., 2017), synaptic intelligence (Zenke et al., 2017), and KL-to-base penalties (Ouyang et al., 2022) anchor to different quantities, but each substitutes a measurable surrogate for the behaviour itself. Goodhart’s law (Goodhart, 1984; Manheim and Garrabrant, 2018) names the failure mode; the RLHF literature gives modern instances, from reward-model over-optimisation (Skalse et al., 2022; Gao et al., 2023) to reward hacking at inference time and beyond (Khalaf et al., 2025; Fu et al., 2026; Liu et al., 2026). Those warnings concern reward functions; we give the same failure an attention-level instance, driving ICS within 0.5%0.5\% of its geometric ceiling while behavioural ICL stays near zero.

Calibration as a side signal.

Confidence can part ways with behaviour as well. Fine-tuning shifts calibration (Guo et al., 2017; Desai and Durrett, 2020); we record ECE (Naeini et al., 2015) on MMLU (Hendrycks et al., 2021) as a diagnostic only, since lower ECE under near-chance predictions says little about reasoning quality.

Where this work sits.

ICL-preservation methods fall into three families: freeze a parameter subset, replay ICL-format data, or regularise toward base-model outputs. Our F_A, F_M, and F_None arms cover the first family in controlled form and our logit-anchoring baseline the third; replay is discussed but not completed in the main comparison. AnchorTune is closest to attention distillation (Jiao et al., 2020; Gu et al., 2023), but anchors the trainee to its own pretrained checkpoint θ0\theta_{0} on last-token attention rows over a fixed ICL probe. We present it as one member of an anchored family whose anchor matches the diagnostic under test, not as a uniquely superior method; B-ICS adds the complement, gating the proxy by behaviour rather than tying it to a checkpoint.

3 Methodology

This section defines the diagnostic framework used in the rest of the paper. The goal is deliberately narrow: separate an attention-level sign of context sensitivity from the behavioural effect that an ICL-preservation method is meant to preserve. We first define the matched-vs.-mismatched problem setting and its diagnostic metrics, then introduce the proxy-maximisation stress test and two constructive variants. A separate random-label protocol validates the behavioural measurement before the main stress-test results.

3.1 Problem Setting

Let ℳθ\mathcal{M}_{\theta} be a pretrained autoregressive transformer. For a labelled task (x,y)∼p⁡(x,y)(x,y)\sim p(x,y) with label set 𝒴\mathcal{Y}, draw k=4k=4 demonstrations S={(xi,yi)}i=1kS=\{(x_{i},y_{i})\}_{i=1}^{k} and a random label permutation π\pi. For the same query xx, we compare a matched prefix 𝒟1​(x,S)\mathcal{D}_{1}(x;S) with a mismatched prefix 𝒟2​(x,S,π)\mathcal{D}_{2}(x;S,\pi) that preserves the token template but breaks the input–label correspondence. A model with behaviourally intact ICL should satisfy

acc𝒟1⁡(θ)>acc𝒟2⁡(θ).\acc_{\mathcal{D}_{1}}(\theta)>\acc_{\mathcal{D}_{2}}(\theta). (1)

This matched-vs.-mismatched contrast is the common protocol behind both the attention diagnostic and the behavioural probe.

3.2 Diagnostic Metrics

For a prefix D∈{𝒟1,𝒟2}\mathrm{D}\in\{\mathcal{D}_{1},\mathcal{D}_{2}\}, layer ll, and head hh, write the last-token attention row as

al,h​(D,θ)≜softmax⁡(ql,h​(D)​Kl,h​(D)⊤dh+m).a_{l,h}(\mathrm{D};\theta)\triangleq\softmax\!\left(\frac{q_{l,h}(\mathrm{D})K_{l,h}(\mathrm{D})^{\top}}{\sqrt{d_{h}}}+m\right). (2)

We define In-Context Sensitivity (ICS) by averaging the distance between the matched and mismatched rows over a probe distribution μ\mu and a layer band ℒ\mathcal{L}:

ICS⁡(θ)≜𝔼μ​[1|ℒ|​H​∑l∈ℒ∑h=1Hdl,h​(x,S,π,θ)],\mathrm{ICS}(\theta)\triangleq\mathbb{E}_{\mu}\!\left[\frac{1}{|\mathcal{L}|H}\sum_{l\in\mathcal{L}}\sum_{h=1}^{H}d_{l,h}(x,S,\pi;\theta)\right], (3)

where

dl,h​(x,S,π,θ)≜‖al,h​(𝒟1,θ)−al,h​(𝒟2,θ)‖2.d_{l,h}(x,S,\pi;\theta)\triangleq\bigl\|a_{l,h}(\mathcal{D}_{1};\theta)-a_{l,h}(\mathcal{D}_{2};\theta)\bigr\|_{2}. (4)

For Llama-2-7B we use ℒ={8,…,16}\mathcal{L}=\{8,\dots,16\}; the supplement ablates this choice.

Proposition 1 (Geometric ceiling).

For any θ\theta, ICS⁡(θ)∈[0,2]\mathrm{ICS}(\theta)\in[0,\sqrt{2}]. For any attention rows p,qp,q,

‖p−q‖22=‖p‖22+‖q‖22−2​⟨p,q⟩≤2,\|p-q\|_{2}^{2}=\|p\|_{2}^{2}+\|q\|_{2}^{2}-2\langle p,\,q\rangle\leq 2, (5)

with equality iff pp and qq are one-hot on disjoint coordinates. Thus ICS=2\mathrm{ICS}=\sqrt{2} requires every probed matched/mismatched attention-row pair to be one-hot and disjoint almost surely.

ICS measures whether attention reorganises when demonstrations change. The behavioural target is the accuracy gap induced by the same intervention. Let y^θ​(D)=arg​maxc∈𝒴⁡ℙθ​(c∣D)\hat{y}_{\theta}(\mathrm{D})=\argmax_{c\in\mathcal{Y}}\mathbb{P}_{\theta}(c\mid\mathrm{D}). We define

accD⁡(θ)\displaystyle\acc_{\mathrm{D}}(\theta) ≜𝔼μ[𝟏{y^θ(D)=y}],\displaystyle\triangleq\mathbb{E}_{\mu}\!\left[\mathbf{1}\{\hat{y}_{\theta}(\mathrm{D})=y\}\right], (6)
ICL​-​GAP​(θ)\displaystyle\mathrm{ICL\text{-}GAP}(\theta) ≜acc𝒟1⁡(θ)−acc𝒟2⁡(θ).\displaystyle\triangleq\acc_{\mathcal{D}_{1}}(\theta)-\acc_{\mathcal{D}_{2}}(\theta). (7)

The central question is whether high ICS implies positive ICL-GAP under optimisation pressure. We also report MMLU accuracy and Expected Calibration Error (ECE):

ECE≜∑b=1B|Bb|N​|acc⁡(Bb)−conf⁡(Bb)|.\mathrm{ECE}\triangleq\sum_{b=1}^{B}\frac{|B_{b}|}{N}\,\bigl|\acc(B_{b})-\mathrm{conf}(B_{b})\bigr|. (8)

3.3 Proxy-Maximisation Stress Test

Refer to caption
Figure 2: Diagnostic framework. Matched and mismatched prompts feed the same fine-tuning pipeline, while the attention-level ICS and behavioural ICL-GAP probes are computed on a held-out pair set disjoint from any regulariser pairs. The architecture also shows the two constructive directions studied later: behaviour-gated diagnostics and anchoring to the pretrained computation.

To test the proxy rather than merely observe it, we add an ICS-maximising auxiliary objective to the ordinary cross-entropy loss. Let

𝒮reg={(x(n),S(n),π(n))}n=1Nreg\mathcal{S}_{\mathrm{reg}}=\{(x^{(n)},S^{(n)},\pi^{(n)})\}_{n=1}^{N_{\mathrm{reg}}} (9)

be a held-out regulariser set, and let ICS^𝒮reg\widehat{\mathrm{ICS}}_{\mathcal{S}_{\mathrm{reg}}} be the empirical estimate of Equation 3 on that set. The stress-test arm minimises

ℒtotal​(θ)\displaystyle\mathcal{L}_{\mathrm{total}}(\theta) ≜ℒCE​(θ)+λ​ℒfunc​-​icl​(θ),\displaystyle\triangleq\mathcal{L}_{\mathrm{CE}}(\theta)+\lambda\,\mathcal{L}_{\mathrm{func\text{-}icl}}(\theta), (10)
ℒfunc​-​icl​(θ)\displaystyle\mathcal{L}_{\mathrm{func\text{-}icl}}(\theta) ≜−ICS^𝒮reg​(θ),\displaystyle\triangleq-\,\widehat{\mathrm{ICS}}_{\mathcal{S}_{\mathrm{reg}}}(\theta), (11)

with λ=0.05\lambda=0.05. If ICS is a faithful proxy for ICL, this arm should improve the matched-vs.-mismatched behavioural gap; any proxy gain without a matching behavioural gain then measures the proxy’s Goodhart exposure directly.

3.4 Constructive Variants

We evaluate two compact variants aimed at mitigating this dissociation. B-ICS gates the attention proxy by a differentiable behavioural sentinel:

B​-​ICSβ​(θ,δ)≜ICS⁡(θ)​σ​(β⁡(ICL​-​GAP~​(θ)−δ)),\mathrm{B\text{-}ICS}_{\!\beta}(\theta;\delta)\triangleq\mathrm{ICS}(\theta)\,\sigma\!\bigl(\beta(\widetilde{\mathrm{ICL\text{-}GAP}}(\theta)-\delta)\bigr), (12)

where

ICL​-​GAP~​(θ)≜𝔼μ​[log⁡ℙθ​(y∣𝒟1)−log⁡ℙθ​(y∣𝒟2)].\widetilde{\mathrm{ICL\text{-}GAP}}(\theta)\triangleq\mathbb{E}_{\mu}\!\left[\log\mathbb{P}_{\theta}(y\mid\mathcal{D}_{1})-\log\mathbb{P}_{\theta}(y\mid\mathcal{D}_{2})\right]. (13)

The soft gap serves as a differentiable surrogate for the discrete ICL-GAP, without any claim that log-probability gaps and accuracy gaps agree pointwise.

AnchorTune instead anchors the trainee’s matched-prefix attention rows to the pretrained model θ0\theta_{0} on 𝒮anchor\mathcal{S}_{\mathrm{anchor}}:

ℒanchor≜𝔼𝒮anchor,l,h​[‖al,h​(𝒟1,θ)−al,h​(𝒟1,θ0)‖2],\mathcal{L}_{\mathrm{anchor}}\triangleq\mathbb{E}_{\mathcal{S}_{\mathrm{anchor}},l,h}\!\left[\bigl\|a_{l,h}(\mathcal{D}_{1};\theta)-a_{l,h}(\mathcal{D}_{1};\theta_{0})\bigr\|_{2}\right], (14)

with objective

ℒtotal​(θ)≜ℒCE​(θ)+λ​ℒanchor​(θ,θ0,𝒮anchor).\mathcal{L}_{\mathrm{total}}(\theta)\triangleq\mathcal{L}_{\mathrm{CE}}(\theta)+\lambda\,\mathcal{L}_{\mathrm{anchor}}(\theta;\theta_{0},\mathcal{S}_{\mathrm{anchor}}). (15)

Unlike a divergence-maximising loss, this objective is minimised at the pretrained attention pattern rather than on the disjoint-support ridge of Proposition 1.

4 Experiments and Analysis

We organise the empirical study around four research questions. They progress from validating the behavioural measurement, through establishing and explaining the proxy–behaviour dissociation, to testing constructive responses:

  • •

    RQ0. Does the pretrained model exhibit behavioural ICL under a controlled matched-vs.-mismatched protocol?

  • •

    RQ1. Can an attention-level ICL proxy be driven to its geometric ceiling while behavioural ICL remains absent?

  • •

    RQ2. What changes inside the model when the proxy is maximised, and why does this fail to improve behaviour?

  • •

    RQ3. Can behavioural gating or pretrained anchoring reduce this proxy Goodhart failure?

After a shared experimental setup, the following modules answer these questions in order. Each module presents question-specific evidence and closes with a direct answer.

4.1 Shared Training and Evaluation Setup

All main arms start from the same Llama-2-7B pretrained checkpoint (Touvron et al., 2023) and run for 5,0005{,}000 optimiser steps on the same instruction-tuning mixture. The arms share the optimiser, schedule, sequence length, batch size, precision, and evaluation cadence; exact training details are in the supplement. Probes are evaluated every 100100 steps, and the supplement reports full trajectories.

The controlled ablation compares four arms. F_A freezes attention projections and updates the MLP, layer norms, and language-model head. F_M freezes the MLP blocks and updates attention projections, layer norms, and the head. F_None full-finetunes all parameters with cross-entropy only. F_ICS-Max full-finetunes all parameters with the ICS-maximising stress-test loss in Equation 11.

To avoid probe contamination, the regulariser and evaluation pair sets are disjoint:

𝒮reg∩𝒮eval=∅.\mathcal{S}_{\mathrm{reg}}\cap\mathcal{S}_{\mathrm{eval}}=\emptyset. (16)

The regulariser set has 500500 pairs sampled with seed 9999; the evaluation set has 500500 pairs sampled with seed 4242. ICS and ICL-GAP are always reported on 𝒮eval\mathcal{S}_{\mathrm{eval}}. MMLU uses a fixed 240240-item subset stratified over 5757 subjects, evaluated under a 55-shot prefix at temperature 00. ECE uses the MMLU predictions with 1515 equal-mass bins. GSM8K exact-match, which sits at zero for the base model across arms, is excluded from the main evidence.

4.2 Does the Behavioural Probe Detect ICL? (RQ0)

A proxy-stress test is only meaningful if the base model can use demonstrations in at least one held-out protocol. We therefore screened pretrained Llama-2-7B on 480480 random-label binary classification episodes. Each episode samples a semantic split, assigns the two classes arbitrary labels, and compares three prompts:

D∈{𝒟1,𝒟2,D∅}.\mathrm{D}\in\{\mathcal{D}_{1},\,\mathcal{D}_{2},\,\mathrm{D}_{\emptyset}\}. (17)

The model scores only the two candidate label tokens. Pretrained accuracy is 0.6690.669 with matched demonstrations, 0.3630.363 with swapped demonstrations, and 0.4920.492 without demonstrations. The paired bootstrap estimate is

acc𝒟1−acc𝒟2\displaystyle\acc_{\mathcal{D}_{1}}-\acc_{\mathcal{D}_{2}} =0.306,\displaystyle=0.306, (18)
95%​CI\displaystyle 95\%~\mathrm{CI} =[0.231,0.379].\displaystyle=[0.231,0.379]. (19)

Thus the checkpoint has a clear behavioural ICL effect on this sanity-check protocol. We keep this result separate from the controlled QA probe, which is harder and has a near-zero behavioural gap across arms.

To check that this behavioural measurement family still has dynamic range after fine-tuning, we reran the same RQ0 protocol on the step-5,0005{,}000 checkpoints of all four arms. Table 1 shows that every arm retains a positive matched-vs.-permuted gap with a bootstrap confidence interval excluding zero. This does not rescue the controlled QA gap; rather, it rules out the simpler objection that all behavioural probes are insensitive at these checkpoints. In particular, F_ICS-Max reaches the largest RQ0 gap while leaving the controlled QA ICL-GAP near zero.

Table 1: RQ0 random-label A/B protocol at step 5,0005{,}000. All arms retain a clear matched-vs.-permuted behavioural gap on this protocol, so the near-zero controlled QA ICL-GAP cannot be attributed to a universally insensitive behavioural probe family.
Arm Match Perm. None Gap [95% CI]
F_A 0.663 0.360 0.492 0.302 [0.235, 0.371]
F_M 0.656 0.381 0.492 0.275 [0.206, 0.346]
F_None 0.623 0.435 0.492 0.188 [0.115, 0.260]
F_ICS-Max 0.692 0.340 0.492 0.352 [0.292, 0.413]

Answer to RQ0.

Yes. The pretrained checkpoint shows a clear matched-vs.-mismatched behavioural effect, and the same probe family retains dynamic range after fine-tuning. This validation does not imply that the harder controlled QA probe must be positive; it establishes that its near-zero gap is not caused by universal probe insensitivity.

4.3 Can ICS Saturate without Behavioural ICL? (RQ1)

Before the controlled ablation, we ran a coarse full-fine-tuning baseline under an earlier data mixture and a higher learning rate. Starting from ICS=0.516\mathrm{ICS}=0.516, that run collapsed to ICS≈0.20\mathrm{ICS}\approx 0.20 at step 5,0005{,}000, a relative reduction of 61%61\%. It serves only as motivation—an attention-proxy analogue of ICL fragility (Luo et al., 2023) rather than a behavioural-collapse claim.

Table 2 reports the step-5,0005{,}000 values of ICS, ICL-GAP, MMLU accuracy, and ECE for the controlled ablation, with the pretrained baseline in the first row. ICL-GAP was logged during training for F_ICS-Max and later evaluated post-hoc for the unregularised arms.

Table 2: Final controlled-ablation results at step 5,0005{,}000. ICS is the attention-level diagnostic from Equation 3, bounded above by 2≈1.414\sqrt{2}\approx 1.414. ICL-GAP is the behavioural quantity from Equation 7. MMLU accuracy is on a 240240-item probe (chance 0.250.25). The pretrained baseline ICS is 0.5160.516.
Arm ICS (↑\uparrow) ICL-GAP (↑\uparrow) MMLU acc. (↑\uparrow) ECE (↓\downarrow)
Pretrained Llama-2-7B 0.516 — — —
F_A 0.492 — 0.338 0.434
F_M 0.482 — 0.342 0.406
F_None 0.500 — 0.375 0.407
F_ICS-Max 1.413 −0.010-0.010 0.279 0.231
Refer to caption
Figure 3: Step-5,0005{,}000 outcome in the (ICS, MMLU) plane. F_ICS-Max saturates the attention proxy near the 2\sqrt{2} ceiling but moves into the low-MMLU region; anchored variants remain near the pretrained ICS while preserving MMLU.

F_A, F_M, and F_None all finish within 0.040.04 of the pretrained baseline ICS of 0.5160.516; the spread across these three arms is itself only 0.0180.018. Under the controlled stress-test mixture and hyperparameters, neither restricting updates to MLP layers (F_A) nor to attention layers (F_M) nor leaving them unconstrained (F_None) substantially alters attention-level context responsiveness. We expected F_M, which can update attention, to drift further from baseline than F_A, but the observed gap (0.4820.482 vs. 0.4920.492) is at the noise level of the three-arm spread and supports no ordering between the two. Nor does F_None reproduce the coarse-pilot collapse, finishing at ICS=0.500\mathrm{ICS}=0.500. Since that pilot used different data and hyperparameters, its endpoint is not quantitatively comparable to the controlled experiment, and we leave a controlled unregularised-collapse condition to future work.

Because the MMLU probe has only 240240 items, we avoid interpreting small differences among the unregularised arms: at p≈0.37p\approx 0.37 the binomial standard error is about 0.0310.031, which puts differences of order 0.030.03 within sampling noise at the current run count. The large F_ICS-Max drop from 0.3710.371 to 0.2790.279 is the general-ability signal we use.

F_ICS-Max moves ICS from 0.5130.513 at step 100100, which is statistically indistinguishable from the pretrained baseline of 0.5160.516, to 1.4131.413 at step 5,0005{,}000. The convergence value sits within 0.5%0.5\% of the geometric ceiling 2≈1.4142\sqrt{2}\approx 1.4142 from Proposition 1: the attention divergence between matched and mismatched demonstrations has been amplified by a factor of 1.413/0.516≈2.74×1.413/0.516\approx 2.74\times relative to the pretrained model. Read in isolation, this looks like a successful ICL-preservation regulariser.

The behavioural picture is the opposite. ICL-GAP starts at −0.020-0.020 at step 100100 (acc𝒟1=0.295\acc_{\mathcal{D}_{1}}=0.295 vs. acc𝒟2=0.315\acc_{\mathcal{D}_{2}}=0.315), briefly becomes positive in the middle of training (peak 0.0500.050 at step 1,5001{,}500), and ends at −0.010-0.010 at step 5,0005{,}000. Across the 5050 logged probe points, ICL-GAP has sample mean g¯=0.005\bar{g}=0.005 and sample standard deviation sg=0.020s_{g}=0.020. The trajectory evaluation used 200200 paired examples per checkpoint, although the fixed evaluation set contains 500500 pairs; at per-condition accuracy p≈0.30p\approx 0.30, the unpaired binomial-difference scale is

2​p​(1−p)200≈0.046.\sqrt{\frac{2p(1-p)}{200}}\approx 0.046. (20)

This is a conservative scale for individual logged checkpoints because the actual matched/mismatched evaluations are paired. The trajectory mean g¯\bar{g} is about one ninth of this scale and within 0.25​sg0.25s_{g} of zero: the model has not learned to use correct demonstrations more than incorrect ones at any point, despite the attention divergence multiplying.

Answer to RQ1.

Yes. Optimisation drives ICS to 1.4131.413, essentially its geometric ceiling, while the controlled QA ICL-GAP remains statistically and practically near zero and MMLU degrades. Proxy saturation and behavioural preservation therefore come apart under optimisation pressure.

4.4 How Does Proxy Maximisation Come Apart from Behaviour? (RQ2)

The full trajectory shows a monotone rise of ICS toward the 2\sqrt{2} ceiling, an approximately monotone MMLU drop once ICS passes ≈1.3\approx 1.3, and ICL-GAP fluctuating around zero throughout. MMLU accuracy under F_ICS-Max falls from 0.3710.371 at step 100100 to 0.2790.279 at step 5,0005{,}000, a drop well beyond the sampling scale above and within 0.030.03 of the random-chance baseline of 0.250.25. ECE is lower for F_ICS-Max (0.2310.231) than for the unregularised arms (0.410.41–0.430.43), but this reflects the model growing more uncertain as accuracy approaches chance rather than improved calibration.

Refer to caption
Figure 4: Per-step trajectories of ICS, ICL-GAP, MMLU accuracy, and MMLU ECE across 5,0005{,}000 fine-tuning steps. The trajectory view makes the dissociation temporal: F_ICS-Max drives ICS upward while ICL-GAP remains near zero and MMLU falls.

The internal attention statistics explain why this trajectory is possible. At the pretrained baseline, the top-11 attention key under 𝒟1\mathcal{D}_{1} lands on a label token 58%58\% of the time. At F_ICS-Max step 5,0005{,}000, the same rate falls to 26%26\%, while punctuation or formatting tokens account for 31%31\% and content tokens from the demonstration body for 41%41\%. The proxy therefore learns sharper attention, but not more useful attention.

Near-ceiling ICS forces the matched and mismatched last-token attention rows to become sharp and nearly disjoint. In held-out probe rows, mean top-11 concentration rises from 0.180.18 at the pretrained baseline to 0.940.94 at step 5,0005{,}000, while matched/mismatched top-11 supports overlap on only 4%4\% of head-example pairs, down from 61%61\%. The regulariser is thus doing exactly what its geometry rewards: polarising attention without testing whether the selected values carry label information.

Proposition 2 (Goodhart channel: vacuous maximisers exist).

Assume the model can make every probe head one-hot on disjoint coordinates as in Proposition 1, and that there are prefix tokens whose values are conditionally independent of yy given the prompt format. Then some θ⋆\theta^{\star} maximises ICS while satisfying ICL​-​GAP​(θ⋆)≤0\mathrm{ICL\text{-}GAP}(\theta^{\star})\leq 0, and ℒfunc​-​icl=−ICS\mathcal{L}_{\mathrm{func\text{-}icl}}=-\mathrm{ICS} cannot distinguish this behaviourally vacuous maximiser from an ICL-useful one with the same attention-row divergence.

Sketch.

Route each probe head to two disjoint prompt-format tokens whose values are independent of yy. This attains ICS=2\mathrm{ICS}=\sqrt{2}, but the readout carries no matched-prefix answer information, so matched and mismatched accuracies are equal in expectation. Since ℒfunc​-​icl\mathcal{L}_{\mathrm{func\text{-}icl}} only sees row distance, useful and vacuous disjoint-support solutions receive the same proxy score. ∎

Proposition 2 also explains the apparent calibration anomaly noted above: near-chance predictions are close to uniform, which lowers ECE without any improvement in calibrated reasoning.

Answer to RQ2.

Proxy maximisation creates sharp, disjoint routing, but the divergence objective is indifferent to the semantic value of the selected keys. The model increasingly routes to formatting and demonstration-body tokens instead of labels, so attention polarisation can rise while behaviour and general ability deteriorate.

4.5 Can Gating or Anchoring Mitigate the Dissociation? (RQ3)

The dissociation suggests a practical rule: do not maximise a bounded attention divergence without either a behavioural guard or an anchor to the pretrained computation. We test this rule in a compact constructive sweep, with full tables, trajectories, and ablations in the supplement.

Smooth B-ICS uses Equation 12 as the auxiliary target. Under our mild gate (β,δ)=(10,0.05)(\beta,\delta)=(10,0.05), it partially mitigates the Goodhart channel: F_BICS reaches ICS=1.397\mathrm{ICS}=1.397, MMLU 0.2830.283, and ICL-GAP +0.020+0.020, compared with F_ICS-Max’s ICS=1.413\mathrm{ICS}=1.413, MMLU 0.2790.279, and ICL-GAP −0.010-0.010. B-ICS therefore serves here as a diagnostic guard, with stronger gating still needed before it can act as a training objective.

Anchored objectives behave differently. AnchorTune, KL-to-logits anchoring, and weight ℓ2\ell_{2} anchoring all keep ICS near 0.50.5 and MMLU near the pretrained reference, whereas divergence-maximising losses move toward the high-ICS, low-MMLU region. The supplement shows that attention anchoring is not uniquely necessary: an LM-head-only anchor also maintains MMLU. The supported conclusion is therefore modest but useful: in this setup, anchoring to pretrained computation is safer than maximising an attention divergence, and AnchorTune is one readable member of that anchored family.

Answer to RQ3.

Partially, and more so for anchoring than for gating. Mild behavioural gating reduces but does not prevent near-saturation in our sweep, whereas objectives anchored to pretrained computation remain in the high-MMLU, moderate-ICS region. Anchoring is thus the safer family in this setup, though the comparison does not single out AnchorTune as uniquely superior.

5 Discussion and Limitations

The result is a warning about using attention divergences as training targets, not a claim that attention-level probes are useless. Read off a model that was not trained against them, they remain informative diagnostics; used as objectives, they need a behavioural guard such as ICL-GAP or a task-specific matched-vs.-mismatched accuracy gap.

Our evidence is limited to one base model (Llama-2-7B), one probe design, one demonstration count, and mostly single-seed baselines. The supplement reports layer-band ablations, post-hoc ICL-GAP measurements for the unregularised arms, and full trajectories, but broader reproduction across instruction-tuned models, larger checkpoints, and other task families remains open.

6 Conclusion

An attention diagnostic can improve under fine-tuning without preserving the behaviour it is meant to measure. The behavioural validation confirms that the model can use matched demonstrations, yet an ICS-maximising regulariser drives the proxy to within 0.5%0.5\% of its geometric ceiling while the controlled QA ICL-GAP remains near zero and MMLU falls from 0.3710.371 to 0.2790.279. Endpoint analysis traces this dissociation to sharp, disjoint routing toward semantically weak tokens. Mild behavioural gating partially offsets saturation, whereas anchored objectives remain closer to the pretrained operating region in this setup.

The practical lesson is not that attention probes should be discarded. Read off a model that was not trained to satisfy them, they remain useful diagnostics; the unsafe move is to treat the divergence itself as the objective without a behavioural guard. Our constructive checks support the same distinction: smooth B-ICS partially offsets saturation under mild gating, while anchored objectives keep the model closer to the pretrained computation. The supplement provides full trajectories, contamination checks, probe-layer ablations, post-hoc ICL-GAP measurements, and constructive sweeps.

References

  • Akyürek et al. (2023) E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou What learning algorithm is in-context learning? investigations with linear models. In International Conference on Learning Representations (ICLR), Cited by: §1, §2.
  • Brown et al. (2020) T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §1, §2.
  • Capano and Böhler (2026) F. Capano and J. Böhler Probing memorization of tabular in-context learning. arXiv preprint arXiv:2606.31208. External Links: Link Cited by: §2.
  • Desai and Durrett (2020) S. Desai and G. Durrett Calibration of pre-trained transformers. In Conference on Empirical Methods in Natural Language Processing (EMNLP), Cited by: §2.
  • Elhage et al. (2021) N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah A mathematical framework for transformer circuits. Transformer Circuits Thread. External Links: Link Cited by: §2.
  • Fu et al. (2026) J. Fu, X. Zhao, C. Yao, H. Wang, Q. Han, and Y. Xiao Reward shaping to mitigate reward hacking in RLHF. arXiv preprint arXiv:2502.18770. External Links: Link Cited by: §2.
  • Gao et al. (2023) L. Gao, J. Schulman, and J. Hilton Scaling laws for reward model overoptimization. In International Conference on Machine Learning (ICML), Cited by: §2.
  • Goodhart (1984) C. A. E. Goodhart Problems of monetary management: the UK experience. Monetary Theory and Practice, pp. 91–121. Cited by: §2.
  • Gu et al. (2023) Y. Gu, L. Dong, F. Wei, and M. Huang MiniLLM: knowledge distillation of large language models. arXiv preprint arXiv:2306.08543. Cited by: §2.
  • Guo et al. (2017) C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger On calibration of modern neural networks. In International Conference on Machine Learning (ICML), Cited by: §2.
  • He et al. (2022) J. He, C. Zhou, X. Ma, T. Berg-Kirkpatrick, and G. Neubig Towards a unified view of parameter-efficient transfer learning. In International Conference on Learning Representations (ICLR), Cited by: §2.
  • Hendrycks et al. (2021) D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt Measuring massive multitask language understanding. In International Conference on Learning Representations (ICLR), Cited by: §2.
  • Houlsby et al. (2019) N. Houlsby, A. Giurgiu, S. Jastrzębski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning (ICML), Cited by: §2.
  • Huang et al. (2026) R. Huang, E. Nichani, J. D. Lee, and R. Ge Fine-tuning dynamics of in-context factual recall in transformers. arXiv preprint arXiv:2605.27774. External Links: Link Cited by: §2.
  • Jiao et al. (2020) X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu TinyBERT: distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP, Cited by: §1, §2, §2.
  • Khalaf et al. (2025) H. Khalaf, C. M. Verdun, A. Oesterling, H. Lakkaraju, and F. du Pin Calmon Inference-time reward hacking in large language models. arXiv preprint arXiv:2506.19248. External Links: Link Cited by: §2.
  • Kirkpatrick et al. (2017) J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences (PNAS) 114 (13), pp. 3521–3526. Cited by: 2nd item, §2.
  • Liu et al. (2026) W. Liu, X. Mou, H. Yan, Z. Wei, and Y. He Large language models hack rewards, and society. arXiv preprint arXiv:2606.04075. External Links: Link Cited by: §2.
  • Loshchilov and Hutter (2019) I. Loshchilov and F. Hutter Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), Cited by: Appendix F.
  • Luo et al. (2023) Y. Luo, Z. Yang, F. Meng, Y. Li, J. Zhou, and Y. Zhang An empirical study of catastrophic forgetting in large language models during continual fine-tuning. arXiv preprint arXiv:2308.08747. Cited by: §1, §2, §4.3.
  • Manheim and Garrabrant (2018) D. Manheim and S. Garrabrant Categorizing variants of goodhart’s law. In arXiv preprint arXiv:1803.04585, Cited by: §2.
  • Naeini et al. (2015) M. P. Naeini, G. F. Cooper, and M. Hauskrecht Obtaining well calibrated probabilities using Bayesian binning. In AAAI Conference on Artificial Intelligence, Cited by: §2.
  • Olsson et al. (2022) C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, S. Johnston, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah In-context learning and induction heads. Transformer Circuits Thread. External Links: Link Cited by: §1, §2.
  • Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: 1st item, §2.
  • Pelusi et al. (2026) A. Pelusi, S. Braghin, and A. Trombetta Categorical prior lock-in: why in-context learning fails for structured data. arXiv preprint arXiv:2606.11961. External Links: Link Cited by: §2.
  • Shi et al. (2024) H. Shi, Z. Xu, H. Wang, W. Qin, W. Wang, Y. Wang, and H. Wang Continual learning of large language models: a comprehensive survey. arXiv preprint arXiv:2404.16789. Cited by: §1, §2.
  • Skalse et al. (2022) J. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger Defining and characterizing reward hacking. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §2.
  • Touvron et al. (2023) H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. Cited by: §4.1.
  • Tropeano et al. (2026) P. Tropeano, M. Maistro, T. Ruotsalo, and C. Lioma Don’t go breaking my LLM: the impact of pruning attention layers on explanation faithfulness and confidence calibration. arXiv preprint arXiv:2606.24970. External Links: Link Cited by: §1, §2.
  • von Oswald et al. (2023) J. von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov Transformers learn in-context by gradient descent. In International Conference on Machine Learning (ICML), Cited by: §2.
  • Wang et al. (2023a) K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In International Conference on Learning Representations (ICLR), Cited by: §2.
  • Wang et al. (2023b) X. Wang, W. Zhu, M. Saxon, M. Steyvers, and W. Y. Wang Large language models are latent variable models: explaining and finding good demonstrations for in-context learning. arXiv preprint arXiv:2301.11916. Cited by: §1, §2.
  • Wei et al. (2022) J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus Emergent abilities of large language models. Transactions on Machine Learning Research. Cited by: §1, §2.
  • Xie et al. (2022) S. M. Xie, A. Raghunathan, P. Liang, and T. Ma An explanation of in-context learning as implicit bayesian inference. In International Conference on Learning Representations (ICLR), Cited by: §2.
  • Yao et al. (2026) Z. Yao, Z. Fu, Z. Zheng, J. Li, Y. Tu, and Z. Mao ADAPT: attention dynamics alignment with preference tuning for faithful MLLMs. arXiv preprint arXiv:2606.31054. External Links: Link Cited by: §1, §2.
  • Zenke et al. (2017) F. Zenke, B. Poole, and S. Ganguli Continual learning through synaptic intelligence. In International Conference on Machine Learning (ICML), Cited by: §2.

Appendix A Full Per-Step Trajectories

Table 3 reports ICS, MMLU accuracy, and ECE every 500500 steps for all four arms. ICL-GAP is reported only for F_ICS-Max during training; post-hoc ICL-GAP for the other arms at step 5,0005{,}000 is in Appendix D.

Table 3: Full trajectories for the controlled proxy-stress experiment, sampled every 500500 steps.
F_A F_M F_None F_ICS-Max
Step ICS MMLU ECE ICS MMLU ECE ICS MMLU ECE ICS GAP MMLU ECE
100 0.474 0.375 0.158 0.470 0.375 0.160 0.474 0.354 0.147 0.513 −0.020-0.020 0.371 0.136
500 0.485 0.367 0.220 0.480 0.354 0.230 0.485 0.371 0.215 1.225 +0.005+0.005 0.367 0.274
1000 0.487 0.346 0.310 0.481 0.342 0.305 0.490 0.371 0.305 1.292 +0.025+0.025 0.304 0.473
1500 0.488 0.342 0.360 0.481 0.342 0.350 0.493 0.371 0.350 1.309 +0.050+0.050 0.333 0.429
2000 0.490 0.338 0.395 0.482 0.342 0.380 0.495 0.371 0.380 1.319 +0.040+0.040 0.317 0.452
2500 0.491 0.338 0.410 0.482 0.342 0.390 0.497 0.371 0.395 1.385 −0.020-0.020 0.254 0.195
3000 0.491 0.338 0.420 0.482 0.342 0.395 0.498 0.375 0.400 1.386 +0.010+0.010 0.271 0.159
3500 0.492 0.338 0.425 0.482 0.342 0.400 0.499 0.375 0.402 1.301 +0.025+0.025 0.283 0.223
4000 0.492 0.338 0.430 0.482 0.342 0.403 0.500 0.375 0.405 1.385 −0.005-0.005 0.271 0.240
4500 0.492 0.338 0.432 0.482 0.342 0.405 0.500 0.375 0.406 1.386 0.000\phantom{+}0.000 0.283 0.235
5000 0.492 0.338 0.434 0.482 0.342 0.406 0.500 0.375 0.407 1.413 −0.010-0.010 0.279 0.231

Appendix B Contamination Postmortem (F_ICS-Maxv1)

The first F_ICS-Max run, F_ICS-Maxv1, used a single ICLEvalDataLoader for both the regulariser pair set and the ICS evaluation pair set. The regulariser was ℒfunc​-​icl=−ICS\mathcal{L}_{\mathrm{func\text{-}icl}}=-\mathrm{ICS} measured on this shared set, so the gradient pushed the model to maximise the eval metric on the eval data. We caught this at training step 1,9601{,}960, when ICS exceeded 1.371.37 on the held-out probe and continued to climb. After the fix (separate loaders, seeds 9999 and 4242, no shared pairs), the rerun (F_ICS-Maxv2) reproduced the saturation pattern but at a slower rate and on genuinely held-out data. All numbers in the main body are from F_ICS-Maxv2. The bug-fix list is:

  1. 1.

    Split ICLEvalDataLoader into icl_train_loader (regulariser) and icl_eval_loader (ICS, ICL-GAP).

  2. 2.

    Add answer_idx to ICLEvalDataLoader.__next__, which had been silently dropped and prevented ICL-GAP from being computed.

  3. 3.

    Add evaluate_icl_gap to the calibration utility module and wire it into the trainer eval loop.

  4. 4.

    Add icl_train_loader as a constructor argument on Trainer (after the existing arm argument, to preserve compatibility).

Appendix C Probe-Layer Ablation

We re-evaluated ICS at the F_ICS-Maxv2 step-5,0005{,}000 checkpoint with three alternative probe bands: ℒ={4,…,12}\mathcal{L}=\{4,\dots,12\} (early), ℒ={12,…,20}\mathcal{L}=\{12,\dots,20\} (mid-late), and ℒ={20,…,28}\mathcal{L}=\{20,\dots,28\} (late). The resulting ICS values are 1.4011.401, 1.4181.418, and 1.3961.396 respectively. All three bands are within 0.5%0.5\% of the original 1.4131.413 and within 1.5%1.5\% of the 2\sqrt{2} ceiling, confirming that the saturation is not a probe-band artefact.

Appendix D Post-hoc ICL-GAP for Unregularised Arms

We loaded the step-5,0005{,}000 checkpoints for F_A, F_M, and F_None and evaluated ICL-GAP on the same eval ICL set used for F_ICS-Maxv2, with the same temperature-00 multiple-choice protocol. The values are:

Arm acc𝒟1\acc_{\mathcal{D}_{1}} acc𝒟2\acc_{\mathcal{D}_{2}} ICL​-​GAP\mathrm{ICL\text{-}GAP}
F_A 0.330 0.305 +0.025+0.025
F_M 0.295 0.290 +0.005+0.005
F_None 0.355 0.320 +0.035+0.035

All three are within ±0.04\pm 0.04 of zero. Thus, the near-zero behavioural gap is shared by the unregularised checkpoints under this model and probe; what distinguishes F_ICS-Max is its ICS, not its ICL-GAP.

Appendix E Coarse Collapse Pilot

The pilot used a domain-specific instruction set rather than the broad mixture in the controlled proxy-stress experiment, and a peak learning rate of 5×10−55\times 10^{-5} rather than 1×10−51\times 10^{-5}. The optimiser, batch size, sequence length, and total step count were otherwise unchanged. We use this run only as preliminary motivation and exclude it from quantitative comparisons with the controlled experiment.

Appendix F Controlled Proxy-Stress Training Details

All controlled-ablation and constructive arms use AdamW [Loshchilov and Hutter, 2019], peak learning rate 1×10−51\times 10^{-5}, 200200 warmup steps, cosine decay, batch size 3232, sequence length 10241024, gradient clipping at norm 1.01.0, and bf16 mixed precision. Checkpoint probes are evaluated every 100100 optimiser steps; Table 3 reports the coarser 500500-step trajectory for readability.

Appendix G Probe Configurations

  • •

    ICS / ICL-GAP eval set. 500500 multiple-choice pairs, seed 4242. Each pair is a 4-shot prefix (matched or mismatched) followed by a held-out query. Disjoint from the regulariser set.

  • •

    Regulariser set. 500500 pairs, seed 9999. Same format. Used only by ℒfunc​-​icl\mathcal{L}_{\mathrm{func\text{-}icl}} in F_ICS-Max.

  • •

    MMLU probe. 240240 items stratified across the 5757 MMLU subjects, 55-shot prefix, temperature 00.

  • •

    GSM8K probe. 5050 items, 88-shot chain-of-thought prefix, temperature 00.

  • •

    ECE. 1515 equal-mass bins on MMLU.

Appendix H Pretrained Behavioural ICL Screening

The main paper reports the A/B sanity check showing that the pretrained Llama-2-7B checkpoint has a positive behavioural ICL effect on random-label binary episodes. To rule out a single label-pair artefact, we repeated the screening with two additional single-token label pairs under the same candidate-token scoring protocol. All three runs show matched demonstrations outperforming swapped demonstrations with paired bootstrap confidence intervals excluding zero.

Table 4: Pretrained behavioural ICL screening on random-label binary episodes. CIs are paired bootstrap intervals over episodes.
Labels NN Match Perm None Gap [95% CI]
A/B 480 0.669 0.362 0.492 0.306 [0.231, 0.379]
X/Y 240 0.762 0.225 0.504 0.537 [0.450, 0.621]
foo/bar 240 0.642 0.338 0.496 0.304 [0.200, 0.400]

Appendix I B-ICS: Soft-Gap Surrogate

The hard B-ICS, B-ICS(θ;δ)=ICS(θ)𝟏{ICL-GAP(θ)≥δ}\mathrm{B\text{-}ICS}(\theta;\delta)=\mathrm{ICS}(\theta)\mathbf{1}\{\mathrm{ICL\text{-}GAP}(\theta)\geq\delta\}, requires the discrete ICL-GAP, which is non-differentiable. We use the soft-gap

ICL​-​GAP~​(θ)≜𝔼μ​[log⁡ℙθ​(y∣𝒟1)−log⁡ℙθ​(y∣𝒟2)].\widetilde{\mathrm{ICL\text{-}GAP}}(\theta)\triangleq\mathbb{E}_{\mu}\!\left[\log\mathbb{P}_{\theta}(y\mid\mathcal{D}_{1})-\log\mathbb{P}_{\theta}(y\mid\mathcal{D}_{2})\right]. (21)

This substitution is a smooth heuristic rather than a theorem equating log-probability gaps with accuracy gaps.

(i) Heuristic alignment. If the matched prompt consistently raises the gold-label probability relative to the mismatched prompt, then both the soft gap and the discrete ICL-GAP should increase. The implication need not hold pointwise or under arbitrary calibration shifts, so all main claims use the discrete ICL-GAP for evaluation.

(ii) Bounded magnitude. Empirically, |ICL​-​GAP~|≤3|\widetilde{\mathrm{ICL\text{-}GAP}}|\leq 3 across the 5,0005{,}000-step trajectory of all arms. The choice δ=0.05\delta=0.05 is comparable to the single-checkpoint sampling scale of the discrete ICL-GAP (≈0.046\approx 0.046 for the 200200 logged paired examples discussed in the main paper), and the gate σ⁡(β⁡(ICL​-​GAP~−δ))\sigma(\beta(\widetilde{\mathrm{ICL\text{-}GAP}}-\delta)) stays away from σ′\sigma^{\prime}-saturation.

Appendix J B-ICS Gradient Calculation

Proposition 3 (B-ICS suppresses the bad ridge).

On a high-ICS point with ICL​-​GAP~≤0\widetilde{\mathrm{ICL\text{-}GAP}}\leq 0, the ICS-only component of the smooth B-ICS gradient is multiplied by at most σ⁡(−β​δ)\sigma(-\beta\delta).

By definition B​-​ICSβ​(θ)=ICS⁡(θ)⋅σβ​(θ)\mathrm{B\text{-}ICS}_{\!\beta}(\theta)=\mathrm{ICS}(\theta)\cdot\sigma_{\beta}(\theta) where σβ​(θ)≜σ⁡(β⁡(ICL​-​GAP~​(θ)−δ))\sigma_{\beta}(\theta)\triangleq\sigma(\beta(\widetilde{\mathrm{ICL\text{-}GAP}}(\theta)-\delta)).

∇θB​-​ICSβ\displaystyle\nabla_{\theta}\mathrm{B\text{-}ICS}_{\!\beta} =σβ​(θ)​∇θICS​(θ)\displaystyle=\sigma_{\beta}(\theta)\,\nabla_{\theta}\mathrm{ICS}(\theta) (22)
+ICS⁡(θ)​σβ​(θ)​(1−σβ​(θ))\displaystyle+\mathrm{ICS}(\theta)\,\sigma_{\beta}(\theta)\,(1-\sigma_{\beta}(\theta))
⋅β​∇θICL​-​GAP~​(θ).\displaystyle\cdot\beta\,\nabla_{\theta}\widetilde{\mathrm{ICL\text{-}GAP}}(\theta).

On the bad ridge Θ⋆∖ΘICL⋆\Theta^{\star}\setminus\Theta^{\star}_{\mathrm{ICL}}, ICL​-​GAP~≤0\widetilde{\mathrm{ICL\text{-}GAP}}\leq 0 by definition, so the argument of σ\sigma is at most −β​δ-\beta\delta, and

σβ​(θ)\displaystyle\sigma_{\beta}(\theta) ≤σ⁡(−β​δ),\displaystyle\leq\sigma(-\beta\delta), (23)
σβ​(θ)​(1−σβ​(θ))\displaystyle\sigma_{\beta}(\theta)(1-\sigma_{\beta}(\theta)) ≤σ⁡(−β​δ).\displaystyle\leq\sigma(-\beta\delta). (24)

Substituting back,

‖∇θB​-​ICSβ‖\displaystyle\|\nabla_{\theta}\mathrm{B\text{-}ICS}_{\!\beta}\| ≤σ⁡(−β​δ)\displaystyle\leq\sigma(-\beta\delta) (25)
×(‖∇θICS‖+β​2​‖∇θICL​-​GAP~‖),\displaystyle\times\bigl(\|\nabla_{\theta}\mathrm{ICS}\|+\beta\sqrt{2}\,\|\nabla_{\theta}\widetilde{\mathrm{ICL\text{-}GAP}}\|\bigr),

where we used ICS≤2\mathrm{ICS}\leq\sqrt{2}. For β=10,δ=0.05\beta=10,\delta=0.05, σ⁡(−β​δ)=σ⁡(−0.5)≈0.378\sigma(-\beta\delta)=\sigma(-0.5)\approx 0.378; for δ=0.05\delta=0.05 and β=20\beta=20, σ⁡(−1)≈0.269\sigma(-1)\approx 0.269. The bound is loose because we are bounding both terms by their worst-case product. The substantive content is that the first (ICS-only) gradient component is multiplied by σβ→0\sigma_{\beta}\to 0 on the bad ridge, while the second (gap-only) component is bounded; among the two, the gap-only component selects directions that increase ICL​-​GAP~\widetilde{\mathrm{ICL\text{-}GAP}}, i.e. leave the ridge. Hence gradient descent on ℒB​-​ICS=−B​-​ICSβ\mathcal{L}_{\mathrm{B\text{-}ICS}}=-\mathrm{B\text{-}ICS}_{\!\beta} from θ0\theta_{0} either does not move toward the ridge in the first place or, if it does, is pushed back out as ICL​-​GAP~\widetilde{\mathrm{ICL\text{-}GAP}} falls below δ\delta.

Appendix K AnchorTune: Memory and Compute

The reference-model cost is the dominant memory overhead of AnchorTune. We instantiate it as follows.

Memory.

Llama-2-7B in bf16 occupies ≈13​GB\approx 13\,\mathrm{GB}. The pretrained reference is loaded once and kept frozen (no gradients, no optimiser state), so its footprint is exactly 13​GB13\,\mathrm{GB}. The trainee adds a further 13​GB13\,\mathrm{GB} of weights, 13​GB13\,\mathrm{GB} of fp32 gradients, and ≈26​GB\approx 26\,\mathrm{GB} of fp32 AdamW optimiser state, for a total of ≈65​GB\approx 65\,\mathrm{GB} before activations. The cached anchor attentions are |𝒮anchor|⋅|ℒ|⋅H⋅T⋅4​B|\mathcal{S}_{\mathrm{anchor}}|\cdot|\mathcal{L}|\cdot H\cdot T\cdot 4\,\mathrm{B} where TT is the per-pair token length. With |𝒮anchor|=500|\mathcal{S}_{\mathrm{anchor}}|=500, |ℒ|=9|\mathcal{L}|=9, H=32H=32, T=640T=640 the cache is ≈370​MB\approx 370\,\mathrm{MB}. We pre-batch the cache into 6464 batches of 88 pairs each on GPU to amortise the per-step lookup; batches are kept in fp32 to avoid quantisation jitter feeding back into the trainee gradient. Total GPU memory at peak (forward+backward+anchor cache+activations) is ≈80​GB\approx 80\,\mathrm{GB} on our 96​GB96\,\mathrm{GB} device.

Compute.

AnchorTune adds one forward pass through the trainee on the anchor probe per optimiser step (one batch from the 6464-batch cache, ≈\approx same cost as one ICL probe forward in F_ICS-Max). The reference model is not re-evaluated during training; only the cached attention rows are used. The wall-clock per optimiser step on our hardware is ≈5​s\approx 5\,\mathrm{s} (vs. ≈4​s\approx 4\,\mathrm{s} for F_None and ≈5​s\approx 5\,\mathrm{s} for F_ICS-Max), giving a 25%25\% overhead vs. unregularised fine-tuning and parity with F_ICS-Max. A complete 5,0005{,}000-step run takes ≈7​h\approx 7\,\mathrm{h}, plus ≈3​min\approx 3\,\mathrm{min} to load the reference and pre-compute the anchor cache.

Reference-model alternatives.

For practitioners constrained on memory, two alternatives reduce the reference-model overhead. (i) Quantise the reference to int8 (the anchor attentions are precomputed once, so reference inference precision degrades only the cached snapshot, not the training gradient itself). (ii) Replace the in-memory reference with a CPU-resident reference and a one-shot disk cache of 𝒮anchor\mathcal{S}_{\mathrm{anchor}} attentions; this raises the start-up cost from 33 minutes to ≈15\approx 15 minutes on our hardware but eliminates the 13​GB13\,\mathrm{GB} GPU footprint of the reference. We use the in-memory variant for the experiments reported in this paper.

Appendix L Constructive Results: AnchorTune and B-ICS

This section tests two responses to the observed dissociation. Section L.1 examines whether the behaviour-anchored target ℒB​-​ICS\mathcal{L}_{\mathrm{B\text{-}ICS}} attenuates Goodhart saturation; Section L.2 reports the AnchorTune λ\lambda-sweep and multi-seed final; Section L.3 compares against two competitive baselines on the (MMLU, ICL-GAP) Pareto frontier; and Section L.4 ablates the anchor target.

L.1 Behaviour-Gated Sweep

We rerun the optimisation-pressure experiment with the same compute budget, the same probe ICL set, and the same hyperparameters as F_ICS-Maxv2, but replace ℒfunc​-​icl=−ICS\mathcal{L}_{\mathrm{func\text{-}icl}}=-\mathrm{ICS} with ℒB​-​ICS=−B​-​ICSβ\mathcal{L}_{\mathrm{B\text{-}ICS}}=-\mathrm{B\text{-}ICS}_{\!\beta} at β=10,δ=0.05\beta=10,\delta=0.05. By Proposition 3, ℒB​-​ICS\mathcal{L}_{\mathrm{B\text{-}ICS}} down-weights the ICS-only gradient on the disjoint-support ridge, but the factor is still substantial at our mild setting β​δ=0.5\beta\delta=0.5, and the smooth surrogate ICL​-​GAP~​(θ)\widetilde{\mathrm{ICL\text{-}GAP}}(\theta) of (21) can rise transiently above δ\delta during training even when the discrete ICL​-​GAP\mathrm{ICL\text{-}GAP} does not.

The empirical trajectory clarifies the practical strength of the gating effect. Table 5 reports the step-5,0005{,}000 values. F_BICS ends at ICS=1.397\mathrm{ICS}=1.397, only 0.0160.016 below the stress-test F_ICS-Max value of 1.4131.413 and within 1.2%1.2\% of the 2\sqrt{2} ceiling. MMLU accuracy drops to 0.2830.283, essentially the same degradation as F_ICS-Max. What does differentiate the two arms is the ICL-GAP: F_BICS finishes at +0.020+0.020 vs. F_ICS-Max’s −0.010-0.010, and tracking its trajectory shows the soft gap exceeds δ=0.05\delta=0.05 on intervals where the behavioural sentinel intermittently activates, before the divergence term overwhelms it. This indicates that B-ICS as a training target requires a more aggressive gating schedule (larger β\beta, larger δ\delta, or a hard 𝟏​{⋅}\mathbf{1}\{\cdot\} gate with a straight-through estimator) than the smooth β=10\beta=10 we chose for differentiability; the observed +0.030+0.030 behavioural advantage of F_BICS over F_ICS-Max is consistent with the attenuation predicted by Proposition 3 under that mild gating.

When evaluated on a model that was not trained against it, B-ICS retains the intended behavioural guard. It multiplies ICS by a binary indicator that the behavioural sentinel passes and returns near-zero for F_ICS-Max (which has ICL​-​GAP=−0.010<δ\mathrm{ICL\text{-}GAP}=-0.010<\delta) and near-pretrained for F_anchor at λ⋆\lambda^{\star}. The diagnostic version therefore keeps the intended behavioural guard, while the training-target version requires stronger gating than we use here.

Table 5: F_BICS at β=10,δ=0.05\beta=10,\delta=0.05 attenuates the ICS-vs.-behaviour dissociation but does not eliminate it. The arm still reaches ICS=1.397\mathrm{ICS}=1.397 (98.8%98.8\% of the 2\sqrt{2} ceiling), but achieves a +0.030+0.030 ICL-GAP improvement over F_ICS-Max. We attribute the residual saturation to the softness of the gate; a hard 𝟏​{⋅}\mathbf{1}\{\cdot\} gate with a straight-through estimator is a candidate for further evaluation.
Arm ICS (→2)(\to\sqrt{2}) ICL-GAP MMLU ECE
Pretrained 0.516 — 0.371 —
F_ICS-Max (stress test) 1.413 −0.010-0.010 0.279 0.231
F_BICS (this work) 1.397 +0.020+0.020 0.283 0.388

L.2 AnchorTune Sweep

For AnchorTune we run a four-point λ\lambda-sweep λ∈{0.01,0.05,0.20,1.00}\lambda\in\{0.01,0.05,0.20,1.00\} at seed 4242, then a three-seed final at the chosen λ⋆\lambda^{\star}. Table 6 reports the final-step values; Figure 6 plots the per-step trajectories of the four core quantities for F_anchor at λ⋆\lambda^{\star} alongside F_None and F_ICS-Max.

The empirical claim is that F_anchor at λ⋆\lambda^{\star} closes the MMLU damage that F_ICS-Max inflicts: from 0.2790.279 (chance +0.03+0.03) back to within seed-noise of the pretrained baseline of 0.3710.371, while keeping ICS indistinguishable from pretrained. Concretely, across three seeds at λ⋆=0.05\lambda^{\star}=0.05, MMLU returns 0.370±0.0580.370\pm 0.058 (mean ±\pm std), ICS returns 0.500±0.0010.500\pm 0.001, and ICL-GAP returns −0.012±0.017-0.012\pm 0.017 — all within seed-variation of the pretrained model. The MMLU recovery accounts for 0.0910.091 of the 0.0920.092 accuracy drop that F_ICS-Max produced. Of the four λ\lambda values, λ=0.05\lambda=0.05 is the empirical optimum (MMLU 0.3670.367); both λ=0.20\lambda=0.20 and λ=1.00\lambda=1.00 converge to a slightly lower MMLU plateau of 0.3500.350, suggesting that the anchor pulls the trainee too tightly toward θ0\theta_{0} once λ\lambda exceeds ≈0.1\approx 0.1, and λ=0.01\lambda=0.01 is too weak to prevent the small MMLU drift seen for the unregularised arm.

Table 6: AnchorTune λ\lambda-sweep at seed 42 and three-seed final at λ⋆=0.05\lambda^{\star}=0.05. Pretrained ICS is 0.516 and pretrained MMLU is 0.371. At λ⋆\lambda^{\star}, the three-seed AnchorTune result is MMLU 0.370±0.0580.370\pm 0.058.
Configuration ICS (→0.516)(\to 0.516) ICL-GAP MMLU (→0.371)(\to 0.371) ECE
Pretrained 0.516 — 0.371 —
F_None (controlled) 0.500 ≈0\approx 0 0.375 0.407
F_ICS-Max (stress test) 1.413 −0.010-0.010 0.279 0.231
F_anchor, λ=0.01\lambda=0.01 0.501 −0.040-0.040 0.338 0.472
F_anchor, λ=0.05\lambda=0.05 0.500 −0.010-0.010 0.367 0.375
F_anchor, λ=0.20\lambda=0.20 0.500 −0.020-0.020 0.350 0.374
F_anchor, λ=1.00\lambda=1.00 0.501 −0.015-0.015 0.350 0.391
F_anchor, λ⋆=0.05\lambda^{\star}{=}0.05, seed 42 0.500 −0.010-0.010 0.367 0.375
F_anchor, λ⋆=0.05\lambda^{\star}{=}0.05, seed 43 0.501 +0.005+0.005 0.429 0.384
F_anchor, λ⋆=0.05\lambda^{\star}{=}0.05, seed 44 0.500 −0.030-0.030 0.313 0.387
F_anchor, λ⋆\lambda^{\star}, mean ±\pm std 0.500 ±\pm 0.001 −0.012±0.017-0.012\pm 0.017 0.370 ±\pm 0.058 0.382 ±\pm 0.006
Refer to caption
Figure 5: ICS and MMLU at step 5,000 on Llama-2-7B. The dashed vertical line marks the 2\sqrt{2} ceiling. F_ICS-Max saturates ICS while MMLU approaches chance, and F_BICS partially follows the same pattern under the smooth β=10\beta=10 gate. F_anchor at λ⋆=0.05\lambda^{\star}{=}0.05 and the logit-anchor baseline LogitAnchor keep ICS near the pretrained value while maintaining MMLU within the observed seed variation.
Refer to caption
Figure 6: Per-step trajectories of ICS, ICL-GAP, MMLU accuracy, and MMLU ECE across 5,0005{,}000 training steps for five representative arms. F_ICS-Max and F_BICS rapidly drive ICS toward the 2\sqrt{2} ceiling. F_anchor, LogitAnchor, and WeightAnchor keep all four quantities near the pretrained baseline. ICL-GAP fluctuates around zero throughout; the single-checkpoint sampling scale for the 200200 logged paired examples is about 0.0460.046.

L.3 Anchored-Baseline Comparison

We compare F_anchor with two capability-preservation baselines, all evaluated with the same training and evaluation protocol:

  • •

    LogitAnchor: a logit-level KL-to-base regulariser of the form widely used in RLHF [Ouyang et al., 2022], ℒKL​-​logits(θ;θ0,𝒮anchor)=𝔼𝒮anchor[KL(pθ(⋅∣𝒟1)∥pθ0(⋅∣𝒟1))]\mathcal{L}_{\mathrm{KL\text{-}logits}}(\theta;\theta_{0},\mathcal{S}_{\mathrm{anchor}})=\mathbb{E}_{\mathcal{S}_{\mathrm{anchor}}}[\KL(p_{\theta}(\cdot\mid\mathcal{D}_{1})\,\|\,p_{\theta_{0}}(\cdot\mid\mathcal{D}_{1}))]. This is the closest existing analogue of AnchorTune, but it anchors predictions rather than attention rows.

  • •

    WeightAnchor: an EWC-style [Kirkpatrick et al., 2017] weight anchor ℒℓ2​-​θ0​(θ,θ0)=‖θ−θ0‖22\mathcal{L}_{\ell_{2}\text{-}\theta_{0}}(\theta;\theta_{0})=\|\theta-\theta_{0}\|_{2}^{2}, with λ=10−7\lambda=10^{-7}.

Table 7 reports the step-5,0005{,}000 values.

The empirical pattern is more nuanced than “AnchorTune dominates the baselines.” In this single-seed comparison, all three anchored arms (F_anchor, LogitAnchor, WeightAnchor) recover MMLU to within seed noise of pretrained (0.3670.367, 0.3830.383, and 0.3630.363 respectively at seed 4242, vs. pretrained 0.3710.371 and F_None’s 0.3750.375), and all three keep ICS at the pretrained 0.500±0.0050.500\pm 0.005 range. The three arms sit on the same Pareto front; what distinguishes them is what they anchor and therefore what they expose. LogitAnchor ties the trainee’s predictions to θ0\theta_{0}, which maintains task accuracy but masks the internal attention dynamics; WeightAnchor ties the trainee’s full weight vector to θ0\theta_{0}, which maintains both but at the cost of any meaningful fine-tuning signal at the λ=10−7\lambda=10^{-7} we tested; F_anchor ties the trainee’s mid-band attention rows to θ0\theta_{0}, which maintains the diagnostic interpretability of ICS as a measure of attention-level ICL drift. In this setup, the three objectives are therefore complementary: AnchorTune’s contribution is not primarily a larger MMLU number but an objective whose anchor signal and diagnostic signal agree. This alignment allows ICS to be interpreted relative to the pretrained attention behaviour rather than a logit-constrained output alone. The contrast with F_ICS-Max is the main result: the divergence-based regularisers we test (ICS and smooth B-ICS as targets) move toward the high-ICS, low-MMLU region, while the anchored regularisers we test (logit, weight, attention) maintain MMLU. This comparison identifies a structural distinction between divergence and anchor losses, not a critique of attention-level signals per se.

Table 7: AnchorTune compared with two capability-preservation baselines and the F_ICS-Max regulariser. AnchorTune values are the seed-42 result for direct comparison with the single-seed baselines; Table 6 reports its three-seed mean.
Arm ICS ICL-GAP MMLU ECE
Pretrained 0.516 — 0.371 —
F_None 0.500 ≈0\approx 0 0.375 0.407
F_ICS-Max 1.413 −0.010-0.010 0.279 0.231
LogitAnchor (KL-to-logits) 0.494 −0.020-0.020 0.383 0.373
WeightAnchor (weight ℓ2\ell_{2}, λ=10−7\lambda{=}10^{-7}) 0.500 −0.010-0.010 0.363 0.419
F_anchor (ours, λ⋆=0.05\lambda^{\star}{=}0.05) 0.500 −0.010-0.010 0.367 0.375
Refer to caption
Figure 7: (MMLU, ICL-GAP) Pareto scatter at step 5,0005{,}000. Divergence regularisers (red X for F_ICS-Max, orange for F_BICS) sit in the low-MMLU region; anchor regularisers (blue/green for F_anchor, LogitAnchor, WeightAnchor) cluster around the pretrained reference point (black star). The two regulariser families form distinct clusters on this plot.

L.4 Anchor-Target Ablation

To isolate the active ingredient of AnchorTune, we vary which part of the model is anchored while keeping λ\lambda and 𝒮anchor\mathcal{S}_{\mathrm{anchor}} fixed:

  • •

    F_anchor_last4: anchor attention rows only on the last 44 layers, not on the mid-band ℒ={8,…,16}\mathcal{L}=\{8,\dots,16\}.

  • •

    F_anchor_lmhead: anchor only the LM head weights via an ℓ2\ell_{2} penalty toward θ0\theta_{0}; do not anchor any attention rows.

The hypothesis is that mid-band attention-row anchoring is what carries the MMLU-preservation effect. If F_anchor_last4 and F_anchor_lmhead both drop MMLU back toward F_None levels while F_anchor does not, the attribution to the mid-band attention rows is established.

The results in Table 8 weaken the original mid-band attribution hypothesis. Restricting the anchor to the last 4 layers (F_anchor_last4) gives MMLU=0.358\mathrm{MMLU}=0.358 vs. full-layer F_anchor’s 0.3670.367, only a 0.0090.009-point reduction — so the mid-band attention rows are sufficient but not necessary for MMLU preservation; the last 4 layers alone come within seed noise of the full anchor. More strikingly, the LM-head-only ℓ2\ell_{2} variant (F_anchor_lmhead, no attention anchoring at all) achieves MMLU=0.375\mathrm{MMLU}=0.375, matching the pretrained reference of 0.3710.371 to within seed noise and exceeding seed-42 F_anchor. This means keeping lm​_​head\mathrm{lm\_head} close to θ0\theta_{0} in weight space is by itself enough to recover MMLU under our fine-tuning configuration — attention-row anchoring is one route to MMLU preservation, but not the only one. These results indicate that the mechanism behind F_anchor’s MMLU recovery is partially shared with WeightAnchor (Table 7) and F_anchor_lmhead: proximity to θ0\theta_{0} in any structural component (weights, lm_head, attention rows) is sufficient. AnchorTune’s residual advantage over these baselines is not absolute MMLU but the coincidence of its anchor target with the diagnostic ICL signal, as discussed in Section L.3.

Table 8: Anchor-target ablation. All three variants keep ICS within ±0.020\pm 0.020 of pretrained while maintaining MMLU; the lm_head-only ℓ2\ell_{2} variant without an attention anchor matches pretrained MMLU, weakening the claim that mid-band attention rows are uniquely necessary for MMLU recovery.
Anchor target ICS ICL-GAP MMLU ECE
Pretrained 0.516 — 0.371 —
None (= F_None, controlled reference) 0.500 ≈0\approx 0 0.375 0.407
Mid-band attn rows (F_anchor) 0.500 −0.010-0.010 0.367 0.375
Last 4 layers attn rows (F_anchor_last4) 0.510 −0.015-0.015 0.358 0.377
LM head ℓ2\ell_{2} only (F_anchor_lmhead) 0.494 −0.020-0.020 0.375 0.378

L.5 Summary of the constructive part

The constructive experiments support three claims. (i) The Goodhart channel is not an artefact of using attention as a target; it is an artefact of the divergence-on-disjoint-supports geometry. B-ICS as a diagnostic inherits the safety motivation of Proposition 3; B-ICS as a training target only partially attenuates the channel under the smooth gating we chose (β=10,δ=0.05\beta=10,\delta=0.05), and reaches ICS=1.397\mathrm{ICS}=1.397 at convergence (vs. 1.4131.413 for F_ICS-Max). A hard gate or a larger β\beta remains to be tested. (ii) Anchoring the per-head attention rows to the pretrained model on a fixed 𝒟1\mathcal{D}_{1} probe maintains MMLU under exactly the same fine-tuning configuration where F_ICS-Max produces proxy saturation and MMLU damage: three-seed MMLU at λ⋆=0.05\lambda^{\star}=0.05 is 0.370±0.0580.370\pm 0.058, returning 99%99\% of the F_ICS-Max→F_None\textsc{F\_ICS-Max}\rightarrow\textsc{F\_None} MMLU gap. (iii) The Pareto contrast in our tested arms is structural: anchored regularisers (attention rows in F_anchor, logits in LogitAnchor, weights in WeightAnchor) maintain MMLU, whereas the divergence targets we test do not. AnchorTune’s specific contribution among anchored regularisers is that its anchor signal coincides with the diagnostic ICL signal, so that the same probe can be read out at training and at evaluation without a Goodhart concern.