Attention Sensitivity Is Not Enough:
Dissociating Attention-Level and Behavioural
In-Context Learning under Fine-Tuning
| Jinyuan Zhang | Peng He∗ |
|---|---|
| 202621116012480@stu.hubu.edu.cn | penghe@hubu.edu.cn |
| Yin Yuan | He Hu |
| 202521120012766@stu.hubu.edu.cn | 202521120012751@stu.hubu.edu.cn |
| ShengShuo Jiao | |
| 202621120012764@stu.hubu.edu.cn |
Hubei University, Wuhan, China
∗Corresponding author: penghe@hubu.edu.cn
Abstract
In-context learning (ICL) lets large language models adapt to new tasks from demonstrations, and fine-tuning can erode this behaviour. Many preservation diagnostics inspect attention: if attention changes when demonstrations change, the model is treated as context-sensitive. This paper asks how far that proxy can be trusted once it is optimised. We formalise In-Context Sensitivity (ICS), the average row distance between last-token attention on matched and mismatched demonstration prefixes, and pair it with ICL-GAP, the behavioural accuracy gap between the same prefixes. In a controlled four-arm ablation on Llama-2-7B, an ICS-maximising regulariser (F_ICS-Max) drives ICS to , within of its geometric ceiling. The behavioural readout tells a different story: ICL-GAP stays near zero and MMLU accuracy moves from to , a Goodhart dissociation of the bounded attention proxy. Endpoint statistics locate the mechanism: attention grows sharp and near-disjoint across prefixes yet routes to formatting and demonstration-body tokens rather than labels. A random-label protocol confirms that the behavioural probe family retains dynamic range at the same checkpoints. In a constructive sweep, behaviour gating partially mitigates the effect, while objectives anchored to pretrained computation hold the high-MMLU, moderate-ICS region that divergence maximisers leave. The main lesson is diagnostic: attention-level ICL proxies earn their place as training targets only after validation against behavioural gaps.
1 Introduction
Large language models can solve new tasks from demonstrations in their input, a capability known as in-context learning (ICL) (Brown et al., 2020; Wei et al., 2022). ICL requires no parameter updates, but subsequent fine-tuning can alter the behaviour on which this adaptation relies. Continued pretraining and instruction tuning have been observed to weaken ICL on held-out tasks (Luo et al., 2023; Shi et al., 2024; Wang et al., 2023b), motivating methods that measure and limit this drift.
Many diagnostics inspect attention because ICL has been associated with mechanisms such as induction heads and copying circuits (Olsson et al., 2022; Akyürek et al., 2023). A common approach compares attention maps under a matched prefix containing correct demonstrations and a mismatched prefix containing shuffled or random demonstrations. The resulting divergence measures whether attention responds to the demonstrations. It is inexpensive and differentiable, which makes it suitable for monitoring and, potentially, regularisation—though on its own it says nothing about whether that response improves task behaviour.
We test whether an attention-level ICL proxy remains faithful when it becomes an optimisation target. Attention maps are already used in distillation, alignment, and faithfulness analysis (Jiao et al., 2020; Tropeano et al., 2026; Yao et al., 2026), making differentiable attention probes plausible auxiliary preservation signals. Without attributing the exact ICS objective to prior work, we use a matched-vs.-mismatched attention diagnostic to test the broader assumption that increasing context-sensitive attention preserves behavioural ICL.
Setup and central finding.
We formalise the proxy as In-Context Sensitivity (ICS): the mean distance, over last-token attention rows in a band of mid-network layers, between a matched and a mismatched demonstration prefix. Because each attention row is a probability distribution, ICS has a fixed geometric ceiling, attained only when the two prefixes induce attention concentrated on disjoint single tokens. Its behavioural counterpart is ICL-GAP: the held-out task accuracy under the matched prefix minus the held-out accuracy under the mismatched prefix. ICL-GAP is the behavioural quantity any ICL-preservation method actually wants to keep above zero.
On Llama-2-7B we run a four-arm controlled ablation, with all arms starting from the same checkpoint and training for 5,000 steps on a fixed instruction-tuning mixture. The first three arms freeze attention, freeze MLPs, or fine-tune all parameters without an auxiliary loss; the fourth, F_ICS-Max, adds an ICS-maximising term to the cross-entropy loss. The unregularised arms remain close to the pretrained ICS of 0.516. F_ICS-Max reaches 1.413, within 0.5 percent of the theoretical ceiling of about 1.414, while ICL-GAP on a held-out QA probe remains near zero and MMLU accuracy falls from 0.371 to 0.279.
Main empirical finding.
Considered alone, the increase in attention divergence could be read as stronger context sensitivity. The behavioural measurement points the other way, and the gap between the two readings is exactly why attention-level ICL diagnostics need validation against a behavioural gap before they are used under optimisation pressure.
Contributions.
We make four contributions. First, we validate a matched-vs.-mismatched behavioural probe and define ICS as a geometrically bounded attention diagnostic paired with behavioural ICL-GAP. Second, a controlled Llama-2-7B ablation shows that explicit ICS maximisation saturates the proxy while leaving ICL-GAP near zero and degrading MMLU. Third, attention endpoint statistics and a structural construction show how sharp, disjoint routing can maximise ICS without selecting answer-relevant tokens. Finally, we test behaviour-gated B-ICS and pretrained anchoring as constructive responses to this dissociation. The supplement gives the contamination postmortem, full trajectories, layer ablations, and constructive sweeps.
2 Related Work
What ICL is, and where it lives.
In-context learning emerged as an empirical property of large transformers (Brown et al., 2020; Wei et al., 2022), and mechanistic work has since tied it to specific circuitry: induction heads that copy an earlier pattern’s continuation forward (Olsson et al., 2022), and attention layers that implement gradient-descent-like or Bayesian updates over the prefix (Akyürek et al., 2023; von Oswald et al., 2023; Xie et al., 2022). Recent probes cut the other way: in tabular and structured-data settings, apparent ICL can reduce to memorisation or label priors (Capano and Böhler, 2026; Pelusi et al., 2026). These accounts disagree about what ICL is, but share one implication: intact ICL should respond to demonstrations in attention. We do not contest that direction; we put the converse under optimisation pressure—that context-sensitive attention implies intact ICL—once the response becomes a training signal.
ICL under fine-tuning.
Continued training moves this behaviour, usually tracked at the output. Instruction tuning, supervised fine-tuning, and continual pretraining can all reduce held-out ICL accuracy (Luo et al., 2023; Shi et al., 2024; Wang et al., 2023b), and finer-grained work follows how fine-tuning reshapes in-context factual recall (Huang et al., 2026). The measurement is behavioural throughout: ICL accuracy, or accuracy gaps between matched and mismatched prefixes. Our ICL-GAP belongs to that family; we depart by instrumenting an attention-level diagnostic, ICS, alongside it and studying the relationship between the two under training.
Attention as evidence.
Attention maps are tempting evidence: cheap, differentiable, already load-bearing. Distillation transfers them between models (Jiao et al., 2020); pruning attention layers changes explanation faithfulness and confidence (Tropeano et al., 2026); alignment work optimises attention dynamics as a preference signal (Yao et al., 2026). Yet attention weights need not track the computation behind a prediction. Our stress test sharpens that caution: not whether attention explains a fixed model, but whether an attention proxy stays meaningful while being optimised.
Component-level attribution.
A parallel line localises capabilities to subnetworks. Mechanistic interpretability isolates circuits at the level of heads and MLP neurons (Elhage et al., 2021; Wang et al., 2023a); parameter-efficient tuning trains only attention or only MLP blocks (Houlsby et al., 2019; He et al., 2022). Our F_A and F_M arms repurpose the second design for a stability question: which subnetwork’s update carries ICL drift? In our setup neither does alone—both preserve ICS near its pretrained value—motivating an explicit proxy regulariser instead of a structural constraint.
Auxiliary objectives under Goodhart pressure.
Regularisers that protect old behaviour assume the auxiliary signal tracks the capability. EWC (Kirkpatrick et al., 2017), synaptic intelligence (Zenke et al., 2017), and KL-to-base penalties (Ouyang et al., 2022) anchor to different quantities, but each substitutes a measurable surrogate for the behaviour itself. Goodhart’s law (Goodhart, 1984; Manheim and Garrabrant, 2018) names the failure mode; the RLHF literature gives modern instances, from reward-model over-optimisation (Skalse et al., 2022; Gao et al., 2023) to reward hacking at inference time and beyond (Khalaf et al., 2025; Fu et al., 2026; Liu et al., 2026). Those warnings concern reward functions; we give the same failure an attention-level instance, driving ICS within of its geometric ceiling while behavioural ICL stays near zero.
Calibration as a side signal.
Confidence can part ways with behaviour as well. Fine-tuning shifts calibration (Guo et al., 2017; Desai and Durrett, 2020); we record ECE (Naeini et al., 2015) on MMLU (Hendrycks et al., 2021) as a diagnostic only, since lower ECE under near-chance predictions says little about reasoning quality.
Where this work sits.
ICL-preservation methods fall into three families: freeze a parameter subset, replay ICL-format data, or regularise toward base-model outputs. Our F_A, F_M, and F_None arms cover the first family in controlled form and our logit-anchoring baseline the third; replay is discussed but not completed in the main comparison. AnchorTune is closest to attention distillation (Jiao et al., 2020; Gu et al., 2023), but anchors the trainee to its own pretrained checkpoint on last-token attention rows over a fixed ICL probe. We present it as one member of an anchored family whose anchor matches the diagnostic under test, not as a uniquely superior method; B-ICS adds the complement, gating the proxy by behaviour rather than tying it to a checkpoint.
3 Methodology
This section defines the diagnostic framework used in the rest of the paper. The goal is deliberately narrow: separate an attention-level sign of context sensitivity from the behavioural effect that an ICL-preservation method is meant to preserve. We first define the matched-vs.-mismatched problem setting and its diagnostic metrics, then introduce the proxy-maximisation stress test and two constructive variants. A separate random-label protocol validates the behavioural measurement before the main stress-test results.
3.1 Problem Setting
Let be a pretrained autoregressive transformer. For a labelled task with label set , draw demonstrations and a random label permutation . For the same query , we compare a matched prefix with a mismatched prefix that preserves the token template but breaks the input–label correspondence. A model with behaviourally intact ICL should satisfy
| (1) |
This matched-vs.-mismatched contrast is the common protocol behind both the attention diagnostic and the behavioural probe.
3.2 Diagnostic Metrics
For a prefix , layer , and head , write the last-token attention row as
| (2) |
We define In-Context Sensitivity (ICS) by averaging the distance between the matched and mismatched rows over a probe distribution and a layer band :
| (3) |
where
| (4) |
For Llama-2-7B we use ; the supplement ablates this choice.
Proposition 1 (Geometric ceiling).
For any , . For any attention rows ,
| (5) |
with equality iff and are one-hot on disjoint coordinates. Thus requires every probed matched/mismatched attention-row pair to be one-hot and disjoint almost surely.
ICS measures whether attention reorganises when demonstrations change. The behavioural target is the accuracy gap induced by the same intervention. Let . We define
| (6) | ||||
| (7) |
The central question is whether high ICS implies positive ICL-GAP under optimisation pressure. We also report MMLU accuracy and Expected Calibration Error (ECE):
| (8) |
3.3 Proxy-Maximisation Stress Test
To test the proxy rather than merely observe it, we add an ICS-maximising auxiliary objective to the ordinary cross-entropy loss. Let
| (9) |
be a held-out regulariser set, and let be the empirical estimate of Equation 3 on that set. The stress-test arm minimises
| (10) | ||||
| (11) |
with . If ICS is a faithful proxy for ICL, this arm should improve the matched-vs.-mismatched behavioural gap; any proxy gain without a matching behavioural gain then measures the proxy’s Goodhart exposure directly.
3.4 Constructive Variants
We evaluate two compact variants aimed at mitigating this dissociation. B-ICS gates the attention proxy by a differentiable behavioural sentinel:
| (12) |
where
| (13) |
The soft gap serves as a differentiable surrogate for the discrete ICL-GAP, without any claim that log-probability gaps and accuracy gaps agree pointwise.
AnchorTune instead anchors the trainee’s matched-prefix attention rows to the pretrained model on :
| (14) |
with objective
| (15) |
Unlike a divergence-maximising loss, this objective is minimised at the pretrained attention pattern rather than on the disjoint-support ridge of Proposition 1.
4 Experiments and Analysis
We organise the empirical study around four research questions. They progress from validating the behavioural measurement, through establishing and explaining the proxy–behaviour dissociation, to testing constructive responses:
- •
RQ0. Does the pretrained model exhibit behavioural ICL under a controlled matched-vs.-mismatched protocol?
- •
RQ1. Can an attention-level ICL proxy be driven to its geometric ceiling while behavioural ICL remains absent?
- •
RQ2. What changes inside the model when the proxy is maximised, and why does this fail to improve behaviour?
- •
RQ3. Can behavioural gating or pretrained anchoring reduce this proxy Goodhart failure?
After a shared experimental setup, the following modules answer these questions in order. Each module presents question-specific evidence and closes with a direct answer.
4.1 Shared Training and Evaluation Setup
All main arms start from the same Llama-2-7B pretrained checkpoint (Touvron et al., 2023) and run for optimiser steps on the same instruction-tuning mixture. The arms share the optimiser, schedule, sequence length, batch size, precision, and evaluation cadence; exact training details are in the supplement. Probes are evaluated every steps, and the supplement reports full trajectories.
The controlled ablation compares four arms. F_A freezes attention projections and updates the MLP, layer norms, and language-model head. F_M freezes the MLP blocks and updates attention projections, layer norms, and the head. F_None full-finetunes all parameters with cross-entropy only. F_ICS-Max full-finetunes all parameters with the ICS-maximising stress-test loss in Equation 11.
To avoid probe contamination, the regulariser and evaluation pair sets are disjoint:
| (16) |
The regulariser set has pairs sampled with seed ; the evaluation set has pairs sampled with seed . ICS and ICL-GAP are always reported on . MMLU uses a fixed -item subset stratified over subjects, evaluated under a -shot prefix at temperature . ECE uses the MMLU predictions with equal-mass bins. GSM8K exact-match, which sits at zero for the base model across arms, is excluded from the main evidence.
4.2 Does the Behavioural Probe Detect ICL? (RQ0)
A proxy-stress test is only meaningful if the base model can use demonstrations in at least one held-out protocol. We therefore screened pretrained Llama-2-7B on random-label binary classification episodes. Each episode samples a semantic split, assigns the two classes arbitrary labels, and compares three prompts:
| (17) |
The model scores only the two candidate label tokens. Pretrained accuracy is with matched demonstrations, with swapped demonstrations, and without demonstrations. The paired bootstrap estimate is
| (18) | ||||
| (19) |
Thus the checkpoint has a clear behavioural ICL effect on this sanity-check protocol. We keep this result separate from the controlled QA probe, which is harder and has a near-zero behavioural gap across arms.
To check that this behavioural measurement family still has dynamic range after fine-tuning, we reran the same RQ0 protocol on the step- checkpoints of all four arms. Table 1 shows that every arm retains a positive matched-vs.-permuted gap with a bootstrap confidence interval excluding zero. This does not rescue the controlled QA gap; rather, it rules out the simpler objection that all behavioural probes are insensitive at these checkpoints. In particular, F_ICS-Max reaches the largest RQ0 gap while leaving the controlled QA ICL-GAP near zero.
| Arm | Match | Perm. | None | Gap [95% CI] |
|---|---|---|---|---|
| F_A | 0.663 | 0.360 | 0.492 | 0.302 [0.235, 0.371] |
| F_M | 0.656 | 0.381 | 0.492 | 0.275 [0.206, 0.346] |
| F_None | 0.623 | 0.435 | 0.492 | 0.188 [0.115, 0.260] |
| F_ICS-Max | 0.692 | 0.340 | 0.492 | 0.352 [0.292, 0.413] |
Answer to RQ0.
Yes. The pretrained checkpoint shows a clear matched-vs.-mismatched behavioural effect, and the same probe family retains dynamic range after fine-tuning. This validation does not imply that the harder controlled QA probe must be positive; it establishes that its near-zero gap is not caused by universal probe insensitivity.
4.3 Can ICS Saturate without Behavioural ICL? (RQ1)
Before the controlled ablation, we ran a coarse full-fine-tuning baseline under an earlier data mixture and a higher learning rate. Starting from , that run collapsed to at step , a relative reduction of . It serves only as motivation—an attention-proxy analogue of ICL fragility (Luo et al., 2023) rather than a behavioural-collapse claim.
Table 2 reports the step- values of ICS, ICL-GAP, MMLU accuracy, and ECE for the controlled ablation, with the pretrained baseline in the first row. ICL-GAP was logged during training for F_ICS-Max and later evaluated post-hoc for the unregularised arms.
| Arm | ICS () | ICL-GAP () | MMLU acc. () | ECE () |
|---|---|---|---|---|
| Pretrained Llama-2-7B | 0.516 | — | — | — |
| F_A | 0.492 | — | 0.338 | 0.434 |
| F_M | 0.482 | — | 0.342 | 0.406 |
| F_None | 0.500 | — | 0.375 | 0.407 |
| F_ICS-Max | 1.413 | 0.279 | 0.231 |
F_A, F_M, and F_None all finish within of the pretrained baseline ICS of ; the spread across these three arms is itself only . Under the controlled stress-test mixture and hyperparameters, neither restricting updates to MLP layers (F_A) nor to attention layers (F_M) nor leaving them unconstrained (F_None) substantially alters attention-level context responsiveness. We expected F_M, which can update attention, to drift further from baseline than F_A, but the observed gap ( vs. ) is at the noise level of the three-arm spread and supports no ordering between the two. Nor does F_None reproduce the coarse-pilot collapse, finishing at . Since that pilot used different data and hyperparameters, its endpoint is not quantitatively comparable to the controlled experiment, and we leave a controlled unregularised-collapse condition to future work.
Because the MMLU probe has only items, we avoid interpreting small differences among the unregularised arms: at the binomial standard error is about , which puts differences of order within sampling noise at the current run count. The large F_ICS-Max drop from to is the general-ability signal we use.
F_ICS-Max moves ICS from at step , which is statistically indistinguishable from the pretrained baseline of , to at step . The convergence value sits within of the geometric ceiling from Proposition 1: the attention divergence between matched and mismatched demonstrations has been amplified by a factor of relative to the pretrained model. Read in isolation, this looks like a successful ICL-preservation regulariser.
The behavioural picture is the opposite. ICL-GAP starts at at step ( vs. ), briefly becomes positive in the middle of training (peak at step ), and ends at at step . Across the logged probe points, ICL-GAP has sample mean and sample standard deviation . The trajectory evaluation used paired examples per checkpoint, although the fixed evaluation set contains pairs; at per-condition accuracy , the unpaired binomial-difference scale is
| (20) |
This is a conservative scale for individual logged checkpoints because the actual matched/mismatched evaluations are paired. The trajectory mean is about one ninth of this scale and within of zero: the model has not learned to use correct demonstrations more than incorrect ones at any point, despite the attention divergence multiplying.
Answer to RQ1.
Yes. Optimisation drives ICS to , essentially its geometric ceiling, while the controlled QA ICL-GAP remains statistically and practically near zero and MMLU degrades. Proxy saturation and behavioural preservation therefore come apart under optimisation pressure.
4.4 How Does Proxy Maximisation Come Apart from Behaviour? (RQ2)
The full trajectory shows a monotone rise of ICS toward the ceiling, an approximately monotone MMLU drop once ICS passes , and ICL-GAP fluctuating around zero throughout. MMLU accuracy under F_ICS-Max falls from at step to at step , a drop well beyond the sampling scale above and within of the random-chance baseline of . ECE is lower for F_ICS-Max () than for the unregularised arms (–), but this reflects the model growing more uncertain as accuracy approaches chance rather than improved calibration.
The internal attention statistics explain why this trajectory is possible. At the pretrained baseline, the top- attention key under lands on a label token of the time. At F_ICS-Max step , the same rate falls to , while punctuation or formatting tokens account for and content tokens from the demonstration body for . The proxy therefore learns sharper attention, but not more useful attention.
Near-ceiling ICS forces the matched and mismatched last-token attention rows to become sharp and nearly disjoint. In held-out probe rows, mean top- concentration rises from at the pretrained baseline to at step , while matched/mismatched top- supports overlap on only of head-example pairs, down from . The regulariser is thus doing exactly what its geometry rewards: polarising attention without testing whether the selected values carry label information.
Proposition 2 (Goodhart channel: vacuous maximisers exist).
Assume the model can make every probe head one-hot on disjoint coordinates as in Proposition 1, and that there are prefix tokens whose values are conditionally independent of given the prompt format. Then some maximises ICS while satisfying , and cannot distinguish this behaviourally vacuous maximiser from an ICL-useful one with the same attention-row divergence.
Sketch.
Route each probe head to two disjoint prompt-format tokens whose values are independent of . This attains , but the readout carries no matched-prefix answer information, so matched and mismatched accuracies are equal in expectation. Since only sees row distance, useful and vacuous disjoint-support solutions receive the same proxy score. ∎
Proposition 2 also explains the apparent calibration anomaly noted above: near-chance predictions are close to uniform, which lowers ECE without any improvement in calibrated reasoning.
Answer to RQ2.
Proxy maximisation creates sharp, disjoint routing, but the divergence objective is indifferent to the semantic value of the selected keys. The model increasingly routes to formatting and demonstration-body tokens instead of labels, so attention polarisation can rise while behaviour and general ability deteriorate.
4.5 Can Gating or Anchoring Mitigate the Dissociation? (RQ3)
The dissociation suggests a practical rule: do not maximise a bounded attention divergence without either a behavioural guard or an anchor to the pretrained computation. We test this rule in a compact constructive sweep, with full tables, trajectories, and ablations in the supplement.
Smooth B-ICS uses Equation 12 as the auxiliary target. Under our mild gate , it partially mitigates the Goodhart channel: F_BICS reaches , MMLU , and ICL-GAP , compared with F_ICS-Max’s , MMLU , and ICL-GAP . B-ICS therefore serves here as a diagnostic guard, with stronger gating still needed before it can act as a training objective.
Anchored objectives behave differently. AnchorTune, KL-to-logits anchoring, and weight anchoring all keep ICS near and MMLU near the pretrained reference, whereas divergence-maximising losses move toward the high-ICS, low-MMLU region. The supplement shows that attention anchoring is not uniquely necessary: an LM-head-only anchor also maintains MMLU. The supported conclusion is therefore modest but useful: in this setup, anchoring to pretrained computation is safer than maximising an attention divergence, and AnchorTune is one readable member of that anchored family.
Answer to RQ3.
Partially, and more so for anchoring than for gating. Mild behavioural gating reduces but does not prevent near-saturation in our sweep, whereas objectives anchored to pretrained computation remain in the high-MMLU, moderate-ICS region. Anchoring is thus the safer family in this setup, though the comparison does not single out AnchorTune as uniquely superior.
5 Discussion and Limitations
The result is a warning about using attention divergences as training targets, not a claim that attention-level probes are useless. Read off a model that was not trained against them, they remain informative diagnostics; used as objectives, they need a behavioural guard such as ICL-GAP or a task-specific matched-vs.-mismatched accuracy gap.
Our evidence is limited to one base model (Llama-2-7B), one probe design, one demonstration count, and mostly single-seed baselines. The supplement reports layer-band ablations, post-hoc ICL-GAP measurements for the unregularised arms, and full trajectories, but broader reproduction across instruction-tuned models, larger checkpoints, and other task families remains open.
6 Conclusion
An attention diagnostic can improve under fine-tuning without preserving the behaviour it is meant to measure. The behavioural validation confirms that the model can use matched demonstrations, yet an ICS-maximising regulariser drives the proxy to within of its geometric ceiling while the controlled QA ICL-GAP remains near zero and MMLU falls from to . Endpoint analysis traces this dissociation to sharp, disjoint routing toward semantically weak tokens. Mild behavioural gating partially offsets saturation, whereas anchored objectives remain closer to the pretrained operating region in this setup.
The practical lesson is not that attention probes should be discarded. Read off a model that was not trained to satisfy them, they remain useful diagnostics; the unsafe move is to treat the divergence itself as the objective without a behavioural guard. Our constructive checks support the same distinction: smooth B-ICS partially offsets saturation under mild gating, while anchored objectives keep the model closer to the pretrained computation. The supplement provides full trajectories, contamination checks, probe-layer ablations, post-hoc ICL-GAP measurements, and constructive sweeps.
References
- What learning algorithm is in-context learning? investigations with linear models. In International Conference on Learning Representations (ICLR), Cited by: §1, §2.
- Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §1, §2.
- Probing memorization of tabular in-context learning. arXiv preprint arXiv:2606.31208. External Links: Link Cited by: §2.
- Calibration of pre-trained transformers. In Conference on Empirical Methods in Natural Language Processing (EMNLP), Cited by: §2.
- A mathematical framework for transformer circuits. Transformer Circuits Thread. External Links: Link Cited by: §2.
- Reward shaping to mitigate reward hacking in RLHF. arXiv preprint arXiv:2502.18770. External Links: Link Cited by: §2.
- Scaling laws for reward model overoptimization. In International Conference on Machine Learning (ICML), Cited by: §2.
- Problems of monetary management: the UK experience. Monetary Theory and Practice, pp. 91–121. Cited by: §2.
- MiniLLM: knowledge distillation of large language models. arXiv preprint arXiv:2306.08543. Cited by: §2.
- On calibration of modern neural networks. In International Conference on Machine Learning (ICML), Cited by: §2.
- Towards a unified view of parameter-efficient transfer learning. In International Conference on Learning Representations (ICLR), Cited by: §2.
- Measuring massive multitask language understanding. In International Conference on Learning Representations (ICLR), Cited by: §2.
- Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning (ICML), Cited by: §2.
- Fine-tuning dynamics of in-context factual recall in transformers. arXiv preprint arXiv:2605.27774. External Links: Link Cited by: §2.
- TinyBERT: distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP, Cited by: §1, §2, §2.
- Inference-time reward hacking in large language models. arXiv preprint arXiv:2506.19248. External Links: Link Cited by: §2.
- Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences (PNAS) 114 (13), pp. 3521–3526. Cited by: 2nd item, §2.
- Large language models hack rewards, and society. arXiv preprint arXiv:2606.04075. External Links: Link Cited by: §2.
- Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), Cited by: Appendix F.
- An empirical study of catastrophic forgetting in large language models during continual fine-tuning. arXiv preprint arXiv:2308.08747. Cited by: §1, §2, §4.3.
- Categorizing variants of goodhart’s law. In arXiv preprint arXiv:1803.04585, Cited by: §2.
- Obtaining well calibrated probabilities using Bayesian binning. In AAAI Conference on Artificial Intelligence, Cited by: §2.
- In-context learning and induction heads. Transformer Circuits Thread. External Links: Link Cited by: §1, §2.
- Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: 1st item, §2.
- Categorical prior lock-in: why in-context learning fails for structured data. arXiv preprint arXiv:2606.11961. External Links: Link Cited by: §2.
- Continual learning of large language models: a comprehensive survey. arXiv preprint arXiv:2404.16789. Cited by: §1, §2.
- Defining and characterizing reward hacking. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §2.
- Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. Cited by: §4.1.
- Don’t go breaking my LLM: the impact of pruning attention layers on explanation faithfulness and confidence calibration. arXiv preprint arXiv:2606.24970. External Links: Link Cited by: §1, §2.
- Transformers learn in-context by gradient descent. In International Conference on Machine Learning (ICML), Cited by: §2.
- Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In International Conference on Learning Representations (ICLR), Cited by: §2.
- Large language models are latent variable models: explaining and finding good demonstrations for in-context learning. arXiv preprint arXiv:2301.11916. Cited by: §1, §2.
- Emergent abilities of large language models. Transactions on Machine Learning Research. Cited by: §1, §2.
- An explanation of in-context learning as implicit bayesian inference. In International Conference on Learning Representations (ICLR), Cited by: §2.
- ADAPT: attention dynamics alignment with preference tuning for faithful MLLMs. arXiv preprint arXiv:2606.31054. External Links: Link Cited by: §1, §2.
- Continual learning through synaptic intelligence. In International Conference on Machine Learning (ICML), Cited by: §2.
Appendix A Full Per-Step Trajectories
Table 3 reports ICS, MMLU accuracy, and ECE every steps for all four arms. ICL-GAP is reported only for F_ICS-Max during training; post-hoc ICL-GAP for the other arms at step is in Appendix D.
| F_A | F_M | F_None | F_ICS-Max | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Step | ICS | MMLU | ECE | ICS | MMLU | ECE | ICS | MMLU | ECE | ICS | GAP | MMLU | ECE |
| 100 | 0.474 | 0.375 | 0.158 | 0.470 | 0.375 | 0.160 | 0.474 | 0.354 | 0.147 | 0.513 | 0.371 | 0.136 | |
| 500 | 0.485 | 0.367 | 0.220 | 0.480 | 0.354 | 0.230 | 0.485 | 0.371 | 0.215 | 1.225 | 0.367 | 0.274 | |
| 1000 | 0.487 | 0.346 | 0.310 | 0.481 | 0.342 | 0.305 | 0.490 | 0.371 | 0.305 | 1.292 | 0.304 | 0.473 | |
| 1500 | 0.488 | 0.342 | 0.360 | 0.481 | 0.342 | 0.350 | 0.493 | 0.371 | 0.350 | 1.309 | 0.333 | 0.429 | |
| 2000 | 0.490 | 0.338 | 0.395 | 0.482 | 0.342 | 0.380 | 0.495 | 0.371 | 0.380 | 1.319 | 0.317 | 0.452 | |
| 2500 | 0.491 | 0.338 | 0.410 | 0.482 | 0.342 | 0.390 | 0.497 | 0.371 | 0.395 | 1.385 | 0.254 | 0.195 | |
| 3000 | 0.491 | 0.338 | 0.420 | 0.482 | 0.342 | 0.395 | 0.498 | 0.375 | 0.400 | 1.386 | 0.271 | 0.159 | |
| 3500 | 0.492 | 0.338 | 0.425 | 0.482 | 0.342 | 0.400 | 0.499 | 0.375 | 0.402 | 1.301 | 0.283 | 0.223 | |
| 4000 | 0.492 | 0.338 | 0.430 | 0.482 | 0.342 | 0.403 | 0.500 | 0.375 | 0.405 | 1.385 | 0.271 | 0.240 | |
| 4500 | 0.492 | 0.338 | 0.432 | 0.482 | 0.342 | 0.405 | 0.500 | 0.375 | 0.406 | 1.386 | 0.283 | 0.235 | |
| 5000 | 0.492 | 0.338 | 0.434 | 0.482 | 0.342 | 0.406 | 0.500 | 0.375 | 0.407 | 1.413 | 0.279 | 0.231 | |
Appendix B Contamination Postmortem (F_ICS-Maxv1)
The first F_ICS-Max run, F_ICS-Maxv1, used a single ICLEvalDataLoader for both the regulariser pair set and the ICS evaluation pair set. The regulariser was measured on this shared set, so the gradient pushed the model to maximise the eval metric on the eval data. We caught this at training step , when ICS exceeded on the held-out probe and continued to climb. After the fix (separate loaders, seeds and , no shared pairs), the rerun (F_ICS-Maxv2) reproduced the saturation pattern but at a slower rate and on genuinely held-out data. All numbers in the main body are from F_ICS-Maxv2. The bug-fix list is:
- 1.
Split ICLEvalDataLoader into icl_train_loader (regulariser) and icl_eval_loader (ICS, ICL-GAP).
- 2.
Add answer_idx to ICLEvalDataLoader.__next__, which had been silently dropped and prevented ICL-GAP from being computed.
- 3.
Add evaluate_icl_gap to the calibration utility module and wire it into the trainer eval loop.
- 4.
Add icl_train_loader as a constructor argument on Trainer (after the existing arm argument, to preserve compatibility).
Appendix C Probe-Layer Ablation
We re-evaluated ICS at the F_ICS-Maxv2 step- checkpoint with three alternative probe bands: (early), (mid-late), and (late). The resulting ICS values are , , and respectively. All three bands are within of the original and within of the ceiling, confirming that the saturation is not a probe-band artefact.
Appendix D Post-hoc ICL-GAP for Unregularised Arms
We loaded the step- checkpoints for F_A, F_M, and F_None and evaluated ICL-GAP on the same eval ICL set used for F_ICS-Maxv2, with the same temperature- multiple-choice protocol. The values are:
| Arm | |||
|---|---|---|---|
| F_A | 0.330 | 0.305 | |
| F_M | 0.295 | 0.290 | |
| F_None | 0.355 | 0.320 |
All three are within of zero. Thus, the near-zero behavioural gap is shared by the unregularised checkpoints under this model and probe; what distinguishes F_ICS-Max is its ICS, not its ICL-GAP.
Appendix E Coarse Collapse Pilot
The pilot used a domain-specific instruction set rather than the broad mixture in the controlled proxy-stress experiment, and a peak learning rate of rather than . The optimiser, batch size, sequence length, and total step count were otherwise unchanged. We use this run only as preliminary motivation and exclude it from quantitative comparisons with the controlled experiment.
Appendix F Controlled Proxy-Stress Training Details
All controlled-ablation and constructive arms use AdamW [Loshchilov and Hutter, 2019], peak learning rate , warmup steps, cosine decay, batch size , sequence length , gradient clipping at norm , and bf16 mixed precision. Checkpoint probes are evaluated every optimiser steps; Table 3 reports the coarser -step trajectory for readability.
Appendix G Probe Configurations
- •
ICS / ICL-GAP eval set. multiple-choice pairs, seed . Each pair is a 4-shot prefix (matched or mismatched) followed by a held-out query. Disjoint from the regulariser set.
- •
Regulariser set. pairs, seed . Same format. Used only by in F_ICS-Max.
- •
MMLU probe. items stratified across the MMLU subjects, -shot prefix, temperature .
- •
GSM8K probe. items, -shot chain-of-thought prefix, temperature .
- •
ECE. equal-mass bins on MMLU.
Appendix H Pretrained Behavioural ICL Screening
The main paper reports the A/B sanity check showing that the pretrained Llama-2-7B checkpoint has a positive behavioural ICL effect on random-label binary episodes. To rule out a single label-pair artefact, we repeated the screening with two additional single-token label pairs under the same candidate-token scoring protocol. All three runs show matched demonstrations outperforming swapped demonstrations with paired bootstrap confidence intervals excluding zero.
| Labels | Match | Perm | None | Gap [95% CI] | |
|---|---|---|---|---|---|
| A/B | 480 | 0.669 | 0.362 | 0.492 | 0.306 [0.231, 0.379] |
| X/Y | 240 | 0.762 | 0.225 | 0.504 | 0.537 [0.450, 0.621] |
| foo/bar | 240 | 0.642 | 0.338 | 0.496 | 0.304 [0.200, 0.400] |
Appendix I B-ICS: Soft-Gap Surrogate
The hard B-ICS, , requires the discrete ICL-GAP, which is non-differentiable. We use the soft-gap
| (21) |
This substitution is a smooth heuristic rather than a theorem equating log-probability gaps with accuracy gaps.
(i) Heuristic alignment. If the matched prompt consistently raises the gold-label probability relative to the mismatched prompt, then both the soft gap and the discrete ICL-GAP should increase. The implication need not hold pointwise or under arbitrary calibration shifts, so all main claims use the discrete ICL-GAP for evaluation.
(ii) Bounded magnitude. Empirically, across the -step trajectory of all arms. The choice is comparable to the single-checkpoint sampling scale of the discrete ICL-GAP ( for the logged paired examples discussed in the main paper), and the gate stays away from -saturation.
Appendix J B-ICS Gradient Calculation
Proposition 3 (B-ICS suppresses the bad ridge).
On a high-ICS point with , the ICS-only component of the smooth B-ICS gradient is multiplied by at most .
By definition where .
| (22) | ||||
On the bad ridge , by definition, so the argument of is at most , and
| (23) | ||||
| (24) |
Substituting back,
| (25) | ||||
where we used . For , ; for and , . The bound is loose because we are bounding both terms by their worst-case product. The substantive content is that the first (ICS-only) gradient component is multiplied by on the bad ridge, while the second (gap-only) component is bounded; among the two, the gap-only component selects directions that increase , i.e. leave the ridge. Hence gradient descent on from either does not move toward the ridge in the first place or, if it does, is pushed back out as falls below .
Appendix K AnchorTune: Memory and Compute
The reference-model cost is the dominant memory overhead of AnchorTune. We instantiate it as follows.
Memory.
Llama-2-7B in bf16 occupies . The pretrained reference is loaded once and kept frozen (no gradients, no optimiser state), so its footprint is exactly . The trainee adds a further of weights, of fp32 gradients, and of fp32 AdamW optimiser state, for a total of before activations. The cached anchor attentions are where is the per-pair token length. With , , , the cache is . We pre-batch the cache into batches of pairs each on GPU to amortise the per-step lookup; batches are kept in fp32 to avoid quantisation jitter feeding back into the trainee gradient. Total GPU memory at peak (forward+backward+anchor cache+activations) is on our device.
Compute.
AnchorTune adds one forward pass through the trainee on the anchor probe per optimiser step (one batch from the -batch cache, same cost as one ICL probe forward in F_ICS-Max). The reference model is not re-evaluated during training; only the cached attention rows are used. The wall-clock per optimiser step on our hardware is (vs. for F_None and for F_ICS-Max), giving a overhead vs. unregularised fine-tuning and parity with F_ICS-Max. A complete -step run takes , plus to load the reference and pre-compute the anchor cache.
Reference-model alternatives.
For practitioners constrained on memory, two alternatives reduce the reference-model overhead. (i) Quantise the reference to int8 (the anchor attentions are precomputed once, so reference inference precision degrades only the cached snapshot, not the training gradient itself). (ii) Replace the in-memory reference with a CPU-resident reference and a one-shot disk cache of attentions; this raises the start-up cost from minutes to minutes on our hardware but eliminates the GPU footprint of the reference. We use the in-memory variant for the experiments reported in this paper.
Appendix L Constructive Results: AnchorTune and B-ICS
This section tests two responses to the observed dissociation. Section L.1 examines whether the behaviour-anchored target attenuates Goodhart saturation; Section L.2 reports the AnchorTune -sweep and multi-seed final; Section L.3 compares against two competitive baselines on the (MMLU, ICL-GAP) Pareto frontier; and Section L.4 ablates the anchor target.
L.1 Behaviour-Gated Sweep
We rerun the optimisation-pressure experiment with the same compute budget, the same probe ICL set, and the same hyperparameters as F_ICS-Maxv2, but replace with at . By Proposition 3, down-weights the ICS-only gradient on the disjoint-support ridge, but the factor is still substantial at our mild setting , and the smooth surrogate of (21) can rise transiently above during training even when the discrete does not.
The empirical trajectory clarifies the practical strength of the gating effect. Table 5 reports the step- values. F_BICS ends at , only below the stress-test F_ICS-Max value of and within of the ceiling. MMLU accuracy drops to , essentially the same degradation as F_ICS-Max. What does differentiate the two arms is the ICL-GAP: F_BICS finishes at vs. F_ICS-Max’s , and tracking its trajectory shows the soft gap exceeds on intervals where the behavioural sentinel intermittently activates, before the divergence term overwhelms it. This indicates that B-ICS as a training target requires a more aggressive gating schedule (larger , larger , or a hard gate with a straight-through estimator) than the smooth we chose for differentiability; the observed behavioural advantage of F_BICS over F_ICS-Max is consistent with the attenuation predicted by Proposition 3 under that mild gating.
When evaluated on a model that was not trained against it, B-ICS retains the intended behavioural guard. It multiplies ICS by a binary indicator that the behavioural sentinel passes and returns near-zero for F_ICS-Max (which has ) and near-pretrained for F_anchor at . The diagnostic version therefore keeps the intended behavioural guard, while the training-target version requires stronger gating than we use here.
| Arm | ICS | ICL-GAP | MMLU | ECE |
|---|---|---|---|---|
| Pretrained | 0.516 | — | 0.371 | — |
| F_ICS-Max (stress test) | 1.413 | 0.279 | 0.231 | |
| F_BICS (this work) | 1.397 | 0.283 | 0.388 |
L.2 AnchorTune Sweep
For AnchorTune we run a four-point -sweep at seed , then a three-seed final at the chosen . Table 6 reports the final-step values; Figure 6 plots the per-step trajectories of the four core quantities for F_anchor at alongside F_None and F_ICS-Max.
The empirical claim is that F_anchor at closes the MMLU damage that F_ICS-Max inflicts: from (chance ) back to within seed-noise of the pretrained baseline of , while keeping ICS indistinguishable from pretrained. Concretely, across three seeds at , MMLU returns (mean std), ICS returns , and ICL-GAP returns — all within seed-variation of the pretrained model. The MMLU recovery accounts for of the accuracy drop that F_ICS-Max produced. Of the four values, is the empirical optimum (MMLU ); both and converge to a slightly lower MMLU plateau of , suggesting that the anchor pulls the trainee too tightly toward once exceeds , and is too weak to prevent the small MMLU drift seen for the unregularised arm.
| Configuration | ICS | ICL-GAP | MMLU | ECE |
| Pretrained | 0.516 | — | 0.371 | — |
| F_None (controlled) | 0.500 | 0.375 | 0.407 | |
| F_ICS-Max (stress test) | 1.413 | 0.279 | 0.231 | |
| F_anchor, | 0.501 | 0.338 | 0.472 | |
| F_anchor, | 0.500 | 0.367 | 0.375 | |
| F_anchor, | 0.500 | 0.350 | 0.374 | |
| F_anchor, | 0.501 | 0.350 | 0.391 | |
| F_anchor, , seed 42 | 0.500 | 0.367 | 0.375 | |
| F_anchor, , seed 43 | 0.501 | 0.429 | 0.384 | |
| F_anchor, , seed 44 | 0.500 | 0.313 | 0.387 | |
| F_anchor, , mean std | 0.500 0.001 | 0.370 0.058 | 0.382 0.006 |
L.3 Anchored-Baseline Comparison
We compare F_anchor with two capability-preservation baselines, all evaluated with the same training and evaluation protocol:
- •
LogitAnchor: a logit-level KL-to-base regulariser of the form widely used in RLHF [Ouyang et al., 2022], . This is the closest existing analogue of AnchorTune, but it anchors predictions rather than attention rows.
- •
WeightAnchor: an EWC-style [Kirkpatrick et al., 2017] weight anchor , with .
Table 7 reports the step- values.
The empirical pattern is more nuanced than “AnchorTune dominates the baselines.” In this single-seed comparison, all three anchored arms (F_anchor, LogitAnchor, WeightAnchor) recover MMLU to within seed noise of pretrained (, , and respectively at seed , vs. pretrained and F_None’s ), and all three keep ICS at the pretrained range. The three arms sit on the same Pareto front; what distinguishes them is what they anchor and therefore what they expose. LogitAnchor ties the trainee’s predictions to , which maintains task accuracy but masks the internal attention dynamics; WeightAnchor ties the trainee’s full weight vector to , which maintains both but at the cost of any meaningful fine-tuning signal at the we tested; F_anchor ties the trainee’s mid-band attention rows to , which maintains the diagnostic interpretability of ICS as a measure of attention-level ICL drift. In this setup, the three objectives are therefore complementary: AnchorTune’s contribution is not primarily a larger MMLU number but an objective whose anchor signal and diagnostic signal agree. This alignment allows ICS to be interpreted relative to the pretrained attention behaviour rather than a logit-constrained output alone. The contrast with F_ICS-Max is the main result: the divergence-based regularisers we test (ICS and smooth B-ICS as targets) move toward the high-ICS, low-MMLU region, while the anchored regularisers we test (logit, weight, attention) maintain MMLU. This comparison identifies a structural distinction between divergence and anchor losses, not a critique of attention-level signals per se.
| Arm | ICS | ICL-GAP | MMLU | ECE |
|---|---|---|---|---|
| Pretrained | 0.516 | — | 0.371 | — |
| F_None | 0.500 | 0.375 | 0.407 | |
| F_ICS-Max | 1.413 | 0.279 | 0.231 | |
| LogitAnchor (KL-to-logits) | 0.494 | 0.383 | 0.373 | |
| WeightAnchor (weight , ) | 0.500 | 0.363 | 0.419 | |
| F_anchor (ours, ) | 0.500 | 0.367 | 0.375 |
L.4 Anchor-Target Ablation
To isolate the active ingredient of AnchorTune, we vary which part of the model is anchored while keeping and fixed:
- •
F_anchor_last4: anchor attention rows only on the last layers, not on the mid-band .
- •
F_anchor_lmhead: anchor only the LM head weights via an penalty toward ; do not anchor any attention rows.
The hypothesis is that mid-band attention-row anchoring is what carries the MMLU-preservation effect. If F_anchor_last4 and F_anchor_lmhead both drop MMLU back toward F_None levels while F_anchor does not, the attribution to the mid-band attention rows is established.
The results in Table 8 weaken the original mid-band attribution hypothesis. Restricting the anchor to the last 4 layers (F_anchor_last4) gives vs. full-layer F_anchor’s , only a -point reduction — so the mid-band attention rows are sufficient but not necessary for MMLU preservation; the last 4 layers alone come within seed noise of the full anchor. More strikingly, the LM-head-only variant (F_anchor_lmhead, no attention anchoring at all) achieves , matching the pretrained reference of to within seed noise and exceeding seed-42 F_anchor. This means keeping close to in weight space is by itself enough to recover MMLU under our fine-tuning configuration — attention-row anchoring is one route to MMLU preservation, but not the only one. These results indicate that the mechanism behind F_anchor’s MMLU recovery is partially shared with WeightAnchor (Table 7) and F_anchor_lmhead: proximity to in any structural component (weights, lm_head, attention rows) is sufficient. AnchorTune’s residual advantage over these baselines is not absolute MMLU but the coincidence of its anchor target with the diagnostic ICL signal, as discussed in Section L.3.
| Anchor target | ICS | ICL-GAP | MMLU | ECE |
| Pretrained | 0.516 | — | 0.371 | — |
| None (= F_None, controlled reference) | 0.500 | 0.375 | 0.407 | |
| Mid-band attn rows (F_anchor) | 0.500 | 0.367 | 0.375 | |
| Last 4 layers attn rows (F_anchor_last4) | 0.510 | 0.358 | 0.377 | |
| LM head only (F_anchor_lmhead) | 0.494 | 0.375 | 0.378 |
L.5 Summary of the constructive part
The constructive experiments support three claims. (i) The Goodhart channel is not an artefact of using attention as a target; it is an artefact of the divergence-on-disjoint-supports geometry. B-ICS as a diagnostic inherits the safety motivation of Proposition 3; B-ICS as a training target only partially attenuates the channel under the smooth gating we chose (), and reaches at convergence (vs. for F_ICS-Max). A hard gate or a larger remains to be tested. (ii) Anchoring the per-head attention rows to the pretrained model on a fixed probe maintains MMLU under exactly the same fine-tuning configuration where F_ICS-Max produces proxy saturation and MMLU damage: three-seed MMLU at is , returning of the MMLU gap. (iii) The Pareto contrast in our tested arms is structural: anchored regularisers (attention rows in F_anchor, logits in LogitAnchor, weights in WeightAnchor) maintain MMLU, whereas the divergence targets we test do not. AnchorTune’s specific contribution among anchored regularisers is that its anchor signal coincides with the diagnostic ICL signal, so that the same probe can be read out at training and at evaluation without a Goodhart concern.