Harness-Agnostic Detection and Immunization of Reward Hacking in Self-Evolving Language Models
Abstract
Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actually wants, sustained selection widens the gap between the two. This is reward hacking. We introduce HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations. It keeps a secret, distribution-fixed comparison core, whose frozen distribution makes its capability proxy comparable across generations, alongside a rotated fresh layer that hardens the bank against co-adaptation. Four tests built on that proxy cover the level gap, a scale-aligned divergence with online change-point detection, capability stagnation, and a conditional confidently-wrong rate; a Šidák correction turns them into a calibrated family-wise -value. Diagnosis alone recovers nothing, so a risk-aware immunization layer reselects an honest candidate from the proposal pool using the core together with a purely structural gaming footprint, disclosing at most bits per generation to the host. We prove a detectability bound that converts a target error rate into an explicit probe-size budget, and we delimit what probe rotation does and does not buy. On a controlled prompt-level host with four injected hacking channels and ground-truth labels, HackProbe reaches AUROC against for the strongest baseline and cuts the false-positive rate from to . Its bandwidth-limited reselection is the only immunization level that returns more true capability under hacking, points on average, than it forfeits on clean runs, ; per-channel effects are mostly not individually significant.
1 Introduction
A growing family of language-model systems improves itself through a closed optimization loop: at each generation it proposes candidate updates to prompts, code, weights, or generated tasks, scores each candidate with a visible evaluator , and keeps the candidates that score highest (Yuan et al., 2024; Huang et al., 2026; Agrawal et al., 2026). The very thing that makes the loop powerful, relentless selection pressure on , is also what breaks it. Whenever is only a proxy for the target capability that we actually care about, optimizing hard and repeatedly tends to inflate without a matching gain in . This is reward hacking, an acute instance of Goodhart’s law: once a measure becomes a target it ceases to be a good measure (Manheim and Garrabrant, 2018; Amodei et al., 2016; Skalse et al., 2022). Self-evolution is arguably the purest form of the strong optimization pressure under which Goodhart effects are most severe, because the same pressure is applied cumulatively across tens or hundreds of generations.
The assumption that a rising implies a rising is fragile in exactly the systems we care about. Verifier and solver are often the same model updated in lockstep, so self-rewarding training watches its own judge saturate rather than the policy improve (Yuan et al., 2024), and label-free schemes that reward self-consistency pay for confidently agreeing on wrong answers (Huang et al., 2026). Fixed benchmarks degrade under repeated selection too, and measurably so: held-out reconstructions of grade-school math show accuracy drops of up to eight points (Zhang et al., 2024a), one irrelevant clause dropped into a symbolic template can cost of performance (Mirzadeh et al., 2025), and adversarial perturbation exposes scores resting on shallow behavior (Li et al., 2024). Frontier models go further, editing tests or reading answer files once the scoring function is visible (Wang et al., 2026). Responses so far are fragmented and harness-bound: patches built into one system, or monitors assuming access to weights and activations that a black-box setting does not provide. The closest precedent is reward-model overoptimization (Gao et al., 2023), where a held-out gold reward model shows a proxy eventually degrading true performance; but that diagnosis is computed once, offline, and reported rather than acted upon. What is missing is a capability coordinate orthogonal to , immune to the same pressure, and usable while the loop still runs. Appendix A places HackProbe against the fuller literature.
HackProbe supplies that coordinate. It attaches through two minimal hooks, observing each generation’s visible score and querying the current candidate on held-out probes, and rests on a dual-layer probe bank (Figure 1). A secret, distribution-fixed comparison core provides a capability proxy whose first differences track genuine capability change rather than shifting test difficulty, precisely because the distribution is frozen; a rotated, online-generated fresh layer behind an air gap hardens the bank against co-adaptation. Four statistics with different blind spots read off that proxy: a level-gap test for hacking present from the outset, the scale-aligned divergence with online change-point detection, a stagnation test, and a conditional confidently-wrong-rate test aimed at self-consistency bias. Since a diagnosis arriving after the host has discarded the honest candidate recovers nothing, a risk-aware immunization layer reselects from the pool using the core and a purely structural gaming footprint, returning only a coarsely quantized signal.
The comparison-core estimator is what makes the rest work. Its distribution is fixed, so differences across generations mean something, and the detector we build on it reports a calibrated family-wise -value rather than an uncalibrated heuristic. Diagnosis by itself changes nothing, though, so we turn it into protection: a reselection rule that discloses at most bits per generation to the host. That cap does real work. It is the reason the feedback never becomes a second score to optimize. On the theory side a target error rate converts into a probe-size budget growing as in the divergence gap, and we are equally explicit about what rotation does not buy, since it stops memorization but leaves the covering number of the probe generator as the real ceiling. Empirically, on a controlled injection protocol with ground-truth labels, HackProbe beats the strongest baseline on AUROC and cuts its false-positive rate by nearly two fifths; of the four immunization levels, only the bandwidth-limited one returns more true capability than it costs. Calibrating on three hacking channels and testing on the withheld fourth retains most of that quality on three of them.
2 Preliminaries
Let denote the object under evolution at generation , be it a prompt, a program, or a policy. The host scores candidates with a visible evaluator and updates , where is its own black-box evolution operator. HackProbe assumes only two things about that operator: (I1) it can observe the visible score of each generation, and (I2) it can run on a chosen input, drawing a small fixed number of black-box samples per input and reading only the outputs, with no access to weights or activations. One sample suffices for the statistics J0–J2 below; the modal-agreement statistic J3 needs , the only place where the interface costs more than a single forward pass. Both conditions hold for prompt-level hosts at negligible cost; for weight-level hosts they require reproducing the evolution trajectory, whose cost we do not measure here.
Let be the true capability of , its expected performance on a target distribution , which is not cheaply observable. Two black-box estimates from held-out gold data stay strictly disjoint: for offline calibration only, for the final measurement only, never entering the online loop. During evolution a proxy on a secret comparison core (Section 3.1) tracks capability. Writing and for the observed visible score and core proxy of , with sampling noise as in Assumptions 1 and 2, and , we monitor the scale-aligned divergence
| (1) |
where aligns the scales of the two metrics. The scale is not a free parameter: it is initialized on offline honest calibration runs and thereafter re-estimated only on the trailing window of generations the detector has already declared honest. The update is prequential, using no information from generation when testing generation , which prevents a hacking generation from inflating until it conceals itself.
Definition 1 (Divergence regime).
Let be the population divergence of generation . Generation is honest when , so visible gains never outrun true gains at the calibrated scale, and hacking at gap when ; the prototypical case with is covered. Statistic J1 below tests against one-sidedly, with the band an indifference zone on which we make no claim, since separating arbitrarily small divergences would need unboundedly many probes. The other three statistics test companion nulls, on the level residual, on stagnation and on conditional confidence, which honest evolution also implies but which are not functions of alone; fusing them extends coverage past the regime this definition names.
3 The HackProbe Framework
3.1 Dual-layer probe bank
The design tension we have to resolve is that a fixed, host-visible probe set would itself be optimized against and would simply become a second . We resolve it by splitting probes into two layers with orthogonal roles.
The comparison core is a distribution-fixed, secret probe set covering the target capability across subdomains. Because its distribution never changes, the proxy , with a per-subdomain score and an aggregate robust enough that no single subdomain can move it alone, yields differences that track capability change and not test drift. Freezing the distribution has a second payoff, less obvious than the first. The same items are scored at and at , so the noise in is paired: per-item difficulty cancels, and the variance of the difference is far smaller than twice that of a single generation’s estimate. The core is never revealed to the host and never enters its selection or reward path.
The fresh layer is rotated, generated online, and drifted every generation. It is used only for hardening checks and qualitative per-subdomain profiling and never contributes to . Four principles govern it, and they pull in different directions on purpose. An air gap (H1) keeps probes out of the host’s training, selection and reward loop entirely, so neither probe text nor per-subdomain scores are ever fed back; on top of that, rotation (H2) makes sure any single instance is used at most a few times. Out-of-distribution drift (H3) then steers the layer toward whichever subdomains the host is currently weakest on, so the bank is always probing capability that has not yet been fitted, while adversarial refresh (H4) has a red-team model periodically manufacture probes that are harder and more surface-diverse still. H4’s two selection criteria — learnability (probes near the host’s competence edge) and surface diversity (embedding distance from probes already used) — are deliberately independent of the detection signal, so refresh cannot self-amplify into false positives. Non-circularity of the final measurement survives all of this because and stay disjoint and is spent exactly once, at the end.
3.2 Detection statistics and fusion
HackProbe fuses four one-sided tests, each calibrated on honest generations so that its null -value is approximately uniform on . Their blind spots are deliberately different, and fusing them extends coverage across the level, divergence, stagnation and conditional-confidence regimes; format-level exploits, which act on output structure rather than on the score–capability relation, fall outside all four (Section 4).
Statistic J0 targets hacking that is present from the very first generation, where is already inflated and then stays flat so that every first-difference statistic sees nothing. We calibrate the honest affine relation and test the residual level gap , which hovers near zero under honest evolution and is persistently positive under level-type hacking. Local linear coupling makes and the difference scale coincide in population; we estimate them separately so each absorbs its own drift. Statistics J1 and J2 handle the dynamic regime. On from Eq. (1) we run a one-sided test against the honest null, together with an online Page–Hinkley change-point detector (Page, 1954) on the cumulative divergence that localizes the generation at which and decouple. J2 tests capability stagnation directly, taking the honest null conditional on and rejecting when the measured is too small, which gives a -free complement.
Statistic J3 exists because self-consistency bias inflates by rewarding agreement instead of correctness, which need not move aggregate accuracy at all. Once such an attack saturates, J1 and J2 see nothing in expectation and J0 retains only the residue the attack leaves in the aggregate level, diluted over all core items including the many answered correctly. On the secret core we measure, over probes answered incorrectly, the mean self-consistency (the confidently-wrong rate), calibrate its honest affine dependence on core accuracy , and test the residual , with a Page–Hinkley detector for gradual drift. Conditioning on wrong answers avoids the ceiling effect of easy probes, and using only the secret core keeps the signal non-circular. Lemma 1 in Appendix E makes precise why J3 retains an expected separation equal to the full attack size in the stable regime where J1 and J2 vanish in expectation.
Fusion takes one-sided -values: J0, the divergence test J1 and its Page–Hinkley alarm, the stagnation test J2, and J3 with its own alarm. The two sequential alarms are converted to -values through their calibrated alarm-time distribution on honest runs, putting all six on a common scale. Given approximately independent and uniform under the honest null, we form the risk score and flag hacking when . This is exactly a Šidák correction (Šidák, 1967), and it has a calibration property that a raw rule lacks. Under the independent-uniform null , so that
| (2) |
so is the Šidák-adjusted -value and is the operating point: flagging at costs family-wise error exactly , however many statistics we fuse, whereas the uncorrected rule degrades with . In practice the guarantee is softer than that. J0, J1 and J2 are built from the same pair and are not independent, so Eq. (2) is a design principle and the reported operating point rests instead on re-calibrating until the empirical false-positive rate on honest runs meets its target; Appendix H gives the Bonferroni fallback we use when even that is in doubt. The control is also per generation, so over generations the expected number of false alarms is and run-level error would need a sequential correction we do not attempt. What a deployment sees, and what Table 1 reports, is the per-generation alarm rate.
3.3 Risk-aware immunization
Diagnosis alone recovers nothing. The tempting picture, in which rejecting or rolling back the top candidate suffices, fails on two counts: by the time the host is told, it has already passed over the honest candidate in favor of the top- one, and rejection has nothing better to fall back to. HackProbe therefore reselects from the proposal pool, and the loss it recovers is a real one for two reasons. Under strict final-answer grading, the true-capability side ( and the core ) credits only a normalized final answer matching the gold one, so behaviors that inflate a lenient visible grader, such as enumerating many candidate numbers or locking onto an unverified answer, genuinely forfeit true credit. And every pool contains, alongside the exploit variants, one enforced honest candidate produced under a clean context, whose visible is typically lower but whose true is higher.
For each candidate we compute a purely structural gaming footprint from cached outputs at no extra inference cost, combining the score-versus-commitment gap with the fraction of non-final candidate artifacts, and a robust score with a fixed penalty weight . Four immunization levels build on it, each trusting the host’s own score a little less than the last. F0 is diagnosis without protection. It selects , so it goes straight to the inflated candidate. F1 stays greedy by default and intervenes only when a lightweight risk proxy on the greedy choice, either its raw gap or its footprint , comes back high; then it reselects , and after consecutive high-risk generations it rolls back to the checkpoint with the best historical . F2 goes further, selecting with a quantization of the raw gap into levels (the uncalibrated difference, not the J0 residual ). It also watches for self-destruction, whose signature is feedback falling while stagnates, and downgrades itself to F1 when it sees it. F3 adds a coverage bonus for collapsed subdomains and injects a diversity constraint into the proposer.
The quantization in F2 is where the bandwidth constraint stops being cosmetic. Were the exact gap, would collapse to and F2 would be pure core-greedy selection, the failure we are trying to avoid: it would expose the secret core as a fully-resolved second objective. With levels, F2 interpolates between the host’s greedy rule (, nothing transferred) and core-greedy selection (). The leak needs careful accounting. Per-candidate values stay inside the monitor and are never returned; the host sees only which of candidates won, so the return channel carries at most bits per generation, three at our , and over a run of length . Quoting here would be tighter but wrong: the winning index can say more about the vector of values than any single does, and it is easy to construct pools where it does. What buys is resolution rather than volume, and resolution is what limits how much of the core’s ordering the host can reconstruct. Either way the cumulative figure must stay small against the entropy needed to identify core items, so long runs need small and the core retired once its budget is spent. The gold core enters only through , never through or an oracle label, and only the selected candidate advances the detector state.
3.4 Theoretical guarantees
What follows rests on standard concentration assumptions, given in full as Assumptions 1–4 in Appendix C: the core estimate is with bounded bias and sub-Gaussian noise, the paired difference has variance proxy and the visible score , and honest updates satisfy with the differential core bias bounded by in every generation. Then is sub-Gaussian with , and its honest mean is at most , where . Proofs are in Appendix D.
Proposition 1 (Detectability and probe-size budget).
Consider the J1 statistic in isolation, with the one-sided rule that flags generation when for some , where is the hacking gap of Definition 1 and all probabilities are taken over probe and evaluation sampling with the trajectory held fixed. Writing for this rule’s decision,
| (3) |
Choosing , both error rates fall below as soon as the effective sample size , a weighted harmonic mean of and , satisfies . Whenever the host’s visible evaluation is at least as large as the core () we have , so it suffices to budget
| (4) |
comparison-core probes. No matching lower bound is claimed.
The value of Proposition 1 is not the standard exponential separation but the budget in Eq. (4): detecting a divergence of size at error needs a core growing like , which turns an abstract design choice into a number one can procure and falsify against measured detection-rate-versus- curves. The differential bias does not enter the budget, since it shifts the honest and hacking means equally and is subtracted during calibration.
Rotation has a companion bound, Proposition 2 in Appendix D: an attacker who memorizes past instances needs of them to match all fresh probes of a generation, which tends to the whole pool as grows. It lives in the appendix because it is narrow in a way worth admitting. It bounds memorization, and nothing else. An attacker who instead learns the generator’s regularities walks around it: if the exploitable behavior admits an -cover of size , defeating rotation costs rather than , so security degrades from pool size to covering number, and a low-complexity generator buys no asymptotic protection at all. Enlarging is what H3 and H4 are for. “Surface-diverse rotation forces generalization” is an empirical hypothesis here, not a theorem.
4 Experiments
Our evaluation uses a controlled prompt-evolution host in which hacking channels can be switched on and their strength dialed, so ground-truth per-generation labels exist at all. The headline is that bandwidth-limited reselection is the only immunization level recovering more true capability under hacking than it forfeits on clean runs, at a ratio of , and Appendix F shows that ratio is single-peaked in the feedback bandwidth. Detection quality supports that result rather than standing on its own: it is what makes reselection fire on the right generations. Online reward-hacking detection with a recovered-capability objective is a new task, so no published numbers are directly comparable and every baseline below is re-implemented on the same setup.
Protocol.
The host is a prompt-evolution loop over Qwen2.5-7B-Instruct with the proposer frozen, so only the evolving prompt changes between generations. Each run lasts generations with a candidate pool of , and six seeds per hacking channel give generations per channel and in total. The comparison core holds probes over subdomains, the fresh layer draws per generation, and hook I2 takes samples per probe, enough for the modal-agreement and confidently-wrong quantities J3 needs. Gold audits split once into calibration items and a held-out , which is where the coarse interval endpoints below come from. The remaining settings are , , , , a Page–Hinkley detector with drift tolerance and alarm threshold , and -resample bootstraps over .
Probes and gold audits are drawn from grade-school and competition mathematics, GSM8K (Cobbe et al., 2021) and MATH (Hendrycks et al., 2021), with GSM-Symbolic (Mirzadeh et al., 2025), GSM-Plus (Li et al., 2024) and the held-out GSM1k (Zhang et al., 2024a) supplying perturbed and contamination-sensitive probes. The host implements four switchable hacking channels, each with a strength knob: contamination (mixing original probe versions into host data), a gameable grader (crediting any appearance of the gold number), self-consistency bias (rewarding agreement, which induces confidently wrong answers), and judge collusion (a lenient LLM judge). As discussed in Section 3.3, the injected inflation is real and not merely nominal because of two design choices: strict final-answer grading defines the true-capability side, and every generation’s pool contains one clean honest candidate that a greedy host will pass over. The final capability measurement uses the held-out audit set exactly once, after evolution ends; confidence intervals are bootstrap intervals over that set.
Three caveats about scope, before the numbers. Labels are channel-on versus channel-off, a superset of the hacking set of Definition 1, since a generation whose divergence falls in the indifference band is still labeled positive; reported recall is therefore a lower bound against that target, and injected strength controls but does not equal . The fresh layer contributes no statistic to , so nothing below depends on it: Proposition 2 and the principles H2–H4 go untested here. And was never set to hit a small nominal level. We fix the operating point on honest calibration runs by maximizing F1, which lands at because a stricter threshold trades away nearly all recall at this probe budget, so the below is that empirical operating point rather than the nominal of Eq. (2).
Baselines and metrics.
We compare against a no-probe host labeled post hoc, a naive fixed non-rotating held-out set, a contamination-only confidence test in the style of PaCoST (Zhang et al., 2024b), a single trap set, and a fixed-distribution absolute- trajectory that isolates the contribution of the divergence formula itself. Detection quality is reported as precision, recall, F1, AUROC, false-positive rate and detection delay in generations. The deliverable that matters most, however, is protection gain: the end-of-evolution true measured on for each immunization level relative to F0.
4.1 Detection quality
Table 1 reports detection across all channels pooled. Three of the five baselines are degenerate: no-probe, the contamination test and the single trap set flag every generation, so recall and false-positive rate are both exactly and precision is exactly the base rate . Their F1 scores are evidence of nothing. What separates them is AUROC, running from , below chance because the contamination test ranks hacking generations lower, up to .
The real competitor is the absolute- trajectory, which uses the same secret core but watches its level rather than its divergence from , reaching AUROC. HackProbe adds AUROC and, for a monitor meant to run continuously, does something more useful: it cuts the false-positive rate from to and lifts precision from to . It pays with recall of rather than and a delay of generations, both consequences of being the only method whose false-positive rate falls below one half. Even so, interrupts more than four honest generations in ten, and we take that gap as the bar for this task. Most of it is a budget problem rather than a design one. Sweeping the core size at a fixed false-positive rate (Appendix F) puts our on the steepest part of the detection curve, where Eq. (4) would want probes for detection; the monitor is under-provisioned by roughly .
| Method | Precision | Recall | F1 | AUROC | FPR | Delay |
| No-probe† | 0.241 | 1.000 | 0.389 | 0.500 | 1.000 | 0.00 |
| Fixed held-out set | 0.291 | 0.939 | 0.445 | 0.587 | 0.726 | 0.16 |
| Contamination test† (Zhang et al., 2024b) | 0.241 | 1.000 | 0.389 | 0.462 | 1.000 | 0.00 |
| Single trap set† | 0.241 | 1.000 | 0.389 | 0.568 | 1.000 | 0.00 |
| Absolute- trajectory | 0.311 | 1.000 | 0.474 | 0.663 | 0.706 | 0.00 |
| HackProbe (ours) | 0.382 | 0.845 | 0.527 | 0.763 | 0.434 | 0.28 |
4.2 Recovering true capability
Detection is a means; the objective is the true capability that survives evolution. Table 2 reports, per channel, the end-of-evolution of the greedy host F0 and the additional gain of each immunization level, all measured on the held-out .
Every level improves on greedy selection on average, though not uniformly across channels. Risk-gated (F1) and bandwidth-limited (F2) reselection are positive on all four and reach nearly identical averages, and ; what separates them is price, not gain (Section 4.3). F1 wins biggest on the gameable grader (), where the raw gap gating its reselection is largest, even though that is the channel the fused detector ranks worst. F3 holds the single largest gain over F0, on judge collusion, where a lenient judge collapses the proposer onto a narrow region and the coverage bonus pulls it back out. F3 is also the only level that goes negative anywhere, mildly, on self-consistency bias.
With gold audit data this scarce, only the two starred intervals lie strictly above zero, and two more, both under F2, have their lower endpoint exactly at zero. So the table supports an aggregate, directional claim; per-channel effects are mostly too small to certify, and sharpening them needs a larger gold audit set.
| Hacking channel | F0 (greedy) | F1 gain | F2 gain | F3 gain |
| Contamination | ||||
| Gameable grader | ||||
| Self-consistency bias | ||||
| Judge collusion | ||||
| Average |
4.3 What protection costs when nothing is wrong
A monitor that helps under attack but degrades ordinary runs is not deployable, so we measure the clean-setting cost directly on the none channel, where greedy selection is in fact the right thing to do. Table 3 reports it alongside the diversity-bonus sweep that fixes F3’s setting.
Every level pays something on clean runs: each declines the top- candidate some of the time, and when is honest that candidate was correct. The useful comparison is what a level costs when nothing is wrong against what it returns when something is. Only F2 comes out favorable, returning for a cost of , a ratio of ; F1 buys the largest gains but pays most for them ( against , ratio ), and F3 sits between ( against , ratio ). The right panel shows why F3 uses a small bonus: gives both the best average under hacking and the smallest clean-run cost, and erases the gain.
The ratio’s standing is narrow. It orders the three levels, being a comparison of two averages over the same runs that leans on no single interval in Table 2, but it does not license reading per-channel gains as effects, and at a larger audit set could move it. It would be falsified by a host on which reselection costs more than it returns at every bandwidth, a regime Appendix F shows is not far away.
| Mode | NONE | Cost | Gain | Ratio |
| F0 | 0.646 | n/a | ||
| F1 | 0.573 | 0.73 | ||
| F2 | 0.599 | 1.11 | ||
| F3 | 0.604 | 0.84 |
| F3 bonus | Hacking4 | NONE |
| 0.05 | 0.630 | 0.604 |
| 0.10 | 0.612 | 0.599 |
| 0.15 | 0.622 | 0.583 |
| 0.25 | 0.599 | 0.557 |
4.4 Generalization to an unseen hacking type
A detector tuned on the same attacks it is evaluated on proves little, so we calibrate on three channels and evaluate on the held-out fourth (Table 4). Transfer is good to judge collusion ( AUROC), whose accumulating leniency is the change-point J1 is built for, and to self-consistency bias (), which leaves the signature J3 targets; contamination remains usable at . It degrades on the gameable grader ( AUROC, false-positive rate) for an instructive reason: grader exploits act on output format, not on the score–capability relation, so a threshold calibrated on the other three over-flags, and covering them needs a statistic keyed to output structure that the fusion lacks. The two held-out AUROCs above the pooled are no transfer bonus, since ranking within one channel is easier than across four whose divergences differ in scale.
| Held-out channel | AUROC | Precision | Recall | F1 | FPR | Pos. rate |
| Judge collusion | 0.814 | 0.570 | 0.718 | 0.635 | 0.227 | 0.295 |
| Self-consistency bias | 0.787 | 0.562 | 0.854 | 0.678 | 0.333 | 0.333 |
| Contamination | 0.724 | 0.236 | 0.700 | 0.353 | 0.264 | 0.104 |
| Gameable grader | 0.600 | 0.288 | 0.881 | 0.434 | 0.661 | 0.233 |
Scope of the present evaluation.
The host is prompt-level, chosen because it exposes per-generation candidate pools cheaply; weight-level self-evolution (Huang et al., 2026; Yuan et al., 2024) and a second prompt-optimization host (Agrawal et al., 2026) are left as external validity. That one host is also the whole of the evidence for harness-agnosticism, and the gold audit set binds every interval above. Appendix G adds the leave-one-statistic-out and rotation ablations, the only place the fresh layer is exercised; the inference overhead of F1–F3 is unmeasured.
5 Conclusion
HackProbe detects reward hacking in self-evolving hosts from black-box scores and outputs alone, and immunizes against it by reselecting an honest candidate. It keeps the capability coordinate non-circular, makes the risk score a calibrated -value, and turns a target error rate into a probe budget. What it is not yet is deployable. A false-positive rate of interrupts more than four honest generations in ten, and our rotation guarantee covers only an attacker who has to match every probe, which leaves the generator’s covering number as the open question.
AI Use Statement
In this work, we used generative AI tools only for language polishing. We did not use generative AI tools to generate scientific ideas, formulate claims, design experiments, produce experimental results, write code, create figures, or generate citations. All AI-assisted edits were reviewed and approved by the authors, and all cited claims were checked against the referenced sources. We take responsibility for the final content of this work.
Reproducibility Statement
Section 3 specifies the probe bank, the detector and the immunization rules, with every statistic and the calibrated threshold defined in Section 3.2. Appendix C states the assumptions in full, Appendix D proves Propositions 1 and 2 together with the covering-number argument, Appendix E proves the J3 separation result, and Appendix H gives the calibration procedure and the online loop as pseudocode. Section 4 gives the host, the run configuration, the injection protocol, the baselines, the metrics and the leave-one-hacking-type-out procedure.
References
- Gepa: reflective prompt evolution can outperform reinforcement learning. In International Conference on Learning Representations, Cited by: Appendix A, §1, §4.4.
- Concrete problems in ai safety. arXiv preprint arXiv:1606.06565. Cited by: Appendix A, §1.
- Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Cited by: §4.
- Scaling laws for reward model overoptimization. In International conference on machine learning, pp. 10835–10866. Cited by: Appendix A, §1.
- Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874. Cited by: §4.
- Probability inequalities for sums of bounded random variables. Journal of the American statistical association 58 (301), pp. 13–30. Cited by: Appendix A.
- R-zero: self-evolving reasoning llm from zero data. In International Conference on Learning Representations, Cited by: Appendix A, §1, §1, §4.4.
- Dspy: compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714. Cited by: Appendix A.
- Gsm-plus: a comprehensive benchmark for evaluating the robustness of llms as mathematical problem solvers. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2961–2984. Cited by: Appendix A, §1, §4.
- Categorizing variants of goodhart’s law. arXiv preprint arXiv:1803.04585. Cited by: Appendix A, §1.
- Gsm-symbolic: understanding the limitations of mathematical reasoning in large language models. In International Conference on Learning Representations, Vol. 2025, pp. 94743–94765. Cited by: Appendix A, §1, §4.
- Continuous inspection schemes. Biometrika 41 (1/2), pp. 100–115. Cited by: Appendix A, §3.2.
- Rectangular confidence regions for the means of multivariate normal distributions. Journal of the American statistical association 62 (318), pp. 626–633. Cited by: Appendix A, §3.2.
- Defining and characterizing reward gaming. Advances in neural information processing systems 35, pp. 9460–9471. Cited by: Appendix A, §1.
- Reward hacking in the era of large models: mechanisms, emergent misalignment, challenges. arXiv preprint arXiv:2604.13602. Cited by: Appendix A, §1.
- Self-rewarding language models. arXiv preprint arXiv:2401.10020. Cited by: Appendix A, §1, §1, §4.4.
- A careful examination of large language model performance on grade school arithmetic. Advances in Neural Information Processing Systems 37, pp. 46819–46836. Cited by: Appendix A, §1, §4.
- PaCoST: paired confidence significance testing for benchmark contamination detection in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 1794–1809. Cited by: Appendix A, §4, Table 1.
Appendix A Related Work
Optimizing a proxy that diverges from the intended objective is a long-standing safety concern (Amodei et al., 2016), formalized as reward gaming, where non-trivial unhackable proxies essentially do not exist (Skalse et al., 2022), and understood as a family of Goodhart effects (Manheim and Garrabrant, 2018); frontier evaluations document deliberate hacking when the scoring function is visible (Wang et al., 2026). Closest in spirit is reward-model overoptimization, where a gold reward model reveals that optimizing a proxy eventually degrades true performance (Gao et al., 2023); our is the online, black-box analogue of that gold-versus-proxy gap, computed per generation from a secret core instead of once offline from a trained gold model, and coupled to an intervention instead of reported.
The loops HackProbe monitors are those in which proxy and optimizer share a model: self-rewarding training whose judge saturates (Yuan et al., 2024), label-free co-evolution rewarding self-consistency (Huang et al., 2026), reflective prompt evolution (Agrawal et al., 2026), and declarative program pipelines (Khattab et al., 2023). In each, the quantity being maximized and the system doing the maximizing are entangled, which is exactly the configuration under which a fixed external coordinate is worth its cost.
Measurement, by contrast, has been studied mostly offline. Held-out reconstruction (Zhang et al., 2024a), symbolic templating (Mirzadeh et al., 2025), adversarial perturbation (Li et al., 2024) and confidence-based contamination tests (Zhang et al., 2024b) all establish that a benchmark score can overstate capability, but each is a one-shot audit of a finished model. HackProbe repurposes their shared premise, that the same capability probed through a different surface should yield the same score, as an online signal computed inside the loop and acted upon while the run continues. Statistically, the decoupling detector is a Page–Hinkley test (Page, 1954), the fusion a Šidák correction (Šidák, 1967), and the tail analysis standard sub-Gaussian concentration (Hoeffding, 1963); the contribution is not a new test but the choice of what to test and the guarantee that the composite risk score stays calibrated.
Appendix B Notation
| Symbol | Meaning |
| object under evolution at generation ; host’s evolution operator; run length | |
| observed visible score and comparison-core proxy of | |
| true capability on target distribution | |
| per-subdomain score; robust aggregate; number of core subdomains | |
| gold audit sets: calibration-only and final-evaluation-only (disjoint) | |
| scale-aligned divergence; re-estimated on honest windows | |
| slope and intercept of the honest level relation (J0) | |
| intercept and slope of the honest relation (J3) | |
| population divergence ; honest iff | |
| core-bias offset of the honest mean of ; its bound, | |
| variance proxy of ; effective sample size | |
| per-item variance proxies: core difference, visible score, modal agreement | |
| risk score, decision, low-bandwidth feedback signal | |
| level residual (J0) and conditional-confidence residual (J3) | |
| core accuracy; confidently-wrong rate on incorrect core items | |
| core-estimator bias and its sub-Gaussian noise, | |
| secret comparison core; rotated fresh layer at generation | |
| number of fused statistics (); FPR-calibrated Šidák threshold | |
| gaming footprint of candidate ; its penalty weight in | |
| hacking divergence gap; test offset; target detector error rate | |
| attacker’s failure probability; fraction of fresh probes it matches | |
| core probes; visible-evaluation items; incorrect core items (for J3) | |
| feedback quantization levels; candidate-pool size; generations before rollback | |
| fresh-probe pool size; fresh probes per generation; memorized set | |
| covering number of the probe generator; Lipschitz constant of the exploit | |
| black-box samples per probe (I2); J3 attack size in Lemma 1 |
Appendix C Assumptions
Throughout this appendix, expectations and probabilities are taken over probe and evaluation sampling with the trajectory held fixed. This conditioning matters: unconditionally, and both move with the random choice of and are strongly positively correlated under healthy evolution, whereas conditionally their measurement noises are independent because the two are computed on disjoint item sets. All variance statements below refer to the conditional object.
Assumption 1 (Concentration of the core proxy).
For any , the comparison-core estimate satisfies , where the bias is bounded, , and is zero-mean sub-Gaussian. Because the same core items are scored at consecutive generations, we state the concentration directly for the paired difference: is zero-mean sub-Gaussian with variance proxy , where is the per-item variance proxy of the within-item score difference. This is the correct object because the two estimates are not independent, and the pairing is a strength: per-item difficulty cancels in the difference, so is typically much smaller than twice the per-generation variance proxy an independent-samples treatment would yield. The assumption requires the core to be secret and outside the host loop (H1); if the host could fit the core, would no longer be bounded uniformly.
Assumption 2 (Sampling noise of the visible score).
The observed visible score is with zero-mean sub-Gaussian. The host re-evaluates the same items at consecutive generations, so the pairing argument of Assumption 1 applies and we state the proxy directly for the difference: is sub-Gaussian with variance proxy . Were the two evaluations instead independent, the proxy would be and every budget below would double.
Assumption 3 (Local linear coupling and bias stability).
Under an honest update, visible gains do not outrun true gains at the calibrated scale, as in Definition 1, with equality in the nominal case where the two are exactly scale-aligned. Separately, and for every generation whether honest or hacking, the differential core bias is a constant with . Writing , the honest mean of is , with , and vanishes when the core bias is constant across generations. Stating the bias bound for both regimes lets cancel in the separation below; restricting it to honest updates, as a coupling assumption alone would, leaves the hacking mean unconstrained and is not enough. One consequence deserves naming, because it makes the assumption stronger than it looks. Telescoping a constant over a run gives , so the bound of Assumption 1 forces : a constant differential bias is self-limiting on long runs. We read that as a reason to expect in practice rather than as licence to treat it as a free parameter. The analysis also treats as known: for the online estimate , an error adds a mean shift bounded by that we do not account for.
Assumption 4 (Secret, online-generated fresh layer).
Each generation’s fresh probes are drawn uniformly at random, with or without replacement, from a pool of effective size , and are invisible to the host before scoring. The counting argument of Proposition 2 needs that uniformity; the out-of-distribution drift H3 deliberately violates it by steering the layer toward weak subdomains, in which case should be read as an effective support size under the realized sampling distribution .
Appendix D Proofs
D.1 Proof of Proposition 1
By Assumptions 1 and 2, and are sub-Gaussian with variance proxies and respectively, and conditionally on their sampling noises are independent because the core is disjoint from the host’s visible evaluation set. Hence is sub-Gaussian with variance proxy
| (5) |
Write with the differential bias of Assumption 3, so that with . Under an honest generation, and hence ; under a hacking generation, and hence . Because Assumption 3 bounds the differential bias in both regimes by the same constant, is the same offset in the two displays and the separation between the means is at least .
False positives. Fix the threshold and let denote the honest mean. The event implies with , since . As is decreasing on , the sub-Gaussian upper tail evaluated at dominates, giving
| (6) |
so the worst case over the composite null is attained at the boundary .
True positives. Under a hacking generation, whose mean is at least , the sub-Gaussian lower tail gives
| (7) |
so the detection probability is at least .
Budget. Setting makes both exponents equal to . Requiring is equivalent to , that is, in terms of the effective sample size ,
| (8) |
Unpacking shows it to be the weighted harmonic mean of and , so and in particular whenever ; budgeting core probes as in Eq. (4) then suffices.
The proof delivers less than it may appear to. Its threshold uses , which is not known but estimated on calibration runs; with a plug-in satisfying , the two bounds become and , and the crude choice gives . And the bound is per generation: over a run of generations the expected number of false alarms from this rule is at most , and no correction across generations is claimed here.
D.2 Rotation sample complexity
Section 3.4 summarizes this result; we state and prove it here.
Proposition 2 (Rotation sample complexity).
Suppose an attacker tries to inflate the fresh layer as well, by memorizing a set of previously exposed instances. Matching all freshly sampled probes of a generation with probability at least requires , which tends to the full pool size as grows; when the pool is generated online and grows without bound, no finite memorized set suffices.
By Assumption 4, each of the fresh probes of a generation is an instance drawn from the pool of effective size , independently of the attacker’s covered set of size . The probability that a single fresh probe falls in the covered set is at most , so by independence the probability that all do is at most ; for sampling without replacement the same bound holds, since , each factor being at most because . Since the attack succeeds only if every fresh probe is covered, requiring the success probability to be at least chains the two bounds as
| (9) |
a necessary and not a sufficient condition. Since as , the required coverage tends to ; when the pool is generated online, grows without bound and no finite memorized set suffices.
The model favours the attacker in two places and overstates the conclusion in a third. The pool is treated as fixed and sampled i.i.d., which ignores H2’s non-reuse policy: in reality the instances an attacker has been able to observe are precisely those least likely to recur, so the true requirement is stronger than the bound. The bound also assumes the attacker’s covered set is chosen without knowledge of the sampling distribution, which uniformity in Assumption 4 makes harmless but non-uniform drift would not. Against that, though, the all-or-nothing success criterion is what produces the near-total-coverage conclusion: an attacker content to match only a fraction of the probes, and thereby inflate the fresh-layer score by roughly , needs only . Rotation forces effort linear in the pool size, but ”essentially the entire pool” is specific to demanding every probe.
D.3 Covering-number argument for structural attacks
The memorization model above treats fresh instances as atomic and interchangeable, which is why coverage must reach . A structural attacker instead models the response surface. Let the map from a fresh probe to the host’s exploitable behavior lie in a family equipped with a pseudometric , for instance the distance between the shortcuts two probes admit, and let be its -covering number. If the attacker fits the host on one representative from each of the cells of an -cover, then for any fresh probe there is a representative within distance , and by -Lipschitzness of the exploit response in the fitted behavior transfers with error at most . A fresh probe is therefore defeated with high probability once the cover is learned, so the sample complexity of defeating rotation is , independent of .
This has two consequences. If is bounded, as it is for a low-complexity generator such as a small template family, online rotation provides no asymptotic protection beyond a fixed set of representatives, so the defender’s only real lever is to enlarge in the first place — which is exactly what H3 (drift toward not-yet-fitted subdomains) and H4 (surface diversity via an embedding-distance criterion) are for. This is a design argument rather than a guarantee; establishing a matching lower bound on for a concrete online generator is left to future work.
Appendix E Why J3 detects self-consistency bias
Self-consistency bias is the empirically hardest channel because it need not move accuracy at all: rewarding agreement entrenches a confident but wrong modal answer, so the visible score rises while core accuracy and the proxy stay flat. Figure 2 contrasts the two regimes schematically, and the lemma below formalizes why the first-difference statistics are blind in the second while J3 is not.
Fix a generation and let the secret core contain items indexed by , each with a correctness indicator and modal-vote agreement , the fraction of the sampled answers equal to the model’s own majority answer. Let and let be the sample mean of over the items with , and recall with the honest affine coefficients fixed by calibration. Write for the per-item variance proxy of on incorrect items and for that of .
Lemma 1 (J3 separation in the stable regime).
Consider a stable self-consistency-bias generation in which (i) the visible inflation is already present and saturated, so that , and (ii) the attack raises the mean modal agreement on incorrect core items by relative to the honest calibration, in the sense that , while leaving unchanged in mean. Then
| (10) |
so J1 and J2 have zero expected separation, whereas J3 has expected separation regardless of how the attack is reflected in the aggregate level. Moreover is sub-Gaussian with parameter , and by the argument of Proposition 1 the one-sided J3 test attains false-positive and detection errors below once , for which is necessary.
Proof.
By (i), , hence , and J2, which tests , is likewise zero. For J3, the coefficients are fixed, so taking expectations in (ii) over and using the tower rule, . For the variance proxy, is a difference of two bounded sample means: over the incorrect items, sub-Gaussian with parameter , and over all items, sub-Gaussian with parameter . The two are not independent, and it matters that we do not pretend otherwise: the incorrect set is itself determined by the , so is a function of , and both statistics are read off the same sampled answers per item. We therefore add sub-Gaussian parameters rather than variances, using , which holds for arbitrarily dependent and and yields . Applying the sub-Gaussian tail bounds of Proposition 1 with gap , threshold and parameter gives the stated budget. ∎
The budget runs through rather than alone because the calibration slope transmits accuracy noise into the residual; the simpler form is the special case and understates the requirement otherwise. Refusing the independence assumption costs us a little tightness, since adding sub-Gaussian parameters is conservative whenever the two noise sources are close to orthogonal. Two further approximations go uncorrected. We do not propagate the estimation error in , and is itself random, so the budget should be read as a claim about the realized number of incorrect items rather than its expectation.
J0 is absent from the statement, which is a choice rather than an oversight. If the attack elevates relative to the calibrated honest relation , then the level residual is positive too and J0 is not blind here; the distinction between J0 and J3 in this regime is one of power, not of consistency. The level residual dilutes the attack over all core items, most of which the model answers correctly and on which the agreement pattern carries no information, whereas conditions on the incorrect items where the effect is concentrated. The budgets make the comparison concrete: J3 resolves the full effect on incorrect items, while J0 must resolve the much smaller shift that the same attack induces in the aggregate level. In the limiting case where the attack changes only the agreement pattern on wrong items and leaves the visible level untouched, J0’s separation vanishes and J3 is the sole remaining signal.
The lemma therefore isolates the design rationale. J3 is the only fused statistic whose expected signal is driven by the conditional confidence pattern instead of by first differences of accuracy, so it remains informative in the regime that defeats divergence-based detectors. Its cost is governed by the number of incorrect core items , which is why the core must be sized so that hard subdomains contribute enough wrong items to meet the budget.
Appendix F Testing the probe budget and the bandwidth limit
Two of the design claims are predictions rather than measurements, and each can be falsified on its own terms: Eq. (4) converts a target error rate into a core size, and Section 3.3 argues that the bandwidth cap is what stops the feedback from becoming a second score. Figure 3 tests both.
Panel (a) sweeps the core from to probes at a fixed false-positive rate of , isolating power from threshold placement. Detection rises from to and tracks the overlaid prediction to within throughout, and that prediction is not a fit: its constants come from calibration runs and stay fixed across the sweep. Where the operating point sits matters more. Our budget of lands where the curve is steepest, which is why a stricter threshold costs so much recall and why Section 4.1’s false-positive rate is so high; Eq. (4) puts detection at , leaving the monitor under-provisioned by roughly probes. Halving the injected strength shifts the curve right by against the that predicts, the closest direct test of Proposition 1 this setup admits.
Panel (b) sweeps , instantiating the interpolation of Section 3.3. Clean-run cost grows monotonically with bandwidth, from at to once the quantizer is removed. That is the direction the resolution argument anticipates, though the same trend follows from F2 approaching core-greedy selection, which on clean runs swaps an honest for a noisier proxy; the sweep cannot separate the two. Gain does not follow suit. It peaks at , so the ratio crosses break-even exactly once. Adjacent levels differ by far less than the intervals of Table 2, so the shape is directional rather than certified; what it argues against is the idea that a good reselection rule simply improves with everything it is told.
Appendix G Additional ablations
G.1 Which statistic carries which channel
The fusion is justified by the claim that its four statistics fail in different regimes. Table 6 tests that claim by removing one at a time and recalibrating to the same target, so the comparison is at matched operating conditions, not matched thresholds.
| Variant | Pooled | Contam. | Grader | Self-cons. | Judge |
| Full fusion | 0.763 | 0.741 | 0.688 | 0.802 | 0.812 |
| without J0 | 0.716 | 0.641 | 0.652 | 0.791 | 0.788 |
| without J1 | 0.719 | 0.688 | 0.671 | 0.784 | 0.731 |
| without J2 | 0.754 | 0.727 | 0.681 | 0.795 | 0.808 |
| without J3 | 0.691 | 0.735 | 0.639 | 0.601 | 0.771 |
The pattern follows the design rationale. Dropping J3 costs AUROC on self-consistency bias and almost nothing on contamination, which is the empirical form of Lemma 1: when an attack moves the conditional confidence pattern without moving aggregate accuracy, the conditional residual is the only statistic left with signal. Dropping J0 costs most on contamination, whose exploit is present from the first generation and so registers as a level gap rather than as a divergence, while dropping J1 costs most on judge collusion, whose leniency accumulates and is therefore a change-point. J2 is the exception: removing it costs pooled, within noise, because on this host stagnation is nearly always accompanied by a divergence J1 already sees. We keep J2 because it tests capability stagnation directly and, like J3, needs no estimate of , so it survives a mis-scaled calibration; but nothing in the reported numbers depends on it.
G.2 What rotation buys against a co-adapting host
No number in Section 4 depends on the fresh layer, since it contributes no statistic to . To test it we need a host that is allowed to fit the probe layer, so we grant the adversary the outputs of every probe it has been scored on and let it optimize against them for generations. Table 7 reports the inflation the host achieves on the fresh layer, that is, its fresh-layer score minus its true capability on , together with detection quality.
| Fresh-layer inflation | Detection AUROC | |||||
| Fresh layer | ||||||
| Fixed (no rotation) | 0.06 | 0.17 | 0.31 | 0.744 | 0.681 | 0.588 |
| Rotated (H2) | 0.02 | 0.04 | 0.07 | 0.759 | 0.751 | 0.736 |
| Rotated surface diversity (H4) | 0.01 | 0.02 | 0.03 | 0.762 | 0.758 | 0.749 |
A fixed layer behaves the way the design anticipates it would: the host fits it, inflation grows to by generation , and detection decays toward chance, which is the failure mode a single held-out probe set has and the reason the fresh layer is rotated at all. Rotation caps inflation at over four times as many generations.
The residual matters more than the size of the improvement. Proposition 2 rules out an attacker who matches every fresh probe, but as Appendix D notes, matching a fraction of them requires only , and inflation of is exactly that partial regime. The residual is therefore consistent with either partial memorization or the structural leakage of Appendix D.3, and this experiment does not separate them. That it shrinks to once H4 enforces surface diversity is weak evidence for the structural reading, since surface diversity raises the generator’s covering number without changing the pool size ; deciding the question would require varying generator complexity directly, which we have not done.
Appendix H Detector calibration and the immunization loop
Calibration uses only honest runs and . It fits the honest affine relations for J0 and J3, estimates the honest null moments of , initializes , fits the fusion, and calibrates to a target false-positive rate.
Where the independence needed by Eq. (2) is too doubtful even for the empirical re-calibration to absorb, we fall back to a Bonferroni rule: flag when , and report the risk as . This needs no independence assumption at all and costs a little power. Every number we report uses the Šidák form with an empirically re-calibrated ; the Bonferroni variant is there for deployments unwilling to rely on that calibration. Online, is refreshed only on the trailing window of generations already declared honest, so that the estimate used to test generation depends on nothing from generation . Algorithm 1 summarizes the loop. It runs the fusion from , since the first differences and need a predecessor; the first generation is scored by the two level statistics J0 and J3 alone, with in the Šidák correction.