Cheap Verifiers, Large Blind Spots:
Measuring the Reliability Cost of Cost-Saving Cascades
Abstract
Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natural extension closes the loop—fine-tune the cheap student on the verifier’s rejections so the escalation rate, and cost, fall each round. We set out to measure this loop on real LLMs, and report four findings. First, the verifier’s blind spot—the fraction of the student’s wrong answers it waves through—is large and moves adversarially: it grows with student capability ( from to as the student scales 0.5B32B) and shrinks with verifier capability, so it is worst exactly in the cheap-student, cheap-verifier configuration cascades exist to create. Second, buying it away returns the saving: a frontier verifier drives to but then escalates on of hard-MATH queries against a true error rate—paying the frontier price on nearly half of all traffic, the very cost the cascade exists to avoid. Third, naive corrective fine-tuning on the verifier-rejected tail does not improve the small student but degrades and ultimately collapses it, across every teacher we tried (cross-family and same-family)—so at this scale the “self-improving” loop is self-defeating. Fourth, throughout all of this the cascade’s own dashboard—every metric computed through the verifier—reads a flat error while true delivered error swings up to : the system is blind to its own degradation by construction. We then give the theory that explains the blindness—a two-population conservation law, , under which every in-loop metric improves while true quality does not—and a synthetic study that validates the mechanism where the blind spot’s dynamics are emergent rather than imposed. The practical conclusion is a measurement discipline: the reliability of a self-improving cascade cannot be read from any metric computed through its own verifier.
Keywords: inference cascades reward overoptimization LLM-as-judge cost-efficient inference Goodhart’s law
1 Introduction
The dominant lever for reducing the cost of large-language-model inference has shifted from the model to the harness — the scaffold of routing, verification, retrieval, and retries wrapped around a fixed set of weights. Among harness techniques, the cascade is the most direct cost play: run a cheap model on every query, and escalate to an expensive frontier model only when some signal says the cheap answer is untrustworthy [1, 2, 3]. The frontier model functions as a verifier; the cascade pays its price only on the escalated tail.
A tempting extension closes the loop. Every escalation yields a frontier-quality answer on exactly the distribution where the cheap model fails — a free training example. Fine-tune the cheap student on these corrections, and its error rate falls, its escalation rate falls with it, and the cascade gets cheaper each round. Recent work explores online, training-free versions of this idea, in which deferred queries produce reusable in-context strategies for the weak model [4, 5]; the parametric version, which updates the student’s weights, is the natural and more aggressive sibling.
The loop invites an analogy to speculative decoding, where a small drafter proposes tokens that a large model verifies in parallel, cheaply and — crucially — losslessly: rejection sampling guarantees the output distribution is exactly the large model’s [6, 7]. If semantic cascades inherited that guarantee, the self-improving loop would be a strict win: lower cost, preserved quality. They do not. Speculative decoding verifies a token against a distribution it has in closed form; a semantic cascade verifies a claim against a judgment the verifier must itself infer, and that judgment is wrong in both directions. Once verification is imperfect, training a student to satisfy the verifier is not distillation toward truth — it is optimization against a proxy, the setting in which reward-model overoptimization is known to arise [8, 9], a modern instance of Goodhart’s law [10].
This paper leads with what we measured. We built the wind tunnel needed to see a verifier’s blind spot — tasks with a cheap oracle from which the verifier is deliberately blinded (Section 4) — and ran the loop on real language models (Qwen2.5 students 0.5B–32B, GSM8K and hard MATH, OpenAI verifiers up to gpt-5-mini). Four findings, all measured rather than assumed, organize the paper.
- •
The blind spot is worst exactly where cascades operate. The verifier’s blind-spot rate — the fraction of the student’s wrong answers it accepts — grows with student capability (from at 0.5B to at 14B, fixed verifier) and shrinks with verifier capability. A more capable student makes subtler, more convincing errors; a cheaper verifier catches fewer of them. The danger zone is therefore the cheap-student, cheap-verifier corner that makes cascades attractive in the first place (Section 5).
- •
Buying the blind spot away gives the cost back. Swapping a frontier verifier in on hard MATH collapses to , but that verifier then escalates on of queries against a true error rate — it buys recall by over-rejecting, paying the frontier price on nearly half of all traffic. Low blind spot and low cost are not simultaneously available from a fixed verifier.
- •
The “self-improving” loop, at this scale, is self-defeating. Naive corrective fine-tuning on the verifier-rejected tail did not improve any small student we tried; it degraded them and, cumulatively, collapsed them — across cross-family and same-family teachers alike. We never instantiated a loop that improved the student, and report that plainly: it is a direct caution against the parametric self-improving loop (fine-tuning the student on the verifier’s rejects), the natural next step beyond today’s training-free deferral-reuse methods, and it means the clean error floor below stays a theoretical result rather than a measured one.
- •
None of this is visible from inside. Every metric a practitioner monitors is computed through the verifier, and reads a flat error while true delivered error swings to — the dashboard cannot distinguish a healthy student from one the loop is actively wrecking.
The rest of the paper explains why the dashboard is blind. We formalize the loop and its one structural property — training signal derived only from verifier-detectable errors (Section 3) — and derive a conservation law: user-facing error asymptotes not to zero but to a floor anchored at the initial confidently-wrong-and-accepted mass, (Section 6, with a two-population proof in Appendix A), under which every verifier-computed metric improves while true quality does not. A synthetic mechanism study, where the blind spot’s dynamics are emergent rather than imposed, validates the law and the mitigation exchange rates (Section 7). The theory is the explanation; the measurements are the result.
2 Background and related work
Cascades and routing. Cascades query models in sequence and decide, from a post-generation signal, whether to accept the cheap answer or escalate [1, 2, 3]. Routers instead decide before generation which model should answer [11]. Both aim at the cost–quality Pareto frontier; both, in their standard form, hold the cheap model fixed. Our object of study is the case where the cheap model is not fixed but is trained on the cascade’s own escalation signal, which changes the dynamics qualitatively.
Verifier-as-training-signal. Using a stronger model or a checker to supervise a weaker generator is the backbone of rejection-sampling fine-tuning and reinforced self-training [12], self-taught reasoning [13], and weak-to-strong supervision studies [14]. Recent cascade work makes the loop explicit and online, storing the strong model’s deferral-time strategies for reuse [4, 5], the latter deferring by self-consistency [15]. This literature reports accuracy-at-cost, typically as a single snapshot. None of it, to our knowledge, characterizes what happens to the composition of the student’s residual errors as the loop runs, which is precisely where the blind-spot effect lives.
Overoptimization and imperfect verifiers. Optimizing a policy against a learned proxy of human preference improves the proxy’s score while eventually degrading the true objective — the gold reward turns over even as the proxy reward climbs [8, 9], a modern Goodhart effect [10]. This is the phenomenon underneath our claim, but the signature we predict is not Gao et al.’s turnover: under the corrective loop the proxy improves while gold error stays flat at a positive floor, and a turnover can appear only when accepted outputs are recycled as labels (self-training, our H3). Closest to us, Stroebl et al. [16] show that inference-time resampling against an imperfect verifier cannot beat the verifier’s false-positive rate — a static, single-shot form of the blind spot. Our claim is its closed-loop counterpart: once the student is trained against the verifier, that same false-accept mass becomes a round-persistent error floor rather than a per-query bound — pushed down only by generalisation spillover (corrective) and held higher when accepted outputs are recycled (self-training) — and turns invisible to every in-loop metric. Beyond the shape difference, our setting differs from reward-model overoptimization in that the proxy is the expensive resource whose invocation the loop is trying to minimize — so the overoptimization dose and the cost saving are the same axis, a coupling absent from RLHF — and that the student is trained on a self-selected slice (only the verifier’s rejects) rather than a fixed preference distribution.
LLM-as-judge and self-preference. Using an LLM to judge another model’s output is now standard [17], as is the finding that judges can favor text from their own family or their own generations [18]. We distinguish that effect from ours. Our correlation corollary is not about a judge preferring its own text it is about shared failure modes — a distractor whose reasoning “looks right” to a student built on certain pretraining priors also looks right to a verifier built on similar priors. Crucially, we make the causal claim on a measured blind-spot rate, not on a family label, so it does not depend on the self-preference mechanism being the cause.
Self-consuming loops and pseudo-labeling. Training a model on its own outputs can cause model collapse — loss of variance and tails under recursive generation [19, 20]. Our self-training loop is a verifier-filtered self-consuming loop, and the difference is the point: there, collapse arises from unfiltered recursion; here, the verifier’s blind spot selects which self-outputs are reinforced, concentrating error precisely where the filter cannot see. At the level of one round this is the classic confirmation bias of pseudo-labeling [21] and a reason intrinsic self-correction stalls without external signal [22]; what those results lack — and what the cascade supplies — is that the filter is an external frontier verifier, so the floor is set by student–verifier blind-mass overlap rather than by the student’s own confidence.
Selective prediction and learning to defer. The verifier’s accept/reject decision is a deferral rule, and its blind-spot rate is the miscoverage of that rule — the quantity selective classification [23] and learning-to-defer [24] are built to control. That literature studies the risk–coverage tradeoff of a static rule; we study the closed-loop dynamics that arise once the deferred slice becomes training data and the rule’s blind region is never corrected.
3 The self-improving cascade
We fix notation. A student maps a query to an output . A verifier maps a query and a candidate output to a binary decision, . A teacher (often the same frontier model as ) produces a replacement output on rejection. An oracle gives ground truth; the oracle exists in analysis but is not available to the loop — if it were, one would simply verify with it.
The cascade delivers when accepts and when rejects (Figure 1). The self-improving variant additionally updates the student. In round :
- 1.
Draw a batch ; the student generates for .
- 2.
The verifier judges each output; let and .
- 3.
Form training targets and fit . The corrective loop trains on . The self-training loop adds the accepted student outputs as positive targets.
Two quantities drive everything. Let
be the student’s raw error rate, and let
be the verifier’s blind-spot rate — the fraction of the student’s genuine errors that the verifier waves through. Its complement is the verifier’s recall on errors.
The structural fact that organizes the rest of the paper is this: the loop’s training signal is a function of , never of . Rejections identify errors can see; the errors cannot see are, by construction, indistinguishable to the loop from correct answers. The loop is therefore a selection process that removes verifier-detectable errors and retains verifier-undetectable ones.
4 Measuring a blind spot: the wind-tunnel method
The obstacle to measuring any of this is that the blind spot is, in genuinely fuzzy tasks, unobservable: seeing it requires ground truth independent of the verifier, and independent ground truth is exactly what fuzzy tasks lack. Our method is a wind tunnel — a controlled setting in which ground truth exists but the verifier is denied it — and the real-model measurements of Section 5 are what it produced.
The wind tunnel. Use tasks that carry a cheap, reliable oracle, and blind the verifier to it. On mathematical problem solving [25, 26], the student emits a reasoning trace and a final answer. The oracle is exact-match of the answer against the gold solution — cheap, reliable, and independent of . The verifier is a frontier model asked whether the solution is correct, given the problem and the trace but not the gold answer. This is a genuinely fuzzy verifier: it false-accepts wrong reasoning that looks convincing and false-rejects correct reasoning that looks unusual. Its blind spot is real, and — because we hold the oracle in reserve — now measurable. The oracle is used for measurement only; it never routes or trains, exactly as in a real deployment where it would be absent.
This design pre-empts the two obvious objections. “Just use the deterministic verifier” is answered by including the oracle-verifier as a control condition (): the point is to characterize the fuzzy regime where no such verifier exists. “Math is not a fuzzy task” is answered by treating it as a model system: math carries a cheap gold oracle and a frontier verifier that visibly false-accepts and false-rejects, which is exactly what makes the blind spot measurable here. Extending the wind tunnel to a genuinely fuzzy task — long-form claim verification or code-review quality, whose oracle would have to be a super-verifier (an expensive ensemble validated against gold on the math tasks, then trusted where gold is unavailable) — is external-validity work we did not run; Section 9 names it as the main open scope.
Loop protocol. Fix disjoint splits: a pool for round batches, a held-out evaluation set used for all reported curves and never trained on, and a fixed probe set for tracking error composition. Each round: the student generates on a fresh batch; the verifier judges, blinded to the oracle; rejects receive a fresh teacher correction; the student is refit. To isolate data composition as the independent variable and remove optimizer path-dependence and catastrophic forgetting as confounds, we retrain cumulatively from the base model on all data collected through round , rather than incrementally from (incremental training is retained only as a robustness check). The verifier is pinned — same model, prompt, temperature, and seed, cached by output hash — so that any change across rounds originates from the student’s inputs, never from drift in .
Measurement and the decomposition that must not be skipped. Every round, on the held-out set and via the oracle, we record raw error , escalation rate , verifier recall and blind-spot rate , false-reject rate on correct answers, verifier-estimated accuracy (the dashboard number), and the user-facing error . Critically, is decomposed into its accepted-and-wrong term and its teacher-error term at all times; the conservation claim is about the former alone, and conflating the two invites exactly the critique a careful reviewer will raise. A sanity gate tracks whether actually falls: if the loop does not improve the student, the conservation-floor reading is unavailable — an outcome we report plainly (it is what the real-model runs of Section 5 in fact exhibit) rather than papering over, since the dashboard-blindness result holds whether the student improves or degrades.
5 Real-model measurements
This section is the paper’s empirical core.11 1 All code and data are public: github.com/AltSlate-Labs/cascade-blindspot. We ran the wind-tunnel method of Section 4 on real language models: students are Qwen2.5-Instruct (0.5B–32B), fine-tuned with LoRA on a single H100; verifiers and teachers are OpenAI models (gpt-4o-mini, gpt-4.1, gpt-5-mini), always blinded to the gold answer; the oracle is exact-match on GSM8K [26] and symbolic equivalence on the hard subset (levels 4–5) of MATH [25]. Nothing is assumed — whether a blind spot exists, how large it is, and how it moves with student and verifier strength are all measured. Error bars throughout are Wilson intervals; the blind-spot rate is a proportion over the wrong-answer subset only, so its intervals (–) are wider than those on the raw error rate (). These are single-seed runs; Section 9 treats seed variance.
None of it shows on the dashboard (Figure 2). Running the corrective loop with a Qwen2.5-7B student and a gpt-4o-mini verifier on GSM8K, the verifier-estimated error of the delivered stream holds flat near across all rounds, while the true (gold) user-facing error swings from to — a – gap whose intervals stop overlapping from round 1 on, so no in-loop metric reveals it. The frozen control stays near ; the loop itself, trained on the frontier teacher’s corrections, degrades the student (raw error rises), and the dashboard is blind to the degradation exactly as it is blind to the level. This is the alarming finding, and it does not depend on any conjecture: the gold and dashboard curves are both directly measured. This is a single training seed, so the exact per-round trajectory could shift; but the level gap — a flat dashboard against a – truth — is far larger than any plausible LoRA seed variance, so the result is the gap, not the particular curve (Section 9).
The blind spot grows with student capability (Figure 3). Holding the verifier fixed (gpt-4o-mini) and sweeping the student 0.5B→32B on GSM8K, rises from (CI ) at 0.5B to () at 14B: a stronger student makes subtler errors, so the fixed verifier is fooled more often. The small-vs-large separation sits well outside the intervals; the top three sizes (7B–32B, all ) form a plateau within noise, so the honest reading is a monotone rise that saturates, not a peak-and-fall. This is a different mechanism from Corollary 1 and worth separating from it. Corollary 1 is about correlation — student and verifier finding the same wrong answer convincing because they share failure modes; here the verifier is held fixed while only the student varies, so what moves is the capability gap: a stronger student’s errors are intrinsically subtler and harder for any fixed verifier to catch, correlated or not. This capability-gap route to a high blind spot is not part of the self-preference literature, and it carries a sharper warning — the blind spot gets worse as students improve, so the problem grows rather than shrinks with progress on the cheap model. Note the finding is a static characterization — a property of the student–verifier pair before any training — so it holds for a plain cascade and does not depend on the loop working; that is exactly why it is robust.
The blind spot is the price of a cheap verifier (Figure 4). On hard MATH, holding the 7B student fixed and sweeping the verifier, collapses from (gpt-4o-mini, CI ) to – for the two strong verifiers (gpt-4.1 and gpt-5-mini, intervals and — statistically indistinguishable from each other): a strong verifier catches almost every error on a task it can largely solve. But that low blind spot is not free — gpt-4.1’s escalation rate (, CI ) meets or exceeds the true error rate (), so it buys recall by over-rejecting correct answers too, paying the frontier price on nearly half of queries — the very cost the cascade exists to avoid. The blind spot is thus largest precisely in the cost-saving configuration — a capable cheap student checked by a cheap verifier — and buying it away gives the cost back. Like the student sweep, this is a static operating-point characterization, independent of any training.
The loop degrades the student — every teacher we tried. The one thing we set out to build, a loop that improves the small student, we could not. The two figures above are inference-only sweeps; the training loop we ran on two students, the 1.5B and the 7B (GSM8K corrective, and the teacher-family comparison of the next paragraph), and on both it fails the same way. Across every teacher — cross-family (gpt-4o-mini and gpt-4o) and same-family (Qwen2.5-32B) — naive LoRA fine-tuning on the verifier-rejected tail degrades the student rather than improving it, and cumulatively collapses it (raw error rising to as training on the hardest, most style-shifted examples breaks its output format) — an effect adjacent to the model-collapse literature [19], and one the dashboard is equally blind to. This is worth stating sharply: the parametric self-improving loop — the weight-updating sibling of today’s training-free deferral-reuse methods [4, 5] — fine-tunes on exactly the hard-tail rejects most likely to destabilise a small student, so its naive form is not merely suboptimal but actively harmful at this scale. So the real runs establish the blind spot’s structure — real, scaling up with student and down with verifier, invisible to every in-loop metric — but not the clean conservation scissors (raw error falling while user-facing error floors at ), which needs a loop that improves the student. The floor therefore stays a theoretical result — derived in Section 6 and Appendix A, validated synthetically in Section 7 — and its real-model demonstration (a stronger student, or a training recipe that resists the hard-tail distribution shift) is open. The next section explains why, whether the loop improves or degrades the student, none of it registers on the dashboard. Configuration and scripts are in the artifact (Section B).
6 Why the dashboard is blind: a conservation-law account
The measurements raise one question above the rest: why does no metric a practitioner can compute move, whether the loop is holding steady or actively collapsing the student? The answer is structural, and this section derives it. The same structure predicts that even a loop that worked — that improved the student round over round — would leave a positive error floor rather than driving delivered error to zero. We could not instantiate that improving loop at real-model scale (Section 5), so the floor is a theoretical prediction; the dashboard blindness it explains is not.
Consider the student’s errors as two populations: those the verifier detects (mass ) and those it does not (mass ). The loop applies a correction pressure to the first population and no direct pressure to the second: the undetectable errors receive no targeted training signal, though parameter updates driven by the detectable ones can still spill into them (Figure 6).
The user-facing error rate — the quantity a deployment actually ships — is the error that survives verification:
The second term is the teacher’s fallibility on the rejected items; its conditional rate is bounded by ‘s error rate, but its mass shrinks with the escalation rate as the loop proceeds, and when and share priors, wrong teacher labels that would also accept can enter training — blurring the corrective/self-training distinction (Section 9). The first term is the blind spot, and our central claim concerns it.
Conjecture (Blind-spot conservation). Under the corrective loop with a fixed, imperfect verifier, and absent generalization spillover into the blind region, the accepted-and-wrong error mass is conserved across rounds even as the detectable mass is driven down, so the user-facing error does not vanish but asymptotes to
Spillover relaxes this equality to an upper anchor, , when is large relative to the student’s capacity floor. Appendix A makes both statements precise (Proposition 1).
The intuition is that the loop removes error mass only where it has signal. Detectable error shrinks because it is trained against; undetectable error persists because it is not. The anchor reaches the raw student too: with no direct signal on the blind region, to first order inherits the same floor, and any drop below is spillover (self-healing). Only when the loop trains on every error — the oracle-in-loop control C of Section 4 — does raw error fall to the student’s capacity limit (the dashed trajectory in Figure 6) while delivered error reaches zero. The synthetic study of Section 7 plots this decomposition directly.
Two second-order effects perturb the anchor, and they push in opposite directions:
- •
Self-healing (corrective loop). Gains in general competence may incidentally fix some blind-spot cases the loop never explicitly targeted, so the observed floor falls below . The size of this gap measures how much reliability the loop delivers “for free” beyond what the verifier can see.
- •
Self-reinforcement (self-training loop). When accepted student outputs are fed back as positive targets, false accepts are injected as labels, actively teaching the student to reproduce the errors the verifier cannot catch. The floor then sits strictly above the corrective floor, and — when reinforcement outweighs self-healing — above itself, with user-facing quality able to degrade while every in-loop signal improves.
The falsifiable core is narrower than this interpretive frame and should not be confused with it. What can be refuted is H1 — a strictly positive floor, against the null — and H2 — the floor increasing in the measured , against the null of no relation. The anchor is a reference scale for the corrective loop, not a conserved quantity; a floor of zero, or a floor uncorrelated with , would count against the framework. The self-healing / self-reinforcement split then interprets where a given loop lands relative to the anchor.
Two corollaries sharpen the practical stakes.
Corollary 1 (Correlation scaling). The floor is increasing in the verifier’s blind-spot rate , and is large when student and verifier share failure modes — when a wrong answer convincing to is also convincing to . (Shared failure modes are one sufficient cause of a high , not the only one — a lazy rubber-stamping verifier has high with no shared bias, and, as Section 5 measures directly, so does a large capability gap: a strong student’s errors are intrinsically subtle, so even an uncorrelated verifier catches fewer of them. Correlation and capability-gap are distinct routes to the same high .) Verifiers from the same family or pretraining lineage as the student should therefore yield higher floors than independent verifiers; a deterministic oracle-verifier () yields no floor. The engineering reading is a design rule: decorrelate the verifier from the student — and, from the capability-gap route, expect the blind spot to grow as the cheap student improves.
Corollary 2 (Dashboard blindness). Every metric a practitioner naturally monitors — the escalation rate (a proxy for cost) and the verifier’s estimated accuracy on the accepted stream — is computed through . As the loop concentrates error into ‘s blind spot, these metrics improve or hold steady — fewer escalations, near-perfect apparent accept-stream quality (Figure 5), the accepted-stream component being structural since cannot flag its own accepts. The gold-truth accuracy of the accepted stream, which no in-loop instrument reports, fails to improve in step — it stays floored. The system is thus blind to its own degradation by construction, not by oversight; the cost dashboard and the quality dashboard are the same instrument, and it is the compromised one.
The account’s five predictions. The conservation-law account makes five directional, falsifiable predictions, tested across the rest of the paper. One is already confirmed on real models: the dashboard-blindness prediction (H4) is directly measured (Section 5). A second is addressed only at the level of its precondition: H2 asks whether the error floor scales with , and while we never measure a floor on real models, we do measure that itself rises with the student–verifier capability gap (Section 5) — the necessary precursor, not the floor-scaling itself. The floor (H1), the loop-design hazard (H3), and the mitigation exchange rates (H5) all need a loop that improves the student, so they are tested in the controlled model of Section 7. Each is stated with its null.
- H1 (Positive floor)
-
User-facing error does not vanish but asymptotes to a strictly positive floor ; to first order raw error inherits the same floor, and a matched control isolates the verifier-induced excess, while the oracle-in-loop control drives delivered error to zero. Null: — no floor.
- H2 (Blind-spot scaling)
-
The asymptotic floor is increasing in the measured initial blind-spot rate . Sweeping the verifier to vary traces a positive relation; an oracle-verifier sits at the origin. Null: the floor is independent of .
- H3 (Loop-design hazard)
-
The self-training loop yields a strictly higher floor than the corrective loop; under strong enough reinforcement it can push the floor above and render non-monotone while in-loop metrics improve. Null: the two loops floor at the same level.
- H4 (Dashboard blindness)
-
A large, persistent gap separates the verifier-estimated error of the delivered stream from its gold error, invisible to every in-loop metric. A further prediction, untested here, is that the gap widens over rounds whenever self-training grows the blind mass. Null: estimated and gold error track each other.
- H5 (Mitigation exchange rate)
-
A decorrelated verifier ensemble, and spending a fixed fraction of budget on random oracle audits that re-inject blind-spot cases into training, each lower the floor toward the oracle-in-loop control, at a quantifiable cost. Null: neither intervention moves the floor.
7 Synthetic validation
Because the real-model loop degrades rather than improves the student (Section 5), the conservation floor — the behaviour of a loop that works — cannot be read off the real runs. We therefore turn to a controlled model in which the loop provably improves the student, and ask whether the law appears when the blind spot’s dynamics are emergent rather than imposed. A positive answer is necessary, not sufficient — it cannot speak to real language models — but a negative answer would refute the mechanism outright. This is where H1, H3, and H5 are tested.
Model. The task is -way classification standing in for “produce the right answer”; ground truth is a fixed random two-layer network. The student is a logistic model on random features, retrained each round on a pool the loop grows; it improves with data. The verifier is fixed with two regimes: on a blind region — a fixed random half-space whose threshold is set so the region covers input-space mass — it rubber-stamps whatever the student says; elsewhere it re-derives its own class with a strong classifier and accepts only on agreement. So sets the verifier’s blind-spot rate; the measured slightly exceeds (e.g. at ) because the strong classifier occasionally agrees with a wrong answer outside the blind region. This models blind mass, not correlation per se — H2 tests floor-against-, and Corollary 1’s correlation reading is one account of what makes large. Crucially, nothing tells the loop to spare blind-spot errors: blind-region items are accepted, so they never enter the training pool; whether that yields a conserved floor, self-healing, or self-reinforcement is left to the learning dynamics. We run five variants — corrective (A), self-training (B), oracle-in-loop (C, perfect verifier trained on all items), frozen (D), and a matched control (perfect verifier, trained on wrong items only, isolating the verifier-induced excess) — for eight rounds over 20 seeds, the teacher supplying true labels on rejected items. The dashboard metric is the verifier-estimated error of the delivered stream — the fraction the pinned verifier rejects on a re-pass — which a practitioner computes without gold.
| variant | ||||||
|---|---|---|---|---|---|---|
| corrective (A) | .438 | .764 | .334 | .046 | ||
| self-training (B) | .438 | .764 | .334 | .045 | ||
| oracle-in-loop (C) | .438 | .000 | — | .000 | ||
| matched | .438 | .000 | — | .000 | ||
| frozen (D) | .438 | .764 | .334 | .048 |
Results. Four of the five predictions are supported as stated, and H3 is supported in a corrected form (Table 1). H1 (positive floor): under the fuzzy verifier, gold user-facing error falls but floors — from to — while the oracle-in-loop control reaches ; a matched control (perfect verifier, same reject-and-retrain rule) isolates the verifier’s effect on raw error, — this excess conflates two verifier-caused effects, never labelling the blind region and the smaller training pool that results, so it is verifier-induced but not blind-mass censoring alone (Figure 6). H4 (dashboard blindness): the dashboard reads while the truth is — a large level gap. It does not widen here: under both loops gold error falls while the dashboard stays flat, so the gap narrows; the widening form needs the blind mass to grow, which our readily-generalising student does not exhibit, so it remains a real-model prediction (Figure 5). H2 (scaling): the floor increases monotonically with the measured , from to as sweeps , with tight per-seed bands (Figure 7b); at the smallest the floor slightly exceeds because the capacity floor binds, so the reading holds for large relative to that capacity limit. H3 (loop-design hazard): self-training floors strictly higher than corrective ( vs ; vs ); its raw error falls over rounds (from to ), not rises — the predicted absolute rise did not occur in this instantiation, so the effect here is reinforcement relative to the corrective loop, and the non-monotone regime remains a real-model prediction (Figure 7a).
(a) loop-design fork
(b) floor vs (H2)
H5 (mitigation exchange rate). Both proposed remedies move the floor toward the oracle bound, and they occupy different regions of the cost–reliability plane (Figure 8). A decorrelated verifier ensemble — verifiers with independent blind regions, rejecting on any dissent — drives the effective blind-spot rate toward plus a residual agreement term (measured : at , above because the strong classifier still false-accepts occasionally outside the blind region) and cuts the floor from to , but pays verifier calls per query. Random oracle audits are far cheaper yet weaker: auditing of items lowers the floor only to , because most of the audit budget lands on non-blind items — an untargeted audit spends most of its checks where the verifier already sees. Neither reaches the oracle floor within the tested budget; the ensemble is the stronger lever, and the obvious refinement — auditing where the verifier is least certain rather than at random — is left to the real-model study.
The one honest correction. The conjecture anchored the corrective floor at . Both learning variants land below it — corrective at , self-training at — and only the frozen control sits on the anchor (, with no learning to spill). The predicted above-anchor regime for self-training was not observed: self-healing dominates even when accepted outputs are recycled as labels, so self-training floors above corrective rather than above . The ordering that held is frozen > self-training > corrective > oracle, which still separates the mechanisms, and the equality should be read as for the corrective loop, the gap quantifying self-healing. Pushing self-training across the anchor would need reinforcement strong enough to outweigh this spillover — a plausible real-model regime our synthetic student, which generalises readily, does not reach.
What this does and does not establish. Two confirmations are close to structural: because blind-region items are censored from the training pool, a persistent floor (H1) and its growth with blind mass (H2) follow almost arithmetically — they show the mechanism is self-consistent, not that it is large in practice. The genuinely informative outcomes are those the censoring does not force: the size of the self-healing gap, the H3 ordering (and the absence of an absolute rise), the verifier-induced excess against the matched control, and the H5 exchange rates. In all cases the blind spot’s existence is modelled (through ) while only its dynamics under the loop are emergent; none of it evidences the effect’s magnitude on real language models — which is why the real-model measurements of Section 5, not this study, carry the paper’s empirical weight. What the synthetic model adds is the one thing the real runs could not supply: the behaviour of a loop that improves the student, and thus direct evidence for the conservation floor the theory predicts.
8 Threats to validity
The measurements and the synthetic study are each exposed to characteristic failures, and we address them in turn. The floor claim depends on the loop improving the student; on real models it did not, and rather than treat that as a void premise we report it as a finding (Section 5) and fall back to the controlled model, where the sanity gate on confirms the loop does improve. In that controlled study, so that forgetting or optimizer instability cannot masquerade as decoupling, we train cumulatively from the base model and include the frozen-student control (D) to separate the two. If the blind-spot mass self-heals, that is not a failure but the negative result the conjecture already anticipates, named by the sign of the deviation from . The model-system critique — that math is not a genuinely fuzzy task — is met partly by the real-model measurements of Section 5, where the frontier verifier visibly false-accepts and false-rejects on GSM8K and hard MATH; extending to a fully fuzzy task via a validated super-verifier is named as open scope (Section 9), not claimed. Verification cost — the dominant expense, since is a frontier model — is controlled by deterministic caching of the pinned verifier.
The correlation corollary is the claim most exposed to challenge, precisely because it borders the self-preference literature [18]. The defense is built into the measurement: H2 is stated over the measured , with model family used only as a manipulation to spread across a range. The relation between the floor and therefore stands whether or not self-preference is the reason is high for same-family pairs.
9 Limitations
Distinct from the threats above — which defend the proposed design — these bound the paper’s own claims. (1) The synthetic blind region is a static input-space set, whereas real verifier blind spots are output-dependent and move with the student’s error distribution (assumption A1 of Appendix A); a soft or moving blind region could weaken conservation. (2) The anchor assumes a stationary query distribution. (3) When the teacher is the verifier , the teacher-error term is not independent of and the floor formula is optimistic — wrong labels would also accept enter training. (4) A -way classification task with a logistic student may not transfer to open-ended generation, where “the same wrong answer” is itself ill-defined. (5) The H5 exchange rates are properties of the synthetic geometry, not portable constants. (6) The real-model runs are single-seed: point estimates carry Wilson intervals (Section 5), but seed-to-seed variance in LoRA training is not bounded, so the per-round trajectory of the degrading loop should be read as one representative run, not an averaged curve — the static blind-spot sweeps, being training-free, do not share this caveat. The real-model measurements (Section 5) and the appendix model each resolve a subset of these.
10 Implications
Three consequences follow for anyone building cost-saving cascades with verifier feedback; the first two rest on what we measured, the third on the theory. First — measured — the cost dashboard lies: escalation rate and accept-stream accuracy read healthy (a flat ) while true delivered error swings to , because both are computed through the verifier and improve or hold steady as error hides in its blind spot. An independent audit channel — a small, periodic gold-labeled sample — is not optional instrumentation but the only instrument that can see the effect. Second — measured — verifier choice is a reliability decision, not only a cost decision: the blind spot grows with student capability and shrinks with verifier capability, so a cheaper verifier trades reliability for cost on an axis no dashboard shows, and buying the blind spot away with a frontier verifier returns the cost saving it was meant to provide. Third — from the theory — the loop-design choice between corrective and self-training is a safety choice: feeding accepted outputs back as positive labels converts a passive blind spot into an actively reinforced one.
None of this argues against cascades, which remain the most direct cost lever available, nor against closing the loop, which genuinely lowers cost. It argues that the reliability of a self-improving cascade must be measured outside the verifier that defines it, because a system optimized against a proxy will, given the chance, satisfy the proxy rather than the goal — and a cascade that retrains on its own verifier is given exactly that chance, round after round.
11 Conclusion
We measured the blind spot of a cost-saving cascade on real LLMs and found it moves adversarially: it grows with student capability and shrinks with verifier capability, so it is largest exactly in the cheap-student, cheap-verifier regime that makes cascades attractive, and buying it away with a frontier verifier returns the cost saving by escalating on nearly half of queries. We found that closing the loop with naive corrective fine-tuning does not improve a small student but degrades and collapses it, across every teacher — a direct caution against the parametric self-improving loop that fine-tunes the student on the verifier’s rejects. And we found that none of this registers on any metric a practitioner can compute: the dashboard reads a flat while true delivered error swings to . To explain that blindness we gave a two-population conservation law — the loop trains only on verifier-detectable errors, so user-facing error asymptotes to a floor rather than vanishing, and every in-loop metric improves while true quality does not — derived in a linear model (Appendix A) and validated in a synthetic study where the loop provably improves the student (Section 7). The clean floor itself we leave as theory, because no real-model loop we ran reached it. The through-line is practical: the reliability of a self-improving cascade cannot be read from any metric computed through its own verifier — it must be measured outside.
Appendix A Two-population model of the error floor
Track the two error masses (detectable) and (blind), with raw error and delivered error , where is the teacher’s error on rejected items. We model one round of each loop as a linear update under idealised assumptions: (A1) the verifier’s blind region is fixed; (A2) the query distribution is stationary; (A3) each round the loop corrects a fraction of the detectable mass, transfers a fraction of that reduction to the blind mass by generalisation (spillover), and — under self-training — recycles false accepts as labels, adding back of the blind mass:
Proposition 1. Under (A1)–(A3), with :
- 1.
(Frozen, , .) , so : exact conservation.
- 2.
(Corrective, , .) and
with equality iff ; the slack is self-healing.
- 3.
(Self-training, .) at every finite horizon ,
strictly once , and once the accumulated reinforcement outweighs . The linear reinforcement has no finite limit (blind mass would grow without bound as ), so this is a finite-horizon statement; a realistic saturating reinforcement — acting only on the not-yet-reinforced mass, capping total error at — gives a genuine elevated fixed point.
Proof sketch. For (i)–(ii), when (and when ); since with depends on the blind mass , not on , the frozen case gives even though never shrinks — persistent detectable mass only keeps escalation cost high. For (ii) the detectable reductions telescope, , so . In (iii) each step adds , so strictly exceeds the corrective floor at every ; the linear term diverges, hence the finite-horizon phrasing.
The synthetic study (Section 7) instantiates this with , , and corrective , giving a fitted spillover : generalisation transfers most of the detectable correction into the blind region, which is why the corrective floor sits well below the anchor. Self-training’s delivered error at , , lands between the corrective floor and — over this horizon its reinforcement offsets, but does not overcome, the spillover. Assumption (A1), the static blind region, is the one Section 9 flags as least realistic; a moving blind region is precisely what a real-model study with output-dependent verifiers would probe.
Appendix B Reproducibility
All code, data, and figure scripts are public at github.com/AltSlate-Labs/cascade-blindspot. The committed measurement files regenerate every figure without a GPU or API key; re-running the measurements from scratch needs one GPU and an OpenAI key.
All synthetic results regenerate from two self-contained scripts — phase0_synthetic.py (H1–H4) and h5_mitigation.py (H5) — in a few minutes of CPU each, at fixed seeds; the five figures are their direct output. Configuration: input dimension , classes, target a fixed random two-layer network (64 hidden, ); student = multinomial logistic regression on random ReLU features (); verifier = the same on features fit on 5000 labelled points (), with a blind region set at the -quantile of a fixed random projection; cumulative-from-base training with base pool , batch /round, held-out eval 3000, rounds, seeds; the reported floor is the last-three-round mean of . The ensemble of Figure 8 uses independent blind half-spaces. numpy 2.3, scikit-learn 1.8.
The real-model results (Section 5) regenerate from run_phase0.py (the loop), sweep_capability.py (student sweep), and math_sweep.py (verifier sweep on hard MATH), with make_real_figures.py rendering the figures. Students are Qwen2.5-Instruct (0.5B–32B) via transformers + LoRA (rank 16 on all seven linear modules, lr with warmup) on one H100; verifiers and teachers are the OpenAI API (gpt-4o-mini, gpt-4.1, gpt-5-mini), called concurrently and blinded to the gold answer; the oracle is GSM8K exact-match and math-verify symbolic equivalence on MATH-500 levels 4–5; held-out problems per condition. torch 2.13, transformers, peft.
References
- [1] (2023) FrugalGPT: how to use large language models while reducing cost and improving performance. arXiv:2305.05176. External Links: 2305.05176 Cited by: §1, §2.
- [2] (2022) Language model cascades. arXiv:2207.10342. External Links: 2207.10342 Cited by: §1, §2.
- [3] (2024) AutoMix: automatically mixing language models. In Advances in Neural Information Processing Systems (NeurIPS), External Links: 2310.12963 Cited by: §1, §2.
- [4] (2025) From deferral to learning: online in-context knowledge distillation for llm cascades. arXiv:2509.22984. External Links: 2509.22984 Cited by: §1, §2, §5.
- [5] (2025) In-context distillation with self-consistency cascades: a simple, training-free way to reduce llm agent costs. arXiv:2512.02543. External Links: 2512.02543 Cited by: §1, §2, §5.
- [6] (2023) Fast inference from transformers via speculative decoding. In International Conference on Machine Learning (ICML), External Links: 2211.17192 Cited by: §1.
- [7] (2023) Accelerating large language model decoding with speculative sampling. arXiv:2302.01318. External Links: 2302.01318 Cited by: §1.
- [8] (2023) Scaling laws for reward model overoptimization. In International Conference on Machine Learning (ICML), External Links: 2210.10760 Cited by: §1, §2.
- [9] (2022) Defining and characterizing reward hacking. In Advances in Neural Information Processing Systems (NeurIPS), External Links: 2209.13085 Cited by: §1, §2.
- [10] (2018) Categorizing variants of goodhart’s law. arXiv:1803.04585. External Links: 1803.04585 Cited by: §1, §2.
- [11] (2024) RouteLLM: learning to route llms with preference data. arXiv:2406.18665. External Links: 2406.18665 Cited by: §2.
- [12] (2023) Reinforced self-training (rest) for language modeling. arXiv:2308.08998. External Links: 2308.08998 Cited by: §2.
- [13] (2022) STaR: bootstrapping reasoning with reasoning. In Advances in Neural Information Processing Systems (NeurIPS), External Links: 2203.14465 Cited by: §2.
- [14] (2023) Weak-to-strong generalization: eliciting strong capabilities with weak supervision. arXiv:2312.09390. External Links: 2312.09390 Cited by: §2.
- [15] (2023) Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations (ICLR), External Links: 2203.11171 Cited by: §2.
- [16] (2024) Inference scaling flaws: the limits of llm resampling with imperfect verifiers. arXiv:2411.17501. External Links: 2411.17501 Cited by: §2.
- [17] (2023) Judging llm-as-a-judge with mt-bench and chatbot arena. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, External Links: 2306.05685 Cited by: §2.
- [18] (2024) LLM evaluators recognize and favor their own generations. In Advances in Neural Information Processing Systems (NeurIPS), External Links: 2404.13076 Cited by: §2, §8.
- [19] (2024) AI models collapse when trained on recursively generated data. Nature 631, pp. 755–759. Cited by: §2, §5.
- [20] (2024) Self-consuming generative models go mad. In International Conference on Learning Representations (ICLR), External Links: 2307.01850 Cited by: §2.
- [21] (2020) Pseudo-labeling and confirmation bias in deep semi-supervised learning. In International Joint Conference on Neural Networks (IJCNN), External Links: 1908.02983 Cited by: §2.
- [22] (2024) Large language models cannot self-correct reasoning yet. In International Conference on Learning Representations (ICLR), External Links: 2310.01798 Cited by: §2.
- [23] (2010) On the foundations of noise-free selective classification. Journal of Machine Learning Research 11, pp. 1605–1641. Cited by: §2.
- [24] (2020) Consistent estimators for learning to defer to an expert. In International Conference on Machine Learning (ICML), External Links: 2006.01862 Cited by: §2.
- [25] (2021) Measuring mathematical problem solving with the math dataset. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, External Links: 2103.03874 Cited by: §4, §5.
- [26] (2021) Training verifiers to solve math word problems. arXiv:2110.14168. External Links: 2110.14168 Cited by: §4, §5.