arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00482v1 [cs.CL] 31 Aug 2026

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

Qiaoyuan Zheng Affiliation: ETH Zurich Affiliation: Zurich, Switzerland Email: zqiaoyuan@ethz.ch    Yiqu Yang Affiliation: ETH Zurich Affiliation: Zurich, Switzerland Email: yangyiq@ethz.ch
Abstract

Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). In owner-disjoint folds, one owner half identifies items with low residual differential item functioning across model families (low-DIF); the resulting frozen, source- and easiness-balanced weights score models in the other half, while equally short matched-random subtests control for generic subtest variation. Full-benchmark and low-DIF rankings remain strongly correlated (τb=.900\tau_{b}=.900–.948.948). Yet in four of five benchmarks, 30.9–47.1% of cross-family pairs initially within one percentage point reverse order, exceeding their matched-random medians by 16.9–28.6 percentage points (all p=.001p=.001). The fifth benchmark shows no reliable excess (−0.9-0.9 points, p=.689p=.689). The pattern survives all pre-specified population perturbations, and residual item–family signatures replicate across owner halves; however, no family shows a consistent advantage across benchmarks. Thus, globally stable rankings can still leave individual near-tie orderings sensitive to benchmark composition, and sub-one-point leaderboard gaps should be accompanied by evidence that the implied ordering is composition-robust.

1 Introduction

Public leaderboards collapse thousands of item-level responses into a single score, often separating models by less than one percentage point. Such gaps do not establish that the implied ordering is robust to benchmark item composition. Rankings can change under prompt weighting or evaluation perturbations (Siska et al., 2024; Alzahrani et al., 2024), while conventional uncertainty analyses leave some close comparisons unresolved (Kotawala, 2026). We ask whether a near-tied ordering is robust to which benchmark items are included.

We study this question through family-dependent item functioning (family-DIF), adapting differential item functioning from psychometrics (Holland and Wainer, 1993). After a family-blind multidimensional adjustment, residual item–family effects measure systematic differences across observational model families. In owner-disjoint folds, one owner half selects source- and easiness-balanced low-DIF anchors whose frozen weights score the other half, after which the halves are swapped. Equally short matched-random subtests separate targeted composition sensitivity from generic shorter-subtest variation. DIF here is a diagnostic, not a causal or intrinsic family property.

Despite stable aggregate rankings, four of five benchmarks show excess near-tie reversals over matched-random controls, with no consistent family advantage.

2 Related work

Psychometrics for model evaluation.

IRT has been used to construct evaluation scales for NLP systems (Lalor et al., 2016), assess leaderboard and test-set discriminability and identify informative examples (Rodriguez et al., 2021; Vania et al., 2021), detect invalid or mislabeled benchmark items (Truong et al., 2025; Land and Bikel, 2026), and reduce evaluation cost (Polo et al., 2024; Hofmann et al., 2025; Zhou et al., 2026a). Recent methods learn text-conditioned IRT representations for model routing and benchmark prediction (Chen et al., 2025), or use fixed-parameter MIRT anchors to preserve comparability as benchmark suites evolve (Habba et al., 2026). In contrast, our spectral representation controls for multidimensional response structure, and our low-DIF anchors diagnose composition sensitivity rather than predict performance, link test forms, or detect mislabeled items.

Prior DIF work examines human–chatbot differences (Zeinfeld et al., 2026) and scoring procedures that downweight DIF items (Halpin, 2024). A course report closest to our setting finds that DIF-selected BBH pools alter held-out Llama–Qwen group gaps (Liu, 2026). We extend this baseline to owner-disjoint rankings across five benchmarks, individual near-tied pairs, and source- and easiness-matched random controls.

Ranking uncertainty and benchmark robustness.

Model rankings are sensitive to item weighting, prompt composition, and evaluation protocols (Mishra and Arunkumar, 2021; Siska et al., 2024; Kim et al., 2026; Alzahrani et al., 2024). Related work augments accuracy with model-output uncertainty (Ye et al., 2024), quantifies variance across training and evaluation choices (Madaan et al., 2024), shows that clustered benchmark structure can widen rank uncertainty (Neuhof and Benjamini, 2026), and tests whether adjacent model comparisons have adequate paired statistical resolution (Kotawala, 2026). Heineman et al. (2025) analyze ranking fidelity under training and checkpoint noise, while Qian et al. (2026) use discriminability and within-family inversions for subset selection. Our estimand instead isolates excess cross-family reversals among models initially within one percentage point, relative to equally short, composition-matched random subtests.

Measurement validity and interpretation.

Benchmark validity depends on the inferences supported by the measurement protocol (Bean et al., 2025). Multidimensional capability structure can provide an alternative explanation for apparent DIF (Zhou et al., 2026b), while conventional IRT estimators can be unreliable in AI benchmark regimes with few models or non-normal ability distributions (Jiang et al., 2026). Diagnostic conclusions can also be fragile (Aribandi et al., 2021), and textual correlates of DIF may reflect intended subdomains rather than construct-irrelevant content (Maeda and Lu, 2025). We therefore separate ranking consequences, owner-disjoint stability, and blinded semantic interpretation. Our residual effects remain conditional on a family-blind, prediction-selected spectral response representation and do not establish that every capability difference has been removed.

A criterion-level comparison is provided in Appendix A.

3 Audit design

A near-tied ordering is composition-robust if it survives downweighting items with residual family dependence. We estimate item–family effects in one owner half, construct a source- and easiness-balanced low-DIF subtest, and apply its frozen weights to the other half. Blueprint-matched random subtests separate targeted composition sensitivity from ordinary short-subtest variation.

3.1 Data and evaluation population

For each benchmark, we have a binary response matrix Y∈{0,1}I×MY\in\{0,1\}^{I\times M}, where Yi​m=1Y_{im}=1 when model mm answers item ii correctly. We audit MMLU-Pro (Wang et al., 2024), BIG-Bench Hard (Suzgun et al., 2023), MMLU (Hendrycks et al., 2021), HellaSwag (Zellers et al., 2019), and WinoGrande (Sakaguchi et al., 2021), using the item-level response matrices released with RouterEval (Huang et al., 2025).

Models are assigned to Qwen, Llama, Gemma, Mistral/Mixtral, or Phi via a frozen, unique model-name match. We cap each owner at 5 checkpoints and split each owner into two approximately equal halves. For ranking comparisons, we retain one checkpoint per owner–family pair, chosen as the checkpoint whose full-benchmark accuracy is closest to that pair’s median.

Exact model-name matching, owner extraction, owner capping, and outer-fold construction details are specified in Appendix C.1.

3.2 Family DIF with a spectral MIRT approximation

We use a family-label-free spectral approximation to the multidimensional item–model term in compensatory MIRT (Reckase, 2006). This is a response-derived spectral working model, rather than a traditional marginal maximum-likelihood MIRT estimator. The choice matches our goal of predictive nuisance adjustment rather than interpretation of latent traits. It permits computationally tractable repeated owner-/item-held-out fitting and anchor purification without specifying a parametric population distribution for model abilities. We do not claim that this approximation is generally superior to likelihood-based MIRT.

Let ηi​m(K)\eta^{(K)}_{im} denote the rank-KK reconstruction, which approximates the conventional MIRT term 𝐚i⊤​𝜽m\mathbf{a}_{i}^{\top}\boldsymbol{\theta}_{m}. For item ii and model mm from observational family f⁡(m)f(m), we define the linear predictor

ξi​m=ηi​m(K)+bi+δi,f⁡(m),\xi_{im}=\eta^{(K)}_{im}+b_{i}+\delta_{i,f(m)}, (1)

and response probability

Pr⁡(Yi​m=1)=σ⁡(ξi​m),σ⁡(x)=11+e−x,\Pr(Y_{im}=1)=\sigma(\xi_{im}),\qquad\sigma(x)=\frac{1}{1+e^{-x}}, (2)

where bib_{i} is item easiness and δi,f⁡(m)\delta_{i,f(m)} is the residual item-by-family effect.

We ran all downstream analyses using the common candidate grid K∈{1,2,4,8,16,32,64,128,256}K\in\{1,2,4,8,16,32,64,128,256\}, omitting dimensions that exceeded the rank of the corresponding fitting matrix. Within each outer owner half, we select KK using nested owner-held-out and item-held-out Bernoulli log loss and the one-standard-error rule (Hastie et al., 2009). Family labels are not used to learn the response directions or select KK, and the opposite outer owner half does not select the directions, dimension, or anchor items. Appendix C.3 gives the spectral construction, held-out coordinate estimation, and dimension-selection details.

With ηi​m(K)\eta^{(K)}_{im} fixed, we estimate item easiness and residual family effects by

(b^,δ^)=arg⁡minb,δ​{∑i,m[log⁡(1+eξi​m)−Yi​m​ξi​m]+λ2​∑i,fδi​f2},λ=1.(\widehat{b},\widehat{\delta})=\arg\min_{b,\delta}\left\{\sum_{i,m}\left[\log(1+e^{\xi_{im}})-Y_{im}\xi_{im}\right]+\frac{\lambda}{2}\sum_{i,f}\delta_{if}^{2}\right\},\qquad\lambda=1. (3)

For identifiability, family effects are centered within every item:

∑fδi​f=0for every item ​i.\sum_{f}\delta_{if}=0\qquad\text{for every item }i. (4)

Thus, δi​f\delta_{if} is family ff’s conditional deviation from that item’s across-family mean; the constraint does not imply that the item has no family effects.

We summarize residual family dependence by

Di=maxf⁡δ^i​f−minf⁡δ^i​f.D_{i}=\max_{f}\widehat{\delta}_{if}-\min_{f}\widehat{\delta}_{if}. (5)

A large DiD_{i} indicates that observational families remain separated on item ii after conditioning on the selected spectral MIRT approximation. It does not imply item unfairness or a causal family mechanism. Optimization and convergence details appear in Appendix C.4.

3.3 Cross-fitted residual low-DIF anchors

We use two-fold, owner-disjoint cross-fitting in the sample-splitting sense (Chernozhukov et al., 2018). In each direction, one owner half selects the spectral MIRT dimension, learns the response directions, estimates residual family DIF, and constructs the anchor weights. These weights are then frozen and applied to representative models from the opposite owner half. We swap the two halves so that every evaluated model is scored using anchors selected without its owner’s responses.

Anchor construction uses three fixed purification rounds, following the general logic of iterative DIF purification (Candell and Drasgow, 1988). Starting with all items, each round learns the spectral MIRT response directions from the current anchors, projects all items onto those directions, and re-estimates residual family DIF on all items. Within each native source, items are divided into four strata using the fitted residual item easiness bib_{i}. In every source-by-easiness cell cc, we retain the fraction α\alpha with the smallest DiD_{i}. The primary analysis uses α=.50\alpha=.50; α=.30\alpha=.30 and α=.70\alpha=.70 are sensitivity analyses. The selection and weights produced by the third round are then frozen.

Here, “anchor” denotes a diagnostic scoring subset rather than a scale-linking anchor: unlike fixed-parameter calibration (Habba et al., 2026), it is not used to preserve an IRT scale across evolving test forms.

If cell cc contains ncn_{c} items and retains kck_{c} anchors, the item weights are

wi={nc/kc,i​ is a selected anchor in cell ​c,0,otherwise.w_{i}=\begin{cases}n_{c}/k_{c},&i\text{ is a selected anchor in cell }c,\\ 0,&\text{otherwise}.\end{cases} (6)

Consequently,

∑i∈cwi=nc,\sum_{i\in c}w_{i}=n_{c}, (7)

so every source-by-easiness cell contributes exactly the same total mass as in the original benchmark. The anchor therefore changes residual-DIF composition without changing the benchmark’s source and coarse easiness blueprint.

For a held-out model mm, the full-benchmark and weighted anchor scores are

Am=1I​∑iYi​m,A~m=∑iwi​Yi​m∑iwi.A_{m}=\frac{1}{I}\sum_{i}Y_{im},\qquad\widetilde{A}_{m}=\frac{\sum_{i}w_{i}Y_{im}}{\sum_{i}w_{i}}. (8)

Appendix C.5 specifies the complete three-round procedure, deterministic within-cell selection, and final refitting step.

3.4 Ranking estimand and matched control

The primary comparison set comprises different-owner, different-family model pairs satisfying |Am−An|≤.01|A_{m}-A_{n}|\leq.01. A pair (m,n)(m,n) undergoes a strict rank reversal when

(Am−An)​(A~m−A~n)<0.(A_{m}-A_{n})(\widetilde{A}_{m}-\widetilde{A}_{n})<0. (9)

Pairs tied under either scoring rule are excluded from the corresponding denominator. We pool eligible-pair and reversal counts across the two held-out owner folds before computing the reversal rate. Secondary summaries include mean fold-specific Kendall τb\tau_{b} (Kendall, 1945), the all-cross-family reversal rate, and cumulative results at score-gap thresholds {0.25,0.5,1,2,3,5}\{0.25,0.5,1,2,3,5\} percentage points; one point is primary.

To isolate variation caused by using fewer items, we compare the residual low-DIF anchor with R=1,000R=1{,}000 blueprint-matched random subtests. Within each outer fold, every replicate samples without replacement the same number of items from each source-by-easiness cell and applies the same cell-restoring weights as the anchor. The controls therefore match its length, source composition, and coarse easiness composition without selecting items by DiD_{i}. Replicate reversal counts are pooled across folds using the same procedure as for the observed low-DIF rate.

Let TDIFT_{\mathrm{DIF}} denote the observed pooled reversal rate and TrT_{r} the rate for random replicate rr. We report the excess reversal rate

Δpp=100​(TDIF−median1≤r≤R⁡Tr),\Delta_{\mathrm{pp}}=100\left(T_{\mathrm{DIF}}-\operatorname{median}_{1\leq r\leq R}T_{r}\right), (10)

and the one-sided empirical randomization pp-value (Phipson and Smyth, 2010)

p=1+∑r=1R𝟏{Tr≥TDIF}R+1.p=\frac{1+\sum_{r=1}^{R}\mathbf{1}\{T_{r}\geq T_{\mathrm{DIF}}\}}{R+1}. (11)

Thus, Δpp>0\Delta_{\mathrm{pp}}>0 means that low-DIF scoring reverses near-tied pairs more often than equally short, blueprint-matched random subtests.

Separately, 2,000 source-block bootstrap resamples (Efron and Tibshirani, 1993) assess sensitivity to benchmark recomposition; single-source benchmarks use 100 deterministic contiguous item clusters. We report the 2.5th–97.5th percentile range of the recomposed statistic, not a confidence interval around the fixed-benchmark estimate. The near-tie pair set is frozen from the observed full-benchmark scores throughout all random-control and resampling analyses.

Appendices C.6 and C.7 give the exact sampling, aggregation, and fold pooling procedures. All implementation and dimension-selection settings are summarized in Appendix Table C.8.

4 Validation and Diagnostic Analyses

We test whether the ranking effect survives population changes, replicates across owner halves, and aligns with broad content categories.

4.1 Population robustness

We test sensitivity to owner concentration, checkpoint selection, lineage definition, and family composition. Relative to the owner-cap-5 baseline, the nine perturbations use owner caps of 1 and 3, a score-blind lexicographic representative, a conservative name-based lineage exclusion, or omit Gemma, Llama, Mistral/Mixtral, Phi, or Qwen in turn.

Each variant preserves the owner-disjoint cross-fit and refits the spectral response directions, residual DIF, anchors, and matched controls; only the benchmark- and fold-specific dimension KK from the primary analysis is frozen. The frozen criteria require every perturbation to retain positive excess reversals in at least three of five benchmarks, with positive matched-random p≤.05p\leq.05 results in at least three. New variants use 200 random controls; the baseline uses 1,000. Full specifications appear in Appendix D.2.

4.2 Item-signature replication and source attribution

We refit residual item–family signatures separately in the two owner halves and construct shared source-by-easiness cells from their averaged item intercepts. After within-cell adjustment, we report three cross-half replication statistics: the median family-wise Spearman correlation, top-20% high-DIF overlap, and advantaged-family agreement among overlapping items. Under independence, the correlation is centered at zero and expected top-20% overlap is 20%; agreement has no fixed baseline because family frequencies differ. We therefore assess all three statistics using 500 one-sided within-cell permutations and report their effect sizes and pp-values.

For MMLU-Pro, BBH, and MMLU, secondary analyses test whether discovery-selected extreme source effects reproduce in validation and whether exact source contributions to family score shifts retain their signs across folds. Full definitions appear in Appendix D.3.

4.3 Blinded content audit

To interpret the reproducible item–family signatures, we sampled 50 stable high-DIF items per benchmark and paired each with a control from the same source-by-easiness cell, yielding 250 matched pairs. High-DIF items were in the within-cell top 20% in both owner halves; controls were in neither top-20% set.

Two independently prompted language-model annotators, blinded to benchmark, group, DIF, and family information, labeled ten binary and three ordinal content axes. Their labels were analyzed separately.

A binary axis was considered confirmatory only if both annotators showed the same pooled direction, BH q≤.05q\leq.05, an absolute paired difference of at least eight percentage points, and the same direction in at least three benchmarks. Family-direction analyses were post-confirmatory and exploratory. Full details appear in Appendix D.4.

5 Results

We report the ranking effect, its robustness, and the replication and semantic audit of residual item–family signatures.

5.1 Global rankings are stable, but near-tie orderings are composition-sensitive

Table 1: Primary near-tie ranking results using the 50% residual low-DIF anchors. KK gives the selected spectral-MIRT dimensions in the two audit folds. τb\tau_{b} is the mean fold-specific Kendall correlation between the full-benchmark and low-DIF rankings. Close nn is the pooled number of different-owner, different-family model pairs separated by at most one percentage point on the full benchmark. Low-DIF is the pooled strict reversal rate for these pairs under anchor scoring; Random is the median reversal rate across 1,000 source-by-easiness-matched random subtests. Excess is Low-DIF minus Random in percentage points, and prandp_{\mathrm{rand}} is the corresponding one-sided empirical matched-random pp-value.
Benchmark KK τb\tau_{b} Close nn Low-DIF (%) Random (%) Excess (pp) prandp_{\mathrm{rand}}
MMLU-Pro 64/32 .924 308 47.1 18.5 +28.6 .001
BBH 128/128 .915 2,226 42.1 25.2 +16.9 .001
MMLU 128/128 .930 3,235 40.7 16.3 +24.4 .001
HellaSwag 16/64 .948 4,533 30.9 11.4 +19.5 .001
WinoGrande 8/8 .900 5,574 33.8 34.7 −0.9-0.9 .689

Table 1 reports the primary result after adjustment with the family-label-free spectral MIRT approximation. Full-benchmark and low-DIF anchor rankings remain strongly correlated (τb=.900\tau_{b}=.900–.948.948), with only 2.5–4.2% of all cross-family comparisons reversing. Among eligible pairs initially within one percentage point, however, four benchmarks show reversal rates of 30.9–47.1%, exceeding matched-random subtests by 16.9–28.6 percentage points (all matched-random p=.001p=.001).

For WinoGrande, low-DIF and matched-random reversal rates are similar (33.8% vs. 34.7%; excess −0.9-0.9 points, p=.689p=.689). Thus, the observed sensitivity is localized to near ties in four benchmarks, not wholesale leaderboard reranking.

Anchor-fraction sensitivity is reported in Appendix B.2.

5.2 Excess reversals persist across wider score bands

Figure 1 extends the maximum full-score gap from 0.25 to 5 percentage points. At the five-point threshold, all four primary-positive benchmarks still exceed matched-random controls by 5.2–19.1 percentage points, whereas WinoGrande shows no significant excess at any tested threshold. Results outside the pre-specified one-point band are descriptive sensitivity analyses.

Figure 1: Excess cross-family reversal rates, relative to the median matched-random subtest, across full-score gap thresholds. The vertical line marks the pre-specified one-percentage-point threshold; filled markers denote one-sided matched-random p≤.05p\leq.05. Curves outside the primary threshold are descriptive sensitivity analyses.

5.3 The result is robust to population specification

Across the baseline and all nine population perturbations—two owner caps, score-blind checkpoint selection, conservative lineage filtering, and five leave-one-family-out analyses shown in figure2—MMLU-Pro, BBH, MMLU, and HellaSwag consistently show positive excess reversals with matched-random p≤.05p\leq.05. WinoGrande reaches p≤.05p\leq.05 only when Mistral/Mixtral models are excluded.

Figure 2: Population robustness of excess near-tie reversals. Horizontal position gives the excess reversal rate in percentage points. Diamonds mark the primary specification; filled markers denote one-sided matched-random p≤.05p\leq.05. Score-blind checkpoint selects one model per owner–family by model-ID order without using benchmark accuracy; conservative lineage filtering excludes name-flagged merged, adapted, distilled, preference-tuned, hybrid, or uncensored derivatives.

Across the five benchmarks, every family changes the direction of its mean anchor-minus-full score shift at least once. Exact-common-model comparisons confirm several such reversals when the evaluated model population is held fixed. The shifts are therefore benchmark-dependent and should not be interpreted as intrinsic family rankings.

5.4 Residual item signatures replicate across owners

Table 2: Owner-disjoint replication of residual item–family signatures after source-by-easiness control. ρ\rho is the median family-wise cross-half Spearman correlation; overlap is discovery top-20% precision in the validation top-20% (independence baseline: 20%); agreement conditions on overlapping top items. All 15 one-sided within-cell permutation tests give p=.002p=.002.
Benchmark Median ρ\rho Top-20% overlap Advantaged-family agreement
MMLU-Pro .407 31.5% 56.4%
BBH .589 36.1% 72.5%
MMLU .501 38.5% 71.9%
HellaSwag .308 37.4% 44.3%
WinoGrande .465 38.5% 66.0%

All three statistics show above-chance owner-disjoint replication on every benchmark (Table 2). Median family-wise correlations are positive (ρ=.308\rho=.308–.589.589), while top-20% overlap is 31.5–38.5%, compared with a 20% independence baseline. Among overlapping top items, the same advantaged family is recovered in 44.3–72.5% of cases. Every statistic exceeds all 500 corresponding within-cell permutation values (one-sided p=.002p=.002; Bonferroni-adjusted p=.030p=.030 over 15 comparisons). WinoGrande remains informative: its signatures replicate without excess near-tie reversals, so replication alone does not imply a ranking effect.

Among the three source-rich benchmarks, discovery-selected source-effect signs replicate in 58 of 60 held-out checks. Exact source contributions to the anchor-minus-full family shifts retain their signs in 43 of 60 checks (both pooled exact p<.001p<.001). These analyses localize reproducible source structure but do not identify a cognitive or training-data mechanism.

5.5 The blinded audit finds no confirmatory semantic enrichment

Annotation agreement is κ=.715\kappa=.715 on the binary axes and ρ=.800\rho=.800 on the ordinal axes, but none of the ten pre-specified binary axes satisfies all dual-annotator confirmatory criteria. The largest effect is annotator A’s −9.6-9.6-percentage-point spatial/temporal difference (q=.106q=.106); annotator B estimates −0.8-0.8 points on the same axis (q=.902q=.902). Thus, the effect is neither significant after correction nor replicated across annotators. No exploratory family-direction association survives BH correction in either annotator.

6 Limitations and claim boundaries

Observational grouping.

Family labels are inferred from model identifiers and may conflate architecture, training, data, and uploader practices. The owner-disjoint design and population perturbations reduce specific artifacts but cannot identify a causal family mechanism.

Scope and estimand.

The five benchmarks are static, primarily closed-form, and drawn from one response collection; MMLU and MMLU-Pro share content lineage and are not independent replications (Hendrycks et al., 2021; Wang et al., 2024). Interactions between family grouping and other evaluation choices remain untested (Madaan et al., 2024; Alzahrani et al., 2024). Low-DIF anchors reweight observed items: they neither estimate future-item performance nor define a superior construct or item-removal rule. The one-point band is operational, and DIF does not imply unfairness.

Residual interpretation.

The family-blind spectral MIRT approximation remains linear. We evaluate powers-of-two dimensions through K=256K=256, subject to fitting-matrix rank; no one-SE choice reaches its feasible ceiling, although five of ten raw held-out-loss minima occur at that ceiling. Residual effects may therefore still include nonlinear, higher-dimensional, or unmeasured capability differences. Likewise, an audit using two model annotators and broad content axes cannot rule out semantic explanations. Our results establish conditional item–family interactions and their ranking consequences, not intrinsic capabilities, contamination, or broad leaderboard invalidity.

7 Conclusion

We audited whether cross-family leaderboard orderings separated by at most one percentage point are robust to benchmark item composition. Although aggregate rankings remain stable, four of five benchmarks show excess near-tie reversals under residual low-DIF scoring; the fifth does not exceed matched-random variation. Owner-disjoint estimation, matched controls, and population robustness localize this sensitivity to near-tied comparisons rather than a universal family advantage or broad leaderboard invalidity. Small score gaps should therefore be accompanied by composition-robustness checks, not treated as self-interpreting evidence of model superiority.

References

  • Alzahrani et al. (2024) N. Alzahrani, H. Alyahya, Y. Alnumay, S. Alrashed, S. Alsubaie, Y. Almushayqih, F. Mirza, N. Alotaibi, N. Al-Twairesh, A. Alowisheq, et al. When benchmarks are targets: revealing the sensitivity of large language model leaderboards. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 13787–13805. Cited by: §1, §2, §6.
  • Aribandi et al. (2021) V. Aribandi, Y. Tay, and D. Metzler How reliable are model diagnostics?. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 1778–1785. Cited by: §2.
  • Bean et al. (2025) A. M. Bean, R. O. Kearns, A. Romanou, F. S. Hafner, H. Mayne, J. Batzner, N. Foroutan Eghlidi, C. Schmitz, K. Korgul, H. Batra, et al. Measuring what matters: construct validity in large language model benchmarks. Advances in Neural Information Processing Systems 38. Cited by: §2.
  • Benjamini and Hochberg (1995) Y. Benjamini and Y. Hochberg Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society. Series B (Methodological) 57 (1), pp. 289–300. External Links: ISSN 00359246, Link Cited by: §D.4.
  • Candell and Drasgow (1988) G. L. Candell and F. Drasgow An iterative procedure for linking metrics and assessing item bias in item response theory. Applied Psychological Measurement 12 (3), pp. 253–260. External Links: Document Cited by: §3.3.
  • Chen et al. (2025) J. Chen, C. Wang, G. Zhang, P. Ye, L. Bai, W. Hu, Y. Qu, and S. Hu Learning compact representations of LLM abilities via item response theory. arXiv preprint arXiv:2510.00844. Cited by: §2.
  • Chernozhukov et al. (2018) V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal 21 (1), pp. C1–C68. External Links: ISSN 1368-4221, Document, Link, https://academic.oup.com/ectj/article-pdf/21/1/C1/27684918/ectj00c1.pdf Cited by: §3.3.
  • Cohen (1960) J. Cohen A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20 (1), pp. 37–46. External Links: Document Cited by: §D.4.
  • Efron and Tibshirani (1993) B. Efron and R. J. Tibshirani An introduction to the bootstrap. Chapman & Hall/CRC, New York, NY. Cited by: §3.4.
  • Habba et al. (2026) E. Habba, I. Itzhak, A. Yehudai, Y. Perlitz, E. Bandel, M. Shmueli-Scheuer, L. Choshen, and G. Stanovsky Growing pains: extensible and efficient LLM benchmarking via fixed parameter calibration. arXiv preprint arXiv:2604.12843. Cited by: Table 3, §2, §3.3.
  • Halko et al. (2011) N. Halko, P. G. Martinsson, and J. A. Tropp Finding structure with randomness: probabilistic algorithms for constructing approximate matrix decompositions. SIAM Review 53 (2), pp. 217–288. External Links: Document, Link, https://doi.org/10.1137/090771806 Cited by: §C.2.
  • Halpin (2024) P. F. Halpin Differential item functioning via robust scaling. Psychometrika 89 (3), pp. 796–821. Cited by: Table 3, §2.
  • Hastie et al. (2009) T. Hastie, R. Tibshirani, and J. Friedman Model assessment and selection. In The Elements of Statistical Learning: Data Mining, Inference, and Prediction, pp. 219–259. External Links: ISBN 978-0-387-84858-7, Document, Link Cited by: §3.2.
  • Heineman et al. (2025) D. Heineman, V. Hofmann, I. Magnusson, Y. Gu, N. Smith, H. Hajishirzi, K. Lo, and J. Dodge Signal and noise: a framework for reducing uncertainty in language model evaluation. Advances in Neural Information Processing Systems 38, pp. 17073–17114. Cited by: §2.
  • Hendrycks et al. (2021) D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt Measuring massive multitask language understanding. In International Conference on Learning Representations, External Links: Link Cited by: §3.1, §6.
  • Hofmann et al. (2025) V. Hofmann, D. Heineman, I. Magnusson, K. Lo, J. Dodge, M. Sap, P. W. Koh, C. Wang, H. Hajishirzi, and N. A. Smith Fluid language model benchmarking. arXiv preprint arXiv:2509.11106. Cited by: §2.
  • P. W. Holland and H. Wainer (Eds.) (1993) P. W. Holland and H. Wainer (Eds.) Differential item functioning. Lawrence Erlbaum Associates, Hillsdale, NJ. Cited by: §1.
  • Huang et al. (2025) Z. Huang, G. Ling, Y. Lin, Y. Chen, S. Zhong, H. Wu, and L. Lin RouterEval: a comprehensive benchmark for routing LLMs to explore model-level scaling up in LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 3860–3887. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §3.1.
  • Jiang et al. (2026) H. Jiang, S. Kwon, J. Luo, Z. Xiao, and S. Zhang Can we trust item response theory for AI evaluation?. External Links: 2607.15190, Document, Link Cited by: §2.
  • Kendall (1945) M. G. Kendall The treatment of ties in ranking problems. Biometrika 33 (3), pp. 239–251. External Links: Document Cited by: §3.4.
  • Kim et al. (2026) E. Kim, H. Yoo, G. Son, H. L. Patel, A. Agarwal, and A. Oh Benchmarks are not atomic: composition-aware LLM evaluation using BenchHub. In ICML 2026 Workshop on Combining Theory and Benchmarks: Towards A Virtuous Cycle to Understand and Guarantee Foundation Model Performance, Cited by: Table 3, §2.
  • Kotawala (2026) A. Kotawala Resolution diagnostics for paired LLM evaluation. arXiv preprint arXiv:2605.30315. Cited by: Table 3, §1, §2.
  • Lalor et al. (2016) J. P. Lalor, H. Wu, and H. Yu Building an evaluation scale using item response theory. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 648–657. Cited by: §2.
  • Land and Bikel (2026) S. Land and D. M. Bikel Auditing LLM benchmarks with item response theory. arXiv preprint arXiv:2605.30504. Cited by: §2.
  • Liu (2026) Y. Liu Benchmark items as measurement instruments: DIF diagnostics for response-derived LLM regimes on BBH. Note: CS321M Final Project, Stanford University External Links: Link Cited by: Table 3, §2.
  • Madaan et al. (2024) L. Madaan, A. K. Singh, R. Schaeffer, A. Poulton, S. Koyejo, P. Stenetorp, S. Narang, and D. Hupkes Quantifying variance in evaluation benchmarks. arXiv preprint arXiv:2406.10229. Cited by: §2, §6.
  • Maeda and Lu (2025) H. Maeda and Y. Lu Finding words associated with DIF: predicting differential item functioning using LLMs and explainable AI. Journal of Educational Measurement 62 (4), pp. 883–906. Cited by: §2.
  • McNemar (1947) Q. McNemar Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12 (2), pp. 153–157. External Links: Document Cited by: §D.4.
  • Mishra and Arunkumar (2021) S. Mishra and A. Arunkumar How robust are model rankings: a leaderboard customization approach for equitable evaluation. In Proceedings of the AAAI conference on Artificial Intelligence, Vol. 35, pp. 13561–13569. Cited by: §2.
  • Neuhof and Benjamini (2026) B. Neuhof and Y. Benjamini Quantifying ranking uncertainty in LLM benchmarks. arXiv preprint arXiv:2607.16259. Cited by: §2.
  • OpenAI (2026a) OpenAI GPT-5.4 model. Note: OpenAI API DocumentationAccessed: August 28, 2026 External Links: Link Cited by: §D.4.
  • OpenAI (2026b) OpenAI GPT-5.5 model. Note: OpenAI API DocumentationAccessed: August 28, 2026 External Links: Link Cited by: §D.4.
  • Phipson and Smyth (2010) B. Phipson and G. K. Smyth Permutation P-values should never be zero: calculating exact P-values when permutations are randomly drawn. Statistical Applications in Genetics and Molecular Biology 9 (1), pp. Article 39. External Links: Document Cited by: §3.4.
  • Polo et al. (2024) F. M. Polo, L. Weber, L. Choshen, Y. Sun, G. Xu, and M. Yurochkin tinyBenchmarks: evaluating LLMs with fewer examples. arXiv preprint arXiv:2402.14992. Cited by: §2.
  • Qian et al. (2026) Q. Qian, C. Huang, J. Xu, C. Lv, M. Wu, W. Liu, X. Wang, Z. Wang, Z. Huang, M. Tian, et al. Benchmark2{}^{2}: systematic evaluation of LLM benchmarks. arXiv preprint arXiv:2601.03986. Cited by: Table 3, §2.
  • Reckase (2006) M. D. Reckase Multidimensional item response theory. Handbook of statistics 26, pp. 607–642. Cited by: §3.2.
  • Rodriguez et al. (2021) P. Rodriguez, J. Barrow, A. M. Hoyle, J. P. Lalor, R. Jia, and J. L. Boyd-Graber Evaluation examples are not equally informative: how should that change NLP leaderboards?. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 4486–4503. Cited by: §2.
  • Sakaguchi et al. (2021) K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi WinoGrande: an adversarial Winograd schema challenge at scale. Communications of the ACM 64 (9), pp. 99–106. Cited by: §3.1.
  • Siska et al. (2024) C. Siska, K. Marazopoulou, M. Ailem, and J. Bono Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 10406–10421. Cited by: Table 3, §1, §2.
  • Spearman (1904) C. Spearman The proof and measurement of association between two things. The American Journal of Psychology 15 (1), pp. 72–101. External Links: Document Cited by: §D.3, §D.4.
  • Suzgun et al. (2023) M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. Le, E. H. Chi, D. Zhou, et al. Challenging BIG-Bench tasks and whether Chain-of-Thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, pp. 13003–13051. Cited by: §3.1.
  • Truong et al. (2025) S. Truong, Y. Tu, M. Hardy, A. Reuel-Lamparth, Z. Tang, J. Burapacheep, J. Perera, C. Uwakwe, B. Domingue, N. Haber, et al. Fantastic bugs and where to find them in AI benchmarks. Advances in Neural Information Processing Systems 38. Cited by: §2.
  • Vania et al. (2021) C. Vania, P. M. Htut, W. Huang, D. Mungra, R. Y. Pang, J. Phang, H. Liu, K. Cho, and S. Bowman Comparing test sets with item response theory. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 1141–1158. Cited by: §2.
  • Wang et al. (2024) Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, et al. MMLU-Pro: a more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems 37, pp. 95266–95290. Cited by: §3.1, §6.
  • Wilcoxon (1945) F. Wilcoxon Individual comparisons by ranking methods. Biometrics Bulletin 1 (6), pp. 80–83. External Links: ISSN 00994987, Link Cited by: §D.4.
  • Xu (2026) K. Xu IRT parameter drift across LLM generations and architectures on Open LLM Leaderboard v2. Note: CS321M Final Project, Stanford University External Links: Link Cited by: Table 3.
  • Ye et al. (2024) F. Ye, M. Yang, J. Pang, L. Wang, D. F. Wong, E. Yilmaz, S. Shi, and Z. Tu Benchmarking LLMs via uncertainty quantification. Advances in Neural Information Processing Systems 37, pp. 15356–15385. Cited by: §2.
  • Zeinfeld et al. (2026) L. Zeinfeld, A. Strugatski, Z. Bar-Dov, R. Blonder, S. Rap, and G. Alexandron Assessment design in the AI era: a method for identifying items functioning differentially for humans and chatbots. arXiv preprint arXiv:2603.23682. Cited by: §2.
  • Zellers et al. (2019) R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi HellaSwag: can a machine really finish your sentence?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4791–4800. Cited by: §3.1.
  • Zhou et al. (2026a) H. Zhou, H. Huang, Z. Zhao, L. Han, H. Wang, K. Chen, M. Yang, W. Bao, J. Dong, B. Xu, et al. Lost in benchmarks? rethinking large language model benchmarking with item response theory. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 35085–35093. Cited by: §2.
  • Zhou et al. (2026b) L. Zhou, L. Pacchiardi, F. Martínez-Plumed, K. M. Collins, Y. Moros-Daval, S. Zhang, Q. Zhao, Y. Huang, L. Sun, J. E. Prunty, et al. General scales unlock AI evaluation with explanatory and predictive power. Nature 652 (8108), pp. 58–67. Cited by: §2.

Appendix A Comparison with the closest related work

Table 3 shows the comparison with the closest related work.

Table 3: Comparison with the closest related work. ✓\checkmark indicates direct coverage; △\triangle indicates partial or adjacent coverage; ×\times indicates that the criterion is not studied.
Work Ability control family DIF Held-out validation Item-set intervention Matched controls Near-tie replication
Liu [Liu, 2026] ✓\checkmark ✓\checkmark ✓\checkmark ×\times ×\times
Siska et al. [Siska et al., 2024] ×\times ×\times ✓\checkmark ×\times △\triangle
Halpin [Halpin, 2024] △\triangle △\triangle ✓\checkmark ×\times ×\times
BenchHub [Kim et al., 2026] ×\times ×\times ✓\checkmark ×\times △\triangle
Kotawala [Kotawala, 2026] ×\times △\triangle ×\times ×\times △\triangle
Benchmark2 [Qian et al., 2026] △\triangle ✓\checkmark ✓\checkmark ×\times ×\times
Xu [Xu, 2026] △\triangle ×\times ×\times ×\times ×\times
Habba et al. [Habba et al., 2026] △\triangle ✓\checkmark ✓\checkmark △\triangle ×\times
Ours ✓\checkmark ✓\checkmark ✓\checkmark ✓\checkmark ✓\checkmark

Appendix B Additional Results

B.1 All score-gap thresholds

Table 4: Low-DIF reversal rate / matched-random median (%) at cumulative full-score-gap thresholds. The frozen primary threshold is one percentage point.
Benchmark 0.25pp\,\mathrm{pp} 0.5pp\,\mathrm{pp} 1pp\,\mathrm{pp} 2pp\,\mathrm{pp} 3pp\,\mathrm{pp} 5pp\,\mathrm{pp}
MMLU-Pro 43.8/37.1 45.8/28.2 47.1/18.5 40.0/9.8 33.1/7.0 23.6/4.5
BBH 49.1/42.3 46.6/36.0 42.1/25.2 33.9/14.1 26.9/9.6 18.0/5.9
MMLU 48.0/38.0 46.3/28.8 40.7/16.3 31.7/7.8 24.1/5.1 15.3/3.1
HellaSwag 43.5/32.8 39.3/21.1 30.9/11.4 19.0/5.7 12.4/3.7 7.4/2.2
WinoGrande 46.0/45.1 40.5/41.5 33.8/34.7 22.6/23.1 16.4/16.3 10.6/10.3

At the 0.25-point threshold, the eligible set falls to 89 pairs for MMLU-Pro, and only BBH, MMLU, and HellaSwag pass the matched-random test. Results outside the frozen one-point threshold are sensitivity analyses rather than additional confirmatory tests.

Details are shown in Table 4.

B.2 Anchor-fraction sensitivity

Table 5 reports descriptive sensitivity to the fraction of low-DIF items retained within each source-by-easiness cell. The primary specification retains 50% of items; 30% and 70% were pre-specified sensitivity settings.

Table 5: Sensitivity to the retained anchor fraction. Each cell reports mean fold-specific Kendall τb\tau_{b} / pooled strict near-tie reversal rate. Near-tied pairs are different-owner, different-family pairs within one percentage point on the full benchmark.
Benchmark 30% retained 50% retained 70% retained
MMLU-Pro .885 / 48.4% .924 / 47.1% .949 / 39.3%
BBH .882 / 44.3% .915 / 42.1% .951 / 33.9%
MMLU .903 / 41.9% .930 / 40.7% .955 / 36.7%
HellaSwag .912 / 40.1% .948 / 30.9% .970 / 21.6%
WinoGrande .855 / 37.7% .900 / 33.8% .939 / 28.5%

As expected, retaining fewer items produces more ranking changes, but near-tie reversals remain substantial across all three fractions. These results are descriptive: matched-random controls were computed for the primary 50% specification only, so confirmatory excess-reversal inference remains attached to that specification.

B.3 Family/owner robustness matrix

Table 6: Excess one-point reversal rate, in percentage points, under the primary specification and nine frozen owner/family robustness variants. Stars indicate nominal randomization p≤.05p\leq.05.
Variant MMLU-Pro BBH MMLU HellaSwag WinoGrande
Owner cap 1 +31.2∗ +16.4∗ +25.4∗ +18.6∗ −1.2-1.2
Owner cap 3 +28.2∗ +15.2∗ +22.1∗ +20.0∗ −0.8-0.8
Baseline cap 5 +28.6∗ +16.9∗ +24.4∗ +19.5∗ −0.9-0.9
Score-blind checkpoint +25.8∗ +15.1∗ +23.3∗ +19.3∗ +1.4
Conservative lineage filter +28.1∗ +19.1∗ +24.3∗ +17.9∗ +0.1
Omit Gemma +28.8∗ +21.1∗ +23.5∗ +23.3∗ +2.0
Omit Llama +33.3∗ +16.1∗ +34.1∗ +31.6∗ −1.5-1.5
Omit Mistral/Mixtral +27.6∗ +20.7∗ +25.3∗ +10.0∗ +6.4∗
Omit Phi +21.0∗ +15.5∗ +27.2∗ +25.7∗ +3.4
Omit Qwen +25.2∗ +13.7∗ +26.0∗ +27.7∗ −1.2-1.2

The baseline row uses 1,000 matched-random controls, so its minimum attainable randomization pp-value is 1/10011/1001; each of the other nine variants uses 200 controls, giving a minimum of 1/2011/201. Under the frozen robustness rule, a specification passes the direction criteria when at least three of the five benchmarks have positive excess and passes the significance criteria when at least three have nominal p≤.05p\leq.05. All ten specifications pass both criteria. More strongly, MMLU-Pro, BBH, MMLU, and HellaSwag—the four primary-positive benchmarks—remain positive and nominally significant in every specification. WinoGrande is significant only under omission of Mistral/Mixtral, so it remains an unstable boundary case rather than a robust positive result. The details are shown in Table 6.

B.4 Exact-common-model directional checks

Table 7: Three illustrative changes in family mean anchor-minus-full score shifts across benchmarks, evaluated on exactly the same family-specific models in each benchmark pair. Shifts and differences are reported in percentage points; 95% confidence intervals are obtained from 2,000 owner-bootstrap replicates.
Exact-common pair Family Left shift Right shift Right−-left [95% CI]
MMLU / HellaSwag Gemma +2.04 −0.51-0.51 −2.55-2.55 [−3.07-3.07, −1.91-1.91]
HellaSwag / WinoGrande Llama −1.48-1.48 +0.14 +1.63 [+1.37, +1.88]
MMLU-Pro / BBH Gemma −0.28-0.28 +0.82 +1.10 [+0.38, +1.94]

Table 7 reports three compact examples used in the main text rather than an exhaustive search for the largest contrasts. Each comparison holds the family-specific model population fixed across the two benchmarks, ruling out changes in model membership as the explanation for the directional change. The released file contains all 15 family estimates for these three exact-common benchmark pairs, including sample sizes and confidence intervals. These shifts are descriptive consequences of benchmark composition and should not be interpreted as intrinsic or causal family rankings.

B.5 Blinded content-audit results

Using the 250 exact-cell matched pairs and two blinded annotators defined in Appendix D.4, Figure 3 reports the pooled paired prevalence differences for the ten prespecified confirmatory binary axes. Annotation reliability passes the frozen criteria. Across the non-degenerate binary axes, the median Cohen’s κ\kappa is .715; across the three ordinal axes, the median Spearman ρ\rho is .800.

No binary axis passes the dual-annotator confirmatory enrichment criteria. The largest absolute difference for annotator A is −9.6-9.6 percentage points on spatial/temporal content (q=.106q=.106), while annotator B estimates −0.8-0.8 points on the same axis (q=.902q=.902). The largest absolute difference for annotator B is +4.0 points on distractor discrimination (q=.902q=.902). Thus, no axis is both significant after correction and replicated across annotators.

A post-confirmatory diagnostic is restricted to the 158 of 250 high-residual-DIF items for which the advantaged-family label agrees across owner halves. No benchmark-stratified family-direction association survives BH correction in either annotator, and no dual-annotator exploratory signal is detected. These analyses remain exploratory and are not used as a semantic explanation of the ranking effect.

Figure 3: Paired prevalence differences between stable high-DIF items and exact-cell matched controls on the ten confirmatory binary content axes. Dashed lines mark the frozen 8-percentage-point effect-size requirement. Annotator A crosses the magnitude threshold only for spatial/temporal content (−9.6-9.6 points), but the effect is not significant after BH correction (q=.106q=.106) and is not replicated by annotator B (−0.8-0.8 points). No axis passes the dual-annotator confirmatory criteria.

Appendix C Implementation details for the MIRT-adjusted audit

This appendix specifies the numerical implementation underlying Section 3. The main text defines the statistical estimand; here we describe the spectral MIRT working representation, dimension selection, residual-DIF optimization, anchor purification, and resampling procedures.

C.1 Model-family and owner construction

Family membership is assigned by a unique, case-insensitive match on model ID. The frozen matching rules are:

Family Model-ID pattern
Qwen qwen
Llama llama
Gemma gemma
Mistral/Mixtral mistral or mixtral
Phi a token beginning with phi

Models matching zero or multiple family patterns are excluded. The owner is the uploader component of the model ID. If an owner contributes more than the frozen cap of five checkpoints to one family, five are selected using the frozen population-selection seed.

Entire owners are assigned to one of two outer halves using a deterministic greedy procedure that approximately balances, within every family, both the number of retained checkpoints and aggregate full-benchmark accuracy. No owner may appear in both halves.

In benchmark order (MMLU-Pro, BBH, MMLU, HellaSwag, and WinoGrande), the complete-response matrices contain respectively 12,032, 5,761, 14,042, 10,042, and 1,267 items and 1,823, 3,811, 5,000, 5,000, and 5,000 models. The corresponding ranking populations contain 204, 497, 673, 673, and 673 owner–family representatives.

C.2 Spectral working-response construction

Let 𝒜\mathcal{A} denote the models in the current audit half and let M𝒜=|𝒜|M_{\mathcal{A}}=|\mathcal{A}|. For every item, we calculate a smoothed marginal success probability

p^i=∑m∈𝒜Yi​m+1/2M𝒜+1,si=p^i​(1−p^i).\widehat{p}_{i}=\frac{\sum_{m\in\mathcal{A}}Y_{im}+1/2}{M_{\mathcal{A}}+1},\qquad s_{i}=\sqrt{\widehat{p}_{i}(1-\widehat{p}_{i})}. (12)

We then form the standardized working-response matrix

Xi​m=Yi​m−p^isi.X_{im}=\frac{Y_{im}-\widehat{p}_{i}}{s_{i}}. (13)

Suppose 𝒮\mathcal{S} is the item set currently used to estimate the model coordinates. We compute a rank-KK randomized singular-value decomposition [Halko et al., 2011]

𝐗𝒮≈𝐔K​𝚺K​𝐕K⊤.\mathbf{X}_{\mathcal{S}}\approx\mathbf{U}_{K}\mathbf{\Sigma}_{K}\mathbf{V}_{K}^{\top}. (14)

The model-coordinate matrix is

𝚯=𝐕K,\mathbf{\Theta}=\mathbf{V}_{K}, (15)

and loadings for all items, including those outside 𝒮\mathcal{S}, are obtained by projection:

𝐋=𝐗​𝚯.\mathbf{L}=\mathbf{X}\mathbf{\Theta}. (16)

Thus the rank-KK reconstruction of the standardized response residual is 𝐋​𝚯⊤\mathbf{L}\mathbf{\Theta}^{\top}. We convert it into the fixed offset used in the logistic residual-DIF model:

oi​m=clip(𝐋i,:𝚯m,:⊤si,−6,6).o_{im}=\operatorname{clip}\left(\frac{\mathbf{L}_{i,:}\mathbf{\Theta}_{m,:}^{\top}}{s_{i}},-6,6\right). (17)

The truncation at six logits is a numerical safeguard and is not tuned using family labels or ranking outcomes.

This construction is a spectral approximation to a multidimensional logistic response surface. It is not a full maximum-likelihood calibration of a traditional 2PL MIRT model; we therefore refer to it as a spectral MIRT working model.

C.3 Nested selection of the MIRT dimension

For outer audit half hh, the feasible candidate set is

𝒦h={K∈{1,2,4,8,16,32,64,128,256}:K≤rank⁡(XS,h)}.\mathcal{K}_{h}=\left\{K\in\{1,2,4,8,16,32,64,128,256\}:K\leq\operatorname{rank}(X_{S,h})\right\}.

The feasible grid reaches K=128K=128 in both MMLU-Pro folds and K=256K=256 in all folds of the other four benchmarks. Dimension selection is performed separately inside each outer audit half.

We first assign 25% of audit-half owners to an inner validation set. Using only the remaining owners, we compute Equation 13 and a rank-KK decomposition. For dimension KK, define the training-owner item factor matrix

𝐅K=𝐔K​𝚺K.\mathbf{F}_{K}=\mathbf{U}_{K}\mathbf{\Sigma}_{K}. (18)

Within every native source, items are randomly divided into a coordinate- calibration set 𝒞\mathcal{C} and an evaluation set ℰ\mathcal{E}, with approximately half assigned to each. For inner-validation model mm, its coordinates are estimated only from 𝒞\mathcal{C}:

𝜽^m,K=(𝐅𝒞,K⊤​𝐅𝒞,K+λθ​𝐈K)−1​𝐅𝒞,K⊤​𝐗𝒞,m,λθ=1.\widehat{\boldsymbol{\theta}}_{m,K}=\left(\mathbf{F}_{\mathcal{C},K}^{\top}\mathbf{F}_{\mathcal{C},K}+\lambda_{\theta}\mathbf{I}_{K}\right)^{-1}\mathbf{F}_{\mathcal{C},K}^{\top}\mathbf{X}_{\mathcal{C},m},\qquad\lambda_{\theta}=1. (19)

For evaluation item i∈ℰi\in\mathcal{E}, the predicted logit is

η^i​m,K=logit⁡(p^i)+clip⁡(𝐅i,K⊤​𝜽^m,Ksi,−6,6).\widehat{\eta}_{im,K}=\operatorname{logit}(\widehat{p}_{i})+\operatorname{clip}\left(\frac{\mathbf{F}_{i,K}^{\top}\widehat{\boldsymbol{\theta}}_{m,K}}{s_{i}},-6,6\right). (20)

The model-level validation loss is

Lm,K=1|ℰ|​∑i∈ℰ[log⁡(1+exp⁡(η^i​m,K))−Yi​m​η^i​m,K].L_{m,K}=\frac{1}{|\mathcal{E}|}\sum_{i\in\mathcal{E}}\left[\log(1+\exp(\widehat{\eta}_{im,K}))-Y_{im}\widehat{\eta}_{im,K}\right]. (21)

Let L¯K\overline{L}_{K} be the mean of Lm,KL_{m,K} over inner-validation models, and define

SEK=sdm⁡(Lm,K)Mval.\operatorname{SE}_{K}=\frac{\operatorname{sd}_{m}(L_{m,K})}{\sqrt{M_{\mathrm{val}}}}. (22)

Let Kbest=arg⁡minK∈𝒦h⁡L¯KK_{\mathrm{best}}=\arg\min_{K\in\mathcal{K}_{h}}\overline{L}_{K}. Following the one-standard-error rule, we select

K⋆=min⁡{K∈𝒦h:L¯K≤L¯Kbest+SEKbest}.K^{\star}=\min\left\{K\in\mathcal{K}_{h}:\overline{L}_{K}\leq\overline{L}_{K_{\mathrm{best}}}+\mathrm{SE}_{K_{\mathrm{best}}}\right\}. (23)

The selected K⋆K^{\star} is frozen before family labels are introduced and is used for every anchor fraction within that outer fold.

The one-SE selection remains below the feasible candidate ceiling in every fold. The raw held-out-loss minimum occurs at the feasible ceiling in five of the ten folds, so the selected representation should not be interpreted as exhausting all possible capability structure.

C.4 Residual-DIF optimization

For a fixed spectral MIRT offset oi​mo_{im}, the response logit is

ξi​m=bi+oi​m+δi,f⁡(m).\xi_{im}=b_{i}+o_{im}+\delta_{i,f(m)}. (24)

Let 𝒜\mathcal{A} denote the models in the current fitting owner half and M𝒜=|𝒜|M_{\mathcal{A}}=|\mathcal{A}|. We initialize

p~i\displaystyle\widetilde{p}_{i} =clip⁡(1M𝒜​∑m∈𝒜Yi​m,10−4,1−10−4),\displaystyle=\operatorname{clip}\!\left(\frac{1}{M_{\mathcal{A}}}\sum_{m\in\mathcal{A}}Y_{im},10^{-4},1-10^{-4}\right), (25)
bi(0)\displaystyle b_{i}^{(0)} =logit⁡(p~i)−1M𝒜​∑m∈𝒜oi​m,\displaystyle=\operatorname{logit}(\widetilde{p}_{i})-\frac{1}{M_{\mathcal{A}}}\sum_{m\in\mathcal{A}}o_{im}, (26)

and set all family effects to zero. Here clip⁡(x,ℓ,u)=max⁡{ℓ,min⁡(x,u)}\operatorname{clip}(x,\ell,u)=\max\{\ell,\min(x,u)\}.

At the current parameter values, let

pi​m=σ⁡(ξi​m).p_{im}=\sigma(\xi_{im}).

For fixed family effects, the score and information for item intercept bib_{i} are

gbi\displaystyle g_{b_{i}} =∑m(Yi​m−pi​m),\displaystyle=\sum_{m}(Y_{im}-p_{im}), (27)
Hbi\displaystyle H_{b_{i}} =∑mpi​m​(1−pi​m).\displaystyle=\sum_{m}p_{im}(1-p_{im}). (28)

We apply the clipped Newton update

bi←bi+clip⁡(gbiHbi+10−9,−2,2).b_{i}\leftarrow b_{i}+\operatorname{clip}\!\left(\frac{g_{b_{i}}}{H_{b_{i}}+10^{-9}},-2,2\right). (29)

For family ff, let

ℳf={m:f⁡(m)=f}.\mathcal{M}_{f}=\{m:f(m)=f\}.

Holding the item intercept fixed, the penalized score and information for δi​f\delta_{if} are

gδi​f\displaystyle g_{\delta_{if}} =∑m∈ℳf(Yi​m−pi​m)−λ​δi​f,\displaystyle=\sum_{m\in\mathcal{M}_{f}}(Y_{im}-p_{im})-\lambda\delta_{if}, (30)
Hδi​f\displaystyle H_{\delta_{if}} =∑m∈ℳfpi​m​(1−pi​m)+λ,\displaystyle=\sum_{m\in\mathcal{M}_{f}}p_{im}(1-p_{im})+\lambda, (31)

with λ=1\lambda=1. We update

δi​f←δi​f+clip⁡(gδi​fHδi​f,−2,2).\delta_{if}\leftarrow\delta_{if}+\operatorname{clip}\!\left(\frac{g_{\delta_{if}}}{H_{\delta_{if}}},-2,2\right). (32)

The probabilities pi​mp_{im} are recomputed at the current parameter values during each Newton iteration.

After every complete family-effect sweep, we impose the identifying constraint by setting

δ¯i=1F​∑fδi​f,δi​f←δi​f−δ¯i,bi←bi+δ¯i.\overline{\delta}_{i}=\frac{1}{F}\sum_{f}\delta_{if},\qquad\delta_{if}\leftarrow\delta_{if}-\overline{\delta}_{i},\qquad b_{i}\leftarrow b_{i}+\overline{\delta}_{i}. (33)

The simultaneous adjustment of bib_{i} preserves every linear predictor while ensuring ∑fδi​f=0\sum_{f}\delta_{if}=0.

We use at most five alternating intercept/family-effect cycles. Each Newton subproblem uses at most 30 iterations and terminates when the largest absolute update is below 10−810^{-8}. The outer alternation also terminates early when the largest family-effect update is below this threshold.

C.5 Anchor purification algorithm

For each outer fold and anchor fraction α\alpha, purification proceeds as follows.

  1. 1.

    Initialize

    𝒮(0)={1,…,I}.\mathcal{S}^{(0)}=\{1,\ldots,I\}.
  2. 2.

    For rounds t=1,2,3t=1,2,3:

    1. (a)

      Fit the rank-K⋆K^{\star} spectral MIRT representation using the response rows in 𝒮(t−1)\mathcal{S}^{(t-1)}.

    2. (b)

      Project all items onto the fitted model coordinates and construct offsets oi​m(t)o_{im}^{(t)}.

    3. (c)

      Fit bi(t)b_{i}^{(t)} and δi​f(t)\delta_{if}^{(t)} on all items conditional on oi​m(t)o_{im}^{(t)}.

    4. (d)

      Compute

      Di(t)=maxf⁡δi​f(t)−minf⁡δi​f(t).D_{i}^{(t)}=\max_{f}\delta_{if}^{(t)}-\min_{f}\delta_{if}^{(t)}.
    5. (e)

      Within each native source, order items by bi(t)b_{i}^{(t)} and divide them into four near-equal strata. Ties are resolved by original item index.

    6. (f)

      Within every source-by-easiness cell cc, retain

      kc=max⁡{1,round⁡(α​nc)}k_{c}=\max\{1,\operatorname{round}(\alpha n_{c})\}

      items with the smallest Di(t)D_{i}^{(t)}. Their union defines 𝒮(t)\mathcal{S}^{(t)}.

  3. 3.

    Assign final weights

    wi={nc/kc,i∈𝒮(3)∩c,0,otherwise.w_{i}=\begin{cases}n_{c}/k_{c},&i\in\mathcal{S}^{(3)}\cap c,\\ 0,&\text{otherwise}.\end{cases}
  4. 4.

    Refit the spectral representation and residual-DIF parameters once on 𝒮(3)\mathcal{S}^{(3)} to obtain the stored final item signatures. Ranking scores use the weights selected at the end of round three.

Purification always runs for three rounds; convergence of the anchor set is recorded through successive Jaccard overlap but is not used as a stopping rule.

C.6 Matched-random subtests

For each outer fold, let 𝒮c\mathcal{S}_{c} be the final low-DIF anchors in cell cc and kc=|𝒮c|k_{c}=|\mathcal{S}_{c}|. Random-control replicate rr independently samples

𝒮c(r)⊆{i:c⁡(i)=c},|𝒮c(r)|=kc,\mathcal{S}_{c}^{(r)}\subseteq\{i:c(i)=c\},\qquad\left|\mathcal{S}_{c}^{(r)}\right|=k_{c}, (34)

uniformly without replacement. Sampled items receive weight nc/kcn_{c}/k_{c}.

The same replicate index is pooled across the two outer target halves:

Tr=∑hNflip,h(r)∑hNeligible,h(r).T_{r}=\frac{\sum_{h}N_{\mathrm{flip},h}^{(r)}}{\sum_{h}N_{\mathrm{eligible},h}^{(r)}}. (35)

The reported matched-random median, interval, excess, and randomization pp-value are calculated from {Tr}r=11000\{T_{r}\}_{r=1}^{1000}.

C.7 Composition bootstrap

For benchmarks with at least five native sources, native source labels define bootstrap units. Otherwise, the original item order is divided deterministically into 100 contiguous, near-equal clusters.

Let BB be the number of bootstrap units and let

(N1(r),…,NB(r))∼Multinomial⁡(B,1B,…,1B).\left(N_{1}^{(r)},\ldots,N_{B}^{(r)}\right)\sim\operatorname{Multinomial}\left(B;\frac{1}{B},\ldots,\frac{1}{B}\right). (36)

be the unit multiplicities for replicate rr.

For unit bb, let Sb​mS_{bm} and nbn_{b} be the model’s full-score successes and item mass, and let S~b​m\widetilde{S}_{bm} and n~b\widetilde{n}_{b} be their weighted anchor equivalents. Bootstrap scores are

Am(r)=∑bNb(r)​Sb​m∑bNb(r)​nb,A~m(r)=∑bNb(r)​S~b​m∑bNb(r)​n~b.A_{m}^{(r)}=\frac{\sum_{b}N_{b}^{(r)}S_{bm}}{\sum_{b}N_{b}^{(r)}n_{b}},\qquad\widetilde{A}_{m}^{(r)}=\frac{\sum_{b}N_{b}^{(r)}\widetilde{S}_{bm}}{\sum_{b}N_{b}^{(r)}\widetilde{n}_{b}}. (37)

The close-pair mask is defined once from the observed full scores and held fixed across the 2,000 bootstrap replicates.

C.8 Implementation settings

Table C.8 summarizes the frozen primary-audit specification.

Table C.8 summarizes the frozen primary-audit specification. Component Frozen Setting Families Qwen, Llama, Gemma, Mistral/Mixtral, Phi Owner cap 5 checkpoints per owner–family Minimum capped family size 20 models Outer split owner-disjoint two-fold cross-fit MIRT candidate dimensions {1,2,4,8,16,32,64,128,256}\{1,2,4,8,16,32,64,128,256\};
rank-infeasible values omitted
Selected KK (two folds) MMLU-Pro 64/32; BBH 128/128; MMLU 128/128;
HellaSwag 16/64; WinoGrande 8/8
Dimension rule minimum KK within one SE of best loss Inner validation owners 25% Coordinate-calibration items 50% within source Coordinate ridge λθ\lambda_{\theta} 1 Maximum MIRT offset 6 logits Randomized SVD iterations 5 DIF penalty λδ\lambda_{\delta} 1 DIF fitting cycles 5 Newton iterations per update at most 30 Newton tolerance 10−810^{-8} Maximum Newton step 2 logits Purification rounds 3 Easiness strata 4 within each source Primary anchor fraction 50% Anchor sensitivity 30%, 70% Primary close-pair band 1 percentage point Gap sensitivity 0.25, 0.5, 1, 2, 3, 5 points Matched-random replicates 1,000 Composition-bootstrap replicates 2,000 Single-source bootstrap units 100 contiguous clusters

Appendix D Validation and Robustness: Implementation Details

D.1 Controls and uncertainty

Eligible pairs and pooled reversal rate.

Let h∈{D,V}h\in\{D,V\} index the target owner half and let aa denote an alternative scoring rule, either the low-DIF anchor score or one matched-random score. For threshold ϵ\epsilon, define

𝒫h(a)​(ϵ)={(m,m′):m<m′,om≠om′,fm≠fm′,0<|smfull−sm′full|≤ϵ,sm​h(a)≠sm′​h(a)},\mathcal{P}_{h}^{(a)}(\epsilon)=\left\{(m,m^{\prime}):\begin{array}[]{l}m<m^{\prime},\quad o_{m}\neq o_{m^{\prime}},\quad f_{m}\neq f_{m^{\prime}},\\[2.0pt] 0<\left|s_{m}^{\mathrm{full}}-s_{m^{\prime}}^{\mathrm{full}}\right|\leq\epsilon,\quad s_{mh}^{(a)}\neq s_{m^{\prime}h}^{(a)}\end{array}\right\}, (38)

where omo_{m} and fmf_{m} denote the owner and family of model mm. Thus, eligible models must come from different owners and different families, must be separated by no more than ϵ\epsilon on the full benchmark, and must not be tied under either scoring rule.

For an eligible pair, the strict reversal indicator is

Rh​m​m′(a)=[(smfull−sm′full)(sm​h(a)−sm′​h(a))<0].R_{hmm^{\prime}}^{(a)}=\mathbf{1}\!\left[\left(s_{m}^{\mathrm{full}}-s_{m^{\prime}}^{\mathrm{full}}\right)\left(s_{mh}^{(a)}-s_{m^{\prime}h}^{(a)}\right)<0\right]. (39)

The pooled reversal rate sums reversal counts and eligible-pair counts across the two cross-fitting directions:

T^ϵ(a)=∑h∑(m,m′)∈𝒫h(a)​(ϵ)Rh​m​m′(a)∑h|𝒫h(a)​(ϵ)|.\widehat{T}_{\epsilon}^{(a)}=\frac{\displaystyle\sum_{h}\sum_{(m,m^{\prime})\in\mathcal{P}_{h}^{(a)}(\epsilon)}R_{hmm^{\prime}}^{(a)}}{\displaystyle\sum_{h}\left|\mathcal{P}_{h}^{(a)}(\epsilon)\right|}. (40)

The primary analysis sets ϵ=0.01\epsilon=0.01, corresponding to a one-percentage-point full-benchmark score gap.

Resampling controls.

Appendix C.6 specifies the matched-random construction and fold pooling, while Appendix C.7 specifies the composition resampling. The present subsection defines eligible pairs and the additional threshold and owner-bootstrap analyses.

Threshold sensitivity.

The primary specification uses a 50% anchor fraction and defines near-tied models as those separated by at most one percentage point on the full benchmark. We additionally evaluate

ϵ∈{0.0025,0.005,0.01,0.02,0.03,0.05},\epsilon\in\{0.0025,0.005,0.01,0.02,0.03,0.05\}, (41)

corresponding to full-score gaps of {0.25,0.5,1,2,3,5}\{0.25,0.5,1,2,3,5\} percentage points. Pair membership at every threshold is frozen from the observed full-benchmark scores. The matched-random distribution is recomputed using 1,000 replicates, and composition uncertainty is recomputed using 2,000 bootstrap replicates. We also repeat anchor construction using the lowest-DIF 30% and 70% of items within each blueprint cell. These specifications are sensitivity analyses and are not used to select the primary threshold or anchor fraction.

Owner bootstrap.

Family-specific score and percentile shifts may be correlated within a model owner. We therefore resample owners rather than checkpoints. Let

dm=smanchor−smfulld_{m}=s_{m}^{\mathrm{anchor}}-s_{m}^{\mathrm{full}} (42)

be the score shift for representative model mm. In bootstrap replicate bb, we sample the observed owners with replacement, using a sample size equal to the number of unique owners, and include all representative records associated with each sampled owner. For family ff, the bootstrap mean is

μ^f(b)=1|ℳf(b)|​∑m∈ℳf(b)dm.\widehat{\mu}_{f}^{(b)}=\frac{1}{|\mathcal{M}_{f}^{(b)}|}\sum_{m\in\mathcal{M}_{f}^{(b)}}d_{m}. (43)

We analogously calculate the mean percentile-rank shift and report empirical 95% intervals over 2,000 replicates. Owner-bootstrap intervals are used only for family-specific shifts; uncertainty in the primary reversal statistic is handled by the composition bootstrap and matched-random control.

D.2 Population robustness

The primary population retains at most five checkpoints per owner–family and selects one ranking representative nearest the within-owner–family median full-benchmark accuracy. Entire owners, including all of their family-specific records, remain assigned to a single outer half. We evaluate nine perturbations of this construction:

  1. 1.

    reduce the owner cap from five checkpoints to one;

  2. 2.

    reduce the owner cap from five checkpoints to three;

  3. 3.

    select one checkpoint per owner–family lexicographically by model identifier, without using benchmark accuracy;

  4. 4.

    exclude model identifiers indicating merged, adapter-based, distilled, hybrid, preference-tuned, or uncensored lineages; and

  5. 5.

    omit each of Gemma, Llama, Mistral/Mixtral, Phi, and Qwen in turn.

The conservative lineage filter excludes identifiers matching terms including merge, slerp, franken, lora, qlora, adapter, distill, hybrid, fusion, dpo, ppo, orpo, kto, abliterat, and uncensor. Each retained family must contain at least 15 models in these robustness populations. Owner-disjoint discovery and validation halves are then reconstructed using the frozen population seed.

For every perturbation, the spectral MIRT dimension KK is fixed to the value selected for the corresponding benchmark and audit fold in the primary population. We refit the MIRT representation, residual family-DIF model, and three-round anchor-purification procedure under the perturbed population, but do not reselect KK. All remaining parameters are held at their primary values, including the 50% anchor fraction and one-percentage-point close-pair threshold.

The owner-cap-five baseline retains its original 1,000 matched-random replicates. Each new perturbation uses 200 matched-random replicates. For benchmark bb and perturbation vv, define Δ^b​v\widehat{\Delta}_{bv} and pb​vp_{bv} as the matched-random excess and one-sided randomization pp-value. The prespecified robustness counts are

Gvdir\displaystyle G_{v}^{\mathrm{dir}} =∑b=15[Δ^b​v>0],\displaystyle=\sum_{b=1}^{5}\mathbf{1}\!\left[\widehat{\Delta}_{bv}>0\right], (44)
Gvsig\displaystyle G_{v}^{\mathrm{sig}} =∑b=15[pb​v≤0.05].\displaystyle=\sum_{b=1}^{5}\mathbf{1}\!\left[p_{bv}\leq 0.05\right]. (45)

A perturbation passes when

Gvdir≥3andGvsig≥3.G_{v}^{\mathrm{dir}}\geq 3\qquad\text{and}\qquad G_{v}^{\mathrm{sig}}\geq 3. (46)

The population-selection audit, including family counts, owner counts, and cross-fitting-half counts, was written before the perturbed DIF analyses were executed.

D.3 Item-signature replication and source attribution

Common blueprint cells.

Item-signature replication is evaluated by refitting the residual family-DIF model independently in the two owner-disjoint halves. These analyses reuse the owner-cap-five baseline population and its frozen owner split. Let bi(D)b_{i}^{(D)} and bi(V)b_{i}^{(V)} denote the item intercepts estimated in the discovery and validation halves. We define the common item intercept

b¯i=bi(D)+bi(V)2\bar{b}_{i}=\frac{b_{i}^{(D)}+b_{i}^{(V)}}{2} (47)

and assign each item to a common blueprint cell

c⁡(i)=(source⁡(i),Q4​(b¯i)),c(i)=\left(\operatorname{source}(i),Q_{4}(\bar{b}_{i})\right), (48)

where Q4Q_{4} denotes the source-specific easiness quartile. Because the same cells are used in both halves, stability cannot be induced by independently changing the stratification boundaries.

Family-effect replication.

Let δi​f(h)\delta_{if}^{(h)} be the residual effect of family ff on item ii in owner half h∈{D,V}h\in\{D,V\}. Before calculating correlations, we remove the mean effect within every common cell:

δ~i​f(h)=δi​f(h)−1|c⁡(i)|∑j:c⁡(j)=c⁡(i)δj​f(h).\widetilde{\delta}_{if}^{(h)}=\delta_{if}^{(h)}-\frac{1}{|c(i)|}\sum_{j:c(j)=c(i)}\delta_{jf}^{(h)}. (49)

For each family, we calculate the Spearman rank correlation [Spearman, 1904]

ρf=ρS​(δ~⋅f(D),δ~⋅f(V))\rho_{f}=\rho_{S}\left(\widetilde{\delta}^{(D)}_{\cdot f},\widetilde{\delta}^{(V)}_{\cdot f}\right) (50)

and use medianf⁡ρf\operatorname{median}_{f}\rho_{f} as the benchmark-level replication statistic.

Within every common cell, we separately identify the top 20% of items by DIF magnitude in each owner half. Let HDH_{D} and HVH_{V} denote these sets. We report the discovery-to-validation overlap precision

O=|HD∩HV||HD|O=\frac{|H_{D}\cap H_{V}|}{|H_{D}|} (51)

and, among stable high-DIF items, the agreement in the maximally advantaged family:

A=1|HD∩HV|∑i∈HD∩HV𝟏[argmaxfδi​f(D)=argmaxfδi​f(V)].A=\frac{1}{|H_{D}\cap H_{V}|}\sum_{i\in H_{D}\cap H_{V}}\mathbf{1}\left[\arg\max_{f}\delta_{if}^{(D)}=\arg\max_{f}\delta_{if}^{(V)}\right]. (52)

The null distribution is constructed from 500 permutations. In each replicate, the complete validation-half item signature is permuted among items within the same common blueprint cell. For statistic SS, the one-sided permutation pp-value is

pperm​(S)=1+∑r=1500𝟏[S(r)≥Sobs]501.p_{\mathrm{perm}}(S)=\frac{1+\sum_{r=1}^{500}\mathbf{1}[S^{(r)}\geq S^{\mathrm{obs}}]}{501}. (53)

Interpretation relative to chance.

We interpret these statistics using their null behavior rather than additional study-defined cutoffs. Under no cross-half association, the family-wise Spearman correlations are centered at zero, while independent top-20% selections have expected overlap precision 20%. The chance level of advantaged-family agreement depends on the empirical family frequencies and is therefore obtained from the within-cell permutation distribution. Across the five benchmarks and three statistics, all 15 observed values exceed all 500 corresponding permutation values:

pperm=1501=.002.p_{\mathrm{perm}}=\frac{1}{501}=.002.

Even a Bonferroni correction over the 15 comparisons gives padj=.030p_{\mathrm{adj}}=.030.

Source-level family effects.

Source attribution is restricted to MMLU-Pro, BBH, and MMLU, which provide sufficient native source variation. To remove global offsets, each family’s item effects are centered by that family’s all-item mean separately in each owner half:

δi​f∗(h)=δi​f(h)−1I​∑j=1Iδj​f(h).\delta_{if}^{*(h)}=\delta_{if}^{(h)}-\frac{1}{I}\sum_{j=1}^{I}\delta_{jf}^{(h)}. (54)

For source ss, the source-level effect is

γs​f(h)=1|Is|​∑i∈Isδi​f∗(h).\gamma_{sf}^{(h)}=\frac{1}{|I_{s}|}\sum_{i\in I_{s}}\delta_{if}^{*(h)}. (55)

Only sources containing at least 20 items are eligible. For each family, the two most positive and two most negative discovery-half source effects are selected, and their signs are tested in the validation half. Pooled support requires at least 70% sign replication with a one-sided exact binomial p≤0.05p\leq 0.05, together with at least 65% replication in at least two of the three source-rich benchmarks.

Exact source contribution to score shifts.

As a secondary attribution analysis, we exactly decompose each cross-fitted family score shift by source. If anchors fitted in half hh are evaluated on target half h¯\bar{h}, the contribution of source ss to family ff is

Cs​f(h→h¯)=100I​|ℳf,h¯|​∑m∈ℳf,h¯∑i∈Is(wi(h)−1)​Yi​m.C_{sf}^{(h\rightarrow\bar{h})}=\frac{100}{I\,|\mathcal{M}_{f,\bar{h}}|}\sum_{m\in\mathcal{M}_{f,\bar{h}}}\sum_{i\in I_{s}}\left(w_{i}^{(h)}-1\right)Y_{im}. (56)

Because the anchor weights sum to II,

∑sCs​f(h→h¯)=100​(s¯fanchor−s¯ffull)\sum_{s}C_{sf}^{(h\rightarrow\bar{h})}=100\left(\bar{s}_{f}^{\mathrm{anchor}}-\bar{s}_{f}^{\mathrm{full}}\right) (57)

up to numerical precision. For each family, the two largest positive and two largest negative contributions in the discovery-to-validation direction are selected, and their signs are checked after swapping the two folds. This decomposition is a secondary, mechanism-adjacent analysis and is not used to establish owner-disjoint item-signature replication.

D.4 Blinded content-audit protocol and statistical tests

Matched item sample.

The content audit is conducted only after establishing that item-family signatures replicate across owner halves. Within every benchmark, stable high-DIF items are defined as items that fall in the top 20% of DIF magnitude within their common blueprint cell in both owner halves. Controls must fall outside the top 20% in both halves. We sample 50 stable high-DIF items per benchmark and match each one without replacement to a control from the exact same source-by-easiness cell. This produces 50 matched pairs, or 100 items, per benchmark and 500 items in total.

Each item is assigned a random blinded identifier. Annotators receive only this identifier and the question text. Benchmark, source, matched-pair identity, experimental group, family effects, advantaged family, and DIF magnitudes are stored separately and are not included in the annotation prompt.

Annotation procedure.

Annotator A used GPT-5.5 [OpenAI, 2026b], and Annotator B used GPT-5.4 [OpenAI, 2026a]; each annotated all 500 items. Questions are processed in batches of 20 with low reasoning effort. The annotation protocol and output schema were frozen before either annotation run. Annotators are not asked to solve the benchmark items or predict which models answer them correctly. Instead, they label visible content and reasoning demands along ten binary axes:

quantitative or symbolic reasoning; formal rule reasoning; factual domain knowledge; contextual reading; commonsense or narrative reasoning; spatial–temporal reasoning; linguistic wordplay; negation or exception handling; code or structured representation; and distractor discrimination.

They additionally assign ordinal scores from 0 to 3 for reasoning steps, context burden, and ambiguity. The complete rubric and structured annotation schema are included in the released artifact.

Reliability criteria.

For each non-degenerate binary axis, agreement between annotators is measured using Cohen’s κ\kappa [Cohen, 1960]. For each ordinal axis, agreement is measured using Spearman’s ρ\rho [Spearman, 1904]. Annotation reliability passes when

medianj⁡κj≥0.45andmediank⁡ρk≥0.50.\operatorname{median}_{j}\kappa_{j}\geq 0.45\qquad\text{and}\qquad\operatorname{median}_{k}\rho_{k}\geq 0.50. (58)

The median for binary axes excludes axes for which κ\kappa is undefined because both annotators assign a constant label.

Paired semantic tests.

Each annotator is analyzed separately; annotations are not averaged or adjudicated. For binary axis jj and matched pair pp, define

dp​j=Xp​jhigh−Xp​jcontrol.d_{pj}=X_{pj}^{\mathrm{high}}-X_{pj}^{\mathrm{control}}. (59)

The pooled prevalence difference is

d^j=1250​∑p=1250dp​j.\widehat{d}_{j}=\frac{1}{250}\sum_{p=1}^{250}d_{pj}. (60)

We apply the exact two-sided McNemar test [McNemar, 1947], equivalently an exact binomial test on the two types of discordant pairs. Benjamini–Hochberg correction [Benjamini and Hochberg, 1995] is applied over the ten binary axes separately for each annotator. Ordinal axes are evaluated with paired Wilcoxon signed-rank tests [Wilcoxon, 1945] and separately corrected, but are not included in the confirmatory semantic criteria.

A binary axis provides a replicated content explanation only if all of the following conditions hold:

  1. 1.

    the pooled difference has the same nonzero direction for both annotators;

  2. 2.

    the BH-adjusted value satisfies q≤0.05q\leq 0.05 for both annotators;

  3. 3.

    |d^j|≥0.08|\widehat{d}_{j}|\geq 0.08 for both annotators; and

  4. 4.

    the pooled direction occurs in at least three of the five benchmarks for both annotators.

The confirmatory content audit passes if at least one prespecified binary axis meets all four requirements.

Post-confirmatory family analysis.

Among the 250 stable high-residual-DIF items, 158 receive the same advantaged-family label in both owner halves. For each annotator, associations between these stable labels and the ten binary content axes are evaluated using benchmark-stratified label-permutation tests with 10,000 permutations, followed by Benjamini–Hochberg correction across axes. Because this analysis followed the confirmatory audit, it is exploratory and is not part of the semantic criteria.

Appendix E Reproducibility and artifact boundary

The anonymized artifact contains the analysis code, final experimental protocols, aggregate outputs, figure-generation code, and automated validation tests. The tests cover model-family assignment, owner-disjoint splitting, representative selection, family-blind spectral-MIRT dimension selection, residual-DIF estimation, cross-fitted anchor construction, matched-random controls, score-gap sensitivity, population robustness, item-signature stability, source attribution, and blinded content-audit analysis.