Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
Abstract
Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). In owner-disjoint folds, one owner half identifies items with low residual differential item functioning across model families (low-DIF); the resulting frozen, source- and easiness-balanced weights score models in the other half, while equally short matched-random subtests control for generic subtest variation. Full-benchmark and low-DIF rankings remain strongly correlated (–). Yet in four of five benchmarks, 30.9–47.1% of cross-family pairs initially within one percentage point reverse order, exceeding their matched-random medians by 16.9–28.6 percentage points (all ). The fifth benchmark shows no reliable excess ( points, ). The pattern survives all pre-specified population perturbations, and residual item–family signatures replicate across owner halves; however, no family shows a consistent advantage across benchmarks. Thus, globally stable rankings can still leave individual near-tie orderings sensitive to benchmark composition, and sub-one-point leaderboard gaps should be accompanied by evidence that the implied ordering is composition-robust.
1 Introduction
Public leaderboards collapse thousands of item-level responses into a single score, often separating models by less than one percentage point. Such gaps do not establish that the implied ordering is robust to benchmark item composition. Rankings can change under prompt weighting or evaluation perturbations (Siska et al., 2024; Alzahrani et al., 2024), while conventional uncertainty analyses leave some close comparisons unresolved (Kotawala, 2026). We ask whether a near-tied ordering is robust to which benchmark items are included.
We study this question through family-dependent item functioning (family-DIF), adapting differential item functioning from psychometrics (Holland and Wainer, 1993). After a family-blind multidimensional adjustment, residual item–family effects measure systematic differences across observational model families. In owner-disjoint folds, one owner half selects source- and easiness-balanced low-DIF anchors whose frozen weights score the other half, after which the halves are swapped. Equally short matched-random subtests separate targeted composition sensitivity from generic shorter-subtest variation. DIF here is a diagnostic, not a causal or intrinsic family property.
Despite stable aggregate rankings, four of five benchmarks show excess near-tie reversals over matched-random controls, with no consistent family advantage.
2 Related work
Psychometrics for model evaluation.
IRT has been used to construct evaluation scales for NLP systems (Lalor et al., 2016), assess leaderboard and test-set discriminability and identify informative examples (Rodriguez et al., 2021; Vania et al., 2021), detect invalid or mislabeled benchmark items (Truong et al., 2025; Land and Bikel, 2026), and reduce evaluation cost (Polo et al., 2024; Hofmann et al., 2025; Zhou et al., 2026a). Recent methods learn text-conditioned IRT representations for model routing and benchmark prediction (Chen et al., 2025), or use fixed-parameter MIRT anchors to preserve comparability as benchmark suites evolve (Habba et al., 2026). In contrast, our spectral representation controls for multidimensional response structure, and our low-DIF anchors diagnose composition sensitivity rather than predict performance, link test forms, or detect mislabeled items.
Prior DIF work examines human–chatbot differences (Zeinfeld et al., 2026) and scoring procedures that downweight DIF items (Halpin, 2024). A course report closest to our setting finds that DIF-selected BBH pools alter held-out Llama–Qwen group gaps (Liu, 2026). We extend this baseline to owner-disjoint rankings across five benchmarks, individual near-tied pairs, and source- and easiness-matched random controls.
Ranking uncertainty and benchmark robustness.
Model rankings are sensitive to item weighting, prompt composition, and evaluation protocols (Mishra and Arunkumar, 2021; Siska et al., 2024; Kim et al., 2026; Alzahrani et al., 2024). Related work augments accuracy with model-output uncertainty (Ye et al., 2024), quantifies variance across training and evaluation choices (Madaan et al., 2024), shows that clustered benchmark structure can widen rank uncertainty (Neuhof and Benjamini, 2026), and tests whether adjacent model comparisons have adequate paired statistical resolution (Kotawala, 2026). Heineman et al. (2025) analyze ranking fidelity under training and checkpoint noise, while Qian et al. (2026) use discriminability and within-family inversions for subset selection. Our estimand instead isolates excess cross-family reversals among models initially within one percentage point, relative to equally short, composition-matched random subtests.
Measurement validity and interpretation.
Benchmark validity depends on the inferences supported by the measurement protocol (Bean et al., 2025). Multidimensional capability structure can provide an alternative explanation for apparent DIF (Zhou et al., 2026b), while conventional IRT estimators can be unreliable in AI benchmark regimes with few models or non-normal ability distributions (Jiang et al., 2026). Diagnostic conclusions can also be fragile (Aribandi et al., 2021), and textual correlates of DIF may reflect intended subdomains rather than construct-irrelevant content (Maeda and Lu, 2025). We therefore separate ranking consequences, owner-disjoint stability, and blinded semantic interpretation. Our residual effects remain conditional on a family-blind, prediction-selected spectral response representation and do not establish that every capability difference has been removed.
A criterion-level comparison is provided in Appendix A.
3 Audit design
A near-tied ordering is composition-robust if it survives downweighting items with residual family dependence. We estimate item–family effects in one owner half, construct a source- and easiness-balanced low-DIF subtest, and apply its frozen weights to the other half. Blueprint-matched random subtests separate targeted composition sensitivity from ordinary short-subtest variation.
3.1 Data and evaluation population
For each benchmark, we have a binary response matrix , where when model answers item correctly. We audit MMLU-Pro (Wang et al., 2024), BIG-Bench Hard (Suzgun et al., 2023), MMLU (Hendrycks et al., 2021), HellaSwag (Zellers et al., 2019), and WinoGrande (Sakaguchi et al., 2021), using the item-level response matrices released with RouterEval (Huang et al., 2025).
Models are assigned to Qwen, Llama, Gemma, Mistral/Mixtral, or Phi via a frozen, unique model-name match. We cap each owner at 5 checkpoints and split each owner into two approximately equal halves. For ranking comparisons, we retain one checkpoint per owner–family pair, chosen as the checkpoint whose full-benchmark accuracy is closest to that pair’s median.
Exact model-name matching, owner extraction, owner capping, and outer-fold construction details are specified in Appendix C.1.
3.2 Family DIF with a spectral MIRT approximation
We use a family-label-free spectral approximation to the multidimensional item–model term in compensatory MIRT (Reckase, 2006). This is a response-derived spectral working model, rather than a traditional marginal maximum-likelihood MIRT estimator. The choice matches our goal of predictive nuisance adjustment rather than interpretation of latent traits. It permits computationally tractable repeated owner-/item-held-out fitting and anchor purification without specifying a parametric population distribution for model abilities. We do not claim that this approximation is generally superior to likelihood-based MIRT.
Let denote the rank- reconstruction, which approximates the conventional MIRT term . For item and model from observational family , we define the linear predictor
| (1) |
and response probability
| (2) |
where is item easiness and is the residual item-by-family effect.
We ran all downstream analyses using the common candidate grid , omitting dimensions that exceeded the rank of the corresponding fitting matrix. Within each outer owner half, we select using nested owner-held-out and item-held-out Bernoulli log loss and the one-standard-error rule (Hastie et al., 2009). Family labels are not used to learn the response directions or select , and the opposite outer owner half does not select the directions, dimension, or anchor items. Appendix C.3 gives the spectral construction, held-out coordinate estimation, and dimension-selection details.
With fixed, we estimate item easiness and residual family effects by
| (3) |
For identifiability, family effects are centered within every item:
| (4) |
Thus, is family ’s conditional deviation from that item’s across-family mean; the constraint does not imply that the item has no family effects.
We summarize residual family dependence by
| (5) |
A large indicates that observational families remain separated on item after conditioning on the selected spectral MIRT approximation. It does not imply item unfairness or a causal family mechanism. Optimization and convergence details appear in Appendix C.4.
3.3 Cross-fitted residual low-DIF anchors
We use two-fold, owner-disjoint cross-fitting in the sample-splitting sense (Chernozhukov et al., 2018). In each direction, one owner half selects the spectral MIRT dimension, learns the response directions, estimates residual family DIF, and constructs the anchor weights. These weights are then frozen and applied to representative models from the opposite owner half. We swap the two halves so that every evaluated model is scored using anchors selected without its owner’s responses.
Anchor construction uses three fixed purification rounds, following the general logic of iterative DIF purification (Candell and Drasgow, 1988). Starting with all items, each round learns the spectral MIRT response directions from the current anchors, projects all items onto those directions, and re-estimates residual family DIF on all items. Within each native source, items are divided into four strata using the fitted residual item easiness . In every source-by-easiness cell , we retain the fraction with the smallest . The primary analysis uses ; and are sensitivity analyses. The selection and weights produced by the third round are then frozen.
Here, “anchor” denotes a diagnostic scoring subset rather than a scale-linking anchor: unlike fixed-parameter calibration (Habba et al., 2026), it is not used to preserve an IRT scale across evolving test forms.
If cell contains items and retains anchors, the item weights are
| (6) |
Consequently,
| (7) |
so every source-by-easiness cell contributes exactly the same total mass as in the original benchmark. The anchor therefore changes residual-DIF composition without changing the benchmark’s source and coarse easiness blueprint.
For a held-out model , the full-benchmark and weighted anchor scores are
| (8) |
Appendix C.5 specifies the complete three-round procedure, deterministic within-cell selection, and final refitting step.
3.4 Ranking estimand and matched control
The primary comparison set comprises different-owner, different-family model pairs satisfying . A pair undergoes a strict rank reversal when
| (9) |
Pairs tied under either scoring rule are excluded from the corresponding denominator. We pool eligible-pair and reversal counts across the two held-out owner folds before computing the reversal rate. Secondary summaries include mean fold-specific Kendall (Kendall, 1945), the all-cross-family reversal rate, and cumulative results at score-gap thresholds percentage points; one point is primary.
To isolate variation caused by using fewer items, we compare the residual low-DIF anchor with blueprint-matched random subtests. Within each outer fold, every replicate samples without replacement the same number of items from each source-by-easiness cell and applies the same cell-restoring weights as the anchor. The controls therefore match its length, source composition, and coarse easiness composition without selecting items by . Replicate reversal counts are pooled across folds using the same procedure as for the observed low-DIF rate.
Let denote the observed pooled reversal rate and the rate for random replicate . We report the excess reversal rate
| (10) |
and the one-sided empirical randomization -value (Phipson and Smyth, 2010)
| (11) |
Thus, means that low-DIF scoring reverses near-tied pairs more often than equally short, blueprint-matched random subtests.
Separately, 2,000 source-block bootstrap resamples (Efron and Tibshirani, 1993) assess sensitivity to benchmark recomposition; single-source benchmarks use 100 deterministic contiguous item clusters. We report the 2.5th–97.5th percentile range of the recomposed statistic, not a confidence interval around the fixed-benchmark estimate. The near-tie pair set is frozen from the observed full-benchmark scores throughout all random-control and resampling analyses.
4 Validation and Diagnostic Analyses
We test whether the ranking effect survives population changes, replicates across owner halves, and aligns with broad content categories.
4.1 Population robustness
We test sensitivity to owner concentration, checkpoint selection, lineage definition, and family composition. Relative to the owner-cap-5 baseline, the nine perturbations use owner caps of 1 and 3, a score-blind lexicographic representative, a conservative name-based lineage exclusion, or omit Gemma, Llama, Mistral/Mixtral, Phi, or Qwen in turn.
Each variant preserves the owner-disjoint cross-fit and refits the spectral response directions, residual DIF, anchors, and matched controls; only the benchmark- and fold-specific dimension from the primary analysis is frozen. The frozen criteria require every perturbation to retain positive excess reversals in at least three of five benchmarks, with positive matched-random results in at least three. New variants use 200 random controls; the baseline uses 1,000. Full specifications appear in Appendix D.2.
4.2 Item-signature replication and source attribution
We refit residual item–family signatures separately in the two owner halves and construct shared source-by-easiness cells from their averaged item intercepts. After within-cell adjustment, we report three cross-half replication statistics: the median family-wise Spearman correlation, top-20% high-DIF overlap, and advantaged-family agreement among overlapping items. Under independence, the correlation is centered at zero and expected top-20% overlap is 20%; agreement has no fixed baseline because family frequencies differ. We therefore assess all three statistics using 500 one-sided within-cell permutations and report their effect sizes and -values.
For MMLU-Pro, BBH, and MMLU, secondary analyses test whether discovery-selected extreme source effects reproduce in validation and whether exact source contributions to family score shifts retain their signs across folds. Full definitions appear in Appendix D.3.
4.3 Blinded content audit
To interpret the reproducible item–family signatures, we sampled 50 stable high-DIF items per benchmark and paired each with a control from the same source-by-easiness cell, yielding 250 matched pairs. High-DIF items were in the within-cell top 20% in both owner halves; controls were in neither top-20% set.
Two independently prompted language-model annotators, blinded to benchmark, group, DIF, and family information, labeled ten binary and three ordinal content axes. Their labels were analyzed separately.
A binary axis was considered confirmatory only if both annotators showed the same pooled direction, BH , an absolute paired difference of at least eight percentage points, and the same direction in at least three benchmarks. Family-direction analyses were post-confirmatory and exploratory. Full details appear in Appendix D.4.
5 Results
We report the ranking effect, its robustness, and the replication and semantic audit of residual item–family signatures.
5.1 Global rankings are stable, but near-tie orderings are composition-sensitive
| Benchmark | Close | Low-DIF (%) | Random (%) | Excess (pp) | |||
|---|---|---|---|---|---|---|---|
| MMLU-Pro | 64/32 | .924 | 308 | 47.1 | 18.5 | +28.6 | .001 |
| BBH | 128/128 | .915 | 2,226 | 42.1 | 25.2 | +16.9 | .001 |
| MMLU | 128/128 | .930 | 3,235 | 40.7 | 16.3 | +24.4 | .001 |
| HellaSwag | 16/64 | .948 | 4,533 | 30.9 | 11.4 | +19.5 | .001 |
| WinoGrande | 8/8 | .900 | 5,574 | 33.8 | 34.7 | .689 |
Table 1 reports the primary result after adjustment with the family-label-free spectral MIRT approximation. Full-benchmark and low-DIF anchor rankings remain strongly correlated (–), with only 2.5–4.2% of all cross-family comparisons reversing. Among eligible pairs initially within one percentage point, however, four benchmarks show reversal rates of 30.9–47.1%, exceeding matched-random subtests by 16.9–28.6 percentage points (all matched-random ).
For WinoGrande, low-DIF and matched-random reversal rates are similar (33.8% vs. 34.7%; excess points, ). Thus, the observed sensitivity is localized to near ties in four benchmarks, not wholesale leaderboard reranking.
Anchor-fraction sensitivity is reported in Appendix B.2.
5.2 Excess reversals persist across wider score bands
Figure 1 extends the maximum full-score gap from 0.25 to 5 percentage points. At the five-point threshold, all four primary-positive benchmarks still exceed matched-random controls by 5.2–19.1 percentage points, whereas WinoGrande shows no significant excess at any tested threshold. Results outside the pre-specified one-point band are descriptive sensitivity analyses.
5.3 The result is robust to population specification
Across the baseline and all nine population perturbations—two owner caps, score-blind checkpoint selection, conservative lineage filtering, and five leave-one-family-out analyses shown in figure2—MMLU-Pro, BBH, MMLU, and HellaSwag consistently show positive excess reversals with matched-random . WinoGrande reaches only when Mistral/Mixtral models are excluded.
Across the five benchmarks, every family changes the direction of its mean anchor-minus-full score shift at least once. Exact-common-model comparisons confirm several such reversals when the evaluated model population is held fixed. The shifts are therefore benchmark-dependent and should not be interpreted as intrinsic family rankings.
5.4 Residual item signatures replicate across owners
| Benchmark | Median | Top-20% overlap | Advantaged-family agreement |
|---|---|---|---|
| MMLU-Pro | .407 | 31.5% | 56.4% |
| BBH | .589 | 36.1% | 72.5% |
| MMLU | .501 | 38.5% | 71.9% |
| HellaSwag | .308 | 37.4% | 44.3% |
| WinoGrande | .465 | 38.5% | 66.0% |
All three statistics show above-chance owner-disjoint replication on every benchmark (Table 2). Median family-wise correlations are positive (–), while top-20% overlap is 31.5–38.5%, compared with a 20% independence baseline. Among overlapping top items, the same advantaged family is recovered in 44.3–72.5% of cases. Every statistic exceeds all 500 corresponding within-cell permutation values (one-sided ; Bonferroni-adjusted over 15 comparisons). WinoGrande remains informative: its signatures replicate without excess near-tie reversals, so replication alone does not imply a ranking effect.
Among the three source-rich benchmarks, discovery-selected source-effect signs replicate in 58 of 60 held-out checks. Exact source contributions to the anchor-minus-full family shifts retain their signs in 43 of 60 checks (both pooled exact ). These analyses localize reproducible source structure but do not identify a cognitive or training-data mechanism.
5.5 The blinded audit finds no confirmatory semantic enrichment
Annotation agreement is on the binary axes and on the ordinal axes, but none of the ten pre-specified binary axes satisfies all dual-annotator confirmatory criteria. The largest effect is annotator A’s -percentage-point spatial/temporal difference (); annotator B estimates points on the same axis (). Thus, the effect is neither significant after correction nor replicated across annotators. No exploratory family-direction association survives BH correction in either annotator.
6 Limitations and claim boundaries
Observational grouping.
Family labels are inferred from model identifiers and may conflate architecture, training, data, and uploader practices. The owner-disjoint design and population perturbations reduce specific artifacts but cannot identify a causal family mechanism.
Scope and estimand.
The five benchmarks are static, primarily closed-form, and drawn from one response collection; MMLU and MMLU-Pro share content lineage and are not independent replications (Hendrycks et al., 2021; Wang et al., 2024). Interactions between family grouping and other evaluation choices remain untested (Madaan et al., 2024; Alzahrani et al., 2024). Low-DIF anchors reweight observed items: they neither estimate future-item performance nor define a superior construct or item-removal rule. The one-point band is operational, and DIF does not imply unfairness.
Residual interpretation.
The family-blind spectral MIRT approximation remains linear. We evaluate powers-of-two dimensions through , subject to fitting-matrix rank; no one-SE choice reaches its feasible ceiling, although five of ten raw held-out-loss minima occur at that ceiling. Residual effects may therefore still include nonlinear, higher-dimensional, or unmeasured capability differences. Likewise, an audit using two model annotators and broad content axes cannot rule out semantic explanations. Our results establish conditional item–family interactions and their ranking consequences, not intrinsic capabilities, contamination, or broad leaderboard invalidity.
7 Conclusion
We audited whether cross-family leaderboard orderings separated by at most one percentage point are robust to benchmark item composition. Although aggregate rankings remain stable, four of five benchmarks show excess near-tie reversals under residual low-DIF scoring; the fifth does not exceed matched-random variation. Owner-disjoint estimation, matched controls, and population robustness localize this sensitivity to near-tied comparisons rather than a universal family advantage or broad leaderboard invalidity. Small score gaps should therefore be accompanied by composition-robustness checks, not treated as self-interpreting evidence of model superiority.
References
- When benchmarks are targets: revealing the sensitivity of large language model leaderboards. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 13787–13805. Cited by: §1, §2, §6.
- How reliable are model diagnostics?. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 1778–1785. Cited by: §2.
- Measuring what matters: construct validity in large language model benchmarks. Advances in Neural Information Processing Systems 38. Cited by: §2.
- Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society. Series B (Methodological) 57 (1), pp. 289–300. External Links: ISSN 00359246, Link Cited by: §D.4.
- An iterative procedure for linking metrics and assessing item bias in item response theory. Applied Psychological Measurement 12 (3), pp. 253–260. External Links: Document Cited by: §3.3.
- Learning compact representations of LLM abilities via item response theory. arXiv preprint arXiv:2510.00844. Cited by: §2.
- Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal 21 (1), pp. C1–C68. External Links: ISSN 1368-4221, Document, Link, https://academic.oup.com/ectj/article-pdf/21/1/C1/27684918/ectj00c1.pdf Cited by: §3.3.
- A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20 (1), pp. 37–46. External Links: Document Cited by: §D.4.
- An introduction to the bootstrap. Chapman & Hall/CRC, New York, NY. Cited by: §3.4.
- Growing pains: extensible and efficient LLM benchmarking via fixed parameter calibration. arXiv preprint arXiv:2604.12843. Cited by: Table 3, §2, §3.3.
- Finding structure with randomness: probabilistic algorithms for constructing approximate matrix decompositions. SIAM Review 53 (2), pp. 217–288. External Links: Document, Link, https://doi.org/10.1137/090771806 Cited by: §C.2.
- Differential item functioning via robust scaling. Psychometrika 89 (3), pp. 796–821. Cited by: Table 3, §2.
- Model assessment and selection. In The Elements of Statistical Learning: Data Mining, Inference, and Prediction, pp. 219–259. External Links: ISBN 978-0-387-84858-7, Document, Link Cited by: §3.2.
- Signal and noise: a framework for reducing uncertainty in language model evaluation. Advances in Neural Information Processing Systems 38, pp. 17073–17114. Cited by: §2.
- Measuring massive multitask language understanding. In International Conference on Learning Representations, External Links: Link Cited by: §3.1, §6.
- Fluid language model benchmarking. arXiv preprint arXiv:2509.11106. Cited by: §2.
- P. W. Holland and H. Wainer (Eds.) Differential item functioning. Lawrence Erlbaum Associates, Hillsdale, NJ. Cited by: §1.
- RouterEval: a comprehensive benchmark for routing LLMs to explore model-level scaling up in LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 3860–3887. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §3.1.
- Can we trust item response theory for AI evaluation?. External Links: 2607.15190, Document, Link Cited by: §2.
- The treatment of ties in ranking problems. Biometrika 33 (3), pp. 239–251. External Links: Document Cited by: §3.4.
- Benchmarks are not atomic: composition-aware LLM evaluation using BenchHub. In ICML 2026 Workshop on Combining Theory and Benchmarks: Towards A Virtuous Cycle to Understand and Guarantee Foundation Model Performance, Cited by: Table 3, §2.
- Resolution diagnostics for paired LLM evaluation. arXiv preprint arXiv:2605.30315. Cited by: Table 3, §1, §2.
- Building an evaluation scale using item response theory. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 648–657. Cited by: §2.
- Auditing LLM benchmarks with item response theory. arXiv preprint arXiv:2605.30504. Cited by: §2.
- Benchmark items as measurement instruments: DIF diagnostics for response-derived LLM regimes on BBH. Note: CS321M Final Project, Stanford University External Links: Link Cited by: Table 3, §2.
- Quantifying variance in evaluation benchmarks. arXiv preprint arXiv:2406.10229. Cited by: §2, §6.
- Finding words associated with DIF: predicting differential item functioning using LLMs and explainable AI. Journal of Educational Measurement 62 (4), pp. 883–906. Cited by: §2.
- Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12 (2), pp. 153–157. External Links: Document Cited by: §D.4.
- How robust are model rankings: a leaderboard customization approach for equitable evaluation. In Proceedings of the AAAI conference on Artificial Intelligence, Vol. 35, pp. 13561–13569. Cited by: §2.
- Quantifying ranking uncertainty in LLM benchmarks. arXiv preprint arXiv:2607.16259. Cited by: §2.
- GPT-5.4 model. Note: OpenAI API DocumentationAccessed: August 28, 2026 External Links: Link Cited by: §D.4.
- GPT-5.5 model. Note: OpenAI API DocumentationAccessed: August 28, 2026 External Links: Link Cited by: §D.4.
- Permutation P-values should never be zero: calculating exact P-values when permutations are randomly drawn. Statistical Applications in Genetics and Molecular Biology 9 (1), pp. Article 39. External Links: Document Cited by: §3.4.
- tinyBenchmarks: evaluating LLMs with fewer examples. arXiv preprint arXiv:2402.14992. Cited by: §2.
- Benchmark: systematic evaluation of LLM benchmarks. arXiv preprint arXiv:2601.03986. Cited by: Table 3, §2.
- Multidimensional item response theory. Handbook of statistics 26, pp. 607–642. Cited by: §3.2.
- Evaluation examples are not equally informative: how should that change NLP leaderboards?. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 4486–4503. Cited by: §2.
- WinoGrande: an adversarial Winograd schema challenge at scale. Communications of the ACM 64 (9), pp. 99–106. Cited by: §3.1.
- Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 10406–10421. Cited by: Table 3, §1, §2.
- The proof and measurement of association between two things. The American Journal of Psychology 15 (1), pp. 72–101. External Links: Document Cited by: §D.3, §D.4.
- Challenging BIG-Bench tasks and whether Chain-of-Thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, pp. 13003–13051. Cited by: §3.1.
- Fantastic bugs and where to find them in AI benchmarks. Advances in Neural Information Processing Systems 38. Cited by: §2.
- Comparing test sets with item response theory. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 1141–1158. Cited by: §2.
- MMLU-Pro: a more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems 37, pp. 95266–95290. Cited by: §3.1, §6.
- Individual comparisons by ranking methods. Biometrics Bulletin 1 (6), pp. 80–83. External Links: ISSN 00994987, Link Cited by: §D.4.
- IRT parameter drift across LLM generations and architectures on Open LLM Leaderboard v2. Note: CS321M Final Project, Stanford University External Links: Link Cited by: Table 3.
- Benchmarking LLMs via uncertainty quantification. Advances in Neural Information Processing Systems 37, pp. 15356–15385. Cited by: §2.
- Assessment design in the AI era: a method for identifying items functioning differentially for humans and chatbots. arXiv preprint arXiv:2603.23682. Cited by: §2.
- HellaSwag: can a machine really finish your sentence?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4791–4800. Cited by: §3.1.
- Lost in benchmarks? rethinking large language model benchmarking with item response theory. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 35085–35093. Cited by: §2.
- General scales unlock AI evaluation with explanatory and predictive power. Nature 652 (8108), pp. 58–67. Cited by: §2.
Appendix A Comparison with the closest related work
Table 3 shows the comparison with the closest related work.
| Work | Ability control family DIF | Held-out validation | Item-set intervention | Matched controls | Near-tie replication |
| Liu [Liu, 2026] | |||||
| Siska et al. [Siska et al., 2024] | |||||
| Halpin [Halpin, 2024] | |||||
| BenchHub [Kim et al., 2026] | |||||
| Kotawala [Kotawala, 2026] | |||||
| Benchmark2 [Qian et al., 2026] | |||||
| Xu [Xu, 2026] | |||||
| Habba et al. [Habba et al., 2026] | |||||
| Ours |
Appendix B Additional Results
B.1 All score-gap thresholds
| Benchmark | 0.25 | 0.5 | 1 | 2 | 3 | 5 |
|---|---|---|---|---|---|---|
| MMLU-Pro | 43.8/37.1 | 45.8/28.2 | 47.1/18.5 | 40.0/9.8 | 33.1/7.0 | 23.6/4.5 |
| BBH | 49.1/42.3 | 46.6/36.0 | 42.1/25.2 | 33.9/14.1 | 26.9/9.6 | 18.0/5.9 |
| MMLU | 48.0/38.0 | 46.3/28.8 | 40.7/16.3 | 31.7/7.8 | 24.1/5.1 | 15.3/3.1 |
| HellaSwag | 43.5/32.8 | 39.3/21.1 | 30.9/11.4 | 19.0/5.7 | 12.4/3.7 | 7.4/2.2 |
| WinoGrande | 46.0/45.1 | 40.5/41.5 | 33.8/34.7 | 22.6/23.1 | 16.4/16.3 | 10.6/10.3 |
At the 0.25-point threshold, the eligible set falls to 89 pairs for MMLU-Pro, and only BBH, MMLU, and HellaSwag pass the matched-random test. Results outside the frozen one-point threshold are sensitivity analyses rather than additional confirmatory tests.
Details are shown in Table 4.
B.2 Anchor-fraction sensitivity
Table 5 reports descriptive sensitivity to the fraction of low-DIF items retained within each source-by-easiness cell. The primary specification retains 50% of items; 30% and 70% were pre-specified sensitivity settings.
| Benchmark | 30% retained | 50% retained | 70% retained |
|---|---|---|---|
| MMLU-Pro | .885 / 48.4% | .924 / 47.1% | .949 / 39.3% |
| BBH | .882 / 44.3% | .915 / 42.1% | .951 / 33.9% |
| MMLU | .903 / 41.9% | .930 / 40.7% | .955 / 36.7% |
| HellaSwag | .912 / 40.1% | .948 / 30.9% | .970 / 21.6% |
| WinoGrande | .855 / 37.7% | .900 / 33.8% | .939 / 28.5% |
As expected, retaining fewer items produces more ranking changes, but near-tie reversals remain substantial across all three fractions. These results are descriptive: matched-random controls were computed for the primary 50% specification only, so confirmatory excess-reversal inference remains attached to that specification.
B.3 Family/owner robustness matrix
| Variant | MMLU-Pro | BBH | MMLU | HellaSwag | WinoGrande |
|---|---|---|---|---|---|
| Owner cap 1 | +31.2∗ | +16.4∗ | +25.4∗ | +18.6∗ | |
| Owner cap 3 | +28.2∗ | +15.2∗ | +22.1∗ | +20.0∗ | |
| Baseline cap 5 | +28.6∗ | +16.9∗ | +24.4∗ | +19.5∗ | |
| Score-blind checkpoint | +25.8∗ | +15.1∗ | +23.3∗ | +19.3∗ | +1.4 |
| Conservative lineage filter | +28.1∗ | +19.1∗ | +24.3∗ | +17.9∗ | +0.1 |
| Omit Gemma | +28.8∗ | +21.1∗ | +23.5∗ | +23.3∗ | +2.0 |
| Omit Llama | +33.3∗ | +16.1∗ | +34.1∗ | +31.6∗ | |
| Omit Mistral/Mixtral | +27.6∗ | +20.7∗ | +25.3∗ | +10.0∗ | +6.4∗ |
| Omit Phi | +21.0∗ | +15.5∗ | +27.2∗ | +25.7∗ | +3.4 |
| Omit Qwen | +25.2∗ | +13.7∗ | +26.0∗ | +27.7∗ |
The baseline row uses 1,000 matched-random controls, so its minimum attainable randomization -value is ; each of the other nine variants uses 200 controls, giving a minimum of . Under the frozen robustness rule, a specification passes the direction criteria when at least three of the five benchmarks have positive excess and passes the significance criteria when at least three have nominal . All ten specifications pass both criteria. More strongly, MMLU-Pro, BBH, MMLU, and HellaSwag—the four primary-positive benchmarks—remain positive and nominally significant in every specification. WinoGrande is significant only under omission of Mistral/Mixtral, so it remains an unstable boundary case rather than a robust positive result. The details are shown in Table 6.
B.4 Exact-common-model directional checks
| Exact-common pair | Family | Left shift | Right shift | Rightleft | [95% CI] |
|---|---|---|---|---|---|
| MMLU / HellaSwag | Gemma | +2.04 | [, ] | ||
| HellaSwag / WinoGrande | Llama | +0.14 | +1.63 | [+1.37, +1.88] | |
| MMLU-Pro / BBH | Gemma | +0.82 | +1.10 | [+0.38, +1.94] |
Table 7 reports three compact examples used in the main text rather than an exhaustive search for the largest contrasts. Each comparison holds the family-specific model population fixed across the two benchmarks, ruling out changes in model membership as the explanation for the directional change. The released file contains all 15 family estimates for these three exact-common benchmark pairs, including sample sizes and confidence intervals. These shifts are descriptive consequences of benchmark composition and should not be interpreted as intrinsic or causal family rankings.
B.5 Blinded content-audit results
Using the 250 exact-cell matched pairs and two blinded annotators defined in Appendix D.4, Figure 3 reports the pooled paired prevalence differences for the ten prespecified confirmatory binary axes. Annotation reliability passes the frozen criteria. Across the non-degenerate binary axes, the median Cohen’s is .715; across the three ordinal axes, the median Spearman is .800.
No binary axis passes the dual-annotator confirmatory enrichment criteria. The largest absolute difference for annotator A is percentage points on spatial/temporal content (), while annotator B estimates points on the same axis (). The largest absolute difference for annotator B is +4.0 points on distractor discrimination (). Thus, no axis is both significant after correction and replicated across annotators.
A post-confirmatory diagnostic is restricted to the 158 of 250 high-residual-DIF items for which the advantaged-family label agrees across owner halves. No benchmark-stratified family-direction association survives BH correction in either annotator, and no dual-annotator exploratory signal is detected. These analyses remain exploratory and are not used as a semantic explanation of the ranking effect.
Appendix C Implementation details for the MIRT-adjusted audit
This appendix specifies the numerical implementation underlying Section 3. The main text defines the statistical estimand; here we describe the spectral MIRT working representation, dimension selection, residual-DIF optimization, anchor purification, and resampling procedures.
C.1 Model-family and owner construction
Family membership is assigned by a unique, case-insensitive match on model ID. The frozen matching rules are:
| Family | Model-ID pattern |
|---|---|
| Qwen | qwen |
| Llama | llama |
| Gemma | gemma |
| Mistral/Mixtral | mistral or mixtral |
| Phi | a token beginning with phi |
Models matching zero or multiple family patterns are excluded. The owner is the uploader component of the model ID. If an owner contributes more than the frozen cap of five checkpoints to one family, five are selected using the frozen population-selection seed.
Entire owners are assigned to one of two outer halves using a deterministic greedy procedure that approximately balances, within every family, both the number of retained checkpoints and aggregate full-benchmark accuracy. No owner may appear in both halves.
In benchmark order (MMLU-Pro, BBH, MMLU, HellaSwag, and WinoGrande), the complete-response matrices contain respectively 12,032, 5,761, 14,042, 10,042, and 1,267 items and 1,823, 3,811, 5,000, 5,000, and 5,000 models. The corresponding ranking populations contain 204, 497, 673, 673, and 673 owner–family representatives.
C.2 Spectral working-response construction
Let denote the models in the current audit half and let . For every item, we calculate a smoothed marginal success probability
| (12) |
We then form the standardized working-response matrix
| (13) |
Suppose is the item set currently used to estimate the model coordinates. We compute a rank- randomized singular-value decomposition [Halko et al., 2011]
| (14) |
The model-coordinate matrix is
| (15) |
and loadings for all items, including those outside , are obtained by projection:
| (16) |
Thus the rank- reconstruction of the standardized response residual is . We convert it into the fixed offset used in the logistic residual-DIF model:
| (17) |
The truncation at six logits is a numerical safeguard and is not tuned using family labels or ranking outcomes.
This construction is a spectral approximation to a multidimensional logistic response surface. It is not a full maximum-likelihood calibration of a traditional 2PL MIRT model; we therefore refer to it as a spectral MIRT working model.
C.3 Nested selection of the MIRT dimension
For outer audit half , the feasible candidate set is
The feasible grid reaches in both MMLU-Pro folds and in all folds of the other four benchmarks. Dimension selection is performed separately inside each outer audit half.
We first assign 25% of audit-half owners to an inner validation set. Using only the remaining owners, we compute Equation 13 and a rank- decomposition. For dimension , define the training-owner item factor matrix
| (18) |
Within every native source, items are randomly divided into a coordinate- calibration set and an evaluation set , with approximately half assigned to each. For inner-validation model , its coordinates are estimated only from :
| (19) |
For evaluation item , the predicted logit is
| (20) |
The model-level validation loss is
| (21) |
Let be the mean of over inner-validation models, and define
| (22) |
Let . Following the one-standard-error rule, we select
| (23) |
The selected is frozen before family labels are introduced and is used for every anchor fraction within that outer fold.
The one-SE selection remains below the feasible candidate ceiling in every fold. The raw held-out-loss minimum occurs at the feasible ceiling in five of the ten folds, so the selected representation should not be interpreted as exhausting all possible capability structure.
C.4 Residual-DIF optimization
For a fixed spectral MIRT offset , the response logit is
| (24) |
Let denote the models in the current fitting owner half and . We initialize
| (25) | ||||
| (26) |
and set all family effects to zero. Here .
At the current parameter values, let
For fixed family effects, the score and information for item intercept are
| (27) | ||||
| (28) |
We apply the clipped Newton update
| (29) |
For family , let
Holding the item intercept fixed, the penalized score and information for are
| (30) | ||||
| (31) |
with . We update
| (32) |
The probabilities are recomputed at the current parameter values during each Newton iteration.
After every complete family-effect sweep, we impose the identifying constraint by setting
| (33) |
The simultaneous adjustment of preserves every linear predictor while ensuring .
We use at most five alternating intercept/family-effect cycles. Each Newton subproblem uses at most 30 iterations and terminates when the largest absolute update is below . The outer alternation also terminates early when the largest family-effect update is below this threshold.
C.5 Anchor purification algorithm
For each outer fold and anchor fraction , purification proceeds as follows.
- 1.
Initialize
- 2.
For rounds :
- (a)
Fit the rank- spectral MIRT representation using the response rows in .
- (b)
Project all items onto the fitted model coordinates and construct offsets .
- (c)
Fit and on all items conditional on .
- (d)
Compute
- (e)
Within each native source, order items by and divide them into four near-equal strata. Ties are resolved by original item index.
- (f)
Within every source-by-easiness cell , retain
items with the smallest . Their union defines .
- (a)
- 3.
Assign final weights
- 4.
Refit the spectral representation and residual-DIF parameters once on to obtain the stored final item signatures. Ranking scores use the weights selected at the end of round three.
Purification always runs for three rounds; convergence of the anchor set is recorded through successive Jaccard overlap but is not used as a stopping rule.
C.6 Matched-random subtests
For each outer fold, let be the final low-DIF anchors in cell and . Random-control replicate independently samples
| (34) |
uniformly without replacement. Sampled items receive weight .
The same replicate index is pooled across the two outer target halves:
| (35) |
The reported matched-random median, interval, excess, and randomization -value are calculated from .
C.7 Composition bootstrap
For benchmarks with at least five native sources, native source labels define bootstrap units. Otherwise, the original item order is divided deterministically into 100 contiguous, near-equal clusters.
Let be the number of bootstrap units and let
| (36) |
be the unit multiplicities for replicate .
For unit , let and be the model’s full-score successes and item mass, and let and be their weighted anchor equivalents. Bootstrap scores are
| (37) |
The close-pair mask is defined once from the observed full scores and held fixed across the 2,000 bootstrap replicates.
C.8 Implementation settings
Table C.8 summarizes the frozen primary-audit specification.
Table C.8 summarizes the frozen primary-audit
specification.
Component
Frozen Setting
Families
Qwen, Llama, Gemma, Mistral/Mixtral, Phi
Owner cap
5 checkpoints per owner–family
Minimum capped family size
20 models
Outer split
owner-disjoint two-fold cross-fit
MIRT candidate dimensions
;
rank-infeasible values omitted
Selected (two folds)
MMLU-Pro 64/32; BBH 128/128; MMLU 128/128;
HellaSwag 16/64; WinoGrande 8/8
Dimension rule
minimum within one SE of best loss
Inner validation owners
25%
Coordinate-calibration items
50% within source
Coordinate ridge
1
Maximum MIRT offset
6 logits
Randomized SVD iterations
5
DIF penalty
1
DIF fitting cycles
5
Newton iterations per update
at most 30
Newton tolerance
Maximum Newton step
2 logits
Purification rounds
3
Easiness strata
4 within each source
Primary anchor fraction
50%
Anchor sensitivity
30%, 70%
Primary close-pair band
1 percentage point
Gap sensitivity
0.25, 0.5, 1, 2, 3, 5 points
Matched-random replicates
1,000
Composition-bootstrap replicates
2,000
Single-source bootstrap units
100 contiguous clusters
Appendix D Validation and Robustness: Implementation Details
D.1 Controls and uncertainty
Eligible pairs and pooled reversal rate.
Let index the target owner half and let denote an alternative scoring rule, either the low-DIF anchor score or one matched-random score. For threshold , define
| (38) |
where and denote the owner and family of model . Thus, eligible models must come from different owners and different families, must be separated by no more than on the full benchmark, and must not be tied under either scoring rule.
For an eligible pair, the strict reversal indicator is
| (39) |
The pooled reversal rate sums reversal counts and eligible-pair counts across the two cross-fitting directions:
| (40) |
The primary analysis sets , corresponding to a one-percentage-point full-benchmark score gap.
Resampling controls.
Threshold sensitivity.
The primary specification uses a 50% anchor fraction and defines near-tied models as those separated by at most one percentage point on the full benchmark. We additionally evaluate
| (41) |
corresponding to full-score gaps of percentage points. Pair membership at every threshold is frozen from the observed full-benchmark scores. The matched-random distribution is recomputed using 1,000 replicates, and composition uncertainty is recomputed using 2,000 bootstrap replicates. We also repeat anchor construction using the lowest-DIF 30% and 70% of items within each blueprint cell. These specifications are sensitivity analyses and are not used to select the primary threshold or anchor fraction.
Owner bootstrap.
Family-specific score and percentile shifts may be correlated within a model owner. We therefore resample owners rather than checkpoints. Let
| (42) |
be the score shift for representative model . In bootstrap replicate , we sample the observed owners with replacement, using a sample size equal to the number of unique owners, and include all representative records associated with each sampled owner. For family , the bootstrap mean is
| (43) |
We analogously calculate the mean percentile-rank shift and report empirical 95% intervals over 2,000 replicates. Owner-bootstrap intervals are used only for family-specific shifts; uncertainty in the primary reversal statistic is handled by the composition bootstrap and matched-random control.
D.2 Population robustness
The primary population retains at most five checkpoints per owner–family and selects one ranking representative nearest the within-owner–family median full-benchmark accuracy. Entire owners, including all of their family-specific records, remain assigned to a single outer half. We evaluate nine perturbations of this construction:
- 1.
reduce the owner cap from five checkpoints to one;
- 2.
reduce the owner cap from five checkpoints to three;
- 3.
select one checkpoint per owner–family lexicographically by model identifier, without using benchmark accuracy;
- 4.
exclude model identifiers indicating merged, adapter-based, distilled, hybrid, preference-tuned, or uncensored lineages; and
- 5.
omit each of Gemma, Llama, Mistral/Mixtral, Phi, and Qwen in turn.
The conservative lineage filter excludes identifiers matching terms including merge, slerp, franken, lora, qlora, adapter, distill, hybrid, fusion, dpo, ppo, orpo, kto, abliterat, and uncensor. Each retained family must contain at least 15 models in these robustness populations. Owner-disjoint discovery and validation halves are then reconstructed using the frozen population seed.
For every perturbation, the spectral MIRT dimension is fixed to the value selected for the corresponding benchmark and audit fold in the primary population. We refit the MIRT representation, residual family-DIF model, and three-round anchor-purification procedure under the perturbed population, but do not reselect . All remaining parameters are held at their primary values, including the 50% anchor fraction and one-percentage-point close-pair threshold.
The owner-cap-five baseline retains its original 1,000 matched-random replicates. Each new perturbation uses 200 matched-random replicates. For benchmark and perturbation , define and as the matched-random excess and one-sided randomization -value. The prespecified robustness counts are
| (44) | ||||
| (45) |
A perturbation passes when
| (46) |
The population-selection audit, including family counts, owner counts, and cross-fitting-half counts, was written before the perturbed DIF analyses were executed.
D.3 Item-signature replication and source attribution
Common blueprint cells.
Item-signature replication is evaluated by refitting the residual family-DIF model independently in the two owner-disjoint halves. These analyses reuse the owner-cap-five baseline population and its frozen owner split. Let and denote the item intercepts estimated in the discovery and validation halves. We define the common item intercept
| (47) |
and assign each item to a common blueprint cell
| (48) |
where denotes the source-specific easiness quartile. Because the same cells are used in both halves, stability cannot be induced by independently changing the stratification boundaries.
Family-effect replication.
Let be the residual effect of family on item in owner half . Before calculating correlations, we remove the mean effect within every common cell:
| (49) |
For each family, we calculate the Spearman rank correlation [Spearman, 1904]
| (50) |
and use as the benchmark-level replication statistic.
Within every common cell, we separately identify the top 20% of items by DIF magnitude in each owner half. Let and denote these sets. We report the discovery-to-validation overlap precision
| (51) |
and, among stable high-DIF items, the agreement in the maximally advantaged family:
| (52) |
The null distribution is constructed from 500 permutations. In each replicate, the complete validation-half item signature is permuted among items within the same common blueprint cell. For statistic , the one-sided permutation -value is
| (53) |
Interpretation relative to chance.
We interpret these statistics using their null behavior rather than additional study-defined cutoffs. Under no cross-half association, the family-wise Spearman correlations are centered at zero, while independent top-20% selections have expected overlap precision 20%. The chance level of advantaged-family agreement depends on the empirical family frequencies and is therefore obtained from the within-cell permutation distribution. Across the five benchmarks and three statistics, all 15 observed values exceed all 500 corresponding permutation values:
Even a Bonferroni correction over the 15 comparisons gives .
Source-level family effects.
Source attribution is restricted to MMLU-Pro, BBH, and MMLU, which provide sufficient native source variation. To remove global offsets, each family’s item effects are centered by that family’s all-item mean separately in each owner half:
| (54) |
For source , the source-level effect is
| (55) |
Only sources containing at least 20 items are eligible. For each family, the two most positive and two most negative discovery-half source effects are selected, and their signs are tested in the validation half. Pooled support requires at least 70% sign replication with a one-sided exact binomial , together with at least 65% replication in at least two of the three source-rich benchmarks.
Exact source contribution to score shifts.
As a secondary attribution analysis, we exactly decompose each cross-fitted family score shift by source. If anchors fitted in half are evaluated on target half , the contribution of source to family is
| (56) |
Because the anchor weights sum to ,
| (57) |
up to numerical precision. For each family, the two largest positive and two largest negative contributions in the discovery-to-validation direction are selected, and their signs are checked after swapping the two folds. This decomposition is a secondary, mechanism-adjacent analysis and is not used to establish owner-disjoint item-signature replication.
D.4 Blinded content-audit protocol and statistical tests
Matched item sample.
The content audit is conducted only after establishing that item-family signatures replicate across owner halves. Within every benchmark, stable high-DIF items are defined as items that fall in the top 20% of DIF magnitude within their common blueprint cell in both owner halves. Controls must fall outside the top 20% in both halves. We sample 50 stable high-DIF items per benchmark and match each one without replacement to a control from the exact same source-by-easiness cell. This produces 50 matched pairs, or 100 items, per benchmark and 500 items in total.
Each item is assigned a random blinded identifier. Annotators receive only this identifier and the question text. Benchmark, source, matched-pair identity, experimental group, family effects, advantaged family, and DIF magnitudes are stored separately and are not included in the annotation prompt.
Annotation procedure.
Annotator A used GPT-5.5 [OpenAI, 2026b], and Annotator B used GPT-5.4 [OpenAI, 2026a]; each annotated all 500 items. Questions are processed in batches of 20 with low reasoning effort. The annotation protocol and output schema were frozen before either annotation run. Annotators are not asked to solve the benchmark items or predict which models answer them correctly. Instead, they label visible content and reasoning demands along ten binary axes:
quantitative or symbolic reasoning; formal rule reasoning; factual domain knowledge; contextual reading; commonsense or narrative reasoning; spatial–temporal reasoning; linguistic wordplay; negation or exception handling; code or structured representation; and distractor discrimination.
They additionally assign ordinal scores from 0 to 3 for reasoning steps, context burden, and ambiguity. The complete rubric and structured annotation schema are included in the released artifact.
Reliability criteria.
For each non-degenerate binary axis, agreement between annotators is measured using Cohen’s [Cohen, 1960]. For each ordinal axis, agreement is measured using Spearman’s [Spearman, 1904]. Annotation reliability passes when
| (58) |
The median for binary axes excludes axes for which is undefined because both annotators assign a constant label.
Paired semantic tests.
Each annotator is analyzed separately; annotations are not averaged or adjudicated. For binary axis and matched pair , define
| (59) |
The pooled prevalence difference is
| (60) |
We apply the exact two-sided McNemar test [McNemar, 1947], equivalently an exact binomial test on the two types of discordant pairs. Benjamini–Hochberg correction [Benjamini and Hochberg, 1995] is applied over the ten binary axes separately for each annotator. Ordinal axes are evaluated with paired Wilcoxon signed-rank tests [Wilcoxon, 1945] and separately corrected, but are not included in the confirmatory semantic criteria.
A binary axis provides a replicated content explanation only if all of the following conditions hold:
- 1.
the pooled difference has the same nonzero direction for both annotators;
- 2.
the BH-adjusted value satisfies for both annotators;
- 3.
for both annotators; and
- 4.
the pooled direction occurs in at least three of the five benchmarks for both annotators.
The confirmatory content audit passes if at least one prespecified binary axis meets all four requirements.
Post-confirmatory family analysis.
Among the 250 stable high-residual-DIF items, 158 receive the same advantaged-family label in both owner halves. For each annotator, associations between these stable labels and the ten binary content axes are evaluated using benchmark-stratified label-permutation tests with 10,000 permutations, followed by Benjamini–Hochberg correction across axes. Because this analysis followed the confirmatory audit, it is exploratory and is not part of the semantic criteria.
Appendix E Reproducibility and artifact boundary
The anonymized artifact contains the analysis code, final experimental protocols, aggregate outputs, figure-generation code, and automated validation tests. The tests cover model-family assignment, owner-disjoint splitting, representative selection, family-blind spectral-MIRT dimension selection, residual-DIF estimation, cross-fitted anchor construction, matched-random controls, score-gap sensitivity, population robustness, item-signature stability, source attribution, and blinded content-audit analysis.