Human-Anchored Factuality Evaluation with Strategic Annotation
Abstract
LLM-based factuality judges provide scalable evaluation signals, but their metrics are often systematically biased relative to human judgments. We study human-anchored factuality evaluation under limited annotation budgets, where judge predictions on the full dataset are combined with human labels on a small selectively sampled subset to obtain statistically valid estimates. The efficiency of this approach depends critically on which examples receive human annotation: in factuality evaluation, judge-human misalignment is not driven solely by low confidence, but also by structured failure modes such as incomplete evidence, temporal mismatch, unverifiable claims, and rubric misalignment. To exploit this structure, we introduce a factuality-specific annotation policy design pipeline that uses failure-space analysis (FSA) to derive diverse predictive signals for modeling human-judge misalignment. On an internal reference-based factuality evaluation system (AutoFA) and RAGTruth, where judge-predicted estimates substantially underestimate human-annotated factual accuracy, our FSA-guided policy improves annotation efficiency over uniform sampling and uncertainty-driven baselines, achieving effective-sample-size gains of 40.3% on AutoFA and 27.1% on RAGTruth.
1 Introduction
Factuality evaluation is a recurring concern when deploying large language models in user-facing systems. In production settings, teams rely on factual accuracy to measure system quality (e.g., LLM model, prompts, grounding retrieval), monitor regressions, and make release decisions from evaluation data that are both timely and statistically reliable. Human annotation remains the most trusted way to measure factual accuracy, but it is expensive and slow. Collecting sufficient labels to obtain narrow confidence intervals requires substantial annotation budget and introduces operational latency. This creates a practical tension between evaluation quality and evaluation efficiency.
LLM-as-a-judge (LLMaaJ) systems offer an attractive way to scale factuality evaluation (Min et al., 2023; Chen et al., 2023; Niu et al., 2024; Wei et al., 2024; Song et al., 2024). Given a query, a response, and reference evidence, the judge can decompose the response into claims, verify them against evidence, and produce factuality labels at much larger scale than human annotators. However, judge-predicted factuality estimates are not necessarily human-anchored (Zheng et al., 2023; Wang et al., 2024). In reference-based factuality evaluation, judge errors are often systematic rather than random. A judge may reject reasonable answers when references are incomplete, stale, conflicting, or indirectly supportive. It may also apply a more literal verification rubric than human annotators. As a result, treating judge predictions as ground truth can produce biased factual accuracy estimates, even when the judge is useful as a proxy signal.
This paper studies factuality evaluation under a budgeted annotation setting. Given a large set of evaluation items, an LLMaaJ provides proxy factuality predictions for all items, while only a small subset can receive human labels. The goal is to estimate the human-defined factual accuracy rate (FAR) with valid uncertainty quantification and lower variance than uniformly sampled human annotation alone. We adopt Active Statistical Inference (ASI) (Zrnic and Candes, 2024), which combines judge predictions with human corrections from a sampled subset of examples. ASI preserves the human-anchored target, but its efficiency depends critically on the annotation policy – human labels should be allocated to examples where the judge is most likely to disagree with humans.
Factuality residuals are often structured rather than uniformly distributed across examples. Existing confidence-based approaches (Gligorić et al., 2025) use the judge’s expressed uncertainty as a sampling signal, but judge-human misalignment is not driven solely by low confidence. A judge can be confidently misaligned when evidence is incomplete, temporally inconsistent, or only indirectly supportive, and when responses require paraphrase, inference, or rubric tolerance that humans handle differently from the judge. To exploit this structure, we construct the annotation policy using failure-space analysis (FSA), a factuality-specific policy-design pipeline that derives predictive signals from judge-output, evidence quality, task and input strata, and answer-rubric alignment before converting them into ASI sampling policies.
We evaluate the framework on two benchmarks: an internal Automatic Factual Evaluation system (AutoFA) for virtual assistant responses, and RAGTruth, a public hallucination benchmark with human span-level annotations. In both settings, judge-predicted FAR substantially underestimates the human-annotated FAR and adopting FSA-derived policies preserve desired coverage while improving annotation efficiency over uniform sampling and confidence-based sampling. Aggregated across budgets, FSA-guided ASI achieves an average effective sample size (ESS) gain of 40.3% on AutoFA and 27.1% on RAGTruth.
To summarize, our contributions are threefold. First, we formulate reference-based factuality evaluation as a human-anchored, judge-assisted inference problem under limited annotation budget. Second, we introduce failure-space analysis as a practical policy-design layer for identifying factuality-specific judge-human residual risk. Third, we demonstrate on both internal and public factuality settings that FSA-guided ASI recovers the human-defined evaluation target while improving annotation efficiency over uniform and confidence-based alternatives. Together, these contributions enable reliable and cost-efficient factuality measurement in practical evaluation pipelines, where fully manual annotation is infeasible and judge-only metrics can be systematically biased.
2 Background
2.1 Problem Setup
We first formulate the problem of reference-based factuality evaluation for short-form query-response pairs, while noting the proposed methodology can be naturally extended to long form settings. Each evaluation item consists of a query , a model response , and the reference evidence obtained from retrieval, curated sources, or an existing evaluation pipeline. We write the observable input as and denote the human-annotated factuality label of each by . In practice, is usually not a direct verdict, but verified atomic claims, or detected hallucinated spans.
In this paper, we focus on Factual Accuracy Rate (FAR)11 1 If the LLMaaJ’s evidence is the grounding provided to the answer generator, FAR measures faithfulness rate., which is defined as:
| (1) |
where denotes the factual correctness for the item.
2.2 Judge-Assisted Factual Evaluation under Budgeted Annotation
While FAR from human annotation is usually treated as the ground truth, acquiring human labels is expensive and slow, making it difficult to scale for drawing statistically significant conclusions. On the other hand, Automatic Factual Evaluation can produce judge predictions at scale but the obtained FAR is generally biased. This motivates us to study a budgeted, judge-assisted factual evaluation problem where we combine the abundant judge predictions with limited human labels to form a statistically valid and more efficient estimator.
Given data points, an LLMaaJ provides predictions, denoted by , for all items, while only items can be sent for human annotation. We adopt Active Statistical Inference (ASI) (Zrnic and Candes, 2024), which selects the labeled subset from an annotation policy, denoted by .
Let indicate whether item receives a human label with the inclusion probability . ASI estimates FAR by:
where and denote the judge predicted and human annotated factual correctness respectively.
| Signal family | Representative signals |
| Judge-output artifacts (F1) | Claim/span structure, verification or hallucination fraction, unverifiable outputs, self-reflection or confidence. |
| Evidence and reference quality (F2) | Reference support strength, reference conflict, source freshness, direct vs. indirect support. |
| Task and input strata (F3) | Task or domain type, query category, time-sensitive or location context indicators. |
| Answer-rubric alignment (F4) | Claim importance, completeness vs. correctness, overprecision or auxiliary details, paraphrase/inference tolerance. |
2.3 Statistical Efficiency
Drawing statistically meaningful conclusions requires not only targeting the correct human-defined metric, but also estimating it with low variance under a limited annotation budget. We compare three estimators that are central to this paper. The classical estimator, , uses only the human-labeled examples sampled uniformly at random. It is unbiased for the human-defined FAR, but can be noisy under small annotation budget. The judge-predicted estimator, , uses judge predictions on all examples and therefore has low variance, but it is biased when the judge is systematically misaligned with human.
The ASI estimator remains human-anchored. However, its variance is policy dependent and scales with , where is the judge-human residual (Zrnic and Candes, 2024). Therefore, efficient annotation requires assigning higher sampling probabilities to examples with larger expected residuals, which motivates the failure-space analysis in the next section.
3 FSA-guided Policy Design
Efficient ASI requires allocating human labels to examples with large expected judge-human residuals. In the oracle case, the optimal policy is proportional to the conditional residual magnitude,
However, since is unavailable before annotation, the practical policy-design problem is to predict residual risk from pre-labeling information: the query , response , reference evidence , and judge output .
Directly modeling residual risk from raw text is difficult, and generic uncertainty is often insufficient for factuality evaluation. We therefore introduce failure-space analysis (FSA), a factuality-specific procedure for deriving structured features that expose where judge-human misalignment is likely to occur. FSA organizes candidate policy signals into four families, summarized in Table 1: judge-output artifacts, evidence and reference quality, task and input strata, and answer-rubric alignment. Together, these signals capture not only whether the judge is uncertain, but also whether the evidence is incomplete or stale, whether the query belongs to a difficult slice, and whether the response requires inference or rubric tolerance that the judge may handle differently from humans.
Operationally, FSA uses a historical labeled set where both judge outputs and human annotations are available. We compute the residual target , extract numerical features from pre-labeling artifacts, and train a policy model to predict residual risk. The resulting scores are converted into sampling probabilities so that the annotation budget is concentrated on examples where the judge is most likely to disagree with humans. A pipeline overview is summarized and illustrated in Figure 1. We then provide a comprehensive demonstration of this pipeline in Section 4.2 and Section 4.3.
4 Experiments
4.1 Evaluation Settings
We study two factuality-judge settings with different judge architectures, data formats, and judge output structures.
Automatic Factual Accuracy (AutoFA)
AutoFA is an internal reference-based factuality evaluation system for voice assistant responses. Given a user query and model response, three reference answers are produced by search-enabled LLMs. The judge decomposes the response into verifiable claims, fact-checks each claim against the references, and produces an overall factuality verdict, cf. SAFE Wei et al. (2024) and VeriScore Song et al. (2024). Both claim-level and response-level outputs can be true, false, or unknown. We use responses and FAR is defined as the fraction of responses rated factually correct among examples with a definitive verdict.
RAGTruth.
RAGTruth (Niu et al., 2024) is a public hallucination benchmark with human span-level annotations. We use GPT-4-0613 responses across question answering, summarization, and data-to-text generation. We evaluate Claude Sonnet 4.6 as a faithfulness judge using the benchmark prompt, which asks the judge to identify hallucinated spans. For RAGTruth, FAR is the faithfulness rate: the fraction of responses without hallucination.
Self-Confidence Score.
Following prior work on confidence-driven inference (Gligorić et al., 2025), we collect a self-reflected confidence score for each record by prompting the judge to assess its confidence in its own previous verdict, yielding a scalar uncertainty signal in .
4.2 FSA of Judge Residuals
We first compare judge FAR against human FAR to characterize the residuals that an annotation policy should target. On AutoFA, the judge underestimates human FAR by percentage points. On RAGTruth, the gap is larger: the judge FAR is , compared with a human FAR of .
Typical judge failures.
In both datasets, judge-human disagreements are dominated by judge over-rejection. On AutoFA, of disagreements occur when the judge marks human-correct responses as false () or unknown (). Manual inspection identifies two common mechanisms: evidence gaps, where references do not directly state a reasonable answer that humans accept, and overly-strict verification, where the judge rejects an otherwise correct response due to unsupported peripheral details. On RAGTruth, of disagreements occur when the judge flags hallucinations that humans do not annotate, mainly caused by overly-literal grounding: the judge penalizes valid paraphrases or reasonable inferences because they are not explicitly stated in the reference.
| Dataset | Stratum | Prev. | Lift | |
| AutoFA | Partial Verification | 0.32 | 0.41 | 2.58 |
| Time-sensitive Query | 0.40 | 0.20 | 1.29 | |
| RAGTruth | Low Self-confidence | 0.32 | 0.78 | 2.24 |
| Judge flagged output | 0.43 | 0.70 | 2.02 |
Non-uniform residual structure.
Table 2 shows that residuals concentrate in observable strata derived from judge outputs and task metadata. We report each stratum’s prevalence, average squared residual , and lift, defined as the ratio between stratum-level and population-level . On AutoFA, partial verification, i.e., cases where at least one extracted claim is marked unknown, has lift, while time-sensitive queries have lift. On RAGTruth, low self-confidence has lift, and judge-flagged outputs, i.e., cases where the judge returns one or more hallucination spans rather than an empty list, have lift. These concentrations show that judge-human residuals are structured rather than uniformly distributed, motivating the FSA-derived features in Section 4.3.
| Family | Feature | Spearman | |
| AutoFA | RAGTruth | ||
| Judge | self-confidence | ||
| flag-reliability | – | ||
| Evidence | ref-support-strength | ||
| ref-temporal-consistency | – | ||
| Strata | time-sensitive / task | ||
| Rubric | inference-required | ||
| qualitative-language | – | ||
4.3 Residual-Risk Feature Modeling
Guided by Section 3 and the residual patterns in Section 4.2, we instantiate dataset-specific features from the artifacts available before human annotation. For both datasets, we use the judge’s self-confidence score as a generic uncertainty signal.
For RAGTruth, where the judge may over-detect hallucinated spans, we additionally define flag-reliability to assess whether the judge-flagged spans are plausibly unsupported by the reference.
For AutoFA, the dominant residuals are tied to reference quality, so we extract ref-support-strength, which measures how directly the references support the response’s core claims, and ref-temporal-consistency, which measures whether search-generated references agree on time-dependent facts. We also include task-level strata, such as time-sensitive queries in AutoFA and task type in RAGTruth.
Finally, we extract rubric-alignment features, including inference-required, which measures whether verification requires reasoning beyond literal evidence, and qualitative-language, which captures interpretive or subjective phrasing. LLM-extracted features are produced with structured prompts that return integer scores on a 0–3 scale. These feature-extraction calls are applied before human annotation and are substantially cheaper than human labeling in our setting. We provide more details in Appendix E.
Table 3 reports the absolute Spearman correlation between each feature and the residual magnitude . This provides a univariate diagnostic of whether each feature monotonically tracks residual risk. On AutoFA, ref-support-strength has the strongest correlation, consistent with evidence sufficiency being the dominant failure mode. On RAGTruth, self-confidence dominates, while flag-reliability provides an additional signal on the judge-flagged subset.
The correlations in Table 3 are diagnostic rather than the final policy criterion. A feature may correlate with but still be less useful for sampling if it is sparse, redundant, or unstable under limited calibration data. We therefore split the calibration set into an inner training split for fitting residual-risk models and an inner validation split for comparing candidate feature configurations. This procedure selects ref-support-strength for AutoFA, and self-confidence together with flag-reliability for RAGTruth, reflecting the different residual structures of the two factuality settings.
4.4 Debiasing and Efficiency Evaluation
We evaluate the method along two dimensions: whether ASI recovers the human-anchored FAR despite judge bias, and whether FSA-derived policies improve annotation efficiency.
Policy learning.
Each dataset is split 40/60 into calibration and evaluation sets. On the calibration set, we train a gradient-boosted tree (XGBoost) to predict from the selected features, producing a residual-risk score for each record. The scores are converted into sampling probabilities by normalization with uniform mixing (Li et al., 2025; Zrnic and Candes, 2024), and we apply power-tuning (Angelopoulos et al., 2024) to set the judge contribution in the ASI estimator. Uniform-mixing and power-tuning parameters are estimated on the calibration set, with details in Appendix A. All reported results are evaluated on the evaluation split.
Estimators for comparison.
We compare five estimators. Classical uniformly samples human labels and estimates FAR with only human annotations. Judge-predicted reports the judge FAR over all records without human correction. Uniform ASI applies the power-tuned ASI correction under a uniform policy. Confidence ASI uses only self-confidence as the sampling signal. FSA ASI uses the FSA-selected features from Section 4.3. The ASI variants differ only in how they allocate the human-labeling budget.
| Dataset | Policy | Cov. | ESS Gain (%) |
| AutoFA | Classical | — | |
| Judge-predicted | — | ||
| Uniform ASI | |||
| Confidence ASI | |||
| FSA ASI | |||
| RAGTruth | Classical | — | |
| Judge-predicted | — | ||
| Uniform ASI | |||
| Confidence ASI | |||
| FSA ASI |
Metrics and simulation.
We report coverage, confidence interval width, and effective sample size (ESS). Coverage is the fraction of confidence intervals containing the human FAR computed on the full evaluation set. ESS measures efficiency on the scale of uniform annotation: an ESS gain of means that policy-sampled labels achieve the same variance as uniformly sampled labels. For each budget , we run 300 Monte Carlo trials on the evaluation split. In each trial, we bootstrap the evaluation set, sample human labels according to the corresponding policy, compute the estimator, and construct a Wald confidence interval at .
Results.
Table 4 and Figure 2 summarize the simulation results. The judge-predicted estimator has zero coverage on both datasets, confirming that judge FAR is biased relative to the human-anchored target. In contrast, all ASI variants maintain near-nominal coverage across budgets, showing that human-labeled correction successfully debiases the judge estimate.
Among the valid estimators, all ASI variants improve efficiency over the classical baseline. Uniform ASI already provides ESS gains of on AutoFA and on RAGTruth while Confidence ASI further improves efficiency to and , respectively, by allocating more labels to low-confidence examples.
FSA ASI achieves the largest gains, improving ESS by on AutoFA and on RAGTruth when aggregated across budgets. The larger gap over Confidence ASI on AutoFA reflects that self-confidence is a weaker residual-risk signal in this setting, while FSA identifies reference support strength as a more informative evidence-quality feature. On RAGTruth, the smaller gap is consistent with self-confidence already capturing much of the residual structure, with flag-reliability providing an additional gain.
5 Related Work
Factuality evaluation is central to assessing LLM outputs when correctness must be verified against retrieved evidence, source documents, or external references. Prior work has developed automatic factuality and hallucination evaluation methods that decompose responses into claims, atomic facts, or hallucinated spans, and assess whether these units are supported by evidence (Min et al., 2023; Chen et al., 2023; Niu et al., 2024; Iqbal et al., 2024; Song et al., 2024; Wei et al., 2024). These methods improve scalability, but LLM judges can systematically diverge from human labels due to incomplete evidence, over-literal grounding, rubric mismatch, and other biases in the judging procedure (Zheng et al., 2023; Wang et al., 2024; Gu et al., 2025).
Another line of works including Prediction-Powered Inference and its variants studies how to combine abundant machine predictions with limited human labels while preserving statistically valid inference (Angelopoulos et al., 2023; Angelopoulos et al., 2024). Active Statistical Inference extends this idea to budgeted annotation, where examples are sampled according to a policy designed to reduce estimator variance (Zrnic and Candes, 2024). Recent work adapts these ideas to LLM evaluation, showing how LLMaaJ metrics can be reported with valid uncertainty quantification (Gligorić et al., 2025; Wu et al., 2026; Lee et al., 2025). Cost optimal AI evaluation (Angelopoulos et al., 2025) extends ASI by incorporating the cost of human annotator and LLMaaJ into the optimization objective to derive cost-aware annotation policy. Meanwhile, MultiPPI (Cowen-Breen et al., 2026) and AM-PPI (Brawand et al., 2026) explore settings where multiple predictors are available and how to route between predictors to balance the cost-performance trade-off.
Closest to our work, confidence-driven inference (Gligorić et al., 2025) uses verbalized LLM confidence to guide human annotation, which provides a strong generic sampling signal. However, factuality residuals are not always explained by uncertainty alone: judges can be confidently misaligned when evidence is incomplete or only indirectly supportive. Our work adds a factuality-specific policy-design layer: FSA derives residual-risk features from judge artifacts, evidence quality, task strata, and rubric alignment, improving annotation efficiency while preserving the human-anchored ASI target.
6 Limitations
Our proposed framework is effective but is constrained by a few limitations:
First, FSA-guided ASI relies on residual structure learned from historical or calibration data. Its efficiency depends on this structure remaining stable across the evaluation population, i.e., model updates, retrieval or grounding changes, shifts in user queries, or new failure modes can degrade the policy. Silent high-confidence errors that are not captured under distribution shift will be under-sampled by the policy, thus inflate inverse-probability weights and increase variance. Uniform mixing can mitigate but does not remove the need for distribution-shift monitoring and periodic policy recalibration in production deployment. Specifically, a small uniformly sampled audit stream can test whether recent LLM–human residuals remain concentrated in the regions predicted by the policy. If this relationship degrades, the system can increase uniform mixing or revert temporarily to Uniform ASI, while acquiring additional human annotations to recalibrate the policy.
Second, while the four failure families provide useful pre-annotation signals for predicting disagreement in the studied settings, they do not exhaustively cover all error categories. Rare but complex cases, for example, failures in multi-hop reasoning across multiple pieces of evidence, may still occur and may not be captured by FSA. Depending on the application setting, the FSA rubric may therefore need to be further adapted or extended.
Third, our formulation also simplifies the annotation process. We treat the final adjudicated human label as the target and do not explicitly model annotator-level noise or disagreement. We also focus on a single human-annotator rather than a multi-annotator setting with different costs and reliabilities, such as cheap annotators, expert annotators, or multiple judge models. Extending FSA-guided ASI to jointly allocate budget across heterogeneous annotators is an important direction for future work.
Finally, ASI improves efficiency for a given annotation budget, but it does not determine the budget required for a particular production decision. In practice, the choice of depends on cost, latency, acceptable uncertainty, and release-risk tolerance. Our results quantify the variance reduction achieved by FSA-guided sampling, while budget selection itself remains an operational decision.
References
- Prediction-Powered Inference. Science 382 (6671), pp. 669–674. External Links: Document, Link, https://www.science.org/doi/pdf/10.1126/science.adi6000 Cited by: §5.
- PPI++: Efficient Prediction-Powered Inference. External Links: 2311.01453, Link Cited by: Appendix A, §4.4, §5.
- Cost-Optimal Active AI Model Evaluation. External Links: 2506.07949, Link Cited by: §5.
- Active Multiple-Prediction-Powered Inference. External Links: 2605.08429, Link Cited by: §5.
- FELM: Benchmarking Factuality Evaluation of Large Language Models. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §1, §5.
- XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, pp. 785–794. External Links: Link, Document Cited by: Appendix A.
- Multiple-Prediction-Powered Inference. External Links: 2603.27414, Link Cited by: §5.
- Can Unconfident LLM Annotations Be Used for Confident Conclusions?. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, pp. 3514–3533. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §1, §4.1, §5, §5.
- A Survey on LLM-as-a-Judge. External Links: 2411.15594, Link Cited by: §5.
- OpenFactCheck: a Unified Framework for Factuality Evaluation of LLMs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, D. I. Hernandez Farias, T. Hope, and M. Li (Eds.), Miami, Florida, USA, pp. 219–229. External Links: Link, Document Cited by: §5.
- How to Correctly Report LLM-as-a-Judge Evaluations. External Links: 2511.21140 Cited by: §5.
- Robust Sampling for Active Statistical Inference. External Links: 2511.08991, Link Cited by: Appendix A, §4.4.
- A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, Red Hook, NY, USA, pp. 4768–4777. External Links: ISBN 9781510860964 Cited by: Appendix B.
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 12076–12100. External Links: Link, Document Cited by: §1, §5.
- RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 10862–10878. External Links: Link, Document Cited by: Appendix C, §1, §4.1, §5.
- VeriScore: evaluating the Factuality of Verifiable Claims in Long-form Text Generation. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 9447–9474. External Links: Link, Document Cited by: §1, §4.1, §5.
- Large Language Models are not Fair Evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 9440–9450. External Links: Link, Document Cited by: §1, §5.
- Long-form Factuality in Large Language Models. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY, USA. External Links: ISBN 9798331314385 Cited by: §1, §4.1, §5.
- Efficient Evaluation of LLM Performance with Statistical Guarantees. External Links: 2601.20251, Link Cited by: §5.
- Judging LLM-as-a-judge with MT-bench and Chatbot Arena. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Cited by: §1, §5.
- Active Statistical Inference. In Proceedings of the 41st International Conference on Machine Learning, R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 62993–63010. External Links: Link Cited by: §1, §2.2, §2.3, §4.4, §5.
Appendix A Policy Learning Details
This appendix provides experimental details on implementing FSA-policy and running simulations.
Regression target.
The policy learning objective is to predict the residual magnitude from pre-labeling features. On both datasets, this target is binary: when the judge and human disagree, and when they agree. Figure 3 shows the class distribution. The disagreement rate is on AutoFA and on RAGTruth.
Score model.
We train an XGBoost regressor (Chen and Guestrin, 2016) with objective=reg:squarederror to predict from the selected features. On a binary target, this is equivalent to learning , producing a residual-risk score in . We use 200 boosting rounds, learning rate 0.05, and maximum tree depth 3. The model is trained on the 40% calibration split and its output serves as the raw score for policy construction.
Score normalization and uniform mixing.
Following the oracle ASI policy in Section 3, we first transform the predicted residual risk as , and then convert it into inclusion probabilities:
with water-filling to ensure . To prevent extreme weights that destabilize the ASI variance estimate, we also apply uniform mixing (Li et al., 2025):
where interpolates between the learned policy () and uniform sampling (). Larger improves stability at the cost of policy concentration.
Mixing coefficient selection.
We select by grid search over , evaluating the ASI variance on the calibration split for each candidate. Figure 4 shows that both datasets exhibit a concave ESS-gain curve: too-small creates extreme inverse-probability weights that inflate variance, while too-large dilutes the policy signal. The selected values are for AutoFA and for RAGTruth. The higher on RAGTruth reflects its larger ratio, which requires a larger mixing coefficient for stability.
Power tuning.
The ASI estimator uses a power-tuning coefficient that controls how strongly the judge contribution enters the correction (Angelopoulos et al., 2024). At , the estimator uses the full judge signal; at , it ignores the judge entirely. Power-tuning selects to minimize the policy-dependent variance term:
Since is quadratic in , the optimum has a closed-form solution:
where . In practice, we estimate from the calibration split using sample averages. Because is convex, the tuned estimator always has variance no greater than the untuned () estimator.
Figure 5 illustrates the effect on AutoFA: without power-tuning (), the ASI variance exceeds the labeled baseline at small budgets because the judge residual is large. Power-tuning () shrinks the judge contribution to the level most appropriate for the residual structure, consistently reducing variance below the baseline.
Appendix B Feature Selection Details
This appendix describes the feature selection procedure and provides post-hoc validation.
Selection procedure.
Feature candidates are derived from the qualitative FSA in Section 4.2. We evaluate each candidate’s residual-ranking quality via 5-fold cross-validated Spearman correlation on the calibration set. We then select the final feature set using two criteria: (1) high CV Spearman, and (2) parsimony: we prefer a smaller feature set because, with only calibration samples, using many features can introduce spurious correlations and cause the score model to overfit.
Joint feature importance.
Figure 6 shows SHAP values (Lundberg and Lee, 2017) for the score model trained on all candidate features, confirming which features drive the model’s predictions. On AutoFA, ref-support-strength dominates. On RAGTruth, self-confidence and flag-reliability jointly dominate. Features from other families contribute marginally.


Post-hoc policy validation.
To further evaluate how the selected features translate into effective policies, we compare held-out ESS gain across feature configurations in Table 5. On AutoFA, ref-support-strength () alone outperforms self-confidence () alone by a large margin, validating that FSA identifies a stronger signal. On RAGTruth, self-confidence () is already strong; the addition of flag-reliability provides a modest improvement (). In both cases, adding more features beyond the selected set does not improve ESS, consistent with the observation of overfitting at limited calibration size.
| Dataset | Feature configuration | CV | ESS gain |
| AutoFA | Self-confidence only | 0.12 | |
| Ref-support-strength | 0.50 | ||
| All candidates (5) | 0.52 | ||
| RAGTruth | Self-confidence only | 0.63 | |
| Self-confidence + flag-reliability | 0.69 | ||
| All candidates (6) | 0.71 |
The key observation is that higher univariate or multivariate correlation does not guarantee better policy performance. On AutoFA, the “all features” model achieves CV but only ESS, while the single-feature model achieves ESS. This occurs because the multi-feature model requires more aggressive uniform mixing () to stabilize extreme weights, which dilutes the policy concentration. The single feature provides smoother, full-population ranking that translates more effectively into sampling probabilities.
Appendix C Human Annotation Details
Both datasets leverage human annotators to acquire reference labels and we summarize the details as follows: RAGTruth (Niu et al., 2024) reports two independent annotators per response and a third review for substantial disagreements, and consistency rates of at the response level and at the span level. Annotators were recruited through a professional vendor and compensated at per hour. For AutoFA, all records received a human annotation, and received a second, independent annotation. For records where two annotators disagree with each other on the final response factual verdict, a third human moderator was involved to adjudicate and made the final judgment. The human annotators are instructed to extract core, verifiable claims from the model responses and search the web to judge the factual correctness of the claims and final overall response level factual correctness.
Appendix D Qualitative Analysis
We provide two qualitative examples from RAGTruth illustrating LLM judge errors and how FSA identifies them.
In record 168, the response states that Walter Scott was “shot whilst running from Officer Michael Slager.” The reference states that Scott was “fatally shot in the back by a police officer” and separately describes video showing Scott running from the officer while the officer fires eight shots. The response therefore combines two closely connected facts into a concise, well-supported statement, but the LLM judge flags “shot whilst running” as hallucinated, apparently requiring the temporal relation to be stated verbatim. Importantly, the LLM judge assigns 0.85 self-confidence, so confidence-based sampling would not identify this as a high-risk record. In contrast, the FSA feature flag-reliability receives a score of 2 (higher means LLM-judge is more likely to be wrong), indicating elevated over-detection risk and allowing the residual-risk policy to prioritize the record for human annotation.
In record 5718, the structured reference lists the same hours, 12:00–17:00, for every day from Monday through Sunday. The response accurately summarizes this as operating “seven days a week from 12:00 to 17:00,” yet the LLM judge flags the statement, apparently failing to aggregate the seven JSON entries. The LLM judge reports a low self-confidence score of 0.30 while the FSA assigns the maximum unreliability (flag-reliability=3).
Appendix E Feature Extraction Prompts
This appendix lists the prompts used for self-confidence scoring and FSA feature extraction. All prompts are written as Jinja2 templates: variables enclosed by {{ }} are replaced with record-specific content at inference time. The prompts are used only to extract auxiliary features or confidence scores, not to determine the final human label.
E.1 Self-Confidence Prompt
We use the prompt in Figure 7 to elicit the judge’s confidence in its own previous verdict. The same structure is used for both datasets, with minor wording changes depending on whether the underlying task is factuality verification or hallucination detection.
E.2 AutoFA Feature Extraction Prompt
The prompt in Figure 8 extracts five FSA features for AutoFA. The judge assesses the verification difficulty of a response given the query and three reference answers. It is explicitly instructed not to make the final factuality decision.
E.3 RAGTruth Feature Extraction Prompt
The prompt in Figure 9 extracts general verification-difficulty features for RAGTruth records. The general verification-difficulty features characterize the response–reference relationship, while flag_reliability additionally assesses the judge’s flagged spans.