arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00805v1 [cs.AI] 01 Sep 2026

Towards a Reliable and Practical Eval Pipeline

Emma Thuong Nguyen Affiliation: Salesforce Email: emmanguyen@salesforce.com    Abhishek Ghose Affiliation: Salesforce Email: aghose@salesforce.com
Abstract

LLM-based software systems increasingly require effective “evals” as quality gates in the development lifecycle. However, existing work typically addresses individual aspects of eval reliability rather than the full set of practical requirements. We present an end-to-end eval pipeline that combines eval checklist creation, with learned aggregation for checklist responses, to improve agreement across LLM judges and accuracy against human judgments. The framework additionally provides self-consistency, explanations, and prediction uncertainty, and we empirically demonstrate its effectiveness.

1 Introduction

Over the past few years Large Language Models (LLM) have become a common ingredient in software, where they assist users in a variety of tasks such as drafting emails, writing and summarizing documents. This has engendered novel challenges in the software development process, where such product features are required to be released only after passing through rigorous quality tests. Here, the traditional practice of writing purely programmatic tests no longer suffices, and this step needs to utilize LLMs as well. For example, if a particular LLM-driven feature produces summaries, a corresponding test that checks for its coherence or fluency also typically needs to use an LLM.

This creates an interesting problem: the fundamental behaviors we want our tests to check, e.g., effects of LLM non-determinism, are also possessed by the tests themselves. How do we then ensure that our tests precisely measure what we intend them to, and also provide guarantees similar to a traditional quality assurance framework?

Effective LLM-as-judge evaluations (“evals” henceforth) are a topic of active research, e.g., G-Eval (Liu et al., 2023), CheckEval (Lee et al., 2025), Prometheus-2 (Kim et al., 2024), FActScore (Min et al., 2023), dynamic rubrics based on information gain (Xu et al., 2026), hybrid judges for image caption evaluation (Matsuda et al., 2025), GraphJudge for evaluating knowledge graphs (Huang et al., 2025). However, most research focuses on specific aspects of evals; usually certain forms of alignment with human judges. We believe that this state of the current literature leaves open many questions wrt practical deployment .

These gaps exist both in form of (a) not examining different aspects of alignment, and (b) not covering the full gamut of practical use-cases. In a sense, recent research leans towards addressing challenges that may be deemed necessary, but are not sufficient for a practical setup. Our focus is the latter.

Our primary contributions are: (a) first, we present the desiderata of a practical eval framework, outlining risks and benefits, and (b) then, we propose such a framework, and empirically show that it effectively satisfies those needs. Rather than a single novel idea, this work offers an assembly of strategies to create an end-to-end eval pipeline that can augment a standard quality assurance pipeline.

The paper is organized as follows: we begin by listing the desiderata for a practical eval framework in §2, and then we present an overview of our framework in §3. Sections 4 - 7 present our empirical analyses. Our final section, §8, discusses limitations of the current study, and our conclusions.

2 A Practical Eval Pipeline

We believe that a practical eval system must at least ensure the following properties:

  • P1.

    Inter-LLM Agreement (or agreement in short): An eval output must exhibit minimal variance across LLMs. The lack of this property is a business risk. LLM A might clear all tests today, but when evals are switched to use LLM B (maybe due to an organization’s policy) that’s a harsher critic, we might end up with high failure rates.

  • P2.

    Accuracy: Eval outcomes must closely match human judgments. This is in contrast to most empirical analyses where some form of correlation, e.g., Kendall’s τ\tau or Spearman correlation coefficient ρ\rho, is measured, e.g., Lee et al. (2025); Liu et al. (2023). Product quality gates are almost always hard thresholds, e.g., “the coherence score for this summary must be ≥4\geq 4”. For such scenarios, there is a budgeting risk in productizing research with reported success in correlation-type metrics.

  • P3.

    Self-consistency: An eval output should exhibit minimal variation across multiple executions against the same LLM. This directly impacts how many times a test needs to be run for it to certify that a feature works. Thus, this impacts development cost and time-to-ship.

  • P4.

    Explainability: In traditional software testing, it is easy to trace failures to pinpoint a cause, by looking at a stack trace or logs. This is difficult when evals are complex and are written as large text prompts. What’s the equivalent of a stack trace in this new world?

  • P5.

    Confidence scores: Since LLM-based evals are probabilistic, we need to provide some representation of confidence for eval outcomes. We need to move away from hard predictions, e.g., “the eval score is 5” to soft estimates, e.g., “the eval score has a 90% prediction interval of 3-5”.

P1-P3 characterize forms of alignment, while P4 and P5 represent additional critical use-cases.

Refer to caption
Figure 1: Our eval pipeline. A checklist of YES/NO eval questions are sent to an LLM, along with inputs for the eval. In this example, the eval assesses the quality of summarization, so a document and its corresponding summary constitute our additional inputs. The outputs from the LLM are sent to a tabular model, e.g., Gradient Boosted Decision Trees, to produce the final judgment score. An explanation, and some form of confidence score, are also produced. See §3 for details.

3 Methodology

Our eval pipeline is shown in Figure 1. All evals are represented as a checklist of dd atomic YES/NO questions. The checklist, along with eval inputs, are sent to an LLM, which produces a binary vector of responses x∈{0,1}dx\in\{0,1\}^{d}. A model f:{0,1}d↦ℝf:\{0,1\}^{d}\mapsto\mathbb{R} learned earlier is used to “aggregate” these responses to predict a judgment y=f⁡(x)y=f(x). We produce two additional outputs: (1) explanations using SHAP, which help a user trace back a judgment to specific questions, and (2) a representation of confidence, either in the form of a calibrated probability (Platt, 1999; Guo et al., 2017; Niculescu-Mizil and Caruana, 2005) or a high-confidence interval via conformal prediction (Shafer and Vovk, 2008).

The checklist and the model are created in a prior eval-authoring step. This requires human inputs in the form of: (a) an eval prompt - we’ll refer to this as the seed prompt, and (b) sample inputs and outputs for their eval. The step produces, as output, (1) an eval checklist, and (2) a model ff.

The seed prompt is decomposed into multiple simpler YES/NO questions. The idea is that (a) having targeted atomic questions leaves little room for ambiguity, and (b) having multiple of them acts as an “error buffer”: even if the LLM misinterprets one question, it is unlikely that it will misinterpret every related question, which leads to a stable fraction of YES responses. This strategy leads to both improved agreement and self-consistency (empirically shown in §5).

We are motivated to follow the above strategy due to its success in a variety of eval settings (Min et al., 2023; Lee et al., 2025; Liu et al., 2024; Li et al., 2025; Xu et al., 2026). We specifically use CheckEval (Lee et al., 2025) because of its good results wrt inter-LLM agreement.

Since our goal also is accuracy wrt human judgment, we use the eval input-output examples to construct our model ff. The technique of using a model in the final step has successful precedents (Hashemi et al., 2024; Liu et al., 2024). We use Gradient Boosted Decision Trees (GBDT) (Friedman, 2001; Ke et al., 2017) in this work, which gives us positive results (see §5).

In addition to being a powerful model family, GBDTs offer other benefits: (a) they may be easily learned for a specific output type, e.g., yy may be boolean, ordinal, categorical or continuous; (b) we use SHAP (Lundberg and Lee, 2017) to generate prediction explanations, which permits exact explanations for tree-based models such as GBDTs.

4 Experiment Setup

In our experiments, we rigorously measure agreement, self-consistency and accuracy, and present persuasive results for confidence scoring and explanations (these rely on standard techniques).

Dataset and LLMs: The SummEval (Fabbri et al., 2021) dataset (MIT License) is used for experiments. It consists of documents, their corresponding summaries (more than one summary per document), and human-provided scores per summary that assess their quality across different axes. These axes are coherence, consistency, fluency and relevance, and we have one eval corresponding to each axis11 1 We’ll use these interchangeably - a quality axis is assessed by an eval.. Each summary, per quality axis, is rated by multiple human raters, with the provided scores being in {1,2,3,4,5}\{1,2,3,4,5\}. The summaries are our measurement units, and for summary sis_{i} and an eval axis, we average the human ratings to obtain the ground truth (GT) label yi∈[1,5]y_{i}\in[1,5].

We generate LLM scores for two additional configuration parameters :

  1. 1.

    LLMs: We use L=4L=4 different LLMs to measure agreement. The LLMs used are Opus 4.8 (abbreviated to opus), Sonnet 4.6 (sonnet), GPT 5.6 Sol (sol), Grok 4.6 (grok). For sonnet the temperature was set to 00; the others don’t support this parameter.

  2. 2.

    Number of trials: Each eval execution is repeated TT times to measure statistical significance. In our experiments T=5T=5.

These additional configurations will be denoted by the indices ll and tt respectively; the predicted scores are denoted as y^i​l​t\hat{y}_{ilt} for summary sis_{i} for a particular eval axis.

Assuming a checklist size of dd questions for an eval, a response vector generated for various combinations is denoted by xi​l​t∈{0,1}dx_{ilt}\in\{0,1\}^{d} . We train one GBDT model per eval, on tuples of the form (xi​l​t,yi)(x_{ilt},y_{i}) to minimize R​M​S​E​(yi,y^i​l​t)RMSE(y_{i},\hat{y}_{ilt}) on a held-out set. Thus, for an eval, each GT label is paired with L×TL\times T inputs. Essentially, the model is trained to map from various noisy realizations of xx to a GT label. We use a train-conformal-test split of 50:10:4050:10:40 (the conformal split is used in §7), and to prevent data leakage, we ensure that data related to a specific sis_{i} is in exactly one split. For additional details, please see §A.

Metrics: As mentioned above, accuracy is measured using the RMSE score. Lower is better.

For measuring agreement and self-consistency (note, these scores don’t use GT, only y^i​l​t\hat{y}_{ilt} ), we do not use the popular Krippendorff’s-α\alpha score because: (a) we have multiple trials per rater/LLM, which α\alpha doesn’t naturally accommodate, and (b) α\alpha might be negative (Artstein and Poesio, 2008), and we prefer a range of [0,1][0,1] for easy interpretation.

Instead we map our scores from a range of [1,5][1,5] to [0,1][0,1], and define the disagreement between two scores aa and bb as D=|a−b|D=|a-b|. For an eval, this is averaged over trials for a measurement unit (here, a summary) to produce D¯i\bar{D}_{i}. For agreement, the averaging for two given LLMs22 2 In our setup, to claim that an output is from a specific LLM A means that the checklist responses x∈{0,1}dx\in\{0,1\}^{d} were obtained using A, and then the learned aggregator model ff for an eval was used to produce the final judgment. is over their T2T^{2} possible trial combinations, and for self-consistency the T⁡(T−1)/2T(T-1)/2 unique combinations for an LLM are used. Finally we report 1−D¯i1-\bar{D}_{i} as our score (agreement or self-consistency, depending on the averaging) for this unit. Further averaging is called out in our analyses in §5.

Of course, since these scores are not chance-adjusted, they should be interpreted jointly with accuracy.

Hardware: The LLMs were accessed through an Amazon Bedrock API. All experiments were run on a M3 Max Apple Macbook Pro laptop.

5 Results

This section presents results from our setup as compared to a baseline. The latter uses seed prompts from prior literature - see §A.1.

Agreement: The average agreement for a pair of LLMs is shown as 4×44\times 4 heatmap in Figure 2. A particular cell further averages 1−D¯i1-\bar{D}_{i} over all summaries and evals in a common held-out set. Figure 2(a) shows baseline results, and both (b) and (c) come from our setup.

In Figure 2(b), the aggregator model ff has seen data from all LLMs during training, and a cell indexed by [A,B][A,B] compares the agreement between LLMs AA and BB on a held-out set. We notice that (b) produces much higher scores than (a), and the overall average (in the plot title) improves: 0.84→0.960.84\rightarrow 0.96.

Perhaps more interesting is Figure 2(c). Here an LLM is held-out as well, and we cycle through all the LLMs, denoted by rows. For example, the first row indicates that the model ff only saw data from opus, sol and sonnet in the training data, but the agreement is computed between each of these and grok, on a held-out set. Even here, we note high scores for all pairs.

Refer to caption
(a) Baseline agreements.
Refer to caption
(b) Agreements - all LLMs in train/test.
Refer to caption
(c) Agreements - with one held-out LLM.
Figure 2: LLM agreements on held-out test data (higher is better). (a) shows baseline agreements, (b) and (c) are from our setup. In (b) data from all LLMs were present both in train and test splits. In (c), for a row, data from the LLM indexing the row was absent in train, but agreements are reported with the other LLMs on a held-out set. See text for details.

Self-consistency: The self-consistency scores on the baseline prompts, and then on the output from our framework, are shown in Table 1. The score for an LLM are averaged over summaries and evals. Here, the scores look good to start with, although Sol and Grok seem to benefit most from our approach, and the overall average slightly improves.

grok opus sol sonnet Average
Baseline 0.94 0.97 0.96 0.99 0.97
Our setup 0.97 0.99 0.98 0.99 0.98
Table 1: Self-consistency averaged across all evals, per LLM. Higher is better.

Accuracy: Accuracy scores, per eval, are shown as grouped bar plots in Figure 3. Each group shows a bar for the baseline RMSE (labeled Baseline), aggregation via just reporting the fraction of YES answers (labeled CheckEval), and the output from the entire pipeline (labeled CheckEval+ML).

We show the CheckEval bars to emphasize that reducing variability - which CheckEval is effective at when the fraction of YESes are reported as the final score (as reported in the original work Lee et al. (2025)) is insufficient to obtain good accuracy. In fact, for the coherence eval, CheckEval is worse than Baseline. In all cases, CheckEval+ML produces the lowest RMSE score.

Figure 3: Eval accuracy measured using RMSE (lower is better).

6 Explanations

It is often not enough to know why a test produced a certain result; the ability to trace an outcome back to causes provides an actionable path to improving a product feature. For a typical LLM eval driven by one prompt, since both the prompt and input might be long, pinpointing an exact fragment in the input, and the corresponding violated rubric in the eval prompt, is much harder. The simple strategy of asking the LLM to explain its judgment has been shown to be unreliable: LLMs tend to favor plausibility over faithfulness (Turpin et al., 2023; Madsen et al., 2024; Chen et al., 2025).

Our framework is able to sidestep this issue:

Figure 4: A standard SHAP waterfall chart for a particular input - for the consistency eval, on opus - showing influences of various questions in a checklist, in decreasing order from top-to-bottom.

7 Prediction Confidence

We discuss how we quantify uncertainty, which arises due to various factors such as variability across LLMs, non-determinism, and eval inputs that are poorly represented in the training data.

While this is convenient with a model like Gaussian Process (Rasmussen and Williams, 2005), as mentioned earlier, we prefer a tree-based model due to high-fidelity SHAP explanations. We use conformal predictions (CP) here to report prediction intervals that are likely to contain the true outcome 90%90\% of the time. Being model-agnostic, CP also makes this step invariant to future model changes. We use the residual-normalized version (Lei et al., 2018; Cordier et al., 2023) to make the intervals input-adaptive.

To validate, we generate synthetic data from valid xi​l​t∈{0,1}dx_{ilt}\in\{0,1\}^{d} vectors by flipping b∈{0.2,0.4,0.6}b\in\{0.2,0.4,0.6\} fraction of bits, and visualize the distribution of the interval widths; shown in Figure 5, for the (a) consistency, and (b) coherence evals. The “inliers” KDE plot is on the test data, and is expected to have relatively small widths; the legend shows its empirical coverage. In both cases, we observe that higher values bb lead to wider intervals (expected behavior). See §A.4 for more plots.

(a) consistency eval
(b) coherence eval
Figure 5: Distribution of interval widths on actual (inliers) vs synthetic noisy data.

8 Limitations and Conclusion

In this work we presented a reliable eval pipeline that meets multiple practical requirements, listed in §2. We think this work is important in bringing academic research into the industry; even if the exact pipeline isn’t used, the general architecture and specific components may be built upon. An obvious limitation of this work is that it is validated on a single dataset for the task of summarization, against proprietary LLMs; something that the authors are working to address.

9 Acknowledgements

This work was possible with the support of the Q3 Salescloud team in Salesforce, specifically Ronnie Fong, Nitin Jain and William Hackett. We are grateful to Harsha Medikonda, Prasad Maderamitla, Hinita Patel, Aaron Quan and Soumya Mittapalli, for pointing us to multiple real-world challenges. We also would like to thank Pratik Gupte, Kimberley Lee and Daniel May for various technical suggestions.

References

  • Artstein and Poesio (2008) R. Artstein and M. Poesio Inter-coder agreement for computational linguistics. Computational Linguistics 34 (4), pp. 555–596. External Links: ISSN 0891-2017, Document, Link, https://direct.mit.edu/coli/article-pdf/34/4/555/1808947/coli.07-034-r2.pdf Cited by: §4.
  • Chen et al. (2025) Y. Chen, J. Benton, A. Radhakrishnan, J. Uesato, C. Denison, J. Schulman, A. Somani, P. Hase, M. Wagner, F. Roger, V. Mikulik, S. R. Bowman, J. Leike, J. Kaplan, and E. Perez Reasoning models don’t always say what they think. External Links: 2505.05410, Link Cited by: §6.
  • Cordier et al. (2023) T. Cordier, V. Blot, L. Lacombe, T. Morzadec, A. Capitaine, and N. Brunel Flexible and systematic uncertainty estimation with conformal prediction via the mapie library. In Proceedings of the Twelfth Symposium on Conformal and Probabilistic Prediction with Applications, H. Papadopoulos, K. A. Nguyen, H. Boström, and L. Carlsson (Eds.), Proceedings of Machine Learning Research, Vol. 204, pp. 549–581. External Links: Link Cited by: §7.
  • Fabbri et al. (2021) A. R. Fabbri, W. Kryściński, B. McCann, C. Xiong, R. Socher, and D. Radev SummEval: re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics 9, pp. 391–409. External Links: ISSN 2307-387X, Document, Link, https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl_a_00373/1923949/tacl_a_00373.pdf Cited by: §4.
  • Friedman (2001) J. H. Friedman Greedy function approximation: A gradient boosting machine.. The Annals of Statistics 29 (5), pp. 1189 – 1232. External Links: Document, Link Cited by: §3.
  • Guo et al. (2017) C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, pp. 1321–1330. Cited by: §3.
  • Hashemi et al. (2024) H. Hashemi, J. Eisner, C. Rosset, B. Van Durme, and C. Kedzie LLM-rubric: a multidimensional, calibrated approach to automated evaluation of natural language texts. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 13806–13834. External Links: Link, Document Cited by: §3.
  • Huang et al. (2025) H. Huang, C. Chen, Z. Sheng, Y. Li, and W. Zhang Can LLMs be good graph judge for knowledge graph construction?. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 10929–10948. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1.
  • Ke et al. (2017) G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T. Liu LightGBM: a highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30, pp. . External Links: Link Cited by: §3.
  • Kim et al. (2024) S. Kim, J. Suk, S. Longpre, B. Y. Lin, J. Shin, S. Welleck, G. Neubig, M. Lee, K. Lee, and M. Seo Prometheus 2: an open source language model specialized in evaluating other language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 4334–4353. External Links: Link, Document Cited by: §1.
  • Lee et al. (2025) Y. Lee, J. Kim, J. Kim, H. Cho, J. Kang, P. Kang, and N. Kim CheckEval: a reliable LLM-as-a-judge framework for evaluating text generation using checklists. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 15771–15798. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: Table 2, Table 2, §1, item P2., §3, §5.
  • Lei et al. (2018) J. Lei, M. G’Sell, A. Rinaldo, R. J. Tibshirani, and L. Wasserman Distribution-free predictive inference for regression. Journal of the American Statistical Association 113 (523), pp. 1094–1111. External Links: Document, Link, https://doi.org/10.1080/01621459.2017.1307116 Cited by: §7.
  • Li et al. (2025) M. Li, Z. Liu, S. Deng, S. Joty, N. Chen, and M. Kan DnA-eval: enhancing large language model evaluation through decomposition and aggregation. In Proceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Schockaert (Eds.), Abu Dhabi, UAE, pp. 2277–2290. External Links: Link Cited by: §3.
  • Liu et al. (2023) Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu G-eval: NLG evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 2511–2522. External Links: Link, Document Cited by: §A.1, §1, item P2..
  • Liu et al. (2024) Y. Liu, T. Yang, S. Huang, Z. Zhang, H. Huang, F. Wei, W. Deng, F. Sun, and Q. Zhang HD-eval: aligning large language model evaluators through hierarchical criteria decomposition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 7641–7660. External Links: Link, Document Cited by: §3, §3.
  • Lundberg et al. (2019) S. M. Lundberg, G. G. Erion, and S. Lee Consistent individualized feature attribution for tree ensembles. External Links: 1802.03888, Link Cited by: 2nd item.
  • Lundberg and Lee (2017) S. M. Lundberg and S. Lee A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, Red Hook, NY, USA, pp. 4768–4777. External Links: ISBN 9781510860964 Cited by: §3.
  • Madsen et al. (2024) A. Madsen, S. Chandar, and S. Reddy Are self-explanations from large language models faithful?. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 295–337. External Links: Link, Document Cited by: §6.
  • Matsuda et al. (2025) K. Matsuda, Y. Wada, S. Hirano, S. Otsuki, and K. Sugiura VELA: an LLM-hybrid-as-a-judge approach for evaluating long image captions. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 8680–8696. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1.
  • Min et al. (2023) S. Min, K. Krishna, X. Lyu, M. Lewis, W. Yih, P. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi FActScore: fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 12076–12100. External Links: Link, Document Cited by: §1, §3.
  • Niculescu-Mizil and Caruana (2005) A. Niculescu-Mizil and R. Caruana Obtaining calibrated probabilities from boosting. In Proceedings of the Twenty-First Conference on Uncertainty in Artificial Intelligence, UAI’05, Arlington, Virginia, USA, pp. 413–420. External Links: ISBN 0974903914 Cited by: §3.
  • Platt (1999) J. C. Platt Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. In Advances in Large Margin Classifiers, pp. 61–74. Cited by: §3.
  • Rasmussen and Williams (2005) C. E. Rasmussen and C. K. I. Williams Gaussian processes for machine learning. The MIT Press. External Links: ISBN 9780262256834, Document, Link, https://direct.mit.edu/book-pdf/2514321/book_9780262256834.pdf Cited by: §7.
  • Shafer and Vovk (2008) G. Shafer and V. Vovk A tutorial on conformal prediction. Journal of Machine Learning Research 9, pp. 371–421 (English). External Links: ISSN 1532-4435 Cited by: §3.
  • Turpin et al. (2023) M. Turpin, J. Michael, E. Perez, and S. Bowman Language models don't always say what they think: unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, pp. 74952–74965. External Links: Document, Link Cited by: §6.
  • Xu et al. (2026) W. Xu, S. Zhao, H. Chen, T. Yan, C. Wang, S. Chen, and Q. Wan Beyond drift: stabilizing subjective LLM evaluation with information-theoretic rubrics. In Forty-third International Conference on Machine Learning, External Links: Link Cited by: §1, §3.

Appendix A Additional Experiment Details

A.1 Baseline Prompts

We use the G-Eval [Liu et al., 2023] prompts without the autogenerated chain-of-thoughts evaluation steps as the baseline prompts in our experiment.

Coherence You will be given one summary written for a news article. Your task is to rate the summary on one metric. Please make sure you read and understand these instructions carefully. Please keep this document open while reviewing, and refer to it as needed. Evaluation Criteria: Coherence (1-5) - the collective quality of all sentences. We align this dimension with the DUC quality question of structure and coherence whereby "the summary should be well-structured and well-organized. The summary should not just be a heap of related information, but should build from sentence to a coherent body of information about a topic." Source Text: {{Document}} Summary: {{Summary}} Evaluation Form (scores ONLY): - Coherence:
Consistency You will be given a news article. You will then be given one summary written for this article. Your task is to rate the summary on one metric. Please make sure you read and understand these instructions carefully. Please keep this document open while reviewing, and refer to it as needed. Evaluation Criteria: Consistency (1-5) - the factual alignment between the summary and the summarized source. A factually consistent summary contains only statements that are entailed by the source document. Annotators were also asked to penalize summaries that contained hallucinated facts. Source Text: {{Document}} Summary: {{Summary}} Evaluation Form (scores ONLY): - Consistency:
Fluency You will be given one summary written for a news article. Your task is to rate the summary on one metric. Please make sure you read and understand these instructions carefully. Please keep this document open while reviewing, and refer to it as needed. Evaluation Criteria: Fluency (1-5): the quality of individual sentences. Drawing again from the DUC quality guidelines, sentences in the summary "should have no formatting problems, capitalization errors or obviously ungrammatical sentences (e.g., fragments, missing components) that make the text difficult to read." Example: Source Text: {{Document}} Summary: {{Summary}} Evaluation Form (scores ONLY): - Fluency:
Relevance You will be given one summary written for a news article. Your task is to rate the summary on one metric. Please make sure you read and understand these instructions carefully. Please keep this document open while reviewing, and refer to it as needed. Evaluation Criteria: Relevance (1-5) - selection of important content from the source. The summary should include only important information from the source document. Annotators were instructed to penalize summaries which contained redundancies and excess information. Example: Source Text: {{Document}} Summary: {{Summary}} Evaluation Form (scores ONLY): - Relevance:

A.2 Generated Checklists

Dimension # Sub-dimensions # Seed Questions # Final Questions
Coherence 3 3 20
Consistency 3 3 16
Fluency 4 4 24
Relevance 5 5 21
Table 2: Checklist size per SummEval dimension. The sub-dimensions and seed questions were taken from CheckEval [Lee et al., 2025]. We used Opus 4.8 to generate the final questions.

A.3 GBDT Hyperparameters

The GBDT in §4 was learned using cross-validation of hyperparameters from the following search space:

    "n_estimators": [20, 50, 100, 200]
    "learning_rate": [0.001, 0.01, 0.1, 1.0]
    "max_depth": [3, 5, 10]
    "min_child_samples": [5, 10, 20, 40]
    "reg_lambda": [0.0, 1.0, 5.0]
  

A.4 Additional Conformal Prediction Results

Continuing from §7, we present the plots for all evals in Figure 6.

(a) consistency eval
(b) coherence eval
(c) fluency eval
(d) relevance eval
Figure 6: Distribution of interval widths for points in the data or “inliers” and synthetic noisy data, which is constructed by flipping bit for different fractions of bits b∈{0.2,0.4,0.6}b\in\{0.2,0.4,0.6\}. The empirical coverage of the inliers is shown in the legend, which may be observed to be quite close to intended coverage of 90%90\%.