Towards a Reliable and Practical Eval Pipeline
Abstract
LLM-based software systems increasingly require effective “evals” as quality gates in the development lifecycle. However, existing work typically addresses individual aspects of eval reliability rather than the full set of practical requirements. We present an end-to-end eval pipeline that combines eval checklist creation, with learned aggregation for checklist responses, to improve agreement across LLM judges and accuracy against human judgments. The framework additionally provides self-consistency, explanations, and prediction uncertainty, and we empirically demonstrate its effectiveness.
1 Introduction
Over the past few years Large Language Models (LLM) have become a common ingredient in software, where they assist users in a variety of tasks such as drafting emails, writing and summarizing documents. This has engendered novel challenges in the software development process, where such product features are required to be released only after passing through rigorous quality tests. Here, the traditional practice of writing purely programmatic tests no longer suffices, and this step needs to utilize LLMs as well. For example, if a particular LLM-driven feature produces summaries, a corresponding test that checks for its coherence or fluency also typically needs to use an LLM.
This creates an interesting problem: the fundamental behaviors we want our tests to check, e.g., effects of LLM non-determinism, are also possessed by the tests themselves. How do we then ensure that our tests precisely measure what we intend them to, and also provide guarantees similar to a traditional quality assurance framework?
Effective LLM-as-judge evaluations (“evals” henceforth) are a topic of active research, e.g., G-Eval (Liu et al., 2023), CheckEval (Lee et al., 2025), Prometheus-2 (Kim et al., 2024), FActScore (Min et al., 2023), dynamic rubrics based on information gain (Xu et al., 2026), hybrid judges for image caption evaluation (Matsuda et al., 2025), GraphJudge for evaluating knowledge graphs (Huang et al., 2025). However, most research focuses on specific aspects of evals; usually certain forms of alignment with human judges. We believe that this state of the current literature leaves open many questions wrt practical deployment .
These gaps exist both in form of (a) not examining different aspects of alignment, and (b) not covering the full gamut of practical use-cases. In a sense, recent research leans towards addressing challenges that may be deemed necessary, but are not sufficient for a practical setup. Our focus is the latter.
Our primary contributions are: (a) first, we present the desiderata of a practical eval framework, outlining risks and benefits, and (b) then, we propose such a framework, and empirically show that it effectively satisfies those needs. Rather than a single novel idea, this work offers an assembly of strategies to create an end-to-end eval pipeline that can augment a standard quality assurance pipeline.
2 A Practical Eval Pipeline
We believe that a practical eval system must at least ensure the following properties:
- P1.
Inter-LLM Agreement (or agreement in short): An eval output must exhibit minimal variance across LLMs. The lack of this property is a business risk. LLM A might clear all tests today, but when evals are switched to use LLM B (maybe due to an organization’s policy) that’s a harsher critic, we might end up with high failure rates.
- P2.
Accuracy: Eval outcomes must closely match human judgments. This is in contrast to most empirical analyses where some form of correlation, e.g., Kendall’s or Spearman correlation coefficient , is measured, e.g., Lee et al. (2025); Liu et al. (2023). Product quality gates are almost always hard thresholds, e.g., “the coherence score for this summary must be ”. For such scenarios, there is a budgeting risk in productizing research with reported success in correlation-type metrics.
- P3.
Self-consistency: An eval output should exhibit minimal variation across multiple executions against the same LLM. This directly impacts how many times a test needs to be run for it to certify that a feature works. Thus, this impacts development cost and time-to-ship.
- P4.
Explainability: In traditional software testing, it is easy to trace failures to pinpoint a cause, by looking at a stack trace or logs. This is difficult when evals are complex and are written as large text prompts. What’s the equivalent of a stack trace in this new world?
- P5.
Confidence scores: Since LLM-based evals are probabilistic, we need to provide some representation of confidence for eval outcomes. We need to move away from hard predictions, e.g., “the eval score is 5” to soft estimates, e.g., “the eval score has a 90% prediction interval of 3-5”.
P1-P3 characterize forms of alignment, while P4 and P5 represent additional critical use-cases.
3 Methodology
Our eval pipeline is shown in Figure 1. All evals are represented as a checklist of atomic YES/NO questions. The checklist, along with eval inputs, are sent to an LLM, which produces a binary vector of responses . A model learned earlier is used to “aggregate” these responses to predict a judgment . We produce two additional outputs: (1) explanations using SHAP, which help a user trace back a judgment to specific questions, and (2) a representation of confidence, either in the form of a calibrated probability (Platt, 1999; Guo et al., 2017; Niculescu-Mizil and Caruana, 2005) or a high-confidence interval via conformal prediction (Shafer and Vovk, 2008).
The checklist and the model are created in a prior eval-authoring step. This requires human inputs in the form of: (a) an eval prompt - we’ll refer to this as the seed prompt, and (b) sample inputs and outputs for their eval. The step produces, as output, (1) an eval checklist, and (2) a model .
The seed prompt is decomposed into multiple simpler YES/NO questions. The idea is that (a) having targeted atomic questions leaves little room for ambiguity, and (b) having multiple of them acts as an “error buffer”: even if the LLM misinterprets one question, it is unlikely that it will misinterpret every related question, which leads to a stable fraction of YES responses. This strategy leads to both improved agreement and self-consistency (empirically shown in §5).
We are motivated to follow the above strategy due to its success in a variety of eval settings (Min et al., 2023; Lee et al., 2025; Liu et al., 2024; Li et al., 2025; Xu et al., 2026). We specifically use CheckEval (Lee et al., 2025) because of its good results wrt inter-LLM agreement.
Since our goal also is accuracy wrt human judgment, we use the eval input-output examples to construct our model . The technique of using a model in the final step has successful precedents (Hashemi et al., 2024; Liu et al., 2024). We use Gradient Boosted Decision Trees (GBDT) (Friedman, 2001; Ke et al., 2017) in this work, which gives us positive results (see §5).
In addition to being a powerful model family, GBDTs offer other benefits: (a) they may be easily learned for a specific output type, e.g., may be boolean, ordinal, categorical or continuous; (b) we use SHAP (Lundberg and Lee, 2017) to generate prediction explanations, which permits exact explanations for tree-based models such as GBDTs.
4 Experiment Setup
In our experiments, we rigorously measure agreement, self-consistency and accuracy, and present persuasive results for confidence scoring and explanations (these rely on standard techniques).
Dataset and LLMs: The SummEval (Fabbri et al., 2021) dataset (MIT License) is used for experiments. It consists of documents, their corresponding summaries (more than one summary per document), and human-provided scores per summary that assess their quality across different axes. These axes are coherence, consistency, fluency and relevance, and we have one eval corresponding to each axis11 1 We’ll use these interchangeably - a quality axis is assessed by an eval.. Each summary, per quality axis, is rated by multiple human raters, with the provided scores being in . The summaries are our measurement units, and for summary and an eval axis, we average the human ratings to obtain the ground truth (GT) label .
We generate LLM scores for two additional configuration parameters :
- 1.
LLMs: We use different LLMs to measure agreement. The LLMs used are Opus 4.8 (abbreviated to opus), Sonnet 4.6 (sonnet), GPT 5.6 Sol (sol), Grok 4.6 (grok). For sonnet the temperature was set to ; the others don’t support this parameter.
- 2.
Number of trials: Each eval execution is repeated times to measure statistical significance. In our experiments .
These additional configurations will be denoted by the indices and respectively; the predicted scores are denoted as for summary for a particular eval axis.
Assuming a checklist size of questions for an eval, a response vector generated for various combinations is denoted by . We train one GBDT model per eval, on tuples of the form to minimize on a held-out set. Thus, for an eval, each GT label is paired with inputs. Essentially, the model is trained to map from various noisy realizations of to a GT label. We use a train-conformal-test split of (the conformal split is used in §7), and to prevent data leakage, we ensure that data related to a specific is in exactly one split. For additional details, please see §A.
Metrics: As mentioned above, accuracy is measured using the RMSE score. Lower is better.
For measuring agreement and self-consistency (note, these scores don’t use GT, only ), we do not use the popular Krippendorff’s- score because: (a) we have multiple trials per rater/LLM, which doesn’t naturally accommodate, and (b) might be negative (Artstein and Poesio, 2008), and we prefer a range of for easy interpretation.
Instead we map our scores from a range of to , and define the disagreement between two scores and as . For an eval, this is averaged over trials for a measurement unit (here, a summary) to produce . For agreement, the averaging for two given LLMs22 2 In our setup, to claim that an output is from a specific LLM A means that the checklist responses were obtained using A, and then the learned aggregator model for an eval was used to produce the final judgment. is over their possible trial combinations, and for self-consistency the unique combinations for an LLM are used. Finally we report as our score (agreement or self-consistency, depending on the averaging) for this unit. Further averaging is called out in our analyses in §5.
Of course, since these scores are not chance-adjusted, they should be interpreted jointly with accuracy.
Hardware: The LLMs were accessed through an Amazon Bedrock API. All experiments were run on a M3 Max Apple Macbook Pro laptop.
5 Results
This section presents results from our setup as compared to a baseline. The latter uses seed prompts from prior literature - see §A.1.
Agreement: The average agreement for a pair of LLMs is shown as heatmap in Figure 2. A particular cell further averages over all summaries and evals in a common held-out set. Figure 2(a) shows baseline results, and both (b) and (c) come from our setup.
In Figure 2(b), the aggregator model has seen data from all LLMs during training, and a cell indexed by compares the agreement between LLMs and on a held-out set. We notice that (b) produces much higher scores than (a), and the overall average (in the plot title) improves: .
Perhaps more interesting is Figure 2(c). Here an LLM is held-out as well, and we cycle through all the LLMs, denoted by rows. For example, the first row indicates that the model only saw data from opus, sol and sonnet in the training data, but the agreement is computed between each of these and grok, on a held-out set. Even here, we note high scores for all pairs.
Self-consistency: The self-consistency scores on the baseline prompts, and then on the output from our framework, are shown in Table 1. The score for an LLM are averaged over summaries and evals. Here, the scores look good to start with, although Sol and Grok seem to benefit most from our approach, and the overall average slightly improves.
| grok | opus | sol | sonnet | Average | |
|---|---|---|---|---|---|
| Baseline | 0.94 | 0.97 | 0.96 | 0.99 | 0.97 |
| Our setup | 0.97 | 0.99 | 0.98 | 0.99 | 0.98 |
Accuracy: Accuracy scores, per eval, are shown as grouped bar plots in Figure 3. Each group shows a bar for the baseline RMSE (labeled Baseline), aggregation via just reporting the fraction of YES answers (labeled CheckEval), and the output from the entire pipeline (labeled CheckEval+ML).
We show the CheckEval bars to emphasize that reducing variability - which CheckEval is effective at when the fraction of YESes are reported as the final score (as reported in the original work Lee et al. (2025)) is insufficient to obtain good accuracy. In fact, for the coherence eval, CheckEval is worse than Baseline. In all cases, CheckEval+ML produces the lowest RMSE score.
6 Explanations
It is often not enough to know why a test produced a certain result; the ability to trace an outcome back to causes provides an actionable path to improving a product feature. For a typical LLM eval driven by one prompt, since both the prompt and input might be long, pinpointing an exact fragment in the input, and the corresponding violated rubric in the eval prompt, is much harder. The simple strategy of asking the LLM to explain its judgment has been shown to be unreliable: LLMs tend to favor plausibility over faithfulness (Turpin et al., 2023; Madsen et al., 2024; Chen et al., 2025).
Our framework is able to sidestep this issue:
- •
Because our eval is distilled into a checklist, there are specific contributing questions to point back to.
- •
Using a tabular model as the aggregator allows us to use a popular technique like SHAP to identify questions most influential to a specific outcome. Additionally for GBDTs, TreeSHAP (Lundberg et al., 2019) provides exact SHAP attribution values. Figure 4 shows a waterfall plot33 3 https://shap.readthedocs.io/en/latest/example_notebooks/api_examples/plots/waterfall.html of question influences, produced by SHAP.
7 Prediction Confidence
We discuss how we quantify uncertainty, which arises due to various factors such as variability across LLMs, non-determinism, and eval inputs that are poorly represented in the training data.
While this is convenient with a model like Gaussian Process (Rasmussen and Williams, 2005), as mentioned earlier, we prefer a tree-based model due to high-fidelity SHAP explanations. We use conformal predictions (CP) here to report prediction intervals that are likely to contain the true outcome of the time. Being model-agnostic, CP also makes this step invariant to future model changes. We use the residual-normalized version (Lei et al., 2018; Cordier et al., 2023) to make the intervals input-adaptive.
To validate, we generate synthetic data from valid vectors by flipping fraction of bits, and visualize the distribution of the interval widths; shown in Figure 5, for the (a) consistency, and (b) coherence evals. The “inliers” KDE plot is on the test data, and is expected to have relatively small widths; the legend shows its empirical coverage. In both cases, we observe that higher values lead to wider intervals (expected behavior). See §A.4 for more plots.
8 Limitations and Conclusion
In this work we presented a reliable eval pipeline that meets multiple practical requirements, listed in §2. We think this work is important in bringing academic research into the industry; even if the exact pipeline isn’t used, the general architecture and specific components may be built upon. An obvious limitation of this work is that it is validated on a single dataset for the task of summarization, against proprietary LLMs; something that the authors are working to address.
9 Acknowledgements
This work was possible with the support of the Q3 Salescloud team in Salesforce, specifically Ronnie Fong, Nitin Jain and William Hackett. We are grateful to Harsha Medikonda, Prasad Maderamitla, Hinita Patel, Aaron Quan and Soumya Mittapalli, for pointing us to multiple real-world challenges. We also would like to thank Pratik Gupte, Kimberley Lee and Daniel May for various technical suggestions.
References
- Inter-coder agreement for computational linguistics. Computational Linguistics 34 (4), pp. 555–596. External Links: ISSN 0891-2017, Document, Link, https://direct.mit.edu/coli/article-pdf/34/4/555/1808947/coli.07-034-r2.pdf Cited by: §4.
- Reasoning models don’t always say what they think. External Links: 2505.05410, Link Cited by: §6.
- Flexible and systematic uncertainty estimation with conformal prediction via the mapie library. In Proceedings of the Twelfth Symposium on Conformal and Probabilistic Prediction with Applications, H. Papadopoulos, K. A. Nguyen, H. Boström, and L. Carlsson (Eds.), Proceedings of Machine Learning Research, Vol. 204, pp. 549–581. External Links: Link Cited by: §7.
- SummEval: re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics 9, pp. 391–409. External Links: ISSN 2307-387X, Document, Link, https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl_a_00373/1923949/tacl_a_00373.pdf Cited by: §4.
- Greedy function approximation: A gradient boosting machine.. The Annals of Statistics 29 (5), pp. 1189 – 1232. External Links: Document, Link Cited by: §3.
- On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, pp. 1321–1330. Cited by: §3.
- LLM-rubric: a multidimensional, calibrated approach to automated evaluation of natural language texts. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 13806–13834. External Links: Link, Document Cited by: §3.
- Can LLMs be good graph judge for knowledge graph construction?. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 10929–10948. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1.
- LightGBM: a highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30, pp. . External Links: Link Cited by: §3.
- Prometheus 2: an open source language model specialized in evaluating other language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 4334–4353. External Links: Link, Document Cited by: §1.
- CheckEval: a reliable LLM-as-a-judge framework for evaluating text generation using checklists. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 15771–15798. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: Table 2, Table 2, §1, item P2., §3, §5.
- Distribution-free predictive inference for regression. Journal of the American Statistical Association 113 (523), pp. 1094–1111. External Links: Document, Link, https://doi.org/10.1080/01621459.2017.1307116 Cited by: §7.
- DnA-eval: enhancing large language model evaluation through decomposition and aggregation. In Proceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Schockaert (Eds.), Abu Dhabi, UAE, pp. 2277–2290. External Links: Link Cited by: §3.
- G-eval: NLG evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 2511–2522. External Links: Link, Document Cited by: §A.1, §1, item P2..
- HD-eval: aligning large language model evaluators through hierarchical criteria decomposition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 7641–7660. External Links: Link, Document Cited by: §3, §3.
- Consistent individualized feature attribution for tree ensembles. External Links: 1802.03888, Link Cited by: 2nd item.
- A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, Red Hook, NY, USA, pp. 4768–4777. External Links: ISBN 9781510860964 Cited by: §3.
- Are self-explanations from large language models faithful?. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 295–337. External Links: Link, Document Cited by: §6.
- VELA: an LLM-hybrid-as-a-judge approach for evaluating long image captions. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 8680–8696. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1.
- FActScore: fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 12076–12100. External Links: Link, Document Cited by: §1, §3.
- Obtaining calibrated probabilities from boosting. In Proceedings of the Twenty-First Conference on Uncertainty in Artificial Intelligence, UAI’05, Arlington, Virginia, USA, pp. 413–420. External Links: ISBN 0974903914 Cited by: §3.
- Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. In Advances in Large Margin Classifiers, pp. 61–74. Cited by: §3.
- Gaussian processes for machine learning. The MIT Press. External Links: ISBN 9780262256834, Document, Link, https://direct.mit.edu/book-pdf/2514321/book_9780262256834.pdf Cited by: §7.
- A tutorial on conformal prediction. Journal of Machine Learning Research 9, pp. 371–421 (English). External Links: ISSN 1532-4435 Cited by: §3.
- Language models don't always say what they think: unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, pp. 74952–74965. External Links: Document, Link Cited by: §6.
- Beyond drift: stabilizing subjective LLM evaluation with information-theoretic rubrics. In Forty-third International Conference on Machine Learning, External Links: Link Cited by: §1, §3.
Appendix A Additional Experiment Details
A.1 Baseline Prompts
We use the G-Eval [Liu et al., 2023] prompts without the autogenerated chain-of-thoughts evaluation steps as the baseline prompts in our experiment.
A.2 Generated Checklists
| Dimension | # Sub-dimensions | # Seed Questions | # Final Questions |
|---|---|---|---|
| Coherence | 3 | 3 | 20 |
| Consistency | 3 | 3 | 16 |
| Fluency | 4 | 4 | 24 |
| Relevance | 5 | 5 | 21 |
A.3 GBDT Hyperparameters
The GBDT in §4 was learned using cross-validation of hyperparameters from the following search space:
"n_estimators": [20, 50, 100, 200]
"learning_rate": [0.001, 0.01, 0.1, 1.0]
"max_depth": [3, 5, 10]
"min_child_samples": [5, 10, 20, 40]
"reg_lambda": [0.0, 1.0, 5.0]