Toward Workflow-Aware Benchmarking for Healthcare NLP Agents
Abstract
Large language model (LLM) agents are increasingly proposed for healthcare tasks such as clinical documentation, evidence retrieval, patient messaging, and care coordination. Yet many evaluations remain limited to static medical question answering or one-shot generation, under-representing longitudinal state, interruptions, and human handoffs. We introduce an episode-level evaluation protocol for healthcare NLP agents. The protocol separates evidence across model, agent, and simulated-workflow behavior; specifies a five-field episode schema; and defines annotation and scoring for state continuity, evidence traceability, and escalation decisions. It is instantiated as four task templates: documentation update, evidence retrieval, patient messaging, and triage handoff. The protocol does not claim to measure clinical outcomes or deployment value. Instead, it supplies a reproducible intermediate evaluation layer between static benchmarks and prospective workflow studies, with an explicit cost-sensitive treatment of missed versus unnecessary escalation.
Index Terms:
healthcare NLP, LLM agents, benchmark design, clinical AI, evaluation, agent memoryI Introduction
Healthcare is one of the clearest domains in which the distinction between a fluent language model and a useful language agent becomes operationally important. Clinical and care workflows are multi-step, stateful, and high stakes. A system may need to summarize a chart, retrieve guideline evidence, maintain dialogue context, defer when information is missing, and support a human handoff when uncertainty increases. In such settings, strong one-turn performance is not sufficient evidence of practical usefulness.
This gap matters because the healthcare AI literature is moving quickly toward agentic framing. Recent work discusses planning, tool use, retrieval, memory, and multi-step interaction as central features of LLM systems in medicine and care [1, 2, 3, 4]. At the same time, several surveys and perspective papers warn that generic benchmark gains can be mistaken for clinical readiness [5, 6, 7, 8, 9]. The main failure is not always incorrect text in isolation. Often it is a failure to preserve state, retrieve the right evidence, surface uncertainty, or stay within workflow scope.
This short paper takes a narrow position: a useful healthcare-agent benchmark needs an operational layer between static prompt-response tests and prospective clinical workflow studies. This concern is consistent with broader agent-evaluation arguments that apparent system performance can depend heavily on hidden execution assumptions rather than architecture alone [10]. Static benchmarks remain useful, but they should be treated as one layer in a broader evidence stack.
Our contribution is threefold. First, we distinguish claims supported by evidence at the model, agent, and simulated-workflow levels. Second, we specify a reproducible episode protocol, including reference decisions, annotation units, and aggregation. Third, we provide four task templates and a cost-sensitive escalation score. The contribution is a design and measurement protocol, not a released dataset or clinical validation study.
II Why Healthcare Agent Evaluation Is Different
Healthcare work is sequential rather than atomic. Documentation, triage, patient follow-up, discharge communication, and evidence-assisted decision support all unfold over time. The user goal may change mid-task. New information may arrive after an initial output. A system may need to ask a clarifying question rather than answer immediately. These features make workflow quality at least as important as local answer quality.
This point is especially visible for agents. A standalone model can appear strong if it gives a fluent answer to a static question. A healthcare agent, by contrast, may fail even when its language is polished. It can omit a contraindication, carry forward stale chart information, or retrieve irrelevant evidence while sounding persuasive. In other words, healthcare failures are often state or process failures rather than pure generation failures. This is why recent work on agent memory, retrieval-grounded generation, and dynamic medical evaluation is directly relevant to healthcare AI assessment [11, 12, 13, 14, 15].
The practical implication is straightforward: evaluation should track whether an agent preserves relevant context, uses evidence appropriately, recovers from interruption, and escalates when scope boundaries are reached. Recent work on realistic static and live evaluation of agent routing further motivates treating routing as an explicit behavior rather than an incidental text output [16]. A benchmark that ignores these behaviors may be useful for capability screening while still being weak as evidence for deployment.
III Limits of Static Benchmarks
Static medical QA benchmarks remain useful for testing factual recall, explanation, multilingual coverage, and specialty breadth [17, 18, 19]. Hallucination benchmarks similarly help characterize unsupported claims and safety-sensitive failure modes [20, 21]. However, these tasks freeze interaction into a narrow slice. They rarely test whether the model should have deferred, queried for missing context, or tracked prior state across turns.
Clinical phenotype concept recognition is another useful capability target [22], but strong performance on a prompting-based extraction task likewise does not establish that an agent can manage state changes, evidence constraints, or escalation within a workflow.
This limitation becomes more serious once systems are framed as assistants embedded in care workflows. Recent benchmark directions begin to address the problem. MedAgentBench evaluates agentic tasks in EHR-like environments rather than only text prompts [23]. PhysicianBench and standardized-patient-style evaluations move further toward long-horizon clinical behavior [24, 25, 26]. EHR-grounded and consultation-centered benchmarks also improve realism by anchoring evaluation in records, physician chats, and evolving patient trajectories [27, 28, 29, 30]. Medical summarization and RAG-focused evaluations add another important axis by testing provenance retention, evidence selection, and grounded generation [31, 15, 32].
The field is therefore already moving in the right direction. The remaining challenge is conceptual integration. Different benchmarks capture different layers of useful behavior, but they are often reported side by side without a common evaluation logic. This can obscure what kind of claim a strong score actually supports and makes reusable benchmark design harder.
IV A Three-Level Evidence Protocol
We organize healthcare NLP agent assessment according to three evidence levels. The distinction is not a new claim that evaluation is layered; its added value is a reporting rule: every result must identify its evidence level and must not be promoted to a stronger deployment claim.
Table I summarizes the framework in a compact form. The table is intentionally practical: each level is tied to a deployment question, a typical metric family, and the main blind spot if that level is used alone.
| Level | Main question | Representative metrics | What it misses if used alone |
|---|---|---|---|
| Model | Does the system know relevant medicine and avoid obvious unsafe errors? | Factuality, calibration, robustness, harmful error rate, multilingual coverage | Longitudinal state, tool use, handoff quality, interruption recovery |
| Agent | Can the system execute a multi-step task with retrieval, memory, and bounded planning? | Retrieval fidelity, tool correctness, memory consistency, update behavior, abstention quality | Real workload effects, review burden, institutional constraints |
| Simulated workflow | Does the system route, hand off, and recover correctly in a designed episode? | Escalation utility, handoff completeness, state recovery, reviewable output burden | Clinical outcomes, time saved, institutional adoption |
IV-A Model Level
The first level measures core model behavior: factuality, calibration, fairness, robustness, and harmful error patterns. This is where static QA, explanation, and hallucination benchmarks remain appropriate. They answer questions such as whether the model knows relevant medicine, whether it makes unsupported claims, and whether it degrades under distributional or linguistic variation [5, 6, 20].
IV-B Agent Level
The second level measures agent behavior during multi-step task execution. Relevant constructs include planning quality, tool-use correctness, retrieval fidelity, memory consistency, update behavior, and uncertainty handling. This is where general LLM-agent and memory research becomes useful to healthcare because it provides a vocabulary for testing write, retrieve, update, forget, and reflect operations rather than only end responses [3, 4, 11, 12]. Healthcare-specific retrieval and EHR-like benchmarks fit naturally at this level [23, 14, 32].
IV-C Simulated Workflow Level
The third level evaluates routing, interruption recovery, handoff completeness, and the reviewability of outputs in a controlled episode. It is deliberately not a measure of time saved, clinician workload, patient safety, or care outcomes. Those claims require prospective workflow studies with local governance and human participants. A system may look strong at the model and agent levels but still fail a simulated episode if it creates an incomplete handoff, ignores a state update, or treats a routing decision as ordinary text generation [8, 9].
The protocol reports a vector, not a single leaderboard score: model competence, agent execution, and simulated-workflow performance. Strong results at the last level support only claims about the defined episodes. They do not establish clinical integration value.
V Task-Class Differences
Documentation support, evidence retrieval, patient messaging, and triage do not fail in the same way. A benchmark package that is adequate for one task may be poorly matched to another. Table II restores this task-sensitive view: a universal healthcare-agent leaderboard can obscure the changing balance between correctness, grounding, escalation, and continuity.
| Task class | Primary risk | Evaluation emphasis |
|---|---|---|
| Documentation support | Omission, distortion, propagation | Update fidelity, provenance, correction review |
| Evidence retrieval and decision support | Unsupported inference, misplaced confidence | Retrieval quality, attribution, abstention |
| Patient messaging | Over-reassurance, unsafe advice, scope drift | Escalation, safe communication, uncertainty marking |
| Triage and coordination | Missed escalation, dropped state, broken handoff | Priority, interruption recovery, handoff completeness |
For example, note transformation can be locally accurate yet still be weak for triage, where the dominant risk shifts from factuality to escalation and continuity. Likewise, evidence retrieval does not establish safety for patient-facing communication without testing boundaries and harmful-advice mitigation.
VI Episode Specification and Scoring
The evaluation unit is a short, artifact-grounded workflow episode rather than a single prompt. Each episode has five required fields: (i) an initial context packet, (ii) a state-changing event or interruption, (iii) an allowed action space, (iv) a reference decision profile, and (v) an adjudication rubric. The context uses de-identified or synthetic artifacts; the state change identifies facts that must be retained, revised, or discarded. The action space makes clear whether the agent may answer, retrieve, revise, ask a question, abstain, or escalate. This fixed schema makes episode construction auditable and prevents post-hoc scoring criteria.
Two qualified annotators independently produce reference decisions using an episode guide. They label required facts, prohibited claims, acceptable evidence sources, the appropriate action set, and handoff fields. A third reviewer adjudicates disagreements, and a benchmark release should retain both the final label and disagreement type. This supports reproducibility without claiming that a single reference response is uniquely correct.
| Episode | State-changing input | Reference decision | Primary score |
|---|---|---|---|
| Documentation update | Note draft plus late chart detail | Revise facts; preserve unaffected content | update fidelity |
| Evidence retrieval | Query plus evidence set and new constraint | Retrieve admissible sources; attribute claims | grounded-decision score |
| Patient messaging | Portal message plus risk cue | Reply within scope or escalate | escalation utility |
| Triage handoff | Longitudinal thread plus interruption | Preserve state; route with required fields | handoff completeness |
For every episode, scores are calculated from atomic annotations rather than global impressions. State continuity is scored as the fraction of required facts retained or correctly revised, minus contradictions and stale facts. Evidence traceability is scored as the proportion of substantive claims linked to an admissible source, with penalties for unsupported or mismatched citations. Handoff completeness is scored as the fraction of required routing fields that are present and correct. All score components and weights are reported by task class.
Escalation is evaluated as a decision, not as a generic safety preference. Let denote the reference escalation label and the agent action. We report sensitivity, specificity, and a configurable utility, , where when missed escalation is judged more harmful than unnecessary escalation. The cost matrix must be declared before evaluation and reviewed by the task’s clinical governance group. This makes the trade-off visible instead of silently rewarding either excessive deferral or unsafe completion.
VII Illustrative Failure Analysis
Consider a patient message reporting worsening shortness of breath after a recent medication change. A static benchmark may reward a correct explanation or summary, but it can miss the clinically salient failure: producing a polished reply instead of initiating escalation. The episode protocol instead asks whether warning signals are recognized, prior context is retained, and the handoff contains the required information. In this setting, a brief acknowledgment plus escalation can be preferable to a detailed answer.
VIII Design Requirements and Boundaries
The protocol operationalizes four requirements and sets a clear boundary on its claims.
VIII-A State Continuity
Episodes test whether relevant information is preserved, refreshed, and discarded across turns. The annotation guide distinguishes retained facts, revised facts, and facts that must no longer be used. Memory is therefore evaluated as a governance problem, not only as a context-window problem [11, 13, 33].
VIII-B Evidence Traceability
The rubric records admissible sources and requires claim-level links to them. This permits separate scoring of retrieval, attribution, and unsupported inference, rather than treating citations as a cosmetic feature [14, 15, 32]. This audit-oriented treatment is aligned with broader work on evidence review for LLM-generated policies, plan-guided retrieval with reranking, and execution provenance in LLM agents [34, 35, 36].
VIII-C Escalation Sensitivity
VIII-D Institutional Realism
Episodes can represent fragmented records, missing information, interruptions, and specialty-specific conventions. They do not reproduce institutional implementation or patient outcomes. Such claims require prospective evaluation; the present protocol is intended to make later studies more targeted.
IX Conclusion
Healthcare NLP agents should be evaluated as workflow-bound systems rather than as isolated response generators. Static medical benchmarks remain useful, but they do not justify claims about deployable assistance by themselves. We provide an episode-level protocol that separates evidence levels and makes state continuity, traceability, escalation, and handoff behavior reproducibly scorable.
The immediate next step is a small, clinically governed pilot that releases episode guides, adjudication decisions, score distributions, and failure analyses. Until then, this paper should be read as a testable measurement protocol rather than evidence that any agent improves healthcare work.
References
- [1] (2025) Large language models in healthcare. External Links: 2503.04748 Cited by: §I.
- [2] (2024) A survey on medical large language models: technology, application, trustworthiness, and future directions. External Links: 2406.03712 Cited by: §I.
- [3] (2025) Large language model agent: a survey on methodology, applications and challenges. External Links: 2503.21460 Cited by: §I, §IV-B.
- [4] (2025) Agentic large language models, a survey. External Links: 2503.23037 Cited by: §I, §IV-B.
- [5] (2024) Evaluating large language models in medical applications: a survey. External Links: 2405.07468 Cited by: §I, §IV-A.
- [6] (2025) A comprehensive survey on the trustworthiness of large language models in healthcare. External Links: 2502.15871 Cited by: §I, §IV-A, §VIII-C.
- [7] (2023) Creation and adoption of large language models in medicine. JAMA. Cited by: §I.
- [8] (2023) The shaky foundations of clinical foundation models: a survey of large language models and foundation models for electronic medical records. npj Digital Medicine. Cited by: §I, §IV-C.
- [9] (2024) Large language models in medicine: a review of current clinical trials across healthcare applications. PLOS Digital Health. Cited by: §I, §IV-C.
- [10] (2026) Beyond agent architecture: execution assumptions and reproducibility in llm-based trading systems. External Links: 2606.08285 Cited by: §I.
- [11] (2026) Memory for autonomous LLM agents: mechanisms, evaluation, and emerging frontiers. External Links: 2603.07670 Cited by: §II, §IV-B, §VIII-A.
- [12] (2026) Agentic memory: learning unified long-term and short-term memory management for large language model agents. External Links: 2601.01885 Cited by: §II, §IV-B.
- [13] (2025) Memory-r1: enhancing large language model agents to manage and utilize memories via reinforcement learning. External Links: 2508.19828 Cited by: §II, §VIII-A.
- [14] (2025) Retrieval-augmented generation in medicine: a scoping review of technical implementations, clinical applications, and ethical considerations. External Links: 2511.05901 Cited by: §II, §IV-B, §VIII-B.
- [15] (2025) Retrieval augmented generation evaluation for health documents. External Links: 2505.04680 Cited by: §II, §III, §VIII-B.
- [16] (2026) TwinRouterBench: fast static and live dynamic evaluation for realistic agentic LLM routing. External Links: 2605.18859 Cited by: §II.
- [17] (2023) MedBench: a large-scale chinese benchmark for evaluating medical large language models. External Links: 2312.12806 Cited by: §III.
- [18] (2024) MedExpQA: multilingual benchmarking of large language models for medical question answering. External Links: 2404.05590 Cited by: §III.
- [19] (2024) MedExQA: medical question answering benchmark with multiple explanations. External Links: 2406.06331 Cited by: §III.
- [20] (2024) MedHalu: hallucinations in responses to healthcare queries by large language models. External Links: 2409.19492 Cited by: §III, §IV-A, §VIII-C.
- [21] (2025) MedHallu: a comprehensive benchmark for detecting medical hallucinations in large language models. External Links: 2502.14302 Cited by: §III.
- [22] (2026) AutoPCR: automated phenotype concept recognition by prompting. Bioinformatics 42 (Supplement_1), pp. btag304. External Links: Document Cited by: §III.
- [23] (2025) MedAgentBench: a realistic virtual EHR environment to benchmark medical LLM agents. External Links: 2501.14654 Cited by: §III, §IV-B.
- [24] (2026) Evaluating large language models in dynamic clinical decision-making with standardized patient cases. External Links: 2606.05112 Cited by: §III.
- [25] (2026) PhysicianBench: evaluating LLM agents in real-world EHR environments. External Links: 2605.02240 Cited by: §III.
- [26] (2026) EHRBench: an automated and reliable EHR-based benchmark for clinical decision making with LLMs. External Links: 2605.30637 Cited by: §III.
- [27] (2024) EHRNoteQA: an LLM benchmark for real-world clinical practice using discharge summaries. External Links: 2402.16040 Cited by: §III.
- [28] (2025) BRIDGE: benchmarking large language models for understanding real-world clinical practice text. External Links: 2504.19467 Cited by: §III.
- [29] (2026) MedDialogRubrics: a comprehensive benchmark and evaluation framework for multi-turn medical consultations in large language models. External Links: 2601.03023 Cited by: §III.
- [30] (2026) HealthBench professional: evaluating large language models on real clinician chats. External Links: 2604.27470 Cited by: §III.
- [31] (2023) Adapted large language models can outperform medical experts in clinical text summarization. External Links: 2309.07430 Cited by: §III.
- [32] (2025) MedRAG: enhancing retrieval-augmented generation with knowledge graph-elicited reasoning for healthcare copilot. External Links: 2502.04413 Cited by: §III, §IV-B, §VIII-B.
- [33] (2026) The memory curse: how expanded recall erodes cooperative intent in LLM agents. arXiv preprint arXiv:2605.08060. Cited by: §VIII-A.
- [34] (2026) CARE: controlling LLM-generated policies through auditable review of evidence in scientific experimentation. External Links: 2606.14581 Cited by: §VIII-B.
- [35] (2026) GRASP: plan-guided graph retrieval with adaptive fusion and reranking on semi-structured knowledge bases. arXiv preprint arXiv:2605.30237. Cited by: §VIII-B.
- [36] (2026) From agent traces to trust: A survey of evidence tracing and execution provenance in LLM agents. CoRR abs/2606.04990. Cited by: §VIII-B.