Screen Before You Serve:
Simulation for Production Customer Experience AI Agents at 140M Scale
Abstract
Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization’s products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust.
We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank’s Card Delivery agent and its expanded successor, Card Management - Nubank’s highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.
Keywords:
Customer experience agents, large language models, user simulation, agent evaluation, open-weight models
1 Introduction
Recent advancements in large language models (LLMs), including complex reasoning (Guo et al., 2025; OpenAI, 2025), multi-step tool use (Qin et al., 2024; Schick et al., 2023), and long-context understanding (Gemini Team and others, 2024), have improved the real-world applicability of LLMs. By perceiving their environment and using tools to address users’ queries, LLM-based agents are being deployed for both personal and business applications (Luo et al., 2025; Gupta et al., 2026; Steinberger and Contributors, 2026; Research et al., 2026). While AI-assisted workflows have become easier to adopt in the era of agents, building reliable benchmarks for domain-specific, open-ended tasks remains challenging (Lynch et al., 2025; Zhu et al., 2026b).
This challenge is especially pronounced for customer experience (CX) agents, which are among the most widely adopted agentic applications across businesses, with deployments at companies such as Airbnb (Zhao et al., 2025), Amazon (Luo et al., 2025), and Nubank (Gupta et al., 2026). We define CX agents to include both reactive agents that address customer questions, complaints, and support needs, and proactive agents that anticipate customer needs and guide customers through conversational workflows to accomplish tasks using a company’s products and services. Although CX agents may not require the same depth of reasoning as coding or mathematical agents, they must meet complex, multidimensional requirements across diverse user groups. More specifically, CX agents should support personalization and localization (Kirk et al., 2024), respond appropriately to user intent across multi-turn conversations (Tack et al., 2026), and handle sensitive data (Vijayvargiya et al., 2026; Tien et al., 2026), among other requirements. Thus, the bars for safety and qualitative performance are higher than those for general agentic applications. Such unique expectations for CX agents raise a central question: how can we establish that an agent is operational, handles edge cases, and is ready for safe online deployment?
Figure 1 contrasts three approaches: (1) authoring test cases by hand, (2) exposing real customers in a live A/B test, and (3) simulating customers on-policy at the tool boundary. Although the first two options incorporate human feedback, they are slow and can provide inconsistent signals with small sample sizes. User simulation with synthetic personas has therefore been studied as an alternative for approximating the human preference distributions with LLMs (Ge et al., 2025; Naous et al., 2026; Li and others, 2026). However, limited controllability and the risk of score inflation raise concerns about discrepancies between simulated and real-world scenarios (Zhu et al., 2024; Mehri et al., 2026).
In this paper, we explore user simulation as an intermediate evaluation layer in which candidate agents are exercised through synthetic customer interactions before live deployment. Through case studies of the Card Delivery and Card Management agents at Nubank, we propose a hypothesis-driven, tool-boundary simulation recipe for screening candidate agents and demonstrate its real-world business impact. Simulation has become an essential part of our CX agent development lifecycle, enabling broad exploration of models, prompts, and reasoning settings that would be impractical to pursue through live experimentation alone. Our contributions are as follows:
- 1.
User simulation as screening-before-deployment: We propose a simulation-aided agent deployment recipe with a hierarchical workflow connecting offline evaluation to online deployment.
- 2.
Similarity analysis against real-world conversations: We demonstrate moderate conversation-level similarity between real-world and simulated conversations, with the lowest cosine distance of 0.035. Across four deployed versions, simulated and production version-level evaluator scores show high correlation.
- 3.
Efficient development-to-production lifecycle: The user simulation layer allows 4.8 times faster iteration of the development, screening, and deployment cycle compared with a workflow without a user simulation layer.
- 4.
Production impact across agent and model changes: The Card Management agent developed through simulation-guided iteration improved self-service rate (SSR) by 4.9 percentage points and transactional net promoter score (tNPS) by 36.69 points relative to Card Delivery in a live A/B test.
- 5.
Generalizability across models: We use the proposed simulation recipe to select Qwen3.5-122B-A10B as a replacement for the incumbent Card Management model. In a subsequent live A/B test, the replacement improved SSR by 8.82 percentage points and reduced p95 latency by 25%, with no statistically significant change in tNPS.
2 Related Work
Customer experience agents in production.
An increasing number of companies across different business categories are adopting agentic workflows for customer support, including Amazon (Luo et al., 2025), Airbnb (Zhao et al., 2025; Su et al., 2025), Kakao (Park et al., 2025), Alibaba Group (Jiang et al., 2025), Thomson Reuters (Juclà et al., 2026), and Nubank (Gupta et al., 2026). Zhao et al. (2025) and Su et al. (2025) propose human-in-the-loop feedback collection strategies and synthetic data generation pipelines for sustainably aligning customer experience agents with human preferences at Airbnb. Gupta et al. (2026) present a hierarchical pipeline spanning offline validation and online deployment for customer experience agents at Nubank, improving agent self-service rate (SSR) and transactional net promoter score (tNPS) (Reichheld, 2003).
| Card Delivery (CD) | Card Management (CM) | ||
|---|---|---|---|
| Version | Configuration | Version | Configuration |
| GPT-4.1; tools for customer-specific logistics queries. | GPT-5.2; CD prompts and tools with ReAct-style tool definitions (Yao et al., 2023); baseline. | ||
| GPT-4.1; card selection and prevention of recent duplicate reissues. | GPT-5.2; Manual revision of ’s prompt and tool-use policy. | ||
| GPT-5.1; prompt rules for routing, personalization, and concise responses. | GPT-5.2; ’s prompt and tools with a different architecture. | ||
| GPT-5.2; tools to retrieve and update the registered delivery address. | GPT-5.2; Manual revision of ’s prompt and tool-use policy. | ||
| GPT-5.2; Return to ReAct; prompt revised using simulated examples. | |||
User simulation in benchmarking agents.
To benchmark agents across a wide range of domains, synthetic personas have been studied as a method for simulating diverse human demands, either at scale through verbalized prompts (Ge et al., 2025; Li and others, 2026) or by fine-tuning the language model used for simulation (Naous et al., 2026). Recent agent benchmarks, in particular, leverage prompt-guided user personas to simulate real-world scenarios (Yao et al., 2025; Tan et al., 2025; Qian et al., 2025; Barres et al., 2025; Huang et al., 2026). Specifically, -Bench defines multi-turn scenarios in which agents interact with prompt-driven user simulators in retail and airline customer experience (CX) domains (Yao et al., 2025), while -Bench extends this setting so that both agents and users can invoke tools, enabling more human-like actions by simulated users (Barres et al., 2025). Building on increasing efforts to expose CX agents to more realistic environments, we address the complementary question of how to characterize and use a simulator as a restricted screening layer in agent development for real-world deployment.
Reliability in user simulation.
Despite the increasing use of user simulation in agent benchmarking, a simulator of unknown fidelity can actively mislead (Zhu et al., 2024; Mehri et al., 2026). Prompt-driven user simulators are often prone to score inflation through data leakage (Zhu et al., 2024) or may fail to adhere to their assigned goals in multi-turn scenarios (Mehri et al., 2026). Accordingly, rigorously validating the robustness of user simulation pipelines and their alignment with real human behavior remains challenging (Dou et al., 2025; Zhu et al., 2026a). SimulatorArena explicitly measures agreement between human judgments and assistant ratings on public tasks such as math tutoring (Dou et al., 2025), while RealUserSim grounds simulator personas in behavioral profiles extracted from real conversations and tests fidelity using an LLM-judged paired-trajectory Turing test (Zhu et al., 2026a). We characterize the reliability and robustness of a user simulation recipe by comparing simulated and real-world conversations, and demonstrate how simulation-based findings can inform changes that are subsequently evaluated through real-world business outcomes at Nubank. Most importantly, our results demonstrate that imperfect simulation can still guide directionally correct improvements to production agents, with benefits confirmed with live A/B evaluation.
3 Simulation Setup
This section consolidates the experimental setup for multi-turn agent trajectory generation and evaluation (Figure 2). Our pipeline selects a target agentic task, simulates on-policy conversations with synthetic personas and tool responses, and evaluates the resulting trajectories offline.
Target agents.
We study Nubank’s Card Delivery (CD) agent and its expanded successor, Card Management (CM), which replaced CD and supports a broader range of card-lifecycle requests. We characterize the simulation pipeline using CD traces and apply the resulting workflow to CM.
3.1 Simulation pipelines
Proposed simulation pipeline.
Our proposed trajectory-generation pipeline uses Snowglobe as its primary simulator (Guardrails AI, 2025). Snowglobe instantiates the persona-driven simulation pattern of Park et al. (2023) and takes an agent description and tool definitions, optionally augmented with a simulation prompt and historical conversations. From these inputs, it constructs use cases, user profiles, tool relationships, and trajectory plans (Appendix A). An orchestration agent allocates conversations across use-case and interaction-style conditions, after which the generated personas interact with the target agent over multiple turns. Each synthetic user turn is conditioned on the agent’s preceding response, while tool calls are intercepted and answered with scenario-conditioned synthetic results rather than invoking production backends. A conversation ends when the persona’s objective is achieved, judged unreachable, or subject to a turn or tool-specification limit. We adopt Snowglobe for its structured orchestration, controllable scenario coverage, and cross-call tool-state consistency support for on-policy, hypothesis-driven agent screening.
Naive LLM baseline.
To assess what can be achieved through direct prompting without Snowglobe’s structured orchestration, we implement a naive LLM-based simulator. A persona language model generates informal Brazilian Portuguese customer turns, while a second language model produces synthetic results for the agent’s tool calls. Both roles use gpt-5.6-sol with high reasoning effort and receive the same categories of non-test-set context as Snowglobe: the agent description, tool schemas, product-requirements meta-knowledge, and tool-call examples. Both simulators also use the same seeded scenario distribution. A baseline conversation ends when the persona signals completion, the agent transfers the conversation to a human, or the customer has sent ten messages. Unlike Snowglobe, the baseline does not orchestrate use-case and interaction-style coverage, generate structured trajectory plans, or use an agent-profile model to maintain cross-call consistency. Appendix G provides the baseline prompts.
3.2 Card Delivery Agent
Agent versions and data.
8,000 real-world conversations between the users and the CD agent were evenly split into two sets: a development set to configure the simulators (i.e., simulation characterization) and a test set to validate the agent with the characterized simulation profiles. The four deployed CD versions are summarized in Table 1 (left).
The test set contains 1,000 conversations from each agent version. We generated 250 synthetic trajectories via the simulator for each version.
Offline evaluation.
We use five semantic categories to evaluate agent trajectories offline. They are evaluated via LLM-as-a-Judge, using GPT-4.1-Mini11 1 https://openai.com/index/gpt-4-1/: (E1) Card reissue failure; (E2) Customer input verification; (E3) Card delivery data check; (E4) Response conciseness; (E5) Resolution conciseness following Gupta et al. (2026). We evaluate each category with separate system prompts optimized via GEPA (Agrawal et al., 2026), which return binary pass/fail scores.
3.3 Card Management Agent
Agent versions and data.
Offline evaluation.
We used a single binary judgment from LLM-as-a-Judge to test the explicit hypothesis in each version as described above. Given the CM agent’s execution traces and available tools, LLM-as-a-Judge provides binary feedback on whether a transfer to human support was unnecessary. We compare each candidate’s failure rate with that of to decide whether to accept the change.
Online evaluation.
We assess production performance using two online metrics: transactional net promoter score (tNPS) and self-service rate (SSR). tNPS is the percentage of promoters minus the percentage of detractors in a post-interaction recommendation survey (Reichheld, 2003). tNPS is thus a measure of customer satisfaction and happiness. Self-service rate (SSR) is the proportion of sessions in which users completed the case with the agent without asking for human support.
4 User Simulation Recipe for CX Agents
Algorithm 1 uses tool-boundary trajectory simulation for hypothesis-driven screening (Balog and Zhai, 2024). Its purpose is to expose beneficial or harmful candidate–incumbent differences large enough to affect a deployment decision. Therefore, exact replication of production metrics is unnecessary.
The algorithm first generates on-policy incumbent trajectories using simulator configuration , which specifies the target use-case mix. These trajectories remain the shared baseline throughout the screening round. Configuration uses the agent description and tool definitions, optionally supplemented by a simulation prompt and historical data. Record scenario composition and keep distribution runs separate from targeted probes. Scenario-conditioned tool outputs must match declared schemas and remain consistent across calls within a conversation. For example, a record queried twice should not change state without an intervening action. Simulation exercises dialogue–tool interactions without testing backend implementations (Appendix B).
Each screening round is guided by a falsifiable hypothesis and a prespecified screening criterion . The hypothesis determines the behavioral effect of interest, while defines how evaluator results support eligibility for live testing. The evaluation suite comprises evaluators relevant to the task and the hypothesis, including LLM-as-a-judge evaluators when appropriate. It covers selected standing criteria and the proposed behavioral effect. The same suite is applied to incumbent and candidate trajectories; any newly introduced evaluator must also score the baseline trajectories. The standing CD suite was calibrated against human annotations (Gupta et al., 2026); calibration of each new hypothesis-specific judge must be reported separately.
Candidate simulations hold fixed. Candidates meeting become eligible for live A/B testing; others are reassessed or revised against the same incumbent. The baseline is updated only after a candidate becomes the incumbent. For candidates meeting , smaller or ambiguous differences are left for live experiments to resolve. This loop complements item-level evaluation and live A/B tests; Figure 3 summarizes it alongside the separate simulator-characterization diagnostics.
4.1 Simulator characterization
We introduce four different diagnostics for the characterized simulators, including three automated evaluations and one human annotation session, namely P1 to P4. For three synthetic evaluations, we use the data and offline evaluation pipeline introduced in Section 3.2, i.e., each uses 1,000 held-out production conversations with real-world users and 250 Snowglobe traces per CD version. The diagnostics address complementary properties (Figure 3b), none as a standalone criterion.
P1: Conversation-length statistics.
We compare the number of words in user messages and the number of user turns per conversation. We report cumulative threshold percentages with full distributions in Appendix C.
P2: Embedding diagnostics.
We embed each conversation using text-embedding-3-large22 2 https://developers.openai.com/api/docs/models/text-embedding-3-large with user: and agent: prefixes. Then, we measure cosine and Euclidean distances between the group centroids, and plot the two-dimensional Uniform Manifold Approximation and Projection (McInnes et al., 2018, UMAP) visualization.
P3: Evaluator-score association.
We apply five canonical evaluation categories, E1 to E5, in Section 3.2 to the production, Snowglobe, and baseline pools. For each agent version, we report the average scores per category and their 95% confidence intervals. Across the versions, we report average rank, Pearson , and Kendall relative to the production ordering.
P4: Blinded human source discrimination.
We sampled conversations, comprising 50 production and 50 Snowglobe trajectories. was selected because it was the most recent version and the version most familiar to the annotators. Seven domain experts contributed labels. After the same source-blind normalization was applied to both groups, annotators saw dialogue text without tool-call sequences, classified each item as production or simulated, and reported confidence on a 1–5 scale.
4.2 Open-weight model screening
Models.
We test 29 configurations of the card management agent with over 16,000 simulated conversations, across model family, reasoning effort, and numerical format, including GPT-OSS-120B (OpenAI and others, 2025), Nemotron-3-Super-120B-A12B (NVIDIA and others, 2026), Qwen3.5-122B-A10B (Qwen Team, 2026a), and Qwen3.6-35B-A3B (Qwen Team, 2026b).
Evaluation.
Building on the offline evaluation approach in Section 3.2, we use LLM-as-a-Judge to evaluate the following categories: (1) gathering proper inputs, (2) validity of card reissue, (3) information retrieval status, and (4) completeness of conversation, which are noted as “Input”, “Reissue”, “Status”, and “Complete” in Table 3.
5 Results
We characterize the simulator on CD versions, then report the offline CM analysis and the live A/B comparison, followed by a predeployment screen of open-weight models.
5.1 Simulator characterization
P1: Conversation length.
Appendix C reports both selected cumulative thresholds for the four CD versions (Table 4) and the complete distributions (Figure 9). In the pooled samples, of simulated conversations contain at most 50 user-message words, compared with in production; contain at least 200 words, compared with in production. Similarly, of simulated conversations have at most two user turns, compared with in production. The shares with at least six and eight turns are and in simulation, versus and in production.
P2: Transcript-embedding proximity.
For each CD version, the cosine distance between the production and simulated transcript centroids is smaller than the distance from either centroid to the off-topic control centroid (Figure 4). The same ordering holds for Euclidean distance (Appendix D, Figure 10). Compared with the baseline, Snowglobe is closer to production in two versions, tied in one at the reported precision, and farther in one by cosine distance; by Euclidean distance, it is closer in three versions and farther in one. Both simulated groups remain closer to production than to the off-topic control. Figure 6 provides the two-dimensional UMAP projections.
P3: Evaluator-score association.
Figure 5 shows that Snowglobe preserves the production extremes, best and worst, although and exchange positions. The baseline instead ranks best and worst. is best in of Snowglobe resamples and of baseline resamples; is worst in and , respectively. Table 5 in Appendix E reports the aggregate scores and ranking-agreement metrics. The average-rank statistic is 1.38 for Snowglobe and 1.62 for the baseline. The Snowglobe and production scores have Pearson and Kendall ; their largest absolute difference is for (0.556 versus 0.437).
P4: Blinded human source discrimination.
Annotators correctly identified 42 of 50 production conversations () and 35 of 50 simulated conversations (). Thus, 15 simulations were classified as production. Reported confidence was approximately 3 on the 1–5 scale, including for correctly classified items.
5.2 Card Management: offline and live results
Applying the simulation recipe, had the lowest unnecessary-transfer failure rate (), compared with for , for , and for and . This represents a 6.4-percentage-point reduction relative to . In the subsequent A/B test, the CM arm had SSR percentage points above CD (95% CI ) and tNPS points above CD (95% CI ) (Table 2). Following the test, the agent became available to Nubank’s customer base in Brazil. A later test replaced the incumbent model in CM with the open-weight candidate screened in Section 5.3, Qwen3.5-122B-A10B with reasoning enabled. SSR rose by percentage points (95% CI , ) while tNPS was statistically unchanged ( points, 95% CI , ), and p95 latency fell by . The screened open-weight configuration therefore matched the incumbent on satisfaction while improving SSR and latency.
A batch of 100 simulated trajectories completed in under 10 minutes on average; this measures generation runtime rather than end-to-end iteration time. The ten-version CD development cycle, including the four versions studied here, spanned 212 calendar days. The five CM versions spanned 22 days. This corresponds to and days per version, respectively, a ratio of .
| Outcome | Difference | 95% CI | |
|---|---|---|---|
| Card Management vs. Card Delivery | |||
| SSR (pp) | |||
| tNPS (pp) | |||
| Qwen3.5-122B-A10B (reasoning on) vs. incumbent | |||
| SSR (pp) | |||
| tNPS (points) | |||
| p95 latency (%) | — | — | |
5.3 Open-weight model screening
Table 3 reports four evaluator failure rates for the open-weight configurations. Performance by criterion, with no configuration uniformly outperforming the others. Qwen3.5 achieves the lowest input gathering (“Input”) failure rate with reasoning enabled (0.4%) and the lowest conversation completeness (“Complete”) failure rate with reasoning disabled (16.%). Qwen3.6 with reasoning enabled leads on card reissue validity (“Reissue”) (14.0%), while GPT-OSS with medium reasoning has the lowest information retrieval status (“Status”) failure rate among open-weight configurations (17.6%). The incumbent retains a substantial advantage on Status (2.4%).
The effect of reasoning also varies by model and criterion. For Qwen3.6, enabling reasoning lowers all four reported mean failure rates. For Nemotron, moving from no to low reasoning reduces Complete failures from 60.8% to 18.0%, but increases Status failures from 21.2% to 36.4%. For Qwen3.5, reasoning reduces Input and Status failures while increasing Reissue and Complete failures. These results imply model-specific benefits from additional reasoning.
We selected Qwen3.5-122B-A10B with reasoning enabled as a starting point for prompt optimization based on its low Input failure rate and competitive Reissue and Complete results. Relative to the incumbent, its Input and Complete failure rates are lower by 3.2 and 15.0 percentage points, respectively, while Reissue is similar (25.4% versus 25.2%). The higher Status failure rate (35.2% versus 2.4%) identified a weakness for further refinement. Subsequent prompt optimization reduced failures on this criterion. The offline results therefore supported Qwen3.5-122B-A10B as a candidate for live evaluation, without establishing it as the strongest configuration across all criteria. Table 2 reports the subsequent live A/B outcomes.
| Configuration | Reasoning | Canonical Evaluations () | |||
|---|---|---|---|---|---|
| Input | Reissue | Status | Complete | ||
| GPT-5.2 (Incumbent) | — | ||||
| Qwen3.6-35B-A3B | Off | ||||
| On | |||||
| Qwen3.5-122B-A10B | Off | ||||
| On | |||||
| Nemotron-3-Super- 120B-A12B | Off | ||||
| Low | |||||
| High | |||||
| GPT-OSS-120B | Low | ||||
| Medium | |||||
| High | |||||
5.4 Deployment lessons and limitations
Simulation supported safer, faster iteration: batches of 100 trajectories ran in under ten minutes, evaluators surfaced behavioral regressions, and the screen caught tool-call failures and a serving configuration without a tool-call parser. Most engineering effort went into integration: the staging agent reused live MCP schemas while stateful tools returned schema-compatible synthetic responses. The agent retained tool control and sequencing, but schemas, examples, authentication, and simulator profiles required ongoing alignment.
The approach deliberately stops at the tool boundary. Stateful or side-effecting tools are mocked, while read-only dependencies such as knowledge-base retrieval may remain live; backend behavior, latency, persistent state, and side effects thus remain untested. Simulation traces must also match production schemas, identifiers, and telemetry, or export and reconciliation work can erase the iteration gains.
6 Conclusion
In this paper, we present a hypothesis-driven approach to using simulation as a pre-deployment screening layer for CX agents. At Nubank, simulation has become an integral part of our development life cycle, enabling teams to explore changes, identify behavioral failures, and refine candidates before customer exposure. Its value lies in supporting development decisions, not in perfectly reproducing production. By combining simulation with live evaluation, this workflow supports broader experimentation while keeping deployment decisions grounded in observed customer outcomes.
References
- GEPA: reflective prompt evolution can outperform reinforcement learning. In International Conference on Learning Representations (ICLR), Note: Oral External Links: Link Cited by: §3.2.
- User simulation for evaluating information access systems. Foundations and Trends in Information Retrieval 18 (1–2), pp. 1–261. External Links: Document, Link Cited by: §4.
- -Bench: evaluating conversational agents in a dual-control environment. arXiv preprint arXiv:2506.07982. External Links: Link Cited by: §2.
- SimulatorArena: are user simulators reliable proxies for multi-turn evaluation of AI assistants?. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 35212–35290. External Links: Document, Link Cited by: §2.
- Scaling synthetic data creation with 1,000,000,000 personas. External Links: 2406.20094, Link Cited by: §1, §2.
- Gemini 1.5: unlocking multimodal understanding across millions of tokens of context. External Links: 2403.05530, Link Cited by: §1.
- SnowGlobe: the simulation engine for AI agents and chatbots. Note: Product whitepaperAccessed 2026-05-21 External Links: Link Cited by: §3.1.
- DeepSeek-r1 incentivizes reasoning in llms through reinforcement learning. Nature 645 (8081), pp. 633–638. Cited by: §1.
- Building customer support agents at 100M-user scale: an evaluation-driven framework. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Industrial Track), Note: To appear Cited by: §1, §1, §2, §3.2, §4.
- CRMArena-pro: holistic assessment of LLM agents across diverse business scenarios and interactions. Transactions on Machine Learning Research. Note: External Links: ISSN 2835-8856, Link Cited by: §2.
- ChatMap: mining human thought processes for customer service chatbots via multi-agent collaboration. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 11927–11947. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §2.
- Retrieval enhancements for RAG: insights from a deployed customer support chatbot. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track), Y. Matusevych, G. Eryiğit, and N. Aletras (Eds.), Rabat, Morocco, pp. 169–180. External Links: Link, Document, ISBN 979-8-89176-384-5 Cited by: §2.
- The PRISM alignment dataset: what participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §1.
- MatrAIx: simulating the world with 8.3 billion persona agents. External Links: 2608.04205, Link Cited by: §1, §2.
- Language model alignment for conversational shopping at amazon. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’25, New York, NY, USA, pp. 4314–4318. External Links: ISBN 9798400715921, Link, Document Cited by: §1, §1, §2.
- Agentic misalignment: how llms could be an insider threat. Anthropic Research. Note: https://www.anthropic.com/research/agentic-misalignment Cited by: §1.
- UMAP: uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426. Cited by: §4.1.
- Goal alignment in LLM-based user simulators for conversational AI. Transactions of the Association for Computational Linguistics. External Links: Document, Link Cited by: §1, §2.
- Flipping the dialogue: training and evaluating user language models. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §2.
- Nemotron 3 super: open, efficient mixture-of-experts hybrid mamba-transformer model for agentic reasoning. External Links: 2604.12374, Link Cited by: §4.2.
- Gpt-oss-120b & gpt-oss-20b model card. External Links: 2508.10925, Link Cited by: §4.2.
- OpenAI o3 and o4-mini system card. Technical report OpenAI. External Links: Link Cited by: §1.
- A practical approach for building production-grade conversational agents with workflow graphs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), G. Rehm and Y. Li (Eds.), Vienna, Austria, pp. 1508–1519. External Links: Link, Document, ISBN 979-8-89176-288-6 Cited by: §2.
- Generative agents: interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), Note: Best Paper Award External Links: Document, Link Cited by: §3.1.
- UserBench: an interactive gym environment for user-centric agents. In Workshop on Scaling Environments for Agents, External Links: Link Cited by: §2.
- ToolLLM: facilitating large language models to master 16000+ real-world APIs. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1.
- Qwen3.5: towards native multimodal agents. External Links: Link Cited by: §4.2.
- Qwen3.6-35B-A3B: agentic coding power, now open to all. External Links: Link Cited by: §4.2.
- The one number you need to grow. Harvard Business Review 81 (12), pp. 46–54. Cited by: §2, §3.3.
- Composer 2 technical report. External Links: 2603.24477, Link Cited by: §1.
- Toolformer: language models can teach themselves to use tools. In Thirty-seventh Conference on Neural Information Processing Systems, External Links: Link Cited by: §1.
- OpenClaw: your own personal ai assistant. GitHub. Note: https://github.com/openclaw/openclawAccessed: 2026-09-03 Cited by: §1.
- LLM-friendly knowledge representation for customer support. In Proceedings of the 31st International Conference on Computational Linguistics: Industry Track, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, S. Schockaert, K. Darwish, and A. Agarwal (Eds.), Abu Dhabi, UAE, pp. 496–504. External Links: Link Cited by: §2.
- LLMs get lost in evolving user intent. External Links: 2607.20734, Link Cited by: §1.
- PersonaBench: evaluating AI models on understanding personal information through accessing (synthetic) private user data. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 878–893. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §2.
- ROGUE: misaligned agent behavior arising from ordinary computer use. External Links: 2606.00341, Link Cited by: §1.
- OpenAgentSafety: a comprehensive framework for evaluating real-world AI agent safety. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1.
- -bench: a benchmark for tool-agent-user interaction in real-world domains. In The Thirteenth International Conference on Learning Representations (ICLR), External Links: Link Cited by: §2.
- ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), Cited by: Table 1.
- Agent-in-the-loop: a data flywheel for continuous improvement in LLM-based customer support. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, S. Potdar, L. Rojas-Barahona, and S. Montella (Eds.), Suzhou (China), pp. 1919–1930. External Links: Link, Document, ISBN 979-8-89176-333-3 Cited by: §1, §2.
- How reliable is your simulator? analysis on the limitations of current LLM-based user simulators for conversational recommendation. In Companion Proceedings of the ACM Web Conference 2024 (WWW ’24), pp. 1726–1732. External Links: Document, Link Cited by: §1, §2.
- RealUserSim: bridging the reality gap in agent benchmarking via grounded user simulation. arXiv preprint arXiv:2605.20204. External Links: Link Cited by: §2.
- Establishing best practices in building rigorous agentic benchmarks. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §1.
Appendix A How Snowglobe infers simulation properties
Before it can interact with an agent in a meaningful way, Snowglobe needs an approximation of that agent’s objectives, its users, and its behavior. The minimal input is a description of the agent together with the definitions of its tools. From these, a set of agents infers the properties a simulation depends on (Figure 7): the use cases the agent serves, profiles of synthetic users along with the data those users would carry, how the tools relate to one another and to the surrounding system, and the user trajectories that would call on each tool. Some of these properties are resolved once during agent onboarding and others are resolved when a simulation begins.
Two optional inputs reorient these properties. A simulation prompt narrows a run to a chosen set of use cases or behaviors while leaving the rest of the inferred picture intact. Real historical conversations, when supplied, are processed offline to extract use cases, linguistic styles, and their distributions, which are stored and used to condition later runs.
Appendix B Tool boundary mocking
Exercising the agent’s tools inside a simulation raises a problem of its own. The simulated environment must not access production data through stateful tools, yet it has to return data through those tools so the agent can run end to end, and that data must stay consistent with the state of the conversation and with the results of earlier tool calls. Testing the tools themselves is a separate concern, better served by other methods, and is not an aim here.
Snowglobe meets these constraints without ever maintaining a database of state (Figure 8). During onboarding, it absorbs the tool definitions and builds a functional model of how the tools relate and how data must be shaped to satisfy each request. This model is called an agent profile. Using an agent profile, an orchestrator produces persona seeds. These are partial seed states paired with a trajectory for how that data should evolve over a conversation. Pesrona seeds are projected onto personas at runtime to ground their objectives.
At simulation time, a call to one of the agent’s tools is routed to the persona agent handling that conversation rather than to the real implementation. Using its partial state and the consistency rules in the agent profile, the persona agent returns a response that fits the tool’s output shape, and that response is passed back to the agent under test. Because no shared world state is instantiated, each conversation runs in its own sandbox. Snowglobe refers to this as tool-boundary mocking: only the individual tool calls at the boundary are answered, never an entire backend or database.
Appendix C Conversation length distributions
Figure 9 shows the full P1 distributions of user-message words and user turns, complementing the thresholds in Table 4. Length is user-side only. Groups are production traffic pooled over – (), Snowglobe simulations of the same variants (), and an off-topic control (). The retained 226-item control is encoded in the plotted artefact and historical experiment code, but the filtering from a documented upstream 1,000-item random control sample is not recorded. Table 4 shows summarized results in bins for easy comparison of extreme values.
Production user text is short (median 19 words; mean 25.8). The control matches that scale (median 20; mean 27.5); simulated users do not (median 111; mean 119). Control shares at the Table 4 word cuts are 85.8% () and 0% (). Because the control is production traffic rather than LLM-generated user text, it cannot determine whether the heavy simulated tail is specific to this configuration or common to LLM user simulators.
In production, of conversations have a single user turn, and the median is three turns. Control never has a one-turn chat and peaks at three turns (33.2%). Simulated conversations last longer (median 6) and hit a hard cap: 25.2% have exactly eight user turns, and none have more—which is why every simulated conversation has fewer than 10 user turns, the 95th-percentile value in production, despite the higher simulated median. Control shares at the table cuts are 18.6% (), 17.7% (), and 4.0% ().
| User-message words per conversation | User turns per conversation | ||||
|---|---|---|---|---|---|
| Criterion | Production | Simulated | Criterion | Production | Simulated |
| 89.0 | 22.4 | 36.8 | 16.4 | ||
| words | 0.1 | 14.9 | 6 | 24.0 | 55.0 |
| 8 | 10.6 | 25.2 | |||
Appendix D Euclidean transcript-centroid distances
Appendix E Evaluator-score association details
Table 5 reports the P3 aggregate scores and ranking-association statistics. Each conversation contributes five binary failure outcomes, one per evaluator. Within each source–version pool, we first average each evaluator across conversations and then average the five evaluator means; lower aggregate scores indicate fewer failures.
The 95% intervals use conversation-clustered bootstrap replicates. Each replicate resamples conversation rows with replacement, preserving the five outcomes for each sampled conversation, and recomputes the aggregate score. Production defines the reference ordering. Average rank summarizes closeness to that order (lower is better), Pearson measures linear association between version-level scores, and Kendall measures rank agreement.
The extreme-rank analysis uses the same resampling procedure. Each replicate ranks the four versions by aggregate score, with ties sharing credit equally. The reported frequencies estimate only how often is best and is worst.
| Source | Avg. rank | Pearson | Kendall | ||||
|---|---|---|---|---|---|---|---|
| Baseline | |||||||
| Simulated | 1.38 | 0.74 | 0.67 | ||||
| Production |
Appendix F Open-weight screening: reasoning and quantization
Because a simulated arm costs no customer exposure, the screen can afford to sweep configuration settings that would otherwise be argued from intuition. This appendix reports the sweep summarised in Section 5.3, run on Nemotron-3-Super-120B-A12B across three quantizations and three reasoning settings. The figure omits the incumbent: the composite is a mean over evaluators built for this agent, and is used to compare a model against itself under different settings rather than to rank candidates against production.
Figure 11 shows the result. Reasoning effort moves the composite monotonically and by a large margin, from with reasoning off to with it on at nvfp4, and the mechanism is retrieval: the rate at which the model fetches the company record the task requires rises from to over the same range. Quantization does not move it. Across nvfp4, fp8, and bf16 the spread is , , and points at reasoning off, low, and on, against a -point standard deviation measured over six replicate runs of a fixed configuration. At this sample size the three quantizations are indistinguishable, while the reasoning setting moves the composite by roughly points.
Appendix G Baseline simulator prompts
Both simulator models (user and tools) are gpt-5.6-sol with reasoning_effort=high. Brace placeholders mark run-specific context concatenated at inference time: agent description, tool JSON schemas, product-requirements text, and tool-call examples.
G.1 Persona model
The persona is given one function tool, end_conversation, so it can trigger the end of the simulation. The first user message is shown below. Later turns reuse the same system prompt and append the dialogue (user and assistant text only).
Customer cards are sampled from a seeded generator.
G.2 Tool-result model
The mock backend is a second completion (no tools). The user message is the pending tool call.