arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.30137v1 [cs.AI] 24 Sep 2026

Screen Before You Serve:
Simulation for Production Customer Experience AI Agents at 140M Scale

Edesio Alcobaça Affiliation: Nubank    Kevin Rossell Affiliation: Nubank    Aman Gupta Affiliation: Nubank Correspondence to: aman.gupta@nubank.com.br    Shao Tang Affiliation: Nubank Correspondence to: tang.shao@nubank.com.br    Jiwoo Hong Affiliation: Nubank    Pabel Carrillo-Mendoza Affiliation: Nubank    Wanderson Conceição Ferreira Affiliation: Nubank    Alvaro Tedeschi Affiliation: Nubank    Zayd Simjee Affiliation: Guardrails AI    Shreya Rajpal Affiliation: Guardrails AI    Bruno Finardi Hime Affiliation: Nubank    Christian Sousa Affiliation: Nubank    Luis Moneda Affiliation: Nubank    Herbert Fei Affiliation: Nubank    Daniel Silva Affiliation: Nubank    Rohan Ramanath Affiliation: Nubank
Abstract

Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization’s products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust.

We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank’s Card Delivery agent and its expanded successor, Card Management - Nubank’s highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.

Keywords: 
Customer experience agents, large language models, user simulation, agent evaluation, open-weight models
††affiliationnotice: *Joint first authors.
Refer to caption
Figure 1: Three ways to decide whether an agent revision is safe to ship. Manual authoring (A) exercises the scenarios someone thought to write down; some tool states that real customers encounter may be missing. A live A/B test (B) measures outcomes directly in production but exposes customers to the agent changes being tested. On-policy, tool-boundary simulation (C) answers the tool call synthetically and lets the agent act on its own policy, so multi-turn failures can surface before customers are exposed. In production (D) the screened agent moved self-service rate and transactional NPS against a matched comparison.

1 Introduction

Figure 2: End-to-end trajectory generation with Snowglobe: an orchestrator reads the inputs and plans coverage over use-case and style slices; each slice generates personas carrying style, state, topics, and trajectory; each persona holds one or more multi-turn conversations with the agent under test, whose tool calls are answered by a mock tool layer.

Recent advancements in large language models (LLMs), including complex reasoning (Guo et al., 2025; OpenAI, 2025), multi-step tool use (Qin et al., 2024; Schick et al., 2023), and long-context understanding (Gemini Team and others, 2024), have improved the real-world applicability of LLMs. By perceiving their environment and using tools to address users’ queries, LLM-based agents are being deployed for both personal and business applications (Luo et al., 2025; Gupta et al., 2026; Steinberger and Contributors, 2026; Research et al., 2026). While AI-assisted workflows have become easier to adopt in the era of agents, building reliable benchmarks for domain-specific, open-ended tasks remains challenging (Lynch et al., 2025; Zhu et al., 2026b).

This challenge is especially pronounced for customer experience (CX) agents, which are among the most widely adopted agentic applications across businesses, with deployments at companies such as Airbnb (Zhao et al., 2025), Amazon (Luo et al., 2025), and Nubank (Gupta et al., 2026). We define CX agents to include both reactive agents that address customer questions, complaints, and support needs, and proactive agents that anticipate customer needs and guide customers through conversational workflows to accomplish tasks using a company’s products and services. Although CX agents may not require the same depth of reasoning as coding or mathematical agents, they must meet complex, multidimensional requirements across diverse user groups. More specifically, CX agents should support personalization and localization (Kirk et al., 2024), respond appropriately to user intent across multi-turn conversations (Tack et al., 2026), and handle sensitive data (Vijayvargiya et al., 2026; Tien et al., 2026), among other requirements. Thus, the bars for safety and qualitative performance are higher than those for general agentic applications. Such unique expectations for CX agents raise a central question: how can we establish that an agent is operational, handles edge cases, and is ready for safe online deployment?

Figure 1 contrasts three approaches: (1) authoring test cases by hand, (2) exposing real customers in a live A/B test, and (3) simulating customers on-policy at the tool boundary. Although the first two options incorporate human feedback, they are slow and can provide inconsistent signals with small sample sizes. User simulation with synthetic personas has therefore been studied as an alternative for approximating the human preference distributions with LLMs (Ge et al., 2025; Naous et al., 2026; Li and others, 2026). However, limited controllability and the risk of score inflation raise concerns about discrepancies between simulated and real-world scenarios (Zhu et al., 2024; Mehri et al., 2026).

In this paper, we explore user simulation as an intermediate evaluation layer in which candidate agents are exercised through synthetic customer interactions before live deployment. Through case studies of the Card Delivery and Card Management agents at Nubank, we propose a hypothesis-driven, tool-boundary simulation recipe for screening candidate agents and demonstrate its real-world business impact. Simulation has become an essential part of our CX agent development lifecycle, enabling broad exploration of models, prompts, and reasoning settings that would be impractical to pursue through live experimentation alone. Our contributions are as follows:

  1. 1.

    User simulation as screening-before-deployment: We propose a simulation-aided agent deployment recipe with a hierarchical workflow connecting offline evaluation to online deployment.

  2. 2.

    Similarity analysis against real-world conversations: We demonstrate moderate conversation-level similarity between real-world and simulated conversations, with the lowest cosine distance of 0.035. Across four deployed versions, simulated and production version-level evaluator scores show high correlation.

  3. 3.

    Efficient development-to-production lifecycle: The user simulation layer allows 4.8 times faster iteration of the development, screening, and deployment cycle compared with a workflow without a user simulation layer.

  4. 4.

    Production impact across agent and model changes: The Card Management agent developed through simulation-guided iteration improved self-service rate (SSR) by 4.9 percentage points and transactional net promoter score (tNPS) by 36.69 points relative to Card Delivery in a live A/B test.

  5. 5.

    Generalizability across models: We use the proposed simulation recipe to select Qwen3.5-122B-A10B as a replacement for the incumbent Card Management model. In a subsequent live A/B test, the replacement improved SSR by 8.82 percentage points and reduced p95 latency by 25%, with no statistically significant change in tNPS.

2 Related Work

Customer experience agents in production.

An increasing number of companies across different business categories are adopting agentic workflows for customer support, including Amazon (Luo et al., 2025), Airbnb (Zhao et al., 2025; Su et al., 2025), Kakao (Park et al., 2025), Alibaba Group (Jiang et al., 2025), Thomson Reuters (Juclà et al., 2026), and Nubank (Gupta et al., 2026). Zhao et al. (2025) and Su et al. (2025) propose human-in-the-loop feedback collection strategies and synthetic data generation pipelines for sustainably aligning customer experience agents with human preferences at Airbnb. Gupta et al. (2026) present a hierarchical pipeline spanning offline validation and online deployment for customer experience agents at Nubank, improving agent self-service rate (SSR) and transactional net promoter score (tNPS) (Reichheld, 2003).

Table 1: Card Delivery (CD) versions used to characterize the simulator and Card Management (CM) configurations evaluated through the simulation workflow.
Card Delivery (CD) Card Management (CM)
Version Configuration Version Configuration
V​1CDV1_{\mathrm{CD}} GPT-4.1; tools for customer-specific logistics queries. V​1CMV1_{\mathrm{CM}} GPT-5.2; CD prompts and tools with ReAct-style tool definitions (Yao et al., 2023); baseline.
V​2CDV2_{\mathrm{CD}} GPT-4.1; card selection and prevention of recent duplicate reissues. V​2CMV2_{\mathrm{CM}} GPT-5.2; Manual revision of V​1CMV1_{\mathrm{CM}}’s prompt and tool-use policy.
V​3CDV3_{\mathrm{CD}} GPT-5.1; prompt rules for routing, personalization, and concise responses. V​3CMV3_{\mathrm{CM}} GPT-5.2; V​1CMV1_{\mathrm{CM}}’s prompt and tools with a different architecture.
V​4CDV4_{\mathrm{CD}} GPT-5.2; tools to retrieve and update the registered delivery address. V​4CMV4_{\mathrm{CM}} GPT-5.2; Manual revision of V​3CMV3_{\mathrm{CM}}’s prompt and tool-use policy.
V​5CMV5_{\mathrm{CM}} GPT-5.2; Return to ReAct; prompt revised using simulated examples.

User simulation in benchmarking agents.

To benchmark agents across a wide range of domains, synthetic personas have been studied as a method for simulating diverse human demands, either at scale through verbalized prompts (Ge et al., 2025; Li and others, 2026) or by fine-tuning the language model used for simulation (Naous et al., 2026). Recent agent benchmarks, in particular, leverage prompt-guided user personas to simulate real-world scenarios (Yao et al., 2025; Tan et al., 2025; Qian et al., 2025; Barres et al., 2025; Huang et al., 2026). Specifically, τ\tau-Bench defines multi-turn scenarios in which agents interact with prompt-driven user simulators in retail and airline customer experience (CX) domains (Yao et al., 2025), while τ2\tau^{2}-Bench extends this setting so that both agents and users can invoke tools, enabling more human-like actions by simulated users (Barres et al., 2025). Building on increasing efforts to expose CX agents to more realistic environments, we address the complementary question of how to characterize and use a simulator as a restricted screening layer in agent development for real-world deployment.

Reliability in user simulation.

Despite the increasing use of user simulation in agent benchmarking, a simulator of unknown fidelity can actively mislead (Zhu et al., 2024; Mehri et al., 2026). Prompt-driven user simulators are often prone to score inflation through data leakage (Zhu et al., 2024) or may fail to adhere to their assigned goals in multi-turn scenarios (Mehri et al., 2026). Accordingly, rigorously validating the robustness of user simulation pipelines and their alignment with real human behavior remains challenging (Dou et al., 2025; Zhu et al., 2026a). SimulatorArena explicitly measures agreement between human judgments and assistant ratings on public tasks such as math tutoring (Dou et al., 2025), while RealUserSim grounds simulator personas in behavioral profiles extracted from real conversations and tests fidelity using an LLM-judged paired-trajectory Turing test (Zhu et al., 2026a). We characterize the reliability and robustness of a user simulation recipe by comparing simulated and real-world conversations, and demonstrate how simulation-based findings can inform changes that are subsequently evaluated through real-world business outcomes at Nubank. Most importantly, our results demonstrate that imperfect simulation can still guide directionally correct improvements to production agents, with benefits confirmed with live A/B evaluation.

3 Simulation Setup

This section consolidates the experimental setup for multi-turn agent trajectory generation and evaluation (Figure 2). Our pipeline selects a target agentic task, simulates on-policy conversations with synthetic personas and tool responses, and evaluates the resulting trajectories offline.

Target agents.

We study Nubank’s Card Delivery (CD) agent and its expanded successor, Card Management (CM), which replaced CD and supports a broader range of card-lifecycle requests. We characterize the simulation pipeline using CD traces and apply the resulting workflow to CM.

3.1 Simulation pipelines

Proposed simulation pipeline.

Our proposed trajectory-generation pipeline uses Snowglobe as its primary simulator (Guardrails AI, 2025). Snowglobe instantiates the persona-driven simulation pattern of Park et al. (2023) and takes an agent description and tool definitions, optionally augmented with a simulation prompt and historical conversations. From these inputs, it constructs use cases, user profiles, tool relationships, and trajectory plans (Appendix A). An orchestration agent allocates conversations across use-case and interaction-style conditions, after which the generated personas interact with the target agent over multiple turns. Each synthetic user turn is conditioned on the agent’s preceding response, while tool calls are intercepted and answered with scenario-conditioned synthetic results rather than invoking production backends. A conversation ends when the persona’s objective is achieved, judged unreachable, or subject to a turn or tool-specification limit. We adopt Snowglobe for its structured orchestration, controllable scenario coverage, and cross-call tool-state consistency support for on-policy, hypothesis-driven agent screening.

Naive LLM baseline.

To assess what can be achieved through direct prompting without Snowglobe’s structured orchestration, we implement a naive LLM-based simulator. A persona language model generates informal Brazilian Portuguese customer turns, while a second language model produces synthetic results for the agent’s tool calls. Both roles use gpt-5.6-sol with high reasoning effort and receive the same categories of non-test-set context as Snowglobe: the agent description, tool schemas, product-requirements meta-knowledge, and tool-call examples. Both simulators also use the same seeded scenario distribution. A baseline conversation ends when the persona signals completion, the agent transfers the conversation to a human, or the customer has sent ten messages. Unlike Snowglobe, the baseline does not orchestrate use-case and interaction-style coverage, generate structured trajectory plans, or use an agent-profile model to maintain cross-call consistency. Appendix G provides the baseline prompts.

3.2 Card Delivery Agent

Agent versions and data.

8,000 real-world conversations between the users and the CD agent were evenly split into two sets: a development set to configure the simulators (i.e., simulation characterization) and a test set to validate the agent with the characterized simulation profiles. The four deployed CD versions are summarized in Table 1 (left).

Figure 3: Simulation recipe and characterization: (a) candidate agent changes are simulated and compared with an incumbent baseline using standing and hypothesis-specific criteria to screen candidates, (b) four diagnostics compare variant-aligned production and simulated Card Delivery samples, characterizing observed differences without establishing backend fidelity.

The test set contains 1,000 conversations from each agent version. We generated 250 synthetic trajectories via the simulator for each version.

Offline evaluation.

We use five semantic categories to evaluate agent trajectories offline. They are evaluated via LLM-as-a-Judge, using GPT-4.1-Mini11 1 https://openai.com/index/gpt-4-1/: (E1) Card reissue failure; (E2) Customer input verification; (E3) Card delivery data check; (E4) Response conciseness; (E5) Resolution conciseness following Gupta et al. (2026). We evaluate each category with separate system prompts optimized via GEPA (Agrawal et al., 2026), which return binary pass/fail scores.

3.3 Card Management Agent

Agent versions and data.

We apply the simulation workflow shown in Figure 3(a) to guide iterative improvements to the CM agent, which uses GPT-5.2 as its backbone model. Table 1 (right) summarizes the baseline and four additional versions developed through the proposed simulation workflow.

Offline evaluation.

We used a single binary judgment from LLM-as-a-Judge to test the explicit hypothesis in each version as described above. Given the CM agent’s execution traces and available tools, LLM-as-a-Judge provides binary feedback on whether a transfer to human support was unnecessary. We compare each candidate’s failure rate with that of V​1CMV1_{\mathrm{CM}} to decide whether to accept the change.

Online evaluation.

We assess production performance using two online metrics: transactional net promoter score (tNPS) and self-service rate (SSR). tNPS is the percentage of promoters minus the percentage of detractors in a post-interaction recommendation survey (Reichheld, 2003). tNPS is thus a measure of customer satisfaction and happiness. Self-service rate (SSR) is the proportion of sessions in which users completed the case with the agent without asking for human support.

Refer to caption
Figure 4: P2 – In all versions, both simulated groups are closer to production than the off-topic control; Snowglobe is no farther from production than the baseline in three versions. Cosine distances between whole-conversation embedding centroids. Underlined values mark the nearest non-production centroid.

4 User Simulation Recipe for CX Agents

Algorithm 1 uses tool-boundary trajectory simulation for hypothesis-driven screening (Balog and Zhai, 2024). Its purpose is to expose beneficial or harmful candidate–incumbent differences large enough to affect a deployment decision. Therefore, exact replication of production metrics is unnecessary.

Algorithm 1 Hypothesis-driven candidate screening
0:  Incumbent AA; simulator configuration SS; evaluation suite EE; prespecified screening criterion CC
1:  TA←Simulate⁡(A,S)T_{A}\leftarrow\mathrm{Simulate}(A;S)
2:  for each candidate revision Δ\Delta do
3:   A′←A+ΔA^{\prime}\leftarrow A+\Delta
4:   TA′←Simulate⁡(A′,S)T_{A^{\prime}}\leftarrow\mathrm{Simulate}(A^{\prime};S)
5:   Evaluate every trajectory in TAT_{A} and TA′T_{A^{\prime}} with EE
6:   if the candidate–incumbent comparison meets CC then
7:    Mark A′A^{\prime} eligible for live A/B testing
8:   else
9:    Reassess or revise Δ\Delta
10:   end if
11:  end for

The algorithm first generates on-policy incumbent trajectories using simulator configuration SS, which specifies the target use-case mix. These trajectories remain the shared baseline throughout the screening round. Configuration uses the agent description and tool definitions, optionally supplemented by a simulation prompt and historical data. Record scenario composition and keep distribution runs separate from targeted probes. Scenario-conditioned tool outputs must match declared schemas and remain consistent across calls within a conversation. For example, a record queried twice should not change state without an intervening action. Simulation exercises dialogue–tool interactions without testing backend implementations (Appendix B).

Each screening round is guided by a falsifiable hypothesis and a prespecified screening criterion CC. The hypothesis determines the behavioral effect of interest, while CC defines how evaluator results support eligibility for live testing. The evaluation suite EE comprises evaluators relevant to the task and the hypothesis, including LLM-as-a-judge evaluators when appropriate. It covers selected standing criteria and the proposed behavioral effect. The same suite is applied to incumbent and candidate trajectories; any newly introduced evaluator must also score the baseline trajectories. The standing CD suite was calibrated against human annotations (Gupta et al., 2026); calibration of each new hypothesis-specific judge must be reported separately.

Candidate simulations hold SS fixed. Candidates meeting CC become eligible for live A/B testing; others are reassessed or revised against the same incumbent. The baseline is updated only after a candidate becomes the incumbent. For candidates meeting CC, smaller or ambiguous differences are left for live experiments to resolve. This loop complements item-level evaluation and live A/B tests; Figure 3 summarizes it alongside the separate simulator-characterization diagnostics.

4.1 Simulator characterization

We introduce four different diagnostics for the characterized simulators, including three automated evaluations and one human annotation session, namely P1 to P4. For three synthetic evaluations, we use the data and offline evaluation pipeline introduced in Section 3.2, i.e., each uses 1,000 held-out production conversations with real-world users and 250 Snowglobe traces per CD version. The diagnostics address complementary properties (Figure 3b), none as a standalone criterion.

P1: Conversation-length statistics.

We compare the number of words in user messages and the number of user turns per conversation. We report cumulative threshold percentages with full distributions in Appendix C.

P2: Embedding diagnostics.

We embed each conversation using text-embedding-3-large22 2 https://developers.openai.com/api/docs/models/text-embedding-3-large with user: and agent: prefixes. Then, we measure cosine and Euclidean distances between the group centroids, and plot the two-dimensional Uniform Manifold Approximation and Projection (McInnes et al., 2018, UMAP) visualization.

P3: Evaluator-score association.

We apply five canonical evaluation categories, E1 to E5, in Section 3.2 to the production, Snowglobe, and baseline pools. For each agent version, we report the average scores per category and their 95% confidence intervals. Across the versions, we report average rank, Pearson rr, and Kendall τ\tau relative to the production ordering.

P4: Blinded human source discrimination.

We sampled N=𝟏𝟎𝟎N{=}\boldsymbol{100} V​4CDV4_{\mathrm{CD}} conversations, comprising 50 production and 50 Snowglobe trajectories. V​4CDV4_{\mathrm{CD}} was selected because it was the most recent version and the version most familiar to the annotators. Seven domain experts contributed labels. After the same source-blind normalization was applied to both groups, annotators saw dialogue text without tool-call sequences, classified each item as production or simulated, and reported confidence on a 1–5 scale.

4.2 Open-weight model screening

Models.

We test 29 configurations of the card management agent with over 16,000 simulated conversations, across model family, reasoning effort, and numerical format, including GPT-OSS-120B (OpenAI and others, 2025), Nemotron-3-Super-120B-A12B (NVIDIA and others, 2026), Qwen3.5-122B-A10B (Qwen Team, 2026a), and Qwen3.6-35B-A3B (Qwen Team, 2026b).

Evaluation.

Building on the offline evaluation approach in Section 3.2, we use LLM-as-a-Judge to evaluate the following categories: (1) gathering proper inputs, (2) validity of card reissue, (3) information retrieval status, and (4) completeness of conversation, which are noted as “Input”, “Reissue”, “Status”, and “Complete” in Table 3.

5 Results

We characterize the simulator on CD versions, then report the offline CM analysis and the live A/B comparison, followed by a predeployment screen of open-weight models.

Figure 5: P3 – Snowglobe better preserves the production ordering than the baseline. Aggregate E1–E5 failure scores for simulated and baseline conversations versus production (lower is better). Labels 1–4 identify CD versions. Error bars show 95% conversation-clustered bootstrap confidence intervals.
Figure 6: P2 – Snowglobe samples are concentrated near production, while the baseline spans a broader region extending toward the off-topic control. Joint UMAP projections compare Snowglobe with production (top) and the baseline with production (bottom); the off-topic control is overlaid in both rows. Haloed stars mark projected group means.

5.1 Simulator characterization

P1: Conversation length.

Appendix C reports both selected cumulative thresholds for the four CD versions (Table 4) and the complete distributions (Figure 9). In the pooled samples, 22.4%\boldsymbol{22.4\%} of simulated conversations contain at most 50 user-message words, compared with 89.0%\boldsymbol{89.0\%} in production; 14.9%\boldsymbol{14.9\%} contain at least 200 words, compared with 0.1%\boldsymbol{0.1\%} in production. Similarly, 16.4%\boldsymbol{16.4\%} of simulated conversations have at most two user turns, compared with 36.8%\boldsymbol{36.8\%} in production. The shares with at least six and eight turns are 55.0%\boldsymbol{55.0\%} and 25.2%\boldsymbol{25.2\%} in simulation, versus 24.0%\boldsymbol{24.0\%} and 10.6%\boldsymbol{10.6\%} in production.

P2: Transcript-embedding proximity.

For each CD version, the cosine distance between the production and simulated transcript centroids is smaller than the distance from either centroid to the off-topic control centroid (Figure 4). The same ordering holds for Euclidean distance (Appendix D, Figure 10). Compared with the baseline, Snowglobe is closer to production in two versions, tied in one at the reported precision, and farther in one by cosine distance; by Euclidean distance, it is closer in three versions and farther in one. Both simulated groups remain closer to production than to the off-topic control. Figure 6 provides the two-dimensional UMAP projections.

P3: Evaluator-score association.

Figure 5 shows that Snowglobe preserves the production extremes, V​4CDV4_{\mathrm{CD}} best and V​2CDV2_{\mathrm{CD}} worst, although V​1CDV1_{\mathrm{CD}} and V​3CDV3_{\mathrm{CD}} exchange positions. The baseline instead ranks V​3CDV3_{\mathrm{CD}} best and V​4CDV4_{\mathrm{CD}} worst. V​4CDV4_{\mathrm{CD}} is best in 96.69%\boldsymbol{96.69\%} of Snowglobe resamples and 𝟎%\boldsymbol{0\%} of baseline resamples; V​2CDV2_{\mathrm{CD}} is worst in 𝟏𝟎𝟎%\boldsymbol{100\%} and 1.56%\boldsymbol{1.56\%}, respectively. Table 5 in Appendix E reports the aggregate scores and ranking-agreement metrics. The average-rank statistic is 1.38 for Snowglobe and 1.62 for the baseline. The Snowglobe and production scores have Pearson 𝒓=0.74\boldsymbol{r=0.74} and Kendall 𝝉=0.67\boldsymbol{\tau=0.67}; their largest absolute difference is for V​2CDV2_{\mathrm{CD}} (0.556 versus 0.437).

P4: Blinded human source discrimination.

Annotators correctly identified 42 of 50 production conversations (84.0%\boldsymbol{84.0\%}) and 35 of 50 simulated conversations (70.0%\boldsymbol{70.0\%}). Thus, 15 simulations were classified as production. Reported confidence was approximately 3 on the 1–5 scale, including for correctly classified items.

5.2 Card Management: offline and live results

Applying the simulation recipe, V​5CMV5_{\mathrm{CM}} had the lowest unnecessary-transfer failure rate (16.0%\boldsymbol{16.0\%}), compared with 22.4%\boldsymbol{22.4\%} for V​1CMV1_{\mathrm{CM}}, 28.8%\boldsymbol{28.8\%} for V​3CMV3_{\mathrm{CM}}, and 34.0%\boldsymbol{34.0\%} for V​2CMV2_{\mathrm{CM}} and V​4CMV4_{\mathrm{CM}}. This represents a 6.4-percentage-point reduction relative to V​1CMV1_{\mathrm{CM}}. In the subsequent A/B test, the CM arm had SSR 4.90\boldsymbol{4.90} percentage points above CD (95% CI [4.08, 5.71]\boldsymbol{[4.08,\,5.71]}) and tNPS 36.69\boldsymbol{36.69} points above CD (95% CI [32.81, 40.57]\boldsymbol{[32.81,\,40.57]}) (Table 2). Following the test, the agent became available to Nubank’s customer base in Brazil. A later test replaced the incumbent model in CM with the open-weight candidate screened in Section 5.3, Qwen3.5-122B-A10B with reasoning enabled. SSR rose by +8.82\boldsymbol{+8.82} percentage points (95% CI [7.95, 9.69]\boldsymbol{[7.95,\,9.69]}, n=8.4​𝐊n=\boldsymbol{8.4\mathrm{K}}) while tNPS was statistically unchanged (−1.21\boldsymbol{-1.21} points, 95% CI [−3.97, 1.55]\boldsymbol{[-3.97,\,1.55]}, n=2.3​𝐊n=\boldsymbol{2.3\mathrm{K}}), and p95 latency fell by 𝟐𝟓%\boldsymbol{25\%}. The screened open-weight configuration therefore matched the incumbent on satisfaction while improving SSR and latency.

A batch of 100 simulated trajectories completed in under 10 minutes on average; this measures generation runtime rather than end-to-end iteration time. The ten-version CD development cycle, including the four versions studied here, spanned 212 calendar days. The five CM versions spanned 22 days. This corresponds to 21.2\boldsymbol{21.2} and 4.4\boldsymbol{4.4} days per version, respectively, a ratio of 4.8\boldsymbol{4.8}.

Table 2: Live A/B outcomes for the two screened changes. Each panel reports differences against its own control. The p95 latency change is a point estimate from the serving stack.
Outcome Difference 95% CI nn
Card Management vs. Card Delivery
SSR (pp) +4.90+4.90 [4.08, 5.71][4.08,\,5.71] 27.8​K27.8\mathrm{K}
tNPS (pp) +36.69+36.69 [32.81, 40.57][32.81,\,40.57] 2.0​K2.0\mathrm{K}
Qwen3.5-122B-A10B (reasoning on) vs. incumbent
SSR (pp) +8.82+8.82 [7.95, 9.69][7.95,\,9.69] 8.4​K8.4\mathrm{K}
tNPS (points) −1.21-1.21 [−3.97, 1.55][-3.97,\,1.55] 2.3​K2.3\mathrm{K}
p95 latency (%) −25-25 — —

5.3 Open-weight model screening

Table 3 reports four evaluator failure rates for the open-weight configurations. Performance by criterion, with no configuration uniformly outperforming the others. Qwen3.5 achieves the lowest input gathering (“Input”) failure rate with reasoning enabled (0.4%) and the lowest conversation completeness (“Complete”) failure rate with reasoning disabled (16.%). Qwen3.6 with reasoning enabled leads on card reissue validity (“Reissue”) (14.0%), while GPT-OSS with medium reasoning has the lowest information retrieval status (“Status”) failure rate among open-weight configurations (17.6%). The incumbent retains a substantial advantage on Status (2.4%).

The effect of reasoning also varies by model and criterion. For Qwen3.6, enabling reasoning lowers all four reported mean failure rates. For Nemotron, moving from no to low reasoning reduces Complete failures from 60.8% to 18.0%, but increases Status failures from 21.2% to 36.4%. For Qwen3.5, reasoning reduces Input and Status failures while increasing Reissue and Complete failures. These results imply model-specific benefits from additional reasoning.

We selected Qwen3.5-122B-A10B with reasoning enabled as a starting point for prompt optimization based on its low Input failure rate and competitive Reissue and Complete results. Relative to the incumbent, its Input and Complete failure rates are lower by 3.2 and 15.0 percentage points, respectively, while Reissue is similar (25.4% versus 25.2%). The higher Status failure rate (35.2% versus 2.4%) identified a weakness for further refinement. Subsequent prompt optimization reduced failures on this criterion. The offline results therefore supported Qwen3.5-122B-A10B as a candidate for live evaluation, without establishing it as the strongest configuration across all criteria. Table 2 reports the subsequent live A/B outcomes.

Table 3: Canonical evaluation failure rates (%) on simulated card management conversations (↓\downarrow). Bold and underlined values indicate the lowest and second-lowest means among open-weight configurations, respectively.
Configuration Reasoning Canonical Evaluations (↓\downarrow)
Input Reissue Status Complete
GPT-5.2 (Incumbent) — 3.63.03.6_{3.0} 25.27.925.2_{7.9} 2.43.32.4_{3.3} 37.610.137.6_{10.1}
Qwen3.6-35B-A3B Off 1.61.71.6_{1.7} 24.03.724.0_{3.7} 29.29.029.2_{9.0} 25.24.125.2_{4.1}
On 1.21.81.2_{1.8} 14.03.7\mathbf{14.0}_{3.7} 25.63.325.6_{3.3} 17.6¯7.4\underline{17.6}_{7.4}
Qwen3.5-122B-A10B Off 0.8¯1.1\underline{0.8}_{1.1} 24.87.724.8_{7.7} 38.86.938.8_{6.9} 16.03.7\mathbf{16.0}_{3.7}
On 0.40.9\mathbf{0.4}_{0.9} 25.47.725.4_{7.7} 35.210.135.2_{10.1} 22.612.222.6_{12.2}
Nemotron-3-Super- 120B-A12B Off 6.42.26.4_{2.2} 32.07.632.0_{7.6} 21.25.021.2_{5.0} 60.812.560.8_{12.5}
Low 3.22.33.2_{2.3} 20.06.320.0_{6.3} 36.42.636.4_{2.6} 18.04.218.0_{4.2}
High 1.22.71.2_{2.7} 24.04.924.0_{4.9} 30.06.330.0_{6.3} 19.25.419.2_{5.4}
GPT-OSS-120B Low 2.02.42.0_{2.4} 18.4¯3.6\underline{18.4}_{3.6} 28.46.228.4_{6.2} 17.6¯4.3\underline{17.6}_{4.3}
Medium 3.22.33.2_{2.3} 26.07.726.0_{7.7} 17.67.9\mathbf{17.6}_{7.9} 20.84.620.8_{4.6}
High 2.40.92.4_{0.9} 28.07.728.0_{7.7} 19.6¯3.0\underline{19.6}_{3.0} 18.86.118.8_{6.1}

5.4 Deployment lessons and limitations

Simulation supported safer, faster iteration: batches of 100 trajectories ran in under ten minutes, evaluators surfaced behavioral regressions, and the screen caught tool-call failures and a serving configuration without a tool-call parser. Most engineering effort went into integration: the staging agent reused live MCP schemas while stateful tools returned schema-compatible synthetic responses. The agent retained tool control and sequencing, but schemas, examples, authentication, and simulator profiles required ongoing alignment.

The approach deliberately stops at the tool boundary. Stateful or side-effecting tools are mocked, while read-only dependencies such as knowledge-base retrieval may remain live; backend behavior, latency, persistent state, and side effects thus remain untested. Simulation traces must also match production schemas, identifiers, and telemetry, or export and reconciliation work can erase the iteration gains.

6 Conclusion

In this paper, we present a hypothesis-driven approach to using simulation as a pre-deployment screening layer for CX agents. At Nubank, simulation has become an integral part of our development life cycle, enabling teams to explore changes, identify behavioral failures, and refine candidates before customer exposure. Its value lies in supporting development decisions, not in perfectly reproducing production. By combining simulation with live evaluation, this workflow supports broader experimentation while keeping deployment decisions grounded in observed customer outcomes.

References

  • Agrawal et al. (2026) L. A. Agrawal, S. Tan, D. Soylu, N. Ziems, R. Khare, K. Opsahl-Ong, A. Singhvi, H. Shandilya, M. J. Ryan, M. Jiang, C. Potts, K. Sen, A. G. Dimakis, I. Stoica, D. Klein, M. Zaharia, and O. Khattab GEPA: reflective prompt evolution can outperform reinforcement learning. In International Conference on Learning Representations (ICLR), Note: Oral External Links: Link Cited by: §3.2.
  • Balog and Zhai (2024) K. Balog and C. Zhai User simulation for evaluating information access systems. Foundations and Trends in Information Retrieval 18 (1–2), pp. 1–261. External Links: Document, Link Cited by: §4.
  • Barres et al. (2025) V. Barres, H. Dong, S. Ray, X. Si, and K. Narasimhan τ2\tau^{2}-Bench: evaluating conversational agents in a dual-control environment. arXiv preprint arXiv:2506.07982. External Links: Link Cited by: §2.
  • Dou et al. (2025) Y. Dou, M. Galley, B. Peng, C. Kedzie, W. Cai, A. Ritter, C. Quirk, W. Xu, and J. Gao SimulatorArena: are user simulators reliable proxies for multi-turn evaluation of AI assistants?. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 35212–35290. External Links: Document, Link Cited by: §2.
  • Ge et al. (2025) T. Ge, X. Chan, X. Wang, D. Yu, H. Mi, and D. Yu Scaling synthetic data creation with 1,000,000,000 personas. External Links: 2406.20094, Link Cited by: §1, §2.
  • Gemini Team et al. (2024) Gemini Team et al. Gemini 1.5: unlocking multimodal understanding across millions of tokens of context. External Links: 2403.05530, Link Cited by: §1.
  • Guardrails AI (2025) Guardrails AI SnowGlobe: the simulation engine for AI agents and chatbots. Note: Product whitepaperAccessed 2026-05-21 External Links: Link Cited by: §3.1.
  • Guo et al. (2025) D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, et al. DeepSeek-r1 incentivizes reasoning in llms through reinforcement learning. Nature 645 (8081), pp. 633–638. Cited by: §1.
  • Gupta et al. (2026) A. Gupta, K. Rossell, E. Alcobaça, J. C. L. Pacheco, C. B. de Lima, S. Tang, L. P. Rabachini, L. Moneda, H. Fei, D. Silva, and R. Ramanath Building customer support agents at 100M-user scale: an evaluation-driven framework. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Industrial Track), Note: To appear Cited by: §1, §1, §2, §3.2, §4.
  • Huang et al. (2026) K. Huang, A. Prabhakar, O. Thorat, D. Agarwal, P. K. Choubey, Y. Mao, S. Savarese, C. Xiong, and C. Wu CRMArena-pro: holistic assessment of LLM agents across diverse business scenarios and interactions. Transactions on Machine Learning Research. Note: External Links: ISSN 2835-8856, Link Cited by: §2.
  • Jiang et al. (2025) X. Jiang, T. Hu, Y. Qin, G. Wang, Z. Huan, K. Chen, G. Huang, R. Lu, and S. Tang ChatMap: mining human thought processes for customer service chatbots via multi-agent collaboration. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 11927–11947. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §2.
  • Juclà et al. (2026) D. G. Juclà, M. Tuteja, M. E. Casademunt, K. Unnikrishnan, Y. Usmani, and A. Roshaan Retrieval enhancements for RAG: insights from a deployed customer support chatbot. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track), Y. Matusevych, G. Eryiğit, and N. Aletras (Eds.), Rabat, Morocco, pp. 169–180. External Links: Link, Document, ISBN 979-8-89176-384-5 Cited by: §2.
  • Kirk et al. (2024) H. R. Kirk, A. Whitefield, P. Röttger, A. M. Bean, K. Margatina, R. Mosquera, J. M. Ciro, M. Bartolo, A. Williams, H. He, B. Vidgen, and S. A. Hale The PRISM alignment dataset: what participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §1.
  • Li et al. (2026) X. Li et al. MatrAIx: simulating the world with 8.3 billion persona agents. External Links: 2608.04205, Link Cited by: §1, §2.
  • Luo et al. (2025) C. Luo, D. Papadimitriou, H. Muralidharan, D. Ramasubbu, A. Kolekar, W. Xu, C. Xu, A. Srinivasan, M. Jain, and Q. He Language model alignment for conversational shopping at amazon. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’25, New York, NY, USA, pp. 4314–4318. External Links: ISBN 9798400715921, Link, Document Cited by: §1, §1, §2.
  • Lynch et al. (2025) A. Lynch, B. Wright, C. Larson, K. K. Troy, S. J. Ritchie, S. Mindermann, E. Perez, and E. Hubinger Agentic misalignment: how llms could be an insider threat. Anthropic Research. Note: https://www.anthropic.com/research/agentic-misalignment Cited by: §1.
  • McInnes et al. (2018) L. McInnes, J. Healy, and J. Melville UMAP: uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426. Cited by: §4.1.
  • Mehri et al. (2026) S. Mehri, X. Yang, T. Kim, G. Tur, S. Mehri, and D. Hakkani-Tür Goal alignment in LLM-based user simulators for conversational AI. Transactions of the Association for Computational Linguistics. External Links: Document, Link Cited by: §1, §2.
  • Naous et al. (2026) T. Naous, P. Laban, W. Xu, and J. Neville Flipping the dialogue: training and evaluating user language models. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §2.
  • NVIDIA et al. (2026) NVIDIA et al. Nemotron 3 super: open, efficient mixture-of-experts hybrid mamba-transformer model for agentic reasoning. External Links: 2604.12374, Link Cited by: §4.2.
  • OpenAI et al. (2025) OpenAI et al. Gpt-oss-120b & gpt-oss-20b model card. External Links: 2508.10925, Link Cited by: §4.2.
  • OpenAI (2025) OpenAI OpenAI o3 and o4-mini system card. Technical report OpenAI. External Links: Link Cited by: §1.
  • Park et al. (2025) C. Park, W. Jang, D. Kim, A. Ahn, K. Yang, W. Hwang, J. Roh, H. Park, H. Wang, M. S. Kim, and J. Kang A practical approach for building production-grade conversational agents with workflow graphs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), G. Rehm and Y. Li (Eds.), Vienna, Austria, pp. 1508–1519. External Links: Link, Document, ISBN 979-8-89176-288-6 Cited by: §2.
  • Park et al. (2023) J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein Generative agents: interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), Note: Best Paper Award External Links: Document, Link Cited by: §3.1.
  • Qian et al. (2025) C. Qian, Z. Liu, A. Prabhakar, Z. Liu, J. Zhang, H. Chen, H. Ji, W. Yao, S. Heinecke, S. Savarese, and H. Wang UserBench: an interactive gym environment for user-centric agents. In Workshop on Scaling Environments for Agents, External Links: Link Cited by: §2.
  • Qin et al. (2024) Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, L. Hong, R. Tian, R. Xie, J. Zhou, M. Gerstein, dahai li, Z. Liu, and M. Sun ToolLLM: facilitating large language models to master 16000+ real-world APIs. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1.
  • Qwen Team (2026a) Qwen Team Qwen3.5: towards native multimodal agents. External Links: Link Cited by: §4.2.
  • Qwen Team (2026b) Qwen Team Qwen3.6-35B-A3B: agentic coding power, now open to all. External Links: Link Cited by: §4.2.
  • Reichheld (2003) F. F. Reichheld The one number you need to grow. Harvard Business Review 81 (12), pp. 46–54. Cited by: §2, §3.3.
  • Research et al. (2026) C. Research, :, A. Chan, A. Shalaby, A. Wettig, A. Sanger, A. Zhai, A. Ajay, A. Nair, C. Snell, C. Lu, C. Shen, E. Jia, F. Cassano, H. Liu, H. Chen, H. Wildermuth, J. Jackson, J. Li, J. Katz, J. Yao, J. Hejna, J. Warner, J. Vering, K. Frans, L. Danilek, L. Wright, L. Cen, L. Melas-Kyriazi, M. Truell, M. de Jong, N. Jain, N. Schmidt, N. Wang, N. Muennighoff, O. Rybkin, P. Loh, P. Kravtsov, R. Yadav, S. Shah, S. Kottler, A. M. Rush, S. Zhang, S. Jain, S. Sankar, S. Heule, S. H. Sul, S. Asif, V. Rong, W. Zhu, W. Lin, Y. Wu, Y. Volkov, Y. Zemlyanskiy, Z. Holbrook, and Z. Zhang Composer 2 technical report. External Links: 2603.24477, Link Cited by: §1.
  • Schick et al. (2023) T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom Toolformer: language models can teach themselves to use tools. In Thirty-seventh Conference on Neural Information Processing Systems, External Links: Link Cited by: §1.
  • Steinberger and Contributors (2026) P. Steinberger and O. F. Contributors OpenClaw: your own personal ai assistant. GitHub. Note: https://github.com/openclaw/openclawAccessed: 2026-09-03 Cited by: §1.
  • Su et al. (2025) H. Su, W. Luo, Y. Mehdad, W. Han, E. Liu, W. Zhang, M. Zhao, and J. Zhang LLM-friendly knowledge representation for customer support. In Proceedings of the 31st International Conference on Computational Linguistics: Industry Track, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, S. Schockaert, K. Darwish, and A. Agarwal (Eds.), Abu Dhabi, UAE, pp. 496–504. External Links: Link Cited by: §2.
  • Tack et al. (2026) J. Tack, P. Laban, and J. Neville LLMs get lost in evolving user intent. External Links: 2607.20734, Link Cited by: §1.
  • Tan et al. (2025) J. Tan, L. Yang, Z. Liu, Z. Liu, R. R N, T. M. Awalgaonkar, J. Zhang, W. Yao, M. Zhu, S. Kokane, S. Savarese, H. Wang, C. Xiong, and S. Heinecke PersonaBench: evaluating AI models on understanding personal information through accessing (synthetic) private user data. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 878–893. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §2.
  • Tien et al. (2026) J. Tien, A. Anand, Y. Tuan, Y. Shen, J. Z. Kolter, and A. Nayebi ROGUE: misaligned agent behavior arising from ordinary computer use. External Links: 2606.00341, Link Cited by: §1.
  • Vijayvargiya et al. (2026) S. Vijayvargiya, A. B. Soni, X. Zhou, Z. Z. Wang, N. Dziri, G. Neubig, and M. Sap OpenAgentSafety: a comprehensive framework for evaluating real-world AI agent safety. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1.
  • Yao et al. (2025) S. Yao, N. Shinn, P. Razavi, and K. Narasimhan τ\tau-bench: a benchmark for tool-agent-user interaction in real-world domains. In The Thirteenth International Conference on Learning Representations (ICLR), External Links: Link Cited by: §2.
  • Yao et al. (2023) S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), Cited by: Table 1.
  • Zhao et al. (2025) C. Zhao, T. Zhang, H. Su, Y. Zhang, S. Su, M. Xu, Y. Liu, W. Han, J. Werner, C. N. Cheng, and Y. Mehdad Agent-in-the-loop: a data flywheel for continuous improvement in LLM-based customer support. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, S. Potdar, L. Rojas-Barahona, and S. Montella (Eds.), Suzhou (China), pp. 1919–1930. External Links: Link, Document, ISBN 979-8-89176-333-3 Cited by: §1, §2.
  • Zhu et al. (2024) L. Zhu, X. Huang, and J. Sang How reliable is your simulator? analysis on the limitations of current LLM-based user simulators for conversational recommendation. In Companion Proceedings of the ACM Web Conference 2024 (WWW ’24), pp. 1726–1732. External Links: Document, Link Cited by: §1, §2.
  • Zhu et al. (2026a) M. Zhu, J. Tan, R. Murthy, J. Qiu, L. Yang, W. Zhao, S. Savarese, S. Heinecke, and H. Wang RealUserSim: bridging the reality gap in agent benchmarking via grounded user simulation. arXiv preprint arXiv:2605.20204. External Links: Link Cited by: §2.
  • Zhu et al. (2026b) Y. Zhu, T. Jin, Y. Pruksachatkun, A. K. Zhang, S. Liu, S. Cui, S. Kapoor, S. Longpre, K. Meng, R. Weiss, F. Barez, R. Gupta, J. Dhamala, J. Merizian, M. Giulianelli, H. Coppock, C. Ududec, A. Kellermann, J. S. Sekhon, J. Steinhardt, S. Schwettmann, A. Narayanan, M. Zaharia, I. Stoica, P. Liang, and D. Kang Establishing best practices in building rigorous agentic benchmarks. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §1.

Appendix A How Snowglobe infers simulation properties

Before it can interact with an agent in a meaningful way, Snowglobe needs an approximation of that agent’s objectives, its users, and its behavior. The minimal input is a description of the agent together with the definitions of its tools. From these, a set of agents infers the properties a simulation depends on (Figure 7): the use cases the agent serves, profiles of synthetic users along with the data those users would carry, how the tools relate to one another and to the surrounding system, and the user trajectories that would call on each tool. Some of these properties are resolved once during agent onboarding and others are resolved when a simulation begins.

Two optional inputs reorient these properties. A simulation prompt narrows a run to a chosen set of use cases or behaviors while leaving the rest of the inferred picture intact. Real historical conversations, when supplied, are processed offline to extract use cases, linguistic styles, and their distributions, which are stored and used to condition later runs.

Agent description required Tool definitions required Simulation prompt optional Historical data optional Snowglobe inference a set of agents infer simulation properties Use cases the agent serves User profiles, with seed data Tool interaction model User trajectories →\rightarrow tools offline distributions steered by the prompt; matched to historical distributions inputs
Figure 7: Property inference in Snowglobe. From a required agent description and tool definitions, optionally augmented with a simulation prompt and historical data, a set of agents infers the use cases, user profiles and their seed data, the tool interaction model, and the user trajectories over the tools. Historical data is processed offline into features that can condition later runs.

Appendix B Tool boundary mocking

Exercising the agent’s tools inside a simulation raises a problem of its own. The simulated environment must not access production data through stateful tools, yet it has to return data through those tools so the agent can run end to end, and that data must stay consistent with the state of the conversation and with the results of earlier tool calls. Testing the tools themselves is a separate concern, better served by other methods, and is not an aim here.

Snowglobe meets these constraints without ever maintaining a database of state (Figure 8). During onboarding, it absorbs the tool definitions and builds a functional model of how the tools relate and how data must be shaped to satisfy each request. This model is called an agent profile. Using an agent profile, an orchestrator produces persona seeds. These are partial seed states paired with a trajectory for how that data should evolve over a conversation. Pesrona seeds are projected onto personas at runtime to ground their objectives.

At simulation time, a call to one of the agent’s tools is routed to the persona agent handling that conversation rather than to the real implementation. Using its partial state and the consistency rules in the agent profile, the persona agent returns a response that fits the tool’s output shape, and that response is passed back to the agent under test. Because no shared world state is instantiated, each conversation runs in its own sandbox. Snowglobe refers to this as tool-boundary mocking: only the individual tool calls at the boundary are answered, never an entire backend or database.

Simulated user persona-driven Agent under test user-provided turns Mock tool layer answers one tool call at a time tool callmocked result Persona from the sim — the world it implies Agent profile rules each tool must follow Real backend ⋅\cdot databases ⋅\cdot services ⋅\cdot persistent state ×\timesnever invoked
Figure 8: Tool-boundary mocking. As the simulated user and the agent under test exchange turns, each tool call the agent makes is answered by a mock tool layer parameterised by the persona and the agent profile. Only the tool boundary is mocked.

Appendix C Conversation length distributions

Figure 9 shows the full P1 distributions of user-message words and user turns, complementing the thresholds in Table 4. Length is user-side only. Groups are production traffic pooled over V​1CDV1_{\mathrm{CD}}–V​4CDV4_{\mathrm{CD}} (n=𝟒𝟎𝟎𝟎n=\boldsymbol{4000}), Snowglobe simulations of the same variants (n=𝟏𝟎𝟎𝟎n=\boldsymbol{1000}), and an off-topic control (n=𝟐𝟐𝟔n=\boldsymbol{226}). The retained 226-item control is encoded in the plotted artefact and historical experiment code, but the filtering from a documented upstream 1,000-item random control sample is not recorded. Table 4 shows summarized results in bins for easy comparison of extreme values.

Production user text is short (median 19 words; mean 25.8). The control matches that scale (median 20; mean 27.5); simulated users do not (median 111; mean 119). Control shares at the Table 4 word cuts are 85.8% (≤𝟓𝟎\leq\boldsymbol{50}) and 0% (≥𝟐𝟎𝟎\geq\boldsymbol{200}). Because the control is production traffic rather than LLM-generated user text, it cannot determine whether the heavy simulated tail is specific to this configuration or common to LLM user simulators.

In production, 24.3%\boldsymbol{24.3\%} of conversations have a single user turn, and the median is three turns. Control never has a one-turn chat and peaks at three turns (33.2%). Simulated conversations last longer (median 6) and hit a hard cap: 25.2% have exactly eight user turns, and none have more—which is why every simulated conversation has fewer than 10 user turns, the 95th-percentile value in production, despite the higher simulated median. Control shares at the table cuts are 18.6% (≤𝟐\leq\boldsymbol{2}), 17.7% (≥𝟔\geq\boldsymbol{6}), and 4.0% (≥𝟖\geq\boldsymbol{8}).

Figure 9: P1 – Snowglobe simulated conversations are substantially longer than production and off-topic conversations in both user-message words and user turns. The top row shows word count and the bottom row user turns: ECDFs (a,d), density or frequency (b,e), and violins with nested boxplots (c,f). Panel (e) pools values ≥𝟏𝟎\geq\boldsymbol{10} into the 10+ bin.
Table 4: P1 – Selected cumulative conversation-length thresholds for V​1CDV1_{\mathrm{CD}}–V​4CDV4_{\mathrm{CD}}. Values are percentages of conversations meeting each criterion; rows may overlap. Word counts include user messages only.
User-message words per conversation User turns per conversation
Criterion Production Simulated Criterion Production Simulated
≤50\leq 50 89.0 22.4 ≤2\leq 2 36.8 16.4
≥200\geq 200 words 0.1 14.9 ≥\geq 6 24.0 55.0
≥\geq 8 10.6 25.2

Appendix D Euclidean transcript-centroid distances

Figure 10 reports the Euclidean-distance counterpart to the cosine-distance result in Figure 4.

Refer to caption
Figure 10: P2 – In all versions, both simulated groups are closer to production than the off-topic control; Snowglobe is no farther from production than the baseline in three versions. Euclidean distances between whole-conversation embedding centroids. Underlined values mark the nearest non-production centroid.

Appendix E Evaluator-score association details

Table 5 reports the P3 aggregate scores and ranking-association statistics. Each conversation contributes five binary failure outcomes, one per evaluator. Within each source–version pool, we first average each evaluator across conversations and then average the five evaluator means; lower aggregate scores indicate fewer failures.

The 95% intervals use 10,000\boldsymbol{10{,}000} conversation-clustered bootstrap replicates. Each replicate resamples conversation rows with replacement, preserving the five outcomes for each sampled conversation, and recomputes the aggregate score. Production defines the reference ordering. Average rank summarizes closeness to that order (lower is better), Pearson rr measures linear association between version-level scores, and Kendall τ\tau measures rank agreement.

The extreme-rank analysis uses the same resampling procedure. Each replicate ranks the four versions by aggregate score, with ties sharing credit equally. The reported frequencies estimate only how often V​4CDV4_{\mathrm{CD}} is best and V​2CDV2_{\mathrm{CD}} is worst.

Table 5: P3 – Aggregate E1–E5 failure scores by version (lower is better), reported as means ±\pm 95% conversation-clustered bootstrap half-widths. Production defines the reference ranking; average rank, Pearson rr, and Kendall τ\tau measure agreement with it.
Source 𝑽​𝟏𝐂𝐃V1_{\mathrm{CD}} 𝑽​𝟐𝐂𝐃V2_{\mathrm{CD}} 𝑽​𝟑𝐂𝐃V3_{\mathrm{CD}} 𝑽​𝟒𝐂𝐃V4_{\mathrm{CD}} Avg. rank Pearson rr Kendall τ\tau
Baseline 0.453±0.0150.453_{\pm 0.015} 0.432±0.0200.432_{\pm 0.020} 0.379±0.0200.379_{\pm 0.020} 0.457±0.0200.457_{\pm 0.020} 1.621.62 −0.09-0.09 −0.33-0.33
Simulated 0.331±0.0220.331_{\pm 0.022} 0.556±0.0160.556_{\pm 0.016} 0.338±0.0190.338_{\pm 0.019} 0.303±0.0180.303_{\pm 0.018} 1.38 0.74 0.67
Production 0.408±0.0080.408_{\pm 0.008} 0.437±0.0110.437_{\pm 0.011} 0.363±0.0100.363_{\pm 0.010} 0.297±0.0090.297_{\pm 0.009} 1.001.00 1.001.00 1.001.00

Appendix F Open-weight screening: reasoning and quantization

Because a simulated arm costs no customer exposure, the screen can afford to sweep configuration settings that would otherwise be argued from intuition. This appendix reports the sweep summarised in Section 5.3, run on Nemotron-3-Super-120B-A12B across three quantizations and three reasoning settings. The figure omits the incumbent: the composite is a mean over evaluators built for this agent, and is used to compare a model against itself under different settings rather than to rank candidates against production.

Figure 11 shows the result. Reasoning effort moves the composite monotonically and by a large margin, from 42.9\boldsymbol{42.9} with reasoning off to 31.3\boldsymbol{31.3} with it on at nvfp4, and the mechanism is retrieval: the rate at which the model fetches the company record the task requires rises from 𝟏𝟔%\boldsymbol{16\%} to 𝟗𝟒%\boldsymbol{94\%} over the same range. Quantization does not move it. Across nvfp4, fp8, and bf16 the spread is 2.2\boldsymbol{2.2}, 0.8\boldsymbol{0.8}, and 4.1\boldsymbol{4.1} points at reasoning off, low, and on, against a 2.2\boldsymbol{2.2}-point standard deviation measured over six replicate runs of a fixed configuration. At this sample size the three quantizations are indistinguishable, while the reasoning setting moves the composite by roughly 𝟏𝟏\boldsymbol{11} points.

Figure 11: Reasoning and quantization sweep for Nemotron-3-Super-120B-A12B: composite evaluator score (lower is better) against reasoning setting, one series per quantization. The bar is ±𝟏\pm\boldsymbol{1} standard deviation of the composite measured over six replicate runs of a fixed configuration, and is a scale reference rather than a comparison against any arm.

Appendix G Baseline simulator prompts

Both simulator models (user and tools) are gpt-5.6-sol with reasoning_effort=high. Brace placeholders mark run-specific context concatenated at inference time: agent description, tool JSON schemas, product-requirements text, and tool-call examples.

G.1 Persona model

The persona is given one function tool, end_conversation, so it can trigger the end of the simulation. The first user message is shown below. Later turns reuse the same system prompt and append the dialogue (user and assistant text only).

You are a Brazilian Nubank customer chatting
in the Nubank app.
Stay in character. Write in informal Brazilian
Portuguese, like a real customer on a phone
keyboard: short messages, occasional typos,
no markdown.
You are talking to a support chatbot. You do
not work at Nubank. You do not know internal
tool names, IDs, or policies beyond what is
listed below.
Your persona:
{persona_card}
Rules:
- Send one customer message at a time.
- If the bot resolved your issue, or you are
done, call end_conversation.
- If the bot transfers you to a human, call
end_conversation.
- Do not invent that a human already joined.
- Maximum 10 of your messages in the whole
conversation.
## Agent description
{description}
## Tools the agent can use (description + schema)
{tools_schema}
## Product requirements (meta-knowledge)
{prd}
## Tool examples
{tool_examples}
Write your first message to the Nubank chatbot now.

Customer cards are sampled from a seeded generator.

Persona card Nome: {name}. Cidade: {city}. Problema: {issue}.
Tom: {tone}. Voc^e é cliente Nubank no Brasil e
está no chat do app.

G.2 Tool-result model

The mock backend is a second completion (no tools). The user message is the pending tool call.

You are an LLM tool-result simulator for a
Nubank card-delivery conversation.
Generate a realistic JSON result for the
given tool call.
Return ONLY the tool output (JSON if
possible). No preamble.
Stay consistent with earlier results in this
conversation when the same IDs appear.
## Agent description
{description}
## Tool description + schema
{tools_schema}
## Product requirements (meta-knowledge)
{prd}
## Tool examples
{tool_examples}
Generate a result for this tool call.
Tool: {tool_name}
Arguments: {arguments}