arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.40198v1 [cs.CL] 30 Sep 2026

SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Model

Kanpat Vesessook ††thanks: This work was primarily conducted during a 2024 internship at SCBX R&D. Affiliation: SCBX R&D Email: t_kanpat.v@scbx.com    Saksorn Ruangtanusak Affiliation: SCB DataX, SCBX Group Email: saksorn.ruangtanusak@data-x.ai
Abstract

Speech-to-speech systems must solve tasks whose requirements emerge across conversational turns. We introduce SpeechConversationBench (SCB), a focused evaluation of spoken mathematical reasoning using a pool of 103 sharded GSM8K problems. The framework compares the original problem delivered in one turn (Full), its concatenated information shards delivered together (Concat), and incremental spoken disclosure across turns (Sharded). We report final-answer accuracies for four commercial speech systems and LEGO, a proprietary speech pipeline developed internally by the SCBX Innovation Lab team, with explicit conversational context management. Relative to Concat, Sharded accuracy decreases by 5.0-25.3 percentage points across the four commercial systems. LEGO attains 77.5% in all three conditions, compared with 76.6% Sharded accuracy for GPT-4o Realtime. Both single-turn baselines are needed to distinguish sensitivity to reformulation from the challenges of incremental interaction.

1 Introduction and Related Work

Solving a complete spoken request does not establish that a system can solve the same task when relevant facts arrive over several turns. Mathematical word problems provide a focused test: success requires integrating incrementally disclosed quantities and constraints into a verifiable final answer.

Speech and dialogue evaluation.

VoiceBench evaluates voice assistants across content, speaker, and environmental variations [1], while Dynamic-SUPERB organizes diverse instruction-based speech tasks [6]. Multi-turn speech evaluation is also an established research direction. MTalk-Bench combines pairwise and rubric-based assessment across semantic, paralinguistic, and ambient-sound dimensions [4]. Audio MultiChallenge evaluates natural spoken interactions involving memory, instruction retention, self-coherence, and spoken repairs [5]. SpeechConversationBench complements these broader evaluations with a narrower comparison: numerical task completion under three presentations of corresponding problem information. It does not measure their full range of acoustic or interactional capabilities.

Incremental reasoning and memory.

We adapt the Full/Concat/Sharded protocol of Laban et al. [7] to spoken mathematical tasks. LoCoMo instead evaluates memory over extended, multi-session conversations [9]; our setting concerns integrating the facts of one problem within an interaction.

Architectural context.

Moshi models full-duplex speech-text dialogue [3], while MemGPT manages information across memory tiers [10]. Context-aware decoding strengthens adherence to spoken history [8]. These approaches motivate our comparison of five speech systems, including the context-managing cascade LEGO; they do not validate LEGO itself.

2 Evaluation Framework

2.1 Task and Input Conditions

The problem pool contains 103 mathematical word problems from the sharded GSM8K resource of Laban et al. [7]. GSM8K contains grade-school problems requiring multi-step numerical reasoning [2]. A shard is a partial statement of a problem’s information; the shards jointly specify the task. Spoken prompts provide the input, and correctness of the final spoken answer is the evaluation target. Figure 1 summarizes the conditions; Figure 2 shows the original simulator.

In Full, the original complete problem is presented in one spoken turn. In Concat, the corresponding shards are combined and presented together in a single turn. In Sharded, information is disclosed incrementally through a multi-turn spoken exchange. Thus, Full and Concat differ in formulation, while Concat and Sharded differ in how the shard content is distributed through interaction. This follows the conceptual controls of the original text protocol [7] without assuming that its model settings transfer to speech.

The Concat baseline is especially useful because reformulation can itself change difficulty. Comparing Sharded only with Full can conflate that change with the effect of incremental disclosure. Nevertheless, Concat versus Sharded is not an isolation of memory alone: assistant responses, speech processing, and the evolving conversational trajectory may also contribute to the observed difference.

GSM8K problem Full: original problem one spoken turn Concat: combined shards one spoken turn Sharded: incremental shards multi-turn spoken exchange Speech system evaluated in each condition Final spoken answer correctness
Figure 1: Spoken evaluation conditions. Full and Concat present information together; Sharded distributes it across an exchange. Final spoken answers are assessed for numerical correctness.
Refer to caption
Figure 2: Sharded-conversation simulator reproduced from Laban et al. [7]. This reference architecture motivates the spoken adaptation; it does not specify our audio-processing or scoring components.

2.2 Systems and Outcome Measures

We evaluate GPT-4o Realtime, GPT-4o Mini Realtime, Gemini 2.5 Flash Live, and Gemini 2.5 Flash Preview Native Audio Dialog, alongside LEGO. These names identify the systems in the reported comparison; they do not establish exact API snapshots. LEGO is a proprietary pipeline developed internally by the SCBX Innovation Lab team. It combines automatic speech recognition (ASR), a large language model (LLM) with explicit context management, and text-to-speech (TTS) synthesis, with all component models self-hosted internally. The pipeline incorporates the Thai semantic end-of-turn detection method described by Popit et al. [11]. Its context mechanism summarizes conversational history and reintroduces relevant information at successive turns. This is a system-level comparison rather than a controlled test of any individual component.

Let AFA_{F}, ACA_{C}, and ASA_{S} denote the reported percentages of correct final answers under Full, Concat, and Sharded. We summarize the multi-turn contrast by an absolute change and a normalized retention ratio:

ΔS−C=AS−AC,RS/C=100​ASAC.\Delta_{S-C}=A_{S}-A_{C},\qquad R_{S/C}=100\,\frac{A_{S}}{A_{C}}. (1)

The change is measured in percentage points (pp); the ratio expresses Sharded accuracy as a percentage of Concat accuracy. A retention value of 100% means equal aggregate accuracies, not that every problem has the same outcome. All derived quantities use the displayed accuracies and are rounded to one decimal place. The 103-problem pool is distinct from the per-condition trial denominators and repeat counts, which are not specified in the aggregate results. Consequently, we do not infer correct-answer counts or uncertainty intervals from that pool size.

3 Results

Table 1 reports all five systems. Among the commercial systems, Sharded accuracy ranges from 46.6% to 76.6%, and every system scores lower in Sharded than in Concat. The absolute decreases are 5.0 pp for GPT-4o Realtime, 25.3 pp for GPT-4o Mini Realtime, 15.5 pp for Gemini 2.5 Flash Live, and 25.0 pp for Gemini 2.5 Flash Preview Native Audio Dialog. These correspond to retention ratios of 93.9%, 67.4%, 75.0%, and 68.8%, respectively.

Table 1: Final-answer accuracy (%) under each input condition. ΔS−C\Delta_{S-C} is the Sharded-minus-Concat change in percentage points; RS/CR_{S/C} is aggregate accuracy retention (%). Derived columns use the displayed accuracies. Bold marks the highest Sharded accuracy without implying significance.
System Full Concat Sharded ΔS−C\Delta_{S-C} RS/CR_{S/C}
GPT-4o Realtime 70.0 81.6 76.6 −5.0-5.0 93.9
GPT-4o Mini Realtime 75.5 77.7 52.4 −25.3-25.3 67.4
Gemini 2.5 Flash Live 54.4 62.1 46.6 −15.5-15.5 75.0
Gemini 2.5 Flash Preview† 85.0 80.0 55.0 −25.0-25.0 68.8
LEGO 77.5 77.5 77.5 0.00.0 100.0

†Gemini 2.5 Flash Preview Native Audio Dialog.

The baseline changes the interpretation.

GPT-4o Realtime improves from 70.0% in Full to 76.6% in Sharded, even though it falls from 81.6% in Concat. A Full-only comparison would therefore conceal its disadvantage relative to the concatenated-shard baseline. More generally, the Concat-minus-Full changes are +11.6+11.6, +2.2+2.2, +7.7+7.7, and −5.0-5.0 pp for the four commercial systems, respectively. Both single-turn controls are therefore informative; degradation is not universal across baseline choices.

Single-turn rankings do not carry over.

Gemini 2.5 Flash Preview Native Audio Dialog has the highest Full accuracy, 85.0%, but achieves 55.0% in Sharded. GPT-4o Realtime has lower Full accuracy yet the strongest Sharded result among the commercial systems. Selecting a system by single-turn accuracy can therefore favor a different system than selecting by multi-turn accuracy.

LEGO preserves aggregate accuracy.

LEGO obtains 77.5% in all three conditions and the highest observed Sharded score, exceeding GPT-4o Realtime by 0.9 pp. Its zero aggregate change makes explicit context management worth investigating, but does not show that summarization caused the result. The systems differ in their components, and no memory ablation is available. Equal aggregate scores could also conceal problems that become correct and others that become incorrect. We therefore interpret LEGO as a promising system-level observation, not proof that a cascade is generally superior to native audio modeling.

4 Discussion, Limitations, and Conclusion

What the comparison measures.

The observed multi-turn gaps do not identify a failure mechanism. A wrong numerical answer can arise from misperception, loss or misuse of earlier information, arithmetic error, or answer rendering. Without intermediate transcripts and matched interventions, these explanations cannot be separated. The table establishes neither forgetting nor monotonic degradation with turn count.

Scope and reproducibility.

This small mathematical evaluation does not measure open-domain or multilingual dialogue, interruptions, prosody, noise robustness, latency, speech naturalness, or memory-management cost. The aggregate record does not specify exact model snapshots, audio-generation settings, sampling parameters, per-condition denominators, repeat counts, or answer-extraction and scoring details. LEGO’s component identities and memory-update implementation are also unspecified. These omissions limit reproduction and statistical interpretation.

Architectural interpretation.

A matched text-only baseline is needed to quantify an audio-specific penalty. To test whether LEGO’s context mechanism is responsible for its observed stability, the informative comparison would hold recognition, reasoning, and synthesis components fixed while varying context summarization. Problem-level outcomes and repeated runs would then support paired comparisons and uncertainty estimates. Such evidence would distinguish improvements in remembering information from improvements in using it, a distinction also raised by work on context-aware decoding [8]. Broader spoken benchmarks remain necessary to assess whether gains in numerical correctness extend to natural interaction.

Conclusion.

SpeechConversationBench compares spoken mathematical task completion across original, consolidated, and incrementally disclosed inputs. All four commercial systems lose accuracy from Concat to Sharded, whereas LEGO preserves its reported aggregate accuracy. The strongest supported lesson is methodological: retain both single-turn controls and report absolute task accuracy alongside the multi-turn change.

Acknowledgments and Disclosure of Funding

We thank Monthol Charattrakool, Peerawat Rojratchadakorn, and Natthapath Rungseesiripak for their engineering contributions to LEGO. We also thank Weerin Chantaroje, Head of Innovation Lab, SCBX, and Tutanon Sinthuprasith, Head of SCBX R&D, for their leadership and support.

References

  • [1] Y. Chen, X. Yue, C. Zhang, X. Gao, R. T. Tan, and H. Li (2024) VoiceBench: benchmarking LLM-based voice assistants. Note: arXiv:2410.17196 External Links: 2410.17196, Link Cited by: §1.
  • [2] K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman (2021) Training verifiers to solve math word problems. Note: arXiv:2110.14168 External Links: 2110.14168, Link Cited by: §2.1.
  • [3] A. Défossez, L. Mazaré, M. Orsini, A. Royer, P. Pérez, H. Jégou, E. Grave, and N. Zeghidour (2024) Moshi: a speech-text foundation model for real-time dialogue. Note: arXiv:2410.00037 External Links: Link Cited by: §1.
  • [4] Y. Du, Q. Huang, G. Zhu, Z. Dai, S. Chen, Q. Zhu, L. Pan, M. Chen, Y. Zhang, L. Zhou, B. Wang, and H. Li (2025) MTalk-Bench: evaluating speech-to-speech models in multi-turn dialogues via arena-style and rubrics protocols. Note: arXiv:2508.18240 External Links: Link Cited by: §1.
  • [5] A. Gosai, T. Vuong, U. Tyagi, S. Li, W. You, M. Bavare, A. Uçar, Z. Fang, B. Jang, B. Liu, and Y. He (2025) Audio MultiChallenge: a multi-turn evaluation of spoken dialogue systems on natural human interaction. Note: arXiv:2512.14865 External Links: Link Cited by: §1.
  • [6] C. Huang, K. Lu, S. Wang, C. Hsiao, C. Kuan, H. Wu, S. Arora, K. Chang, J. Shi, Y. Peng, R. Sharma, S. Watanabe, B. Ramakrishnan, S. Shehata, and H. Lee (2023) Dynamic-SUPERB: towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech. Note: arXiv:2309.09510 External Links: Link Cited by: §1.
  • [7] P. Laban, H. Hayashi, Y. Zhou, and J. Neville (2025) LLMs get lost in multi-turn conversation. Note: arXiv:2505.06120 External Links: 2505.06120, Link Cited by: §1, Figure 2, §2.1, §2.1.
  • [8] C. H. Lee, H. Kim, and S. Yoon (2026) From awareness to adherence: bridging the context gap in spoken dialogue systems via context-aware decoding. Note: arXiv:2606.16472 External Links: Link Cited by: §1, §4.
  • [9] A. Maharana, D. Lee, S. Tulyakov, M. Bansal, F. Barbieri, and Y. Fang (2024) Evaluating very long-term conversational memory of LLM agents. Note: arXiv:2402.17753 External Links: Link Cited by: §1.
  • [10] C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez (2023) MemGPT: towards LLMs as operating systems. Note: arXiv:2310.08560 External Links: Link Cited by: §1.
  • [11] T. Popit, N. Rungseesiripak, M. Charattrakool, and S. Ruangtanusak (2025) Thai semantic end-of-turn detection for real-time voice agents. Note: arXiv:2510.04016 External Links: 2510.04016, Link Cited by: §2.2.