Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
Abstract
Long-horizon tasks remain uncommon in large language model (LLM) evaluation, and for a reason: when each step depends on the last, per-step accuracy that looks excellent in isolation decays catastrophically, as errors cascade and the end-to-end failure probability grows sharply with length. Existing agentic benchmarks report end-to-end success but confound this state-tracking difficulty with instruction interpretation, give no control group that isolates it, and are vulnerable to shortcuts such as a hallucinated final answer, so they cannot say why a long run fails. Whether an LLM can carry exact intermediate state across many tool calls at all is itself not well established. We test this cleanly by having the model compute a cryptographic hash, MD5, step by step: a sequence of dependent tool calls over rounds while it carries four -bit words in its own context from one call to the next. Interpretation is trivial and, because we implement MD5 from scratch (RFC 1321), we align every call to the ground-truth trace and check the digest to the bit, so any failure is pure bookkeeping. gpt-oss-120b, a mixture-of-experts model with only 5.5B active parameters per token, at temperature with a short fixed prompt, carries the full state across all calls and returns the correct digest on a majority of completed runs. In the strongest setting we replace every primitive tool with a second LLM, so a driver and a worker compute the whole hash from scratch with no exact-arithmetic oracle in the loop. Two ingredients decide success and neither changes the weights: keeping the model’s own reasoning in its context each turn (stripping it makes success collapse), and voting over a thinking-enabled worker to remove its modular-arithmetic slips. We localize the residual failures by origin, separating state-carrying from arithmetic and from serving.
1 Introduction
Long, deterministic, rule-governed procedures are widely regarded as a weak spot for large language models (LLMs). The empirical case is strong. Single-step accuracy that looks excellent in isolation decays sharply once a task requires many dependent steps, because per-step errors compound and models “self-condition” on their own earlier mistakes (Sinha et al. 2025); and quality degrades as the context grows past a few tens of thousands of tokens or a few tens of turns (Liu et al. 2024; Modarressi et al. 2025). The usual engineering response is to take the model out of the driver’s seat and hand control to a hand-authored controller, a decision tree, a workflow graph, a finite-state machine, that owns the procedure while the LLM is confined to local text tasks.
This paper asks whether that trade-off is actually forced, and answers: no. Given the right context, an LLM can drive a long, exact, multi-step procedure to completion by itself. We make the claim under conditions chosen to be as unforgiving as possible. Interpretation is trivial and every step is provably right or wrong, so nothing can hide. The procedure is long and every step depends on the last, so there is nowhere to get lucky. And in the strongest configuration there is no exact-arithmetic oracle anywhere in the loop: even the primitive computations are performed by an LLM.
Our testbed is executing MD5 step by step. MD5 over a single -bit block (Rivest 1992) is rounds maintaining four -bit words . We expose each primitive (padding, per-round constants, the round functions, the mixing step, the final addition) as a tool, so the model’s job is to call the tools in order and carry from one call to the next, a canonical sequence of dependent tool calls. This isolates exactly the capability in question: relaying exact intermediate state across a long horizon, and taking many mechanical turns that a computer finds trivial, consistently. Because we implement MD5 from scratch (no hashlib) as an exact reference, we can align the agent’s calls to the ground-truth trace and check the digest to the bit.
We then remove the safety net. In the strongest configuration the primitive tools are not CPU code at all: each is replaced by a worker LLM that computes the operation, -bit bitwise , modular addition, left rotation. The driver LLM and the worker LLM together compute the entire hash from scratch, with a from-scratch CPU implementation used only to verify values for our analysis, never to supply them.
Driving gpt-oss-120b (OpenAI 2025a), a mixture-of-experts model with only 5.5B active parameters per token, at temperature with a short fixed system prompt, we find it carries the full state across all calls and reproduces the correct digest on non-memorized inputs on a majority of completed runs. Success turns on two conditions, neither a change to the weights: keeping the model’s own reasoning in its context each turn, and using self-consistency (majority voting) on the worker that computes the arithmetic. We analyze both, and characterize the residual failures, in the body.
Contributions.
-
1.
A controlled, ground-truthed benchmark in which an LLM executes MD5 step by step through tools, turning long-horizon exactness into a measurable, bit-checkable task with a canonical -call reference trace.
-
2.
A positive result: a 5.5B-active-parameter LLM executes the full -step procedure end-to-end and returns the correct digest, and, with a worker LLM replacing every primitive, the driver/worker pair computes the hash from scratch, arithmetic included.
-
3.
A driver/worker method with self-consistent workers that separates state tracking (driver) from primitive arithmetic (worker), and shows the arithmetic is made reliable by majority voting while the state is carried by the driver.
-
4.
An identification of the two context conditions that make it work, reasoning-in-context and self-consistency, and a characterization of the residual failures by origin (driver state, worker arithmetic, serving).
2 Background
Two threads of prior work make our setting look, a priori, nearly impossible.
Error accumulation over horizon.
A high per-step accuracy yields an end-to-end success rate of roughly over dependent steps; at even predicts a 14% success rate, and predicts . Sinha et al. (2025) formalize this: strong single-step accuracy systematically overstates long-horizon capability, models degrade as they condition on their own prior outputs, and larger models and chain-of-thought (Wei et al. 2022) push the horizon out but do not remove the wall.
Long-context degradation.
Independently, retrieval and reasoning quality drop as context length grows, the “lost in the middle” positional effect (Liu et al. 2024) and length-driven degradation (Modarressi et al. 2025), with many models falling below half their short-context score by 32k tokens, and function-calling accuracy degrading in long contexts specifically (Adhikary, Contractor et al. 2025).
A -step exact procedure sits squarely in the intersection of both hazards: long horizon and a growing transcript of tool calls and results. That is exactly why success here is informative, and why the conditions that produce it (reasoning-in-context, self-consistency, a driver/worker split) are the point.
MD5 (Rivest 1992) maps a message to a -bit digest. The message is padded and split into -bit blocks; each block is read as sixteen -bit words . A running state of four -bit words is initialized to fixed constants. Each block is processed by rounds. Round computes a nonlinear function of three of the state words, where is bitwise selection for and the analogous for the later quarters, together with a message index . It then forms
where is a per-round sine constant, a per-round rotation amount, is addition modulo , and is a left rotation; finally the four words are rotated . After rounds the block’s result is added, again modulo , into the running state, and the final is emitted little-endian as the digest. In our harness one block is a canonical tool calls: two to set up, three per round (constants, function, mix) across rounds, and two to finalize.
Two properties make this a demanding test. First, it is a long, strictly dependent computation: the steps must each be exactly right, and by design MD5 has no exploitable structure (its add-rotate-xor construction targets the avalanche effect, so the digest cannot be guessed or interpolated, only computed step by step). This makes it a clean probe of sustained exact execution. Second, the per-step arithmetic, -bit modular addition with carries and bit rotations, is exactly the kind of large-number exact computation on which LLMs are known to be brittle (Lee et al. 2024), so the worker path stresses a second, independent weakness.
3 Related Work
Long-horizon execution and long context.
Our result pushes against two strands of pessimism. Sinha et al. (2025) show that high single-step accuracy overstates multi-step capability because errors compound and models self-condition on their own outputs, the analytical wall. Separately, quality decays with context length: the “lost in the middle” positional effect (Liu et al. 2024), length-driven degradation past roughly k tokens (Modarressi et al. 2025), and tool-calling specifically degrading in long contexts (Adhikary, Contractor et al. 2025). A -step transcript exercises both.
Reasoning traces and exact arithmetic.
That intermediate work must be externalized to enable multi-step computation is the scratchpad finding (Nye et al. 2022) and, at prompting time, chain-of-thought (Wei et al. 2022); our reasoning-channel result is its serving-time analogue. The worker path stresses a second known weakness, exact arithmetic on large numbers, where LLM competence is surface-form and magnitude dependent (Lee et al. 2024); we make it reliable with primitive-level self-consistency (Wang et al. 2023) rather than a new method.
Tool-using agents and self-managed memory.
Tool-augmented agents such as ReAct (Yao et al. 2023) and Toolformer (Schick et al. 2023) raise capability but leave workflow state implicit in context, which is exactly what we stress. Our forward direction, a model that edits its own context, connects to self-editing memory architectures like MemGPT (Packer et al. 2023); our MD5 harness offers a bit-checkable testbed for whether such self-management preserves exact state across a long horizon.
4 Experimental Setup
Reference.
We implement MD5 from scratch (RFC 1321; no hashlib), validated against the RFC test vectors. This is ground truth: for any input we can produce the exact digest and the exact canonical sequence of primitive operations.
Tools.
Seven primitives are exposed as tools: prepare_message (pad and split into sixteen -bit words), initial_state, get_step_constants(), compute_f_and_g() (round function plus the message index), mix_step, add_states, and to_hex_digest. Each is a thin exact wrapper over the reference (CPU mode) or an LLM twin (worker mode). Figure 1 shows a single driver run with deterministic tools, and Figure 2 shows how a tool call is computed by voting workers.
Task.
A single-block input ( bytes, one -bit block, rounds). The driver agent reads a short, fixed instruction library and must call the tools in order while carrying between calls and applying the state-update rotation itself. A correct hash is a canonical tool calls ( setup finalize); the driver’s final answer is a -hex digest compared for exact equality with .
Input: message text (single block, bytes)
Output: -hex digest
MD5 is a clean difficulty dial: interpretation is trivial and every primitive is provably correct, so what remains to be tested is purely (i) whether the model can relay exact state between tools without dropping any, and (ii) whether it can take many mechanical turns consistently. We fix a single block deliberately, it keeps the instruction library linear, and note that scaling the horizon is as easy as adding blocks (each block is another rounds of the same loop), so the benchmark extends to arbitrarily long horizons without adding interpretive difficulty. We never use famous vectors (e.g. "abc") whose digests may be memorized; inputs are varied non-memorized single-block strings, including random ones, so a correct digest cannot be recalled and must be computed.
4.1 Models and configuration
Models.
Driver and worker are both gpt-oss-120b, an open-weight mixture-of-experts model with 117B total and 5.1–5.5B active parameters per token, shipped with the MoE expert weights natively quantized to MXFP4 (4.25 bits), which is what all our endpoints serve.
Parameters.
The driver runs at temperature with a short fixed system prompt and context caching enabled (the large invariant prefix, the instructions and the constant tables, is cached so only the growing tail is re-processed). In swap mode the worker runs at temperature with -sample majority voting. We repeat runs on multiple non-memorized single-block inputs.
Serving.
One provider serves the serverless and dedicated gpt-oss endpoints; a second serves the fast, thinking-enabled worker with its own low-precision kernels. We pin dates and log the endpoint per run; “temperature ” is not bit-deterministic on quantized serving, so we treat every cell as a distribution and report .
4.2 Verification and metrics
Ground-truth alignment.
We build the ground-truth call sequence for any input and align it call-for-call against what the agent actually did, rendering a green/red trace, green where the agent’s tool inputs match ground truth, red where they diverge. We inject each first divergence into the real algorithm to produce a counterfactual digest (“if only this one slip had occurred, the hash would be ”), separating a fresh error from its downstream propagation. A live mode flags the first state slip in real time, keyed by round index so benign duplicate lookups do not misalign the comparison.
Worker verification.
When a primitive is computed by a worker LLM, each swapped op is checked live against its CPU twin. This verification is for measurement and human legibility only, it is not a corrective oracle in the reported “from scratch” runs. A policy knob can either let a wrong worker value cascade (to observe end-to-end damage) or correct it in place (to attribute a failure), and we report both.
Metrics.
Success rate (primary): final digest equals true , exact, reported as successes / runs. First-divergence round: where the carried state first departs ground truth, as a distribution over runs. Calls-to-completion vs. the canonical (skips / duplicates / early stop). Per-op worker pass rate: worker vs. CPU twin, per primitive, with and without -sample voting. Latency (context only). Every request, response, tool call, result, and worker sample (including full reasoning) is logged as structured JSONL, indexed into SQLite with a per-session success / wrong / no-digest label computed against .
5 Results
5.1 State-carrying driver
The seven primitives run as exact CPU tools; the driver carries between them. The driver emits the full canonical -call sequence and returns the correct -hex digest, verified as a match, on multiple distinct inputs: a 5.5B-active-parameter MoE model carries exact -bit state from step to step . Against the prior, the horizon wall is not fundamental once the context is managed. Serving, not weights, then governs the rest. For the same weights, correctness and speed differ sharply by route (Table 1): a fast reasoning endpoint that is excellent single-shot is a poor multi-turn driver, looping or stopping early and rarely emitting a digest at all, which is itself evidence that the difficulty is sustained state across turns rather than any single step. When a run does fail, the tools are never wrong: the green/red trace goes red exactly where the driver feeds a tool a mis-carried -bit word and stays red as the corruption propagates. The most common fresh error is the state-update rotation (); failed runs also show dropped/duplicated calls and early stops (e.g. finalizing in rather than calls).
| Driver endpoint | Succ. | Wrong | Compl. | Rate |
|---|---|---|---|---|
| (gpt-oss-120b, echo on) | ||||
| Groq | ||||
| Fireworks |
5.2 Removing the arithmetic oracle
Design.
Every primitive is swapped onto a worker LLM, also gpt-oss-120b, but served on a fast, thinking-enabled endpoint as a single-shot, stateless computer. The worker runs at temperature with -sample majority voting; the majority value is returned to the driver, which continues its -call journey. The driver stays at temperature . The CPU twin records a pass/fail for each op for our analysis but does not supply values.
Findings.
With majority voting, per-op worker pass rates are high across all six computed primitives, so the driver/worker pair computes the hash from scratch with no exact oracle in the loop. The characteristic worker error is in the modular reduction of mix_step/add_states: it reasons its way to the correct unreduced sum but subtracts the wrong multiple of (e.g. the modulus where it needs ). The worker frequently catches and fixes this itself mid-reasoning by comparing magnitudes, and what it misses in one sample is corrected by the majority vote across three. When a full run does fail under swap-all it is overwhelmingly the driver mis-carrying state, not the worker computing a primitive wrong: correcting every worker op in place (removing arithmetic error entirely) does not by itself make a driver-slipped run succeed, so the two error sources are independent and state tracking is the harder one.
5.3 What makes it work
Reasoning in context.
The single largest lever is whether the model’s own reasoning trace is kept in its context each turn. gpt-oss emits its intermediate work in a dedicated reasoning channel (the “analysis” channel of OpenAI’s Harmony format (OpenAI 2025b)); the question is whether that channel is fed back on the next turn. With it, the driver tracks state across calls; stripped, success collapses on the identical model, prompt, and input. Crucially, the failure is not fixed by relocating the same text: when we place the reasoning content in the visible assistant message instead of the Harmony reasoning channel, keeping the tokens byte-for-byte identical, the driver still fails, and does so early, usually within the first to of steps, looping or finalizing prematurely. The effect is therefore specific to the reasoning channel being present as the model’s working memory, not to the tokens being visible somewhere. This mirrors the scratchpad result (Nye et al. 2022) that multi-step computation depends on where intermediate work lives, not merely that it exists. It is also provider plumbing, not weights: some stacks keep the reasoning field on input, others reject it (returning an HTTP on the echoed field), so the same weights behave differently by route. Consistent with this, most of the runs that never produce a digest are early collapses: of the logged Fireworks gpt-oss-120b sessions, emit no digest at all, and nearly half of those halt within the first of calls. It is a clean ablation with a large success delta and, we argue, the most actionable finding: manage the reasoning channel and you get the horizon.
With high thinking effort the driver’s state-carrying is reliable and the worker’s arithmetic, including modular reduction, is almost always right; the remaining failures are knocked down further by increasing worker samples. The two knobs that matter are: keep the thinking, and vote.
The skip.
A striking failure recurs at a fixed location: among the runs that get past round , several skip exactly rounds through , so the driver’s per-round tool calls jump straight from round to round and it finalizes with a short block. In our corpus this appears in long runs and, tellingly, in none of the successful ones. It is contiguous and position-locked rather than random: rounds to are the back half of the third () quarter, and a couple of runs skip rounds to as well (the back half of the quarter), so the model appears to “round off” the second half of a -round block. We read this as a candidate model fingerprint rather than a claim that it happens on every run.
For the same serverless weights, route changes latency by 5 (a gateway path 2.5 s per step vs. a direct path 13 s per step) and, historically, correctness (dedicated vs. shared; quantization / precision). We therefore control the endpoint per experimental cell and report speed only as context.
5.4 Failure analysis
Because the reference is exact, every failure has a precise location and cause. Table 2 is the consolidated taxonomy; the counterfactual injection (Sec. Verification) lets us confirm that each failed run has a single originating error whose corruption then propagates. Three points are worth drawing out. First, the errors partition cleanly by origin: driver (state), worker (arithmetic), or serving (context), and these are independent, so removing one class (e.g. voting out worker slips) does not remove another. Second, the driver errors dominate: the rotation slip and the dropped/duplicated step account for most failed runs, consistent with state tracking, not arithmetic, being the bottleneck. Third, one driver error is not random in where it lands. The skip recurs at the same location (and only in failing runs), a position-locked fingerprint rather than a stochastic slip, and a candidate diagnostic for the model family. By contrast the worker’s modular wrap-around is stochastic and self-limiting, the worker often repairs it mid-reasoning by comparing magnitudes, and majority voting removes almost all of what remains.
| Failure mode | Origin | Mechanism / effect |
|---|---|---|
| State slip (rotation) | Driver | mis-applies ; wrong from that round on |
| Dropped / duplicated step | Driver | skips or repeats a call; trace misaligns |
| Early stop / no digest | Driver | finalizes early (e.g. vs calls) or loops without a digest |
| skip (signature) | Driver | jumps rounds – at a fixed location (failing runs only); recurring fingerprint |
| Modular wrap-around | Worker | subtracts wrong multiple of ; one bad primitive, caught by vote / twin |
| Reasoning stripped | Serving | analysis channel not echoed; collapse within – steps |
6 Discussion
Practical reach.
Long-horizon exactness is not a toy concern. Accounting and bookkeeping close-out, legal and compliance workflows, provisioning and reconciliation pipelines, all require an intermediate state to be preserved verbatim across many steps, where a single dropped or mis-copied value silently corrupts the outcome. The prevailing assumption is that such workflows must be driven by a hand-built controller with the LLM confined to leaf tasks. Our result says the LLM can hold the state itself when its context is managed for it, reasoning kept in context, arithmetic made reliable by voting, which widens where a model may be trusted to drive, not just assist.
Toward a self-managing agent.
The driver/worker split is deliberately a two-agent decomposition: one agent owns the long-horizon state and sequencing, the other owns bounded, verifiable computation. We use the same model for both roles on purpose. The split is scaffolding, not the destination: because driver and worker are one model, the natural next step is to collapse them into a single agent that carries its own state, computes its own primitives, and, critically, curates its own context.
Self-editing context.
The most direct version of that next step is to give the model a tool that edits its own context: operations to write a value into a durable scratch region, read it back, summarize or compress the transcript so far, and drop material it no longer needs. In our task the state to be managed is tiny and exact, the four working words and the round index, so a self-editing policy has an unusually clean success criterion: we can check to the bit whether the model preserved the right state after every edit, rather than judging memory quality qualitatively. This connects to self-editing memory architectures such as MemGPT (Packer et al. 2023), which expose paging between in-context and external memory as tool calls; our contribution would be a ground-truthed testbed for whether such self-management actually keeps exact state across a long horizon. Would this be useful? We expect so: a model that can offload to a scratch tool and reload it on demand should be far more robust to the length-driven degradation of long transcripts, and the same mechanism generalizes to the branching, real-world procedures (accounting close-outs, compliance workflows) where the state that must survive is larger and messier than four words. Whether a single model can own its state and its computation and its context is the question this setup is designed to open.
7 Conclusion
Executing a deterministic algorithm step by step is a clean, bit-checkable probe of long-horizon reliability, and it delivers a positive answer: given the right context, a mixture-of-experts LLM with only 5.5B active parameters computes an MD5 digest across dependent tool calls, and, with a self-consistent worker replacing every primitive, does the whole thing, arithmetic included, from scratch. Success turns on two levers that are about context, not weights: keep the model’s reasoning in its context, and make the arithmetic reliable by voting. The residual failures are few, mechanical, and, in the skip, a reproducible signature of the model. The broader lesson runs opposite to the usual prescription: rather than assume long, exact workflows must be pulled out of the model and handed to a hand-built controller, give the model the context it needs and it can drive them, a foundation for agents that eventually manage that context themselves.
Acknowledgments
We thank Google Cloud for research grant support, which provided the GPU access used to run the experiments reported in this paper.
References
- Adhikary, Contractor et al. (2025) Adhikary, B.; Contractor, D.; et al. 2025. LongFuncEval: Measuring the Effectiveness of Long Context Models for Function Calling. arXiv:2505.10570.
- Lee et al. (2024) Lee, N.; Sreenivasan, K.; Lee, J. D.; Lee, K.; and Papailiopoulos, D. 2024. Teaching Arithmetic to Small Transformers. arXiv:2307.03381.
- Liu et al. (2024) Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12: 157–173.
- Modarressi et al. (2025) Modarressi, A.; Deilamsalehy, H.; Dernoncourt, F.; Bui, T.; Rossi, R. A.; Yoon, S.; and Schütze, H. 2025. NoLiMa: Long-Context Evaluation Beyond Literal Matching. arXiv:2502.05167.
- Nye et al. (2022) Nye, M.; Andreassen, A.; Gur-Ari, G.; Michalewski, H.; Austin, J.; Bieber, D.; Dohan, D.; Lewkowycz, A.; Bosma, M.; Luan, D.; Sutton, C.; and Odena, A. 2022. Show Your Work: Scratchpads for Intermediate Computation with Language Models. In Deep Learning for Code Workshop, ICLR. ArXiv:2112.00114.
- OpenAI (2025a) OpenAI. 2025a. gpt-oss-120b and gpt-oss-20b Model Card. arXiv:2508.10925.
- OpenAI (2025b) OpenAI. 2025b. OpenAI Harmony Response Format. https://github.com/openai/harmony. Accessed: 2026-07-29.
- Packer et al. (2023) Packer, C.; Wooders, S.; Lin, K.; Fang, V.; Patil, S. G.; Stoica, I.; and Gonzalez, J. E. 2023. MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560.
- Rivest (1992) Rivest, R. L. 1992. The MD5 Message-Digest Algorithm. Request for Comments RFC 1321, Internet Engineering Task Force.
- Schick et al. (2023) Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. In Advances in Neural Information Processing Systems (NeurIPS).
- Sinha et al. (2025) Sinha, A.; Arun, A.; Goel, S.; Staab, S.; and Geiping, J. 2025. The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs. arXiv:2509.09677.
- Wang et al. (2023) Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In International Conference on Learning Representations (ICLR).
- Wei et al. (2022) Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; and Zhou, D. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems (NeurIPS).
- Yao et al. (2023) Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In International Conference on Learning Representations (ICLR).