Deep-TOON

The cost of tokens can be showstoppers for Agentic applications. Lately the TOON – Token-Oriented Object Notation get some buzz as it’s “token compressing” JSON data which is one of the most popular dataformat used with LLMs and AI Agents.

Inspired by the original toon-format I’ve created my own flavor deep-toon which is better able to compress embedded structures too which are quite common in real life examples.

What is Deep-TOON?

Deep-TOON is a compact, token-optimized representation format for hierarchical data (nested objects, arrays, mixed types) that is particularly suited for large-language-model (LLM) and AI use-cases. It supports perfect round-trip fidelity (so you can encode then decode back to the original structure – NOTE order can change) and typically yields significant token savings compared to raw JSON. PyPI+1
It’s implemented in Python and available via pip install deep-toon. PyPI+1

Why use Deep-TOON?

  • Reduced token usage – In one example, a JSON blob of 1,675 tokens was compressed to 1,065 tokens (≈ 36 % reduction) using Deep-TOON. GitHub+1
  • Better for nested / repeated structures – When your data includes arrays of objects, deep hierarchies (e.g., metadata with nested sub-objects), Deep-TOON shines. PyPI+1
  • LLM-friendly and prompt budget conscious – Because many AI/LLM systems charge (directly or indirectly) by token usage, formats that reduce token load yet preserve semantics are valuable.
  • Perfect fidelity – The original structure and content are preserved when decoding, so you aren’t trading off correctness for compression. PyPI+1
  • Simple integration – It works in Python (requires Python 3.8+). PyPI

Key features & format at a glance

  • Schema declaration up-front – Deep-TOON explicitly declares the structure once (fields, nested objects) then provides values in a compact form. PyPI+1
  • Tuple-grouping of related fields – Nested objects become tuples in the encoded form, helping reduce repetition and flatten value lists. GitHub
  • Lists/arrays optimized when schema is consistent – Repeated object-type arrays benefit most in terms of token reduction. PyPI+1
  • Graceful fallback – If the array schema is inconsistent (objects differ in fields), Deep-TOON will revert to the JSON style fallback rather than risk incorrect encoding. GitHub
  • Custom delimiter support – If your strings include commas or other characters that would conflict with the default delimiter, Deep-TOON allows customizing the delimiter for safe encoding. PyPI

Here’s an example:
Original JSON:

{
"users": [
{
"id": 1,
"firstName": "Emily",
"lastName": "Johnson",
"age": 28,
"address": {
"address": "626 Main Street",
"city": "Phoenix",
"state": "Mississippi",
"coordinates": {"lat": -77.16213, "lng": -92.084824}
},
"bank": {
"cardNumber": "9289760655481815",
"cardType": "Elo"
}
}
],
"total": 208,
"skip": 0,
"limit": 3
}

Encoded with Deep-TOON:

users[1,]{id,firstName,lastName,age,address{address,city,state,coordinates{lat,lng}},bank{cardNumber,cardType}}:
1,Emily,Johnson,28,("626 Main Street",Phoenix,Mississippi,(-77.16213,-92.084824)),("9289760655481815",Elo)
total: 208
skip: 0
limit: 3

NOTE: the example above do not demonstrate the compression just the data format. Token saving would happen in the case if it would be multiple elements in the “users” array.

Getting started

Installation:

pip install deep-toon

Basic usage:

import deep_toon

data = {
"users": [
{
"id": 1,
"name": "Alice",
"address": {
"street": "123 Main St",
"city": "NYC",
"coordinates": {"lat": 40.7, "lng": -74.0}
}
},
{
"id": 2,
"name": "Bob",
"address": {
"street": "456 Oak Ave",
"city": "LA",
"coordinates": {"lat": 34.0, "lng": -118.2}
}
}
]
}

compressed = deep_toon.encode(data)
print("Compressed:", compressed)

restore = deep_toon.decode(compressed)
print("Round-trip equals original?", restore == data)

Evaluation

Compression

Compression (how much we can reduce the tokens in a given JSON) varies, highly dependenton the structure of the JSON. In ideal cases we can see up to ~50% toke reduction, while in some corner cases the converted data might have more tokens.

To avoid the negative effect use the smaert encoder that targets a given saving percentage and if not accomplihs it returns the original JSON.

from deep_toon import smart_encode

# Only use Deep-TOON if it saves > 10% tokens
# Otherwise returns minified JSON
encoded = smart_encode(data, threshold=0.1)

Based on a 20 test case evaluation real life savings on the input tokens:

Test Cost (Queries):      $0.1204
  - JSON Cost:            $0.0642
  - Deep-TOON Cost:       $0.0562
  - Net Savings:          $0.0080 (12.5%)

LLM comprehension

While token savings are important, we also want to guarantee that the LLM can still able to work with the toon formatted data. To test this I’ve built an LLM comprehension test that use realistic JSON examples and perform data retreival from the test data both with the original JSON and the deep-toon representation then uses LLM as judge to compare the two results and also compare them with the ground truth. Model used: GPT5-mini.

Example question:

  Q2: Look at all users and find the one with the longest 'firstName' (most characters). If multiple users have the same longest length, provide the first one found. Provide only that person's exact 'firstName'.
     🎯 Expected: Alexander
     JSON: ✅ | Deep-TOON: ✅ | Equivalent: ✅
     💾 Deep-TOON savings: 6804 tokens
     💰 Cost: $0.01054 (Savings: $0.00157)
     📝 JSON: Alexander...
     📝 Deep-TOON: Alexander...
     🔍 Notes: Responses are identical...
     🔍 API calls used: 4/150

Result:

Total questions tested: 20
JSON Accuracy:      18/20 (90.0%)
Deep-TOON Accuracy: 19/20 (95.0%)
Equivalence Rate:   19/20 (95.0%)

Based on multiple executions of the test showed minimal variations but the results above I found representative. The result shows comparable performance.