Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LLM Transport Layer Benchmarks

Benchmarks comparing transport protocols for LLM API interactions, with a focus on agentic coding workflows where multi-turn tool-calling conversations amplify protocol overhead.

Agentic Benchmark (agentic-benchmark/)

Compares HTTP vs WebSocket for multi-turn agentic coding workflows using OpenAI's Responses API. The key difference: HTTP must resend the full conversation context each turn (stateless), while WebSocket uses previous_response_id with server-side in-memory state (stateful continuation).

Key Results

Model Approach Avg bytes sent/task Avg time/task vs HTTP
GPT-5.4 HTTP 176 KB 40.8 s —
GPT-5.4 WebSocket 32 KB 28.9 s 82% less data, 29% faster
GPT-4o-mini HTTP 153 KB 53.6 s —
GPT-4o-mini WebSocket 21 KB 45.5 s 86% less data, 15% faster

How It Works

Three simulated coding tasks (fix a test, add a feature, refactor an API) drive a multi-turn tool-calling loop against the OpenAI Responses API. The model makes real API calls and decides which tools to call; tool responses are simulated with realistic file contents and command outputs.

  • Cell 1 (HTTP): Each turn resends the full conversation history as a new POST request
  • Cell 2 (WebSocket): Each turn sends only previous_response_id + new tool outputs over a persistent connection

Prerequisites

  • Python 3.12+
  • OpenAI API key with access to the Responses API

Setup

python3 -m venv .venv
source .venv/bin/activate
pip install -r agentic-benchmark/requirements.txt
export OPENAI_API_KEY=sk-...

Running

# Quick run (1 iteration, both cells)
python3 agentic-benchmark/benchmark.py --model gpt-4o-mini --runs 1 --cells 1,2

# Full run (3 iterations for statistical averaging)
python3 agentic-benchmark/benchmark.py --model gpt-5.4 --runs 3 --cells 1,2

# With network shaping (macOS, requires sudo)
sudo ./agentic-benchmark/benchmark.sh --model gpt-5.4 --runs 1

Results are saved as JSON to agentic-benchmark/results-<timestamp>.json.

Options

Flag Default Description
--model gpt-4o-mini OpenAI model to use
--runs 1 Number of runs per approach
--max-turns 15 Max turns per task
--cells 1,2,3,4 Cells to run (1=HTTP, 2=WS, 3=WS-parallel, 4=HTTP-parallel)
--output auto Output JSON file path

WebTransport Benchmark (webtransport-benchmark/)

A separate Go-based benchmark comparing per-token wire efficiency across three response streaming approaches using a local LLM (Ollama):

Approach Wire bytes/token vs OpenAI-compatible
HTTP SSE (OpenAI-compatible JSON) 372 bytes/token —
HTTP SSE (tokens only) 194 bytes/token 48% less
WebTransport (tokens only, binary framing) 132 bytes/token 65% less

See webtransport-benchmark/ for Go server/client code and the network-conditioned benchmark runner.

Prerequisites

  • Go 1.25+
  • Ollama running locally with gemma3:12b

Running

cd webtransport-benchmark
./generate_cert.sh

# Start both servers
go run ./server &       # WebTransport (port 4433)
go run ./httpserver &   # HTTP SSE (port 8080)

# Run benchmark
go run ./benchmark -reuse

# Network-conditioned (macOS, requires sudo)
./benchmark/benchmark.sh

About

Benchmark for comparing HTTP vs WebSocket for agentic coding workflows

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages