Benchmarks comparing transport protocols for LLM API interactions, with a focus on agentic coding workflows where multi-turn tool-calling conversations amplify protocol overhead.
Compares HTTP vs WebSocket for multi-turn agentic coding workflows using OpenAI's Responses API. The key difference: HTTP must resend the full conversation context each turn (stateless), while WebSocket uses previous_response_id with server-side in-memory state (stateful continuation).
| Model | Approach | Avg bytes sent/task | Avg time/task | vs HTTP |
|---|---|---|---|---|
| GPT-5.4 | HTTP | 176 KB | 40.8 s | — |
| GPT-5.4 | WebSocket | 32 KB | 28.9 s | 82% less data, 29% faster |
| GPT-4o-mini | HTTP | 153 KB | 53.6 s | — |
| GPT-4o-mini | WebSocket | 21 KB | 45.5 s | 86% less data, 15% faster |
Three simulated coding tasks (fix a test, add a feature, refactor an API) drive a multi-turn tool-calling loop against the OpenAI Responses API. The model makes real API calls and decides which tools to call; tool responses are simulated with realistic file contents and command outputs.
- Cell 1 (HTTP): Each turn resends the full conversation history as a new POST request
- Cell 2 (WebSocket): Each turn sends only
previous_response_id+ new tool outputs over a persistent connection
- Python 3.12+
- OpenAI API key with access to the Responses API
python3 -m venv .venv
source .venv/bin/activate
pip install -r agentic-benchmark/requirements.txt
export OPENAI_API_KEY=sk-...# Quick run (1 iteration, both cells)
python3 agentic-benchmark/benchmark.py --model gpt-4o-mini --runs 1 --cells 1,2
# Full run (3 iterations for statistical averaging)
python3 agentic-benchmark/benchmark.py --model gpt-5.4 --runs 3 --cells 1,2
# With network shaping (macOS, requires sudo)
sudo ./agentic-benchmark/benchmark.sh --model gpt-5.4 --runs 1Results are saved as JSON to agentic-benchmark/results-<timestamp>.json.
| Flag | Default | Description |
|---|---|---|
--model |
gpt-4o-mini |
OpenAI model to use |
--runs |
1 |
Number of runs per approach |
--max-turns |
15 |
Max turns per task |
--cells |
1,2,3,4 |
Cells to run (1=HTTP, 2=WS, 3=WS-parallel, 4=HTTP-parallel) |
--output |
auto | Output JSON file path |
A separate Go-based benchmark comparing per-token wire efficiency across three response streaming approaches using a local LLM (Ollama):
| Approach | Wire bytes/token | vs OpenAI-compatible |
|---|---|---|
| HTTP SSE (OpenAI-compatible JSON) | 372 bytes/token | — |
| HTTP SSE (tokens only) | 194 bytes/token | 48% less |
| WebTransport (tokens only, binary framing) | 132 bytes/token | 65% less |
See webtransport-benchmark/ for Go server/client code and the network-conditioned benchmark runner.
- Go 1.25+
- Ollama running locally with
gemma3:12b
cd webtransport-benchmark
./generate_cert.sh
# Start both servers
go run ./server & # WebTransport (port 4433)
go run ./httpserver & # HTTP SSE (port 8080)
# Run benchmark
go run ./benchmark -reuse
# Network-conditioned (macOS, requires sudo)
./benchmark/benchmark.sh