Reproduction package for the paper "Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery" (arXiv:2606.05037).
We test whether returning structured, context-dependent recovery hints in API error responses improves LLM-agent task completion compared to (a) traditional terse errors and (b) verbose plain-English errors. Two domains: a recipe-conversion API (main study) and an Acme-billing API backed by Stripe sandbox (§5.5.3 Domain Portability replication).
Post-leakage-audit, N=30 per (model, mode) cell:
| Model | Reflective vs verbose | Reflective vs traditional |
|---|---|---|
| claude-haiku-4-5 | +36.7pp (p=0.0011) | +86.7pp (p=2.4e-12) |
| claude-sonnet-4-6 | +40.0pp (p=0.0022) | +70.0pp (p=7.0e-8) |
| gpt-4o-mini | +13.3pp (n.s., p=0.435) | +43.3pp (p=0.0014) |
The non-significant gpt-4o-mini verbose-vs-reflective cell is reported honestly; see paper §5.5.2 External Validity.
apis/
recipe/ # main-study domain (FastAPI)
acme_billing/ # §5.5.3 portability replication (FastAPI + Stripe sandbox)
experiments/
run_adversarial_experiment.py # sweep driver
audit_prompt_leakage.py # GATE — run before any sweep
compute_paper_stats.py # significance tests, effect sizes
generate_paper_report.py # tables/markdown for the paper
agents/ # base / simple / langchain agents
tasks/ # task libraries (JSON) + loader
instrumentation/ # token/cost tracking, metrics
config/ # settings + .env.example
results/
paper_runs/ # 21 JSON sweep outputs backing the paper
analysis/ # figure-generation scripts
paper/ # LaTeX source + compiled PDF + figures
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r experiments/requirements.txt
cp experiments/config/.env.example experiments/config/.env
# edit .env with ANTHROPIC_API_KEY, OPENAI_API_KEY, STRIPE_API_KEY (test key only)-
Audit the task library for answer leakage (required gate):
python experiments/audit_prompt_leakage.py
This is the methodological fix that produced the headline numbers above. Do not skip.
-
Run the sweep (~30 minutes; spend depends on provider pricing — budget a few USD for the recipe sweep):
python experiments/run_adversarial_experiment.py \ --tasks 10 \ --task-file task_library_adversarial_enhanced.json \ --models claude-haiku-4-5,claude-sonnet-4-6,gpt-4o-mini \ --modes traditional,verbose,reflective \ --n 30 -
Compute statistics:
python experiments/compute_paper_stats.py results/
-
Regenerate figures (consume
results/paper_runs/):python analysis/make_figures.py
-
Build the paper:
cd paper && pdflatex main.tex && pdflatex main.tex
The Acme API is a thin policy layer on top of Stripe test mode. It uses synthetic approval tokens (acme_appr_*) that look like secrets but are not. To run:
python experiments/setup_acme_refund_fixtures.py # one-time, creates test PaymentIntents
python experiments/run_adversarial_experiment.py --domain acme --n 30A Stripe test key (sk_test_*) is required; the code refuses to run with a live key.
@article{canedo2026selfreflective,
title = {Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery},
author = {Canedo, Arquimedes and Grama, Chethan},
year = {2026},
eprint = {2606.05037},
archivePrefix = {arXiv}
}License pending. This repository is published as a reference implementation accompanying the paper. A patent application covering the underlying technique has been filed by Siemens. The code is not yet licensed for redistribution or production use. For licensing inquiries, contact the corresponding author.