Skip to content

Repository files navigation

Self-Reflective APIs

Reproduction package for the paper "Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery" (arXiv:2606.05037).

We test whether returning structured, context-dependent recovery hints in API error responses improves LLM-agent task completion compared to (a) traditional terse errors and (b) verbose plain-English errors. Two domains: a recipe-conversion API (main study) and an Acme-billing API backed by Stripe sandbox (§5.5.3 Domain Portability replication).

Headline result

Post-leakage-audit, N=30 per (model, mode) cell:

Model Reflective vs verbose Reflective vs traditional
claude-haiku-4-5 +36.7pp (p=0.0011) +86.7pp (p=2.4e-12)
claude-sonnet-4-6 +40.0pp (p=0.0022) +70.0pp (p=7.0e-8)
gpt-4o-mini +13.3pp (n.s., p=0.435) +43.3pp (p=0.0014)

The non-significant gpt-4o-mini verbose-vs-reflective cell is reported honestly; see paper §5.5.2 External Validity.

Repo layout

apis/
  recipe/            # main-study domain (FastAPI)
  acme_billing/      # §5.5.3 portability replication (FastAPI + Stripe sandbox)
experiments/
  run_adversarial_experiment.py   # sweep driver
  audit_prompt_leakage.py         # GATE — run before any sweep
  compute_paper_stats.py          # significance tests, effect sizes
  generate_paper_report.py        # tables/markdown for the paper
  agents/                         # base / simple / langchain agents
  tasks/                          # task libraries (JSON) + loader
  instrumentation/                # token/cost tracking, metrics
  config/                         # settings + .env.example
results/
  paper_runs/                     # 21 JSON sweep outputs backing the paper
analysis/                         # figure-generation scripts
paper/                            # LaTeX source + compiled PDF + figures

Quickstart

python -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install -r experiments/requirements.txt
cp experiments/config/.env.example experiments/config/.env
# edit .env with ANTHROPIC_API_KEY, OPENAI_API_KEY, STRIPE_API_KEY (test key only)

Reproducing the paper

  1. Audit the task library for answer leakage (required gate):

    python experiments/audit_prompt_leakage.py

    This is the methodological fix that produced the headline numbers above. Do not skip.

  2. Run the sweep (~30 minutes; spend depends on provider pricing — budget a few USD for the recipe sweep):

    python experiments/run_adversarial_experiment.py \
        --tasks 10 \
        --task-file task_library_adversarial_enhanced.json \
        --models claude-haiku-4-5,claude-sonnet-4-6,gpt-4o-mini \
        --modes traditional,verbose,reflective \
        --n 30
  3. Compute statistics:

    python experiments/compute_paper_stats.py results/
  4. Regenerate figures (consume results/paper_runs/):

    python analysis/make_figures.py
  5. Build the paper:

    cd paper && pdflatex main.tex && pdflatex main.tex

Acme billing replication

The Acme API is a thin policy layer on top of Stripe test mode. It uses synthetic approval tokens (acme_appr_*) that look like secrets but are not. To run:

python experiments/setup_acme_refund_fixtures.py   # one-time, creates test PaymentIntents
python experiments/run_adversarial_experiment.py --domain acme --n 30

A Stripe test key (sk_test_*) is required; the code refuses to run with a live key.

Citation

@article{canedo2026selfreflective,
  title  = {Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery},
  author = {Canedo, Arquimedes and Grama, Chethan},
  year   = {2026},
  eprint = {2606.05037},
  archivePrefix = {arXiv}
}

License

License pending. This repository is published as a reference implementation accompanying the paper. A patent application covering the underlying technique has been filed by Siemens. The code is not yet licensed for redistribution or production use. For licensing inquiries, contact the corresponding author.

About

Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages