Measuring silent failure in production APIs used as agent tools.
An LLM agent calling a third-party API cannot tell these two situations apart:
- the query was well formed and no records match it, and
- the query was not understood, the server discarded part of it, and the response describes a different question.
Both arrive as HTTP 200 with a parsable body. There is no exception to catch, no status code to branch on, and no field in the response that marks the difference. This repository holds the code, data and results for a study of what predicts which one occurred, and what it does to the agent.
A parameter whose controlled vocabulary is only exemplified rather than stated is missed by every model tested, silently, and writing the vocabulary into the schema fixes it.
Two separable properties govern this:
- Disclosure decides whether the model can comply. A vocabulary written out
in prose is used correctly 88 to 91% of the time, with no schema support at
all. One shown by example (
"e.g. 'executive'", 1 of 18 real values) is used correctly 0% of the time, across twelve models. - Machine-readability decides whether a wrong value is caught. Only a schema constraint lets a generic validator reject the call and name the legal values: 0 of 111 machine-checkable perturbations failed silently, against 44 of 61 prose-only ones (Fisher exact, p = 2×10⁻¹³).
Downstream, in full agent loops: models detected the resulting silent failure 12% of the time, repaired it 0% of the time, asserted a false negative to the user in 41% of cases, and invented a figure in 12%.
The counterintuitive practical rule: e.g. is more dangerous than an
incomplete closed list. An example invites a model to invent values; a
closed-list phrasing keeps it inside the documented set even when that set is
incomplete.
The study uses two aggregation layers, one on each axis it needs to vary. An aggregation layer sits between a caller and many independent providers, exposing them through a single contract, a single credential and a uniform request format.
Monid — the OpenRouter for agent tools
Monid aggregates data and tool endpoints: web search, scraping, people and company enrichment, social and commercial data, and similar services from many independent vendors, reached with one key and priced per call. The endpoints under test sit behind it.
Four of its properties are load-bearing for this study, and a replication needs an equivalent for each:
| Property | Why the method needs it |
|---|---|
| Schemas in one format, retrievable without charge or auth | Makes the static audit free, and lets one classifier run over every vendor instead of one parser per vendor |
| A run identifier, price and timing on every execution | Each of the 219 perturbations can be re-fetched individually, so any single measurement can be checked rather than trusted |
| One contract and one key across 27 vendors | Otherwise this is a procurement exercise before it is a research one |
| A single schema validator in front of every vendor | Gateway rejections and vendor behaviour become separable, which turns the intermediary from a confound into the mechanism |
OpenRouter — the same idea, for models
OpenRouter aggregates language models: models from many providers behind one key and one OpenAI-compatible request format. The twelve models evaluated here are reached through it.
The symmetry is why the design is feasible. Comparing failure behaviour across vendors requires that vendors be callable uniformly; comparing model behaviour across families requires that models be callable uniformly. Without the first, the perturbation study is a procurement exercise; without the second, the agent study is twelve separate integrations.
probe/ the measurement code
monid.py CLI wrapper; records runId, price and timing per call
pool.py build the endpoint pool (free)
fetch_schemas.py fetch every published JSON Schema (free)
gap_audit.py find constraints stated in prose but absent from schema
corpus_audit.py the same audit over 2,501 public OpenAPI documents
perturb.py deterministic, schema-derived perturbation families
experiment.py the triple probe: signatures, tolerance, classification
run_main.py the measurement campaign
rq4_multi.py agent experiment, twelve models, AS_IS vs LIFTED
rq4_cross.py cross-vendor replication
rq6_downstream.py full agent loop: what the user is told after a silent zero
hidden_vocab.py probe free-text params for undisclosed vocabularies
report*.py every table in the paper
data/ schemas, perturbation records, agent transcripts, judge labels
out/ generated reports (FINAL-*.txt)
paper/ LaTeX source, figures and build script
python probe/pool.py # free: discovery
python probe/fetch_schemas.py # free: schema retrieval
python probe/gap_audit.py # free: the prevalence numbers
python probe/corpus_audit.py # free: the public-corpus replication
python probe/select_endpoints.py 70
SF_BUDGET=9 SF_WORKERS=3 python probe/run_main.py # needs a Monid key
SF_SLIM=1 SF_THROTTLE=4 python probe/rq4_multi.py # needs OPENROUTER_API_KEY
python probe/rq6_downstream.py
python probe/report.py && python probe/report_multi.pyThe whole campaign cost under US$8: US$4.67 of data-API spend and US$1.89 of model inference. Discovery and schema inspection are free, so the static audit over 721,320 parameters costs nothing at all.
Caching matters at this scale: models converge on identical arguments, so 815 agent tool calls reduced to 44 distinct API calls and 88 executions instead of 2,445.
One vendor throttles by returning a bare HTTP 400, not 429. Status alone
cannot separate throttling from a genuine parameter rejection, and scoring an
unconfirmed 4xx converts your own request rate into a finding about the endpoint.
Pace at 4 seconds and confirm every 4xx by repeating it: a real rejection is
deterministic, throttling is not.
Expose few parameters to the model. With a full 20-parameter schema visible, some models fill fields they cannot know, including the pagination cursor. Two of twelve did so in 100% of their calls; five did so in none. The call then fails, and the parameter under test becomes unobservable.
- The executed measurements cover endpoints reachable through one aggregation layer, read-only, priced at or below $0.02 per call. The public-corpus audit covers the static claim; the execution claim does not have that cover.
- The parameter-inert exclusion cannot distinguish "this filter does not discriminate on this query" from "this filter is always discarded", so the reported silent rate is a lower bound.
- Semantic downgrade, the fourth failure mode, is not detectable by this method at all. It requires comparing a result against the caller's intent, which no mechanical oracle has.
- Every recovered vocabulary is a lower bound: we learn only about terms we thought to try.
paper/ builds the write-up with ./build.sh, which needs a LaTeX toolchain.
Code under MIT. Data released for research use.