Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SilentProbe

Measuring silent failure in production APIs used as agent tools.

An LLM agent calling a third-party API cannot tell these two situations apart:

  1. the query was well formed and no records match it, and
  2. the query was not understood, the server discarded part of it, and the response describes a different question.

Both arrive as HTTP 200 with a parsable body. There is no exception to catch, no status code to branch on, and no field in the response that marks the difference. This repository holds the code, data and results for a study of what predicts which one occurred, and what it does to the agent.


The finding in one line

A parameter whose controlled vocabulary is only exemplified rather than stated is missed by every model tested, silently, and writing the vocabulary into the schema fixes it.

Two separable properties govern this:

  • Disclosure decides whether the model can comply. A vocabulary written out in prose is used correctly 88 to 91% of the time, with no schema support at all. One shown by example ("e.g. 'executive'", 1 of 18 real values) is used correctly 0% of the time, across twelve models.
  • Machine-readability decides whether a wrong value is caught. Only a schema constraint lets a generic validator reject the call and name the legal values: 0 of 111 machine-checkable perturbations failed silently, against 44 of 61 prose-only ones (Fisher exact, p = 2×10⁻¹³).

Downstream, in full agent loops: models detected the resulting silent failure 12% of the time, repaired it 0% of the time, asserted a false negative to the user in 41% of cases, and invented a figure in 12%.

The counterintuitive practical rule: e.g. is more dangerous than an incomplete closed list. An example invites a model to invent values; a closed-list phrasing keeps it inside the documented set even when that set is incomplete.


Instruments

The study uses two aggregation layers, one on each axis it needs to vary. An aggregation layer sits between a caller and many independent providers, exposing them through a single contract, a single credential and a uniform request format.

Monid — the OpenRouter for agent tools

Monid aggregates data and tool endpoints: web search, scraping, people and company enrichment, social and commercial data, and similar services from many independent vendors, reached with one key and priced per call. The endpoints under test sit behind it.

Four of its properties are load-bearing for this study, and a replication needs an equivalent for each:

Property Why the method needs it
Schemas in one format, retrievable without charge or auth Makes the static audit free, and lets one classifier run over every vendor instead of one parser per vendor
A run identifier, price and timing on every execution Each of the 219 perturbations can be re-fetched individually, so any single measurement can be checked rather than trusted
One contract and one key across 27 vendors Otherwise this is a procurement exercise before it is a research one
A single schema validator in front of every vendor Gateway rejections and vendor behaviour become separable, which turns the intermediary from a confound into the mechanism

OpenRouter — the same idea, for models

OpenRouter aggregates language models: models from many providers behind one key and one OpenAI-compatible request format. The twelve models evaluated here are reached through it.

The symmetry is why the design is feasible. Comparing failure behaviour across vendors requires that vendors be callable uniformly; comparing model behaviour across families requires that models be callable uniformly. Without the first, the perturbation study is a procurement exercise; without the second, the agent study is twelve separate integrations.


What is here

probe/          the measurement code
  monid.py            CLI wrapper; records runId, price and timing per call
  pool.py             build the endpoint pool                 (free)
  fetch_schemas.py    fetch every published JSON Schema       (free)
  gap_audit.py        find constraints stated in prose but absent from schema
  corpus_audit.py     the same audit over 2,501 public OpenAPI documents
  perturb.py          deterministic, schema-derived perturbation families
  experiment.py       the triple probe: signatures, tolerance, classification
  run_main.py         the measurement campaign
  rq4_multi.py        agent experiment, twelve models, AS_IS vs LIFTED
  rq4_cross.py        cross-vendor replication
  rq6_downstream.py   full agent loop: what the user is told after a silent zero
  hidden_vocab.py     probe free-text params for undisclosed vocabularies
  report*.py          every table in the paper
data/           schemas, perturbation records, agent transcripts, judge labels
out/            generated reports (FINAL-*.txt)
paper/          LaTeX source, figures and build script

Reproducing

python probe/pool.py            # free: discovery
python probe/fetch_schemas.py   # free: schema retrieval
python probe/gap_audit.py       # free: the prevalence numbers
python probe/corpus_audit.py    # free: the public-corpus replication

python probe/select_endpoints.py 70
SF_BUDGET=9 SF_WORKERS=3 python probe/run_main.py     # needs a Monid key
SF_SLIM=1 SF_THROTTLE=4 python probe/rq4_multi.py     # needs OPENROUTER_API_KEY
python probe/rq6_downstream.py

python probe/report.py && python probe/report_multi.py

The whole campaign cost under US$8: US$4.67 of data-API spend and US$1.89 of model inference. Discovery and schema inspection are free, so the static audit over 721,320 parameters costs nothing at all.

Caching matters at this scale: models converge on identical arguments, so 815 agent tool calls reduced to 44 distinct API calls and 88 executions instead of 2,445.

Two hazards worth knowing before you replicate

One vendor throttles by returning a bare HTTP 400, not 429. Status alone cannot separate throttling from a genuine parameter rejection, and scoring an unconfirmed 4xx converts your own request rate into a finding about the endpoint. Pace at 4 seconds and confirm every 4xx by repeating it: a real rejection is deterministic, throttling is not.

Expose few parameters to the model. With a full 20-parameter schema visible, some models fill fields they cannot know, including the pagination cursor. Two of twelve did so in 100% of their calls; five did so in none. The call then fails, and the parameter under test becomes unobservable.

Caveats

  • The executed measurements cover endpoints reachable through one aggregation layer, read-only, priced at or below $0.02 per call. The public-corpus audit covers the static claim; the execution claim does not have that cover.
  • The parameter-inert exclusion cannot distinguish "this filter does not discriminate on this query" from "this filter is always discarded", so the reported silent rate is a lower bound.
  • Semantic downgrade, the fourth failure mode, is not detectable by this method at all. It requires comparing a result against the caller's intent, which no mechanical oracle has.
  • Every recovered vocabulary is a lower bound: we learn only about terms we thought to try.

Paper

paper/ builds the write-up with ./build.sh, which needs a LaTeX toolchain.

License

Code under MIT. Data released for research use.

About

Measuring silent failure in production APIs used as agent tools. Calls that return HTTP 200 with a parsable body and a wrong answer.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages