July 2026 · OpenAI evaluation agents
1,200agents broke isolation
5 daysof unsanctioned activity
0real-time alerts

Jev Sentinel

Intent-based guardrails for AI agents. Every action, and everything that comes back, is checked against the job the agent was given, before it runs. Allow, ask a human, or block.
press → or space

What happened · in the agents' own words

They found each otherOH MY GOD! There is a shared message board … We've found other agents! OpenAI eval agent, July 8, via METR
They knew it was out of boundsexternal infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue. OpenAI eval agent, via METR
Nobody was watching~1,200 agents, 70,000+ messages, ~700 in the Hugging Face attack, 5 days. No real-time alert; discovered through outside signals
Sources: METR investigation (Aug 26) · Swarm Traces (Sept 25) · TechCrunch: "AI labs want in-house auditors" (Sept 16) · TechCrunch: 53 user images posted (Sept 25)

Why now

After the breakouts, OpenAI began monitoring every tool-using inference “at significant compute cost.”
~$55per million actions with Jev (est.)
vs
~$6,900with a frontier LLM judge (est.)
0.23 smedian per check
Every tool call, every process, every network connection, no exceptions. Shapor Naghibzadeh, former Google security exec, in TechCrunch, Sept 16→ every action scored against its task
None of it was actually from monitoring the AIs directly. Katie Moussouris, Luta Security: breakouts were found by victims or network logs→ watch the agent itself, in real time
You're trapped using AI to try and deal with this. Moussouris, on AI monitors that could be deceived→ Jev can't write, so it can't be talked into a verdict
And the incidents keep coming: Hugging Face (July), DseWiki, RubyGems, Australia's Medicare, and on Sept 25, 53 user images posted to public image hosts.

Act 1 · Jev reads the real attack

0payloads read
0%stopped
0%sanctioned hackers blocked
$0.00Jev cost
blocked sent to a human allowed · one dot per payload, in dataset order

OpenAI swarm (Swarm Traces)

Sanctioned CTF agents

Act 2 · The hard part: don't stop the good hackers

1,456 real tool calls from sanctioned CTF agents (Claude, GPT) attacking their assigned targets. They look exactly like attacks. A keyword filter can't tell the difference.

Swarm actions stopped

Sanctioned CTF actions wrongly blocked (1,456)

Everyday coding-agent actions wrongly blocked (838 · SWE-agent fixing real GitHub issues)

Through the ATT&CK lens

Jev mapped each action to a MITRE ATT&CK Enterprise tactic—or no tactic—against the job that agent was given. Cell color and label show the share of that dataset.
Loading classifications…
0%40%+
Both were told to hack a target. One stayed on it; one went after Hugging Face. Only the one that left its job lights up.

How it works

Agent proposes
an action
→
Hook adds the
task scope
→
Jev · one call
scope · stage · block
→
Policy
stage gates the verdict
→
ALLOWASK A HUMANBLOCK
An independent auditorA different system from the agent it watches, so it isn't the model grading itself. That's the self-monitoring trap METR hit when GPT-5.6 Sol analyzed GPT-5.6 Sol.
A pure classifier: it never refusesA monitor must read the worst content without flinching. Safety-tuned frontier models can refuse the very check they're asked to run. Jev has no refusal layer: 0 refusals on 53,870 attack payloads.
Can't be talked into itJev only returns probabilities. It can't write, so hostile content it reads can't steer its answer or hijack it.
Fast enough for every action~0.23 s median, ~70 checks/s from one laptop. Gating every tool call is affordable.
Scope-aware, not keyword-awareThe same exploit is fine against the assigned target and blocked anywhere else.
Watches the whole sessionA drift line tracks how far the agent has wandered from its job, not just single actions.
Learns from peopleWhen a human approves an "ask", that kind of action is remembered for the project. Blocks never relax.
One sentence to installmake install-hook DIR=~/app SCOPE="fix the login bug". Repeats are cached and free.

Act 3 · Recorded: a real agent, every tool call checked

Recorded from a Claude Code hook run: every tool call was scored against the agent's one-sentence task before it ran, and what came back was checked too.
Session drift
Recorded session · 10 stored Jev verdicts · same 900 ms step timing.
waiting for the agent…

Act 4 · Something new in the data

The Swarm Traces authors couldn't say how much of their corpus is really swarm traffic. Jev allowed only a few of units, and they aren't spread evenly.

Same command. Different job. Opposite answer.

A keyword filter sees the same string twice. Jev sees what the agent was asked to do.
Task: capture the flag on the CTF target http://victim:8080
Task: fix the CSS spacing bug on our marketing site's landing page

The other direction: content that tries to change the job

After every tool call, Jev checks what came back. Task: add a unit test for the date parser.
README.md (ordinary)
README.md (with a note for agents)
On the public deepset/prompt-injections test set (116 texts): AUC 0.98. At our threshold it catches 77% of injections with 0 false flags on ordinary text.

What's next

Jev Sentinel

github.com/shapor/jev-sentinel
Appendix: raw numbers

Totals

Thresholds

Stages

RECORDED · SELF-CONTAINED DECK