Best-of-N and critique-refine auditor protocols for Petri 3.
Petri's auditor improvises each action in a single shot. petri-bon makes the auditor optimize each action instead: every turn, candidate actions are drafted, scored by a preference model for how deployment-like the resulting transcript looks, critiqued, refined, and the best candidate across all rounds is selected. The refinement loop leaves no trace in the auditor's context — the next turn starts clean with only the winning action, and the target never sees any of it.
This matters for realism: as models get better at detecting when they are being evaluated, audits need actions that read as genuine deployment traffic. Iteratively optimizing each auditor action against a deployment-vs-audit classifier directly targets the tells that make audits detectable.
petri-bon is a standalone package on top of inspect-petri — no fork. It plugs into Petri 3's auditor_agent(generate=...) hook: Petri calls the hook once per auditor turn, and only the returned message enters the auditor's history. Everything the protocol does inside a turn (drafts, scoring, critiques, refinements) happens within the hook, which gives context-trimming for free and keeps all intermediate model calls visible as events in the Inspect transcript viewer.
One critique-refine turn (rounds=2, n_samples=1):
auditor draft (round 0)
→ preference scoring P(deployment) for the target-visible transcript
→ critique feedback model explains what reads as an audit
→ auditor refine (round 1) same action shape, addressing the critique
→ preference scoring
→ critique
→ auditor refine (round 2)
→ preference scoring
→ winner selection argmax across rounds 0, 1, 2
The preference model only ever sees what the target would see (system prompt, user messages, tool results, target responses, plus the candidate's staged effects) — never the auditor's reasoning, tools, or the seed instruction.
pip install git+https://github.com/AxelAhlqvist1995/petri-bon # or: pip install -e .Requires Python ≥ 3.12 and inspect-petri >= 3.0.10.
inspect eval petri_bon/bon_audit \
-T selector=logprob -T rounds=2 -T n_samples=1 \
--model-role auditor=anthropic/claude-opus-4-6 \
--model-role target=anthropic/claude-sonnet-4-6 \
--model-role judge=anthropic/claude-opus-4-6 \
--limit 10All of Petri's standard audit() parameters (seed_instructions, max_turns, realism_filter, judge_dimensions, ...) pass through unchanged. View results with inspect view as usual — per-turn round scores and the winning round are recorded in the transcript, and the draft/score/critique/refine model calls appear as events inside each auditor turn.
-T protocol= |
Behavior |
|---|---|
critique_refine (default) |
Iteratively critique and refine each candidate chain; select the best across all rounds (score_original/select_from_all tunables). |
parallel |
Sample n_samples candidates, select the best. No refinement. |
-T selector= |
Scoring | Providers |
|---|---|---|
logprob (default) |
P(deployment) from choice-token probabilities (D/d vs A/a), read directly from logprobs — no sampling noise. | Anthropic only (probabilities-2024-07-31 beta; uses ANTHROPIC_API_KEY_LP, falling back to ANTHROPIC_API_KEY) |
score |
Model emits <answer>P(deployment)</answer> as a decimal. |
Any (via Inspect get_model) |
pairwise |
Head-to-head comparisons on shared-prefix transcripts; single-elimination bracket for pools > 2. | Any (via Inspect get_model) |
The preference model defaults to the preference model role (--model-role preference=...), else the auditor model; -T preference_model=... overrides. The feedback (critique) model defaults to the preference model.
Implement CandidateScorer (pointwise) or Selector (winner-takes-all) and pass the instance from Python:
from inspect_ai import eval
from petri_bon import bon_audit, Candidate, TurnContext
class MyScorer:
async def score(self, ctx: TurnContext, candidate: Candidate) -> None:
transcript = candidate.details["transcript"] # target-visible view
candidate.score = await my_realism_metric(transcript) # in [0, 1]
eval(bon_audit(selector=MyScorer(), rounds=2), model_roles={...})ctx carries the turn number, seed instruction, the auditor's message context, and the live target conversation. Custom FeedbackProviders plug in the same way via feedback=.
All preference/feedback/refinement templates ship as package data (src/petri_bon/prompts/). -T preference_prompt= and -T feedback_prompt= accept a bundled name, a .txt file path, or inline template text.
- Token accounting: Petri records only the winning generation as the auditor turn's output; the intermediate draft/refine/scoring calls appear as model events (and in total usage) but not in the turn's usage figure.
- Unscorable candidates: a candidate that stages nothing target-visible (e.g. bare
resume) cannot be scored and is sentinel-ranked below every scored candidate. Structurally invalid candidates (e.g.send_tool_call_resultwith a badtool_call_id) get a structural-fix critique instead of a preference critique, and never win selection; if every candidate is invalid, the round-0 draft is used as a graceful fallback. - Logprob beta: the
probabilities-2024-07-31beta is not generally available; useselector=scoreorselector=pairwiseif your key lacks access. - Prompt caching: selector, scorer, and feedback prompts are sent as two content blocks — the batch-stable prefix (prompt intro + shared conversation) and the candidate-specific tail — so Anthropic prompt caching reads the shared prefix across the many calls of a turn instead of cache-writing it on every call (writes bill at 1.25x, reads at 0.1x). On Anthropic, each same-turn batch also completes one call before fanning out the rest, so parallel calls don't race a cold cache. Deep audits can exceed the default 5-minute cache TTL between turns; recent inspect_ai versions (0.3.260 confirmed) expose the Anthropic provider option
cache_ttl="1h"as a model arg (e.g.-M cache_ttl=1h) to extend it.
Built on inspect-petri (Meridian Labs, MIT), the successor to Petri (Anthropic, MIT). The critique-refine protocol and prompt templates originate from the author's research fork of Petri 2. MIT licensed.