Skip to content

Commit 876a0c8

Browse files
committed
Add tiered iterative reproduce strategy and retries
TAG=agy CONV=dc7c7ce1-1d98-4b34-8b02-6b6c97c4523d Change-Id: Id5383c205c389ff6b6d6413572d38d050f540a22
1 parent e69f447 commit 876a0c8

5 files changed

Lines changed: 179 additions & 10 deletions

File tree

‎README_AGENTS.md‎

Lines changed: 17 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -164,49 +164,64 @@ when unavailable. Consumers query via
164164
vulnerabilities from `workspace/historical_learnings.jsonl` to enrich
165165
summaries and provide a quick reference map to optimize downstream planning
166166
and research.
167+
167168
04. **`/mantis-architecture` (Knowledge Base Architect):** Analyzes the codebase
168169
and clears the `workspace/learnings.jsonl` inbox to synthesize a permanent,
169170
interlinked Markdown Knowledge Base (`workspace/kb/`) detailing entities,
170171
data flows, and historical vulnerability classes.
172+
171173
05. **`/mantis-threat-model` (Threat Modeler):** Evaluates the entities and
172174
architecture defined in the KB to establish or refine a living
173175
`workspace/kb/THREAT_MODEL.md`, focusing on trust boundaries and attacker
174176
profiles.
177+
175178
06. **`/mantis-plan` (Strategist):** Scans workspace boundaries and reads the KB
176179
indices to output a targeted review strategy into `workspace/plan.json`,
177180
injecting specific `kb_references` file paths for context.
181+
178182
07. **`/mantis-researcher` (Mantis Researcher):** Executes file-by-file triage
179183
and deep security flaw reviews, outputting hotspots as individual JSON files
180184
in `workspace/findings/`.
185+
181186
08. **`/mantis-dedupe` (Deduplicator):** Groups index-based duplicate findings,
182187
merging records and deleting redundancies within `workspace/findings/`.
188+
183189
09. **`/mantis-review` (Validator):** Filters out false positives using strict
184190
pragmatic constraints, updating the status in
185191
`workspace/findings/<id>.json`.
192+
186193
10. **`/mantis-critic` (Critic):** Verifies release-build crash reproducibility
187194
(ignoring debug/assert checks), updates production viability in
188195
`workspace/findings/<id>.json`, and appends false positives/non-viable paths
189196
to `workspace/learnings.jsonl`.
197+
190198
11. **`/mantis-reproduce` (Proof-of-Concept Developer):** Writes
191199
Proof-of-Concept Reproduction Scripts (Repros) or raw payloads, executes
192-
them in isolated environments such as gVisor or Virtual Machines, and
193-
updates reproduction status in `workspace/findings/<id>.json`.
200+
them using a Tiered Iterative Reproduction strategy (unit micro-harness ->
201+
functional subsystem -> full sandboxed service) with
202+
intra/inter-conversation retries in isolated environments (gVisor, VMs,
203+
QEMU), and updates reproduction status in `workspace/findings/<id>.json`.
204+
194205
12. **`/mantis-chain` (Vulnerability Chainer):** Analyzes individual validated
195206
findings and knowledge base primitives to identify and construct complex
196207
multi-step exploit chains, creating new "Super Findings" in
197208
`workspace/findings/`.
209+
198210
13. **`/mantis-patch` (Patcher):** Generates and applies code fixes, runs
199211
post-patch validation tests inside the sandbox, updates patch status in
200212
`workspace/findings/<id>.json`, and appends logs to
201213
`workspace/learnings.jsonl`.
214+
202215
14. **`/mantis-calibrate` (Risk Calibrator):** Calculates a final numerical
203216
Mantis Risk Score (1-10) for each finding in the workspace directory based
204217
on impact, evidence, and viability, appending the results directly to each
205218
`workspace/findings/<id>.json` file.
219+
206220
15. **`/mantis-reflect` (Reflector):** Parses the execution trajectories of the
207221
agents from the current round, extracting false assumptions, tool failures,
208222
and successes, and appends these structured insights to the
209223
`workspace/learnings.jsonl` inbox.
224+
210225
16. **`/mantis-report` (Reporter):** Generates a human-readable security review
211226
packet containing verified/reproduced findings, evidence, risk rationales,
212227
and patch information at `workspace/report/review_packet-latest.md` (and

‎mantis-critic/SKILL.md‎

Lines changed: 5 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -245,7 +245,11 @@ Execute the critic evaluation as follows:
245245
CODE_ROOT and read at least **15 lines of preceding context** and **15 lines
246246
of succeeding context** around the designated line numbers. This targeted
247247
window is necessary to analyze surrounding structures and macro definitions.
248-
Proceed to Steps 4-5.
248+
Additionally, inspect `repro_hints` and `history` for empirical execution
249+
telemetry recorded by `mantis-reproduce` (e.g. `build_profile`,
250+
`sanitizers_used`, `assertions_disabled`, `ingress_blocked`). Use this
251+
empirical execution telemetry to corroborate release-build viability. Proceed
252+
to Steps 4-5.
249253

250254
4. **Evaluate Domain-Specific Viability Constraints:**
251255

‎mantis-pipeline-adapter/SKILL.md‎

Lines changed: 78 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -105,6 +105,10 @@ Follow these guidelines during the consultation:
105105
LLM reasoning, runs after the snapshot is pinned and before the first
106106
code-reading analysis stage, and degrades gracefully to grep when
107107
unavailable.
108+
11. **Advise on Tiered Iterative Reproduction & Multi-Conversation Retries:** If
109+
the user is targeting complex services where single-shot repro is brittle,
110+
walk them through the tiered iterative reproduction strategy and
111+
multi-conversation retry pattern in Reference Architecture Guideline 10.
108112

109113
## Reference Architecture Guidelines
110114

@@ -1069,3 +1073,77 @@ behave exactly as they do today:
10691073
slices) by using function boundaries from the structural index.
10701074
- The structural index and the RAG index can share the same vector embedding
10711075
infrastructure if both are implemented.
1076+
1077+
______________________________________________________________________
1078+
1079+
### 10. Tiered Iterative Reproduction & Multi-Conversation Retry Strategy
1080+
1081+
For complex services, attempting a single-shot reproduction directly against a
1082+
full sandboxed service often suffers from high search entropy, brittle
1083+
configuration, and hard-to-debug failures. A tiered strategy breaks reproduction
1084+
into incremental milestones, while an inter-conversation retry architecture
1085+
prevents reasoning deadlocks and context bloat.
1086+
1087+
#### A. Orchestration & Inter-Conversation Retries
1088+
1089+
1. **Intra-Conversation Retries (Local Agent Trajectory):**
1090+
1091+
- The active subagent conversation retries 2–3 times locally within its
1092+
context window to adjust parameters, fix setup bugs, or refine payloads.
1093+
1094+
2. **Inter-Conversation Retries (Fresh Context + Accumulated Artifacts):**
1095+
1096+
- **Trigger:** If intra-conversation retries fail to reach Tier 3
1097+
(`reproduced`), the orchestrator terminates the stalled conversation and
1098+
launches a **new subagent conversation** (a fresh context window).
1099+
- **Context Provisioning:** The orchestrator populates the new prompt with
1100+
structured attempt data from
1101+
`state_root/workspace/archive/.repro_attempts.json` and trajectory
1102+
learnings from `workspace/learnings.jsonl` (e.g., *"Attempt 1 failed due to
1103+
missing auth header X; Attempt 2 proved parser strips unescaped quotes"*).
1104+
- **Benefit:** Eliminates context bloat and reasoning inertia ("hallucination
1105+
traps"), enabling a fresh agent to solve the problem using prior empirical
1106+
observations without repeating past mistakes.
1107+
1108+
3. **Attempt Cap Accounting & Arithmetic:**
1109+
1110+
- The orchestrator maintains
1111+
`state_root/workspace/archive/.repro_attempts.json`.
1112+
- **Tier 3 Increment Only:** Only Tier-3 full sandboxed service executions
1113+
(or full end-to-end reproducer runs) increment the per-finding attempt
1114+
counter toward the absolute hard ceiling of 6.
1115+
- **Stepping-Stone Sub-Budget:** Internal Tier-1 and Tier-2 trial runs are
1116+
bounded by a local sub-budget (max 3 trial executions per conversation) and
1117+
do not consume the absolute 6-attempt cap.
1118+
1119+
#### B. Harness-Enforced Tier 4 (Staging & Live Pre-Production Execution)
1120+
1121+
To prevent un-gated exploit execution on live infrastructure, Tier 4 (Staging /
1122+
Pre-Production verification) is **strictly owned and enforced by the
1123+
programmatic orchestrator harness**, never by LLM discretion:
1124+
1125+
- **Sandbox Boundary:** `mantis-reproduce` executes strictly within isolated
1126+
local sandboxes (Tiers 1–3) ending at Tier 3 (`reproduced` /
1127+
`failed_to_reproduce`).
1128+
- **Deterministic Gate:** The harness intercepts a Tier 3 `reproduced` verdict.
1129+
If live/staging validation (Tier 4) is configured, the harness **MUST NOT**
1130+
automatically invoke remote execution. It must enforce a **programmatic
1131+
Human-in-the-Loop gate**:
1132+
- Prompt the human operator for explicit interactive confirmation, OR
1133+
- Require a cryptographically signed approval token / authorization callback.
1134+
- **Fail-Closed Default:** If human approval is missing or denied, Tier 4 is
1135+
skipped and the Tier 3 sandboxed verdict remains authoritative.
1136+
1137+
#### C. Local Ingress / Middlebox Edge Annotation (Post-Tier-3 Sandbox Verification)
1138+
1139+
To eliminate false-positive findings caused by default edge filters ("Works on
1140+
localhost:8080, but dies at the WAF/proxy"), the orchestrator can optionally
1141+
execute a Tier 3 PoC through a local reverse proxy or API gateway (e.g. NGINX,
1142+
Envoy, ModSecurity) running inside the local sandbox:
1143+
1144+
- **Annotation Only:** A payload blocked by a local middlebox MUST NOT downgrade
1145+
a Tier 3 `reproduced` verdict to `failed_to_reproduce`.
1146+
- **Critic Integration:** The harness records `ingress_blocked: true` (or edge
1147+
filter details) in the finding's `repro_hints` or history. This provides
1148+
empirical evidence for `/mantis-critic` to classify `production_viability` as
1149+
`"CONDITIONAL_VIABLE"` (mitigated by default edge proxy configuration).

‎mantis-reproduce/SKILL.md‎

Lines changed: 77 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -198,6 +198,18 @@ Execute the reproduction stage under these constraints:
198198
viability, but always check status).
199199
- If no applicable findings exist, notify the user and exit.
200200

201+
**Tier 0 — Structural Reachability Pre-Check (Advisory Queue Sorting):** If a
202+
structural code index (`mantis-structural-index`) is available, you MAY query
203+
`query_structural_index.py` (`find_callers`) before authoring code to check
204+
whether an AST call path exists from a public entrypoint to the vulnerable
205+
sink. Use this query to **prioritize candidate execution order** (process
206+
findings with verified AST reachability first).
207+
208+
- **CRITICAL HINT-ONLY GUARDRAIL:** AST reachability is a ranking HINT ONLY.
209+
Call graphs miss macros, function pointers, dynamic dispatch, and interface
210+
tables. An absent call path MUST NEVER reject a finding, skip reproduction,
211+
or set `failed_to_reproduce`.
212+
201213
**Snapshot drift check:** For each loaded finding, if it already has a
202214
`repro_snapshot_id` and Block B (Step 0) returns NOT_MATCHED, treat any
203215
stored PoC/offsets as STALE: regenerate the reproducer from scratch against
@@ -285,11 +297,57 @@ Execute the reproduction stage under these constraints:
285297
standard execution. Use your best judgment to construct a working harness for
286298
the artifact.
287299

288-
- *Optional Parallel Trajectory Search:* If your environment or agent
289-
framework supports spawning subagents, you can deploy multiple concurrent
290-
agents to attempt writing the reproducer via different logical approaches.
291-
If any trajectory succeeds, immediately adopt its payload and discard the
292-
others to escape potential "give up" loops.
300+
- *Parallel Trajectory Search vs. Tiered Iterative Reproduction:*
301+
302+
- **Parallel Trajectory Search (Breadth-First):** When subagents are
303+
available, deploy concurrent workers taking diverse logical approaches to
304+
reproduce the bug. If any trajectory succeeds, immediately adopt its
305+
payload and discard the others to escape potential "give up" loops and
306+
prune compute costs.
307+
308+
- **Tiered Iterative Reproduction (Depth-First Payload Refinement):** Each
309+
trajectory worker (or a single agent) uses a tiered escalation ladder
310+
(Tier 1 -> Tier 2 -> Tier 3) to refine its trigger payload incrementally
311+
rather than attempting a single-shot end-to-end launch.
312+
313+
- *Tiered Iterative Execution Ladder:*
314+
315+
- **Tier 1 (Micro-Harness / Sink Logic Validation):** Construct a
316+
lightweight test calling the vulnerable function/module directly to
317+
verify that the core bug hypothesis is sound in isolation.
318+
- **Tier 2 (Subsystem / Interface Validation):** Pass the payload through
319+
input serialization, parsers, routing, and auth wrappers to verify the
320+
input survives intermediate processing without sanitization or
321+
truncation.
322+
- **Tier 3 (Full Sandboxed Service / E2E Validation):** Execute the
323+
self-contained PoC against the target service via public APIs inside the
324+
isolated sandbox (Docker, QEMU, VM). Yields the authoritative
325+
`reproduced` verdict per Block F.
326+
327+
- **CRITICAL STEP-4 TIER-1 HARD GATE (Fail-Closed):**
328+
329+
- Tiers 1 and 2 are **internal stepping stones only**. You **MUST NEVER**
330+
record `repro_status = "reproduced"` or `"statically_confirmed"` based on
331+
a Tier-1 or Tier-2 execution.
332+
- If a crash can **ONLY** be achieved by compiling a direct-call harness
333+
that feeds a private/static function or bypasses the public API (Step 4),
334+
and the payload cannot be escalated to trigger through Tier 3 (the public
335+
API / sandboxed service), you **MUST TERMINATE AND RECORD**
336+
`repro_status = "failed_to_reproduce"` with details citing
337+
`"Internal Invariant Protection"`.
338+
339+
- *Attempt Cap Accounting & Local Retries:*
340+
341+
- **Sub-Tier-3 Stepping-Stone Sub-Budget:** Internal Tier-1 and Tier-2
342+
trial runs are bounded local execution steps (max 3 trial executions per
343+
conversation) and **DO NOT** increment the absolute per-finding attempt
344+
counter in `state_root/workspace/archive/.repro_attempts.json`.
345+
- **Absolute Attempt Cap Counting:** Only Tier-3 full sandboxed service
346+
executions (or full end-to-end reproducer runs) increment the absolute
347+
attempt counter toward the hard ceiling of 6 (Section 6).
348+
- **Intra-Conversation Retries:** When a tier fails, inspect logs, adjust
349+
payload parameters, fix harness setup, and retry up to 2-3 times within
350+
the active conversation before reporting back to the orchestrator.
293351

294352
### Step 3a: Variant Hunting (re-attack only, MANDATORY)
295353

@@ -534,17 +592,25 @@ and apply INV-1 (downgrade `VERIFIED_SECURE` → `VERIFICATION_FAILED`).
534592
* Open the lock file `state_root/workspace/archive/.repro_attempts.lock`
535593
(creating it if missing) and acquire an exclusive lock (`fcntl.flock`
536594
with `fcntl.LOCK_EX`) inside a context manager (`with` statement).
595+
537596
* Read the current contents of the cache file
538597
`state_root/workspace/archive/.repro_attempts.json` (treating it as `{}`
539598
if missing or empty).
599+
540600
* Increment the `count` field of this finding's cache-key entry — keyed by
541601
`signature` if present, else `stable_key`, the SAME key selection defined
542-
above (the `{count, last_snapshot}` object) — by 1.
602+
above (the `{count, last_snapshot}` object) — by 1 ONLY for Tier-3 (full
603+
end-to-end sandboxed service) executions. Internal Tier-1 and Tier-2
604+
stepping-stone trials MUST NOT increment `count` (they are governed by
605+
the sub-budget rule in Step 3).
606+
543607
* Write the updated JSON to a temporary file in the same directory (e.g.,
544608
`state_root/workspace/archive/.repro_attempts.json.tmp`).
609+
545610
* Atomically replace the target cache file with the temporary file (e.g.,
546611
`os.replace` in Python) to ensure readers never see a truncated or
547612
incomplete file.
613+
548614
* Close the lock file descriptor to release the lock (automatically handled
549615
by exiting the `with` context manager).
550616

@@ -564,6 +630,11 @@ and apply INV-1 (downgrade `VERIFIED_SECURE` → `VERIFICATION_FAILED`).
564630

565631
- `"repro_snapshot_id"`: the current SNAPSHOT_ID this run executed against.
566632

633+
- `"repro_hints"`: Record compilation and sandbox execution telemetry
634+
(e.g., `sanitizers_used: ASan+UBSan`, `assertions_disabled: true`,
635+
`build_profile: release`) to provide empirical execution evidence for
636+
`/mantis-critic`.
637+
567638
- If reproduction succeeds (`repro_status` is evaluated as `"reproduced"`
568639
or `"statically_confirmed"`) and the finding's current `"status"` is
569640
`"PROVISIONALLY_VALID"`: BEFORE upgrading, scan the finding's

‎schema.json‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -115,8 +115,9 @@
115115
},
116116
"repro_hints": {
117117
"type": "string",
118-
"description": "Instructions for the reproducer agent on how to trigger the bug."
118+
"description": "Instructions for triggering the bug, reached-sink evidence channels, and execution telemetry (e.g. build_profile, sanitizers_used, assertions_disabled, ingress_blocked)."
119119
},
120+
120121
"production_viability": {
121122
"type": "string",
122123
"enum": ["VIABLE", "NON_VIABLE", "SAMPLE_OR_TEST", "CONDITIONAL_VIABLE"],

0 commit comments

Comments
 (0)