Skip to content

Commit 69eea7a

Browse files
committed
Mantis: Improve reproduction decisiveness, quality filtering, and reporting
- Update mantis_reproduce/SKILL.md to instruct the reproducer to dynamically adjust finding details on failure or abandon if unsuccessful. - Update mantis_meta_agent/SKILL.md to exclude un-reproduced, low-priority, or non-viable findings from confirmed reports. - Update mantis_review/SKILL.md to add nuance for race conditions, ensuring they are not dismissed if they are automatable/brute-forceable. - Update mantis_critic/SKILL.md to mark findings blocked by non-configurable environmental controls as non-viable. - Update mantis_calibrate/SKILL.md to cap the impact of internal admin privilege escalation/lateral movement at 2, and add a "Critical Sanity Triage" step to force-downgrade weak, speculative, or minor hygiene findings to LOW priority. TAG=agy CONV=c12a2092-bd2e-4b83-8237-51876a76909a Change-Id: I9ee9bce1a1b3aa499a9c93e9b0e42c8f9b796336
1 parent a8a4538 commit 69eea7a

5 files changed

Lines changed: 68 additions & 11 deletions

File tree

‎mantis_calibrate/SKILL.md‎

Lines changed: 33 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -64,6 +64,13 @@ Execute the calibration as follows:
6464
the type "the code is fragile", "lack of defense-in-depth", or
6565
purely theoretical hygiene issues MUST have an Impact score of 1,
6666
ensuring they are rated LOW at most.
67+
- *Note on Internal Admin & Lateral Movement:* If a finding requires
68+
pre-existing administrative privileges (internal privilege
69+
escalation, e.g., admin to super-admin) or only allows lateral
70+
movement/pivoting between internal components from an already
71+
compromised state, cap its individual Impact score at 2. Its
72+
escalated risk will be captured separately if it is successfully
73+
chained into a "Super Finding" by the chainer.
6774
- **Likelihood (1-5):** Evaluate the probability of occurrence based on
6875
proven exploitability rather than theoretical difficulty.
6976
- 5: Actively exploited in the wild, OR the agent successfully
@@ -114,7 +121,31 @@ Execute the calibration as follows:
114121
damage, user sentiment fallout) is taken into account. Do *not* include the
115122
outrage factor in the final numerical score.
116123

117-
3. **Determine Priority:**
124+
3. **Critical Sanity Triage (Downgrading Weak Findings):** Before determining
125+
the final priority, perform a second-level sanity check on the quality of
126+
the finding and its accumulated evidence. You **MUST** force-downgrade the
127+
finding's priority to **LOW** (and cap its final score at **2.0**) if it
128+
meets any of the following "weak finding" criteria:
129+
130+
- **Speculative Viability on Repro Failure:** The reproduction failed
131+
(`repro_status: "failed_to_reproduce"`), and the reasoning for why it is
132+
still viable in production is highly speculative, theoretical, or relies
133+
on unverified assumptions about downstream systems.
134+
- **Minor Configuration Hygiene:** The issue represents a minor deviation
135+
from best-practice configuration (e.g., slightly loose permissions on an
136+
internal directory, lack of modern encryption on low-value internal
137+
transport) but does not lead to a clear exploit path, privilege
138+
escalation, or data exposure.
139+
- **Vague Code Paths / Fragile Assumptions:** The finding's description or
140+
reasoning relies on unverified assumptions about caller behavior or
141+
adjacent system components that are not documented in the active
142+
Knowledge Base.
143+
- **Unreliable/Noisy Triggers:** The finding represents an issue that can
144+
technically be triggered but is highly likely to be ignored in practice
145+
due to high noise, or is indistinguishable from normal system operations
146+
without causing real harm.
147+
148+
4. **Determine Priority:**
118149

119150
- **CRITICAL (8.0 - 10.0):** Immediate action required. Very high hazard
120151
(e.g. high impact and likelihood). **Must NOT be used unless it
@@ -130,7 +161,7 @@ Execute the calibration as follows:
130161
exclusively affects a single user's own data MUST be capped at LOW
131162
priority regardless of the calculated score.**
132163

133-
4. **Token-Optimized File Updates:** To minimize LLM output tokens, **do not
164+
5. **Token-Optimized File Updates:** To minimize LLM output tokens, **do not
134165
re-emit or manually rewrite the entire JSON object in your output.**
135166
Instead, write a reusable helper script (e.g.,
136167
`workspace/helpers/append_calibrate.py`) during your first finding update.

‎mantis_critic/SKILL.md‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -72,6 +72,11 @@ Execute the critic evaluation as follows:
7272
- **Debug-Only Features:** Security flaws that exist inside files or
7373
sections conditionally compiled with debug flags (e.g. `#ifdef DEBUG`)
7474
are NON_VIABLE.
75+
- **Blocked by Environmental Controls:** If the exploit path is blocked by
76+
standard, non-configurable production environmental controls (e.g.,
77+
OS-level permissions, kernel-level sandboxing, read-only filesystems, or
78+
hardware-enforced write protections) that cannot be bypassed, mark it
79+
NON_VIABLE.
7580
- **Harnesses, Mocks, & Examples:** Issues residing in example code, test
7681
suites, fuzzing harnesses, or validation frameworks are technically not
7782
deployed to production. However, because developers often copy sample

‎mantis_meta_agent/SKILL.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -116,6 +116,14 @@ Execute your orchestration duties in a continuous loop:
116116
subagent), or when `@mantis_calibrate` finishes its scoring, you may output
117117
a brief text summary to the user.
118118

119+
- **Do NOT report findings that failed to reproduce (`repro_status:
120+
"failed_to_reproduce"`) as confirmed vulnerabilities.**
121+
- **Do NOT report findings that are marked `LOW` priority or are
122+
`NON_VIABLE` as confirmed vulnerabilities.** These are considered
123+
low-quality, fragile, or non-actionable. You may list them separately at
124+
the bottom of your summary under a "Hygiene & Low Priority Notes"
125+
section, but do not present them as active security flaws.
126+
119127
4. **Human-in-the-Loop Steering & Collaboration:** While you are designed for
120128
autonomy, remain responsive to user input. The user may interrupt the loop
121129
to ask for progress updates, collaboratively debug environmental issues, or

‎mantis_reproduce/SKILL.md‎

Lines changed: 14 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -45,14 +45,20 @@ Execute the reproduction stage under these constraints:
4545

4646
3. **Writing and Launching the Reproducer:** Write a self-contained test script
4747
(e.g., `poc.py` or a C reproducer file) or write a raw crash input data
48-
payload (e.g., `crash.payload`) that triggers the target bug. To run your
49-
script or payload, use the execution or containerization tools available in
50-
your environment to execute the code safely. Select the most appropriate
51-
runtime image and flags for the target. **Execute your reproduction using
52-
the appropriate environment:** If the target is firmware, you may write a
53-
script to boot it via `qemu`, `unicorn`, or Firmadyne. If it's a binary, you
54-
may use dynamic instrumentation or standard execution. Use your best
55-
judgment to construct a working harness for the artifact.
48+
payload (e.g., `crash.payload`) that triggers the target bug. Analyze the
49+
code path and constraints carefully. If your initial reproduction attempt
50+
fails, evaluate if the finding details (such as input paths, parameters, or
51+
assumptions) are slightly incorrect based on your observations, and adjust
52+
the finding details dynamically to attempt a fix. If you cannot find a
53+
triggerable path after trying multiple approaches and adjustments, abandon
54+
the attempt and mark it as `failed_to_reproduce`. To run your script or
55+
payload, use the execution or containerization tools available in your
56+
environment to execute the code safely. Select the most appropriate runtime
57+
image and flags for the target. **Execute your reproduction using the
58+
appropriate environment:** If the target is firmware, you may write a script
59+
to boot it via `qemu`, `unicorn`, or Firmadyne. If it's a binary, you may
60+
use dynamic instrumentation or standard execution. Use your best judgment to
61+
construct a working harness for the artifact.
5662

5763
- *Optional Parallel Trajectory Search:* If your environment or agent
5864
framework supports spawning subagents, you can deploy multiple

‎mantis_review/SKILL.md‎

Lines changed: 8 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -52,7 +52,14 @@ Execute your validation as follows:
5252
as security flaws.
5353
3. **Require Strict Reproducibility:** Only mark a finding as VALID if a
5454
direct, unambiguous, and triggerable flaw exists within the boundaries
55-
of the code logic.
55+
of the code logic. If the finding is extremely fragile (e.g., relies on
56+
unstable timing that cannot be automated or brute-forced, or requires
57+
unrealistic environmental conditions to trigger), mark it as
58+
FALSE_POSITIVE. *Note on Race Conditions:* Do NOT dismiss race
59+
conditions or timing bugs simply because they have a low success
60+
probability (e.g., 1 in a million), provided the attack path can be
61+
automated and repeatedly attempted by an attacker to eventually trigger
62+
the exploit.
5663
4. **Avoid Pedantic Linting:** If the code uses standard safe libraries
5764
(such as `json.loads`, parameterised SQL queries, or secure standard
5865
library hashes) but lacks extreme paranoia, mark it as FALSE_POSITIVE.

0 commit comments

Comments
 (0)