Skip to content

Commit 79538a2

Browse files
committed
feat(skills): evaluate global intent for SAMPLE_OR_TEST viability
Change-Id: Idcec29bebf7b4ecacd3e280dbad55271c2b77e7e
1 parent fba313b commit 79538a2

8 files changed

Lines changed: 53 additions & 28 deletions

File tree

‎SCHEMA.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -39,7 +39,7 @@ evolves sequentially as different skills process it.
3939

4040
- **`production_viability`** (Enum): Whether the bug is triggerable in a
4141
release build.
42-
- Values: `"VIABLE"`, `"NON_VIABLE"`
42+
- Values: `"VIABLE"`, `"NON_VIABLE"`, `"SAMPLE_OR_TEST"`
4343
- **`critic_reasoning`** (String): Rationale for viability (e.g., "Not
4444
protected by allocator padding").
4545

@@ -127,8 +127,8 @@ trajectory.
127127
remove", "target_entity": "[e.g., auth_module.py]", "insight":
128128
"[description]", "source_stage": "[e.g., mantis_researcher]"}`
129129
- Alternatively for findings: `{"title": "[finding_title]", "code_paths":
130-
["[path1:line1]"], "status": "[VIABLE / NON_VIABLE / FALSE_POSITIVE /
131-
VERIFIED_SECURE / VERIFICATION_FAILED / ERROR]"}`
130+
["[path1:line1]"], "status": "[VIABLE / NON_VIABLE / SAMPLE_OR_TEST /
131+
FALSE_POSITIVE / VERIFIED_SECURE / VERIFICATION_FAILED / ERROR]"}`
132132

133133
This file serves as an ephemeral inbox/queue for new learnings. It is
134134
periodically synthesized into the permanent Markdown Knowledge Base

‎mantis_calibrate/SKILL.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -97,6 +97,14 @@ Execute the calibration as follows:
9797
paths), drastically reduce the multiplier to 0.2.
9898
- If `workspace/kb/THREAT_MODEL.md` does not exist or does not
9999
mention the components, default the Multiplier to 1.0.
100+
- If `production_viability` is **SAMPLE_OR_TEST**:
101+
- Set the Context Multiplier to a reduced value (e.g. `0.4`) so
102+
that severe bugs in sample code typically land in the MEDIUM
103+
bucket rather than HIGH or CRITICAL.
104+
- In the `executive_summary`, explicitly state that this is not a
105+
production bug. The recommendation MUST focus on fixing the
106+
example/test so that developers do not copy insecure patterns
107+
into production code.
100108

101109
**Final Score (Hazard) = (Impact + Likelihood) * Multiplier** (Capped at
102110
10.0).

‎mantis_chain/SKILL.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -32,7 +32,7 @@ Execute the chaining stage as follows:
3232

3333
- Read the JSON files in the `workspace/findings/` directory. Filter for
3434
findings that have passed validation (e.g., status is `"VALID"` or
35-
viability is `"VIABLE"`).
35+
viability is `"VIABLE"` or `"SAMPLE_OR_TEST"`).
3636
- Read the Markdown Knowledge Base (`workspace/kb/entities/` and
3737
`workspace/kb/vulnerabilities/`) to identify architectural primitives
3838
that might not be bugs on their own, but could serve as stepping stones

‎mantis_critic/SKILL.md‎

Lines changed: 28 additions & 16 deletions
Original file line numberDiff line numberDiff line change
@@ -34,14 +34,21 @@ Execute the critic evaluation as follows:
3434
that false positives can be properly logged to long-term memory. If none
3535
exist, notify the user.
3636

37-
2. **Acquire Targeted Code Snippets:** For each finding where `status` is
37+
2. **Evaluate Global Repository Intent:** Read `workspace/kb/THREAT_MODEL.md`
38+
(if it exists). Check the **Deployment Intent** section. If the threat model
39+
explicitly states the entire repository is exclusively a tutorial, sample
40+
project, or test suite (e.g., `Intent: SAMPLE_OR_TEST_ONLY`), you MUST mark
41+
all findings as **`SAMPLE_OR_TEST`** regardless of where they are located in
42+
the file structure, and skip the remaining per-finding viability checks.
43+
44+
3. **Acquire Targeted Code Snippets:** For each finding where `status` is
3845
`"VALID"`, read the target file. Read at least **15 lines of preceding
3946
context** and **15 lines of succeeding context** around the designated line
4047
numbers. This targeted window is necessary to analyze surrounding structures
4148
and macro definitions. (Skip this and the following evaluation steps for
4249
`"FALSE POSITIVE"` findings).
4350

44-
3. **Evaluate Domain-Specific Viability Constraints:**
51+
4. **Evaluate Domain-Specific Viability Constraints:**
4552

4653
- **For Memory Safety Flaws:** Locate the allocation source of the
4754
affected buffer. Determine if it is allocated with safety margins or
@@ -52,7 +59,7 @@ Execute the critic evaluation as follows:
5259
deployments. If the flaw relies on a debug-only backdoor, a mock
5360
authentication provider, or a test-only route, mark it **NON_VIABLE**.
5461

55-
4. **Check Critical Non-Viable Criteria:** Mark a finding as **NON_VIABLE** if
62+
5. **Check Critical Non-Viable Criteria:** Mark a finding as **NON_VIABLE** if
5663
it meets any of the following criteria to ensure we do not waste patching
5764
resources on unreachable or compiled-out code:
5865

@@ -65,37 +72,42 @@ Execute the critic evaluation as follows:
6572
- **Debug-Only Features:** Security flaws that exist inside files or
6673
sections conditionally compiled with debug flags (e.g. `#ifdef DEBUG`)
6774
are NON_VIABLE.
68-
- **Harnesses & Mocks:** Issues residing in the helper configurations of
69-
testing libraries, fuzzing suites, or validation frameworks rather than
70-
the production library core are NON_VIABLE.
71-
72-
5. **Token-Optimized File Updates:** To minimize LLM output tokens, **do not
75+
- **Harnesses, Mocks, & Examples:** Issues residing in example code, test
76+
suites, fuzzing harnesses, or validation frameworks are technically not
77+
deployed to production. However, because developers often copy sample
78+
code, these should NOT be marked NON_VIABLE. Instead, mark them as
79+
**SAMPLE_OR_TEST** so the pipeline can properly adjust their risk
80+
severity.
81+
82+
6. **Token-Optimized File Updates:** To minimize LLM output tokens, **do not
7383
re-emit or manually rewrite the entire JSON object in your output.**
7484
Instead, use in-place editing tools (like a short script in your preferred
7585
language, or `jq`) to programmatically append the new fields to the existing
7686
`workspace/findings/<id>.json` file.
7787

7888
You must append the following to the existing object:
7989

80-
- A `"production_viability"` field (either `"VIABLE"` or `"NON_VIABLE"`).
90+
- A `"production_viability"` field (`"VIABLE"`, `"NON_VIABLE"`, or
91+
`"SAMPLE_OR_TEST"`).
8192
- A `"critic_reasoning"` field explaining your evaluation.
8293
- An entry to the `"history"` array:
8394

8495
```json
8596
{
8697
"stage": "critic",
8798
"action": "evaluated",
88-
"details": "Determined production viability as [VIABLE/NON_VIABLE] because [reason]"
99+
"details": "Determined production viability as [VIABLE/NON_VIABLE/SAMPLE_OR_TEST] because [reason]"
89100
}
90101
```
91102

92-
6. **Append to Long-Term Memory:** For each finding you loaded (including both
93-
`NON_VIABLE` ones and `FALSE POSITIVE`s), append a single structured JSON
94-
line to a workspace database file named `learnings.jsonl` (using append
95-
mode). This ensures false positives and non-viable paths are remembered
96-
across runs, helping the strategist avoid re-scanning them.
103+
7. **Append to Long-Term Memory:** For each finding you loaded (including
104+
`NON_VIABLE`, `SAMPLE_OR_TEST`, and `FALSE POSITIVE`s), append a single
105+
structured JSON line to a workspace database file named `learnings.jsonl`
106+
(using append mode). This ensures false positives and non-viable paths are
107+
remembered across runs, helping the strategist avoid re-scanning them.
97108

98109
- **Memory Entry Format:** `{"title": "[finding_title]", "code_paths":
99-
["[path1:line1]"], "status": "[NON_VIABLE / FALSE_POSITIVE / VIABLE]"}`
110+
["[path1:line1]"], "status": "[NON_VIABLE / SAMPLE_OR_TEST /
111+
FALSE_POSITIVE / VIABLE]"}`
100112

101113
When complete, notify the user.

‎mantis_dedupe/SKILL.md‎

Lines changed: 4 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -45,9 +45,10 @@ Execute your task as follows:
4545
`learnings.jsonl` entries (using `code_paths` and `title` similarities). If
4646
a current finding exactly matches a flaw that was already processed in this
4747
loop (regardless of whether its status was `FALSE_POSITIVE`, `NON_VIABLE`,
48-
or `VERIFIED_SECURE`), you must **delete the new finding file entirely** and
49-
drop it from your active list. This prevents the pipeline from getting stuck
50-
re-evaluating the same issues in the current pass.
48+
`SAMPLE_OR_TEST`, or `VERIFIED_SECURE`), you must **delete the new finding
49+
file entirely** and drop it from your active list. This prevents the
50+
pipeline from getting stuck re-evaluating the same issues in the current
51+
pass.
5152

5253
3. **Filter Duplicate Findings in Current Batch:** Check the current findings
5354
against each other to find duplicates (using `code_paths` and `title`

‎mantis_meta_agent/SKILL.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -93,8 +93,8 @@ Execute your orchestration duties in a continuous loop:
9393
`workspace/findings/` directory (e.g., move it to
9494
`workspace/archive/findings_pass_N/`) to clear the state for the next
9595
loop. *Crucially*, you must ensure that before archiving, any finalized
96-
findings (especially `FALSE_POSITIVE`, `NON_VIABLE`, or
97-
`VERIFIED_SECURE`) were successfully captured by the
96+
findings (especially `FALSE_POSITIVE`, `NON_VIABLE`, `SAMPLE_OR_TEST`,
97+
or `VERIFIED_SECURE`) were successfully captured by the
9898
`@mantis_architecture` subagent and written into the permanent Markdown
9999
Knowledge Base (`workspace/kb/`). If they are only archived but not in
100100
the KB, the Researcher will just re-find them in the next loop.

‎mantis_reproduce/SKILL.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -27,9 +27,9 @@ that reproduces a confirmed security flaw.
2727
Execute the reproduction stage under these constraints:
2828

2929
1. **Load Viable Findings:** Read the JSON files in the `workspace/findings/`
30-
directory. Filter for findings where `production_viability` is `"VIABLE"`
31-
(or skip the filter if you're not checking viability). If no applicable
32-
findings exist, notify the user.
30+
directory. Filter for findings where `production_viability` is `"VIABLE"` or
31+
`"SAMPLE_OR_TEST"` (or skip the filter if you're not checking viability). If
32+
no applicable findings exist, notify the user.
3333

3434
2. **Strict Host Isolation Constraint:**
3535

‎mantis_threat_model/SKILL.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -55,6 +55,10 @@ Execute the threat modeling process as follows:
5555

5656
- **System Overview Summary:** A concise summary derived from
5757
`architecture.md`.
58+
- **Deployment Intent:** Determine if the entire repository is intended
59+
for production use, or if it is exclusively a tutorial, sample project,
60+
or test suite. State this clearly (e.g., `Intent: SAMPLE_OR_TEST_ONLY`
61+
or `Intent: PRODUCTION`).
5862
- **Trust Boundaries:** Clear, rigorous definitions of where untrusted
5963
inputs meet internal trusted states. Reference the specific entities
6064
(e.g., `[Auth Module](entities/auth_module.md)`).

0 commit comments

Comments
 (0)