@@ -37,15 +37,15 @@ useful, it is even more important to use this in a suitably isolated environment
3737to prevent impacting production systems. See the notes on unattended cloud
3838deployment later in this guide.
3939
40- --------------------------------------------------------------------------------
40+ ______________________________________________________________________
4141
4242## Architecture and Sequential Flow
4343
4444For a detailed breakdown of the pipeline stages, sequential flow, and
4545inter-stage contracts, please refer to the
4646[ Agent Reference Guide] ( README_AGENTS.md ) .
4747
48- --------------------------------------------------------------------------------
48+ ______________________________________________________________________
4949
5050## Prerequisites and Setup
5151
@@ -56,23 +56,24 @@ CLI, among others. Any coding agent framework should work. We have used these
5656skills successfully with both the Google ADK and Antigravity SDK. You might also
5757consider:
5858
59- 1 . ** Docker** For testing containers.
60- 2 . ** gVisor (runsc)** : For enhanced security when executing untrusted
61- AI-generated crash reproducer code, register the ` runsc ` runtime in your
62- Docker daemon configuration (` /etc/docker/daemon.json ` ):
59+ 1 . ** Docker** For testing containers.
6360
64- ``` json
65- {
66- "runtimes" : {
67- "runsc" : {
68- "path" : " runsc"
69- }
70- }
71- }
72- ```
61+ 2 . ** gVisor (runsc)** : For enhanced security when executing untrusted
62+ AI-generated crash reproducer code, register the ` runsc ` runtime in your
63+ Docker daemon configuration (` /etc/docker/daemon.json ` ):
7364
74- 3 . **Relevant Cloud SDKs**: If running remote cloud sandboxes instead of local
75- containers.
65+ ``` json
66+ {
67+ "runtimes" : {
68+ "runsc" : {
69+ "path" : " runsc"
70+ }
71+ }
72+ }
73+ ```
74+
75+ 3 . ** Relevant Cloud SDKs** : If running remote cloud sandboxes instead of local
76+ containers.
7677
7778### Installing the Skills
7879
@@ -87,7 +88,7 @@ To install the skills via CLI:
8788npx skills add google/mantis
8889```
8990
90- --------------------------------------------------------------------------------
91+ ______________________________________________________________________
9192
9293## Beginner's Guide & Best Practices
9394
@@ -142,80 +143,80 @@ AI-based scanning will produce
142143[ false positives] ( README_AGENTS.md#understanding-false-positives-the-negative-filter-rule )
143144(the review stage applies negative rules to filter them).
144145
145- * ** Expect Noise** : Customize the negative validation filters in
146- ` /mantis-review ` to match your codebase.
147- * ** Iterate Small** : Start with narrow-scope scans to tune the pipeline rather
148- than running a repository-wide sweep on Day 1.
146+ - ** Expect Noise** : Customize the negative validation filters in
147+ ` /mantis-review ` to match your codebase.
148+ - ** Iterate Small** : Start with narrow-scope scans to tune the pipeline rather
149+ than running a repository-wide sweep on Day 1.
149150
150- --------------------------------------------------------------------------------
151+ ______________________________________________________________________
151152
152153## Running the Pipeline (Manual Mode)
153154
154155You can execute the reviewing stages sequentially from ** inside** your active
155156CLI terminal.
156157
157- 1 . Start your CLI from your terminal.
158+ 1 . Start your CLI from your terminal.
158159
159- 2 . Inside the interactive UI prompt, type the skills sequentially:
160+ 2 . Inside the interactive UI prompt, type the skills sequentially:
160161
161- ``` text
162- # 0. (Optional) Analyze repository's version control system (VCS) history and extract past vulnerabilities
163- /mantis-history
162+ ``` text
163+ # 0. (Optional) Analyze repository's version control system (VCS) history and extract past vulnerabilities
164+ /mantis-history
164165
165- # 1. (Optional) Generate mantis-summary.md directory maps
166- /mantis-summarize
166+ # 1. (Optional) Generate mantis-summary.md directory maps
167+ /mantis-summarize
167168
168- # 2. Synthesize codebase structure and historical learnings into the Markdown Knowledge Base
169- /mantis-architecture
169+ # 2. Synthesize codebase structure and historical learnings into the Markdown Knowledge Base
170+ /mantis-architecture
170171
171- # 3. Iteratively develop the project's living threat model based on the KB
172- /mantis-threat-model
172+ # 3. Iteratively develop the project's living threat model based on the KB
173+ /mantis-threat-model
173174
174- # 4. Map target external boundary and build scanning roadmap, injecting KB references
175- /mantis-plan
175+ # 4. Map target external boundary and build scanning roadmap, injecting KB references
176+ /mantis-plan
176177
177- # 5. Run multi-threaded/sequential security flaw sweep using injected context
178- /mantis-researcher
178+ # 5. Run multi-threaded/sequential security flaw sweep using injected context
179+ /mantis-researcher
179180
180- # 6. Consolidate overlapping files and duplicate bugs
181- /mantis-dedupe
181+ # 6. Consolidate overlapping files and duplicate bugs
182+ /mantis-dedupe
182183
183- # 7. Verify code validity & filter false positives
184- /mantis-review
184+ # 7. Verify code validity & filter false positives
185+ /mantis-review
185186
186- # 8. Eliminate non-viable production issues
187- /mantis-critic
187+ # 8. Eliminate non-viable production issues
188+ /mantis-critic
188189
189- # 9. Generate proof-of-concept crash reproducers and run them in sandboxes
190- /mantis-reproduce
190+ # 9. Generate proof-of-concept crash reproducers and run them in sandboxes
191+ /mantis-reproduce
191192
192- # 10. Combine validated individual findings into multi-step exploit chains
193- /mantis-chain
193+ # 10. Combine validated individual findings into multi-step exploit chains
194+ /mantis-chain
194195
195- # 11. Apply minimal fixes and verify they block the crash reproducer
196- /mantis-patch
196+ # 11. Apply minimal fixes and verify they block the crash reproducer
197+ /mantis-patch
197198
198- # 12. Calculate final matrix risk ratings and append to individual findings
199- /mantis-calibrate
199+ # 12. Calculate final matrix risk ratings and append to individual findings
200+ /mantis-calibrate
200201
201- # 13. Extract insights from execution trajectories and append to the learnings inbox
202- /mantis-reflect
202+ # 13. Extract insights from execution trajectories and append to the learnings inbox
203+ /mantis-reflect
203204
204- # 14. Generate human-readable security review packet report
205- /mantis-report
205+ # 14. Generate human-readable security review packet report
206+ /mantis-report
206207
207- # 15. (Manual Step) Review the report. Optionally, you can apply & commit approved patches to your codebase. To continue analysis, archive workspace/findings/, and loop back to Step 2 to start the next pass.
208- ```
208+ # 15. (Manual Step) Review the report. Optionally, you can apply & commit approved patches to your codebase. To continue analysis, archive workspace/findings/, and loop back to Step 2 to start the next pass.
209+ ```
209210
210- --------------------------------------------------------------------------------
211+ ______________________________________________________________________
211212
212213## Building Deterministic Pipelines & Non-Determinism
213214
214215For information on how to build production-grade deterministic pipelines
215216wrapping these skills, and how to manage LLM non-determinism, see the
216217[ Agent Reference Guide] ( README_AGENTS.md ) .
217218
218- --------------------------------------------------------------------------------
219+ ______________________________________________________________________
219220
220221## Advanced / Unattended Cloud Deployment (GCE)
221222
@@ -224,62 +225,61 @@ hardened environment to mitigate security risks (such as prompt injection). For
224225the step-by-step hardened Google Compute Engine (GCE) deployment guide, see the
225226[ Agent Reference Guide] ( README_AGENTS.md#advanced--unattended-cloud-deployment-gce ) .
226227
227- --------------------------------------------------------------------------------
228+ ______________________________________________________________________
228229
229230## Meta-Agent Orchestration & Evaluation
230231
231232For details on autonomous Meta-Agent execution and how to evaluate/optimize
232233skill performance, see the [ Agent Reference Guide] ( README_AGENTS.md ) .
233234
234- --------------------------------------------------------------------------------
235+ ______________________________________________________________________
235236
236237## Roadmap / Future Work
237238
238- * **Continuous Pipeline:** The current pipeline is designed to be run as a
239- point-in-time review of a codebase, and not as something that is intended to
240- regularly sync with upstream changes mid-run. It should be straightforward
241- (if a little tricky) to tweak the pipeline to better support this, but it
242- probably will not work today.
243- * **Skill Self-Improvement (Meta-Learning):** The current
244- `workspace/learnings.jsonl` and Knowledge Base (KB) architecture tracks
245- codebase-specific empirical outcomes to adapt the `THREAT_MODEL.md` and
246- context pointers. Future iterations of the pipeline could take this a step
247- further and use this historical data to reflect on and automatically rewrite
248- its own `SKILL.md` prompts. For example, if a certain type of hallucination
249- is repeatedly caught by the Critic, a self-improvement meta-agent could
250- update the Researcher's `SKILL.md` instructions to explicitly filter out
251- that specific pattern before it even reaches the Review stage. **Security
252- Note:** Committing automated changes to `SKILL.md` files must always be
253- human-gated to prevent an attacker from using prompt injection (e.g., via a
254- malicious payload in a target file) to trick the meta-agent into ignoring a
255- vulnerability class globally.
256- * **Software Dark Factory:** Integrate this pipeline into an entirely AI
257- driven software development. Instead of vulnerable discovery for action by
258- humans, Mantis would become the autonomous vulnerability research and
259- release gating component of the dark factory. Before the dark factory can
260- push to production, it must have had N hours of adversarial vulnerability
261- research or "red teaming" by a pipeline like Mantis.
262-
263- --------------------------------------------------------------------------------
239+ - ** Continuous Pipeline:** The current pipeline is designed to be run as a
240+ point-in-time review of a codebase, and not as something that is intended to
241+ regularly sync with upstream changes mid-run. It should be straightforward (if
242+ a little tricky) to tweak the pipeline to better support this, but it probably
243+ will not work today.
244+ - ** Skill Self-Improvement (Meta-Learning):** The current
245+ ` workspace/learnings.jsonl ` and Knowledge Base (KB) architecture tracks
246+ codebase-specific empirical outcomes to adapt the ` THREAT_MODEL.md ` and
247+ context pointers. Future iterations of the pipeline could take this a step
248+ further and use this historical data to reflect on and automatically rewrite
249+ its own ` SKILL.md ` prompts. For example, if a certain type of hallucination is
250+ repeatedly caught by the Critic, a self-improvement meta-agent could update
251+ the Researcher's ` SKILL.md ` instructions to explicitly filter out that
252+ specific pattern before it even reaches the Review stage. ** Security Note: **
253+ Committing automated changes to ` SKILL.md ` files must always be human-gated to
254+ prevent an attacker from using prompt injection (e.g., via a malicious payload
255+ in a target file) to trick the meta-agent into ignoring a vulnerability class
256+ globally.
257+ - ** Software Dark Factory:** Integrate this pipeline into an entirely AI driven
258+ software development. Instead of vulnerable discovery for action by humans,
259+ Mantis would become the autonomous vulnerability research and release gating
260+ component of the dark factory. Before the dark factory can push to production,
261+ it must have had N hours of adversarial vulnerability research or "red
262+ teaming" by a pipeline like Mantis.
263+
264+ ______________________________________________________________________
264265
265266## Troubleshooting Guide
266267
267268### 1. Loop Iterations are Re-Evaluating the Same Code
268269
269- * **Symptom:** The loop keeps reviewing the same files and reporting identical
270- bugs.
271- * **Solution:** Ensure `/mantis-architecture` completes successfully and
272- writes its synthesized knowledge to the `workspace/kb/` directory. The
273- `/mantis-plan` strategist checks this Knowledge Base to dynamically skip
274- already analyzed areas. Check that file permissions allow writing to
275- `workspace/kb/`.
270+ - ** Symptom:** The loop keeps reviewing the same files and reporting identical
271+ bugs.
272+ - ** Solution:** Ensure ` /mantis-architecture ` completes successfully and writes
273+ its synthesized knowledge to the ` workspace/kb/ ` directory. The ` /mantis-plan `
274+ strategist checks this Knowledge Base to dynamically skip already analyzed
275+ areas. Check that file permissions allow writing to ` workspace/kb/ ` .
276276
277277### 2. Other Issues
278278
279- * **Symptom:** Something isn't working.
280- * **Solution:** Ask an AI coding tool to review your pipeline and the
281- conversations or trajectories that are leading to the unexpected behavior.
282- They will often give you useful insights.
279+ - ** Symptom:** Something isn't working.
280+ - ** Solution:** Ask an AI coding tool to review your pipeline and the
281+ conversations or trajectories that are leading to the unexpected behavior.
282+ They will often give you useful insights.
283283
284284This is not an officially supported Google product. This project is not eligible
285285for the
0 commit comments