Skip to content

Commit c29af6c

Browse files
authored
Merge pull request #22 from Lolo1222/main
add passnet skills
2 parents ecc4814 + b803a68 commit c29af6c

13 files changed

Lines changed: 3665 additions & 0 deletions

File tree

Lines changed: 145 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,145 @@
1+
---
2+
name: passnet-feedback
3+
description: >
4+
Fast, deterministic feedback for PassNet work: pre-flight pattern/match verification
5+
WITHOUT burning a GPU evaluation, per-node bottleneck analysis of a sample's graphs,
6+
and eval-log parsing into per-variant status + estimated score + failure classification.
7+
Use BEFORE every GPU evaluation (check_pattern), at the START of a sample
8+
(analyze_graph), and AFTER every evaluation (parse_eval_log).
9+
---
10+
11+
GPU evaluations cost minutes and are rate-limited; these tools answer in seconds. The
12+
iteration loop is: `analyze_graph` once → author passes → `check_pattern` until green →
13+
GPU evaluate → `parse_eval_log` → fix the top issue → repeat.
14+
15+
All scripts live in the `scripts/` directory next to this `SKILL.md`. In this repo, prefer
16+
`SCRIPTS=<repo>/.claude/skills/passnet-feedback/scripts`. Invoke
17+
`python3 $SCRIPTS/<name>.py ...` from anywhere. They need: the PassNet repo importable
18+
(auto-detected from the sample dir's `entry.sh` symlink, or set
19+
`PYTHONPATH`/`PASSNET_ROOT`), torch, and for smoke/bench steps a GPU
20+
(`CUDA_VISIBLE_DEVICES` respected; they degrade to CPU structure-checks without one).
21+
22+
## 1. `analyze_graph.py` — where is the bottleneck, what is matchable
23+
24+
```bash
25+
python3 $SCRIPTS/analyze_graph.py --sample-dir <sample_root> [--bench] [--max-variants N]
26+
# default --max-variants 3, dtype-spread
27+
```
28+
- Captures the REAL dynamo graph of each variant (exactly what passes must match).
29+
- Per node: op kind, target, written args/kwargs form, output shape/dtype, and callable
30+
pattern **matchability** (`method`/`C-bound` = mirror exactly; `PY-sig positional` =
31+
matchable; `PY-sig kwargs/partial` = normal callable pattern will normalize differently).
32+
A high-value kwargs-form Python functional region may still be recoverable with an exact
33+
manual FX `GraphModule` pattern; use passnet-pattern-fusion and confirm with
34+
`check_pattern.py`.
35+
- `--bench` (GPU): per-node eager µs via an instrumented FX interpreter + whole-forward
36+
eager e2e estimate → tells you the absorbable time and a ROUGH speedup ceiling. The
37+
ceiling assumes a ~70 µs fixed tax — real tax grows with FX node count (150–230 µs on
38+
13-node graphs); treat the printed ceiling as optimistic and recalibrate after your
39+
first eval.
40+
- Use the region table at the bottom as the starting plan: it lists maximal runs of
41+
matchable nodes with their absorbed-µs totals.
42+
43+
## 2. `check_pattern.py` — pre-flight a pass_dir in seconds (run before EVERY eval)
44+
45+
```bash
46+
python3 $SCRIPTS/check_pattern.py --sample-dir <sample_root> [--pass-dir <dir>] [--smoke [--bench]]
47+
# --max-variants N (default 8, dtype-spread; 0 = all). --bench implies --smoke.
48+
# --smoke/--bench run on ONE variant per dtype; matching runs on all chosen variants.
49+
```
50+
Replicates the harness pipeline faithfully and reports:
51+
1. JSON manifest sanity (names ↔ files).
52+
2. AST validation per pass (the real `validate_pass_source` when repo importable).
53+
3. `replacement_func()` stability + **distinct-function count** (must be 1 when >1 pass —
54+
otherwise passes WILL be silently dropped by the limit).
55+
4. Pattern trace → real `SubgraphMatcher` against each variant's dynamo graph:
56+
match count per pass per variant, single-output check, containment check, and on
57+
mismatch a node-by-node nearest-miss diff (pattern form vs graph form).
58+
5. Verdict per variant: would the run match (≥1 pass), or early-exit at 0.1.
59+
6. `--smoke` (GPU): actually applies the passes (real `PassMgrBackend`) and runs one
60+
poisoned warmup call + numeric comparison vs eager at the dtype's baseline tolerance —
61+
catches Unauthorized-Operator, dtype mismatch, and gross numeric bugs pre-eval.
62+
7. `--bench` (GPU): micro-benchmarks compiled-vs-eager e2e (200 calls) per variant for a
63+
speedup preview (no dynamo guards, so real eval ≈ a few % worse).
64+
65+
CPU-only output can validate structure and matchability triage, but it does not prove poisoned-wrapper legality, dtype behavior, numeric correctness, or real speed. GPU smoke/bench or a completed evaluation is still required for those claims.
66+
67+
`--smoke` is a one-shot semantic check, not a full harness proof. Be extra cautious when a
68+
replacement absorbs `inplace=True` nodes or other side-effectful behavior: a one-shot compare
69+
can pass while repeated benchmark calls expose mutation/aliasing differences. For those
70+
regions, either leave the in-place node outside the pattern or confirm repeated-call behavior
71+
with a harness-style check/completed evaluation.
72+
73+
Green output = safe to spend a GPU evaluation. Any red line tells you which skill to open:
74+
match issues → passnet-pattern-fusion §7; poison/API issues → passnet-pattern-fusion §6;
75+
numeric issues → passnet-triton-opt §3.
76+
77+
## 3. `parse_eval_log.py` — turn an eval into decisions
78+
79+
```bash
80+
python3 $SCRIPTS/parse_eval_log.py <validation.log | evaluate_response.json>
81+
# also reads the "stdout"/"stderr" fields of a saved /evaluate JSON response
82+
```
83+
Prints:
84+
- Per-variant table: dtype | status | e2e/gpu speedup | eager/compiled medians | max_diff |
85+
passes applied/failed.
86+
- Failure classification with the fix pointer:
87+
- localhost connection failure from a Codex managed sandbox → retry normal bounded curl once, then retry through the approved escalation path if available before diagnosing service downtime.
88+
- empty, interrupted, HTTP000-style, status-corrupted, or non-JSON `/evaluate` response → no completed eval; leave score, speed, correctness, and pass-matched metrics null, retry service access/upload state if appropriate, and do not interpret any performance or correctness result.
89+
- `no pass matched` + diagnostic lines → pattern form problem (which pass, which node).
90+
- `AssertionError` in `_replace_pattern` → multi-output pattern (rule 1).
91+
- `Unauthorized Operator (aten.xxx)` → illegal torch op in wrapper (poison).
92+
- `Detected hacking behavior` → AST validation rejected a file (it was NOT loaded).
93+
- `Loaded N passes` < listed → replacement_func limit dropped passes (shared dispatch!).
94+
- dtype mismatch / accuracy with max_diff vs that dtype's baseline tolerance.
95+
- replacement crash mentioning a returning node with no users → the pattern likely matched
96+
userless/dead duplicate work; re-anchor the region on an observable value or use a proven
97+
graph-rewrite approach instead of a normal output replacement.
98+
- evaluator/environment fluctuation messages, timing-stability errors, or rerun requests
99+
with otherwise clean matching and successful-variant correctness → classify separately
100+
from numeric, no-match, unauthorized-operator, timeout, and OOM. If the pass pre-flights
101+
cleanly and successful variants are correct, an unchanged rerun can be a valid use of
102+
remaining evaluation budget.
103+
- timeout/OOM hints.
104+
- **Estimated sample score** using the real ES(t) weights, plus per-variant rectified
105+
speedups — so you know the score impact of fixing each failing variant before re-running.
106+
107+
For round or multi-worker aggregation, keep family labels separate:
108+
109+
- `triage_family`: the static family assigned before implementation.
110+
- `worker_confirmed_family`: the family the worker confirmed after reading the graph and
111+
pre-flight results.
112+
- `actual_winning_region_family`: the family of the best completed-evaluation state.
113+
114+
Do not aggregate a win under the original triage family if the best completed state belongs
115+
to a different region family. Round summaries should report, per family: number of completed
116+
workers, score-above-eager count, correct-but-slow count, numeric-preflight-blocked count,
117+
correctness-regressed-on-larger-region count, larger-region-attempted count, and evaluator
118+
instability count. Keep evaluator/environment fluctuation separate from numeric, no-match,
119+
unauthorized-operator, timeout, and OOM.
120+
121+
## 4. Reading raw logs yourself (when needed)
122+
123+
Key markers in `stdout`/`validation.log`:
124+
`[PassMgrBackend] Loaded/Applied/failed to match/Diagnostic`,
125+
`[Result] status: success|failed`, `[Speedup][e2e]/[gpu]`,
126+
`[Performance][eager|compiled]` (medians), `[Correctness][max_diff|mean_diff|equal]`,
127+
`[Datatype][eager|compiled]`, `debug-model-execution <ExcType>` (crash),
128+
`Has Any pass matched? [True|False]`, `aggregated_speedup=...`.
129+
The service strips `Trial`/`[Profiling]`/`all_close` lines; everything above survives.
130+
131+
## 5. Bottleneck decision guide (after analyze/eval)
132+
133+
| signal | diagnosis | action |
134+
|---|---|---|
135+
| eager e2e < 150 µs, ≤2 matchable cheap nodes | overhead-bound, ceiling < 1 | floor pass, move on |
136+
| big gap between eager e2e and Σ(node µs) | per-call fixed costs dominate | absorb more nodes per launch; nothing else helps |
137+
| one node ≥ 60% of eager time, matchable | real kernel target | fuse it + its neighbors; tune via passnet-triton-opt |
138+
| a `conv` dominates AND `stride == kernel_size`, `kernel_size > 1`, `padding == 0` (non-overlapping) | disjoint windows ⇒ exactly a dense patch matmul; cuDNN pays for a general im2col path | reformulate as a `tl.dot` patch-gather kernel + fused tail (kernel-templates §13), AFTER banking a tail floor. Don't dismiss as "cuDNN, leave alone" or "unmatchable" |
139+
| a `conv`/`matmul` dominates in any OTHER form (1×1, depthwise/grouped, overlapping conv, general dense matmul) | vendor sweet spot | default: leave it in aten; fuse its cheap tail (bias/BN/act/residual). Rewriting usually loses — try only if parameters look off-regime and a completed eval beats the floor |
140+
| one node ≥ 60%, callable-unmatchable kwargs form | try exact manual FX if the region is single-output and valuable; otherwise fuse what's left | expectation depends on whether the exact pattern pre-flight matches |
141+
| compiled gpu ≈ e2e, both < 1 | stream gaps (launch-bound) | fewer launches/allocs |
142+
| compiled gpu > 1, e2e < 1 | host overhead | drop autotune churn, simplify wrapper, fewer passes |
143+
| some dtype variants fail only | numeric fidelity | passnet-triton-opt §3 recipes |
144+
| timeout (600 s) | too many variants × compile/tune cost | remove autotune, single config, fewer passes |
145+
| evaluator reports timing/environment fluctuation while successful variants match and pass correctness | evaluator instability, not proven kernel bug | rerun unchanged if budget remains; record separately from numeric/no-match/unauthorized |

0 commit comments

Comments
 (0)