|
| 1 | +--- |
| 2 | +name: passnet-feedback |
| 3 | +description: > |
| 4 | + Fast, deterministic feedback for PassNet work: pre-flight pattern/match verification |
| 5 | + WITHOUT burning a GPU evaluation, per-node bottleneck analysis of a sample's graphs, |
| 6 | + and eval-log parsing into per-variant status + estimated score + failure classification. |
| 7 | + Use BEFORE every GPU evaluation (check_pattern), at the START of a sample |
| 8 | + (analyze_graph), and AFTER every evaluation (parse_eval_log). |
| 9 | +--- |
| 10 | + |
| 11 | +GPU evaluations cost minutes and are rate-limited; these tools answer in seconds. The |
| 12 | +iteration loop is: `analyze_graph` once → author passes → `check_pattern` until green → |
| 13 | +GPU evaluate → `parse_eval_log` → fix the top issue → repeat. |
| 14 | + |
| 15 | +All scripts live in the `scripts/` directory next to this `SKILL.md`. In this repo, prefer |
| 16 | +`SCRIPTS=<repo>/.claude/skills/passnet-feedback/scripts`. Invoke |
| 17 | +`python3 $SCRIPTS/<name>.py ...` from anywhere. They need: the PassNet repo importable |
| 18 | +(auto-detected from the sample dir's `entry.sh` symlink, or set |
| 19 | +`PYTHONPATH`/`PASSNET_ROOT`), torch, and for smoke/bench steps a GPU |
| 20 | +(`CUDA_VISIBLE_DEVICES` respected; they degrade to CPU structure-checks without one). |
| 21 | + |
| 22 | +## 1. `analyze_graph.py` — where is the bottleneck, what is matchable |
| 23 | + |
| 24 | +```bash |
| 25 | +python3 $SCRIPTS/analyze_graph.py --sample-dir <sample_root> [--bench] [--max-variants N] |
| 26 | +# default --max-variants 3, dtype-spread |
| 27 | +``` |
| 28 | +- Captures the REAL dynamo graph of each variant (exactly what passes must match). |
| 29 | +- Per node: op kind, target, written args/kwargs form, output shape/dtype, and callable |
| 30 | + pattern **matchability** (`method`/`C-bound` = mirror exactly; `PY-sig positional` = |
| 31 | + matchable; `PY-sig kwargs/partial` = normal callable pattern will normalize differently). |
| 32 | + A high-value kwargs-form Python functional region may still be recoverable with an exact |
| 33 | + manual FX `GraphModule` pattern; use passnet-pattern-fusion and confirm with |
| 34 | + `check_pattern.py`. |
| 35 | +- `--bench` (GPU): per-node eager µs via an instrumented FX interpreter + whole-forward |
| 36 | + eager e2e estimate → tells you the absorbable time and a ROUGH speedup ceiling. The |
| 37 | + ceiling assumes a ~70 µs fixed tax — real tax grows with FX node count (150–230 µs on |
| 38 | + 13-node graphs); treat the printed ceiling as optimistic and recalibrate after your |
| 39 | + first eval. |
| 40 | +- Use the region table at the bottom as the starting plan: it lists maximal runs of |
| 41 | + matchable nodes with their absorbed-µs totals. |
| 42 | + |
| 43 | +## 2. `check_pattern.py` — pre-flight a pass_dir in seconds (run before EVERY eval) |
| 44 | + |
| 45 | +```bash |
| 46 | +python3 $SCRIPTS/check_pattern.py --sample-dir <sample_root> [--pass-dir <dir>] [--smoke [--bench]] |
| 47 | +# --max-variants N (default 8, dtype-spread; 0 = all). --bench implies --smoke. |
| 48 | +# --smoke/--bench run on ONE variant per dtype; matching runs on all chosen variants. |
| 49 | +``` |
| 50 | +Replicates the harness pipeline faithfully and reports: |
| 51 | +1. JSON manifest sanity (names ↔ files). |
| 52 | +2. AST validation per pass (the real `validate_pass_source` when repo importable). |
| 53 | +3. `replacement_func()` stability + **distinct-function count** (must be 1 when >1 pass — |
| 54 | + otherwise passes WILL be silently dropped by the limit). |
| 55 | +4. Pattern trace → real `SubgraphMatcher` against each variant's dynamo graph: |
| 56 | + match count per pass per variant, single-output check, containment check, and on |
| 57 | + mismatch a node-by-node nearest-miss diff (pattern form vs graph form). |
| 58 | +5. Verdict per variant: would the run match (≥1 pass), or early-exit at 0.1. |
| 59 | +6. `--smoke` (GPU): actually applies the passes (real `PassMgrBackend`) and runs one |
| 60 | + poisoned warmup call + numeric comparison vs eager at the dtype's baseline tolerance — |
| 61 | + catches Unauthorized-Operator, dtype mismatch, and gross numeric bugs pre-eval. |
| 62 | +7. `--bench` (GPU): micro-benchmarks compiled-vs-eager e2e (200 calls) per variant for a |
| 63 | + speedup preview (no dynamo guards, so real eval ≈ a few % worse). |
| 64 | +
|
| 65 | +CPU-only output can validate structure and matchability triage, but it does not prove poisoned-wrapper legality, dtype behavior, numeric correctness, or real speed. GPU smoke/bench or a completed evaluation is still required for those claims. |
| 66 | +
|
| 67 | +`--smoke` is a one-shot semantic check, not a full harness proof. Be extra cautious when a |
| 68 | +replacement absorbs `inplace=True` nodes or other side-effectful behavior: a one-shot compare |
| 69 | +can pass while repeated benchmark calls expose mutation/aliasing differences. For those |
| 70 | +regions, either leave the in-place node outside the pattern or confirm repeated-call behavior |
| 71 | +with a harness-style check/completed evaluation. |
| 72 | +
|
| 73 | +Green output = safe to spend a GPU evaluation. Any red line tells you which skill to open: |
| 74 | +match issues → passnet-pattern-fusion §7; poison/API issues → passnet-pattern-fusion §6; |
| 75 | +numeric issues → passnet-triton-opt §3. |
| 76 | +
|
| 77 | +## 3. `parse_eval_log.py` — turn an eval into decisions |
| 78 | +
|
| 79 | +```bash |
| 80 | +python3 $SCRIPTS/parse_eval_log.py <validation.log | evaluate_response.json> |
| 81 | +# also reads the "stdout"/"stderr" fields of a saved /evaluate JSON response |
| 82 | +``` |
| 83 | +Prints: |
| 84 | +- Per-variant table: dtype | status | e2e/gpu speedup | eager/compiled medians | max_diff | |
| 85 | + passes applied/failed. |
| 86 | +- Failure classification with the fix pointer: |
| 87 | + - localhost connection failure from a Codex managed sandbox → retry normal bounded curl once, then retry through the approved escalation path if available before diagnosing service downtime. |
| 88 | + - empty, interrupted, HTTP000-style, status-corrupted, or non-JSON `/evaluate` response → no completed eval; leave score, speed, correctness, and pass-matched metrics null, retry service access/upload state if appropriate, and do not interpret any performance or correctness result. |
| 89 | + - `no pass matched` + diagnostic lines → pattern form problem (which pass, which node). |
| 90 | + - `AssertionError` in `_replace_pattern` → multi-output pattern (rule 1). |
| 91 | + - `Unauthorized Operator (aten.xxx)` → illegal torch op in wrapper (poison). |
| 92 | + - `Detected hacking behavior` → AST validation rejected a file (it was NOT loaded). |
| 93 | + - `Loaded N passes` < listed → replacement_func limit dropped passes (shared dispatch!). |
| 94 | + - dtype mismatch / accuracy with max_diff vs that dtype's baseline tolerance. |
| 95 | + - replacement crash mentioning a returning node with no users → the pattern likely matched |
| 96 | + userless/dead duplicate work; re-anchor the region on an observable value or use a proven |
| 97 | + graph-rewrite approach instead of a normal output replacement. |
| 98 | + - evaluator/environment fluctuation messages, timing-stability errors, or rerun requests |
| 99 | + with otherwise clean matching and successful-variant correctness → classify separately |
| 100 | + from numeric, no-match, unauthorized-operator, timeout, and OOM. If the pass pre-flights |
| 101 | + cleanly and successful variants are correct, an unchanged rerun can be a valid use of |
| 102 | + remaining evaluation budget. |
| 103 | + - timeout/OOM hints. |
| 104 | +- **Estimated sample score** using the real ES(t) weights, plus per-variant rectified |
| 105 | + speedups — so you know the score impact of fixing each failing variant before re-running. |
| 106 | +
|
| 107 | +For round or multi-worker aggregation, keep family labels separate: |
| 108 | +
|
| 109 | +- `triage_family`: the static family assigned before implementation. |
| 110 | +- `worker_confirmed_family`: the family the worker confirmed after reading the graph and |
| 111 | + pre-flight results. |
| 112 | +- `actual_winning_region_family`: the family of the best completed-evaluation state. |
| 113 | +
|
| 114 | +Do not aggregate a win under the original triage family if the best completed state belongs |
| 115 | +to a different region family. Round summaries should report, per family: number of completed |
| 116 | +workers, score-above-eager count, correct-but-slow count, numeric-preflight-blocked count, |
| 117 | +correctness-regressed-on-larger-region count, larger-region-attempted count, and evaluator |
| 118 | +instability count. Keep evaluator/environment fluctuation separate from numeric, no-match, |
| 119 | +unauthorized-operator, timeout, and OOM. |
| 120 | +
|
| 121 | +## 4. Reading raw logs yourself (when needed) |
| 122 | +
|
| 123 | +Key markers in `stdout`/`validation.log`: |
| 124 | +`[PassMgrBackend] Loaded/Applied/failed to match/Diagnostic`, |
| 125 | +`[Result] status: success|failed`, `[Speedup][e2e]/[gpu]`, |
| 126 | +`[Performance][eager|compiled]` (medians), `[Correctness][max_diff|mean_diff|equal]`, |
| 127 | +`[Datatype][eager|compiled]`, `debug-model-execution <ExcType>` (crash), |
| 128 | +`Has Any pass matched? [True|False]`, `aggregated_speedup=...`. |
| 129 | +The service strips `Trial`/`[Profiling]`/`all_close` lines; everything above survives. |
| 130 | +
|
| 131 | +## 5. Bottleneck decision guide (after analyze/eval) |
| 132 | +
|
| 133 | +| signal | diagnosis | action | |
| 134 | +|---|---|---| |
| 135 | +| eager e2e < 150 µs, ≤2 matchable cheap nodes | overhead-bound, ceiling < 1 | floor pass, move on | |
| 136 | +| big gap between eager e2e and Σ(node µs) | per-call fixed costs dominate | absorb more nodes per launch; nothing else helps | |
| 137 | +| one node ≥ 60% of eager time, matchable | real kernel target | fuse it + its neighbors; tune via passnet-triton-opt | |
| 138 | +| a `conv` dominates AND `stride == kernel_size`, `kernel_size > 1`, `padding == 0` (non-overlapping) | disjoint windows ⇒ exactly a dense patch matmul; cuDNN pays for a general im2col path | reformulate as a `tl.dot` patch-gather kernel + fused tail (kernel-templates §13), AFTER banking a tail floor. Don't dismiss as "cuDNN, leave alone" or "unmatchable" | |
| 139 | +| a `conv`/`matmul` dominates in any OTHER form (1×1, depthwise/grouped, overlapping conv, general dense matmul) | vendor sweet spot | default: leave it in aten; fuse its cheap tail (bias/BN/act/residual). Rewriting usually loses — try only if parameters look off-regime and a completed eval beats the floor | |
| 140 | +| one node ≥ 60%, callable-unmatchable kwargs form | try exact manual FX if the region is single-output and valuable; otherwise fuse what's left | expectation depends on whether the exact pattern pre-flight matches | |
| 141 | +| compiled gpu ≈ e2e, both < 1 | stream gaps (launch-bound) | fewer launches/allocs | |
| 142 | +| compiled gpu > 1, e2e < 1 | host overhead | drop autotune churn, simplify wrapper, fewer passes | |
| 143 | +| some dtype variants fail only | numeric fidelity | passnet-triton-opt §3 recipes | |
| 144 | +| timeout (600 s) | too many variants × compile/tune cost | remove autotune, single config, fewer passes | |
| 145 | +| evaluator reports timing/environment fluctuation while successful variants match and pass correctness | evaluator instability, not proven kernel bug | rerun unchanged if budget remains; record separately from numeric/no-match/unauthorized | |
0 commit comments