The numbered directories follow the released pipeline:
step1: render and validate model-facing prompts;step2_runner: run the 11 tested models;step3: build judge inputs, run judges, parse, and validate;step4: aggregate three judge decisions and apply documented manual F resolution;step5: validation-stage aggregate and human/judge coherence utilities;step6: main RPCBench metrics and figure;step7: released downstream analyses and evidence ablations.