Official repository for RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation.
RPCBench evaluates whether LLM-based recommendation assistants proactively detect, localize, and appropriately handle faulty premises while remaining faithful to the evidence visible in the recommendation context.
This repository contains the complete 4,623-instance RPCBench benchmark. It releases all 4,623 paired-query records, all rendered tested-model inputs, 50,853 tested-model response records for 11 models, sample-level three-judge aggregates, main metric outputs, and the downstream analyses supported by the released artifacts.
RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation
The public preprint will be linked here upon release.
| Item | Count |
|---|---|
| Final benchmark instances | 4,623 |
| Paired-query records | 4,623 |
| Recommendation domains | 5 |
| Premise-failure types | 10 |
| Evaluated models | 11 |
| Sample-model response records | 50,853 |
| Judge models in the reported main experiment | 3 |
| Numeric sample-model-judge votes represented in the aggregate | 152,559 |
The paper describes an initial pool of 6,250 generated candidates and a final set of 4,623 instances after cross-model filtering and global deduplication.
| Code | Premise-failure type | Instances |
|---|---|---|
| U1 | Missing decision-critical slot | 600 |
| U2 | Insufficient personalization basis | 700 |
| I1 | Internal query contradiction | 416 |
| I2 | Contradiction with visible user-side evidence | 598 |
| I3 | Contradiction with visible candidate-side evidence | 545 |
| I4 | Contradiction with visible snapshot-internal state | 308 |
| X1 | Out-of-schema attribute premise | 590 |
| X2 | Out-of-snapshot or open-world premise | 430 |
| B1 | Capability-boundary request | 218 |
| B2 | Safety or compliance-boundary request | 218 |
| Total | 4,623 |
| Source dataset/domain | Instances |
|---|---|
| MovieLens-1M | 462 |
| MIND-small | 703 |
| Yelp Local | 1,310 |
| Amazon Sports | 1,270 |
| Goodreads Dual-Domain | 878 |
| Total | 4,623 |
The released experiment covers GPT-5.5, Claude Sonnet 4.6, Gemini 3.1 Pro Preview, DeepSeek V4 Pro, DeepSeek V4 Flash, Qwen3.5 Plus, Qwen3.5-397B-A17B, Qwen3.5-122B-A10B, Qwen3.5-35B-A3B, Llama-3.1-8B-Instruct, and Llama-3.1-70B-Instruct.
The three main judge models are GPT-5.5, Claude Sonnet 4.6, and Gemini 3.1 Pro Preview. The exact tested-model and judge prompts are under prompts/.
RPCBench reports four dimensions:
- D — Detection: whether the model recognizes a premise failure;
- L — Localization: whether it identifies the correct root cause;
- S — Strategy: whether it repairs, clarifies, conditions, or blocks appropriately;
- F — Evidence faithfulness: whether its claims remain faithful to the visible evidence.
D is binary. L, S, and F use a 0–2 scale. The complete scoring rules and strategy labels are in prompts/judge_prompt.txt.
benchmark_data/full/ complete benchmark split by 10 error types
configs/ runtime and configuration examples
prompts/ tested-model and judge prompt templates
docs/ data, schema, reproduction, and artifact documentation
src/step1/ prompt rendering and validation
src/step2_runner/ tested-model inference
src/step3/ judge input, execution, parsing, and validation
src/step4/ three-judge aggregation
src/step5/ aggregate validation and human/judge coherence
src/step6/ main metrics and figure
src/step7/ analyses actually mapped to the paper
outputs/step1_recllm_inputs/ all 4,623 rendered inputs and audit mappings
outputs/step3_judge/ all 50,853 tested-model response/judge-input records
outputs/step4_aggregate/ aggregation-stage audit artifact
outputs/step5_judge_validation/ canonical aggregate consumed by metrics
outputs/step6_metrics/ main tables and plots
outputs/step7_analysis/ adopted downstream analysis outputs
tools/ manifest generation and release validation
| Purpose | Canonical path | Records |
|---|---|---|
| Full benchmark | benchmark_data/full/*/merged_true_samples_full.jsonl |
4,623 |
| Rendered tested-model inputs | outputs/step1_recllm_inputs/recllm_prompts.jsonl |
4,623 |
| Input-to-sample audit map | outputs/step1_recllm_inputs/recllm_audit_map.jsonl |
4,623 |
| Tested-model responses embedded in judge inputs | outputs/step3_judge/judge_inputs/full/by_error_type/ |
50,853 |
| Canonical metric input | outputs/step5_judge_validation/aggregate_3judge.jsonl |
50,853 |
| Main metrics | outputs/step6_metrics/ |
11 models plus group breakdowns |
| Adopted paper analyses | outputs/step7_analysis/ |
analysis-specific |
outputs/step4_aggregate/sample_level_3judge_aggregate.jsonl is the adjacent aggregation-stage audit copy. The downstream code consistently reads the validated Step 5 aggregate above; the two are not presented as separate benchmark versions.
Each benchmark JSONL object contains:
sample_index,dataset, andunit_id;target_error_profile;visible_payload;generated_queries.correct_query;generated_queries.corrupted_query;generated_queries.brief_rationale;- construction and reviewer metadata.
The tested model receives neither the clean query nor the gold error metadata. Its exact messages are released in outputs/step1_recllm_inputs/recllm_prompts.jsonl. Field visibility, source-native evidence, fixed windows, truncation, and leakage-only removal are documented in docs/SCHEMA.md and docs/DATA_CARD.md.
The validated release environment uses Python 3.11.
python -m venv .venv
python -m pip install --upgrade pip
python -m pip install -r requirements.txtOn Windows PowerShell, activate with .venv\Scripts\Activate.ps1; on macOS/Linux, use source .venv/bin/activate.
python tools/validate_release.pyThis streams the large JSONL files, verifies exact counts and required outputs, and checks every file against ARTIFACT_MANIFEST.csv. Use --skip-checksums for a faster structural check. Progress is appended to live_console.log.
python src/step6/step6_dual_axis_metrics.py
python src/step6/step6_plot_dual_axis.pyThese commands consume the released 50,853-row Step 5 aggregate and audit map. Exact Step 7 mappings and commands are in docs/PAPER_MAPPING.md and docs/REPRODUCTION.md.
docs/DATA_CARD.md: benchmark purpose, composition, construction, and intended use;docs/SCHEMA.md: field visibility, source-native schema, evidence windows, and truncation;docs/REPRODUCTION.md: stage-by-stage commands and prerequisites;docs/ARTIFACT_GUIDE.md: canonical, supporting, and unavailable artifacts;docs/PAPER_MAPPING.md: adopted analyses mapped to exact code and outputs;docs/REORGANIZATION_NOTES.md: old-to-new structure and organization decisions.
Citation information will be added once the public preprint is available.