Skip to content

Repository files navigation

RPCBench

Official repository for RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation.

RPCBench evaluates whether LLM-based recommendation assistants proactively detect, localize, and appropriately handle faulty premises while remaining faithful to the evidence visible in the recommendation context.

This repository contains the complete 4,623-instance RPCBench benchmark. It releases all 4,623 paired-query records, all rendered tested-model inputs, 50,853 tested-model response records for 11 models, sample-level three-judge aggregates, main metric outputs, and the downstream analyses supported by the released artifacts.

Paper

RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation

The public preprint will be linked here upon release.

At a glance

Item Count
Final benchmark instances 4,623
Paired-query records 4,623
Recommendation domains 5
Premise-failure types 10
Evaluated models 11
Sample-model response records 50,853
Judge models in the reported main experiment 3
Numeric sample-model-judge votes represented in the aggregate 152,559

The paper describes an initial pool of 6,250 generated candidates and a final set of 4,623 instances after cross-model filtering and global deduplication.

Benchmark composition

Code Premise-failure type Instances
U1 Missing decision-critical slot 600
U2 Insufficient personalization basis 700
I1 Internal query contradiction 416
I2 Contradiction with visible user-side evidence 598
I3 Contradiction with visible candidate-side evidence 545
I4 Contradiction with visible snapshot-internal state 308
X1 Out-of-schema attribute premise 590
X2 Out-of-snapshot or open-world premise 430
B1 Capability-boundary request 218
B2 Safety or compliance-boundary request 218
Total 4,623
Source dataset/domain Instances
MovieLens-1M 462
MIND-small 703
Yelp Local 1,310
Amazon Sports 1,270
Goodreads Dual-Domain 878
Total 4,623

Evaluated models

The released experiment covers GPT-5.5, Claude Sonnet 4.6, Gemini 3.1 Pro Preview, DeepSeek V4 Pro, DeepSeek V4 Flash, Qwen3.5 Plus, Qwen3.5-397B-A17B, Qwen3.5-122B-A10B, Qwen3.5-35B-A3B, Llama-3.1-8B-Instruct, and Llama-3.1-70B-Instruct.

The three main judge models are GPT-5.5, Claude Sonnet 4.6, and Gemini 3.1 Pro Preview. The exact tested-model and judge prompts are under prompts/.

Evaluation dimensions

RPCBench reports four dimensions:

  • D — Detection: whether the model recognizes a premise failure;
  • L — Localization: whether it identifies the correct root cause;
  • S — Strategy: whether it repairs, clarifies, conditions, or blocks appropriately;
  • F — Evidence faithfulness: whether its claims remain faithful to the visible evidence.

D is binary. L, S, and F use a 0–2 scale. The complete scoring rules and strategy labels are in prompts/judge_prompt.txt.

Repository layout

benchmark_data/full/             complete benchmark split by 10 error types
configs/                         runtime and configuration examples
prompts/                         tested-model and judge prompt templates
docs/                            data, schema, reproduction, and artifact documentation
src/step1/                       prompt rendering and validation
src/step2_runner/                tested-model inference
src/step3/                       judge input, execution, parsing, and validation
src/step4/                       three-judge aggregation
src/step5/                       aggregate validation and human/judge coherence
src/step6/                       main metrics and figure
src/step7/                       analyses actually mapped to the paper
outputs/step1_recllm_inputs/     all 4,623 rendered inputs and audit mappings
outputs/step3_judge/             all 50,853 tested-model response/judge-input records
outputs/step4_aggregate/         aggregation-stage audit artifact
outputs/step5_judge_validation/  canonical aggregate consumed by metrics
outputs/step6_metrics/           main tables and plots
outputs/step7_analysis/          adopted downstream analysis outputs
tools/                           manifest generation and release validation

Canonical artifacts

Purpose Canonical path Records
Full benchmark benchmark_data/full/*/merged_true_samples_full.jsonl 4,623
Rendered tested-model inputs outputs/step1_recllm_inputs/recllm_prompts.jsonl 4,623
Input-to-sample audit map outputs/step1_recllm_inputs/recllm_audit_map.jsonl 4,623
Tested-model responses embedded in judge inputs outputs/step3_judge/judge_inputs/full/by_error_type/ 50,853
Canonical metric input outputs/step5_judge_validation/aggregate_3judge.jsonl 50,853
Main metrics outputs/step6_metrics/ 11 models plus group breakdowns
Adopted paper analyses outputs/step7_analysis/ analysis-specific

outputs/step4_aggregate/sample_level_3judge_aggregate.jsonl is the adjacent aggregation-stage audit copy. The downstream code consistently reads the validated Step 5 aggregate above; the two are not presented as separate benchmark versions.

Data record structure

Each benchmark JSONL object contains:

  • sample_index, dataset, and unit_id;
  • target_error_profile;
  • visible_payload;
  • generated_queries.correct_query;
  • generated_queries.corrupted_query;
  • generated_queries.brief_rationale;
  • construction and reviewer metadata.

The tested model receives neither the clean query nor the gold error metadata. Its exact messages are released in outputs/step1_recllm_inputs/recllm_prompts.jsonl. Field visibility, source-native evidence, fixed windows, truncation, and leakage-only removal are documented in docs/SCHEMA.md and docs/DATA_CARD.md.

Install

The validated release environment uses Python 3.11.

python -m venv .venv
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

On Windows PowerShell, activate with .venv\Scripts\Activate.ps1; on macOS/Linux, use source .venv/bin/activate.

Verify the complete release

python tools/validate_release.py

This streams the large JSONL files, verifies exact counts and required outputs, and checks every file against ARTIFACT_MANIFEST.csv. Use --skip-checksums for a faster structural check. Progress is appended to live_console.log.

Recompute the main metrics

python src/step6/step6_dual_axis_metrics.py
python src/step6/step6_plot_dual_axis.py

These commands consume the released 50,853-row Step 5 aggregate and audit map. Exact Step 7 mappings and commands are in docs/PAPER_MAPPING.md and docs/REPRODUCTION.md.

Documentation

  • docs/DATA_CARD.md: benchmark purpose, composition, construction, and intended use;
  • docs/SCHEMA.md: field visibility, source-native schema, evidence windows, and truncation;
  • docs/REPRODUCTION.md: stage-by-stage commands and prerequisites;
  • docs/ARTIFACT_GUIDE.md: canonical, supporting, and unavailable artifacts;
  • docs/PAPER_MAPPING.md: adopted analyses mapped to exact code and outputs;
  • docs/REORGANIZATION_NOTES.md: old-to-new structure and organization decisions.

Citation

Citation information will be added once the public preprint is available.

About

Official repository for RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages