Zhenchen Tang1,2,4 · Yang Li1,2,4 · Songlin Yang3,4,† · Bo Peng1,2 · Xiaotong Zhao4 · Shuai Li4 · Haotian Fan4 · Alan Zhao4 · Jing Dong1,2,*
1New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences
2School of Artificial Intelligence, University of Chinese Academy of Sciences
3The Hong Kong University of Science and Technology
4Tencent
† Project Lead
* Corresponding Author
Unconstrained scorers suffer from scalar drift. RewardVerse inserts a dynamic rubric — themes, weights, and scoring tips generated from the query alone — as an intermediate representation that anchors the scoring process.
Video reward models usually map a video straight to one scalar, with no explicit notion of what is being checked. The scoring scale then collapses into a narrow high band or shifts across prompts — a failure mode we call scalar drift.
RewardVerse treats the rubric as a learned intermediate representation:
-
Dynamic rubric generator — from the evaluation query
$(d, p)$ only (never the video), it emits a schema-constrained rubric of themes$t_k$ , weights$w_k$ (summing to 1), and execution tips$T_k$ . -
Rubric-guided scorer — scores each theme with a soft-logits readout, i.e. the expected value over the logits of rating tokens
1–5, then aggregates the per-theme scores by their weights. This avoids fragile free-form score parsing and keeps the score continuous. - Rubric-Guided Policy Optimization (RGPO) — a two-stage GRPO recipe that (i) warms up the scorer against self-evolving seed rubrics and (ii) jointly optimizes the rubric generator, while a human-aligned margin calibration loss keeps the scorer's score gaps faithful to human ratings.
Stage 1 — Seed-guided scorer warm-up. A frontier MLLM proposes → verifies → revises rubrics over 30 preference pairs per dimension; verified rubrics are deduplicated and reduced to
Stage 2 — Joint policy optimization. The generator samples
The implementation is being cleaned up and will be released upon acceptance — training code, preference splits, seed rubrics, trained reward models, and evaluation scripts.
Please watch / star this repository to be notified when the release lands.
@misc{tang2026rewardverserubricguidedpolicyoptimization,
title={RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling},
author={Zhenchen Tang and Yang Li and Songlin Yang and Bo Peng and Xiaotong Zhao and Shuai Li and Haotian Fan and Alan Zhao and Jing Dong},
year={2026},
eprint={2609.22947},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.22947},
}We thank the maintainers of EvalVerse, VideoGen-RewardBench (VGRB), VBench, Qwen2.5-VL, and Wan-2.2, whose datasets, benchmarks and models made this study possible.
The released code and models will be provided under the Apache-2.0 license. The paper text and figures are shared for academic use; please cite the paper if you use them.
