ViTeX-Bench Leaderboard

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

  • Xinghao Chen
  • Xiangbo Gao
  • Jiongze Yu
  • Yuheng Wu
  • Zhengzhong Tu

Texas A&M University

NeurIPS 2026Track on Evaluations and Datasets

Change the word, keep the scene

Each task gives a real video, a mask over its text and a source–target string pair. The editor must render the target string inside the mask and leave the rest of the scene, and its motion, unchanged. Below: raw ViTeX-Edit-14B output on frozen evaluation clips.

ViTeX-Edit-14B Source + mask

The line sweeps from the source to the edit; drag it to compare. Scenes change on their own.

Abstract

Recent video generation is increasingly realistic and controllable, yet video editing remains comparatively underdeveloped, particularly for precise local edits that must preserve the original scene dynamics.

Video scene text editing aims to replace text appearing on scene surfaces in a video, such as storefront signs, whiteboards, and product labels, while preserving the surrounding content, motion, and camera dynamics. Although scene text editing has been extensively studied for static images, video scene text editing that achieves high visual quality, temporal consistency, and edit locality remains largely underexplored. Existing resources offer limited paired real-video data, and general video-editing metrics do not directly measure whether the requested text remains correct over time. We introduce ViTeX-Bench, a benchmark suite comprising ViTeX-Dataset and a three-axis evaluation protocol. The dataset contains 387 real-world 720p videos with text-region masks and editing instructions: 230 provide reviewed, pipeline-generated paired edits for training, and 157 form a frozen evaluation split. The protocol evaluates text correctness, visual and temporal quality, and edit locality through 13 metrics, with one primary metric per axis and a Pareto comparison of their trade-offs. OCR calibration, human evaluation, and annotation-sensitivity analyses support the interpretation of these scores.

Across eight baselines from four editing families, accurate text, temporal stability, and scene preservation remain difficult to achieve together. We also release ViTeX-Edit-14B, an open-source reference editor fine-tuned on the paired training split with motion-aligned glyph-video conditioning. It achieves CharAcc 0.688, the highest mean among the evaluated video-native editors, and the lowest comparable text-crop Warp among raw editor outputs. ViTeX-Bench provides a reproducible foundation for studying these trade-offs in video scene text editing.

Every method, one scene

Eight baselines from four families and the reference editor, run on the same evaluation clips. Pick a method and drag the line, or see all of them at once.

ViTeX-Edit-14B Source + mask

What a single score would miss

Four diagnosed baseline failures, one per editing family. Each looks acceptable on one axis and is caught by another, which is why ViTeX-Bench reports all three.

  1. Stable ≠ edited

    Wan2.1-VACE-14B C, mask-conditioned inpainting

    NEW OLD

    The masked region comes back essentially unchanged and still reads NEW. Frames agree with each other and the scene is preserved, so stability and fidelity alone would score it well.

    • TTS high
    • SeqAcc = 0
    Source, mask highlighted
    Output: text unchanged
  2. Correct per frame ≠ stable

    FLUX-Text A, per-frame image editing

    AWESOME WELCOME

    Each frame is edited on its own by a strong image editor, so single frames read nearly right, but the glyphs change shape from one frame to the next.

    • CharAcc high
    • TTS low
    • Warpc 13.01
    Source, mask highlighted
    Output: glyphs flicker and drift
  3. Looks good ≠ correct

    Kling Video 3.0 Omni D, instruction-guided editing

    SOC COC, rendered as ZOO

    A polished, plausible clip that leads both MUSIQ columns, yet it spells the wrong string. Visual quality and a successful replacement are different properties.

    • MUSIQ highest
    • SeqAcc = 0
    Source, mask highlighted
    Output: reads ZOO, not COC
  4. First frame right ≠ local over time

    TextCtrl + AnyV2V B, first-frame edit + propagation

    NDP PSP

    The first-frame edit is right, but as propagation drifts the text fades and the unmasked scene drifts with it. Correctness and locality collapse together.

    • SeqAcc 0.057
    • PSNRloc 21.08
    Source, mask highlighted
    Output: text and scene drift

Three axes, no single score

Every output is scored with 13 metrics on three axes and reported as the full vector. One primary metric per axis defines the comparison, and methods are compared by Pareto dominance rather than by a weighted sum, whose weights would set an arbitrary exchange rate between wrong characters, flicker and a damaged background.

The Pareto front, in three dimensions

Each star is a method on the three primary metrics, from the live leaderboard. The front is the set no other raw editor beats on all three; the brightest star is the ideal corner.

Loading…

  • Pareto front
  • Other editor
  • Reference
  • Ideal

Drag to rotate. Select a star for its scores.

On the front is not the same as good. FLUX-Text, TextCtrl, RS-STE, ViTeX-Edit-14B and Wan2.1-VACE-14B form the Pareto front. Wan2.1-VACE-14B is on it with SeqAcc 0: it returns the masked region almost unchanged, which keeps its temporal and locality scores strong. The other ten metrics stay visible for exactly this reason.

What is not ranked. Composite post-processing and the Source video are shown for reference only. VideoPainter's Flicker and Warp come from interpolated frames, so it is left out of temporal comparisons and of the front.

Do the numbers mean what they say?

  • ρ = 0.95Spearman agreement between method-blinded human transcription and the OCR-based method ranking.
  • +0.71 · −0.40 · −0.53Correlation of three non-author raters' scores (70 outputs) with SeqAcc, Warpc and DreamSimloc.
  • τ = 0.936Mean Kendall agreement of bootstrap-resampled SeqAcc rankings with the full split.
Open the full leaderboard

What is released

Data, a frozen protocol, and a reference model, released together so that results are reproducible and comparable.

ViTeX-Dataset

Real-world videos from Panda-70M and InternVid with per-frame text-region masks and source–target strings. A human-in-the-loop pipeline produced reviewed paired edits for training; the evaluation split is frozen and withholds its edits.

Videos
387
Paired training
230
Frozen evaluation
157
Frames
1280×720 · 120 · 24 fps
Eval scripts
4: Latin, Chinese, Japanese, Cyrillic
Hugging Face dataset

ViTeX-Bench

The evaluation protocol: OCR-anchored text correctness with PP-OCRv5, full-frame and text-crop temporal quality, and locality measured outside the mask. It writes the eval.json the leaderboard accepts.

Metrics
13
Axes
3
Primaries
SeqAcc · Warpc · DreamSimloc
Comparison
Pareto front
Evaluation code on GitHub

ViTeX-Edit-14B

An open reference editor. It extends Wan2.1-VACE-14B with a glyph video: the target string rendered in a matching typeface and warped along the tracked source text, read by every VACE block through an added cross-attention layer.

Backbone
Wan2.1-VACE-14B
Fine-tuning
576 GPU-h · 8×H100
Inference
one 50-step pass
CharAcc
0.688
ViTeX-Dataset construction pipeline: SAM 3 text masks, Qwen3-VL source–target strings with annotator audit, Removal-1.3B clean background, Nano Banana Pro first-frame patch, composed by alpha composition (Strategy A) or a fine-tuned PISCO inserter (Strategy B).
Dataset construction. Four assets per clip: a dilated SAM 3 text mask, a Qwen3-VL string pair audited by an annotator, a clean background from Removal-1.3B, and a first-frame target patch. Static clips can use alpha composition (Strategy A); every clip can use a fine-tuned PISCO inserter (Strategy B). The final split has 56 Strategy-A and 174 Strategy-B videos. Full-size figure.
ViTeX-Edit-14B architecture: target text via uMT5-XXL, source video and mask via the VACE condition unit, and a glyph video encoded by the Wan VAE and a glyph encoder into tokens that every VACE block reads through condition cross-attention.
ViTeX-Edit-14B. Three streams condition the backbone: the target string through the frozen uMT5-XXL encoder, the source video and mask through the VACE condition unit, and the glyph video, pooled by a glyph encoder into 64 tokens. Only the VACE branch, the glyph encoder and the new cross-attention layers are trained. Full-size figure.

Cite

@inproceedings{chen2026vitexbench,  title     = {ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing},  author    = {Chen, Xinghao and Gao, Xiangbo and Yu, Jiongze and Wu, Yuheng and Tu, Zhengzhong},  booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},  year      = {2026},  url       = {https://vitex-bench.github.io/}}

Contact: Xinghao Chen (cxh4242@gmail.com), Zhengzhong Tu