Skip to content

Latest commit

 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ViTeX-Bench

🌐 Project page  ·  📊 Dataset  ·  🧪 Code  ·  🤖 Model weights  ·  🏆 Leaderboard

This repository holds the benchmark's evaluation code and the code of the reference model, ViTeX-Edit-14B: glyph rendering, inference, Composite post-processing and training.

Evaluation pipeline for video scene text editing. A 13-metric, three-axis protocol (text correctness, visual quality, edit locality) on the frozen 157-clip evaluation split of ViTeX-Dataset. The full thirteen-metric vector is the unit of report. One metric per axis is primary — SeqAcc (text correctness), Warp_crop (temporal quality), DreamSim_loc (edit locality) — and the public Leaderboard marks the Pareto set on these three instead of computing an aggregate score.

Accompanies ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing (NeurIPS 2026 Track on Evaluations and Datasets).

Quickstart

git clone https://github.com/taco-group/ViTeX-Bench.git && cd ViTeX-Bench

# Two envs because PaddleOCR conflicts with PyTorch / pyiqa.
conda create -n paddleocr   python=3.12 -y && conda activate paddleocr   && pip install paddleocr opencv-python && conda deactivate
conda create -n vitex-bench python=3.12 -y && conda activate vitex-bench && pip install -r requirements.txt        && conda deactivate

# Drop your method's predictions in baseline_output_videos/<your_method>/<id>.mp4
# (1280×720, 24 fps, 120 frames; one .mp4 per clip id from parsed_records.json)
bash scripts/run_benchmark.sh <your_method>

The runner auto-downloads the ViTeX-Dataset eval split on first run. Output:

  • outputs/<your_method>/eval.json — per-clip metrics + 13 aggregates with 95 % bootstrap CIs.
  • outputs/summary.tsv — one-row-per-baseline TSV across runs.

Reference model: ViTeX-Edit-14B

vitex_edit/ contains everything needed to run the reference editor on your own clips or to reproduce it on the evaluation split: typeface selection, glyph-video rendering, inference with multi-GPU sharding and low-memory modes, the Composite wrapper, and the two-stage training recipe. Weights are on Hugging Face. See vitex_edit/README.md.

Submitting

Open a submission issue on the leaderboard repository and attach the eval.json; entries are reviewed before they appear on the public Leaderboard. Pre-computed paper baselines and TSV summary live in results/; metric definitions and normalization rules in docs/PROTOCOL.md; reference baselines and reproducibility notes in docs/BASELINES.md and docs/REPRODUCIBILITY.md.

License

Apache-2.0 (this code; see LICENSE). vitex_edit/diffsynth/ is adapted from DiffSynth-Studio (Apache-2.0). The dataset is CC-BY-NC 4.0 (non-commercial research only); see the Dataset card.

Citation

@inproceedings{chen2026vitexbench,
  title     = {ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing},
  author    = {Chen, Xinghao and Gao, Xiangbo and Yu, Jiongze and Wu, Yuheng and Tu, Zhengzhong},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},
  year      = {2026},
  url       = {https://vitex-bench.github.io/}
}

About

ViTeX-Bench (NeurIPS 2026 E&D): evaluation code for video scene text editing and code of the ViTeX-Edit-14B reference model

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages