MM-Eval is a new multilingual benchmark for LLM-as-Judge and reward models! We hope this work can contribute towards better multilingual alignment for LLMs :)
🔥 New multilingual benchmark for testing both reward models & LLM-as-a-Judge 🔥
🌎 MM-Eval covers 18 languages across six subsets and includes language-specific challenges such as linguistics and language hallucinations.


