🔥The debate’s been wild: How does the reward in RLVR actually improve LLM reasoning?🤔
🚀Introducing our new paper👇
💡TL;DR: Just penalizing incorrect rollouts❌ — no positive reward needed — can boost LLM reasoning, and sometimes better than PPO/GRPO!
🧵[1/n]