Pinned
💡 Success under uncertainty is hard to repeat, while confident failures tend to recur. How can we learn from both to improve exploration in LLM reasoning?
Introducing EAPO (Entropic Advantage Policy Optimization), entropy-guided credit assignment for RLVR that treats success


