Topic
reinforcement learning
Latest news
- 6 OctExPPO: exploration-preserving policy optimization for verified-reward RL
Researchers propose Exploration-Preserving Policy Optimization (ExPPO), an advantage-shaping rule that redistributes credit using prompt-relative surprisal and prompt pass rate to preserve verifier polarity. Experiments on multi-answer and out-of-domain reasoning tasks show improved coverage, higher aggregate accuracy, and increased yield of correct modes; code is available on the paper page.
Research · 1 source
Articles
No articles about reinforcement learning yet.