Topic

ExPPO

Sign in to follow ExPPO

Latest news

  1. 6 Oct
    ExPPO: exploration-preserving policy optimization for verified-reward RL

    Researchers propose Exploration-Preserving Policy Optimization (ExPPO), an advantage-shaping rule that redistributes credit using prompt-relative surprisal and prompt pass rate to preserve verifier polarity. Experiments on multi-answer and out-of-domain reasoning tasks show improved coverage, higher aggregate accuracy, and increased yield of correct modes; code is available on the paper page.

    Research · 1 source

Articles

No articles about ExPPO yet.