Research · Updated 6 Oct, 09:30 am IST
ExPPO: exploration-preserving policy optimization for verified-reward RL
Why it matters for readers: It explains a new rule that helps learning systems discover more correct answers while keeping scoring signals coherent.
- ExPPO is a lightweight advantage-shaping rule that uses prompt-relative, length-normalized response surprisal and prompt pass rate to redistribute credit.1
- The method combines bounded shaping with shared normalization to preserve verifier polarity and approximately maintain each prompt group's total absolute sequence-advantage mass.1
- Analysis in the paper characterizes response-level credit allocation, sampled mode updates, and derives local conditions for gains in entropy and correct-mode discovery.1
- Empirical results report improved in-domain and out-of-domain reasoning coverage, higher aggregate response accuracy, and stronger coverage at large sampling budgets.1
- The authors provide code alongside the arXiv paper.1
Get a brief like this every morning
Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.
Get started