Research · Updated 6 Oct, 09:30 am IST
Training policies improves reliability of autonomous coding agents
Why it matters for readers: The study reveals concrete reasons why some automated code fixes fail and how targeted training can make them better.
- Extra inference-time compute helps only when it yields a useful repair and reliable evidence to pick it, driven by three teachable behaviors.1
- Directing search with execution feedback and scoring patches against a reverted tree resolves 52.8% of SWE-bench Verified while using 48.1% of the agent-steps of an eight-sample baseline.1
- Weighted supervised fine-tuning on 270 held-out issues raises pass@1 from 31.9% to 35.2% and pass@8 from 46.7% to 51.1%.1
- A reinforcement objective that trains the verifier against correct and incorrect repairs further raises pass@1 to 43.0%, pass@8 to 60.7%, increases verifier precision from 26.8% to 41.7%, and halves false acceptance.1
- Verification can be misleading because tests the agent writes for its own patch may accept many incorrect repairs, motivating stronger verifier training.1
Get a brief like this every morning
Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.
Get started