← Today's brief

Research · Updated 6 Oct, 09:30 am IST

Training policies improves reliability of autonomous coding agents

Why it matters for readers: The study reveals concrete reasons why some automated code fixes fail and how targeted training can make them better.

  • Extra inference-time compute helps only when it yields a useful repair and reliable evidence to pick it, driven by three teachable behaviors.1
  • Directing search with execution feedback and scoring patches against a reverted tree resolves 52.8% of SWE-bench Verified while using 48.1% of the agent-steps of an eight-sample baseline.1
  • Weighted supervised fine-tuning on 270 held-out issues raises pass@1 from 31.9% to 35.2% and pass@8 from 46.7% to 51.1%.1
  • A reinforcement objective that trains the verifier against correct and incorrect repairs further raises pass@1 to 43.0%, pass@8 to 60.7%, increases verifier precision from 26.8% to 41.7%, and halves false acceptance.1
  • Verification can be misleading because tests the agent writes for its own patch may accept many incorrect repairs, motivating stronger verifier training.1

Get a brief like this every morning

Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.

Get started