← Today's brief

Research · Updated 11 Oct, 07:14 pm IST

Study: AI agents fail to innovate and fall short of autonomous research

Image: The Decoder

Why it matters for readers: Helps set realistic expectations about what AI agents can and cannot invent on their own.

  • Epoch AI created InnovationEval to test whether agents can invent, implement, test and refine a new LM training method independently.1
  • Agents were asked to improve on GRPO by devising a new method and had up to 3,000 hours of offline compute but no internet access.1
  • Claude Fable 5 and GPT-5.6 Sol largely recycled existing ideas instead of generating novel methods.1
  • GPT-5.6 Sol proposed reinforcing successful solutions when all answers are correct, but that idea was not new and scored about 35% of the improvement SDPO delivers over GRPO under generous grading.1
  • Neither model came close to matching the human-designed SDPO reference in the benchmark tests across short-answer and coding tasks.1

Get a brief like this every morning

Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.

Get started