Research · Updated 11 Oct, 07:14 pm IST
Study: AI agents fail to innovate and fall short of autonomous research

Why it matters for readers: Helps set realistic expectations about what AI agents can and cannot invent on their own.
- Epoch AI created InnovationEval to test whether agents can invent, implement, test and refine a new LM training method independently.1
- Agents were asked to improve on GRPO by devising a new method and had up to 3,000 hours of offline compute but no internet access.1
- Claude Fable 5 and GPT-5.6 Sol largely recycled existing ideas instead of generating novel methods.1
- GPT-5.6 Sol proposed reinforcing successful solutions when all answers are correct, but that idea was not new and scored about 35% of the improvement SDPO delivers over GRPO under generous grading.1
- Neither model came close to matching the human-designed SDPO reference in the benchmark tests across short-answer and coding tasks.1
Get a brief like this every morning
Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.
Get started