← Today's brief

Research · Updated 11 Oct, 09:09 pm IST

Agent teams waste tokens for little quality gain, benchmark finds

Image: The Decoder

Why it matters for readers: Demonstrates that agent teams can be expensive and don't always produce noticeably better outputs than single models.

  • Vals AI found agent teams cost between 1.8x and 5.1x more than single agents while showing almost no consistent quality gain.1
  • Out of four comparisons between teams and solo agents, only GPT-6 Sol at medium reasoning showed a statistically significant improvement (7.3 points).1
  • At maximum reasoning effort, team setups gave neither GPT-6 Sol nor Claude Opus 5.5 a meaningful advantage.1
  • Anthropic’s tests with Opus 5.5 showed that larger teams reached performance levels faster but gains from ten to 100 agents were only small after 24 hours.1
  • Separate ProgramBench tests reported speed gains for teams came with higher token usage, and some tasks showed stronger gains for alternative models like Fable 5.1.1

Get a brief like this every morning

Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.

Get started