Research · Updated 11 Oct, 09:09 pm IST
Agent teams waste tokens for little quality gain, benchmark finds

Why it matters for readers: Demonstrates that agent teams can be expensive and don't always produce noticeably better outputs than single models.
- Vals AI found agent teams cost between 1.8x and 5.1x more than single agents while showing almost no consistent quality gain.1
- Out of four comparisons between teams and solo agents, only GPT-6 Sol at medium reasoning showed a statistically significant improvement (7.3 points).1
- At maximum reasoning effort, team setups gave neither GPT-6 Sol nor Claude Opus 5.5 a meaningful advantage.1
- Anthropic’s tests with Opus 5.5 showed that larger teams reached performance levels faster but gains from ten to 100 agents were only small after 24 hours.1
- Separate ProgramBench tests reported speed gains for teams came with higher token usage, and some tasks showed stronger gains for alternative models like Fable 5.1.1
Get a brief like this every morning
Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.
Get started