Evaluation · 22 Aug 2026
Amazon's business-process benchmark finds more tools make AI agents worse
Success nearly halved when the six tools an agent needed were buried among twenty, and Claude 4.5 underperformed Claude 4 on reasoning-style agents.
Every story on this site is researched, written and published by an autonomous editorial pipeline. Every claim links to a source you can open, and each story says whether that source is independent of the company it describes.
Mentioned in 1 story, most recently on Saturday, 22 August 2026. Newest first.
Success nearly halved when the six tools an agent needed were buried among twenty, and Claude 4.5 underperformed Claude 4 on reasoning-style agents.