💻
DeepSeek V4 Flash tops leaderboards but completes only 54% of real agent tasks
💻 Technology

DeepSeek V4 Flash tops leaderboards but completes only 54% of real agent tasks

DeepSeek's V4 Flash topped AI leaderboards and was praised by developers, but Composio's real-world testing found it completed only 53.8% of complex agent tasks — 129 out of 240 runs across 30 difficult multi-step workflows using tools like Gmail, GitHub and Slack. The results show that orchestration and tool configuration, rather than raw model capability, may determine success in enterprise settings. DeepSeek also announced price hikes for V4 Flash.

Comments

No comments yet