Composio benchmarked 5 open-weight models across 30 multi-step agent tasks, finding GLM-5.3 Flash and DeepSeek-V4 Flash delivered the highest cost efficiency at a fraction of full-model pricing.
Key Takeaways
- βFlash-class models achieve near parity with full-scale weights at approximately 75% lower inference cost;
- βMulti-application tool orchestration remains the primary failure mode across all open-source models;
- βProvides empirical cost-per-success metrics to guide architectural selection for enterprise agent deployments.