Composio benchmarked 5 open-weight models across 30 multi-step agent tasks, finding GLM-5.3 Flash and DeepSeek-V4 Flash delivered the highest cost efficiency at a fraction of full-model pricing.

Key Takeaways

  • βœ“Flash-class models achieve near parity with full-scale weights at approximately 75% lower inference cost;
  • βœ“Multi-application tool orchestration remains the primary failure mode across all open-source models;
  • βœ“Provides empirical cost-per-success metrics to guide architectural selection for enterprise agent deployments.