CodeAF is an Apache-2.0 Go binary that runs open-weight models as a local software factory with task fan-out and Pareto crewing. v0.6.0 shipped 2026-10-01; ★259 now. Vendor DeepSWE doc: /senior-dev solved 62/113 (54.9%) on DeepSeek V4 Flash vs nine other harnesses—with stated one-seed and sampling caveats.

Key Takeaways

  • ✓Agent-Field/CodeAF Apache-2.0; ★259 / 34 forks; install via agentfield.ai/get/codeaf
  • ✓v0.6.0 published 2026-10-01T18:59:00Z; README still says early preview
  • ✓Vendor DeepSWE same-model table: senior-dev 62/113 (54.9%) at ~22¢/task vs nine harnesses—one seed; sampling-contract caveat documented
  • ✓Later same 113 tasks: 88/113 (77.9%) on DeepSeek V4.1 Flash; 78/113 (69.0%) on Kimi K3 (senior-dev only)
  • ✓Built-in providers include OpenRouter, DeepSeek, GLM, Kimi, MiniMax, Qwen, Codex, Ollama; default worker/planner/checker crew
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Why it matters

Open-weight models often underperform in bare CLIs. CodeAF packages a local software factory (one Go binary, Apache-2.0) aimed at DeepSeek/Qwen/GLM/Kimi-class models: fan out work, crew seats, track spend. v0.6.0 is tagged but still labeled early preview.

Architecture

Chat seat plus per-task worker/planner/checker crew with Pareto crewing over connected providers. Headless codeaf do/exec for CI. Optional anonymized model-pool stats. Providers: OpenRouter, DeepSeek, GLM, Kimi, MiniMax, Qwen, Codex, Ollama, OpenAI-compatible.

Benchmarks (vendor)

DeepSWE docs: 10 harnesses, same DeepSeek V4 Flash, 113 tasks. senior-dev 62/113 (54.9%) at ~22¢/task (1x cost/solve). Caveats: one seed; sampling-contract difference for senior-dev. Later: 88/113 (77.9%) on V4.1 Flash; 78/113 on Kimi K3—senior-dev only. Treat as vendor-reported until independently reproduced.

Try it

curl -fsSL https://agentfield.ai/get/codeaf | bash, connect a key, run a small tested fix before citing leaderboard numbers in production decisions.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.