On Oct 9 Prime Intellect shipped a Rust rewrite of Prime Agent, its open-source coding agent harness. A root orchestrator ran 2,209 sub-agents across 10,000+ Prime Sandboxes and ~228.7B GLM-5.3 tokens over two weeks, with differential TUI, harness and protocol parity tests gating every merge. Vendor-measured cold time-to-type dropped from 737.8ms to 55.8ms and post-startup memory from 607MB to 106MB; Windows (beta) and Homebrew installs were added.

Key Takeaways

  • ✓Scale: 2,209 agents, 16,758 agent-to-agent messages, 228.70B tokens (192.99B/1,981 agents for the port, 35.70B/228 for perf hillclimbing) across 10,000+ Prime Sandboxes
  • ✓Vendor-measured: cold time-to-type 737.8ms→55.8ms (13.22x), first paint 722.8ms→23.6ms, post-startup RSS 607.4MB→106.0MB, install 172.1MB→59.6MB
  • ✓Each task ran a Planner→Implementer→adversarial Reviewer (different model)→Verifier state machine; failures loop back to the implementer
  • ✓A 3-day target-free hillclimb loop logged 144+ experiment/audit records and merged 69+ performance changes
  • ✓Codebase now 9 crates, largest file ~2,500 lines (was ~15,000); native Windows (beta) and Homebrew installs, still open source
Prime Intellect's Prime Agent used a 2,209-agent swarm to rewrite itself in Rust: time-to-type 738ms→56ms, 80%+ less memory
🖼️Official Media
Click to view high-res
🧭

Heavy Claude Code use: compare subscription limits and API bills

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Prime Intellect rewrote Prime Agent, its open-source long-running coding agent harness, from TypeScript to Rust, with the agent doing most of the work. A root orchestrator that wrote no product code split the port into topologically ordered tasks; each passed through a planner, an implementer in its own worktree, an adversarial reviewer on a different model, and a verifier running parity checks in a fresh Prime Sandbox. Humans mainly built objective verification: terminal-frame TUI diffs, transcript and model-request diffs, daemon protocol checks and a per-component feature audit. In total 2,209 agents exchanged 16,758 messages and used 228.7B GLM-5.3 tokens over about two weeks, followed by a three-day hillclimbing loop that merged 69+ performance changes. Vendor-measured results: cold time-to-type 737.8ms to 55.8ms, first paint 722.8ms to 23.6ms, post-startup memory 607.4MB to 106.0MB, install size 172.1MB to 59.6MB; Prime Intellect cautions that cross-harness comparisons lack a common benchmark. The codebase is now nine crates with per-session worker isolation, native Windows (beta) and Homebrew installs.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.